Artificial intelligence
On pstack: the skills behind 2,500 PRs a month, and the mental models behind the skills
Lauren Tan’s agentic engineering skills, and the six pictures she uses to explain them.
Contents
On pstack: the skills behind 2,500 PRs a month, and the mental models behind the skills
If you are deep into agentic engineering, you will have seen a lot of buzz lately about pstack, the set of agentic engineering skills from Lauren Tan, aka @poteto. A principal engineer at SpaceX AI, she leads development of Grok Bot, probably the fastest-growing cloud agent at the time of writing, and she has used these skills to automate much of its development.
I think understanding how the skills work, and the mental models behind them, is worth the effort for every agentic engineer, including those of us focused on knowledge work. What follows is my attempt to break it down.
Part I: Lauren’s mental models
pstack is the code form of a way of thinking. Lauren explains that thinking through a small set of pictures: a kitchen, a ladder, two loops, a garden. Select a model to jump to it. Each model starts with her own words, then maps the picture to agent work. The sources are her two essays, The Complete Guide to pstack, Part 1 and Part 2, and her interview with Matt Pocock. They are listed at the foot of the page.
Mental model #1: The Michelin Kitchen
Lauren describes engineering with agents as running a professional kitchen. She prefers this picture to the “software factory”. A factory suggests volume. A Michelin kitchen keeps the volume and adds craft: the same ingredients can feed a person or become a meal that they remember. As the kitchen grows, the chef cooks less and organises more, and the chef’s name stays on every plate.

The chef sets the standard and tastes the result. Prepared kitchens let the cooks work without her at every station.
From home cook to restaurant group
Lauren compares her own path to four stages in a cook’s career. Each stage needs the one before it.
One person does all the work: prep, cooking and cleaning.
You prompt one agent and check every line it writes. Lauren calls herself the “meat proxy” at this stage, carrying results between the agent and the browser tools by hand.
More hands arrive, with no system. Nobody knows where the utensils are, and the cook gets stressed.
Several agents without shared tools or rules collide, repeat work and each invent their own method.
The chef stops cooking every dish. The job becomes ordering, storing and preparing ingredients, and setting the standard. Lauren calls the chef “the CEO of the kitchen”.
You build the environment: skills, tools, checks and a codebase that guides each agent. You remain accountable for the result.
The first kitchen runs without the chef. Gordon Ramsay opens a second and a third restaurant, and visits each one to taste.
Parallel projects run at the same time, each with a coordinator agent. You move between them and sample their output.
What each part of the kitchen stands for
You, the engineer.
You set the intent and the standard. “Your name still is associated with the work that you do.”
A coordinator agent: in Lauren’s setup, a Cursor project.
It delegates and supervises a list of tasks. It does no cooking of its own.
Sub-agents that write, run and check the code.
Frontier models are capable cooks. What they lack is context and a prepared station.
Skills, small CLIs and verification tools.
“If you have a dull knife, then everything’s going to take a long time.” Without tools, each cook makes up a method of their own.
The codebase, its types and its lint rules.
Constraints put each cook on a narrow, correct path. A new hire, human or agent, is productive from the first day.
Bug reports and requests from Slack, X, email and Linear.
Grok Bot routines watch these channels and send the work to the right project. Mental model 3 covers this outer loop.
Sampling merged work.
A chef with several restaurants cannot taste every dish. When the same fault appears in several dishes, the chef changes the kitchen. Mental model 4 covers this.
Each engineer’s personal set of skills.
A chef takes their knives to each new restaurant. Lauren encourages readers to build their own set from their past transcripts, and to borrow from pstack only what they trust.
Mental model #2: The Trust Ladder
Matt Pocock opens the interview with Lauren’s “trust ladder”. The more you trust your agents, the larger the tasks you can give them and the more agents you can run at once. Agents earn that trust with evidence: work they can show and check. Lauren also places trust in the tools themselves: “trust to me is really about trust in your own tools.”

Each rung is a tool that lets the agent prove its own work. The human builds and checks the rungs.
How Lauren climbed it
Her own account, in order. She describes no fixed number of rungs. The steps below follow her story.
On a side project after Meta, she spent hours steering one agent and writing skills whose effect she could not measure. “I was kind of flying blind.”
Her first skills, which tried to teach the agent to work like her. They became the basis of pstack.
At Cursor she fixed performance problems by hand, carrying flame graphs and heap snapshots between Chrome DevTools and the agent.
The realisation that she was the bottleneck.
Her first Cursor skill gave the agent “hands and eyes”: it could run the app, use it, debug it and take traces.
The agent closes its own loop. It can hill-climb, improving against a score without her.
She added tools, types, lint rules and constraints that “make the easy thing the right thing”.
Many agents and several projects at the same time.
“Full autopilot” sends verifier agents to fuzz each pull request before it lands. She reads the commit history in the morning and reverts or adds a lint rule where needed.
“I’m sleeping so much better now.” The first night was “very scary”.
Two cautions
Low on the ladder, people get stuck. Without trust you must micromanage, and micromanagement leaves no time to sharpen your knives. And the top rungs depend on how far the work can be verified. Software and some mathematics verify well. Lauren says she has no answer yet for work that is hard to verify, or where a mistake cannot be undone.
Mental model #3: Inner and outer loops
Lauren’s agents work toward “a snapshot of my intent”. The snapshot goes stale: new bug reports, requests and limits arrive in Slack, Linear, email and X. For a long time she carried that context to her agents herself. Her fix is two loops, joined. She borrows a management principle from her time at Netflix: give people the context to work alone, and they need less supervision.

The outer loop brings in fresh orders. The inner loop cooks, tastes and corrects until the dish passes.
Grok Bot routines with connectors to Slack, X, email and Linear.
Watch for new information and send it to the right project.
Cursor projects: a coordinator agent in the cloud, with its own computer.
Split the work, start sub-agents, and drive each task to a checked result.
Grok Bot sends messages straight to a project.
Related reports go to one project, so one coordinator sees the shared cause and each fix happens once.
Set the intent and sample the results.
Ask “where am I the bottleneck?”, then teach the agent to fetch that answer from real data.
Outer loopAcross the projects
Context and work
Results and evidence
Inner loopWithin each Cursor project
This connects with my two-loop essay and When AI Leaves the Chat Window. My essay places human direction in the outer loop. Lauren also automates external context gathering. The diagram combines these roles. Review effort and intervention depend on the work and its risks.
Mental model #4: The Gardener
In Part 1 of her guide, Lauren writes that “every team needs a gardener”: someone who watches the stream of pull requests and notices the smells, such as “the third isRecord this week, the lint suppressions creeping like ivy”. Most of her 2,500 pull requests a month are gardening work: refactors, new lint rules and new checks. A tended garden helps everyone who works in it. A new hire, human or agent, writes good code from the first day.

Pull the weed by the root, then change the bed so the same weed cannot return.
An agent scans the code all the time for banned patterns.
Smells appear faster than one person can read code.
The agent appends each finding to a document and fixes nothing yet.
A buffer shows patterns that a fix for each single bug would hide.
Every few days she reads the queue and finds that many entries share one cause.
In pure execution mode, “you sometimes miss the big picture”.
Each day she reads a sample of merged pull requests closely.
At this volume she cannot taste every dish.
When several agents take the same shortcut, she changes the skills, constraints, lint rules or types. A one-off mistake needs no action.
The environment then guides every later agent to the right path.
Mental model #5: Hands, eyes and a map
Lauren calls verification “the single most important skill” in an agent toolkit. It gives the agent “hands and eyes”: it can run the app, use it as a person would, debug it and take traces. An agent that can see the result of its work can keep going until it succeeds, without you in the middle. She treats a good verification skill as critical infrastructure, and suggests it could have its own on-call rotation.

The CLI is a remote control the agent can hold. The Feature Map tells it which buttons exist.
Three parts
Part 1 of her guide builds the skill from three pieces. Each one removes a reason for the agent to ask you.
A project skill that launches, drives and observes the real app. /create-verification-skill writes it, and /maintain-verification-skill keeps it current each day.
“Verification is all you need.” Without it, you stay the bottleneck and spend the day babysitting agents.
One script inside the skill, with named commands such as doctor, snapshot, screenshot and send.
Code does the mechanical parts and the agent keeps the judgement. Every agent reuses one tested tool, where each used to “rebuild the world each time”. pstack calls this principle “Build the Lever”.
A searchable markdown map of every feature: what it does, how a user reaches it, how to drive it with the CLI, and its traps.
“Materialised memory”. The codebase is the true memory, and the map is a compact copy that saves tokens. It needs daily upkeep, because the agent trusts it.
Mental model #6: Measure a hundred times, cut once
Part 2 of her guide asks the next question: once agents can verify their work, how do you decide what to build? Lauren calls it “the art of supervising someone smarter than you”. Models now write code better than most people, in codebases too large to hold in one head. Her answer has two halves. First draw out understanding. Then plan with code and evidence. “I do plan, but I do so through code.” In her view, a long abstract plan gives “the illusion of progress”.

Define the use first. Build several designs, load-test each, and keep the one that holds.
First, understand
“Restate in your own words and in plain english what you think the underlying issue is.”
It catches a misunderstanding before any code is written, and keeps her own guesses out of the prompt.
/how · /why · /teach/how traces how a subsystem runs. /why searches Git history, tickets, docs and Slack for the reason. /teach explains the result to her.
It keeps her mental model current, and it makes the agent read the code before it states a claim.
/recallLoads past transcripts into a fresh chat.
“Your past transcripts are often a gold mine for rich context.”
Then, plan with code
Write the tutorial or README before the code. /technical-writing keeps tutorial, how-to, reference and explanation apart.
It gives the agent a concrete target to check its work against.
Two or three variants behind a switcher. The verification skill drives each one and measures it.
The two common mistakes are accepting the agent’s first design and overcooking a plan with no evidence.
/architectSeveral model families sketch types and signatures in parallel. A judge on a different model picks. The design is scrapped if workarounds or forced casts appear.
Architecture is now the engineer’s main job. Agents fill in the implementation.
Written only after the design settles. Each task names its proof. The plan is deleted when the work lands.
Work counts as verified only after the code has actually run.
/architect ship inside pstack.Part II: The package
This inventory describes Lauren Tan's published package, pstack 0.15.9 at commit e43c7ee in cursor/plugins. Cursor installs it as a plugin. The copy I run in Claude Code, Codex, Grok and Muse is a direct port of the same files. It maps Cursor's subagent, question and path conventions onto each host.
A router + a library + execution support
e43c7ee. Select a component to inspect it. Manifest · RouterOne router connects the parts
The model reads instructions, selects a playbook and uses the host runtime. Mode · README
Instructions steer the model
Routerpoteto-modeSelect the task sequence. Copy its steps. Load the relevant skills.Selects the playbook and reads relevant leaves
The lead applies the instructions through runtime tools
The lead checks the actual result
Arrows show instruction and execution dependencies at commit e43c7ee. Verification rule
Every skill, in one view
Select a name to see its job, dependencies and exact source. The five families are editorial groupings. Every skill except setup-pstack runs only when it is called. Cursor lists each one under /.
Route, design & delegate7
Understand & recover context4
Verify & review7
Code & writing4
Learn & extend4
/poteto-mode
Selects a task playbook and applies the shared working rules.
Depends on
Reads a playbook, the relevant principles and any named skill.
/setup-pstack
Sets the model for each role at a chosen reasoning budget.
Depends on
Writes an always-applied model rule. Offers create-verification-skill.
/figure-it-out
Designs an auditable workflow when no playbook fits.
Depends on
Uses architect, arena and show-me-your-work.
/architect
Sketches types, signatures and modules before code, then stays with the build.
Depends on
Uses how and why. Can call arena and interrogate.
/arena
Runs parallel candidates, picks a base and grafts in the best parts of the others.
Depends on
Spawns candidates and a cross-judge.
/swarm
Fans out parallel workers and returns one report.
Depends on
Spawns workers on the swarm role model.
/show-me-your-work
Keeps a decision log a reviewer can inspect.
Depends on
Appends rows with its log script.
/how
Explains how a subsystem works, from source and runtime flow.
Depends on
Spawns explorers and one explainer.
/why
Reconstructs why a design is the way it is, with stated confidence.
Depends on
Queries each evidence source it can reach. Uses how.
/recall
Rebuilds your working context from history, live state and the shared record.
Depends on
Uses why.
/blast-radius
Finds what a change can break beyond the diff and proves the safety fact.
Depends on
Uses how and why. Can call arena.
/tdd
Proves a focused test fails before the fix and passes after it.
Depends on
Used by the bug-fix playbook.
/create-verification-skill
Generates a project skill that drives the app as a user does.
Depends on
Produces the project verify skill and feature map.
/maintain-verification-skill
Checks every mapped feature against source and the live app.
Depends on
Keeps the project verify skill current.
/interrogate
Has several model reviewers challenge a change from separate angles.
Depends on
Spawns the interrogate reviewer panel.
/benchmark-checklist
Checks that a measurement tests the work it claims to test.
Depends on
Pairs with the Explain the Number principle.
/no-comments
Runs a comment audit and acts on the accepted findings.
Depends on
Spawns the Comment Sicko agent. Uses architect, how and why.
/correct
Finds mistakes agents repeat in a repo and makes each one impossible.
Depends on
Prefers architecture, then types, lints and tests, then docs.
/unslop
Removes AI writing patterns from any prose.
Depends on
Applied to every reply and document.
/technical-writing
Applies a layered standard for technical documents.
Depends on
Uses how and unslop.
/typescript-best-practices
Applies TypeScript rules when a .ts or .tsx file is read or edited.
Depends on
Loads on TypeScript paths.
/bro
Restates the last message in plain language.
Depends on
Reads the previous reply.
/teach
Explains a body of work so a person understands it.
Depends on
Uses how, why and unslop.
/reflect
Reviews the session and routes each lesson to a skill edit.
Depends on
Spawns three reviewers over the transcript.
/automate-me
Turns your working preferences into a personal mode skill.
Depends on
Builds on poteto-mode. Uses unslop.
/make-bot-ui
Builds a small UI that wakes a Grok Bot through a webhook.
Depends on
Optional exposure on Tailscale.
How skills use each other
Selected dependencies from the source text. Branches show called skills, delegated agents or generated outputs.
architecthowTrace the systemwhyWhen rationale is neededarenaCompare design candidatesinterrogateChallenge the designfigure-it-outarchitectShape the designarenaCompare approachesshow-me-your-workLog each decisionteachhowMechanismwhyRationaleunslopCheck the explanationblast-radiushowTrace callerswhyRecover intentarenaSettle an open forkno-commentscomment-sickoDelegate the comment audithowCheck a claimed constraintcreate-verification-skillproject verifyGenerate the app driver and feature mapmaintain-verification-skillproject verifyCheck every mapped featuretechnical-writinghowGround the documentunslopApply the writing-pattern checks23 routes through a task
The router selects a playbook by task. Large or unusual work goes through figure-it-out, which builds a bespoke workflow. Playbook index
Principles guide the decisions
24 principle skills. poteto-mode names when each applies, and the agent reads the leaf in full before it applies one. Principle index
Core
Laziness ProtocolFoundational ThinkingRedesign from First PrinciplesAttack the PremiseSubtract Before You AddMinimize Reader LoadOutcome-Oriented ExecutionExperience FirstExhaust the Design SpaceBuild the LeverArchitecture
Model the DomainBoundary DisciplineType System DisciplineMake Operations IdempotentMigrate Callers Then Delete Legacy APIsSeparate Before Serializing Shared StateAgent type and model are separate choices
The package defines 2 agents and sets no model in either. The model rule names a model for each role, and the caller passes it on each spawn. poteto-agent · comment-sicko · Setup
2 agent definitions
Role configuration
setup-pstack detects the available models, asks for a budget (unlimited, large, medium or small) and writes one always-applied rule. A role without a line uses the skill's default. Setup workflow
The runtime carries out the work
The selected playbook determines the required tools.
Keep work state
orch · check-planshow-me-your-work log
Units · gates · checkpoints
orch · check-plan · log
The package also contains a read-only worktree audit for disk cleanup. Cursor cloud agents run the long, parallel playbooks. Worktree audit · Orchestrate
Recent upstream changes
A new skill, correct
Finds mistakes that agents repeat in a repo and makes each one impossible.
Designs that resist agent mistakes
architect now shapes a design so later agents cannot misuse it.
Explain the Number
A measurement principle, fresh subagents per task and an hourly autopilot tick.
Complete source catalogue
Lauren Tan created pstack. Source links open the public files at commit e43c7ee, which fixes this October 2026 snapshot. Package README
The package inventory
| Package | Version | Skills | Principles | Playbooks | Agents |
|---|---|---|---|---|---|
| cursor/plugins pstack | 0.15.9 · e43c7ee | 26 | 24 | 23 | 2 |
Counts come from the skill folders, the playbook files and the agent files at that commit. Every skill sets disable-model-invocation: true except setup-pstack, so a skill runs when the user or poteto-mode calls it. Manifest
The skills have distinct jobs
The five groups below are editorial categories. Open a group to inspect its commands and dependencies.
Routing, design and delegation · 7 skills
| Command | Job | Dependencies and conditions | Evidence |
|---|---|---|---|
/poteto-mode | Selects a task playbook and applies the shared working rules. | Reads a playbook, the relevant principles and any named skill. | source |
/setup-pstack | Sets the model for each role at a chosen reasoning budget. | Writes an always-applied model rule. Offers create-verification-skill. | source |
/figure-it-out | Designs an auditable workflow when no playbook fits. | Uses architect, arena and show-me-your-work. | source |
/architect | Sketches types, signatures and modules before code, then stays with the build. | Uses how and why. Can call arena and interrogate. | source |
/arena | Runs parallel candidates, picks a base and grafts in the best parts of the others. | Spawns candidates and a cross-judge. | source |
/swarm | Fans out parallel workers and returns one report. | Spawns workers on the swarm role model. | source |
/show-me-your-work | Keeps a decision log a reviewer can inspect. | Appends rows with its log script. | source |
Understanding and context · 4 skills
| Command | Job | Dependencies and conditions | Evidence |
|---|---|---|---|
/how | Explains how a subsystem works, from source and runtime flow. | Spawns explorers and one explainer. | source |
/why | Reconstructs why a design is the way it is, with stated confidence. | Queries each evidence source it can reach. Uses how. | source |
/recall | Rebuilds your working context from history, live state and the shared record. | Uses why. | source |
/blast-radius | Finds what a change can break beyond the diff and proves the safety fact. | Uses how and why. Can call arena. | source |
Verification and review · 7 skills
| Command | Job | Dependencies and conditions | Evidence |
|---|---|---|---|
/tdd | Proves a focused test fails before the fix and passes after it. | Used by the bug-fix playbook. | source |
/create-verification-skill | Generates a project skill that drives the app as a user does. | Produces the project verify skill and feature map. | source |
/maintain-verification-skill | Checks every mapped feature against source and the live app. | Keeps the project verify skill current. | source |
/interrogate | Has several model reviewers challenge a change from separate angles. | Spawns the interrogate reviewer panel. | source |
/benchmark-checklist | Checks that a measurement tests the work it claims to test. | Pairs with the Explain the Number principle. | source |
/no-comments | Runs a comment audit and acts on the accepted findings. | Spawns the Comment Sicko agent. Uses architect, how and why. | source |
/correct | Finds mistakes agents repeat in a repo and makes each one impossible. | Prefers architecture, then types, lints and tests, then docs. | source |
Code and writing · 4 skills
| Command | Job | Dependencies and conditions | Evidence |
|---|---|---|---|
/unslop | Removes AI writing patterns from any prose. | Applied to every reply and document. | source |
/technical-writing | Applies a layered standard for technical documents. | Uses how and unslop. | source |
/typescript-best-practices | Applies TypeScript rules when a .ts or .tsx file is read or edited. | Loads on TypeScript paths. | source |
/bro | Restates the last message in plain language. | Reads the previous reply. | source |
Learning and extension · 4 skills
| Command | Job | Dependencies and conditions | Evidence |
|---|---|---|---|
/teach | Explains a body of work so a person understands it. | Uses how, why and unslop. | source |
/reflect | Reviews the session and routes each lesson to a skill edit. | Spawns three reviewers over the transcript. | source |
/automate-me | Turns your working preferences into a personal mode skill. | Builds on poteto-mode. Uses unslop. | source |
/make-bot-ui | Builds a small UI that wakes a Grok Bot through a webhook. | Optional exposure on Tailscale. | source |
The playbooks cover 23 task types
Each row quotes the condition poteto-mode uses to select the playbook.
All 23 playbooks
| Playbook | Selected for | Evidence |
|---|---|---|
| Authoring or modifying a skill | Writing or editing a SKILL.md. | source |
| Autonomous run | A long task to drive to completion without stopping ("run until done", "/loop until X"). | source |
| Autopilot-full | A queue of independent PRs run to merged with full autonomy. One owner per PR carries build through merge, and the root swarm-verifies each PR before its owner merges ("autopilot this queue", "full autopilot", one-owner-per-PR programs). | source |
| Autopilot-stack | A queue of changes built and verified with full autonomy, delivered as one linear reviewed base-branch stack the operator lands ("autopilot-stack", "stack them, don't ship", "build the stack, I'll land it"). | source |
| Babysit | Driving a PR or a stack to merge-ready: conflicts, review threads, CI. | source |
| Bug fix | A reported defect to reproduce, root-cause, and fix with runtime evidence. | source |
| Eval | Testing how a skill, structure, or prompt change affects agent behavior before promoting it. | source |
| Feature | New or changed behavior, built from a named data shape. | source |
| Hillclimb | Sustained, scientific improvement of one metric against a target: loop hypotheses with before/after measurement, a decision log, and one commit per accepted win. Distinct from Perf issue, which is a one-off fix. | source |
| Investigation | Read-only question: how does X work, why was Y built this way, are we sure about Z, should we do X or Y. | source |
| Multi-phase or multi-PR plan | Work that spans phases or stacked PRs. | source |
| Opening a PR | Invoked at the end of every other playbook. | source |
| Orchestrate | A standing project handed to one coordinator chat: multi-day, many stacked PRs, dozens to hundreds of subagents, minimal human turns ("run this whole project", "own this migration until it lands"). Distinct from Autonomous run, which drives one task to a predicate. Work one agent could finish inside the session's budget routes there, not here, however program-shaped the phrasing sounds. | source |
| Pause safely | Suspending in-flight work cleanly so it can be resumed, on an explicit pause, going offline, a Cursor restart, or imminent context compaction. The complement to Session pickup. Full steps: | source |
| Perf issue | A measured slowness to trace and improve against a baseline. | source |
| Prototype | A throwaway sketch to make a design or behavioral decision cheaply, or to settle an empirical fork by observing it instead of asking the human ("prototype", "mock it up", "try this layout", "sketch it to decide"). | source |
| Refactoring | A behavior-preserving change to structure or shape (rename, extract, inline, dedupe, move). | source |
| Runtime forensics | Diagnose a runtime symptom (leak, idle-CPU spin, glitch) from live instrumentation. The deliverable is a diagnosis, not a fix. | source |
| Session pickup | Resuming or taking over a prior agent's in-flight work from a transcript, cloud-agent URL, or pushed branch. | source |
| Shipping | The half after Babysit. Independently verifying a green stack, then landing the contiguous verified run bottom-up through gh by default or Origin when its CLI is available. | source |
| Trace forensics | Diagnose a captured profiling artifact (cpuprofile, trace, spindump, heap snapshot) handed to you after the fact. The deliverable is a diagnosis, not a fix. | source |
| Visual parity | Pixel-exact UI equivalence: matching two implementations or migrating a styling system. | source |
| Worktree and simulator cleanup | Reclaiming local disk by pruning merged or abandoned git worktrees and stale iOS simulators ("what's using my disk", "clean up worktrees", "prune safe-to-prune worktrees", "free up space", "delete old simulators"). | source |
The principles apply across the workflows
Each row quotes when the principle applies, from the poteto-mode index.
Core · 10 principles
| Principle | Applies when | Evidence |
|---|---|---|
| Laziness Protocol | Refactoring, sizing a diff, or tempted to add abstractions, layers, or signal threading. Bias to deletion and the smallest change that solves the problem. | source |
| Foundational Thinking | Before writing logic: core types and data structures, scaffold-vs-feature sequencing, what concurrent actors share. | source |
| Redesign from First Principles | Integrating a new requirement into an existing design. Redesign as if it had been foundational from day one. | source |
| Attack the Premise | Two or more fixes that share one premise have failed the same gate. Take a census of which actors hold the imbalance before the next fix, then question the premise instead of writing another fix that assumes it. | source |
| Subtract Before You Add | Sequencing an addition, refactor, or rewrite. Remove dead weight first, then build on the simpler base. | source |
| Minimize Reader Load | Reviewing or shaping code that's hard to trace. Count layers and hidden state, collapse one-caller wrappers, shrink mutable scope. | source |
| Outcome-Oriented Execution | Planned rewrites and migrations with explicit phase boundaries. Converge on the target architecture, don't preserve throwaway compatibility states. | source |
| Experience First | Product, UX, or feature-scope tradeoffs. Choose user delight over implementation convenience. | source |
| Exhaust the Design Space | A novel interaction or architectural decision with no precedent. Build 2-3 competing prototypes and compare before committing. | source |
| Build the Lever | Any non-trivial work. Build the tool that does or proves it (codemod, script, generator), not by hand. The tool is the artifact a reviewer reruns. | source |
Architecture · 6 principles
| Principle | Applies when | Evidence |
|---|---|---|
| Model the Domain | Writing stateful logic, or code that branches a lot or repeats a shape assumption across files. Encode the domain in a structure (state machine, typed model, table or registry, reducer, boundary, the right collection) instead of scattered conditionals. | source |
| Boundary Discipline | Wiring validation, error handling, or framework adapters. Guards at system boundaries, trust internal types, keep business logic pure. | source |
| Type System Discipline | Designing types or a signature in any typed language. Make illegal states unrepresentable, brand primitives, parse external data at boundaries. | source |
| Make Operations Idempotent | Designing commands, lifecycle steps, or loops that run amid crashes and retries. Converge to the same end state. | source |
| Migrate Callers Then Delete Legacy APIs | Introducing a new internal API while old callers exist. Migrate and delete in one wave. | source |
| Separate Before Serializing Shared State | Concurrent actors might write the same file, branch, key, or object. Eliminate the sharing first. | source |
Verification · 5 principles
| Principle | Applies when | Evidence |
|---|---|---|
| Prove It Works | After a task, before declaring done. Verify against the real artifact, not a proxy or "it compiles". | source |
| Fix Root Causes | Debugging. Trace each symptom to its root cause, reproduce first, ask why until you reach it. | source |
| Sequence Work into Verifiable Units | Multi-step work (sweeps, migrations, runs of similar edits) and how you stack commits and PRs. Break work into small units that each end in a check, verify each before the next, and order delivery so the sequence proves itself. | source |
| Test Behavior, Not Implementation | Writing, changing, or keeping a test. Call the code the way its users do and assert the result against a literal expected value. If the test would still pass when every imported function returns undefined, rewrite the assertion or delete the test. | source |
| Explain the Number | Before you trust, report, or act on a number you measured (a speedup, a regression, a throughput, a latency, or an eval result). Find what limits it, and rule out that it measured something other than the work you think. | source |
Delegation · 2 principles
| Principle | Applies when | Evidence |
|---|---|---|
| Guard the Context Window | Context fills up: large outputs, long files, repeated reads, fan-out planning. Route bulk to subagents, keep summaries in the main thread. | source |
| Never Block on the Human | Tempted to ask "should I do X?" on reversible work. Proceed, present the result, let the human course-correct. | source |
Meta · 1 principle
| Principle | Applies when | Evidence |
|---|---|---|
| Encode Lessons in Structure | You catch yourself writing the same instruction a second time. Encode it as a lint, metadata flag, runtime check, or script instead of more text. | source |
The helpers complete the stack
| Helper | Role | Evidence |
|---|---|---|
orch | A work-state store for long projects: units, gates, checkpoints and an inbox. | source |
check-plan | Checks that each planned PR has its unit, live and performance boxes ticked. | source |
watch-pr | Watches PR status, checks and review threads for the babysit and shipping playbooks. | source |
bootstrap | Installs the script dependencies with Bun. | source |
worktree-audit | Classifies git worktrees for cleanup and deletes nothing. | source |
log.sh | Appends one well-formed row to a show-me-your-work decision log. | source |
Source and verification scope
Verification depends on a driver for the real app. The package supplies skills to generate and maintain one. Generation · Maintenance
The skill defaults name Cursor model slugs. On another host, the model rule must name that host's models.
Upstream changes most days. The links above stay fixed at commit e43c7ee. The current package is in cursor/plugins.
Part III: Read and watch Lauren
The mental models above come from these four sources, listed in the order they appeared.
The Complete Guide to pstack, Part 1Verification is all you need. How to build a verification skill, a small CLI and a Feature Map, and why they are critical infrastructure.
The Complete Guide to pstack, Part 2The art of supervising someone smarter than you. Research, prototypes and architecture before a single plan is written.
How I shipped 2,500 PRs last month to productionThe talk that introduced the trust ladder, verification and the software factory to a wide audience. Matt Pocock’s interview works through it as a Q&A.
Poteto on shipping 1,000 PRs a month at SpaceXLauren and Matt Pocock discuss the trust ladder, the Michelin Kitchen, the two loops and sampling instead of review.
The package: pstack on GitHub, cursor/plugins