Barnaby Robson
White line drawing on black: Lauren Tan pulls a weed out by its roots from a raised bed while a small robot waters the plants beside her.

Artificial intelligence

On pstack: the skills behind 2,500 PRs a month, and the mental models behind the skills

Lauren Tan’s agentic engineering skills, and the six pictures she uses to explain them.

On pstack: the skills behind 2,500 PRs a month, and the mental models behind the skills

If you are deep into agentic engineering, you will have seen a lot of buzz lately about pstack, the set of agentic engineering skills from Lauren Tan, aka @poteto. A principal engineer at SpaceX AI, she leads development of Grok Bot, probably the fastest-growing cloud agent at the time of writing, and she has used these skills to automate much of its development.

I think understanding how the skills work, and the mental models behind them, is worth the effort for every agentic engineer, including those of us focused on knowledge work. What follows is my attempt to break it down.

5 October 2026

Part I: Lauren’s mental models

pstack is the code form of a way of thinking. Lauren explains that thinking through a small set of pictures: a kitchen, a ladder, two loops, a garden. Select a model to jump to it. Each model starts with her own words, then maps the picture to agent work. The sources are her two essays, The Complete Guide to pstack, Part 1 and Part 2, and her interview with Matt Pocock. They are listed at the foot of the page.

Mental model #1: The Michelin Kitchen

Lauren describes engineering with agents as running a professional kitchen. She prefers this picture to the “software factory”. A factory suggests volume. A Michelin kitchen keeps the volume and adds craft: the same ingredients can feed a person or become a meal that they remember. As the kitchen grows, the chef cooks less and organises more, and the chef’s name stays on every plate.

Lauren holds a finished dish and a brief between two working kitchens. Robot cooks prepare food inside each kitchen, a courier robot runs in with fresh orders, and a knife roll of tools rests at her feet.

The chef sets the standard and tastes the result. Prepared kitchens let the cooks work without her at every station.

From home cook to restaurant group

Lauren compares her own path to four stages in a cook’s career. Each stage needs the one before it.

1Home cook

One person does all the work: prep, cooking and cleaning.

You prompt one agent and check every line it writes. Lauren calls herself the “meat proxy” at this stage, carrying results between the agent and the browser tools by hand.

2Family in the kitchen

More hands arrive, with no system. Nobody knows where the utensils are, and the cook gets stressed.

Several agents without shared tools or rules collide, repeat work and each invent their own method.

3Head chef

The chef stops cooking every dish. The job becomes ordering, storing and preparing ingredients, and setting the standard. Lauren calls the chef “the CEO of the kitchen”.

You build the environment: skills, tools, checks and a codebase that guides each agent. You remain accountable for the result.

4Restaurant group

The first kitchen runs without the chef. Gordon Ramsay opens a second and a third restaurant, and visits each one to taste.

Parallel projects run at the same time, each with a coordinator agent. You move between them and sample their output.

What each part of the kitchen stands for

Chef

You, the engineer.

You set the intent and the standard. “Your name still is associated with the work that you do.”

Executive chef

A coordinator agent: in Lauren’s setup, a Cursor project.

It delegates and supervises a list of tasks. It does no cooking of its own.

Cooks

Sub-agents that write, run and check the code.

Frontier models are capable cooks. What they lack is context and a prepared station.

Knives and a garlic press

Skills, small CLIs and verification tools.

“If you have a dull knife, then everything’s going to take a long time.” Without tools, each cook makes up a method of their own.

Kitchen layout and prep

The codebase, its types and its lint rules.

Constraints put each cook on a narrow, correct path. A new hire, human or agent, is productive from the first day.

Orders from the floor

Bug reports and requests from Slack, X, email and Linear.

Grok Bot routines watch these channels and send the work to the right project. Mental model 3 covers this outer loop.

Tasting

Sampling merged work.

A chef with several restaurants cannot taste every dish. When the same fault appears in several dishes, the chef changes the kitchen. Mental model 4 covers this.

The chef’s own knives

Each engineer’s personal set of skills.

A chef takes their knives to each new restaurant. Lauren encourages readers to build their own set from their past transcripts, and to borrow from pstack only what they trust.

Lauren explains the kitchen at 11:20, the home cook and the family at 12:52, the chef as organiser at 14:08, sharp knives at 27:09, the restaurant group at 34:14, the executive chef at 42:40, tasting at 49:34 and her own knives at 1:01:33. The “meat proxy” is at 4:52.

Mental model #2: The Trust Ladder

Matt Pocock opens the interview with Lauren’s “trust ladder”. The more you trust your agents, the larger the tasks you can give them and the more agents you can run at once. Agents earn that trust with evidence: work they can show and check. Lauren also places trust in the tools themselves: “trust to me is really about trust in your own tools.”

A robot climbs a tall ladder while Lauren kneels beside it and checks a rung with a spirit level. Books, a mug and a tool roll sit at a workbench behind her.

Each rung is a tool that lets the agent prove its own work. The human builds and checks the rungs.

How Lauren climbed it

Her own account, in order. She describes no fixed number of rungs. The steps below follow her story.

1Micromanage one agent

On a side project after Meta, she spent hours steering one agent and writing skills whose effect she could not measure. “I was kind of flying blind.”

Her first skills, which tried to teach the agent to work like her. They became the basis of pstack.

2The meat proxy

At Cursor she fixed performance problems by hand, carrying flame graphs and heap snapshots between Chrome DevTools and the agent.

The realisation that she was the bottleneck.

3Verification

Her first Cursor skill gave the agent “hands and eyes”: it could run the app, use it, debug it and take traces.

The agent closes its own loop. It can hill-climb, improving against a score without her.

4A prepared environment

She added tools, types, lint rules and constraints that “make the easy thing the right thing”.

Many agents and several projects at the same time.

5Agents merge their own code

“Full autopilot” sends verifier agents to fuzz each pull request before it lands. She reads the commit history in the morning and reverts or adds a lint rule where needed.

“I’m sleeping so much better now.” The first night was “very scary”.

Two cautions

Low on the ladder, people get stuck. Without trust you must micromanage, and micromanagement leaves no time to sharpen your knives. And the top rungs depend on how far the work can be verified. Software and some mathematics verify well. Lauren says she has no answer yet for work that is hard to verify, or where a mistake cannot be undone.

Lauren tells the story at 1:58, the meat proxy at 4:52, verification at 16:14, hill climbing at 18:01, being stuck low on the ladder at 26:14, autopilot at 52:38, verifiability at 56:28 and trust in tools at 1:01:33.

Mental model #3: Inner and outer loops

Lauren’s agents work toward “a snapshot of my intent”. The snapshot goes stale: new bug reports, requests and limits arrive in Slack, Linear, email and X. For a long time she carried that context to her agents herself. Her fix is two loops, joined. She borrows a management principle from her time at Netflix: give people the context to work alone, and they need less supervision.

An outer cycle runs from Lauren setting a goal, to Lauren with an idea at her laptop, to a courier robot carrying orders, to Lauren approving a finished dish. Inside it, robot cooks repeat a smaller cycle of planning, preparing, tasting and correcting.

The outer loop brings in fresh orders. The inner loop cooks, tastes and corrects until the dish passes.

Outer loop

Grok Bot routines with connectors to Slack, X, email and Linear.

Watch for new information and send it to the right project.

Inner loop

Cursor projects: a coordinator agent in the cloud, with its own computer.

Split the work, start sub-agents, and drive each task to a checked result.

The join

Grok Bot sends messages straight to a project.

Related reports go to one project, so one coordinator sees the shared cause and each fix happens once.

You

Set the intent and sample the results.

Ask “where am I the bottleneck?”, then teach the agent to fetch that answer from real data.

Outer loopAcross the projects

Human directionSet intent and standards
Grok BotGather fresh context
Parallel Cursor projectsCoordinate and execute work
Human judgementInspect evidence and adjust
Keep intent currentPrepare and improve the environment
Reports and telemetry supply fresh context.

Context and work

Results and evidence

Inner loopWithin each Cursor project

ImplementUse the current intent
Run and inspectCheck real behaviour
CorrectChange and recheck
Project coordinatorSelect and delegate the work
Repeat until the relevant checks pass.

This connects with my two-loop essay and When AI Leaves the Chat Window. My essay places human direction in the outer loop. Lauren also automates external context gathering. The diagram combines these roles. Review effort and intervention depend on the work and its risks.

Lauren describes the stale snapshot at 35:32, the outer loop at 36:12, the bottleneck question at 37:58, the Netflix principle at 41:20, the two tools at 43:57 and grouping related reports at 44:05.

Mental model #4: The Gardener

In Part 1 of her guide, Lauren writes that “every team needs a gardener”: someone who watches the stream of pull requests and notices the smells, such as “the third isRecord this week, the lint suppressions creeping like ivy”. Most of her 2,500 pull requests a month are gardening work: refactors, new lint rules and new checks. A tended garden helps everyone who works in it. A new hire, human or agent, writes good code from the first day.

Lauren pulls a spreading weed out by its roots from a raised bed while a robot waters healthy plants. A divider separates the beds, and a clipboard of observations stands to one side.

Pull the weed by the root, then change the bed so the same weed cannot return.

1Watch

An agent scans the code all the time for banned patterns.

Smells appear faster than one person can read code.

2Queue

The agent appends each finding to a document and fixes nothing yet.

A buffer shows patterns that a fix for each single bug would hide.

3Group

Every few days she reads the queue and finds that many entries share one cause.

In pure execution mode, “you sometimes miss the big picture”.

4Sample

Each day she reads a sample of merged pull requests closely.

At this volume she cannot taste every dish.

5Change the bed

When several agents take the same shortcut, she changes the skills, constraints, lint rules or types. A one-off mistake needs no action.

The environment then guides every later agent to the right path.

Part 1, the gardener post. Lauren discusses gardening at 45:51, new hires at 46:12, the queue at 47:41, sampling at 49:34 and changing the environment at 50:33.

Mental model #5: Hands, eyes and a map

Lauren calls verification “the single most important skill” in an agent toolkit. It gives the agent “hands and eyes”: it can run the app, use it as a person would, debug it and take traces. An agent that can see the result of its work can keep going until it succeeds, without you in the middle. She treats a good verification skill as critical infrastructure, and suggests it could have its own on-call rotation.

A robot points at a large folded map of pictorial destinations joined by dashed routes. A second robot uses a multitool on the matching button of a product panel while Lauren watches the result and holds a tablet.

The CLI is a remote control the agent can hold. The Feature Map tells it which buttons exist.

Three parts

Part 1 of her guide builds the skill from three pieces. Each one removes a reason for the agent to ask you.

Verification skill

A project skill that launches, drives and observes the real app. /create-verification-skill writes it, and /maintain-verification-skill keeps it current each day.

“Verification is all you need.” Without it, you stay the bottleneck and spend the day babysitting agents.

Small CLI

One script inside the skill, with named commands such as doctor, snapshot, screenshot and send.

Code does the mechanical parts and the agent keeps the judgement. Every agent reuses one tested tool, where each used to “rebuild the world each time”. pstack calls this principle “Build the Lever”.

Feature Map

A searchable markdown map of every feature: what it does, how a user reaches it, how to drive it with the CLI, and its traps.

“Materialised memory”. The codebase is the true memory, and the map is a compact copy that saves tokens. It needs daily upkeep, because the agent trusts it.

Part 1, verification and the Feature Map. Lauren discusses hands and eyes at 16:14, the CLI and determinism at 21:03 and rebuilding the world at 22:47.

Mental model #6: Measure a hundred times, cut once

Part 2 of her guide asks the next question: once agents can verify their work, how do you decide what to build? Lauren calls it “the art of supervising someone smarter than you”. Models now write code better than most people, in codebases too large to hold in one head. Her answer has two halves. First draw out understanding. Then plan with code and evidence. “I do plan, but I do so through code.” In her view, a long abstract plan gives “the illusion of progress”.

Lauren studies a sketch of people crossing a river. On a workbench, robots load-test a truss bridge model and an arch bridge model with weights and a gauge. Both pass, and a third model lies broken.

Define the use first. Build several designs, load-test each, and keep the one that holds.

First, understand

In your own words

“Restate in your own words and in plain english what you think the underlying issue is.”

It catches a misunderstanding before any code is written, and keeps her own guesses out of the prompt.

/how · /why · /teach

/how traces how a subsystem runs. /why searches Git history, tickets, docs and Slack for the reason. /teach explains the result to her.

It keeps her mental model current, and it makes the agent read the code before it states a claim.

/recall

Loads past transcripts into a fresh chat.

“Your past transcripts are often a gold mine for rich context.”

Then, plan with code

README first

Write the tutorial or README before the code. /technical-writing keeps tutorial, how-to, reference and explanation apart.

It gives the agent a concrete target to check its work against.

Prototype

Two or three variants behind a switcher. The verification skill drives each one and measures it.

The two common mistakes are accepting the agent’s first design and overcooking a plan with no evidence.

/architect

Several model families sketch types and signatures in parallel. A judge on a different model picks. The design is scrapped if workarounds or forced casts appear.

Architecture is now the engineer’s main job. Agents fill in the implementation.

Multi-phase plan

Written only after the design settles. Each task names its proof. The plan is deleted when the work lands.

Work counts as verified only after the code has actually run.

Part 2, research, prototypes and architecture. The prototype playbook and /architect ship inside pstack.

Part II: The package

This inventory describes Lauren Tan's published package, pstack 0.15.9 at commit e43c7ee in cursor/plugins. Cursor installs it as a plugin. The copy I run in Claude Code, Codex, Grok and Muse is a direct port of the same files. It maps Cursor's subagent, question and path conventions onto each host.

A router + a library + execution support

pstack 0.15.9, commit e43c7ee. Select a component to inspect it. Manifest · Router

One router connects the parts

The model reads instructions, selects a playbook and uses the host runtime. Mode · README

Arrows show instruction and execution dependencies at commit e43c7ee. Verification rule

Every skill, in one view

Select a name to see its job, dependencies and exact source. The five families are editorial groupings. Every skill except setup-pstack runs only when it is called. Cursor lists each one under /.

Route, design & delegate7

poteto-modesetup-pstackfigure-it-outarchitectarenaswarmshow-me-your-work

Understand & recover context4

howwhyrecallblast-radius

Verify & review7

tddcreate-verification-skillmaintain-verification-skillinterrogatebenchmark-checklistno-commentscorrect

Code & writing4

unsloptechnical-writingtypescript-best-practicesbro

Learn & extend4

teachreflectautomate-memake-bot-ui

How skills use each other

Selected dependencies from the source text. Branches show called skills, delegated agents or generated outputs.

architect
howTrace the system
whyWhen rationale is needed
arenaCompare design candidates
interrogateChallenge the design
architect source
figure-it-out
architectShape the design
arenaCompare approaches
show-me-your-workLog each decision
figure-it-out source
teach
howMechanism
whyRationale
unslopCheck the explanation
teach source
blast-radius
howTrace callers
whyRecover intent
arenaSettle an open fork
blast-radius source
no-comments
comment-sickoDelegate the comment audit
howCheck a claimed constraint
no-comments source
create-verification-skill
project verifyGenerate the app driver and feature map
create-verification-skill source
maintain-verification-skill
project verifyCheck every mapped feature
maintain-verification-skill source
technical-writing
howGround the document
unslopApply the writing-pattern checks
technical-writing source

23 routes through a task

The router selects a playbook by task. Large or unusual work goes through figure-it-out, which builds a bespoke workflow. Playbook index

Principles guide the decisions

24 principle skills. poteto-mode names when each applies, and the agent reads the leaf in full before it applies one. Principle index

Agent type and model are separate choices

The package defines 2 agents and sets no model in either. The model rule names a model for each role, and the caller passes it on each spawn. poteto-agent · comment-sicko · Setup

2 agent definitions

poteto-agentTask worker. Reads poteto-mode in full before any work, and runs in the background.
comment-sickoComment auditor. Deletes comments that restate code and flags workaround code.

Role configuration

setup-pstackModel ruleRole → model · budget
Agent dispatchAgent type + role model

setup-pstack detects the available models, asks for a budget (unlimited, large, medium or small) and writes one always-applied rule. A role without a line uses the skill's default. Setup workflow

The runtime carries out the work

The selected playbook determines the required tools.

Verify the app

Project verify skill
Feature map
Launch · drive · observe · clean up

Generated driver

Keep work state

orch · check-plan
show-me-your-work log
Units · gates · checkpoints

orch · check-plan · log

Use the host tools

Git · GitHub CLI · Bun
watch-pr
PR status · checks · review threads

watch-pr · bootstrap

The package also contains a read-only worktree audit for disk cleanup. Cursor cloud agents run the long, parallel playbooks. Worktree audit · Orchestrate

Recent upstream changes

A new skill, correct

Finds mistakes that agents repeat in a repo and makes each one impossible.

Commit 9511e60

Designs that resist agent mistakes

architect now shapes a design so later agents cannot misuse it.

Commit a586282

Explain the Number

A measurement principle, fresh subagents per task and an hourly autopilot tick.

Commit 23e4138

Complete source catalogue

Lauren Tan created pstack. Source links open the public files at commit e43c7ee, which fixes this October 2026 snapshot. Package README

The package inventory

PackageVersionSkillsPrinciplesPlaybooksAgents
cursor/plugins pstack0.15.9 · e43c7ee2624232

Counts come from the skill folders, the playbook files and the agent files at that commit. Every skill sets disable-model-invocation: true except setup-pstack, so a skill runs when the user or poteto-mode calls it. Manifest

The skills have distinct jobs

The five groups below are editorial categories. Open a group to inspect its commands and dependencies.

Routing, design and delegation · 7 skills
CommandJobDependencies and conditionsEvidence
/poteto-modeSelects a task playbook and applies the shared working rules.Reads a playbook, the relevant principles and any named skill.source
/setup-pstackSets the model for each role at a chosen reasoning budget.Writes an always-applied model rule. Offers create-verification-skill.source
/figure-it-outDesigns an auditable workflow when no playbook fits.Uses architect, arena and show-me-your-work.source
/architectSketches types, signatures and modules before code, then stays with the build.Uses how and why. Can call arena and interrogate.source
/arenaRuns parallel candidates, picks a base and grafts in the best parts of the others.Spawns candidates and a cross-judge.source
/swarmFans out parallel workers and returns one report.Spawns workers on the swarm role model.source
/show-me-your-workKeeps a decision log a reviewer can inspect.Appends rows with its log script.source
Understanding and context · 4 skills
CommandJobDependencies and conditionsEvidence
/howExplains how a subsystem works, from source and runtime flow.Spawns explorers and one explainer.source
/whyReconstructs why a design is the way it is, with stated confidence.Queries each evidence source it can reach. Uses how.source
/recallRebuilds your working context from history, live state and the shared record.Uses why.source
/blast-radiusFinds what a change can break beyond the diff and proves the safety fact.Uses how and why. Can call arena.source
Verification and review · 7 skills
CommandJobDependencies and conditionsEvidence
/tddProves a focused test fails before the fix and passes after it.Used by the bug-fix playbook.source
/create-verification-skillGenerates a project skill that drives the app as a user does.Produces the project verify skill and feature map.source
/maintain-verification-skillChecks every mapped feature against source and the live app.Keeps the project verify skill current.source
/interrogateHas several model reviewers challenge a change from separate angles.Spawns the interrogate reviewer panel.source
/benchmark-checklistChecks that a measurement tests the work it claims to test.Pairs with the Explain the Number principle.source
/no-commentsRuns a comment audit and acts on the accepted findings.Spawns the Comment Sicko agent. Uses architect, how and why.source
/correctFinds mistakes agents repeat in a repo and makes each one impossible.Prefers architecture, then types, lints and tests, then docs.source
Code and writing · 4 skills
CommandJobDependencies and conditionsEvidence
/unslopRemoves AI writing patterns from any prose.Applied to every reply and document.source
/technical-writingApplies a layered standard for technical documents.Uses how and unslop.source
/typescript-best-practicesApplies TypeScript rules when a .ts or .tsx file is read or edited.Loads on TypeScript paths.source
/broRestates the last message in plain language.Reads the previous reply.source
Learning and extension · 4 skills
CommandJobDependencies and conditionsEvidence
/teachExplains a body of work so a person understands it.Uses how, why and unslop.source
/reflectReviews the session and routes each lesson to a skill edit.Spawns three reviewers over the transcript.source
/automate-meTurns your working preferences into a personal mode skill.Builds on poteto-mode. Uses unslop.source
/make-bot-uiBuilds a small UI that wakes a Grok Bot through a webhook.Optional exposure on Tailscale.source

The playbooks cover 23 task types

Each row quotes the condition poteto-mode uses to select the playbook.

All 23 playbooks
PlaybookSelected forEvidence
Authoring or modifying a skillWriting or editing a SKILL.md.source
Autonomous runA long task to drive to completion without stopping ("run until done", "/loop until X").source
Autopilot-fullA queue of independent PRs run to merged with full autonomy. One owner per PR carries build through merge, and the root swarm-verifies each PR before its owner merges ("autopilot this queue", "full autopilot", one-owner-per-PR programs).source
Autopilot-stackA queue of changes built and verified with full autonomy, delivered as one linear reviewed base-branch stack the operator lands ("autopilot-stack", "stack them, don't ship", "build the stack, I'll land it").source
BabysitDriving a PR or a stack to merge-ready: conflicts, review threads, CI.source
Bug fixA reported defect to reproduce, root-cause, and fix with runtime evidence.source
EvalTesting how a skill, structure, or prompt change affects agent behavior before promoting it.source
FeatureNew or changed behavior, built from a named data shape.source
HillclimbSustained, scientific improvement of one metric against a target: loop hypotheses with before/after measurement, a decision log, and one commit per accepted win. Distinct from Perf issue, which is a one-off fix.source
InvestigationRead-only question: how does X work, why was Y built this way, are we sure about Z, should we do X or Y.source
Multi-phase or multi-PR planWork that spans phases or stacked PRs.source
Opening a PRInvoked at the end of every other playbook.source
OrchestrateA standing project handed to one coordinator chat: multi-day, many stacked PRs, dozens to hundreds of subagents, minimal human turns ("run this whole project", "own this migration until it lands"). Distinct from Autonomous run, which drives one task to a predicate. Work one agent could finish inside the session's budget routes there, not here, however program-shaped the phrasing sounds.source
Pause safelySuspending in-flight work cleanly so it can be resumed, on an explicit pause, going offline, a Cursor restart, or imminent context compaction. The complement to Session pickup. Full steps:source
Perf issueA measured slowness to trace and improve against a baseline.source
PrototypeA throwaway sketch to make a design or behavioral decision cheaply, or to settle an empirical fork by observing it instead of asking the human ("prototype", "mock it up", "try this layout", "sketch it to decide").source
RefactoringA behavior-preserving change to structure or shape (rename, extract, inline, dedupe, move).source
Runtime forensicsDiagnose a runtime symptom (leak, idle-CPU spin, glitch) from live instrumentation. The deliverable is a diagnosis, not a fix.source
Session pickupResuming or taking over a prior agent's in-flight work from a transcript, cloud-agent URL, or pushed branch.source
ShippingThe half after Babysit. Independently verifying a green stack, then landing the contiguous verified run bottom-up through gh by default or Origin when its CLI is available.source
Trace forensicsDiagnose a captured profiling artifact (cpuprofile, trace, spindump, heap snapshot) handed to you after the fact. The deliverable is a diagnosis, not a fix.source
Visual parityPixel-exact UI equivalence: matching two implementations or migrating a styling system.source
Worktree and simulator cleanupReclaiming local disk by pruning merged or abandoned git worktrees and stale iOS simulators ("what's using my disk", "clean up worktrees", "prune safe-to-prune worktrees", "free up space", "delete old simulators").source

The principles apply across the workflows

Each row quotes when the principle applies, from the poteto-mode index.

Core · 10 principles
PrincipleApplies whenEvidence
Laziness ProtocolRefactoring, sizing a diff, or tempted to add abstractions, layers, or signal threading. Bias to deletion and the smallest change that solves the problem.source
Foundational ThinkingBefore writing logic: core types and data structures, scaffold-vs-feature sequencing, what concurrent actors share.source
Redesign from First PrinciplesIntegrating a new requirement into an existing design. Redesign as if it had been foundational from day one.source
Attack the PremiseTwo or more fixes that share one premise have failed the same gate. Take a census of which actors hold the imbalance before the next fix, then question the premise instead of writing another fix that assumes it.source
Subtract Before You AddSequencing an addition, refactor, or rewrite. Remove dead weight first, then build on the simpler base.source
Minimize Reader LoadReviewing or shaping code that's hard to trace. Count layers and hidden state, collapse one-caller wrappers, shrink mutable scope.source
Outcome-Oriented ExecutionPlanned rewrites and migrations with explicit phase boundaries. Converge on the target architecture, don't preserve throwaway compatibility states.source
Experience FirstProduct, UX, or feature-scope tradeoffs. Choose user delight over implementation convenience.source
Exhaust the Design SpaceA novel interaction or architectural decision with no precedent. Build 2-3 competing prototypes and compare before committing.source
Build the LeverAny non-trivial work. Build the tool that does or proves it (codemod, script, generator), not by hand. The tool is the artifact a reviewer reruns.source
Architecture · 6 principles
PrincipleApplies whenEvidence
Model the DomainWriting stateful logic, or code that branches a lot or repeats a shape assumption across files. Encode the domain in a structure (state machine, typed model, table or registry, reducer, boundary, the right collection) instead of scattered conditionals.source
Boundary DisciplineWiring validation, error handling, or framework adapters. Guards at system boundaries, trust internal types, keep business logic pure.source
Type System DisciplineDesigning types or a signature in any typed language. Make illegal states unrepresentable, brand primitives, parse external data at boundaries.source
Make Operations IdempotentDesigning commands, lifecycle steps, or loops that run amid crashes and retries. Converge to the same end state.source
Migrate Callers Then Delete Legacy APIsIntroducing a new internal API while old callers exist. Migrate and delete in one wave.source
Separate Before Serializing Shared StateConcurrent actors might write the same file, branch, key, or object. Eliminate the sharing first.source
Verification · 5 principles
PrincipleApplies whenEvidence
Prove It WorksAfter a task, before declaring done. Verify against the real artifact, not a proxy or "it compiles".source
Fix Root CausesDebugging. Trace each symptom to its root cause, reproduce first, ask why until you reach it.source
Sequence Work into Verifiable UnitsMulti-step work (sweeps, migrations, runs of similar edits) and how you stack commits and PRs. Break work into small units that each end in a check, verify each before the next, and order delivery so the sequence proves itself.source
Test Behavior, Not ImplementationWriting, changing, or keeping a test. Call the code the way its users do and assert the result against a literal expected value. If the test would still pass when every imported function returns undefined, rewrite the assertion or delete the test.source
Explain the NumberBefore you trust, report, or act on a number you measured (a speedup, a regression, a throughput, a latency, or an eval result). Find what limits it, and rule out that it measured something other than the work you think.source
Delegation · 2 principles
PrincipleApplies whenEvidence
Guard the Context WindowContext fills up: large outputs, long files, repeated reads, fan-out planning. Route bulk to subagents, keep summaries in the main thread.source
Never Block on the HumanTempted to ask "should I do X?" on reversible work. Proceed, present the result, let the human course-correct.source
Meta · 1 principle
PrincipleApplies whenEvidence
Encode Lessons in StructureYou catch yourself writing the same instruction a second time. Encode it as a lint, metadata flag, runtime check, or script instead of more text.source

The helpers complete the stack

HelperRoleEvidence
orchA work-state store for long projects: units, gates, checkpoints and an inbox.source
check-planChecks that each planned PR has its unit, live and performance boxes ticked.source
watch-prWatches PR status, checks and review threads for the babysit and shipping playbooks.source
bootstrapInstalls the script dependencies with Bun.source
worktree-auditClassifies git worktrees for cleanup and deletes nothing.source
log.shAppends one well-formed row to a show-me-your-work decision log.source

Source and verification scope

Verification depends on a driver for the real app. The package supplies skills to generate and maintain one. Generation · Maintenance

The skill defaults name Cursor model slugs. On another host, the model rule must name that host's models.

Upstream changes most days. The links above stay fixed at commit e43c7ee. The current package is in cursor/plugins.

Part III: Read and watch Lauren

The mental models above come from these four sources, listed in the order they appeared.

The package: pstack on GitHub, cursor/plugins