AI Tools & Skills
A working guide to the AI tools I use. The model reasons; the harness supplies context, tools and a way to work.
AI Tools & Skills
A working guide to the AI tools I use. The model reasons; the harness supplies context, tools and a way to work.
Updated 22 September 2026. Screenshots show earlier product pages and sessions; linked sources carry current product details.
An agent = its model + its harness
The model receives context and proposes responses or tool calls. The harness loads instructions and evidence, runs permitted tools and returns their results to the model. It also manages session history, context limits and approval settings. I shape that harness through the five parts shown here.
01 / 10
Device & getting started
Start with an agent account and a computer that can run its harness.
Start with one agent
Choose an agent and sign in through its supported account or API. The provider links below explain current access and billing.
GitHub
I keep code and project history here. The agents use repositories, issues and pull requests to organise changes.
Account
The device
The machine runs the harness. The model can run in a provider’s data centre or on the machine itself.
Hosted models
Claude Code, Codex, Muse and Grok run locally while calling hosted models. Memory use also depends on the files, browsers and development tools open alongside them.
Internet connection required
My local machine
I use a MacBook Pro with an M5 Pro chip and 48 GB of memory. The local model shares that memory with macOS and the other applications.
My setup · September 2026
Local inference
Model weights and the context window both use memory. I check the model’s requirements and measure it on the machine before choosing a size.
See the local stack below
02 / 10
Models & providers
I choose a model for the task and the allowance available. The harness gives it access to the work.
Hosted providers
These supply the models behind my agents and image tools. Accounts, plans and API billing differ by product.
Anthropic
Claude models, used through Claude Code and the Claude app.
Hosted models
OpenAI
GPT models, Codex and ChatGPT. My Hermes server also uses OpenAI through its supported Codex sign-in.
Hosted models
Meta
Muse models, used through the Muse Code terminal agent.
Hosted models
xAI
Grok models, used through Grok Build on the command line.
Hosted models
Gemini models, including image generation and editing in AI Studio.
Hosted models
My local stack
Model inference runs on the Mac. I use MLX to serve the model and separate clients to work with it.
MLX LM + Qwen
MLX LM serves a Qwen model on my M5 Pro through a local API. My short command, qw, opens a terminal chat client connected to that server.
Local model server
Goose
An agent that can run tools and work with files. My Goose configuration connects to the same local MLX server.
Local agent client
Model comparison
I use these to shortlist models, then test them on my own work. Speed, task performance and cost are separate measures.

Benchmarks
Artificial Analysis
Free
Independent benchmarks that compare models and providers across intelligence, speed and price, so you can pick the right one for a job.

Leaderboard
Arena
Public leaderboard
A leaderboard based on human votes between model responses. I use it alongside task benchmarks.
03 / 10
Interfaces & agents
The agents and applications I work through. Each harness loads context, runs tools and keeps a session moving.
Coding agents
Four daily agents: Claude Code, Codex, Muse Code and Grok Build. My local shortcuts are cc, cx, ms and gk.

Coding agent
Claude Code
My shortcut · cc
Anthropic’s coding agent. I use it to read project files, plan changes, edit code and run checks.

Coding agent
Codex
My shortcut · cx
OpenAI’s coding agent. I use it for coding, research and review, with the same project files and shared instructions.
Muse Code
Meta’s terminal coding agent. It works with project files and shell tools, and supports skills, plugins and MCP connections.
My shortcut · ms
Grok Build
xAI’s terminal coding agent. I use it alongside Claude Code, Codex and Muse, with shared project instructions and skills.
My shortcut · gk
The terminal
The workspace that holds my agent sessions.

Terminal
cmux
Open source · Free
A native macOS terminal built on Ghostty, with vertical tabs, split panes and notification rings - built for running several AI coding agents side by side. It is the piece of this whole stack I touch most.

cmux setup guide
Running agents in parallel
Free
Artem Zhutov’s guide to organising agent sessions in cmux, with an orchestrator and a view of the work in progress.
Desktop apps
I also use desktop chat apps for conversation, images and voice.
Claude
Anthropic’s desktop app. Good for thinking a problem through before any of it reaches a repository.
Desktop chat app
ChatGPT
OpenAI’s desktop app. I use it for conversation, image work and voice.
Desktop chat app
Menu-bar tools
Small native apps that keep an agent fleet in view.

Menu bar
CodexBar
Open source
Shows provider usage and reset times in the macOS menu bar. I check it before assigning work.
RepoBar
Keeps GitHub work in view from the menu bar - CI status, open issues, pull requests, releases and rate-limit health for every repo you follow.
Open source · Free
Desktop utilities
Small tools that support the workspace around the agents.
Magnet
I use Magnet to arrange application windows side by side. It snaps them into place with keyboard shortcuts or by dragging.
Window management
Plash
Displays a web page as the Mac’s desktop wallpaper. I use it for a custom calendar that updates the date.
Desktop calendar
Session controls
Built-in commands control the harness. Claude Code and Codex both expose /model for model selection. Use each agent’s help for its command list.
Aliases
An alias gives a launch command a short name. I use cc for Claude Code, cx for Codex, ms for Muse and gk for Grok. The model comes from the agent’s current settings.
My local shortcuts
04 / 10
Context & knowledge
How I feed and maintain the context the agents work from.
The vault
I keep separate vaults for different work. Each is a folder of Markdown files and attachments.

Knowledge base
Obsidian
Free · optional paid services
My notes and project evidence live in local Markdown files. Agents can read them with file tools. A hosted model receives the content loaded into its context.
Plugins & knowledge bases
The views, capture tools and reference structures I use with my notes.

Obsidian plugin
Dataview
Open source
Turns your notes into a queryable database - list, table and task views generated live from frontmatter and inline fields.

Obsidian core
Graph
Free with Obsidian
Every note a node, every link an edge, so the shape of what you know is something you can see.
Web Clipper
Saves web content into a vault with source details and metadata. I check the saved text before using it as evidence.
Sync
Synchronises my vaults across devices, with end-to-end encryption and version history.
Paid service
Excalidraw
Hand-drawn diagrams and sketches you draw straight inside a note, kept alongside the writing.
Open source
Bases
Obsidian’s built-in way to turn a folder of notes into a database view - filter, sort and group without leaving the vault.
Free with Obsidian
LLM wiki
Andrej Karpathy’s pattern for a Markdown wiki maintained with an agent. I use a wiki for reusable methods and keep links back to the source evidence.
Marker
Converts documents into Markdown and structured text. I check the extracted tables and figures against the source.
Source available
05 / 10
Connections & GitHub
MCP servers, APIs and command-line tools give the harness access to other systems.
MCP tools
An MCP server exposes tools the harness can call. These are the connections I use for files, repositories and browser checks.
QMD
Searches my indexed Markdown collections with lexical and vector retrieval, then reranks the results. Agents can call it through MCP.
Local search
Playwright MCP
Gives an agent browser tools to open pages, click controls and inspect the result. I use browser checks when building websites. Playwright also has a code API for scripted checks.
Browser connection
GitHub MCP server
Gives an agent tools to read repositories and work with GitHub issues and pull requests.
Repository connection
GitHub & the command line
The shell is another way to reach tools. GitHub keeps the project history and the decisions around each change.
gh CLI
The agents use GitHub’s command-line client for issues, pull requests and build checks.
Command-line tool
Git worktrees
Each agent gets a separate checkout and branch. This keeps concurrent edits in separate working directories.
Built into Git
Private network
Tailscale
Connects my machines through a private network. It supports remote access to services and SSH across that network.
Device connections
Browser & X connections
These connect the agents to web pages and the posts I read.
Claude in Chrome
Connects Claude Code to the Chrome extension so it can read pages and interact with browser tabs.
Browser extension
twitterapi.io
Hermes uses this third-party API to fetch posts from the accounts I track and prepare the X digest. Its fallback is X’s official MCP service, reached through xurl.
X capture on my server
06 / 10
Skills & plugins
Ways of working, written down. A skill holds instructions in a SKILL.md file, with scripts and references when needed. The agent reads it when the task calls for that workflow.
Third-party skills I use
These are the skills I return to across my work. Check each project’s installation instructions for the agent you use.
The three I use most are grill-me, wait-what and wayfinder.
grill-me interviews me to sharpen a plan. wait-what asks the agent to explain its last answer again, with context and simpler language.
wayfinder supports my graph engineering: it maps decisions and their dependencies in an issue tracker, then works through them across sessions. I use it to plan work that exceeds one context window.
Available as a Claude Code plugin or editable skill files for Codex and other agents.
A favourite for presentations in the browser. It builds an HTML deck from a brief or converts a PowerPoint file. You choose a direction from visual previews, then refine the slides.
The output is a single HTML file with its styles and scripts included. It ships as a Claude Code plugin; other coding agents can read its skill files.
The design skill I use for websites and interfaces. It helps set the direction, then critique and refine the result. I use it to check the page in the browser at desktop and mobile widths.
Other skills I return to
Smaller workflows for reviewing a session, researching a topic and working with documents.
Artem Zhutov’s skills
retrospective proposes improvements from a completed session. handoff carries its state into a new session. skill-management organises larger workflows into files loaded as needed.
Session and skill management
gstack
Garry Tan’s collection of workflows for planning, design, engineering and quality checks.
Engineering workflows
last30days
Matt Van Horn’s research skill searches recent discussion across social and web sources, then produces a summary with sources.
Research
Codex plugin for Claude Code
Lets Claude Code ask Codex to review code or take on a task.
Agent delegation
hands-on-deck
Helps agents inspect, edit, create and check PowerPoint files.
PowerPoint
Tufte data viz
Caylent’s skill applies Tufte’s principles to charts: remove unnecessary marks, label clearly and keep the evidence readable.
Charts
Skills from model providers
Providers also supply skills. The available set depends on the product, account and installation.
Anthropic
Document skills cover PDF, Word, Excel and PowerPoint. Claude Code also bundles workflows such as /loop and /claude-api. Check the skills available in your session.
Documents and coding workflows
OpenAI
Codex includes system skills to create and install skills. My setup also has OpenAI skills for documents, spreadsheets, presentations, PDFs and images.
System skills and installed plugins
Meta
My Muse installation bundles skills for planning, Git, browser delivery and plugin creation.
Bundled Muse skills
xAI
Grok Build supports provider and third-party skills through its plugin marketplace.
Skills and plugins
How skills, plugins and commands fit together
A skill describes a workflow. A plugin is an installation package: it can contain skills, tool connections and hooks. I group them here by the work they help me do.
Built-in commands control the application. For example, /model changes the model in Claude Code and Codex. A slash can also invoke a skill: Claude Code’s /loop runs a bundled workflow. The prefix alone does not tell you which kind it is.
Claude plugins ↗ · OpenAI plugins ↗
My own skills
I also write skills for my own work. Public examples are being prepared; downloads will appear here once they are published.
07 / 10
The engineering process
A repeatable process gives the agent a sequence to follow and a result to check.
From every.to
The process
Every piece of work should leave the next one easier.

Skills & plugins
Compound engineering
Engineering workflow package
Every’s skills organise work into a cycle of planning, implementation, review and recorded learning. I use the package across Claude Code and Codex.
The loop
Brainstorm → Plan → Work → Review → Compound. Record what worked so the next change can use it.
Brainstorm
Clarifies an idea and agrees the requirements before coding.
Plan
Researches the codebase and writes an implementation plan.
Work
Executes the agreed plan and checks the result. Use an isolated worktree when changes run in parallel.
Review
Reviews the change for correctness and risks, using the specialists needed for the task.
Compound
Writes the solved problem into docs/solutions/, so the next run starts ahead.
LFG
Coordinates the workflow from a requirement through implementation and review.
08 / 10
Loops & hosted agents
The workflow layer also includes repeated tasks and scheduled work on a server.
Loops & autonomy
A loop needs a task, a check and a stopping condition. A schedule controls when it starts.

Loop library
Loop Library
Free
Forward Future’s collection of agent workflows with success checks and stopping conditions. I use it as a source of patterns for repeated work.
looper
Helps design an agent loop and review its scope, checks and stopping conditions before it runs.
Open source · Free
autoresearch
Andrej Karpathy’s research loop runs short model-training experiments and keeps changes that improve the chosen metric. I use the keep-or-revert pattern in other work too.
Open source
Hooks run commands at selected events in an agent session. Scheduled jobs start work at set times. Scripts hold the repeatable steps that those hooks and jobs call.
Hosted agents
The always-on side: an autonomous agent on a rented server, working while the Macs are asleep.

Hosted agent
Hermes
Open source · my server uses Codex sign-in
My server agent runs scheduled work: market data, X and newsletter digests, investment analysis and vault maintenance. I inspect its output and session records.
Hosted server
I run Hermes on a Hostinger VPS so scheduled jobs can continue while the Macs are asleep. The server holds the runtime and its working files.
My always-on device
09 / 10
Voice, image & making things
The specialist tools: pictures, voice, and the things I make with them.
Images
The tools I use to generate and edit images.

Image generation
ChatGPT Images
Image generation and editing
I use it to generate images from a brief and refine them through conversation. The current release is Images 2.5, checked in September 2026.
Google AI Studio
I use Google’s interface for Gemini image generation and editing. The available models and limits are listed in AI Studio.
Image generation and editing
Screenshots
A screenshot helps me show an agent what needs to change.
Shottr
My screenshot tool on the Mac. It captures scrolling pages, adds annotations and extracts text from an image.
Screen capture and markup
Voice
A lot of my first drafting happens out loud.

Dictation
FluidVoice
Open source · dictation
I use this Mac app to dictate prompts and drafts. It supports local transcription. I check names and file paths before sending the text.
ChatGPT voice
I talk a problem through in the ChatGPT app on a walk, then return to the notes at a keyboard.
Included with ChatGPT
Otter
I use Otter to transcribe voice recordings and meetings, then turn the transcript into a structured note.
Voice and meeting transcription
Design & making
For presentation workflows, see frontend-slides and hands-on-deck in Skills. I use the references below for the format and design.

Typography
Tufte CSS
Open source · Free
Tools for styling web articles after Edward Tufte’s books - generous sidenotes, tight text-and-figure integration and careful typography.
Occasional design reference
A smaller part of my setup now.
Figma
I have used it for interface design and shared design files. Most of my current interface work goes directly into code with an agent.
Occasional reference
Reading Markdown
A simple way to read an individual file.
Typora
I use Typora to read Markdown files in a clean, formatted view.
Paid app
10 / 10
People
Who I learn from. Several of the tools above came off these timelines before they reached anywhere else.
People I follow on X
The people whose tools and explanations I return to.
Matt Pocock
I use grill-me, wait-what and wayfinder from his skills collection. His work connects clear requirements to engineering practice.
Peter Steinberger
The developer behind CodexBar and RepoBar, two tools I use to keep agent and repository activity in view.
Artem Zhutov
His session-management skills and cmux guide helped shape how I work with agents.
Andrej Karpathy
The LLM wiki pattern in section four and the autoresearch loop in section eight are both his.
Garry Tan
gstack, his opinionated Claude Code setup, is in section six.
Matt Van Horn
The last30days research skill in section six.
Thariq Shihipar
Engineer at Anthropic working on Claude Code.
Boris Cherny
Head of Claude Code at Anthropic. Shares how the team uses coding agents in its own work.
Matthew Berman
AI educator and YouTube creator covering new models, tools and practical tutorials.