Skip to content

Agent tools (MCP)

Every tool Warpforge hands your agents — what each one is for, when to reach for it, and which ones only an orchestrator gets.

Updated View as Markdown

Most coding agents can read and write files and run shell commands. Inside Warpforge they get more: the running app, its logs, the task board, and a memory that outlives the session — as tools, not as instructions you paste into a prompt.

Warpforge exposes these over the Model Context Protocol. Every agent session gets them automatically; nothing to configure.

Two toolsets

Which tools an agent sees depends on how the task was started:

Regular task Orchestrator task
Runtime, backlog, memory (24 tools) ✅ ✅
Automations (7 tools) ✅ ✅
Factory (2 tools) ✅ ✅
Browser (6 tools, when the task belongs to a project) ✅ ✅
Visual replies (2 tools) ✅ ✅
Advisor (1 tool, when the task has an advisor) ✅ —
Sub-agents and pipelines (11 tools) — ✅

A regular session is one agent doing the work. An orchestrator session is an agent that can also delegate — so it gets 11 extra tools for dispatching sub-agents and driving pipelines. See Choosing your mode for which to reach for.


Runtime — see and drive the running app

Available in every session.

list_runtime

Lists the project’s services and port-forwards with live status, allocated ports, and URLs.

Reach for it first. It’s the discovery call — names, ports, and the logSeq cursor the log tools take. An agent that guesses localhost:3000 is wrong in Warpforge, where every project gets its own port range.

Each service also reports its checkout — the directory it runs from. Services run from the project root, so an agent working in a task worktree is not testing its own edits by restarting them.

list_runtime(project?)

read_service_logs · read_portforward_logs

Read a window of a service’s retained stdout/stderr. These behave like kubectl logs --timestamps | grep | tail:

Argument What it does
service / name Which one to read (from list_runtime)
after Monotonic sequence cursor. Start at 0, then pass the returned nextSeq to cheaply poll for new lines — stable even as old lines age out of the ring buffer
before Optional upper sequence bound: only lines below it. With after it reads an exact range — the one a log chip in the chat points at
limit Max lines, newest kept (default 100)
filter Case-insensitive substring, run over the whole buffer before limit is applied — so a match deep in history isn’t lost
context N lines around each match, like grep -C
timestamps UTC timestamps prepended, Z-suffixed (default on)

When: after a change, to confirm the app actually came back up. After a failure, to find out why. The after cursor is what makes polling cheap enough for an agent to watch a restart.

service_start · service_stop · service_restart · portforward_start · portforward_stop

Control a service or port-forward. All dispatch asynchronously and return immediately — the operation is still in flight when the call returns.

How to use them properly: call, then follow with read_service_logs to watch the outcome. An agent that assumes success because the call returned is an agent that reports a broken app as fixed.


Backlog — file, read and close local backlog items

create_backlog_task

Creates a local backlog item without auto-running an agent. It lands in the backlog with status todo by default.

create_backlog_task(title, project?, body?, priority?, status?, tracker?)

title is required. Older agents that send prompt instead are still accepted; project defaults to the one the session is working in.

Pass tracker (github or linear) to also open the issue in that tracker and link it to the item, the same as New work item does in the backlog. Agents set it only when you ask for a GitHub or Linear issue. If the issue cannot be created, the item is removed too, so nothing claims to live in a tracker it never reached.

When: an agent finds real work that isn’t the work it was asked to do — a bug next door, a missing test, a refactor the change made obvious. Instead of scope-creeping the current diff or silently dropping it, it files it.

Why it doesn’t auto-run: discovered work is exactly the work a human should triage. The deprecated create_task(prompt, ...) alias is retained for compatibility and now creates the same local backlog item.

list_backlog_tasks · get_backlog_task

Read the backlog instead of guessing what is in it.

list_backlog_tasks(project?, status?, priority?, search?, limit?)
get_backlog_task(number | id, project?)

list_backlog_tasks prints one line per item, most recently changed first: #87 [todo] [high] Fix the login redirect, then the total. limit defaults to 50 and tops out at 100; search matches the title and body. get_backlog_task returns one item in full — body, status, priority, and the linked task or tracker URL.

Items are addressed the way you write them: number is the 87 in “#87”. id also works. An unknown number is an error, not a guess.

update_backlog_task · close_backlog_task

Keep an item current as the work moves.

update_backlog_task(number | id, project?, title?, body?, status?, priority?)
close_backlog_task(number | id, project?, status?, note?)

update_backlog_task changes only the fields you pass and returns the updated item. body replaces the whole body, so read the item first if you mean to add to it. close_backlog_task sets the status to done, or to cancelled when you pass status: "cancelled" for work that won’t happen, and appends note to the body as Closed: <note>.

Statuses are todo, in_progress, waiting, done and cancelled; priorities are none, low, medium, high and urgent. Anything else is refused with the list of valid values. Edits change the local item only — they are not pushed to a linked GitHub or Linear issue.

When: an agent finishes work that a backlog item described, or is asked to triage “#87”. Closing the item with a note leaves the trail for whoever reads the backlog next.


Factory — start Factory tasks for backlog items

Available in every session. A Factory task runs a backlog item through a workflow template — implement, an optional check in the running app, review ⇄ fix — and can open a draft pull request when it succeeds.

runner_enqueue · runner_status

runner_enqueue(number | numbers, project?, workflow?, agent?, model?, run_location?, pull_request?)
runner_status(project?, runs?)

runner_enqueue starts one Factory task per item, all with the same configuration: one item (number, as in #87) or several (numbers, started in that order).

Argument What it does
workflow The workflow template id. Defaults to the project’s Factory default.
agent / model The lead agent and model, for every stage the template does not give its own agent. Default to the project’s Factory defaults.
run_location Where each task runs: default (Automatic — your project folder when the template tests the running app, a background copy otherwise), worktree (a background copy of the repository) or checkout (your project folder, one task at a time).
pull_request true (the default): a successful run is committed, pushed and opened as a draft pull request. false: the change is left for a person to review and commit.

Each task appears in the sidebar at once and starts as soon as the project’s Factory limits allow, showing Queued until then. The reply says which tasks started and which are queued, and which items were skipped and why — an item that already has a Factory task, or one that is done or cancelled.

runner_status reports the project’s Factory limits, why queued tasks are waiting, and every Factory task that is queued, running or in review, with its pull request. With runs: true it adds the ten most recent runs with their outcome and cost.

When: you ask a chat to start the next few backlog items in the Factory, or to report how the project’s Factory work is going. The project’s limits, not the agent, decide when each task starts. For one goal that is not a backlog item, an orchestrator uses spawn_workflow with pull_request: true.

Why a Factory task cannot call it: a task that started more work would keep the Factory busy on work no person chose. runner_enqueue from inside a Factory task is refused.


Memory — knowledge that outlives the session

Available in every session. One store, shared across Claude Code, Codex, and opencode — see Cross-harness memory for the full picture.

memory_store

Persists a durable fact.

memory_store(content, scope?, kind?, tags?, project_id?)
  • scope — global (every project) or project (this one). Defaults to project when a project is in play.
  • kind — fact, decision, preference, gotcha, or note. Defaults to note.
  • tags — free-form, and they’re searchable.

What belongs here: decisions and their reasons, gotchas that cost someone an hour, preferences that would otherwise be re-litigated every session. What doesn’t: anything already obvious from the code — memory is for what the codebase cannot tell you.

Full-text, relevance-ranked search. With embeddings enabled it becomes a hybrid search that fuses keyword and semantic ranking, so a query phrased differently than the stored note still finds it.

memory_search(query, scope?, limit?, mode?)

When: at the start of a task, before making assumptions. This is the tool that turns “the new agent doesn’t know why we did that” into “the new agent already knows.”

memory_list · memory_update · memory_delete · memory_stats

  • memory_list(scope?, kind?, limit?, offset?) — browse most-recently-updated first.
  • memory_update(id, content) — rewrite a memory whose facts changed.
  • memory_delete(id) — permanent, and only ever an explicit action. Nothing is auto-deleted.
  • memory_stats() — counts and which scopes are live. Cheap; useful for an agent to check what it’s working with before leaning on memory.

memory_addEdge · memory_edges

Link two memories with a labelled, directed relation, and read a memory’s links back.

memory_addEdge(src_id, dst_id, relation)
memory_edges(id)

When: a decision supersedes an older one, or a gotcha explains a convention. Edges keep the why attached to the what instead of leaving two unrelated notes.

memory_dream · memory_list_compaction · memory_resolve_compaction

Memory that only grows eventually rots. Dreaming is the pass that finds duplicates, contradictions, and stale facts — see Dreaming.

  • memory_dream(dry_run?) — run the pass. Proposals are written to a log, never applied.
  • memory_list_compaction() — the pending proposals: id, type, targets, reason.
  • memory_resolve_compaction(id, approve) — approve or reject one.

The important part: a dream never edits memory on its own. It proposes; a human (or an agent that verified against the code) decides.


Automations — schedule a prompt to run itself

Available in every session. The same automations the desktop screen manages — an agent that notices work worth repeating can schedule it instead of telling you to.

automation_create

automation_create(name, prompt, agent, project?, model?, preset?, cron?, timezone?,
                  precheck?, missed_run_grace_minutes?, reuse_session?, worktree?)
  • project — defaults to the project the session is working in.
  • preset — hourly, daily, weekdays, weekly, or custom. With no preset and no cron it is a daily 09:00 job; a cron with no preset is treated as custom, never as a silent daily.
  • cron — 5-field (min hour dom month dow), required for custom and ignored otherwise. Mind the two ways this cron is not Unix cron. The every-5-minutes preset has no name here — write */5 * * * *.
  • timezone — IANA name; empty resolves the daemon host’s zone and stores it.
  • precheck — shell command run before each scheduled run; a non-zero exit skips the run.
  • missed_run_grace_minutes — catch-up window for an occurrence the daemon was down for, default 720.
  • reuse_session — send every run into the previous run’s task instead of creating one per run.
  • worktree — run each run in an isolated git worktree. The tool always starts them from the current branch; pick another base in the desktop app.

automation_list · automation_get · automation_runs

  • automation_list(project?) — next run time, last status, enabled state.
  • automation_get(id) — one automation in full.
  • automation_runs(id, limit?) — run history newest-first: status, the occurrence it belongs to, the task it used, and an output excerpt.

When: before editing. The last few runs say whether a schedule is doing anything useful — an automation whose history is all Skipped · precheck is a precheck problem, not a prompt problem.

automation_update

automation_update(id, name?, prompt?, agent?, model?, preset?, cron?, timezone?,
                  precheck?, enabled?, missedRunGraceMinutes?, reuseSession?, worktree?)

Every field is optional and absent means unchanged, so flipping enabled is a two-key call. model and precheck also accept null, which clears them — that is the only way to hand an automation back to the agent’s default model or remove a precheck.

automation_run_now

Runs an automation immediately. It skips the precheck (an explicit ask is its own authorization), is refused while a run is already in flight, and does not move the next scheduled occurrence. If the agent’s account is out of quota the run is recorded as Skipped · quota rather than started.

automation_delete

Deletes the automation and its run history permanently, and stops a run in flight. Tasks the automation already created stay.


Browser — act in the in-app browser

Available in every session bound to a project, while the desktop app is open. The tools work in the agent’s own tab in the project’s Browser, signed in as you: browser_navigate opens it and makes it active, and every other tool acts in it, never in the user’s tabs. Each result names the tab it came from.

The project’s own services (every port Warpforge assigned to its services and port-forwards, on localhost) are allowed. Any other site raises a permission prompt in the task’s chat before the action runs: allow for that action, allow always for that site for the rest of the task, deny to stop. The check runs right before every action, on the page actually loaded in the tab. Everything read from a page comes back inside a block marked as untrusted data.

browser_navigate

Opens url in the agent’s tab and waits for it to load, returning the final address and title. A bare localhost:4001/path means http://; only http and https pages open. The first call opens the tab in the background if the Browser pane isn’t showing. A page that doesn’t load within about 20 seconds — a service that isn’t up, or that listens on another port — is reported as not loaded, not as opened.

Reach for it first, with an address from list_runtime.

browser_snapshot

Reads the agent tab’s page as a compact outline: headings, text, and every link, button, field and option, each interactive one tagged with a reference like [e12]. A reference keeps pointing at the same element across snapshots until the page reloads. Long text is cut short and a large page is truncated, with a note saying so.

browser_click

Clicks the element ref names, with the pointer and mouse events a real click produces — links follow, checkboxes toggle, submit buttons submit. Take a new snapshot to see what changed.

browser_type

Replaces the value of the text field, text area, editable element or select ref names with text, firing the input and change events frameworks listen for. For a select, text is an option’s label or value. submit: true presses Enter afterwards, submitting the form.

browser_screenshot

A picture of the visible part of the page, returned as an image. macOS only.

browser_console

The page’s console messages and uncaught errors since it loaded — the newest 200.

Why this exists: the agent can check its own UI change the way you would — load it, click through it, read the error — instead of handing you a diff and hoping.


Visual replies — show a page in the chat

Available in every session. See Visual replies for what the user sees.

render_preview

Loads html — one self-contained page — the way the chat shows it and returns a screenshot, the height the page needs (contentHeight), and its console warnings and errors. width defaults to the reply column (720px); appearance is light or dark, defaulting to the app’s. It shows nothing to the user. Needs the desktop app on macOS; without it the tool says so and render_html still works.

render_html

Shows html inline in the task’s chat above the agent’s reply, under a short title. height is the frame’s starting height (80–2000px); the frame then fits the page. The page gets the app’s theme as CSS variables and runs sandboxed: inline scripts and styles, plus scripts, styles, fonts and images from https URLs, and no other network access. Pages are at most 512 KB.

Reach for it when a chart, table, diagram or mockup says more than prose — after checking the page with render_preview.


Advisor — ask a second agent

Available in a regular task started with an Advisor picked in New Task. See Add an advisor for when that’s worth it.

ask_advisor

Puts question to the task’s advisor — a second agent, often another harness or a stronger model — and returns its answer. Optional context carries what the agent tried or which options it’s weighing. The advisor already gets the task’s goal, the messages since the last question and the list of changed files, and it remembers its earlier answers, so the question can be the question.

Reach for it before a big design decision, when stuck after a couple of failed attempts, and before declaring the task done — the agent is told exactly that when the task starts.

The call waits for the answer. Harnesses that give a tool call a short timeout (Codex) get a “still working” reply after about 50 seconds; calling again with wait: true keeps waiting for the same answer instead of asking twice. A turn allows three questions; the fourth is refused, as is a question while the previous one is still being answered or while the advisor’s account is out of quota. An answer that takes longer than 15 minutes is given up on.

The advisor’s own session is read-only: of these tools it gets only the ones that read — list_runtime, the log tools, memory_search / memory_list / memory_stats / memory_edges, the read-only automation tools, and render_preview / render_html.


Orchestrator only — delegation

These 11 tools appear only in an orchestrator session. They’re what lets one agent act as a lead rather than a worker.

spawn_agent

Dispatches a sub-agent — any configured harness, not just the lead’s own — and returns immediately.

spawn_agent(agent, task)

The point: the lead keeps its context for coordinating instead of burning it on implementation detail, and you can put the right harness on each piece of work in the same task.

Refused on an exhausted account. If the agent’s account is known to be out of quota, the call fails with the reason — which account, and when it resets — so the lead can choose another agent. A low or unchecked account is never refused.

read_inbox

Drains finished sub-agent results delivered since the last call.

When: after dispatching. Because spawn_agent returns immediately, the inbox is where results actually arrive — an orchestrator that never reads its inbox never learns anything came back.

message_agent

Sends a follow-up into an existing sub-agent’s session, with its full history intact.

message_agent(task_id, message)

Use this instead of spawn_agent when you want to continue a conversation — a correction, a clarification, “also handle the null case.” Spawning a fresh agent throws away everything the first one learned.

list_agents · stop_agent · cleanup_agents

  • list_agents(project?) — this orchestrator’s children and their state. Also where a pipeline’s workflowRun.waiting shows up, which the workflow tools below key off.
  • stop_agent(task_id) — hard-stop one child, keeping its history. Also stops an owned pipeline.
  • cleanup_agents(max_age_seconds?, dry_run?, include_active?) — permanently remove finished children and their history. dry_run first is the safe habit; include_active is required to touch anything still running.

spawn_workflow

Starts a Factory pipeline — a workflow template’s plan → implement → review ⇄ fix, with an optional check in the running app — for any goal, instead of a single sub-agent.

spawn_workflow(workflow_id, agent, goal?, model?, pull_request?, run_location?, backlog_item?)
Argument What it does
workflow_id The workflow template id.
agent / model The lead agent and its model, for every stage the template does not give its own agent. The model must be one of the agent’s models (list_agent_models); a stage on another agent uses that agent’s own default.
goal What to build. May be left out when backlog_item is given: the item’s brief is used.
pull_request false (the default): the pipeline is your child, as before — it starts at once, commits nothing, and its outcome arrives in your inbox. true: it runs as a Factory task, like one started from New Task — see below.
run_location auto (the project’s Factory setting: your project folder when the template tests the running app, a background copy otherwise), worktree (a background copy of the repository) or checkout (your project folder). Defaults to auto with a pull request, and to your own checkout without.
backlog_item A backlog item number, as in #87, to link the task to.

With pull_request: true the task waits for the project’s Factory limits, shows Queued in the sidebar until then, and a successful run is committed, pushed and opened as a draft pull request for a person to review; merging it marks a linked backlog item done. It is not your child: nothing arrives in your inbox and its questions go to the person, so follow it with runner_status. A backlog item that already has a Factory task is refused.

spawn_workflow or runner_enqueue? Use spawn_workflow for one goal — your own words, or one backlog item — and when you want the pipeline as your child without a pull request. Use runner_enqueue to start several backlog items at once with one configuration.

pause_workflow · resume_workflow

Soft-pause at the next stage boundary — the running stage finishes its turn, then the pipeline holds. resume_workflow(task_id, note?) continues, and the optional note is delivered to the next stage as extra context.

When: you learned something mid-flight that changes the remaining stages. Pausing is cheaper and less destructive than stopping and re-spawning.

answer_workflow

Answers a stage’s pending question — valid only while list_agents shows workflowRun.waiting.kind == "question". The message is forwarded to the session that asked. Pass the workflowRun.waiting.barrierId you saw as barrier_id; an answer written against a barrier that has since moved on is refused.

decide_workflow

Decides what happens when a pipeline exhausts its review ⇄ fix rounds with findings still open (waiting.kind == "limit") — grant more rounds, or accept and stop.

decide_workflow(task_id, decision, rounds?, note?, barrier_id?)

Why this exists: an unbounded review loop is how an agent pipeline burns your quota all night. The limit is a feature; this tool is how you answer it deliberately.


Project scoping

A session works in one project, and most tools stay inside it:

  • Scoped to the session’s project. The runtime tools (list_runtime, the log readers, service_*, portforward_*) and the orchestrator tools that look at your own children (list_agents, stop_agent, cleanup_agents, and the pipeline tools). Their optional project may only repeat the session’s own project — naming another is refused. spawn_agent and spawn_workflow always start work in the session’s project.
  • Default to the session’s project, but can name another. create_backlog_task and the other backlog tools (project), runner_enqueue and runner_status (project), the memory_* tools (project_id), and automation_create (project). automation_list lists every project’s automations unless you pass a project, and the other automation_* tools address an automation by id, whatever project it belongs to.

A bridge configured outside the daemon, with no project set, resolves the project from the working directory instead — including task worktrees — so one setup covers every registered project.

See also

Navigation

Type to search…

↑↓ navigate↵ selectEsc close