---
title: "The browser"
description: "Open the app you're building next to the agent building it — signed in once, scoped per project, able to hand the agent any element you point at, and one the agent can drive itself."
---

> Documentation Index
> Fetch the complete documentation index at: https://warpforge.app/llms.txt
> Use this file to discover all available pages before exploring further.

# The browser

Agent-built work has a gap in the middle of it: the agent writes the change, and you go somewhere else to look at it. Another window, another login, another hunt for which localhost port it came up on — and when something looks wrong, you're back in the chat describing a button instead of showing it.

**Browser** is a task tab that closes that gap. It is a real browser — tabs, an address bar, back and forward, favicons — but it is pointed at what you're building rather than at the web in general.

## It opens what's running

A new tab doesn't land on a search engine. It lands on the services this project has up, each with the address it's actually listening on:

1. Start services in [Runtime](/guides/projects-and-runtime/) — or let them start with the project.
2. Open **Browser**. Every running service is a row on the start page.
3. Click one. No copying a port out of a log.

Nothing running yet just says so, and the address bar still takes anything you type.

> **Why this and not a search box**
>
> The port a dev service lands on is [assigned by Warpforge](/concepts/ports/), not memorised by you — which makes "which localhost was it this time" a question the tool should answer, not ask. If you pin ports in `.warpforge/workspace.yaml`, these links follow the pin.

## Sign in once, not every session

Logins persist. Sign into a staging environment, a dashboard, or GitHub inside the Browser tab and you are still signed in after you quit and reopen Warpforge, and in every task.

There is one web session for the app, deliberately. Per-task isolation would mean logging into your own staging site once per task, which is the misery this replaces.

> **Passwords come from your password manager**
>
> The Browser tab can't reach Safari's iCloud Keychain autofill — no embedded browser in a third-party app can. A password manager with system-wide autofill (1Password, Bitwarden) fills here like it does anywhere else. Warpforge never reads another browser's saved logins or cookies.

## Tabs belong to the project

Each project keeps its own tabs. Open your docs site, a staging dashboard, and two API responses in one project, and none of it follows you into an unrelated project.

Tabs come back after a restart, into the login that was already there. **Remove a project and its tabs go with it** — the pages it had open are dropped and their views closed. Your logins are not touched; those are the app's, not the project's.

Pages keep their state while you move around. Switch to another tab, another task or another part of the app and come back, and each page is exactly as you left it — same scroll position, same half-filled form, no reload — with its title and back history intact.

## Point at an element, don't describe it

The hardest thing to put in a prompt is a thing you're looking at. The picker turns it into context.

1. Click the **pointer** button in the toolbar. The cursor becomes a crosshair.
2. Click the element you mean — a broken button, a mislaid heading, a wrong number.
3. It lands in the composer as one chip: a **screenshot of that element** with its text underneath.
4. Write your message around it and send.

The agent receives the element's text, its role, its CSS selector, and the page URL, plus the picture. The chat shows it back as a card, not as raw markup, so the conversation stays readable when you scroll it later.

> **Page content is treated as data**
>
> Everything the picker reads comes from the page, and a page is not a trustworthy author. The context is handed over explicitly marked as untrusted data, and the text is neutralised so a page can't close the block early and start issuing instructions to your agent.

Picking ends on its own after one element, and any navigation cancels it. Click the button again to pick another.

## Let the agent drive it

Pointing is for when you know what's wrong. When the agent should find out for itself — sign up with a test account, click through the checkout it just built, read the error the page throws — it can use the browser directly, in a tab of its own right next to yours.

Ask for it in plain words: *"open the signup page, register a test user and tell me what breaks"*. The agent has six tools for it:

| Tool | What it does |
| --- | --- |
| `browser_navigate` | Opens an address in the project's tab and waits for it to load |
| `browser_snapshot` | Reads the page as a compact outline — headings, text, and every link, button and field with a short reference |
| `browser_click` | Clicks one of those elements |
| `browser_type` | Fills a field or picks an option, and can press Enter to submit |
| `browser_screenshot` | Takes a picture of what the page shows |
| `browser_console` | Reads the page's recent console messages and errors |

The agent works in its own **Agent** tab. Its first `browser_navigate` opens that tab and makes it the active one, and every action after that — reading, clicking, typing, screenshots — happens in that tab, even if you switch to one of yours. Your own tabs are never touched. If the Browser pane is on screen you watch it happen; if it isn't, nothing jumps in front of you — the page loads in the background and is waiting in the Agent tab the next time you open **Browser**. Close the Agent tab to take the browser back; the agent opens a new one the next time it navigates.

If a page doesn't load — a service that isn't up, or one listening on a different port than Warpforge gave it — the agent is told the page didn't load rather than being handed whatever the tab showed before.

### Your own services open freely; everything else asks

The browser is signed in as you, so an agent acting in it acts as you. Warpforge draws the line at your project:

- **The project's own services** — every `localhost` port [Warpforge assigned](/concepts/ports/) to its services and port-forwards — are allowed without asking.
- **Any other site** — staging, GitHub, your email — needs your yes first. The request appears in the task's chat like any other permission prompt, naming the site. **Allow** lets that one action through; **Allow always** lets the agent use that site for the rest of the task; **Deny** stops it and tells the agent not to retry. These prompts are answered in the task itself: the notification and toast for one offer **Review**, not a one-click approve, so you always see which site you're allowing. Your agent's permission mode doesn't answer them either — a site grant isn't a tool permission.

The check happens right before each action, on the page that is actually loaded in the tab, not on an address that is still loading, so if a link or redirect takes the page to another site, the agent has to ask before it reads or touches that one. A prompt nobody answers is withdrawn after five minutes and the action doesn't run.

> **Pages are data, not instructions**
>
> Everything the agent reads from a page reaches it marked as untrusted data, the same way picked elements do, so a page can't slip instructions to your agent.

### Testing every change before review

A [Factory](/guides/factory/) task whose workflow has a [**verify** stage](/reference/workflow-files/#verify) makes this routine instead of something you ask for — turn on **Test in the running app** in New Task. After the implementation, and after each fix, a tester agent starts the project's services, works out a test plan from the task — an issue an analyst wrote works fine — and walks through it in its Agent tab. It checks the console for errors, takes a screenshot at each key step, and reports **pass**, **fail** or **blocked** with a checklist. A failure goes back to the fixer; only a pass goes on to the reviewers.

Because the tester only visits the project's own services, it never stops to ask you for access. Its screenshots are kept with the pipeline: open the verify stage in the task's **Pipeline** view to see the verdict, each step, and the pictures, and click one to see it full size. Give the tester what it can't guess — a test login, seed data, the page to start from — with `instructions` in the workflow file.

Leave the desktop app open while it runs. A task in a background copy of the repository can't be verified, because the services run from your project folder; with **Where it runs** on *Automatic*, such a task runs in your project folder for you.

### What it can't do

The agent's clicks and keystrokes are simulated inside the page rather than made with a mouse and keyboard, and a page can tell the difference. Most apps don't care, but a few things are reserved for a real person:

- File pickers, pop-up windows and new-tab links don't open. For a link that opens a new window, the agent is told to open its address directly instead.
- Captchas, payment forms and anything else that insists on a real user gesture ignore it.
- Content inside frames from another site can't be read or clicked.
- One tab, one agent at a time: two agents driving the same project's browser at once will step on each other.

Page screenshots are macOS only, and a tab that isn't on screen may not render one — open the Browser pane and ask again.

## When a page doesn't load

A dev server that isn't up yet, a port that moved, a service that crashed on boot — the tab says **This site can't be reached** with the address it tried and a **Try again** button, rather than showing you a blank rectangle and letting you wonder.

If that happens on a service link, the usual cause is an app that ignores the `PORT` it was given and binds its own default instead. [How ports work](/concepts/ports/) covers the fix.

## Dialogs, menus and notifications

The page is a native view sitting on top of the app, so the app makes room for its own UI instead of letting the page cover it. A dialog — settings, the push dialog, quick open — hides the page while it's open, and the page comes back exactly where it was when the dialog closes. A menu or notification that overlaps the page makes the page shrink out of its way rather than disappear, so you keep seeing the rest of it.

## Shortcuts and chrome

| | |
| --- | --- |
| <kbd>⌘L</kbd> | Focus the address bar and select it |
| <kbd>⌘T</kbd> | New tab |
| **Address bar** | A URL, a bare host like `localhost:5173`, or a search — searches go to DuckDuckGo |
| **Back / forward** | Greyed when there's nowhere to go |
| **Reload** | Becomes **Stop** while a page is loading |
| **Tabs** | Favicons and page titles; the strip scrolls when they overflow |

> **Scope and limits**
>
> Browser needs the desktop app — it doesn't run in a web build, and the agent's browser tools only work while the app is open. Element and page screenshots are macOS today.

## See also

- [Working in a task](/guides/working-in-a-task/) — the other ways to get context into the chat
- [MCP tools](/reference/mcp-tools/) — the full reference for the agent's browser tools
- [Projects and their runtime](/guides/projects-and-runtime/) — starting the services this tab opens
- [How ports work](/concepts/ports/) — why a service is on the port it's on

Source: https://warpforge.app/guides/browser/index.mdx
