
Browser Automation
FreeControl web pages efficiently with OpenClaw.
Free · Opens the source repo
What Browser Automation does
Browser Automation is a skill designed for developers and designers who need to manage complex interactions with web pages using the OpenClaw browser tool. This skill is particularly useful for automating multi-step processes, handling login checks, managing tabs, and recovering from stale references or timeouts. By providing a structured approach to browser interactions, it allows users to automate repetitive tasks and streamline workflows without manual intervention.
The skill operates through a series of well-defined steps. First, it checks the browser state to ensure that the environment is ready for actions, using commands like openclaw browser doctor or action="status". This proactive approach helps avoid errors that may arise from an unstable browser state. Users are encouraged to work with stable tab handles, which can be labeled for easier reference in future actions. This minimizes the risk of errors associated with tab management, especially in scenarios where multiple tabs are open or when navigating between pages.
In addition to managing tabs and browser states, the skill emphasizes the importance of reading the page content before executing actions. By using scoped actions such as action="snapshot", users can capture relevant data without overwhelming the system with unnecessary information. This capability is particularly beneficial when dealing with virtualized lists or ambiguous link texts, ensuring that the automation process remains efficient and accurate.
Overall, Browser Automation is an essential tool for anyone looking to enhance their web automation capabilities with OpenClaw. It provides a comprehensive framework for managing browser interactions, making it easier to execute complex workflows and improve productivity in web-based tasks.
When to use it
Use this skill when automating tasks that involve multiple web pages or require precise control over browser states and actions.
When not to use it
This skill may not be suitable for simple, one-off page checks or when a lightweight solution is sufficient.
What you can build with it
Automating Login Processes
Use this skill to automate the login process across multiple web applications, ensuring that session states are managed correctly.
Managing Complex Workflows
Implement this skill to automate workflows that require navigating through several pages, capturing data, and performing actions based on the content.
Handling Tab Management
Utilize this skill to efficiently manage multiple tabs, ensuring that actions are directed to the correct tab without losing context.
How to install Browser Automation
View source1. Install with the skills CLI
npx skills add openclaw/openclaw/browser-automation --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by openclawBrowser Automation
Use this skill when you need the browser tool for anything beyond a single page check.
Operating Loop
- Check browser state before acting:
openclaw browser doctororaction="status"when the browser/plugin setup itself may be broken.action="status"for availability.action="profiles"if login state or profile choice matters.action="tabs"before opening a new tab if retries/timeouts may have left windows behind.
- Prefer stable tab handles:
- Open important tabs with
label, for examplelabel="meet". - After
action="tabs"oraction="open", storesuggestedTargetIdand pass it astargetIdin later calls. suggestedTargetIdis the label when one exists, otherwise the stabletabIdhandle liket1.- Avoid relying on raw DevTools
targetIdexcept for immediate diagnostics; it can change under Chromium target replacement.
- Open important tabs with
- Read before you click:
- For “read the page and answer X,” use a selector-scoped
action="snapshot"oract:evaluatethat returns only relevant text or structured data. Let the active agent model answer from that bounded result; use efficient snapshots for controls and action discovery because they omit most non-interactive prose. - For virtualized lists, scroll through each segment, capture only the relevant rows, then merge the results.
- Use
action="snapshot"on the intendedtargetId. - Use the same
targetIdfor follow-up actions so refs stay on the same tab. - For durable Playwright refs, request
refs="aria"when supported. If you receiveaxNrefs fromsnapshotFormat="aria", use them only after that same snapshot call; stale or unboundaxNrefs fail fast and need a fresh snapshot. - Use
urls=truewhen link text is ambiguous or a direct navigation target would avoid brittle clicks. - Use
labels=trueon snapshot or screenshot when visual position matters. On Playwright-backed profiles, the response includes anannotationsarray ({ref, number, role, name?, box}) with each ref's bounding box in the captured image's coordinate space, so you can reason about position without re-snapshotting; screenshot labels can also combine withfullPage=true(CLI:--full-page) to label the whole document, orref/elementto clip to one element.profile="user"and other existing-session (chrome-mcp) profiles render an overlay into page screenshots but do not attachannotationsor use the Playwright full-page/ref/element projection helper, so read positions from the labeled image itself on those profiles. The raw-CDP fallback (no Playwright) does not support labeled screenshots at all and returns a 501, so only requestlabelswhen Playwright is available.
- For “read the page and answer X,” use a selector-scoped
- Act narrowly:
- Prefer
action="act"with a ref from the latest snapshot. navigatereturns the loaded page's compact snapshot inline, and batchactresults that report a cross-document navigation include fresh page state; use those refs directly instead of a follow-up snapshot call.- After a single act that triggers navigation, and after modal changes or form submissions, snapshot again before the next action.
- Avoid blind waits. Wait for visible UI state when possible.
- Prefer
- Report real blockers:
- If the page needs login, permission, captcha, 2FA, camera/microphone approval, or another manual step, stop and tell the user exactly what is needed.
- Do not claim the browser is not logged in just because the current page shows a permission or onboarding dialog. Inspect the visible UI first.
Browser batch CLI
openclaw browser batch runs an array of nested /act actions in one /act call (the same kind="batch" runtime reached through the agent tool), so CLI users and scripts can combine actions like wait, click, type, and evaluate into a single replayable plan without per-action round trips. Each entry in actions[] is a BrowserActRequest — the closed union the /act route accepts — not arbitrary openclaw browser subcommands. batch is not supported on profile="user" and other existing-session (chrome-mcp) profiles; send actions individually there.
- CLI:
openclaw browser batch --actions '<json>',--actions-file plan.json, or--actions-file -for stdin.--continuesetsstopOnError=false; default stops on first error. - Ref lifecycle: refs come from a
snapshotrun before the batch (snapshot is not a nested action). A nested action that changes page state — such as aclickthat triggers navigation, or anevaluatethat mutates the DOM — can invalidate earlier refs for the rest of the batch; put state-changing actions first, or split into a follow-up batch after re-snapshotting. Navigation and re-snapshotting happen outside the batch, sinceopen,navigate, andsnapshotare not/actkinds. - Target id: nested actions share the request's tab; an explicit nested
targetIdthat resolves to a different tab is rejected withACT_TARGET_ID_MISMATCH. - Response:
{ "results": [{ "ok": true } | { "ok": false, "error": "..." }, ...] }in order; with defaultstopOnErrorthe array ends at the first failure. Any failed entry exits nonzero; use--jsonto preserve the full response in scripts.
Code Mode Loop
When tools.codeMode is enabled, call the Browser tool from exec cells:
const browserTool = "openclaw:browser:browser";
let previousSnapshot = "";
const callBrowser = async (input) => await tools.call(browserTool, input);
Keep the same labeled tab through the loop, and alternate reads with actions:
const snapshotCall = await callBrowser({
action: "snapshot",
targetId: "task",
refs: "aria",
interactive: true,
});
const details = snapshotCall?.result?.details ?? {};
const snapshot = (snapshotCall?.result?.content ?? []).map((block) => block?.text ?? "").join("\n");
const relevant = snapshot
.split("\n")
.filter((line) => /submit|dialog|error|\[new\]/i.test(line))
.slice(0, 12);
const changed = snapshot !== previousSnapshot;
previousSnapshot = snapshot;
return { targetId: details.targetId, url: details.url, relevant, changed };
- Request interactive-only snapshots and filter them in code before returning.
- Return only the handful of relevant elements; never return the full tree.
- Keep
previousSnapshotbetween cells when a local diff helps explain a change. - Interleave each act with a URL or tabs check before the next dependent act.
- If a batch returns
aborted, take a fresh snapshot before continuing. - If
[new]markers appear, inspect those elements first, then update the saved snapshot. - Use separate act calls when navigation is expected between steps.
Tab Hygiene
Before creating a tab for a named task, list tabs and reuse an existing matching label or URL when it is still usable.
Example:
{ "action": "tabs" }
If no suitable tab exists:
{ "action": "open", "url": "https://example.com", "label": "task" }
Then target it by label:
{ "action": "snapshot", "targetId": "task", "refs": "aria" }
If a retry creates duplicates, close the extras by tabId:
{ "action": "close", "targetId": "t3" }
Do not pass bare numbers like "2" as targetId. Numeric tab positions are only for the CLI openclaw browser tab select 2 helper; browser tool calls need a suggestedTargetId, label, tabId, or raw target id.
Stale Ref Recovery
If an action fails with a missing or stale ref:
- Snapshot the same
targetIdagain. - Find the current visible control.
- Retry once with the new ref.
- If the UI moved to a blocker state, report the blocker instead of looping.
Existing User Browser
Use profile="user" only when existing cookies/login matter. This attaches to the user's running Chromium-based browser.
On macOS, action="importprofile" is the alternative when the agent should use an isolated managed browser with cookies copied from a real Chrome-family profile. First use action="profiles" and inspect systemProfiles, then import into a fresh managed profile name. Import asks for one Keychain/Touch ID consent prompt. It copies cookies, not local storage or IndexedDB; device-bound session credentials (DBSC) mean some Google sessions may still require re-authentication.
For profile="user" and other existing-session profiles, omit timeoutMs on act:type, hover, scrollIntoView, drag, select, and fill; that driver rejects per-call timeout overrides for those actions. act:evaluate accepts timeoutMs.
Google Meet Notes
When creating or joining a Meet:
- Treat camera/microphone permission screens as progress, not login failure.
- If asked whether people can hear you, click the microphone option when voice is required.
- If Google asks for sign-in, 2FA, account chooser confirmation, or permission that needs user approval, report the exact manual action.
- Use one labeled tab per meeting flow, for example
label="meet", and reuse it during retries.
Frequently asked questions about Browser Automation
Similar skills
Agent-Browser Core
Efficient browser automation for AI agents.
Setup My IQ
Effortlessly create and update your personal context portfolio.
CRM Maintenance
Automate HubSpot updates from your calendar and emails.
Zoom MCP
Streamline access to Zoom meeting assets and recordings.
Slack Automation
Automate tasks and extract data from Slack easily.
SMB Onboard
Guides small business owners through initial tool setup.
