New to Claude Skills? Learn how to install them →

Bopenclaw on GitHub

Browser Automation

Free

Control web pages efficiently with OpenClaw.

Get this skill

Free · Opens the source repo

What Browser Automation does

Browser Automation is a skill designed for developers and designers who need to manage complex interactions with web pages using the OpenClaw browser tool. This skill is particularly useful for automating multi-step processes, handling login checks, managing tabs, and recovering from stale references or timeouts. By providing a structured approach to browser interactions, it allows users to automate repetitive tasks and streamline workflows without manual intervention.

The skill operates through a series of well-defined steps. First, it checks the browser state to ensure that the environment is ready for actions, using commands like openclaw browser doctor or action="status". This proactive approach helps avoid errors that may arise from an unstable browser state. Users are encouraged to work with stable tab handles, which can be labeled for easier reference in future actions. This minimizes the risk of errors associated with tab management, especially in scenarios where multiple tabs are open or when navigating between pages.

In addition to managing tabs and browser states, the skill emphasizes the importance of reading the page content before executing actions. By using scoped actions such as action="snapshot", users can capture relevant data without overwhelming the system with unnecessary information. This capability is particularly beneficial when dealing with virtualized lists or ambiguous link texts, ensuring that the automation process remains efficient and accurate.

Overall, Browser Automation is an essential tool for anyone looking to enhance their web automation capabilities with OpenClaw. It provides a comprehensive framework for managing browser interactions, making it easier to execute complex workflows and improve productivity in web-based tasks.

When to use it

Use this skill when automating tasks that involve multiple web pages or require precise control over browser states and actions.

When not to use it

This skill may not be suitable for simple, one-off page checks or when a lightweight solution is sufficient.

What you can build with it

Automating Login Processes

Use this skill to automate the login process across multiple web applications, ensuring that session states are managed correctly.

Managing Complex Workflows

Implement this skill to automate workflows that require navigating through several pages, capturing data, and performing actions based on the content.

Handling Tab Management

Utilize this skill to efficiently manage multiple tabs, ensuring that actions are directed to the correct tab without losing context.

How to install Browser Automation

View source

1. Install with the skills CLI

npx skills add openclaw/openclaw/browser-automation --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by openclaw

Browser Automation

Use this skill when you need the browser tool for anything beyond a single page check.

Operating Loop

  1. Check browser state before acting:
    • openclaw browser doctor or action="status" when the browser/plugin setup itself may be broken.
    • action="status" for availability.
    • action="profiles" if login state or profile choice matters.
    • action="tabs" before opening a new tab if retries/timeouts may have left windows behind.
  2. Prefer stable tab handles:
    • Open important tabs with label, for example label="meet".
    • After action="tabs" or action="open", store suggestedTargetId and pass it as targetId in later calls.
    • suggestedTargetId is the label when one exists, otherwise the stable tabId handle like t1.
    • Avoid relying on raw DevTools targetId except for immediate diagnostics; it can change under Chromium target replacement.
  3. Read before you click:
    • For “read the page and answer X,” use a selector-scoped action="snapshot" or act:evaluate that returns only relevant text or structured data. Let the active agent model answer from that bounded result; use efficient snapshots for controls and action discovery because they omit most non-interactive prose.
    • For virtualized lists, scroll through each segment, capture only the relevant rows, then merge the results.
    • Use action="snapshot" on the intended targetId.
    • Use the same targetId for follow-up actions so refs stay on the same tab.
    • For durable Playwright refs, request refs="aria" when supported. If you receive axN refs from snapshotFormat="aria", use them only after that same snapshot call; stale or unbound axN refs fail fast and need a fresh snapshot.
    • Use urls=true when link text is ambiguous or a direct navigation target would avoid brittle clicks.
    • Use labels=true on snapshot or screenshot when visual position matters. On Playwright-backed profiles, the response includes an annotations array ({ref, number, role, name?, box}) with each ref's bounding box in the captured image's coordinate space, so you can reason about position without re-snapshotting; screenshot labels can also combine with fullPage=true (CLI: --full-page) to label the whole document, or ref / element to clip to one element. profile="user" and other existing-session (chrome-mcp) profiles render an overlay into page screenshots but do not attach annotations or use the Playwright full-page/ref/element projection helper, so read positions from the labeled image itself on those profiles. The raw-CDP fallback (no Playwright) does not support labeled screenshots at all and returns a 501, so only request labels when Playwright is available.
  4. Act narrowly:
    • Prefer action="act" with a ref from the latest snapshot.
    • navigate returns the loaded page's compact snapshot inline, and batch act results that report a cross-document navigation include fresh page state; use those refs directly instead of a follow-up snapshot call.
    • After a single act that triggers navigation, and after modal changes or form submissions, snapshot again before the next action.
    • Avoid blind waits. Wait for visible UI state when possible.
  5. Report real blockers:
    • If the page needs login, permission, captcha, 2FA, camera/microphone approval, or another manual step, stop and tell the user exactly what is needed.
    • Do not claim the browser is not logged in just because the current page shows a permission or onboarding dialog. Inspect the visible UI first.

Browser batch CLI

openclaw browser batch runs an array of nested /act actions in one /act call (the same kind="batch" runtime reached through the agent tool), so CLI users and scripts can combine actions like wait, click, type, and evaluate into a single replayable plan without per-action round trips. Each entry in actions[] is a BrowserActRequest — the closed union the /act route accepts — not arbitrary openclaw browser subcommands. batch is not supported on profile="user" and other existing-session (chrome-mcp) profiles; send actions individually there.

  • CLI: openclaw browser batch --actions '<json>', --actions-file plan.json, or --actions-file - for stdin. --continue sets stopOnError=false; default stops on first error.
  • Ref lifecycle: refs come from a snapshot run before the batch (snapshot is not a nested action). A nested action that changes page state — such as a click that triggers navigation, or an evaluate that mutates the DOM — can invalidate earlier refs for the rest of the batch; put state-changing actions first, or split into a follow-up batch after re-snapshotting. Navigation and re-snapshotting happen outside the batch, since open, navigate, and snapshot are not /act kinds.
  • Target id: nested actions share the request's tab; an explicit nested targetId that resolves to a different tab is rejected with ACT_TARGET_ID_MISMATCH.
  • Response: { "results": [{ "ok": true } | { "ok": false, "error": "..." }, ...] } in order; with default stopOnError the array ends at the first failure. Any failed entry exits nonzero; use --json to preserve the full response in scripts.

Code Mode Loop

When tools.codeMode is enabled, call the Browser tool from exec cells:

const browserTool = "openclaw:browser:browser";
let previousSnapshot = "";
const callBrowser = async (input) => await tools.call(browserTool, input);

Keep the same labeled tab through the loop, and alternate reads with actions:

const snapshotCall = await callBrowser({
  action: "snapshot",
  targetId: "task",
  refs: "aria",
  interactive: true,
});
const details = snapshotCall?.result?.details ?? {};
const snapshot = (snapshotCall?.result?.content ?? []).map((block) => block?.text ?? "").join("\n");
const relevant = snapshot
  .split("\n")
  .filter((line) => /submit|dialog|error|\[new\]/i.test(line))
  .slice(0, 12);
const changed = snapshot !== previousSnapshot;
previousSnapshot = snapshot;
return { targetId: details.targetId, url: details.url, relevant, changed };
  • Request interactive-only snapshots and filter them in code before returning.
  • Return only the handful of relevant elements; never return the full tree.
  • Keep previousSnapshot between cells when a local diff helps explain a change.
  • Interleave each act with a URL or tabs check before the next dependent act.
  • If a batch returns aborted, take a fresh snapshot before continuing.
  • If [new] markers appear, inspect those elements first, then update the saved snapshot.
  • Use separate act calls when navigation is expected between steps.

Tab Hygiene

Before creating a tab for a named task, list tabs and reuse an existing matching label or URL when it is still usable.

Example:

{ "action": "tabs" }

If no suitable tab exists:

{ "action": "open", "url": "https://example.com", "label": "task" }

Then target it by label:

{ "action": "snapshot", "targetId": "task", "refs": "aria" }

If a retry creates duplicates, close the extras by tabId:

{ "action": "close", "targetId": "t3" }

Do not pass bare numbers like "2" as targetId. Numeric tab positions are only for the CLI openclaw browser tab select 2 helper; browser tool calls need a suggestedTargetId, label, tabId, or raw target id.

Stale Ref Recovery

If an action fails with a missing or stale ref:

  1. Snapshot the same targetId again.
  2. Find the current visible control.
  3. Retry once with the new ref.
  4. If the UI moved to a blocker state, report the blocker instead of looping.

Existing User Browser

Use profile="user" only when existing cookies/login matter. This attaches to the user's running Chromium-based browser.

On macOS, action="importprofile" is the alternative when the agent should use an isolated managed browser with cookies copied from a real Chrome-family profile. First use action="profiles" and inspect systemProfiles, then import into a fresh managed profile name. Import asks for one Keychain/Touch ID consent prompt. It copies cookies, not local storage or IndexedDB; device-bound session credentials (DBSC) mean some Google sessions may still require re-authentication.

For profile="user" and other existing-session profiles, omit timeoutMs on act:type, hover, scrollIntoView, drag, select, and fill; that driver rejects per-call timeout overrides for those actions. act:evaluate accepts timeoutMs.

Google Meet Notes

When creating or joining a Meet:

  • Treat camera/microphone permission screens as progress, not login failure.
  • If asked whether people can hear you, click the microphone option when voice is required.
  • If Google asks for sign-in, 2FA, account chooser confirmation, or permission that needs user approval, report the exact manual action.
  • Use one labeled tab per meeting flow, for example label="meet", and reuse it during retries.

Frequently asked questions about Browser Automation

Similar skills