New to Claude Skills? Learn how to install them →

ruvnet on GitHub

Browser Intent

Free

Execute natural-language browser commands effortlessly.

by ruvnet67.6k stars on ruvnet/ruflo
1 views
Updated Aug 10, 2026
Get this skill

Free · Opens the source repo

What Browser Intent does

Browser Intent provides a natural-language interface for interacting with web pages through a low-level selector toolset. Unlike traditional methods that rely on specific selectors, this skill allows users to describe their desired actions in plain language, such as 'Click the login button' or 'Fill the search box with cats and submit'. This is particularly useful in scenarios where elements are difficult to select due to dynamic class names or ambiguous structures. By leveraging the page-agent, the skill translates these natural language commands into actionable browser interactions.

The skill operates by invoking the browser_act function with a task string, and optionally a URL and session identifier. If the page-agent is properly configured, it will execute the command and return a structured response detailing the success of the action, the steps taken, and any relevant history. In cases where the page-agent or an OpenAI-compatible LLM provider is not available, the skill degrades gracefully, allowing users to fall back on traditional selector-based commands without losing functionality.

This tool is ideal for developers and designers who need to automate browser interactions without the overhead of writing complex selector chains. It is particularly beneficial for one-off tasks where the time spent crafting a selector outweighs the simplicity of describing the action in natural language. However, users should be aware that using this skill introduces additional latency due to the LLM processing, so it is best suited for situations where such overhead is acceptable.

In summary, Browser Intent simplifies browser automation by allowing natural language commands, making it easier to interact with web elements that are otherwise cumbersome to select. It is a valuable addition for anyone looking to streamline their web interaction workflows, especially in dynamic environments.

When to use it

Use this skill when you need to perform browser actions that are easier to describe than to select, especially for dynamic or ambiguous elements.

When not to use it

Avoid this skill when you already have precise selectors available, as it introduces additional latency and cost compared to direct calls.

What you can build with it

Quick Login Automation

Automate the login process by simply stating 'Click the login button' without needing to write complex selectors.

Dynamic Content Interaction

Interact with elements that change frequently by describing them instead of relying on potentially outdated selectors.

Simplified Form Filling

Fill out forms by stating your intent, such as 'Fill the search box with cats', making the process faster and more intuitive.

How to install Browser Intent

View source

1. Install with the skills CLI

npx skills add ruvnet/ruflo/browser-intent --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by ruvnet

Browser Intent

Natural-language layer on top of the low-level browser_* selector tools. Where browser-extract and browser-form-fill compose selector-based primitives (browser_click, browser_fill, browser_snapshot), browser-intent lets the caller say what they want ("Click the login button", "Fill the search box with cats and submit") and delegates execution to page-agent — in-page injected JS that turns the DOM into text and drives an LLM tool-call loop against it.

When to use

  • The target element is easier to describe in words than to select reliably (dynamic class names, ambiguous structure, A/B-tested markup).
  • A one-shot interaction where writing out a selector chain isn't worth it.
  • Prefer browser_click / browser_fill / browser_snapshot directly when you already know the exact selector or ref (@e1) — browser_act adds LLM latency + cost that a direct selector call doesn't.

Steps

  1. Call browser_act with a task string, and optionally url (navigates first) and session (default "default"):
    mcp__plugin_ruflo-core_ruflo__browser_act({
      task: "Click the login button",
      url: "https://example.com/account",
      session: "my-session"
    })
    
  2. Read the response contract:
    • { success: true, result, steps, history, contentFlagged, llmSource } — the intent executed. result is the AIDefence-gated final text page-agent produced; history is the full step trace (reflection + action + tool result per step); steps is history.length.
    • { success: true, degraded: true, reason, hint } — page-agent isn't installed, or no OpenAI-compatible LLM provider is configured. Never treat degraded: true as an error to retry — surface the hint and fall back to selector-based browser_* tools instead.
    • { success: false, error, ... } — a real failure (browser open failed, injection failed, execution timed out, or page-agent's own execute() reported success:false).
  3. On contentFlagged: true, the returned result has already been redacted by AIDefence (PII or a prompt-injection/threat pattern was detected in the page-agent output) — do not attempt to recover the original text.
  4. Prefer a recorded session (browser-record) when the interaction matters enough to replay later; browser_act itself does not open an RVF container — it operates on whatever session id you pass (or "default").

Provider requirements (why this degrades so often)

page-agent calls its LLM directly from the browser page context via a plain OpenAI-compatible POST {baseURL}/chat/completions. That means:

  • A bare ANTHROPIC_API_KEY is not sufficient — Anthropic's native API is a different shape (/v1/messages).
  • Configure one of: OPENROUTER_API_KEY (OpenRouter, OpenAI-compatible), OLLAMA_API_KEY (Ollama Cloud, OpenAI-compatible), or CLAUDE_FLOW_PAGE_AGENT_BASE_URL + CLAUDE_FLOW_PAGE_AGENT_API_KEY for a custom OpenAI-compatible endpoint.
  • The real provider key never enters the page: browser_act starts a short-lived loopback HTTP proxy that holds the key server-side and injects the real Authorization header itself. The page only ever sees a 127.0.0.1 URL and a placeholder key string.

Caveats

  • page-agent is an optionalDependencies entry (npm i page-agent if the doctor/degraded hint asks for it) — this plugin stays fully operational without it; you simply lose the natural-language layer and fall back to selector-based tools.
  • The npm bundle's demo auto-init tail (which would otherwise construct a second PageAgent instance against Alibaba's public test endpoint) is stripped before injection — you should never see traffic to a page-ag-testing-* host from this tool.
  • Every successful browser_act call best-effort records the intent + resulting trajectory into the browser memory namespace (ADR-174 distillation loop). This is fire-and-forget — a memory-store failure never fails the tool call.
  • timeoutMs (default 120000) bounds how long browser_act polls for execute() to settle; a slow multi-step intent may need a higher value.

Frequently asked questions about Browser Intent

Similar skills