feat(web-access): 迁移到内置受管浏览器工具(v2.1.0) (#72)

## 变更说明 / Description

### 中文

客户端 v10.0.98(desirecore/desirecore#1596)停用了旧的 `BrowserListTabs` /
`BrowserNavigate` / `BrowserEval` / `BrowserClick` / `BrowserScreenshot`
/ `BrowserScroll` / `BrowserSetFiles` / `BrowserCloseTab` 及其 cdp-proxy
后端,调用会直接返回「该旧 BrowserXxx/cdp-proxy 入口已停用」。而 web-access v2.0.2 的
`provides.tools` 仍声明这批工具——技能激活后注入的是一组必然失败的工具。

本次把 `provides.tools` 换成统一浏览器工具,并同步正文与参考文档:

- `provides.tools`:`BrowserManage` / `BrowserSnapshot` / `BrowserAct` /
`BrowserImport` / `BrowserShare` + 保留 `SitePatternRead` /
`SitePatternWrite` / `LocalBookmarks`
- 中英文 SKILL 正文同步改写(L0 / 能力描述 / 决策树 / 四层策略表 / L3-fast 速查 / 反模式)
- `references/browser-tools.md` 整篇重写为新 API + 实测边界
- 5 份站点经验(小红书 / B站 / 微博 / 知乎 / 飞书)的流程改用新工具
- 版本 2.0.2 → 2.1.0,`updated_at` 更新,i18n `source_hash` 重算

**L3-fast 的定位相应收窄**:内置浏览器负责「到达 + 交互 + 截图 + 隔离」,抽取长正文仍回落 Jina
Reader(公开页)或 Playwright(登录态)——理由见下方实测。

### English

Client v10.0.98 retired the legacy `BrowserXxx` tools and the cdp-proxy
behind them, so web-access v2.0.2 was injecting a set of tools that
always fail. This PR migrates `provides.tools` to the unified browser
tools and rewrites the body, the browser-tools reference, and the five
site-pattern playbooks accordingly. L3-fast is re-scoped to
navigation/interaction/screenshots; bulk text extraction still falls
back to Jina Reader or Playwright.

## 测试方式 / Test Plan

在客户端 v10.0.98 + `electron-embedded` Provider 上实测:

- [x] `provides.tools` 里 8 个工具 ID 全部在 builtin registry 中存在
- [x] 把本 PR 的技能装进 dev 实例,带 `skillIds:['web-access']` 驱动智能体:真实调用
`BrowserManage(create_space)` → `BrowserManage(start_session)` →
`BrowserAct(tab.navigate)` → `BrowserManage(close_session)`,全部 success
- [x] `scripts/i18n/validate-i18n.py` 全仓库通过(中英文标题数一致、source_hash 一致)

文档中记录的边界均来自实测,而非推测:

| 边界 | 实测现象 |
|------|---------|
| 截图前必须 `tab.activate` | 标签页默认停在 `(-10000,-10000,1x1)`,直接截图卡满 30s
deadline 并触发 `browser.host.gone`,之后全部 `BROWSER_TAB_HOST_NOT_FOUND` |
| `page.evaluate` 不是取文通道 | 每次调用需人工审批;字符串/对象返回值被替换为
`[REDACTED:browser-runtime-value]`,仅 number/boolean/null 穿透 |
| `accessibility` 快照真实页面不可用 | example.com 正常返回 StaticText;维基百科条目一律
`BROWSER_RESULT_TOO_LARGE`(2 MB 上限,且 `depth` 参数被宿主忽略) |
| `semantic` 快照不含正文 | 只列 button / input / a 等可交互元素 |

## 风险与回滚 / Risk and rollback

- 纯技能内容变更,无脚本或清单结构改动
- 需要客户端 v10.0.98+;旧客户端装到本版会拿到一组不存在的工具名(旧客户端上原本那批工具也已失效,不构成回退)
- 回滚即 revert 本 PR
This commit is contained in:
2026-08-06 00:49:35 +08:00
committed by yi-ge
parent 413f2cc00b
commit af4176bbd7
8 changed files with 281 additions and 186 deletions

View File

@@ -5,16 +5,18 @@ description: >-
— searching for current information, fetching public web pages, browsing
login-gated sites (微博/小红书/B站/飞书/Twitter), comparing products,
researching topics, gathering documentation, or summarizing news.
This skill orchestrates three complementary layers: (1) WebSearch + WebFetch
This skill orchestrates four complementary layers: (1) WebSearch + WebFetch
for public pages, (2) Jina Reader as the default token-optimization layer for
heavy/JS-rendered pages, and (3) Chrome DevTools Protocol (CDP) via Python
Playwright for login-gated sites that require the user's existing browser
session. Always cite source URLs. Use when 用户提到 联网搜索、上网查、
heavy/JS-rendered pages, (3) the governed built-in browser (isolated
BrowserSpace + Cookie import) to reach and interact with login-gated sites,
and (4) Chrome DevTools Protocol via Python Playwright as the fallback for
bulk text extraction and complex automation. Always cite source URLs.
Use when 用户提到 联网搜索、上网查、
查资料、抓取网页、研究、调研、最新资讯、文档查询、对比、竞品、技术文档、
新闻、网址、URL、找一下、搜一下、查一下、小红书、B站、微博、飞书、Twitter、
推特、X、知乎、公众号、已登录、登录状态。
license: Complete terms in LICENSE.txt
version: 2.0.2
version: 2.1.0
type: procedural
risk_level: low
status: enabled
@@ -29,20 +31,17 @@ tags:
- playwright
provides:
tools:
- BrowserListTabs
- BrowserNavigate
- BrowserEval
- BrowserClick
- BrowserScreenshot
- BrowserScroll
- BrowserSetFiles
- BrowserCloseTab
- BrowserManage
- BrowserSnapshot
- BrowserAct
- BrowserImport
- BrowserShare
- SitePatternRead
- SitePatternWrite
- LocalBookmarks
metadata:
author: desirecore
updated_at: '2026-05-05'
updated_at: '2026-08-06'
i18n:
default_locale: en-US
source_locale: zh-CN
@@ -51,17 +50,17 @@ metadata:
- en-US
zh-CN:
name: 联网访问
short_desc: 联网搜索、网页抓取、登录态浏览器访问CDP、研究调研工作流
description: 层联网访问工具包——搜索公开页面、Jina 优化抓取、CDP 登录态浏览器访问
short_desc: 联网搜索、网页抓取、内置受管浏览器登录态访问、研究调研工作流
description: 层联网访问工具包——搜索公开页面、Jina 优化抓取、内置受管浏览器登录态访问、Python Playwright CDP 兜底
body: ./SKILL.zh-CN.md
source_hash: sha256:0ba170b3126a0823
source_hash: sha256:2e6753fc05142e9a
translated_by: human
en-US:
name: Web Access
short_desc: Web search, page fetching, logged-in browser access via CDP, research workflows
description: A three-layer web-access toolkit — search public pages, fetch heavy pages via Jina Reader, and reach logged-in sites via Chrome CDP.
short_desc: Web search, page fetching, logged-in access via the governed built-in browser, research workflows
description: A four-layer web-access toolkit — search public pages, fetch heavy pages via Jina Reader, reach logged-in sites through the governed built-in browser, and fall back to Chrome CDP.
body: ./SKILL.md
source_hash: sha256:4cda9ca594fbd974
source_hash: sha256:2e6753fc05142e9a
translated_by: human
market:
icon: >-
@@ -90,7 +89,7 @@ market:
## L0: One-line Summary
A three-layer web-access toolkit — search public pages, optimize fetches via Jina Reader, and reach login-gated sites via Chrome CDP.
A four-layer web-access toolkit — search public pages, optimize fetches via Jina Reader, reach login-gated sites through the governed built-in browser, and fall back to Chrome CDP.
## L1: Overview & Use Cases
@@ -100,24 +99,26 @@ web-access is a **procedural skill** that provides four complementary layers of
- **L1** (WebSearch + WebFetch): public, static pages
- **L2** (Jina Reader): JS-rendered heavy pages, saving tokens by default
- **L3-fast** (BrowserXxx builtin tool family**new in v2.0**): preferred for logged-in sites — zero Python dependency, in-process cdp-proxy, supports CDP real-mouse events
- **L3-fast** (governed built-in browser**rewritten in v2.1**): reach and *interact with* logged-in / interactive sites — isolated BrowserSpace per task, zero Python dependency, every action carries a signed receipt. **Reading long article text is still L2/L3-fallback's job** — see the extraction note below
- **L3-fallback** (Chrome CDP + Python Playwright): backup for complex automation (long waits, race conditions, custom in-browser scripts)
### v2.0BrowserXxx tool family (default-hidden, exposed only after Skill activation)
### v2.1governed built-in browser (default-hidden, exposed only after Skill activation)
When you call `Skill('web-access')`, the following 11 tools are injected into the current session so the LLM can drive Chrome directly:
When you call `Skill('web-access')`, the following 8 tools are injected into the current session so the LLM can drive the built-in browser directly:
| Tool | Purpose |
|------|---------|
| BrowserListTabs / BrowserNavigate / BrowserCloseTab | Tab management |
| BrowserEval | Run JS to extract data |
| BrowserClick (`mode: js \| real-mouse`) | Click elements; real-mouse defeats anti-bot |
| BrowserScreenshot / BrowserScroll | Screenshots, scroll to trigger lazy loading |
| BrowserSetFiles | Upload local files (requires user confirmation) |
| BrowserManage | Create/destroy isolated BrowserSpace, start sessions, manage tabs |
| BrowserSnapshot | `semantic` / `accessibility` / `visual` page snapshots — the primary way to read a page |
| BrowserAct | One governed action per call: navigate, click, type, scroll, screenshot, … |
| BrowserImport | Import Cookies from the user's Chrome/Edge/Firefox/Safari profile (human-approved) |
| BrowserShare | Delegate a Space/Session to another Agent (isolated / snapshot / copy-on-write / live) |
| SitePatternRead / SitePatternWrite | Per-domain "site experience" (AgentFS three-layer) |
| LocalBookmarks | Search local Chrome bookmarks / history |
> **Important**: before `Skill('web-access')` is called, none of these tools appear in the LLM tools list — default conversations don't pay their token cost. See [references/browser-tools.md](references/browser-tools.md).
>
> **Removed in v2.1**: `BrowserListTabs` / `BrowserNavigate` / `BrowserEval` / `BrowserClick` / `BrowserScreenshot` / `BrowserScroll` / `BrowserSetFiles` / `BrowserCloseTab` and the cdp-proxy behind them are retired. Calling them now returns "该旧 BrowserXxx/cdp-proxy 入口已停用". Requires client v10.0.98+.
### Use Cases
@@ -130,7 +131,7 @@ When you call `Skill('web-access')`, the following 11 tools are injected into th
- **Four-layer progression**: from lightweight search to heavy JS rendering to logged-in access — pick on demand
- **Token optimization**: Jina Reader cuts token usage by 5080% by default
- **Logged-in session reuse**: connect to the user's already-logged-in Chrome via CDP — no re-login required
- **Logged-in session reuse**: import the user's existing Cookies into an isolated Space via BrowserImport (or attach to their Chrome via CDP in the fallback layer) — no re-login required
## L2: Detailed Specification
@@ -146,7 +147,7 @@ If any fetch fails, explicitly tell the user which URL failed and which fallback
## Prerequisites: Chrome CDP Setup (for login-gated sites)
**Only required when accessing sites that need the user's login session** (Xiaohongshu / Bilibili / Weibo / Feishu / Twitter / Zhihu / WeChat Official Accounts).
**Only required for the L3-fallback layer** (Python Playwright). The L3-fast built-in browser does not need this — use `BrowserImport` to reuse the user's login instead.
### One-time setup
@@ -205,10 +206,11 @@ User intent
│ (Jina Reader = default for JS-rendered content, saves tokens)
├─ "Read this login-gated page" (Xiaohongshu/Bilibili/Weibo/Feishu/Twitter/Zhihu/WeChat)
─→ 1. Verify CDP ready (curl http://localhost:9222/json/version)
2. Bash: python3 script with playwright.connect_over_cdp()
3. Extract content → feed to Jina Reader for clean Markdown
(or use BeautifulSoup directly on the raw HTML)
─→ Reach it: BrowserManage(create_space/start_session) → BrowserImport(cookies for
that domain) → BrowserAct(tab.navigate) → tab.activate + page.screenshot
to confirm you landed on the content (semantic snapshot has no body text)
└─→ Extract the text: verify CDP ready, then python3 playwright.connect_over_cdp()
│ → page.content() → Jina Reader / BeautifulSoup
├─ "API documentation / GitHub / npm package info"
│ └─→ Prefer official API endpoints over scraping HTML:
@@ -217,9 +219,9 @@ User intent
│ - PyPI: curl https://pypi.org/pypi/<pkg>/json
└─ "Real-time interactive task" (click, fill form, scroll, screenshot)
├─→ **Default: BrowserXxx tools** (BrowserNavigate / BrowserEval / BrowserClick / BrowserScreenshot —
├─→ **Default: built-in browser** (BrowserManage → BrowserAct → BrowserSnapshot —
│ see references/browser-tools.md, no Python needed)
└─→ Fallback: CDP + Python Playwright (references/cdp-browser.md) when BrowserXxx is insufficient
└─→ Fallback: CDP + Python Playwright (references/cdp-browser.md) when the built-in browser is insufficient
(e.g., complex race conditions, multi-event waits, long-running in-browser scripts)
```
@@ -229,10 +231,10 @@ User intent
|-------|----------|--------------|------------|
| L1 | Public, static | `WebFetch` | Low |
| L2 | JS-heavy, long articles, token savings | `Bash curl r.jina.ai` | **Lowest** (Markdown pre-cleaned) |
| **L3-fast** | **Login-gated, interactive (PRIMARY)** | **BrowserXxx tool family** | Medium |
| **L3-fast** | **Login-gated navigation & interaction (PRIMARY)** | **built-in browser (BrowserManage / BrowserAct / BrowserSnapshot)** | Medium |
| L3-fallback | Complex automation (race / long-wait / custom scripts) | `Bash + Python Playwright CDP` | Medium |
**Default priority**: L1 for simple public pages → L2 for heavy → **L3-fast for login-gated** → L3-fallback only when BrowserXxx is insufficient.
**Default priority**: L1 for simple public pages → L2 for heavy → **L3-fast for login-gated** → L3-fallback only when the built-in browser is insufficient.
---
@@ -348,33 +350,49 @@ See [references/cdp-browser.md](references/cdp-browser.md) for:
---
## L3-fast: BrowserXxx Tool Cheatsheet (v2.0 recommended)
## L3-fast: Built-in Browser Cheatsheet (v2.1)
**Only after you call `Skill('web-access')` will the following tools appear in `tools[]`.**
| Tool | One-line example |
|------|-----------------|
| `BrowserListTabs()` | List all open tabs |
| `BrowserNavigate({ url })` | Open URL in a new tab |
| `BrowserNavigate({ target, url })` | Navigate an existing tab |
| `BrowserEval({ target, expression })` | Run JS in the tab to extract structured data |
| `BrowserClick({ target, selector, mode: 'real-mouse' })` | Real-mouse mode for anti-bot-strict sites |
| `BrowserScreenshot({ target })` | Saved under ${DESIRECORE_ROOT}/screenshots/ |
| `BrowserScroll({ target, direction: 'bottom' })` | Trigger lazy loading |
| `BrowserSetFiles({ target, selector, files })` | Upload local files (**user confirmation required**) |
| `BrowserCloseTab({ target })` | Clean up temporary tabs at task end |
| `BrowserManage({ action: 'create_space', name, persistence: 'ephemeral' })` | One isolated Space per task |
| `BrowserManage({ action: 'start_session', spaceId, capabilities })` | Start a session, returns sessionId + first tab |
| `BrowserManage({ action: 'list_tabs' \| 'create_tab' \| 'close_session' })` | Tab / lifecycle management |
| `BrowserAct({ action: 'tab.navigate', params: { url } })` | Navigate the current tab |
| `BrowserSnapshot({ mode: 'semantic' })` | Read the page: interactive elements + `ref` handles |
| `BrowserAct({ action: 'input.click', params: { ref } })` | Click by snapshot `ref`, never by raw x/y |
| `BrowserAct({ action: 'input.text', params: { text } })` | Type into the focused element |
| `BrowserAct({ action: 'input.wheel', params: { deltaX: 0, deltaY: 720 } })` | Scroll to trigger lazy loading |
| `BrowserAct({ action: 'tab.activate', params: { bounds } })``page.screenshot` | **activate first, then screenshot** |
| `BrowserImport({ action: 'discover' \| 'plan' \| 'apply' })` | Reuse the user's login cookies (human-approved) |
Full API and edge cases: see [references/browser-tools.md](references/browser-tools.md).
**How to read a page** — pick by what you need:
| You need | Use | Note |
|----------|-----|------|
| Interactive elements + `ref` handles | `BrowserSnapshot({ mode: 'semantic' })` | Buttons / inputs / links only — **no article text** |
| Article text on a small/simple page | `BrowserSnapshot({ mode: 'accessibility' })` | Returns StaticText nodes; fails with `BROWSER_RESULT_TOO_LARGE` on real content pages (2 MB result cap, and the `depth` argument is currently ignored by the host) |
| Article text on a real page | L2 Jina Reader (public) or L3-fallback Playwright (login-gated) | The built-in browser has no working bulk text-extraction channel yet |
| What the page looks like | `tab.activate``BrowserAct({ action: 'page.screenshot' })` | Full-page only; read it visually |
`page.evaluate` is **not** an extraction channel: it needs human approval per call and its string/object return values come back as `[REDACTED:browser-runtime-value]`.
### Recommended flow (Xiaohongshu example)
```
1. BrowserListTabs() → check whether there's an already-logged-in xhs tab
2. If not → BrowserNavigate({ url: "https://www.xiaohongshu.com/explore/abc123" })
3. BrowserEval({ target, expression: "...JSON.stringify({title, content})" })
4. SitePatternRead({ domain: "xiaohongshu.com" }) ← read accumulated experience
5. At task end → BrowserCloseTab({ target })
6. If you find a new pitfall → SitePatternWrite({ domain, scope: "agent", mode: "merge", content })
1. BrowserManage({ action: 'create_space', name: 'xhs-note', persistence: 'ephemeral' })
2. BrowserManage({ action: 'start_session', spaceId, capabilities: [...] })
3. BrowserImport({ action: 'discover' }) → plan → apply for xiaohongshu.com ← reuse login
4. BrowserAct({ action: 'tab.navigate', params: { url: 'https://www.xiaohongshu.com/explore/abc123' } })
5. BrowserSnapshot({ mode: 'semantic' }) ← element refs for interaction (no body text)
tab.activate + page.screenshot ← confirm the note actually rendered
→ extract the note text via L3-fallback Playwright
6. SitePatternRead({ domain: 'xiaohongshu.com' }) ← read accumulated experience
7. At task end → BrowserManage({ action: 'close_session', sessionId })
8. If you find a new pitfall → SitePatternWrite({ domain, scope: 'agent', mode: 'merge', content })
```
---
@@ -446,9 +464,11 @@ Read [references/jina-reader.md](references/jina-reader.md) for Jina Reader posi
-**Forgetting the year in time-sensitive queries** — "best AI models" returns 2023 results; "best AI models 2026" returns current.
-**Hardcoding login credentials in scripts** — always rely on the user's pre-logged CDP session.
-**Citing only after the fact** — collect URLs as you fetch, not from memory afterwards.
-**(v2.0) Writing Python heredoc when BrowserXxx would do** — slow, requires Python+Playwright install, and bloats context. Prefer L3-fast; fall back to Python only when BrowserXxx is insufficient (race / long-wait / custom scripts).
-**(v2.0) Discovering new pitfalls and not writing a site-pattern** — next time the same Agent runs the task, it'll repeat the same mistakes. Anything that took 2+ steps to figure out is worth `SitePatternWrite(scope='agent', mode='merge')`.
-**(v2.0) Writing cookies / phone numbers to scope='agent'** — that layer is Git-tracked and may be published to the marketplace. SitePatternWrite auto-downgrades, but don't deliberately write secrets to the agent layer.
-**(v2.1) Writing Python heredoc when the built-in browser would do** — slow, requires Python+Playwright install, and bloats context. Prefer L3-fast; fall back to Python only when the built-in browser is insufficient (race / long-wait / custom scripts).
-**(v2.1) Calling `page.screenshot` before `tab.activate`** — tabs are parked off-screen at `(-10000,-10000,1x1)` and have no compositing surface, so the capture stalls until the 30s deadline **and the tab host is destroyed**; every later call then fails with `BROWSER_TAB_HOST_NOT_FOUND`. Always `tab.activate` with real bounds first.
-**(v2.1) Reaching for `page.evaluate` to read a page** — it needs human approval on every call, its string/object return values come back as `[REDACTED:browser-runtime-value]` (only number/boolean/null survive), and anti-debug sites stall it for tens of seconds. Read pages with `BrowserSnapshot` instead.
-**(v2.1) Discovering new pitfalls and not writing a site-pattern** — next time the same Agent runs the task, it'll repeat the same mistakes. Anything that took 2+ steps to figure out is worth `SitePatternWrite(scope='agent', mode='merge')`.
-**(v2.1) Writing cookies / phone numbers to scope='agent'** — that layer is Git-tracked and may be published to the marketplace. SitePatternWrite auto-downgrades, but don't deliberately write secrets to the agent layer.
---