---
name: wecom-faq-bot
display_name: 企微百问百答机器人项目
description: 企微百问百答机器人项目(WeCom FAQ Bot)—— 当保险/金融类产品的知识库需要以两种形态对外发布时使用本 skill:(1) 同步为企微(企业微信)客户群里的关键词自动回复规则集;(2) 同步为一份可公开访问的 HTTPS 网页知识库(部署到 Cloudflare Pages)。skill 覆盖完整流水线:解析规则表、按企微单条规则 10 关键词上限做拆分、检测关键词冲突、用持久化 Chrome profile 驱动企微管理后台、批量录入、逐页全量核对,以及通过 API 把静态 HTML 知识库部署到 Cloudflare Pages。触发词:「录入企微自动回复」「同步到企微」「把知识库挂到客户群」「把 HTML 知识库部署上线」「Cloudflare Pages 部署」「verify auto-reply rules」「check keyword conflicts」及类似表述。
agent_created: true
---
# 企微百问百答机器人项目 (WeCom FAQ Bot)
End-to-end pipeline for taking a product knowledge base (currently tuned for 装修险/装修保险 type products) and publishing it in two synchronized surfaces:
1. **WeCom (企业微信) auto-reply rules** — keyword-driven replies inside customer-facing external groups. The bot answers questions, the user can also send one of the navigation words to reach the bot.
2. **A public HTTPS page** — the full knowledge base as a single self-contained HTML file, deployed to Cloudflare Pages so a stable `*.pages.dev` URL can be shared in WeCom replies or sent directly.
The two surfaces stay in sync: every rule in the WeCom bot corresponds to a section/group in the HTML page, and the HTML's `https://<project>.pages.dev` URL is included verbatim in the E-series bot reply.
## When to use
- A new knowledge base (v1 / v2 / v3 …) needs to replace the current one in WeCom and on the public page.
- A new chunk of rules needs to be **added** to the existing 157+ rule set without breaking existing matches.
- A new `*.pages.dev` link needs to be wired into a WeCom auto-reply (e.g., a new "知识库入口" / E-series rule pointing to the latest HTML build).
- The WeCom admin console needs a partial refresh / re-entry after the underlying rule sheet (xlsx) changes.
## Hard constraints (do not violate)
- WeCom allows **≤10 keywords per auto-reply rule** and exposes an exact-match toggle in the UI; rules longer than 10 keywords must be **split** into `(.../N)` parts, in a specific recording order (see `references/wecom_ui.md`).
- WeCom auto-reply is **enterprise-wide** — there is no per-region routing. Multi-region content differences must be handled inside the reply text itself.
- The driver **only works on a persistent Chrome profile** because the WeCom admin session is a QR scan login that expires within a few hours; the profile must live outside the working directory so a re-run can pick up a still-valid cookie.
- The bot reply can be sent to customers in seconds; therefore any new reply text or keyword list must be **conflict-checked** before it touches the console (see "Mandatory keyword check" below).
- The `*.pages.dev` domain is **partially poisoned** by some China-mainland ISPs and **not ICP-registered**. Treat it as a "try-it-out" link, not a customer-facing one. A custom ICP-registered domain is the long-term answer.
## Workflow (do this in order)
### 0. Confirm scope with the user
Use `AskUserQuestion` or a plain text question to confirm three things before touching anything:
- Which surface to update (bot rules / public page / both).
- Whether to add new rules or replace an existing set (replace = delete the current set first, see Step 5 below).
- The region/口径 the answers must follow (default = 广东).
### 1. Prepare the rule sheet
The canonical source is a `.xlsx` file with one row per logical rule and columns for `name`, `group`, `keywords`, `match`, `reply`, `source`, `note`. Read it with openpyxl and normalize into a list of dicts called `vN_clean.json` (see `references/wecom_ui.md` for the schema).
If the source is already JSON, skip the xlsx step.
### 2. Split oversized rules + enforce recording order
WeCom's hard 10-keyword limit means every rule whose `keywords` length > 10 must become N rules named `<base> (1/N)`, `(2/N)`, …, `(N/N)`. Output: `entry_data_vN_split.json`.
**Recording order inside that file matters** because WeCom replies to the **first** matching rule it finds in the list. Use this priority:
1. 细分规则 first (B / C / D / F — concrete sub-questions with specific keywords).
2. G 递进导航 (broad navigation words like "目录", "导航", "菜单").
3. A 总览 (A01 / A05 / A06 generic intros — kept at the bottom of the inclusion group so they don't swallow everything).
4. H 完全匹配 single words (must be **完全匹配**, not 包含; the recording order puts them last so they are only used as a fallback).
5. E 服务 / 工具类 (knowledge-base links, transfer-to-human, etc.) at the very end.
### 3. **Mandatory keyword conflict check** before entry
This is the single highest-risk step. Run `scripts/check_keywords.py` on the **merged** keyword set (existing 1246+ keywords in the live console + the new ones). The script does a **bidirectional** substring check:
- For every new "包含匹配" keyword, does it **contain** any existing 包含 keyword? If yes, a user message containing the new keyword will also trigger the existing rule → dual reply or wrong reply. Reject or rewrite the new keyword.
- For every new "完全匹配" keyword, is it **contained in** any existing 包含 keyword? Same dual-trigger risk. Reject or rewrite.
The script also flags any **exact duplicate** of an existing keyword, which silently replaces rather than coexists. Do not ignore any of these warnings; trim / paraphrase the new keyword, or place it under 完全匹配 if appropriate.
Save the conflict report and attach it to the user-visible summary; only proceed to entry if the report is empty.
### 4. Drive the WeCom console
The driver is `scripts/driver.js` (Node, playwright-core, port 9777). Start it in the background:
```bash
cd <skill>/scripts
NODE_PATH=<node-managed-workspace>/node_modules node driver.js
```
Endpoints (full list in `references/driver_api.md`): `/launch`, `/goto`, `/click`, `/fill`, `/eval`, `/snapshot`, `/screenshot`, `/upload`, `/status`, `/close`.
For each rule in `entry_data_vN_split.json`, `scripts/batch_entry.py` calls:
1. `POST /goto` with `ADD_URL` (`https://work.weixin.qq.com/wework_admin/frame#customer/autoReply/add`).
2. `POST /fill` with `input.ww_inputText_Big >> nth=0` (rule name) — pause 0.4s.
3. If the rule is "完全匹配", call the `select_exact_match()` helper (open `[class*=csAutoReply_inputWidthDropMe] .ww_btn_Dropdown`, click the menu item whose text is "完全匹配", then verify the dropdown text is "完全匹配").
4. Fill keyword #1 into `input.ww_inputText_Big >> nth=1`.
5. For keywords #2..N, **click** `a.ww_inputWidthDropMenuGroup_addLink` (0.9s wait — the new input field renders after a brief delay), then fill into the new `input.ww_inputText_Big >> nth=<i+1>`.
6. `POST /fill` with the `textarea` selector and the reply text.
7. `POST /click` on `a.ww_btn_Blue` (the save button — not a class you might guess, this is the one) and wait 3.5s.
8. `POST /goto` to the list page, then `/eval` to scrape the first 5 row name cells; verify the new rule is in there. If not, the entry is recorded as failed and a retry is run.
The /eval-based verification step catches the common failure mode of WeCom silently refusing the save (e.g., an empty keyword after a missed `addLink`).
If the user has **not** scanned the WeCom QR code recently, the navigation will land on `loginpage_wx` instead of the admin page. Detect this by checking the URL and `document.body.innerText`; if it's a login page, take a screenshot, show the user, and stop until they confirm scan completion.
### 5. Verify by paginating the full list
WeCom list page shows 10 rules per page. The full v5 set is 16 pages. Run `scripts/verify_rules.py`:
1. Goto list page.
2. Loop: scrape all 10 rows, collect names; click the right-arrow `ww_commonImg_PageNavArrowRightNormal` (or its class-hash variant `PageNavArrowRightNormal` — the class uses woff2 sprites so the class name is what matters, not the visual icon); when the arrow is disabled, stop.
3. After 16 pages, if the collected names match `entry_data_vN_split.json` exactly (same count, same set, no extras, no duplicates), report success.
### 6. (When replacing an existing set) Bulk delete
If the new version supersedes an old one, run `scripts/batch_delete.py` **before** the new entry:
1. Goto list.
2. For each row, hover, click the row's "删除" link, wait for the confirm dialog (`qui_dialog`), click the confirm button, wait for the dialog to close.
Delete 10 at a time and re-check the list after each batch — accidentally removing the wrong row is unrecoverable.
### 7. Generate + deploy the public HTML
The HTML knowledge base is one self-contained file. Build it with whatever generator fits the source format; the canonical example for 装修险 is in the project (78.8KB single HTML, no CDN, all CSS inline). The deliverable is `<project>-vN.html`, renamed to `index.html` for Pages.
Deploy via Cloudflare Pages REST API + `wrangler`. Run `scripts/cf_helpers.sh` to:
1. **Verify the token** (`GET /user/tokens/verify`).
2. **Get the account id** (`GET /accounts` — needs `Account Settings: Read` in the token).
3. **Create the Pages project** (`POST /accounts/{id}/pages/projects` with `{"name":"<project>","production_branch":"main"}`).
Then run `scripts/cf_deploy.sh` to push the file:
```bash
# Install wrangler with --ignore-scripts because esbuild's postinstall is blocked by the sandbox:
npm install --prefix <node-managed-workspace> wrangler --ignore-scripts --no-fund --no-audit
# Deploy:
CLOUDFLARE_API_TOKEN=... CLOUDFLARE_ACCOUNT_ID=... \
node <node-managed-workspace>/node_modules/wrangler/bin/wrangler.js \
pages deploy <deploy-dir> --project-name=<project> --branch=main --commit-dirty=true
```
The `pages deploy` call uploads `index.html` and returns the live URL of the form `https://<project>.pages.dev` plus a per-deployment URL `https://<hash>.<project>.pages.dev`. Verify the production URL returns HTTP 200 with the expected title before declaring success.
### 8. Add the URL to the bot
If a new E-series rule ("knowledge-base entry") doesn't yet exist in the live set, add it now. The reply text must contain the just-deployed `*.pages.dev` URL, plus a 收藏 prompt and a fallback for the rare case the page can't load. Reuse `scripts/batch_entry.py` with a 2-rule payload (one 包含匹配, one 完全匹配, ≤10 keywords each — see `references/wecom_ui.md` for the keyword list pattern and the conflict-check pass that goes with it).
## Required environment
- **Chrome for Testing** at a stable path (the driver hard-codes it via `CHROME_PATH` env var, default `C:/Users/向/.agent-browser/browsers/chrome-<ver>/chrome.exe`).
- **Persistent profile directory** at `PROFILE_PATH` (default `C:/Users/向/.wecom-automation-profile`). It must be a directory playwright can `launchPersistentContext` against; do not delete it between runs.
- **playwright-core** in the Node workspace (`NODE_PATH=<workspace>/node_modules`).
- **Cloudflare API token** with the **Cloudflare Pages: Edit** + **Account Settings: Read** permissions.
- **Python 3.13** in the managed venv (`<python-managed>/envs/default/Scripts/python.exe`).
- **openpyxl** in the venv (for xlsx reading).
## Hard "don't"s
- Do **not** rely on the `/snapshot` (ariaSnapshot) endpoint for verification — it 500s on the WeCom list page; use `/eval` with `document.querySelectorAll(...)` instead.
- Do **not** trust a rule was saved just because the click succeeded. WeCom silently no-ops on bad input; always do the list-page re-verify.
- Do **not** put the bot reply URL inside a `pages.dev` link AND trust it for customers. Always include a fallback ("若链接打不开可直接在群里提问").
- Do **not** install `wrangler` without `--ignore-scripts` — the esbuild postinstall is blocked by the local sandbox and the install will hang.
- Do **not** try to install the skill's Node deps globally. Use the managed node workspace.
## Resources
### scripts/
- `driver.js` — persistent Chrome automation server (port 9777). Read it to understand endpoints, or start it and call via HTTP from Python.
- `check_keywords.py` — bidirectional conflict checker.
- `batch_entry.py` — full-pipeline driver of the WeCom console.
- `verify_rules.py` — 16-page pagination verifier.
- `batch_delete.py` — hover-and-delete with dialog handling.
- `cf_helpers.sh` — Cloudflare API calls (verify token, get account id, create project).
- `cf_deploy.sh` — wrangler install (with `--ignore-scripts` workaround) and `pages deploy`.
### references/
- `wecom_ui.md` — WeCom admin UI selectors, recording order, and the v5 schema.
- `driver_api.md` — full driver.js HTTP endpoint reference.
- `cloudflare_pages.md` — Cloudflare Pages REST API quick reference.
- `decision_log.md` — why we chose Pages-direct, why we don't use Gitee, why we don't trust WeCom multi-match, etc.
### assets/
- (none required; the HTML knowledge-base template is project-specific and lives in the project, not the skill)
参考资料 references/
🎯
wecom_ui.md
企微后台 UI 选择器表 + 10 关键词上限 + 录入顺序
MD94 行 · 6.7 KB
# WeCom Auto-Reply Admin UI Reference
This file is the source of truth for everything UI-side. The selectors and limits are not documented by WeCom — they are observed from the live admin console and may change when WeCom ships a UI update.
## Hard limits
- **≤10 keywords per rule.** The add-rule page literally renders 11 `input.ww_inputText_Big` elements on the page (1 for the rule name + 10 keyword slots), so the 11th keyword cannot be entered without page modification. Split the rule into `(1/N)` parts.
- **No per-region routing.** All rules apply enterprise-wide. Multi-region content differences must be handled inside the reply text.
- **No API.** No public endpoint exists for managing auto-reply rules; everything goes through the admin SPA at `work.weixin.qq.com/wework_admin/frame#customer/autoReply`.
- **Match types:** `包含` (contains) and `完全匹配` (exact). The default is `包含`. For H-class single-word rules, switch to `完全匹配` to avoid swallowing any message that happens to contain the word.
- **Multi-match behavior:** WeCom replies to the **first** matching rule in the list. Recording order is therefore the priority mechanism. (See "Recording order" below.)
- **Default reply + 30-minute anti-spam.** A rule can be marked as "default" and WeCom will only send it once per 30 minutes per user. The first rule in the list can be set as default with a generic fallback reply.
## URL routes
| Purpose | URL |
|---|---|
| List | `https://work.weixin.qq.com/wework_admin/frame#customer/autoReply` |
| Add | `https://work.weixin.qq.com/wework_admin/frame#customer/autoReply/add` |
| Edit | Same as Add with a query string in the hash; clicking "编辑" in a row navigates there. |
If the WeCom session has expired, both routes redirect to `https://work.weixin.qq.com/wework_admin/loginpage_wx?...`. Detect this with `location.href.includes("loginpage_wx")` and stop; the user has to re-scan the QR code in the browser window.
## Selectors (observed in the live admin UI as of 2026-10)
| Purpose | Selector | Notes |
|---|---|---|
| Rule name input | `input.ww_inputText_Big >> nth=0` | First big input on the add page. |
| Keyword input #1..N | `input.ww_inputText_Big >> nth=<i+1>` | nth=1 is the first keyword. The page renders 10 keyword slots, so the highest `nth` you'll ever use is 10. |
| Add keyword link | `a.ww_inputWidthDropMenuGroup_addLink` | Click to grow one more keyword input. Wait 0.9s after click — the new input renders after a brief delay; if you fill too fast, the fill lands in the wrong input. |
| Match-type dropdown | `[class*=csAutoReply_inputWidthDropMe] .ww_btn_Dropdown` | The button is the dropdown trigger. Click opens the menu. |
| Match-type menu item | `.ww_dropdownMenu_itemLink_text` | The menu items. Filter by `innerText.trim() === "完全匹配"` and click the matching one. |
| Reply textarea | `textarea` | Single textarea on the add page; safe to use bare `textarea`. |
| Save button | `a.ww_btn_Blue` | The blue "保存" / "Save" link. Click and wait 3.5s for the page to round-trip back to the list. |
| Delete link in a row | hover the row first, then look for a delete link in that row only | Use a JS query inside the row, not a global selector. |
| Delete confirm button | inside `.qui_dialog` (or the active dialog) | WeCom uses its own dialog component (`qui_dialog` class). |
| Pagination — right arrow | `img.ww_commonImg_PageNavArrowRightNormal` | The right-arrow icon (woff2 sprite). `aria-disabled` or its absence signals the last page. |
| Pagination — left arrow | `img.ww_commonImg_PageNavArrowLeftNormal` | Returns to the first page. |
| List row | `.ww_table tbody tr` | Each row's first cell holds the rule name. |
## Recording order (priority)
WeCom replies to the first matching rule. When two rules' keyword sets would both match the same user message, the rule that was **added earlier** (i.e. appears earlier in the list) wins. The recommended recording order for a fresh split dataset:
1. **细分规则 (B, C, D, F)** — concrete sub-questions with very specific keywords. These are the "deep" answers; they should be hit first.
2. **G 递进导航** — broad navigation words (e.g., "目录", "导航", "菜单", "主页"). These should be hit only if no细分 rule matched.
3. **A 总览 (A01, A05, A06)** — generic intros with the most permissive keyword lists. They are kept at the bottom of the inclusion group so they don't swallow everything.
4. **H 完全匹配** — single-word exact-match rules. These are fallbacks; if a细分 rule was going to match, it would have already replied.
5. **E 服务 / 工具** — knowledge-base entry, transfer-to-human, etc. These are operational; the user explicitly asks for them, so they can be last.
The numbering convention `A01 装修险-通用-产品介绍(1/3)` mirrors the structure of the source spreadsheet; `(1/3)` etc. indicates split chunks of one original rule.
## Schema (v5_clean.json)
```json
{
"idx": "A01",
"group": "通用",
"name": "A01 装修险-通用-产品介绍(1/3)",
"keywords": ["产品", "产品介绍", "装修险产品", "装修险介绍", "介绍装修", "装修险产品介绍", "装修保险", "装修保险产品", "装修保险介绍", "买什么险"],
"match": "包含匹配",
"reply": "本产品为「室内装修保险」…",
"source": "v5.xlsx 第 1 行",
"note": "细分:产品总览,第 1 段"
}
```
`match` is one of `"包含匹配"` or `"完全匹配"` (full strings, not abbreviations — the source spreadsheet uses the full names).
## Match-type switch helper
The `select_exact_match()` function in `scripts/batch_entry.py` does:
1. Click `[class*=csAutoReply_inputWidthDropMe] .ww_btn_Dropdown` (the dropdown trigger). If click fails, return False.
2. Sleep 0.6s (menu animation).
3. Find `.ww_dropdownMenu_itemLink_text` whose `innerText.trim()` is exactly `完全匹配` and click it.
4. Sleep 0.4s.
5. Verify by reading the dropdown's current text — must be `完全匹配`; otherwise return False.
The verifier step at the end is critical: WeCom sometimes silently rejects the change if the menu's z-index conflicts with the sticky footer.
## Verification via /eval (don't trust the screenshot)
The list page renders 10 rows per page, and after a successful save WeCom puts the new rule at the top of page 1. To verify, after a save:
```js
(function(){
var rows = document.querySelectorAll('.ww_table tbody tr');
return Array.from(rows).slice(0, 5).map(function(x) { return x.cells[0].innerText; }).join('||');
})()
```
If the new rule's name appears in the first 5, the save succeeded. This catches the common failure mode where a save click returns 200 OK but the form was actually rejected (e.g., a duplicate keyword, an over-length reply, or the exact-match dropdown didn't latch).
🔌
driver_api.md
浏览器自动化驱动的 HTTP 接口与注意事项
MD100 行 · 4.9 KB
# Driver HTTP API (scripts/driver.js)
A tiny Node HTTP server that exposes a long-lived playwright-core session. Started on `127.0.0.1:9777`. All endpoints are POST unless otherwise noted.
## Lifecycle
- Start: `node driver.js` (with `NODE_PATH` set to the managed node workspace).
- The driver keeps a single persistent Chrome context alive across requests. The context retains cookies / localStorage / IndexedDB, so a single QR-scan login keeps the WeCom admin session valid for hours.
- `/close` shuts down the browser and exits the process. After `/close`, the driver must be restarted; the next `/launch` reuses the profile directory.
## Configuration (env vars)
| Var | Default | Purpose |
|---|---|---|
| `CHROME_PATH` | `C:/Users/向/.agent-browser/browsers/chrome-154.0.8037.57/chrome.exe` | Chrome for Testing executable. |
| `PROFILE_PATH` | `C:/Users/向/.wecom-automation-profile` | Persistent profile directory for `launchPersistentContext`. |
| `PORT` | `9777` | HTTP port. |
| `SHOT_DIR` | `<project>/.企微自动回复配置包-室内装修保险百问百答.ref/shots` | Where `/screenshot` writes. The endpoint ignores the `path` argument; this is a known limitation. |
## Endpoints
### POST /launch
Body: `{ "headless": false }`. Launches the persistent browser if not already running. Idempotent.
Response: `{ ok: true, headless: bool }`.
### GET /status
Returns the current browser state without modifying anything.
Response: `{ ok: true, browser: bool, url: string, title: string }`.
### POST /goto
Body: `{ "url": "https://..." }`. Navigates the current page; uses `waitUntil: "domcontentloaded"` with a 60s timeout, then waits 1.5s for the SPA to settle.
Response: `{ ok: true, url: <final URL>, title: <page title> }`.
### POST /click
Body: `{ "selector": "input.foo", "timeout": 10000 }`. Clicks the first element matching the selector. Waits 0.8s after.
Response: `{ ok: true, url: <current URL> }`.
### POST /fill
Body: `{ "selector": "input.foo >> nth=0", "text": "value", "timeout": 10000 }`. Fills the input. **No trailing wait** — chain a `/wait` or include the wait in the calling code's sleep.
Response: `{ ok: true }`.
### POST /press
Body: `{ "selector": "input.foo", "key": "Enter" }`. Sends a key press to the focused / specified element.
### POST /wait
Body: either `{ "text": "立即购买", "timeout": 30000 }` (waits for that text to appear) or `{ "ms": 1000 }` (raw wait).
### POST /eval
Body: `{ "code": "(function(){...})()" }`. Runs the code in the page's main world. Return value is JSON-serialized.
This is the workhorse. Use it to read DOM, click items inside rows, switch dropdowns, etc.
Response: `{ ok: true, result: <JSON> }`.
### POST /snapshot
Returns an aria-snapshot. **Warning: returns 500 on the WeCom list page** (the rendering is too deep for playwright's aria engine). Do not use it for verification — use `/eval` with `querySelectorAll`.
Response on success: `{ ok: true, url: ..., snapshot: "..." }`. On failure, the response includes the first 3000 chars of `document.body.innerText` as a fallback.
### POST /screenshot
Saves a fullPage=false PNG to `SHOT_DIR/shot-<timestamp>.png`. The `path` argument in the body is **ignored** — to find the file, `ls -t` the SHOT_DIR.
Response: `{ ok: true, file: "<absolute path>" }`.
### POST /upload
Body: `{ "selector": "input[type=file] >> nth=0", "files": ["C:/path/to/file.html", "C:/path/to/other.css"] }`. Uses `page.setInputFiles` to upload local files via a hidden `<input type=file>` (used by Cloudflare Pages dashboard drag-and-drop, and the WeCom product import flow if it's ever needed).
### POST /close
Closes the browser context and exits the process after 300ms. After this, the script must be restarted.
## Why a separate HTTP layer?
Two reasons:
1. **Long-lived context.** A direct `node script.js` invocation that does `chromium.launchPersistentContext()` per run would re-launch Chrome and re-do the QR scan every time. The HTTP layer keeps the context warm.
2. **Python orchestration.** The higher-level pipeline (`batch_entry.py`, `verify_rules.py`) is in Python because xlsx parsing, conflict checking, and progress logging are all easier in Python. Driving playwright via HTTP from Python is `urllib.request`; no SDK lock-in.
## Gotchas
- The `page` reference can become stale after a navigation opens a new tab (e.g., the Cloudflare login flow). The driver auto-recovers via `ensurePage()`, which finds the first non-closed page in the context and re-binds `page`. The caller does not need to re-`/launch`.
- A `/goto` that lands on a `loginpage_wx` URL is the universal signal that the WeCom session expired. The calling Python code should detect this and stop, then ask the user to re-scan.
- Chrome's `--no-sandbox` and `--disable-gpu` are set because the sandbox is incompatible with how the driver is started (the `playwright-core` launch is non-elevated). GPU is disabled because the test machine often runs headless on a VM.
☁️
cloudflare_pages.md
Cloudflare Pages 部署速查 + 平台选型对比
MD110 行 · 5.6 KB
# Cloudflare Pages REST API + Wrangler Quick Reference
## API endpoints (used in `scripts/cf_helpers.sh`)
Base URL: `https://api.cloudflare.com/client/v4`. All requests take `Authorization: Bearer <token>` and `Content-Type: application/json`.
### Verify the token
```
GET /user/tokens/verify
```
Response: `{ "result": { "id": "...", "status": "active" }, "success": true, ... }`.
A 403 or `success: false` means the token is invalid or the requested scope is denied. The token must have been created with the right permissions **for the account that owns the project**.
### List accounts
```
GET /accounts
```
Response: `{ "result": [{ "id": "<32-hex>", "name": "...", "type": "standard" }, ...] }`.
Requires `Account Settings: Read` permission on the token. The first account is the default; if there are multiple, the user must pick one.
### Create a Pages project
```
POST /accounts/{account_id}/pages/projects
Body: { "name": "zx-knowledge-base", "production_branch": "main" }
```
Response: includes `subdomain: "<name>.pages.dev"` — that is the production URL. The project is empty until you call `pages deploy`.
Project name rules: lowercase letters, digits, and hyphens; ≤ 58 chars; must be unique within the account. Reusing an existing name returns 409 (project already exists) — in that case skip the create call.
### List projects (sanity check)
```
GET /accounts/{account_id}/pages/projects
```
Returns the project's `subdomain` and `latest_deployment.id`. Useful for confirming a deploy succeeded.
## Wrangler
The official Cloudflare CLI. We use it for `pages deploy`, which takes a local directory and pushes every file as a static asset.
### Install (in the managed node workspace)
```
npm install --prefix <node-managed-workspace> wrangler --ignore-scripts --no-fund --no-audit
```
`--ignore-scripts` is **required** because the esbuild postinstall is blocked by the local sandbox (it spawns a downloaded `esbuild.exe` to verify the binary, and the sandbox kills the spawn). Without this flag, `npm install` exits 1 and leaves the workspace in a half-installed state. Wrangler's main binary is a pre-built JS file under `node_modules/wrangler/bin/wrangler.js` and doesn't need the postinstall.
### Deploy
```
CLOUDFLARE_API_TOKEN=... CLOUDFLARE_ACCOUNT_ID=... \
node <node-managed-workspace>/node_modules/wrangler/bin/wrangler.js \
pages deploy <dir> --project-name=<name> --branch=main --commit-dirty=true
```
`--branch=main` targets the production branch (matching `production_branch` in the create call). `--commit-dirty=true` is needed because the dir is not a git repo.
Output:
```
Uploading... (N/N)
✨ Success! Uploaded N files (X.XX sec)
🌎 Deploying...
✨ Deployment complete! Take a peek over at https://<hash>.<project>.pages.dev
```
The hash-prefixed URL is the per-deployment preview; the canonical production URL is `https://<project>.pages.dev` (no hash).
## Token permissions
The minimal token for this workflow:
| Scope | Permission | Why |
|---|---|---|
| Account | Cloudflare Pages: Edit | Create project, push deployments. |
| Account | Account Settings: Read | List accounts to discover the account id. |
Without `Account Settings: Read`, the `/accounts` call returns 403 and the user has to paste the account id manually.
## What Pages does **not** do
- **No ICP.** Pages serves from Cloudflare's global edge; there is no mainland China node and no way to attach an ICP number. For a customer-facing link in mainland China, a Pages URL is at the mercy of DNS poisoning. The long-term answer is to bind a custom ICP-registered domain (or to move to a domestic provider like Tencent EdgeOne).
- **No env vars / secrets at runtime.** Pages is a static host; the only "build-time" env vars are those passed to the build command. For a pure HTML deploy (no build), env vars are irrelevant.
- **No server-side code.** The HTML must be fully self-contained: all CSS inline, no CDN links (or links to a CDN the user can guarantee is reachable from mainland).
## Decision: why Pages and not Gitee / Vercel / Netlify / EdgeOne
- **Gitee Pages** — shut down by Gitee in 2025. Repository still works, but the public site URL is unavailable.
- **Vercel** — best mainland connectivity of the Western options (uses Vercel China DNS for custom domains), but its free tier aggressively suspends "inactive" projects and we have no reason to trust our use case won't trip that.
- **Netlify** — similar to Vercel but worse in mainland.
- **Tencent EdgeOne** — fast in mainland, but the default `*.edgeone.cool` domain has a 3-hour token-only preview when the acceleration region includes China mainland. Long-term hosting requires binding an ICP-registered custom domain. The 3-hour limit is hard-coded by compliance, not a switch.
- **Cloudflare Pages** — free, no 3-hour limit on the default domain, but `pages.dev` is itself a soft target for DNS poisoning. Acceptable as a "try-it-out" link.
**Recommended path:** start with Cloudflare Pages for fast iteration; when a custom ICP-registered domain becomes available, move the same `index.html` to EdgeOne with that domain bound. The HTML file is portable; the only thing that changes is where it lives.
## Hardening tips
- Make `index.html` self-contained: inline all CSS, inline all JS, no CDN links. Pages serves the file as-is and a CDN outage in mainland would otherwise 404 the whole site.
- Include the production URL inside the bot reply verbatim — do not rely on the user being able to type the project name and find it.
- For a China-friendly deploy, the single biggest win is the ICP-registered custom domain on a domestic edge. Do not try to be clever with reverse proxies on foreign providers; the GFW will eventually find them.
📝
decision_log.md
7 项关键设计决策的依据
MD56 行 · 5.4 KB
# Decision Log
A short, frank list of the calls we made, why, and what the alternatives were. Future instances of WorkBuddy (or a future you) should re-read this before deviating.
## Driver: persistent Chrome profile, not a fresh launch
We launch a single Chrome process via `playwright-core` with `launchPersistentContext` against a profile directory. We keep it alive across HTTP requests and re-bind `page` on every call via `ensurePage()`.
**Why not** just `chromium.launch()` (non-persistent) every time? WeCom requires a QR scan to log in. A non-persistent launch throws away the cookies every time, forcing a fresh scan for every script run. That is intolerable for a 157-rule batch entry that needs multiple sessions.
**Why not** use the user's default Chrome? Chrome 136+ refuses `--remote-debugging-port` against the default profile directory, and copying the whole default profile is wasteful (and the cookies are DPAPI-encrypted to the user, which actually does work for the same Windows user, but the size and sync activity make it flaky). A separate small profile is much simpler.
## Browser: Chrome for Testing, not Edge
The driver uses Chrome for Testing because the chromium that ships with playwright-core does not include an executable; we need to point at a real Chrome binary. Chrome for Testing is the supported, version-pinned Chrome binary that matches what playwright tests against, so we pick that over the user's system Chrome (which may be ahead/behind and may have extensions that confuse selectors).
## Batch entry: Python orchestrates, Node drives the browser
Python is the language of the data (xlsx, openpyxl, JSON manipulation) and the language of the surrounding business logic (conflict checking, recording order). Node is the language of the browser. Forcing either side to do both is a bad trade. The driver is a tiny HTTP server in Node; the pipeline is a Python script that uses `urllib.request` to call it. No SDK is needed on either side.
## WeCom match logic: trust the recording order
The official WeCom docs are silent on what happens when a message matches multiple rules. Empirically (and from the v3/v4 design notes in the original project), WeCom replies to the **first** matching rule. So the priority mechanism is the recording order, not a per-rule priority field.
**Implication:** every batch entry must be planned so that 细分 rules (specific keywords) come first, 通用 rules (broad keywords) come after, and 完全匹配 rules (single-word fallbacks) come last. The `split_rules.py` step encodes this.
## Keyword conflict check: bidirectional substring, not just exact duplicate
The obvious check is "is this keyword already in the live set?" — but that's not enough. The real failure mode is **dual reply**:
- A new 包含 rule with keyword `知识库链接` will match any message containing `知识库链接`, but the existing 包含 rule with keyword `链接` will also match. The user gets two replies. Bad.
- A new 完全 rule with keyword `手册` will match a message that is exactly `手册`, but the existing 包含 rule with keyword `在线手册` will *also* match a message that is exactly `手册`? No — `手册` is a substring of `在线手册`, but the message `手册` is the whole message and does not contain `在线手册`. So in this specific case there is no dual reply. But the inverse case (new 完全 rule with keyword `查询`, existing 包含 rule with keyword `查询` — exact duplicate) does cause a dual reply or a "first match wins" with the wrong rule winning.
The bidirectional check in `check_keywords.py` covers all of these.
## Pages deploy: Cloudflare Pages first, not EdgeOne
We tried EdgeOne first (because the 装修险 knowledge base is targeted at mainland China customers, where EdgeOne has a domestic edge and Cloudflare does not). EdgeOne's `*.edgeone.cool` default domain is **3-hour token-only** when the acceleration region includes mainland China — a hard-coded compliance rule with no switch. The only long-term path is to bind a custom ICP-registered domain, which is a multi-day IT process.
For a "try it out" deploy, the 3-hour token is annoying. Cloudflare Pages has the same mainland reachability problem (DNS poisoning of `pages.dev`) but no token expiry, so it's strictly better as a starter. We default to Pages and document the EdgeOne path as the upgrade.
## HTML knowledge base: one self-contained file
The HTML is a single file with all CSS inlined and all JS inlined, no CDN, no external font. Reasons:
- The file gets included in chat messages and email attachments; smaller and more portable.
- It must work offline for printing / PDF export.
- The Cloudflare Pages free tier does not have to worry about file count limits.
- The deploy step is then literally `wrangler pages deploy <dir>` and the `dir` only contains one file. Predictable.
## The skill's "mandatory keyword check" rule
The WeCom admin console is a dangerous surface: a bad rule with a too-permissive keyword will reply to thousands of customer messages, and a too-restrictive one will leave customers unanswered. Every batch entry should be preceded by a conflict check; every conflict should block the entry until the keyword is paraphrased. This is the only step where the skill can fail safely.
The check has a single output format (a list of `[new_keyword, conflicting_existing_keyword, existing_rule_name]` triples) so it's reviewable in 30 seconds.
# -*- coding: utf-8 -*-
"""
Bidirectional keyword conflict checker for WeCom auto-reply rules.
Usage:
python check_keywords.py <existing.json> <new.json>
Both files are lists of rule dicts with at least `name`, `keywords`, `match`.
`match` is one of "包含匹配" or "完全匹配" (full Chinese strings).
The script detects three failure modes:
1. EXACT DUP: a new keyword is byte-identical to an existing keyword.
WeCom silently drops the duplicate.
2. CONTAINED: a new 包含 keyword contains an existing 包含 keyword.
Any user message containing the new keyword will also trigger the
existing rule -> dual reply.
3. SUPERSET: a new 完全 keyword is contained in an existing 包含 keyword.
A user message that is exactly the new keyword is also a hit for
the existing 包含 rule -> dual reply.
Output (stdout, plain text): a list of conflict triples and a 1-line summary.
Exit code 0 if no conflicts, 1 if any.
The script is parameter-free on purpose: it expects the data, computes
everything, and prints. No flags, no config files. Run it before every
batch entry, read the output, and rewrite the offending keywords.
"""
import json, io, sys
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
sys.stderr = sys.stdout
def load(path):
with open(path, encoding="utf-8") as f:
d = json.load(f)
# Tolerate both list and {rules: [...]} formats.
if isinstance(d, dict) and "rules" in d:
d = d["rules"]
return d
def kw_map(rules):
"""Return ({kw -> rule_name} for 包含, same for 完全)."""
inc, exc = {}, {}
for r in rules:
for k in r.get("keywords", []):
if "完全" in str(r.get("match", "")):
exc[k] = r["name"]
else:
inc[k] = r["name"]
return inc, exc
def check(existing_rules, new_rules):
inc_old, exc_old = kw_map(existing_rules)
inc_new, exc_new = kw_map(new_rules)
conflicts = []
for kw, owner in inc_new.items():
# EXACT DUP
if kw in inc_old:
conflicts.append((kw, "EXACT_DUP_CONTAIN", owner, inc_old[kw]))
if kw in exc_old:
conflicts.append((kw, "EXACT_DUP_EXACT", owner, exc_old[kw]))
# CONTAINED: new kw contains an old 包含 kw
for ok, oname in inc_old.items():
if ok and ok != kw and ok in kw:
conflicts.append((kw, "NEW_CONTAINS_OLD_INC", owner, oname))
# SUPERSET: new kw is a substring of an old 包含 kw
for ok, oname in inc_old.items():
if ok and ok != kw and kw in ok:
conflicts.append((kw, "NEW_IN_OLD_INC", owner, oname))
for kw, owner in exc_new.items():
if kw in inc_old:
conflicts.append((kw, "NEW_EXACT_IN_OLD_INC", owner, inc_old[kw]))
if kw in exc_old:
conflicts.append((kw, "EXACT_DUP_EXACT", owner, exc_old[kw]))
return conflicts
def main():
if len(sys.argv) < 3:
print("usage: python check_keywords.py <existing.json> <new.json>", file=sys.stderr)
sys.exit(2)
existing = load(sys.argv[1])
new = load(sys.argv[2])
conflicts = check(existing, new)
if not conflicts:
print(f"OK · 0 conflicts · {sum(len(r.get('keywords', [])) for r in new)} new keywords checked")
return
print(f"FAIL · {len(conflicts)} conflicts:")
for kw, kind, new_owner, old_owner in conflicts:
print(f" [{kind}] new '{kw}' ({new_owner}) vs old '{old_owner}'")
sys.exit(1)
if __name__ == "__main__":
main()
📥
batch_entry.py
单条规则录入(名→关键词→匹配方式→回复→保存)
PY178 行 · 7.3 KB
# -*- coding: utf-8 -*-
"""
Batch entry of WeCom auto-reply rules via the driver on port 9777.
Usage:
python batch_entry.py <entry_data.json> [start_index]
`entry_data.json` is a list of rule dicts (or a {rules: [...]} envelope).
Each rule needs: name, keywords (list, <=10), reply, and optionally match
(default "包含匹配"). Recording order inside the file IS the priority
order in the live list -- put 细分 rules first, 完全匹配 rules last.
Behavior:
- /goto ADD_URL, fill name, switch to 完全匹配 if needed, fill keywords
one at a time (clicking addLink between them), fill reply, click save.
- After every save, /goto LIST_URL, scrape the first 5 row names, verify
the new rule is there. If not, the entry is recorded as failed.
- Failures do not abort the run; they're collected and printed at the end.
Selectors and timings are tuned to the live WeCom admin UI as of 2026-10.
If WeCom ships a UI change, edit here -- this is the single point of truth.
"""
import io, json, sys, time, urllib.error, urllib.request
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
sys.stderr = sys.stdout
BASE = "http://127.0.0.1:9777"
LIST_URL = "https://work.weixin.qq.com/wework_admin/frame#customer/autoReply"
ADD_URL = "https://work.weixin.qq.com/wework_admin/frame#customer/autoReply/add"
def api(path, payload=None, method="GET", timeout=60):
if payload is None:
req = urllib.request.Request(BASE + path)
else:
body = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(BASE + path, data=body, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
return json.loads(r.read().decode("utf-8"))
except urllib.error.HTTPError as e:
return {"_http_error": e.code, "_body": e.read().decode("utf-8", "ignore")[:200]}
except Exception as e:
return {"_exc": type(e).__name__, "_msg": str(e)}
def api_retry(path, payload=None, method="GET", timeout=60, tries=3):
for i in range(tries):
r = api(path, payload, method, timeout)
if "_http_error" not in r and "_exc" not in r:
return r
if i < tries - 1:
time.sleep(2)
return r
def log(msg):
print(time.strftime("%H:%M:%S ") + msg, flush=True)
def verify_rule(rule):
"""After save, list page re-orders with the new rule at the top of page 1."""
r = api_retry("/goto", {"url": LIST_URL}, "POST")
time.sleep(2.2)
r = api_retry("/eval", {"code": "(function(){var rows=document.querySelectorAll('.ww_table tbody tr');return Array.from(rows).slice(0,5).map(function(x){return x.cells[0].innerText;}).join('||');})()"}, "POST")
if "_http_error" in r or "_exc" in r:
return False
return rule["name"] in r.get("result", "")
def select_exact_match():
"""Click the match-type dropdown, then click the 完全匹配 menu item, then verify."""
r = api_retry("/eval", {"code": "(function(){var b=document.querySelector('[class*=csAutoReply_inputWidthDropMe] .ww_btn_Dropdown');if(!b)return 'no btn';b.click();return 'opened';})()"}, "POST")
if r.get("result") != "opened":
return False
time.sleep(0.6)
r = api_retry("/eval", {"code": "(function(){var el=Array.from(document.querySelectorAll('.ww_dropdownMenu_itemLink_text')).find(function(e){return e.innerText.trim()==='完全匹配';});if(!el)return 'no item';el.click();return 'selected';})()"}, "POST")
if r.get("result") != "selected":
return False
time.sleep(0.4)
r = api_retry("/eval", {"code": "(function(){var b=document.querySelector('[class*=csAutoReply_inputWidthDropMe] .ww_btn_Dropdown');return b?b.innerText.trim():'?';})()"}, "POST")
return r.get("result") == "完全匹配"
def enter_rule(rule):
r = api_retry("/goto", {"url": ADD_URL}, "POST")
if "_http_error" in r or "_exc" in r:
log(f" goto add fail: {r}")
return False
time.sleep(3.0)
r = api_retry("/fill", {"selector": "input.ww_inputText_Big >> nth=0", "text": rule["name"]}, "POST")
if not r.get("ok"):
log(f" fill name fail: {r}")
return False
time.sleep(0.4)
if "完全" in str(rule.get("match", "")):
if not select_exact_match():
log(" select exact match fail")
return False
kws = rule["keywords"]
if len(kws) > 10:
log(f" too many keywords ({len(kws)}), max 10")
return False
r = api_retry("/fill", {"selector": "input.ww_inputText_Big >> nth=1", "text": kws[0]}, "POST")
if not r.get("ok"):
log(f" fill kw1 fail: {r}")
return False
time.sleep(0.25)
for i in range(1, len(kws)):
r = api_retry("/click", {"selector": "a.ww_inputWidthDropMenuGroup_addLink"}, "POST")
if not r.get("ok"):
log(f" click addLink {i} fail: {r}")
return False
time.sleep(0.9) # new input needs a beat to render
r = api_retry("/fill", {"selector": f"input.ww_inputText_Big >> nth={i+1}", "text": kws[i]}, "POST")
if not r.get("ok"):
log(f" fill kw{i+1} fail: {r}")
return False
time.sleep(0.15)
r = api_retry("/fill", {"selector": "textarea", "text": rule["reply"]}, "POST")
if not r.get("ok"):
log(f" fill reply fail: {r}")
return False
time.sleep(0.4)
r = api_retry("/click", {"selector": "a.ww_btn_Blue"}, "POST")
if not r.get("ok"):
log(f" click save fail: {r}")
return False
time.sleep(3.5)
return verify_rule(rule)
def main():
if len(sys.argv) < 2:
print("usage: python batch_entry.py <entry_data.json> [start_index]", file=sys.stderr)
sys.exit(2)
data = json.load(open(sys.argv[1], encoding="utf-8"))
rules = data["rules"] if isinstance(data, dict) and "rules" in data else data
# Pre-flight: make sure the driver is launched and we are NOT on the login page.
r = api_retry("/launch", {"headless": False}, "POST")
if not r.get("ok"):
log(f"driver launch failed: {r}")
sys.exit(1)
r = api_retry("/goto", {"url": LIST_URL}, "POST")
time.sleep(2.0)
r = api_retry("/status", None, "GET")
log(f"driver status: url={r.get('url')}")
if "loginpage_wx" in (r.get("url") or ""):
log("!! WeCom session expired. User must re-scan the QR code in the browser window.")
log(" Take a screenshot, show the user, then run this script again.")
sys.exit(1)
start = int(sys.argv[2]) if len(sys.argv) > 2 else 0
log(f"=== entry start: {len(rules)-start} rules (from index {start}) ===")
ok_n, fail = 0, []
for idx, rule in enumerate(rules):
if idx < start:
continue
try:
t0 = time.time()
ok = enter_rule(rule)
dt = time.time() - t0
if ok:
ok_n += 1
log(f"[{idx+1}/{len(rules)}] {rule['name']} OK kw={len(rule['keywords'])} match={rule.get('match','包含')} {dt:.1f}s")
else:
fail.append(rule["name"])
log(f"[{idx+1}/{len(rules)}] {rule['name']} FAIL")
except Exception as e:
fail.append(rule["name"])
log(f"[{idx+1}/{len(rules)}] {rule['name']} ERROR {type(e).__name__}: {e}")
log(f"=== done: ok={ok_n}, fail={len(fail)} {fail} ===")
if __name__ == "__main__":
main()
🗑️
batch_delete.py
单条规则删除 + 弹窗确认
PY151 行 · 5.9 KB
# -*- coding: utf-8 -*-
"""
Bulk delete WeCom auto-reply rules by name.
Usage:
python batch_delete.py <names.json>
`names.json` is a list of rule names (strings). The script paginates the
list page, and for each page that contains one of the target names, it
clicks the row's "删除" link, confirms the dialog, and moves on.
Hard rules (enforced in code):
- One row at a time, not in parallel.
- After every 10 deletions, re-scrape the list and continue. If the
list has been disturbed, abort.
- Refuses to delete a name that is NOT in the input file (defensive:
catches off-by-one in the row indexing).
Output: a 1-line summary. Exit 0 on success, 1 if any target remained.
"""
import io, json, sys, time, urllib.error, urllib.request
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
sys.stderr = sys.stdout
BASE = "http://127.0.0.1:9777"
LIST_URL = "https://work.weixin.qq.com/wework_admin/frame#customer/autoReply"
def api(path, payload=None, method="GET", timeout=60):
if payload is None:
req = urllib.request.Request(BASE + path)
else:
body = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(BASE + path, data=body, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
return json.loads(r.read().decode("utf-8"))
except urllib.error.HTTPError as e:
return {"_http_error": e.code, "_body": e.read().decode("utf-8", "ignore")[:200]}
except Exception as e:
return {"_exc": type(e).__name__, "_msg": str(e)}
def api_retry(path, payload=None, method="GET", timeout=60, tries=3):
for i in range(tries):
r = api(path, payload, method, timeout)
if "_http_error" not in r and "_exc" not in r:
return r
if i < tries - 1:
time.sleep(2)
return r
def click_next_page():
r = api_retry("/eval", {"code": "(function(){var a=document.querySelector('img.ww_commonImg_PageNavArrowRightNormal');if(!a)return 'no arrow';a.click();return 'clicked';})()"}, "POST")
return r.get("result") == "clicked"
def goto_first_page():
for _ in range(20):
r = api_retry("/eval", {"code": "(function(){var a=document.querySelector('img.ww_commonImg_PageNavArrowLeftNormal');return a?'yes':'no';})()"}, "POST")
if r.get("result") != "yes":
return
api_retry("/click", {"selector": "img.ww_commonImg_PageNavArrowLeftNormal"}, "POST")
time.sleep(1.2)
def scrape_page_with_rows():
"""Return [{name, hasDelete}] for every row on the current page."""
code = """
(function(){
var rows = document.querySelectorAll('.ww_table tbody tr');
return Array.from(rows).map(function(row){
var name = (row.cells[0] ? row.cells[0].innerText.trim() : '');
var hasDelete = !!Array.from(row.querySelectorAll('a')).find(function(a){
return a.innerText && a.innerText.indexOf('删除') >= 0;
});
return {name: name, hasDelete: hasDelete};
}).filter(function(x){return x.name;});
})()
"""
r = api_retry("/eval", {"code": code}, "POST")
return r.get("result") or []
def delete_row_at(idx):
"""Click the delete link in row `idx` of the current page, confirm dialog."""
api_retry("/eval", {"code": f"(function(){{var rows=document.querySelectorAll('.ww_table tbody tr');var row=rows[{idx}];if(!row)return 'no row';var link=Array.from(row.querySelectorAll('a')).find(function(a){{return a.innerText && a.innerText.indexOf('删除')>=0;}});if(!link)return 'no link';link.click();return 'clicked';}})()"}, "POST")
time.sleep(0.8)
api_retry("/eval", {"code": "(function(){var dlg=document.querySelector('.qui_dialog');if(!dlg)return 'no dlg';var btn=Array.from(dlg.querySelectorAll('a,button')).find(function(b){return b.innerText && b.innerText.indexOf('确定')>=0;});if(!btn)return 'no btn';btn.click();return 'confirmed';})()"}, "POST")
time.sleep(1.5)
def main():
if len(sys.argv) < 2:
print("usage: python batch_delete.py <names.json>", file=sys.stderr)
sys.exit(2)
targets = set(json.load(open(sys.argv[1], encoding="utf-8")))
api_retry("/launch", {"headless": False}, "POST")
api_retry("/goto", {"url": LIST_URL}, "POST")
time.sleep(2.5)
r = api_retry("/status", None, "GET")
if "loginpage_wx" in (r.get("url") or ""):
print("ABORT: WeCom session expired, re-scan QR and retry.", file=sys.stderr)
sys.exit(1)
goto_first_page()
time.sleep(1.5)
deleted, still_there = 0, set()
for batch_round in range(50):
progressed = False
page = 1
while True:
rows = scrape_page_with_rows()
if not rows:
break
for idx, row in enumerate(rows):
if row["name"] in targets and row["hasDelete"]:
if row["name"] in still_there:
still_there.discard(row["name"])
print(f" delete: {row['name']}")
delete_row_at(idx)
deleted += 1
progressed = True
if not click_next_page():
break
time.sleep(1.5)
page += 1
if not progressed:
break
goto_first_page()
time.sleep(1.5)
# Final scrape to find anything that didn't get deleted.
goto_first_page()
time.sleep(1.5)
all_live = set()
while True:
rows = scrape_page_with_rows()
all_live.update(r["name"] for r in rows)
if not click_next_page():
break
time.sleep(1.5)
still_there = targets & all_live
print(f"deleted {deleted}, remaining {len(still_there)}")
if still_there:
print("FAIL: " + ", ".join(sorted(still_there)))
sys.exit(1)
print("OK")
if __name__ == "__main__":
main()
✅
verify_rules.py
全量翻页核对(缺失/多余/重复)
PY122 行 · 4.6 KB
# -*- coding: utf-8 -*-
"""
Paginate the WeCom auto-reply list page and verify the live rule set
matches the expected JSON.
Usage:
python verify_rules.py <expected.json>
The expected file is the same shape as batch_entry's input (list of rule
dicts with at least `name`).
Output: prints the names scraped from every page, then a 1-line summary
(success: counts match, set match; failure: the diff). Exit 0 on success,
1 on mismatch.
The script also handles the case where the live list page redirects to
loginpage_wx: it aborts with a clear message instead of producing a false
mismatch.
"""
import io, json, sys, time, urllib.error, urllib.request
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
sys.stderr = sys.stdout
BASE = "http://127.0.0.1:9777"
LIST_URL = "https://work.weixin.qq.com/wework_admin/frame#customer/autoReply"
def api(path, payload=None, method="GET", timeout=60):
if payload is None:
req = urllib.request.Request(BASE + path)
else:
body = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(BASE + path, data=body, method=method,
headers={"Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
return json.loads(r.read().decode("utf-8"))
except urllib.error.HTTPError as e:
return {"_http_error": e.code, "_body": e.read().decode("utf-8", "ignore")[:200]}
except Exception as e:
return {"_exc": type(e).__name__, "_msg": str(e)}
def api_retry(path, payload=None, method="GET", timeout=60, tries=3):
for i in range(tries):
r = api(path, payload, method, timeout)
if "_http_error" not in r and "_exc" not in r:
return r
if i < tries - 1:
time.sleep(2)
return r
def scrape_page_names():
"""Return the list of rule names on the current page (top to bottom)."""
r = api_retry("/eval", {"code": "(function(){var rows=document.querySelectorAll('.ww_table tbody tr');return Array.from(rows).map(function(x){return x.cells[0].innerText;}).filter(function(s){return s && s.trim();});})()"}, "POST")
return r.get("result") or []
def click_next_page():
"""Click the right-arrow pagination control. Return True on success."""
r = api_retry("/eval", {"code": "(function(){var a=document.querySelector('img.ww_commonImg_PageNavArrowRightNormal');if(!a)return 'no arrow';a.click();return 'clicked';})()"}, "POST")
return r.get("result") == "clicked"
def goto_first_page():
"""Click the left-arrow enough times to return to page 1."""
for _ in range(20):
r = api_retry("/eval", {"code": "(function(){var a=document.querySelector('img.ww_commonImg_PageNavArrowLeftNormal');return a?'yes':'no';})()"}, "POST")
if r.get("result") != "yes":
return
api_retry("/click", {"selector": "img.ww_commonImg_PageNavArrowLeftNormal"}, "POST")
time.sleep(1.2)
def main():
if len(sys.argv) < 2:
print("usage: python verify_rules.py <expected.json>", file=sys.stderr)
sys.exit(2)
expected = json.load(open(sys.argv[1], encoding="utf-8"))
exp_names = [r["name"] for r in (expected["rules"] if isinstance(expected, dict) and "rules" in expected else expected)]
exp_set = set(exp_names)
api_retry("/launch", {"headless": False}, "POST")
api_retry("/goto", {"url": LIST_URL}, "POST")
time.sleep(2.5)
r = api_retry("/status", None, "GET")
if "loginpage_wx" in (r.get("url") or ""):
print("ABORT: WeCom session expired, re-scan QR and retry.", file=sys.stderr)
sys.exit(1)
goto_first_page()
time.sleep(1.5)
scraped, page = [], 1
while True:
names = scrape_page_names()
if not names:
break
scraped += names
if not click_next_page():
break
time.sleep(1.5)
page += 1
if page > 50: # safety
print("WARN: 50 pages reached, aborting", file=sys.stderr)
break
live_set = set(scraped)
missing = exp_set - live_set
extra = live_set - exp_set
dups = [n for n in scraped if scraped.count(n) > 1]
print(f"scraped {len(scraped)} names across {page} pages")
print(f"expected {len(exp_names)} names")
if not (missing or extra or dups):
print("OK: counts and sets match.")
return
print("FAIL:")
if missing: print(f" missing ({len(missing)}): {sorted(missing)}")
if extra: print(f" extra ({len(extra)}): {sorted(extra)}")
if dups: print(f" duplicates ({len(dups)}): {sorted(set(dups))}")
sys.exit(1)
if __name__ == "__main__":
main()
🔧
cf_helpers.sh
Cloudflare API 辅助(验证令牌/账户/建项目)
SH66 行 · 2.3 KB
#!/usr/bin/env bash
# Cloudflare Pages API helpers. Use these to verify the token, get the
# account id, and create a Pages project. All four functions are safe to
# re-run: they print the existing state instead of failing if a call would
# be a duplicate.
#
# Usage:
# source cf_helpers.sh
# cf_verify_token "$CF_API_TOKEN"
# ACCOUNT_ID=$(cf_get_account "$CF_API_TOKEN")
# cf_create_project "$CF_API_TOKEN" "$ACCOUNT_ID" "zx-knowledge-base"
#
# All curl calls use --fail-with-body so non-2xx responses print the
# error JSON to stderr and the function exits 1. That makes CI-style
# pipelines able to detect a bad token immediately.
set -euo pipefail
API_BASE="https://api.cloudflare.com/client/v4"
cf_verify_token() {
local token="$1"
curl --fail-with-body -sS \
-H "Authorization: Bearer $token" \
-H "Content-Type: application/json" \
"$API_BASE/user/tokens/verify"
echo
}
# Print the first account id (the user's primary account).
cf_get_account() {
local token="$1"
curl --fail-with-body -sS \
-H "Authorization: Bearer $token" \
-H "Content-Type: application/json" \
"$API_BASE/accounts" | python -c "import sys, json; d=json.load(sys.stdin); print(d['result'][0]['id'])"
}
# Create a Pages project. If it already exists, prints its current state
# instead of failing. The project name must be lowercase-hyphenated.
cf_create_project() {
local token="$1"
local account_id="$2"
local project_name="$3"
local body
body=$(printf '{"name":"%s","production_branch":"main"}' "$project_name")
local resp
resp=$(curl --fail-with-body -sS -X POST \
-H "Authorization: Bearer $token" \
-H "Content-Type: application/json" \
-d "$body" \
"$API_BASE/accounts/$account_id/pages/projects" 2>&1) || {
# 409 means the project already exists -- look it up.
if echo "$resp" | grep -q '"code": 8000003\|already exists'; then
curl --fail-with-body -sS \
-H "Authorization: Bearer $token" \
"$API_BASE/accounts/$account_id/pages/projects" \
| python -c "import sys, json; d=json.load(sys.stdin); [print(p['name'], p['subdomain']) for p in d['result'] if p['name']=='$project_name']"
return 0
fi
echo "$resp" >&2
return 1
}
echo "$resp" | python -c "import sys, json; d=json.load(sys.stdin); print(d['result']['name'], d['result']['subdomain'])"
}
🚀
cf_deploy.sh
wrangler pages deploy 部署脚本
SH59 行 · 2.6 KB
#!/usr/bin/env bash
# Deploy a local directory of static files to a Cloudflare Pages project.
#
# Usage:
# cf_deploy.sh <deploy_dir> <project_name> <cf_api_token> [cf_account_id]
#
# Steps:
# 1. (If <cf_account_id> not given) look it up via cf_helpers.sh.
# 2. Install wrangler into the managed node workspace with --ignore-scripts
# (the esbuild postinstall is blocked by the local sandbox).
# 3. Run `wrangler pages deploy <deploy_dir> --project-name=<project>
# --branch=main --commit-dirty=true`.
#
# Output: prints wrangler's deployment summary. The production URL is
# `https://<project>.pages.dev` and a per-deploy preview is also printed.
#
# Requirements:
# - $NODE_WORKSPACE = the absolute path to the managed Node workspace
# (default: $WORKBUDDY_NODE_WORKSPACE or
# "C:/Users/向/.workbuddy/binaries/node/workspace").
# - $NODE_BIN = the absolute path to the managed Node binary
# (default: "C:/Users/向/.workbuddy/binaries/node/versions/22.22.2-6/node.exe").
set -euo pipefail
DEPLOY_DIR="${1:?usage: cf_deploy.sh <deploy_dir> <project> <token> [account_id]}"
PROJECT="${2:?missing project name}"
TOKEN="${3:?missing CF token}"
ACCOUNT_ID="${4:-}"
NODE_WORKSPACE="${NODE_WORKSPACE:-${WORKBUDDY_NODE_WORKSPACE:-C:/Users/向/.workbuddy/binaries/node/workspace}}"
NODE_BIN="${NODE_BIN:-C:/Users/向/.workbuddy/binaries/node/versions/22.22.2-6/node.exe}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=cf_helpers.sh
source "$SCRIPT_DIR/cf_helpers.sh"
if [[ -z "$ACCOUNT_ID" ]]; then
ACCOUNT_ID=$(cf_get_account "$TOKEN")
fi
echo ">> account_id=$ACCOUNT_ID project=$PROJECT dir=$DEPLOY_DIR"
# Install wrangler if it's not already in the workspace. The
# --ignore-scripts flag is mandatory: esbuild's postinstall is blocked
# by the local sandbox and would otherwise abort the install.
if [[ ! -f "$NODE_WORKSPACE/node_modules/wrangler/bin/wrangler.js" ]]; then
echo ">> installing wrangler (this takes ~30s)"
(cd "$NODE_WORKSPACE" && npm install wrangler --ignore-scripts --no-fund --no-audit)
fi
# Deploy. We invoke the .js entry directly with the managed node binary
# to avoid the npm shim which sometimes tries to fetch the latest wrangler.
CLOUDFLARE_API_TOKEN="$TOKEN" CLOUDFLARE_ACCOUNT_ID="$ACCOUNT_ID" \
"$NODE_BIN" "$NODE_WORKSPACE/node_modules/wrangler/bin/wrangler.js" \
pages deploy "$DEPLOY_DIR" --project-name="$PROJECT" --branch=main --commit-dirty=true
echo
echo ">> Production URL: https://$PROJECT.pages.dev"
echo ">> Verify with: curl -s -o /dev/null -w '%{http_code}\n' https://$PROJECT.pages.dev/"