webextrator

v2026.09.24

Render and extract web page content via AceDataCloud's WebExtrator API. Use when scraping a page's final rendered HTML, or extracting typed structured data (Article, Product, Recipe, Video, Discussion, Job) plus clean markdown/text from any URL. Real headless Chromium with schema.org + LLM extraction.

GitHub
安装命令
npx skhub add acedatacloud/webextrator
Markdown
SKILL.md

WebExtrator Web Render & Extract

Render and extract web content through AceDataCloud's WebExtrator API — real headless Chromium plus a three-tier extraction pipeline (schema.org JSON-LD mapper → LLM typed extractor → Readability/markdown fallback).

Setup: See authentication for token setup.

Quick Start

curl -X POST https://api.acedata.cloud/webextrator/extract \
  -H "Authorization: Bearer $ACEDATACLOUD_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "expected_type": "general"}'

Returns synchronously in seconds — no task polling needed.

Endpoints

PathPurpose
POST /webextrator/renderHeadless Chromium render → raw HTML + clean text + title
POST /webextrator/extractRender + structured extraction (schema.org + LLM types) + markdown
POST /webextrator/tasksLook up historical render/extract task envelopes (7-day retention, free)

Workflows

1. Extract typed content

POST /webextrator/extract
{
  "url": "https://example.com",
  "expected_type": "general"
}

Real response (trimmed):

{
  "success": true,
  "task_id": "604b1cfb-6c5a-42c9-b900-a281e1b9c3c5",
  "trace_id": "f2a7c0b0-c17c-4bc9-b6e7-9c59746dd366",
  "elapsed": 0.003,
  "data": {
    "kind": "extract",
    "url": "https://example.com",
    "finalUrl": "https://example.com/",
    "contentType": "general",
    "title": "Example Domain",
    "description": "This domain is for use in documentation examples without needing permission. Avoid use in operations.",
    "language": "en",
    "images": [],
    "links": [],
    "markdown": "...",
    "text": "...",
    "structured": { "schemaOrg": {}, "openGraph": {}, "jsonLd": [] }
  }
}

2. Render raw HTML

POST /webextrator/render
{
  "url": "https://example.com",
  "wait_until": "networkidle",
  "block_resources": ["image", "media", "font"]
}

Returns data.html, data.text, data.title, data.status, data.finalUrl.

3. Look up a task

POST /webextrator/tasks
{
  "action": "retrieve",
  "id": "604b1cfb-6c5a-42c9-b900-a281e1b9c3c5"
}

Parameters

Render & Extract (shared)

ParameterRequiredDescription
urlYesPage URL to render (http(s)://)
wait_untilNoload / domcontentloaded / networkidle / commit (default networkidle)
timeoutNoNavigation timeout in seconds (default 30)
delayNoExtra wait in seconds after wait_until (for SPAs)
wait_for_selectorNoCSS selector to wait for before ready
block_resourcesNoDrop image/font/media/stylesheet/xhr/fetch
headersNoExtra request headers for the target site
user_agentNoOverride the browser User-Agent
asyncNotrue submits without blocking; poll /webextrator/tasks for the result
callback_urlNoPosted the final envelope when running asynchronously

Extract-only

ParameterRequiredDescription
expected_typeNoproduct / article / general — skips the heuristic
enable_llmNoAllow LLM extractor when schema.org found nothing (default false)

Tasks

ParameterRequiredDescription
actionYesretrieve (single) or retrieve_batch (many)
id / trace_idone ofFor retrieve
ids / trace_idsone ofFor retrieve_batch

Gotchas

  • Parameters use snake_case (wait_until, block_resources), not camelCase
  • Cache hits are still billed; identical URLs return in ~0.003s
  • expected_type only allows product/article/general — typed kinds (recipe/video/job) are detected automatically from schema.org
  • enable_llm has no effect when the page ships schema.org JSON-LD — the deterministic mapper wins for free
  • Tasks API is free and retains records for 7 days only

MCP: pip install mcp-webextrator | Hosted: https://webextrator.mcp.acedata.cloud/mcp | See all MCP servers

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

NOASSERTION

源路径

skills/webextrator

默认分支

main

最新提交

57cc298

Tree SHA

acaa402