scrapingant-web-fetch

v2026.09.24

Fetch live web pages for agents through ScrapingAnt's hosted MCP server (`https://api.scrapingant.com/mcp`) with headless-Chrome rendering, rotating datacenter/residential proxies, Cloudflare and anti-bot handling, and LLM-ready Markdown output. Use when a plain fetch/WebFetch is blocked (403/429, Cloudflare challenge), when a JavaScript/SPA page returns an empty shell, when geo-specific content is needed, or when standing up a local browser scraper costs more than the task is worth. Triggers on: scrapingant, MCP web scraping, fetch blocked page, Cloudflare bypass, anti-bot scraping, JS rendered page, scrape to markdown, residential proxy fetch, geo-targeted scrape, live web access for agents.

GitHub
安装命令
npx skhub add akillness/scrapingant-web-fetch
Markdown
SKILL.md

ScrapingAnt Web Fetch — hosted MCP for live, unblocked web content

ScrapingAnt exposes a hosted MCP server at https://api.scrapingant.com/mcp. An agent that registers it gets three fetch tools backed by headless Chrome and a rotating proxy pool, so blocked or JavaScript-rendered pages come back as clean Markdown instead of a challenge page. Nothing runs locally: no browser binary, no Python environment, no MCP process to supervise.

Sponsor. ScrapingAnt is a partner of jeo-skills. Signing up through scrapingant.com?ref=ztewzmv&tm_source=readme supports this repository at no extra cost to you. The free tier (10,000 credits/month as of signup, no credit card) is enough to evaluate every workflow below.

When to use this skill

  • A normal fetch/WebFetch returns 403/429, a Cloudflare interstitial, or a bot-check page instead of content
  • The target is a SPA (React/Next.js docs, dashboards) whose raw HTML is an empty shell until JavaScript runs
  • You need page content as Markdown for RAG, summarization, or doc reference without writing selectors
  • Content is geo-restricted and must be fetched from a specific country
  • A one-off or low-volume scrape does not justify installing Playwright, Scrapling, or a browser image in CI
  • The agent runtime speaks MCP (Claude Code, Cursor, Windsurf, Cline, VS Code Copilot, Claude Desktop) and you want a fetch tool available in-conversation

When not to use this skill

  • The page is public, static, and unprotected — a plain curl/WebFetch costs zero credits and is faster
  • You need a full crawl, link frontier, or selector-drift healing across many pages — use scrapling (local Python, spiders) instead
  • The target is X/Twitter — x-twitter-scraper handles that platform's specifics
  • The work needs an authenticated session, form filling, or multi-step browser interaction — MCP fetch tools take a URL, not a script; drive a real browser
  • Scraping the target would violate its Terms of Service, robots policy, or applicable law — decline instead of routing around the block

Instructions

Step 1 — Get an API key

  1. Sign up at scrapingant.com?ref=ztewzmv&tm_source=readme (free tier, no card) and copy the key from the dashboard.
  2. Export it in the shell profile — never commit it, never echo it, never paste it into a repo file:
export SCRAPINGANT_API_KEY="<your-key>"
  1. Confirm the environment is ready (read-only, no network writes):
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh doctor

If the key is missing the skill stops here and prints the signup link — do not fall back to fabricated page content.

Step 2 — Register the MCP server

Claude Code (CLI) — one command:

bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh install claude-code

which runs the vendor-documented registration:

claude mcp add scrapingant --transport http https://api.scrapingant.com/mcp \
  -H "x-api-key: $SCRAPINGANT_API_KEY"

Every other documented client uses the same streamable-HTTP block:

{
  "mcpServers": {
    "scrapingant": {
      "url": "https://api.scrapingant.com/mcp",
      "transport": "streamableHttp",
      "headers": {
        "x-api-key": "${SCRAPINGANT_API_KEY}"
      }
    }
  }
}

VS Code / GitHub Copilot is the one exception — it uses servers, requestInit.headers, and a trailing slash on the URL. Per-client file paths and snippets: references/mcp-clients.md, or print one with scrapingant.sh install <client>.

Step 3 — Pick the right tool

MCP toolReturnsUse it forDefault?
get_web_page_markdownLLM-ready MarkdownRAG, summarizing, reading docs✅ default
get_web_page_htmlRaw HTMLselector-based post-processing, DOM checkson request
get_web_page_textPlain textcheapest token footprint, text-only checkson request

Default to Markdown. Only reach for HTML when something downstream actually parses the DOM — raw HTML burns far more context for the same page.

Step 4 — Tune parameters for cost and success

All three tools take the same arguments:

ParameterTypeDefaultNotes
urlstring—required
browserbooleantruefalse = no JS rendering, much cheaper
proxy_typestringdatacenterresidential only after a datacenter block
proxy_countrystringrandomISO-3166 code, e.g. DE, KR

Credit cost is driven by those choices (verified against docs.scrapingant.com/credits-cost):

Request shapeCredits
No browser + datacenter proxy1
Headless browser with JS rendering + datacenter proxy10
No browser + residential proxy25
Headless browser with JS rendering + residential proxy125

So 10,000 free credits ≈ 10,000 static fetches, ≈ 1,000 JS-rendered fetches, or 80 residential+JS fetches. Escalate, never start at the top: try browser=false first, add browser=true when the body is empty, and switch to proxy_type=residential only when a datacenter attempt is actually blocked.

Step 5 — Verify before reporting

# remaining credits on the key (GET /v2/usage)
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh credits

# end-to-end smoke test against the REST twin of the MCP tools
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh probe https://example.com

probe uses the REST endpoint (/v2/markdown) so a key can be validated without an MCP client attached. It reports the credit shape it used.

Examples

Once the server is registered, drive it in plain language:

Fetch https://example.com with scrapingant and summarize it.
Get https://docs.python.org/3/tutorial/index.html as markdown, then list the main topics.
This page 403s for me — refetch it through scrapingant with residential proxies.
Fetch https://example.com through a German proxy and compare it with the US version.

Cheap-first escalation inside one task:

1. get_web_page_markdown(url, browser=false)          → 1 credit
2. body empty/JS-only? retry with browser=true         → 10 credits
3. still 403/Cloudflare? retry proxy_type=residential  → 125 credits, last resort

Shell equivalents for CI or a non-MCP runtime:

scripts/scrapingant.sh probe https://example.com --no-browser              # 1 credit
scripts/scrapingant.sh probe https://spa.example.com                        # 10 credits
scripts/scrapingant.sh probe https://blocked.example.com --proxy residential --country DE

Best practices

  • Try free first. Plain WebFetch/curl costs nothing; route to ScrapingAnt when it actually fails, not by default.
  • Markdown by default. get_web_page_markdown keeps the token footprint small; get_web_page_html is opt-in for DOM work.
  • Escalate one axis at a time (browser → residential → country) and record which shape worked so the next run starts there.
  • Never hardcode the key. It lives in SCRAPINGANT_API_KEY; scripts mask it in output, and MCP config files should reference the env var where the client supports interpolation.
  • Watch the budget. Run scrapingant.sh credits before a batch; the free tier does not roll over between months.
  • Respect the target. Honor robots/ToS and rate limits; anti-bot bypass is for legitimate access, not for evading a site's explicit refusal.
  • Re-verify the vendor surface (tools, parameters, credit table) before editing this skill — see the sourced links below.

Troubleshooting

SymptomCauseFix
SCRAPINGANT_API_KEY is not setkey not exportedStep 1; restart the client after exporting
403 / challenge page still returneddatacenter proxy blockedproxy_type=residential, then a specific proxy_country
Empty or shell-only contentJS-rendered page fetched with browser=falseretry with browser=true
Tools missing in the clientserver not registered or client not restartedrerun Step 2, restart the client, re-check claude mcp list
credits reports 0 remainingmonthly free tier exhaustedwait for renewal or upgrade; credits do not roll over

References

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

未指定

源路径

.agent-skills/scrapingant-web-fetch

默认分支

main

最新提交

f579bfe

Tree SHA

34a09b3