scrapingant-web-fetch

v2026.09.24

Fetch live web pages for agents through ScrapingAnt's hosted MCP server (`https://api.scrapingant.com/mcp`) with headless-Chrome rendering, rotating datacenter/residential proxies, Cloudflare and anti-bot handling, and LLM-ready Markdown output. Use when a plain fetch/WebFetch is blocked (403/429, Cloudflare challenge), when a JavaScript/SPA page returns an empty shell, when geo-specific content is needed, or when standing up a local browser scraper costs more than the task is worth. Triggers on: scrapingant, MCP web scraping, fetch blocked page, Cloudflare bypass, anti-bot scraping, JS rendered page, scrape to markdown, residential proxy fetch, geo-targeted scrape, live web access for agents.

GitHub
Install command
npx skhub add akillness/scrapingant-web-fetch
Markdown
SKILL.md

ScrapingAnt Web Fetch — hosted MCP for live, unblocked web content

ScrapingAnt exposes a hosted MCP server at https://api.scrapingant.com/mcp. An agent that registers it gets three fetch tools backed by headless Chrome and a rotating proxy pool, so blocked or JavaScript-rendered pages come back as clean Markdown instead of a challenge page. Nothing runs locally: no browser binary, no Python environment, no MCP process to supervise.

Sponsor. ScrapingAnt is a partner of jeo-skills. Signing up through scrapingant.com?ref=ztewzmv&tm_source=readme supports this repository at no extra cost to you. The free tier (10,000 credits/month as of signup, no credit card) is enough to evaluate every workflow below.

When to use this skill

  • A normal fetch/WebFetch returns 403/429, a Cloudflare interstitial, or a bot-check page instead of content
  • The target is a SPA (React/Next.js docs, dashboards) whose raw HTML is an empty shell until JavaScript runs
  • You need page content as Markdown for RAG, summarization, or doc reference without writing selectors
  • Content is geo-restricted and must be fetched from a specific country
  • A one-off or low-volume scrape does not justify installing Playwright, Scrapling, or a browser image in CI
  • The agent runtime speaks MCP (Claude Code, Cursor, Windsurf, Cline, VS Code Copilot, Claude Desktop) and you want a fetch tool available in-conversation

When not to use this skill

  • The page is public, static, and unprotected — a plain curl/WebFetch costs zero credits and is faster
  • You need a full crawl, link frontier, or selector-drift healing across many pages — use scrapling (local Python, spiders) instead
  • The target is X/Twitter — x-twitter-scraper handles that platform's specifics
  • The work needs an authenticated session, form filling, or multi-step browser interaction — MCP fetch tools take a URL, not a script; drive a real browser
  • Scraping the target would violate its Terms of Service, robots policy, or applicable law — decline instead of routing around the block

Instructions

Step 1 — Get an API key

  1. Sign up at scrapingant.com?ref=ztewzmv&tm_source=readme (free tier, no card) and copy the key from the dashboard.
  2. Export it in the shell profile — never commit it, never echo it, never paste it into a repo file:
export SCRAPINGANT_API_KEY="<your-key>"
  1. Confirm the environment is ready (read-only, no network writes):
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh doctor

If the key is missing the skill stops here and prints the signup link — do not fall back to fabricated page content.

Step 2 — Register the MCP server

Claude Code (CLI) — one command:

bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh install claude-code

which runs the vendor-documented registration:

claude mcp add scrapingant --transport http https://api.scrapingant.com/mcp \
  -H "x-api-key: $SCRAPINGANT_API_KEY"

Every other documented client uses the same streamable-HTTP block:

{
  "mcpServers": {
    "scrapingant": {
      "url": "https://api.scrapingant.com/mcp",
      "transport": "streamableHttp",
      "headers": {
        "x-api-key": "${SCRAPINGANT_API_KEY}"
      }
    }
  }
}

VS Code / GitHub Copilot is the one exception — it uses servers, requestInit.headers, and a trailing slash on the URL. Per-client file paths and snippets: references/mcp-clients.md, or print one with scrapingant.sh install <client>.

Step 3 — Pick the right tool

MCP toolReturnsUse it forDefault?
get_web_page_markdownLLM-ready MarkdownRAG, summarizing, reading docs✅ default
get_web_page_htmlRaw HTMLselector-based post-processing, DOM checkson request
get_web_page_textPlain textcheapest token footprint, text-only checkson request

Default to Markdown. Only reach for HTML when something downstream actually parses the DOM — raw HTML burns far more context for the same page.

Step 4 — Tune parameters for cost and success

All three tools take the same arguments:

ParameterTypeDefaultNotes
urlstring—required
browserbooleantruefalse = no JS rendering, much cheaper
proxy_typestringdatacenterresidential only after a datacenter block
proxy_countrystringrandomISO-3166 code, e.g. DE, KR

Credit cost is driven by those choices (verified against docs.scrapingant.com/credits-cost):

Request shapeCredits
No browser + datacenter proxy1
Headless browser with JS rendering + datacenter proxy10
No browser + residential proxy25
Headless browser with JS rendering + residential proxy125

So 10,000 free credits ≈ 10,000 static fetches, ≈ 1,000 JS-rendered fetches, or 80 residential+JS fetches. Escalate, never start at the top: try browser=false first, add browser=true when the body is empty, and switch to proxy_type=residential only when a datacenter attempt is actually blocked.

Step 5 — Verify before reporting

# remaining credits on the key (GET /v2/usage)
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh credits

# end-to-end smoke test against the REST twin of the MCP tools
bash .agent-skills/scrapingant-web-fetch/scripts/scrapingant.sh probe https://example.com

probe uses the REST endpoint (/v2/markdown) so a key can be validated without an MCP client attached. It reports the credit shape it used.

Examples

Once the server is registered, drive it in plain language:

Fetch https://example.com with scrapingant and summarize it.
Get https://docs.python.org/3/tutorial/index.html as markdown, then list the main topics.
This page 403s for me — refetch it through scrapingant with residential proxies.
Fetch https://example.com through a German proxy and compare it with the US version.

Cheap-first escalation inside one task:

1. get_web_page_markdown(url, browser=false)          → 1 credit
2. body empty/JS-only? retry with browser=true         → 10 credits
3. still 403/Cloudflare? retry proxy_type=residential  → 125 credits, last resort

Shell equivalents for CI or a non-MCP runtime:

scripts/scrapingant.sh probe https://example.com --no-browser              # 1 credit
scripts/scrapingant.sh probe https://spa.example.com                        # 10 credits
scripts/scrapingant.sh probe https://blocked.example.com --proxy residential --country DE

Best practices

  • Try free first. Plain WebFetch/curl costs nothing; route to ScrapingAnt when it actually fails, not by default.
  • Markdown by default. get_web_page_markdown keeps the token footprint small; get_web_page_html is opt-in for DOM work.
  • Escalate one axis at a time (browser → residential → country) and record which shape worked so the next run starts there.
  • Never hardcode the key. It lives in SCRAPINGANT_API_KEY; scripts mask it in output, and MCP config files should reference the env var where the client supports interpolation.
  • Watch the budget. Run scrapingant.sh credits before a batch; the free tier does not roll over between months.
  • Respect the target. Honor robots/ToS and rate limits; anti-bot bypass is for legitimate access, not for evading a site's explicit refusal.
  • Re-verify the vendor surface (tools, parameters, credit table) before editing this skill — see the sourced links below.

Troubleshooting

SymptomCauseFix
SCRAPINGANT_API_KEY is not setkey not exportedStep 1; restart the client after exporting
403 / challenge page still returneddatacenter proxy blockedproxy_type=residential, then a specific proxy_country
Empty or shell-only contentJS-rendered page fetched with browser=falseretry with browser=true
Tools missing in the clientserver not registered or client not restartedrerun Step 2, restart the client, re-check claude mcp list
credits reports 0 remainingmonthly free tier exhaustedwait for renewal or upgrade; credits do not roll over

References

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Not specified

Source path

.agent-skills/scrapingant-web-fetch

Default branch

main

Latest commit

f579bfe

Tree SHA

34a09b3