Back to Discover

scrape-le

connector

nolindnaidoo

Analyse robots.txt content and report whether a path may be crawled.

View on GitHub
0 starsSynced Aug 5, 2026

Install to Claude Code

/plugin marketplace add nolindnaidoo/scrape-le

README

Scrape-LE Logo

Scrape-LE: Zero Hassle Scrapeability Checks

Load a URL in headless Chromium and see what will block your scraper — before you write it
Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots

Install from VS Code Marketplace Open VSX downloads scrape-le-mcp on npm LE Tools


Scrapeability Check Demo

Useful? A star or rating is how other developers find it — ★ GitHub · ★ Marketplace · ★ Open VSX

What it does

Run Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Code–based editors like Cursor and VSCodium (installable from Open VSX).

One-time setup: run Scrape-LE: Setup Browser to install Chromium (~130MB, into Playwright's browser cache).

Use it from an AI agent

The same engine runs as an MCP server, so an agent can call it directly instead of you running a command.

EditorHow
VS Code 1.101+Nothing to install — the extension registers analyze_robots_txt with agent mode
ZedScrape-LEpending review
Claude Codeclaude mcp add scrape-le -- npx -y scrape-le-mcp
Cursor, Windsurf, anything elsepoint it at npx scrape-le-mcp
analyze_robots_txt(content, path, maxResults?)

Given robots.txt contents and a path, reports whether the generic (User-agent: *) rules permit crawling it, plus the crawl delay, disallowed patterns and any sitemaps.

The server takes content and returns data — it reads no files and makes no network requests of its own. Published as scrape-le-mcp on npm and as io.github.nolindnaidoo/scrape-le in the MCP registry.

Configuring it by hand — any host with an MCP config file

Most hosts read a JSON config. Add one entry:

{
  "mcpServers": {
    "scrape-le": {
      "command": "npx",
      "args": ["-y", "scrape-le-mcp"]
    }
  }
}

-y skips the install prompt on first run. Pin a version if you would rather not track releases — scrape-le-mcp@2.2.1.

Prefer not to go through npx on every launch? Install it once and point at the binary instead:

npm install -g scrape-le-mcp
{
  "mcpServers": {
    "scrape-le": { "command": "scrape-le-mcp" }
  }
}

It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:

echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | npx -y scrape-le-mcp

That prints the tool list and exits — if you see analyze_robots_txt, the server works.

Detections

DetectionHow it works
Anti-bot vendorsResponse headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX
Rate limitingX-RateLimit-* / RateLimit-* / Retry-After response headers, plus HTTP 429
robots.txtFetches <origin>/robots.txt and evaluates the User-agent: * rules against your URL with RFC 9309 semantics — grouped agents, Allow/Disallow longest-match, * wildcards, $ anchors, crawl-delay, sitemaps
AuthenticationHTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL

Honest limitations: signatures are best-effort fingerprints of public integration patterns — a detected widget means the page can challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the * rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.

Commands

CommandDescription
Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S)Prompt for a URL and run the full check
Scrape-LE: Check Selected URLRun the check on the URL in the current selection (also in the right-click menu)
Scrape-LE: Setup BrowserInstall or verify the Chromium browser
Scrape-LE: Open SettingsOpen Scrape-LE settings
Scrape-LE: Help & TroubleshootingBuilt-in documentation

Settings

SettingDefaultDescription
scrape-le.browser.timeout30000Page-load timeout in ms (5000–120000)
scrape-le.browser.viewport.width1280Viewport width
scrape-le.browser.viewport.height720Viewport height
scrape-le.browser.userAgent""Custom User-Agent (empty = Chromium default)
scrape-le.retry.userAgentsfalseOn a blocked or failed check, retry under common User-Agents and report which worked
scrape-le.screenshot.enabledtrueSave a full-page screenshot per check
scrape-le.screenshot.path.vscode/scrape-leScreenshot directory (workspace-relative or absolute)
scrape-le.screenshot.formatpngpng or jpeg
scrape-le.screenshot.quality90JPEG quality 0–100 (ignored for png)
scrape-le.checkConsoleErrorstrueCapture console and page errors while loading
scrape-le.detections.antiBottrueAnti-bot vendor detection
scrape-le.detections.rateLimittrueRate-limit detection
scrape-le.detections.robotsTxttruerobots.txt fetch + evaluation
scrape-le.detections.authenticationtrueAuthentication-wall detection
scrape-le.notificationsLevelimportantall = every notification, important = warnings + errors, silent = errors only
scrape-le.statusBar.enabledtrueShow the status bar item

Languages

Twelve languages besides English:

German · Spanish · French · Indonesian · Italian · Japanese · Korean · Portuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)

Both halves are covered — the manifest (command titles, setting names and descriptions) and everything shown while the extension runs (notifications, the status bar, quick-picks and prompts). The extension follows VS Code's display language, so it matches whatever the editor is already set to; no setting of its own.

Privacy & security

  • Network access is the feature, and it is scoped. A check talks to exactly two things: the URL you enter (loaded in headless Chromium, which fetches that page's own resources like any browser) and that origin's /robots.txt. Nothing is sent anywhere else — no telemetry, no analytics.
  • Screenshots stay local, written to the configured path inside your workspace.
  • The MCP server makes no network request at all — unlike the extension, deliberately. fetchRobotsTxt builds a URL from an arbitrary origin, which inside an agent loop is an SSRF primitive: the caller supplying the URL is the model, not you. The server analyses robots.txt content you already fetched, and a test asserts no tool accepts a url argument.
  • Error notifications redact home directories and credential-shaped fragments.
  • Respect the sites you check: a scrapeability report is information, not permission.

Development

bun install
bun run build            # esbuild bundle -> dist/extension.js
bun run typecheck        # tsc --noEmit (includes tests)
bun run test             # vitest unit suite
bun run test:integration # real VS Code extension host
bun run lint             # biome
bun run package          # VSIX into release/

Architecture and conventions live in AGENTS.md. Changes are tracked in CHANGELOG.md.

Performance

InputSizeFoundTimeRateScan speed
Header signature scan2.83 MB20,0005.19 ms3,852,946/sec544.6 MB/s
robots.txt path match3.32 MB60,0009.64 ms6,223,689/sec344.3 MB/s

Median of 7 runs after warmup, on Apple M5 Pro, 24 GB RAM, Node 24.3.0. Inputs are generated by scripts/benchmark.ts rather than checked in, so the sizes above are exactly what was measured. Reproduce with bun run benchmark.

These are machine-specific and are not asserted in CI — a benchmark that gates a build only tells you how busy the runner was.

Testing

MetricCoverage
Statements92.96%
Branches83.12%
Functions95.13%
Lines94.40%

305 test cases across 28 files, plus an integration suite that runs in a real VS Code extension host and an end-to-end test that installs the built .vsix into a clean profile.

Generated from coverage/coverage-summary.json by scripts/coverage-readme.js; CI fails if this section drifts from a fresh run. Reproduce with bun run test:coverage.

More from the LE Family

Every tool in the family, one page: letools.dev

All ten also ship as MCP servers — npx <name>-mcp gives any agent the same engine.

  • Paths-LE - Extract file paths from JS/TS imports, JSON, HTML, CSS, TOML, CSV, and .env
  • String-LE - Extract string values for i18n from JSON, YAML, CSV, TOML, INI, and .env
  • Numbers-LE - Extract numeric values from JSON, YAML, CSV, TOML, INI, and .env
  • EnvSync-LE - Spot missing keys across your .env files, with a markdown report
  • Regex-LE - Find, test, and validate regular expressions with ReDoS screening
  • Secrets-LE - Detect and sanitize credentials locally, before you commit
  • Colors-LE - Extract and analyze colors from CSS, SCSS, LESS, Stylus, HTML, JS/TS, and SVG
  • URLs-LE - Extract URLs from documentation, configs, and code
  • Dates-LE - Extract and analyze dates from logs, configs, and code

Also by nolindnaidoo

Rust

Contact DeveloperGitHub · LinkedIn

License

MIT © nolindnaidoo

Rendered live from nolindnaidoo/scrape-le's GitHub README — not stored, always reflects the source repo.

1 Install Method

NameDescriptionCategorySource
npm packageInstall via npm (stdio transport)mcp-serverscrape-le-mcp

0 Comments

Login required
Log in to post a comment or update on this repo.

No comments yet — be the first to share an update.