Scrape-LE: Zero Hassle Scrapeability Checks
Load a URL in headless Chromium and see what will block your scraper — before you write it
Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots
Useful? A star or rating is how other developers find it — ★ GitHub · ★ Marketplace · ★ Open VSX
What it does
Run Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Code–based editors like Cursor and VSCodium (installable from Open VSX).
One-time setup: run Scrape-LE: Setup Browser to install Chromium (~130MB, into Playwright's browser cache).
Use it from an AI agent
The same engine runs as an MCP server, so an agent can call it directly instead of you running a command.
| Editor | How |
|---|---|
| VS Code 1.101+ | Nothing to install — the extension registers analyze_robots_txt with agent mode |
| Zed | Scrape-LE — pending review |
| Claude Code | claude mcp add scrape-le -- npx -y scrape-le-mcp |
| Cursor, Windsurf, anything else | point it at npx scrape-le-mcp |
analyze_robots_txt(content, path, maxResults?)
Given robots.txt contents and a path, reports whether the generic (User-agent: *) rules permit crawling it, plus the crawl delay, disallowed patterns and any sitemaps.
The server takes content and returns data — it reads no files and makes no network requests of its own. Published as scrape-le-mcp on npm and as io.github.nolindnaidoo/scrape-le in the MCP registry.
Configuring it by hand — any host with an MCP config file
Most hosts read a JSON config. Add one entry:
{
"mcpServers": {
"scrape-le": {
"command": "npx",
"args": ["-y", "scrape-le-mcp"]
}
}
}
-y skips the install prompt on first run. Pin a version if you would rather not track releases — scrape-le-mcp@2.2.1.
Prefer not to go through npx on every launch? Install it once and point at the binary instead:
npm install -g scrape-le-mcp
{
"mcpServers": {
"scrape-le": { "command": "scrape-le-mcp" }
}
}
It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | npx -y scrape-le-mcp
That prints the tool list and exits — if you see analyze_robots_txt, the server works.
Detections
| Detection | How it works |
|---|---|
| Anti-bot vendors | Response headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX |
| Rate limiting | X-RateLimit-* / RateLimit-* / Retry-After response headers, plus HTTP 429 |
| robots.txt | Fetches <origin>/robots.txt and evaluates the User-agent: * rules against your URL with RFC 9309 semantics — grouped agents, Allow/Disallow longest-match, * wildcards, $ anchors, crawl-delay, sitemaps |
| Authentication | HTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL |
Honest limitations: signatures are best-effort fingerprints of public integration patterns — a detected widget means the page can challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the * rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.
Commands
| Command | Description |
|---|---|
Scrape-LE: Check URL Scrapeability (Ctrl+Alt+S / Cmd+Alt+S) | Prompt for a URL and run the full check |
Scrape-LE: Check Selected URL | Run the check on the URL in the current selection (also in the right-click menu) |
Scrape-LE: Setup Browser | Install or verify the Chromium browser |
Scrape-LE: Open Settings | Open Scrape-LE settings |
Scrape-LE: Help & Troubleshooting | Built-in documentation |
Settings
| Setting | Default | Description |
|---|---|---|
scrape-le.browser.timeout | 30000 | Page-load timeout in ms (5000–120000) |
scrape-le.browser.viewport.width | 1280 | Viewport width |
scrape-le.browser.viewport.height | 720 | Viewport height |
scrape-le.browser.userAgent | "" | Custom User-Agent (empty = Chromium default) |
scrape-le.retry.userAgents | false | On a blocked or failed check, retry under common User-Agents and report which worked |
scrape-le.screenshot.enabled | true | Save a full-page screenshot per check |
scrape-le.screenshot.path | .vscode/scrape-le | Screenshot directory (workspace-relative or absolute) |
scrape-le.screenshot.format | png | png or jpeg |
scrape-le.screenshot.quality | 90 | JPEG quality 0–100 (ignored for png) |
scrape-le.checkConsoleErrors | true | Capture console and page errors while loading |
scrape-le.detections.antiBot | true | Anti-bot vendor detection |
scrape-le.detections.rateLimit | true | Rate-limit detection |
scrape-le.detections.robotsTxt | true | robots.txt fetch + evaluation |
scrape-le.detections.authentication | true | Authentication-wall detection |
scrape-le.notificationsLevel | important | all = every notification, important = warnings + errors, silent = errors only |
scrape-le.statusBar.enabled | true | Show the status bar item |
Languages
Twelve languages besides English:
German · Spanish · French · Indonesian · Italian · Japanese · Korean · Portuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)
Both halves are covered — the manifest (command titles, setting names and descriptions) and everything shown while the extension runs (notifications, the status bar, quick-picks and prompts). The extension follows VS Code's display language, so it matches whatever the editor is already set to; no setting of its own.
Privacy & security
- Network access is the feature, and it is scoped. A check talks to exactly two things: the URL you enter (loaded in headless Chromium, which fetches that page's own resources like any browser) and that origin's
/robots.txt. Nothing is sent anywhere else — no telemetry, no analytics. - Screenshots stay local, written to the configured path inside your workspace.
- The MCP server makes no network request at all — unlike the extension, deliberately.
fetchRobotsTxtbuilds a URL from an arbitrary origin, which inside an agent loop is an SSRF primitive: the caller supplying the URL is the model, not you. The server analyses robots.txt content you already fetched, and a test asserts no tool accepts aurlargument. - Error notifications redact home directories and credential-shaped fragments.
- Respect the sites you check: a scrapeability report is information, not permission.
Development
bun install
bun run build # esbuild bundle -> dist/extension.js
bun run typecheck # tsc --noEmit (includes tests)
bun run test # vitest unit suite
bun run test:integration # real VS Code extension host
bun run lint # biome
bun run package # VSIX into release/
Architecture and conventions live in AGENTS.md. Changes are tracked in CHANGELOG.md.
Performance
| Input | Size | Found | Time | Rate | Scan speed |
|---|---|---|---|---|---|
| Header signature scan | 2.83 MB | 20,000 | 5.19 ms | 3,852,946/sec | 544.6 MB/s |
| robots.txt path match | 3.32 MB | 60,000 | 9.64 ms | 6,223,689/sec | 344.3 MB/s |
Median of 7 runs after warmup, on Apple M5 Pro, 24 GB RAM, Node 24.3.0. Inputs are generated
by scripts/benchmark.ts rather than checked in, so the sizes above are
exactly what was measured. Reproduce with bun run benchmark.
These are machine-specific and are not asserted in CI — a benchmark that gates a build only tells you how busy the runner was.
Testing
| Metric | Coverage |
|---|---|
| Statements | 92.96% |
| Branches | 83.12% |
| Functions | 95.13% |
| Lines | 94.40% |
305 test cases across 28 files, plus an integration suite that runs
in a real VS Code extension host and an end-to-end test that installs the
built .vsix into a clean profile.
Generated from coverage/coverage-summary.json by
scripts/coverage-readme.js; CI fails if this section drifts from a fresh
run. Reproduce with bun run test:coverage.
More from the LE Family
Every tool in the family, one page: letools.dev
All ten also ship as MCP servers — npx <name>-mcp gives any agent the same engine.
- Paths-LE - Extract file paths from JS/TS imports, JSON, HTML, CSS, TOML, CSV, and .env
- String-LE - Extract string values for i18n from JSON, YAML, CSV, TOML, INI, and .env
- Numbers-LE - Extract numeric values from JSON, YAML, CSV, TOML, INI, and .env
- EnvSync-LE - Spot missing keys across your .env files, with a markdown report
- Regex-LE - Find, test, and validate regular expressions with ReDoS screening
- Secrets-LE - Detect and sanitize credentials locally, before you commit
- Colors-LE - Extract and analyze colors from CSS, SCSS, LESS, Stylus, HTML, JS/TS, and SVG
- URLs-LE - Extract URLs from documentation, configs, and code
- Dates-LE - Extract and analyze dates from logs, configs, and code
Also by nolindnaidoo
Rust
- pixelcoords — Freeze your screen, mark regions, get pixel-exact coordinates and crops pixelcoords.dev · crates.io · docs.rs
- pixelactions — Consume human-verified coordinates, perform the interaction, confirm it landed pixelactions.dev · crates.io · docs.rs
Contact Developer — GitHub · LinkedIn
License
MIT © nolindnaidoo