A Playwright-powered web scraper with a FastAPI UI, plus a three-tier programmatic fetch stack (HTTP → browser → stealth). It renders JavaScript-heavy pages and extracts structured content (text, comments, videos, images, metadata), with optional local media download and playback.
Language / 语言: English | 简体中文
Enter a URL → open Advanced Options → start scrape → view Text, Videos, Log, and Selectors tabs.
| Layer | Module | Role |
|---|---|---|
| Web UI / SSE API | app.py + scraper_core.py |
Full scrape pipeline for browsers and video sites |
| Tier 1 — HTTP | fetcher.py |
Fast stealth requests (curl_cffi): TLS fingerprint, headers, optional HTTP/3 |
| Tier 2 — Browser | dynamic_fetcher.py |
Playwright Chromium / Google Chrome for JS/SPA pages |
| Tier 3 — Stealth | stealthy_fetcher.py |
Patchright/Playwright + fingerprint spoofing + Cloudflare UI flow |
| Sessions | sessions.py |
Sync + async Session classes — cookies + state |
| Async | async_*_fetcher.py + AsyncFetcher* |
Full asyncio twins for every Fetcher / Session |
| Proxy | proxy_rotator.py |
Round-robin / random / custom rotation; per-request override |
| DNS | doh.py |
Optional Cloudflare DNS-over-HTTPS (leak protection with proxies) |
| Blocking | request_blocking.py |
blocked_domains + block_ads (~3,500 trackers) |
- Real browser rendering (Playwright) — JS, SPAs, lazy-loaded content
- Auto body/comment detection + optional CSS overrides
- Video & image extraction (incl. lazy-load attrs like
data-src) - Smart image filtering (icons, sprites, junk thumbnails)
- Metadata from
<meta>tags; export to TXT / JSON
- Auto-detects: Bilibili, YouTube, Vimeo, TikTok, Douyin, Twitter/X, Twitch, Dailymotion, Niconico
- Bilibili —
__INITIAL_STATE__,__playinfo__, WBI comment pagination, DASH streams - Others — Open Graph, JSON-LD,
ytInitialPlayerResponse, DOM<video> - Result fields:
platform,platform_data, curated images (cover / avatar / first-frame) - Known video URLs skip the generic auto-selector
- Auto-download to
downloads/; magic-byte validation (rejects HTML error pages) - ffmpeg remux for Bilibili DASH (
.m4s→ playable.mp4) - Videos tab — inline
<video>player via/downloads/...with correct MIME types
- Persistent Chrome profile in
.chrome_profile/ - Log in once (Visible mode); reuse sessions for any site
- Optional Cookie field overrides the saved profile
- Heuristic DOM scoring + stable CSS generation
- AI fallback (OpenAI-compatible: OpenAI, DeepSeek, Ollama)
- Selectors tab — method, confidence, one-click apply
- System Chrome +
playwright-stealth+ fingerprint patches - Human-like mouse/scroll; challenge-page wait; multi-strategy retry
- Proxy via UI /
SCRAPER_PROXY/HTTP_PROXY; dead env proxies skipped - ProxyRotator for all Sessions; domain/ad blocking on browser fetchers
- DNS-over-HTTPS — optional Cloudflare DoH (
dns_over_https=True/SCRAPER_DOH=1) to prevent DNS leaks when using proxies - Async —
AsyncFetcher/AsyncDynamicFetcher/AsyncStealthyFetcherand matchingAsync*Sessionclasses - Port auto-select if 8000 is busy (8001+)
spaider_crawler/
├── app.py # FastAPI + SSE API + /downloads
├── scraper_core.py # Main Playwright scrape pipeline
├── selector_engine.py # Heuristic + AI selector discovery
├── media_downloader.py # Image/video download, ffmpeg, MIME
├── fetcher.py # Stealth HTTP + AsyncFetcher (curl_cffi)
├── dynamic_fetcher.py # Playwright DynamicFetcher
├── async_dynamic_fetcher.py # AsyncDynamicSession / AsyncDynamicFetcher
├── stealthy_fetcher.py # StealthyFetcher + CF challenge flow
├── async_stealthy_fetcher.py # AsyncStealthySession / AsyncStealthyFetcher
├── session_store.py # Cookie / state JSON helpers
├── sessions.py # Unified sync + async session exports
├── proxy_rotator.py # Proxy rotation strategies
├── doh.py # Cloudflare DNS-over-HTTPS helpers
├── request_blocking.py # Domain + ad request blocking
├── ad_domains.py # Loads bundled tracker list
├── image_utils.py # Image URL cleanup / junk filter
├── data/ad_domains.txt # ~3,500 ad/tracker hosts (Peter Lowe)
├── video_platforms/ # Multi-platform video extractors
├── templates/index.html
├── static/css|js/
├── scripts/start.ps1 # Windows launcher
├── scripts/scrape_video.py
├── requirements.txt
└── .env.example
- Python 3.10+
- Google Chrome (optional, recommended)
- ffmpeg (optional, for DASH → MP4)
- patchright (optional, stronger
StealthyFetcher)
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromiumOptional stealth engine:
pip install patchright
python -m patchright install chromeCopy .env.example → .env if needed:
OPENAI_API_KEY=sk-your-key-here
# SCRAPER_PROXY=http://127.0.0.1:7890
# BILI_COOKIE=SESSDATA=...; bili_jct=...Windows:
.\scripts\start.ps1Or:
python app.py
# python -m uvicorn app:app --host 127.0.0.1 --port 8000 --reloadOpen http://127.0.0.1:8000/ (or the port printed). Header should show v1.3.0.
| Option | Value |
|---|---|
| Remember login | On (Visible, first run) |
| JS wait | 8000 ms |
| Auto-scroll / system Chrome / Auto-download | On |
| Smart auto-selector | Off (auto for video sites) |
Same as above; use Visible mode for stream URLs. Set Proxy only if your network needs it.
python scripts/scrape_video.py "https://www.bilibili.com/video/BV1yk7X6KEz4" output.json| Option | Description |
|---|---|
| Text / Comment selector | CSS; empty = auto |
| Remember login | .chrome_profile/ persistent session |
| Cookie | Optional override (k=v; ...) |
| Proxy | http:// / socks5://; empty = direct |
| JS wait (ms) | 500–30000 after load |
| Browser mode | Auto / Headless / Visible |
| Max retries | 0–4 alternate strategies |
| Use system Chrome | Prefer installed Chrome |
| Simulate human | Mouse + scroll noise |
| Block resources | Skip images/fonts (may look bot-like) |
| Auto-download | Save media; play in Videos tab |
| Smart auto-selector / AI | Discover CSS; AI needs API key |
Protected sites: Remember login + Visible + system Chrome.
Video sites: Leave selectors empty; enable Auto-download.
from fetcher import Fetcher, FetcherSession
r = Fetcher.get("https://example.com", stealthy_headers=True, impersonate="chrome")
r = Fetcher.get("https://http3-capable.example", http3=True)
with FetcherSession(session_file=".sessions/api.json") as s:
s.get("https://example.com/login")
s.state["user"] = "alice"
s.post("https://example.com/api", json_body={"q": 1})Uses curl_cffi for TLS/JA3 impersonation; falls back to urllib if missing.
import asyncio
from fetcher import AsyncFetcher, AsyncFetcherSession
async def main():
r = await AsyncFetcher.get("https://example.com", stealthy_headers=True)
async with AsyncFetcherSession(session_file=".sessions/api.json") as s:
await s.get("https://example.com/login")
await s.save()
asyncio.run(main())Or via sync class helpers: await Fetcher.async_get(...).
from dynamic_fetcher import DynamicFetcher, DynamicSession
r = DynamicFetcher.fetch(
"https://spa.example.com",
real_chrome=True,
network_idle=True,
wait=1500,
wait_selector="main",
block_ads=True,
blocked_domains={"metrics.vendor.com"},
dns_over_https=True, # Cloudflare DoH — avoid DNS leaks with proxies
)
with DynamicSession(real_chrome=True, session_file=".sessions/web.json") as s:
s.fetch("https://example.com")
s.fetch("https://example.com/account") # cookies persistfrom stealthy_fetcher import StealthyFetcher, StealthySession
r = StealthyFetcher.fetch(
"https://protected.example",
solve_cloudflare=True,
hide_canvas=True,
block_webrtc=True,
block_ads=True,
dns_over_https=True,
real_chrome=True,
timeout=60000,
)CF flow automates the challenge UI in a realistic browser — it does not cryptographically break CAPTCHAs.
from sessions import FetcherSession, DynamicSession, StealthySession
from sessions import AsyncFetcherSession, AsyncDynamicSession, AsyncStealthySession
from sessions import ProxyRotator, random_rotation
# Shared API: get/set/clear cookies, save/load/snapshot/restore, state={}
with FetcherSession(session_file=".sessions/api.json") as s:
s.set_cookies({"token": "x"}, url="https://example.com")
s.save()
# Async sessions
async with AsyncDynamicSession(real_chrome=True) as s:
await s.fetch("https://example.com")
rotator = ProxyRotator([
"http://1.2.3.4:8080",
{"server": "http://5.6.7.8:8080", "username": "u", "password": "p"},
])
with FetcherSession(proxy_rotator=rotator) as s:
s.get("https://example.com/a") # #1
s.get("https://example.com/b") # #2
s.get("https://example.com/c", proxy=None) # direct this call
print(s.last_proxy)
# Random / custom strategies also supported:
# ProxyRotator(proxies, strategy=random_rotation)
# HTTP layer DoH (libcurl CURLOPT_DOH_URL via curl_cffi):
from fetcher import Fetcher
Fetcher.get("https://example.com", proxy="http://127.0.0.1:7890", dns_over_https=True)Do not pass both static proxy= and proxy_rotator= on the same session. Per-request proxy= always wins.
Set SCRAPER_DOH=1 (or DNS_OVER_HTTPS=true) to enable DoH by default for the Web UI pipeline.
{
"status": "ok",
"version": "1.3.0",
"features": [
"video_platforms", "wbi_comments", "download_media", "saved_profile",
"stealth_fetcher", "dynamic_fetcher", "stealthy_fetcher",
"session_manager", "proxy_rotator", "request_blocking", "dns_over_https",
"async_fetchers"
]
}Serves downloaded media with correct MIME types (e.g. video/mp4).
SSE stream. Body fields:
| Field | Default | Description |
|---|---|---|
url |
required | Target URL |
text_selector / comment_selector |
"" |
CSS overrides |
cookie |
"" |
Auth cookie string |
proxy |
"" |
Proxy URL |
wait_ms |
3500 |
JS settle (500–30000) |
scroll |
true |
Auto-scroll |
use_chrome |
true |
System Chrome |
headless |
"auto" |
auto / hidden / visible |
max_retries |
2 |
0–4 |
simulate_human |
true |
Mouse/scroll |
block_resources |
false |
Skip images/fonts |
dns_over_https |
false |
Cloudflare DoH (also env SCRAPER_DOH) |
auto_selector / auto_selector_ai |
true |
Smart selectors |
ai_api_key / ai_base_url / ai_model |
"" |
LLM overrides |
download_media |
true |
Save to downloads/ |
use_saved_profile |
true |
.chrome_profile/ |
SSE events: log, ping, done, error. Validation errors → HTTP 422.
import json, urllib.request
body = json.dumps({
"url": "https://www.bilibili.com/video/BV1yk7X6KEz4",
"wait_ms": 8000, "use_chrome": True,
"download_media": True, "use_saved_profile": True,
"auto_selector": False,
}).encode()
req = urllib.request.Request(
"http://127.0.0.1:8000/api/scrape", data=body,
headers={"Content-Type": "application/json"}, method="POST",
)
with urllib.request.urlopen(req, timeout=300) as resp:
for line in resp.read().decode().splitlines():
if line.startswith("data: "):
print(line[6:]){
"url": "https://www.bilibili.com/video/BV1yk7X6KEz4",
"title": "Video title",
"platform": "bilibili",
"text_paragraphs": ["播放 ...", "UP主: ..."],
"comments": ["user: comment"],
"videos": ["/downloads/.../videos/video.mp4"],
"images": ["https://.../cover.jpg"],
"meta": { "video_platform": "bilibili", "bilibili_bvid": "BV1yk7X6KEz4" },
"platform_data": {
"platform": "bilibili",
"video_streams": [{ "url": "...", "width": 1920 }],
"audio_streams": [{ "url": "..." }],
"comments": ["user: comment"]
},
"downloads": {
"dir": ".../downloads/...",
"web_dir": "/downloads/...",
"images": [{ "web_path": "/downloads/.../images/001.jpg" }],
"videos": [{ "web_path": "/downloads/.../videos/video.mp4", "playable": true, "mime": "video/mp4" }]
}
}Tabs: Text · Comments · Videos · Images · Selectors · Metadata · Log.
| Problem | Solution |
|---|---|
| Playwright won't launch | python -m playwright install chromium |
| Port 8000 busy | Use .\scripts\start.ps1 or the port shown (8001+) |
ERR_CONNECTION_CLOSED |
Clear dead HTTP_PROXY; try Visible + system Chrome |
| Cloudflare / WAF | Visible + Chrome + proxy; or StealthyFetcher(solve_cloudflare=True) |
| Empty SPA content | Increase JS wait; enable auto-scroll |
| CAPTCHA | Visible + Remember login; solve once manually |
| Video won't play | Auto-download on; install ffmpeg for DASH |
| Download is HTML | Need login / stream expired — Visible + saved profile |
| Few Bilibili comments | Remember login or BILI_COOKIE |
| Profile locked | Close other Chrome/scraper using .chrome_profile/ |
| Patchright missing | Optional: pip install patchright && python -m patchright install chrome |
- Only scrape content you are authorized to access. Respect
robots.txtand site terms. - For learning and legitimate research — not a universal CAPTCHA/WAF bypass.
- Use your own cookies/sessions; never misuse others’ credentials.
- Video streams may be copyrighted — use data responsibly.
BSD 3-Clause © 2026 Nameless

