Skip to content

Opt-in deep crawling via CRAWL4AI_ALLOW_DEEP_CRAWL parameter - #2340

Open
SohamKukreti wants to merge 1 commit into
developfrom
feat/docker-deep-crawl-opt-in
Open

SohamKukreti wants to merge 1 commit into
developfrom
feat/docker-deep-crawl-opt-in

Conversation

@SohamKukreti

Copy link
Copy Markdown
Collaborator

Summary

Fixes #2308

Since 0.9.0 the Docker API rejects any request that sets deep_crawl_strategy (HTTP 400). This PR adds an operator opt-in: with CRAWL4AI_ALLOW_DEEP_CRAWL=true, requests may run BFS / DFS / BestFirst deep crawls with filters and scorers. The default stays closed, and nothing changes for servers that do not set the flag.

Opening the gate alone would expose several holes, so the opt-in comes with these guards:

  • Allowed types only: the 3 strategies, FilterChain, URLPatternFilter, DomainFilter, ContentTypeFilter, and the keyword / composite / domain-authority / freshness / path-depth scorers. SEOFilter and ContentRelevanceFilter stay refused, because they fetch pages with raw httpx outside the egress proxy.
  • No unsafe fields: resume_state (its pending list skips the seed SSRF check), logger, on_state_change and should_cancel are refused.
  • No ReDoS: URLPatternFilter accepts only globs that start with *, /prefix/* or *.ext. Raw regex and ** are refused.
  • Faster pattern matching: *-led globs are now anchored with \A. Before, re.search retried the pattern at every offset, so */blog/* on a 30k-char URL took 3.5 s; now it takes 0.0002 s. Match results are identical, and SDK users get the speedup too.
  • Limits: max_pages / max_depth are clamped to limits.max_pages / limits.max_depth from config.yml. Before, fixed numbers were used and the config keys were never read. The clamp is NaN-safe: "max_pages": NaN used to skip it.
  • Long links skipped: a URLLengthFilter (2048 chars) runs first in the filter chain, so long URLs never reach the patterns, scorers or caches. A null filter_chain is handled too.
  • One start URL per deep crawl request, and no deep crawl inside crawler_configs.
  • No SSL fetch with deep crawl: fetch_ssl_certificate cannot be combined with a deep crawl, because it dials each discovered link directly.
  • List limit: lists (patterns, keywords, filters, scorers) are limited to 100 items.
  • Stream deadline: a deep crawl on /crawl/stream now stops at limits.wall_clock_s. It sends the pages crawled so far, then {"status": "error", "error": "Crawl exceeded the time limit"}. Normal streams are unchanged.

Out of scope, and older library behavior: failed pages do not count toward max_pages, and DFS goes past max_pages. Both happen in the SDK on develop too.

List of files changed and why

  • crawl4ai/async_configs.py: the opt-in: the _deep_crawl_allowed() flag, the UNTRUSTED_DEEP_CRAWL_TYPES allowlist, forbidden strategy fields, glob-only pattern check, list limit, SSL + deep crawl refusal.
  • crawl4ai/deep_crawling/filters.py: \A anchor for *-led globs (1 line, same results, linear time).
  • deploy/docker/governor.py: NaN-safe page/depth clamp, URLLengthFilter, deep_crawl_limits(config) that reads config.yml.
  • deploy/docker/api.py: clamp with the config limits, one-URL rule (also for crawler_configs), deadline for streamed deep crawls.
  • docker-compose.yml: passes CRAWL4AI_ALLOW_DEEP_CRAWL from the shell or .env (same pattern as CRAWL4AI_API_TOKEN).
  • deploy/docker/config.yml: comment on limits.max_pages names the flag.
  • docs/md_v2/core/self-hosting.md, deploy/docker/README.md: new "Deep Crawling (opt-in)" section (flag, example request, server rules). It replaces the *(Keep Deep Crawler Example)* placeholder.
  • deploy/docker/MIGRATION.md: notes that deep_crawl_strategy can be re-enabled.
  • tests/unit/test_config_provenance.py: opt-in works, unsafe patterns / resume_state / SSL + deep crawl are refused, the \A fix gives the same matches and is fast.
  • deploy/docker/tests/test_security_resource_caps.py: NaN clamp, URL-length guard added even with a null filter chain.

How Has This Been Tested?

  • Unit: 7 new tests (8 runs). All fail on develop and pass with this change. Undoing each fix one at a time is caught by its test (12/13; the 100-item list limit is not tested). Flag-off refusal was already covered by deploy/docker/tests/test_security_trust_boundary.py.
  • Existing suites (tests/unit, tests/general filter tests, tests/regression/test_reg_deep_crawl.py, tests/deep_crawling, deploy/docker/tests): same results before and after, apart from the new passes.
  • Filter equivalence: old vs new URLPatternFilter on 110 pattern sets × reverse on/off × 3,357 URLs: 718,398 checks, 0 differences. Real SDK deep crawls (BFS / DFS / BestFirst, 10 pattern sets, 5 live sites) returned identical pages with the old and new filter.
  • Live Docker audit (image built from this branch, hooks and execute_js on):
    • develop image: 32/32. Flag off: 61/61. Flag on (limits 12 pages / depth 3): 62/62.
    • Normal endpoints on 8 real sites (/crawl, multi-URL, stream, job, /md, /html, /screenshot, /pdf, /execute_js, /llm, /config/dump): same results as develop.
    • Deep crawls: BFS / DFS / BestFirst, all allowed filters and scorers, stream, job, limit clamp, NaN, null chain, 4 deep + 4 normal concurrent.
    • 13 unsafe requests rejected with 400.
    • Links to internal IPs from a deep crawl: 0 hits on a host trap server.
    • With wall_clock_s: 8: a deep stream stopped at about 9 s with partial results and the error line, no background fetches after it, normal streams unaffected.
    • 0 tracebacks / 500s in any container.
  • Compose: 8 docker compose up cases (.env true/false/missing, shell export, shell overrides .env, .llm.env only, invalid value) plus docker run --env-file. All behaved as documented.

Checklist:

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • I have added/updated unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Requests may set deep_crawl_strategy when the flag is true, with safe types only,
glob-only URL patterns, config page/depth limits and a deadline on streamed deep crawls.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant