Opt-in deep crawling via CRAWL4AI_ALLOW_DEEP_CRAWL parameter - #2340
Open
SohamKukreti wants to merge 1 commit into
Open
SohamKukreti wants to merge 1 commit into
SohamKukreti wants to merge 1 commit into
Conversation
Requests may set deep_crawl_strategy when the flag is true, with safe types only, glob-only URL patterns, config page/depth limits and a deadline on streamed deep crawls.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #2308
Since 0.9.0 the Docker API rejects any request that sets
deep_crawl_strategy(HTTP 400). This PR adds an operator opt-in: withCRAWL4AI_ALLOW_DEEP_CRAWL=true, requests may run BFS / DFS / BestFirst deep crawls with filters and scorers. The default stays closed, and nothing changes for servers that do not set the flag.Opening the gate alone would expose several holes, so the opt-in comes with these guards:
FilterChain,URLPatternFilter,DomainFilter,ContentTypeFilter, and the keyword / composite / domain-authority / freshness / path-depth scorers.SEOFilterandContentRelevanceFilterstay refused, because they fetch pages with rawhttpxoutside the egress proxy.resume_state(itspendinglist skips the seed SSRF check),logger,on_state_changeandshould_cancelare refused.URLPatternFilteraccepts only globs that start with*,/prefix/*or*.ext. Raw regex and**are refused.*-led globs are now anchored with\A. Before,re.searchretried the pattern at every offset, so*/blog/*on a 30k-char URL took 3.5 s; now it takes 0.0002 s. Match results are identical, and SDK users get the speedup too.max_pages/max_depthare clamped tolimits.max_pages/limits.max_depthfromconfig.yml. Before, fixed numbers were used and the config keys were never read. The clamp is NaN-safe:"max_pages": NaNused to skip it.URLLengthFilter(2048 chars) runs first in the filter chain, so long URLs never reach the patterns, scorers or caches. Anullfilter_chainis handled too.crawler_configs.fetch_ssl_certificatecannot be combined with a deep crawl, because it dials each discovered link directly./crawl/streamnow stops atlimits.wall_clock_s. It sends the pages crawled so far, then{"status": "error", "error": "Crawl exceeded the time limit"}. Normal streams are unchanged.Out of scope, and older library behavior: failed pages do not count toward
max_pages, and DFS goes pastmax_pages. Both happen in the SDK ondeveloptoo.List of files changed and why
crawl4ai/async_configs.py: the opt-in: the_deep_crawl_allowed()flag, theUNTRUSTED_DEEP_CRAWL_TYPESallowlist, forbidden strategy fields, glob-only pattern check, list limit, SSL + deep crawl refusal.crawl4ai/deep_crawling/filters.py:\Aanchor for*-led globs (1 line, same results, linear time).deploy/docker/governor.py: NaN-safe page/depth clamp,URLLengthFilter,deep_crawl_limits(config)that readsconfig.yml.deploy/docker/api.py: clamp with the config limits, one-URL rule (also forcrawler_configs), deadline for streamed deep crawls.docker-compose.yml: passesCRAWL4AI_ALLOW_DEEP_CRAWLfrom the shell or.env(same pattern asCRAWL4AI_API_TOKEN).deploy/docker/config.yml: comment onlimits.max_pagesnames the flag.docs/md_v2/core/self-hosting.md,deploy/docker/README.md: new "Deep Crawling (opt-in)" section (flag, example request, server rules). It replaces the*(Keep Deep Crawler Example)*placeholder.deploy/docker/MIGRATION.md: notes thatdeep_crawl_strategycan be re-enabled.tests/unit/test_config_provenance.py: opt-in works, unsafe patterns /resume_state/ SSL + deep crawl are refused, the\Afix gives the same matches and is fast.deploy/docker/tests/test_security_resource_caps.py: NaN clamp, URL-length guard added even with anullfilter chain.How Has This Been Tested?
developand pass with this change. Undoing each fix one at a time is caught by its test (12/13; the 100-item list limit is not tested). Flag-off refusal was already covered bydeploy/docker/tests/test_security_trust_boundary.py.tests/unit,tests/generalfilter tests,tests/regression/test_reg_deep_crawl.py,tests/deep_crawling,deploy/docker/tests): same results before and after, apart from the new passes.URLPatternFilteron 110 pattern sets × reverse on/off × 3,357 URLs: 718,398 checks, 0 differences. Real SDK deep crawls (BFS / DFS / BestFirst, 10 pattern sets, 5 live sites) returned identical pages with the old and new filter.execute_json):developimage: 32/32. Flag off: 61/61. Flag on (limits 12 pages / depth 3): 62/62./crawl, multi-URL, stream, job,/md,/html,/screenshot,/pdf,/execute_js,/llm,/config/dump): same results asdevelop.nullchain, 4 deep + 4 normal concurrent.wall_clock_s: 8: a deep stream stopped at about 9 s with partial results and the error line, no background fetches after it, normal streams unaffected.docker compose upcases (.envtrue/false/missing, shell export, shell overrides.env,.llm.envonly, invalid value) plusdocker run --env-file. All behaved as documented.Checklist: