Skip to content

[Bug]: Failed CrawlResult has links: {} / media: {} instead of the usual keys #2286

Description

@talelboussetta

crawl4ai version

0.9.4 (also on 0.9.2)

Expected Behavior

A result that failed (robots.txt refusal, fetch error, "All proxies failed") has the same links and media keys as a successful one:

"links": {"internal": [], "external": []},
"media": {"images": [], "videos": [], "audios": []}

so a client can read links["internal"] without special-casing failures.

Current Behavior

CrawlResult.links and CrawlResult.media default to {} (models.py#L136-L137). Successful crawls fill them from the scraper's typed Links / Media models, but the results AsyncWebCrawler.arun builds itself never set them:

So in a mixed batch /crawl returns "links": {}, "media": {} for exactly the entries a client has to handle differently, and in the SDK result.links["internal"] raises KeyError. Since #2134 keeps failed results in /crawl responses, every batch with a refused or failed URL has this shape. In one client that read links.internal.length, the resulting TypeError was taken for a failed request and the healthy pages of the batch were re-crawled several times over.

cc @nightcityblade (#2134) @ntohidi

Is this reproducible?

Yes

Inputs Causing the Bug

- Two URLs on one site, one of them disallowed by its robots.txt
- crawler_config params: {"check_robots_txt": true}

Steps to Reproduce

1. Run the Docker image (0.9.4, or built from develop)
2. POST /crawl with both URLs and check_robots_txt=true
3. Compare `links` / `media` of the two results

Code snippets

from crawl4ai.models import CrawlResult

# What AsyncWebCrawler.arun returns for a URL robots.txt disallows
r = CrawlResult(url="https://example.com/private", html="", success=False,
                status_code=403, error_message="Access denied by robots.txt")
print(r.links, r.media)   # {} {}
r.links["internal"]       # KeyError: 'internal'

OS

Linux (Docker image built from develop @ 1f68e5b)

Python version

3.12 (the image's)

Browser

Chromium (Playwright, bundled)

Browser version

No response

Error logs & Screenshots (if applicable)

POST /crawl, check_robots_txt=true, two URLs of a local test site (CRAWL4AI_ALLOW_INTERNAL_URLS=true), /private disallowed by its robots.txt:
[shape] HTTP 200
/index.html success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
/private/index.html success=False status=403 links={} media={} error='Access denied by robots.txt'

Activity

  1. nightcityblade commented on Oct 9, 2026

    @nightcityblade
    Contributor

    Hi, I'd like to work on this. I'll submit a PR shortly.

  2. nightcityblade commented on Oct 9, 2026

    @nightcityblade
    Contributor

    I see #2289 already addresses this issue, so I'm stepping back to avoid duplicate work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    🐞 BugSomething isn't working🩺 Needs TriageNeeds attention of maintainers

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions