When an untrusted client selects PDFContentScrapingStrategy on the Docker API, the server downloads and parses a remote PDF with no limit on file size, page count, or (by default) wall-clock time. A single request pointing at a large or many-page PDF can exhaust disk, CPU, and bandwidth on shared workers.
Type: Uncontrolled resource consumption / Denial of Service (availability only).
Who is affected: Multi-tenant or shared Crawl4AI Docker deployments that accept crawl requests from untrusted or semi-trusted clients.
What an attacker can do: With one or a few requests, cause unbounded disk writes (the downloaded PDF is streamed to a temp file with no cap), unbounded CPU during page parsing, and bandwidth/cost amplification. No data disclosure or code execution results from this issue alone.
PDFContentScrapingStrategy is in UNTRUSTED_ALLOWED_TYPES , and the non-streaming /crawl handler does not force LXMLWebScrapingStrategy (the streaming handler does), so an untrusted body can select the PDF strategy.
On that path there are no resource bounds:
crawl4ai/processors/pdf/__init__.py _get_pdf_path() streams the full remote body to a temp file; content-length is read only for progress logging, never to abort.
crawl4ai/processors/pdf/processor.py iterates every page ( enumerate(reader.pages) / range(total_pages) ) with no page cap.
deploy/docker/config.yml ships limits.wall_clock_s: 0 (no per-crawl deadline) by default.
Note: the existing BodySizeLimitMiddleware ( limits.max_body_bytes , default 10 MiB) only limits the size of the inbound HTTP request body, not the size of the remote PDF the server fetches, so it does not mitigate this. In-process memory guards reduce the RAM-exhaustion angle but do not bound disk or CPU.
Send POST /crawl with PDFContentScrapingStrategy and a URL pointing at a large (or high-page-count) remote PDF. The worker downloads the entire body to disk and parses all pages with no cap and, by default, no deadline.
Suggested remediation
Enforce a max_pdf_bytes limit in the downloader; abort and delete the temp file when the streamed size exceeds it (do not trust content-length ).
Enforce a max_pdf_pages limit in the processor.
Ship a non-zero default limits.wall_clock_s in the production Docker image.
Cap or disable extract_images for untrusted callers.
Consider gating PDFContentScrapingStrategy behind an admin feature flag for untrusted API bodies.
Reported privately by Nguyen Tran Thanh Lam (c240030).
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
