docs

Crawler

Map the attack surface — links, forms, and endpoints pulled from HTML and JavaScript — into History.

You can't test what you haven't found. The Crawler walks the target and feeds the attack surface into History as flows. Run it first so the rest of your tools have something to work on. It ships in Community — no license needed.

The Crawler view configured against a target
A breadth-first walk of the target that feeds links, forms, and endpoints — pulled from HTML and JavaScript — into History.

What it does

The core is a static HTTP crawl — a breadth-first (BFS) walk that respects robots.txt and parses sitemap.xml, with per-host rate limiting and manual redirect handling (each hop captured as its own flow). Defaults: depth 3, up to 1000 pages, 10 at a time. Every request and response it sends lands in History, so anything it finds is one right-click away from Repeater, Intruder, or the Scanner.

On each page it also extracts links, inline scripts, HTML comments, meta tags, emails, phone numbers, inline event handlers (onclick, onerror, …), and a Content Security Policy (CSP) analysis — all surfaced on the result. JavaScript analysis runs by default; the next sections cover what it digs out and how to point the crawl where you want it.

Run an authenticated crawl

Most of the interesting surface sits behind a login. To walk it as a logged-in user:

  1. Paste your session cookies (name=value pairs) from a logged-in browser, or carry the session in a header (see Route the traffic).
  2. Set an auth-check URL — a page only an authenticated user can reach.
  3. Start the crawl.

Before it spawns a single worker, the Crawler hits the auth-check URL and proves the session actually authenticates. A login redirect, a 401, a 403, or a logged-out body means the proof failed. (No auth-check URL set? It proves against the first seed instead.)

The proof is fail-closed: in authenticated mode, a failed proof produces zero results — nothing reaches History, the Scanner, or access-control testing. A silently logged-out crawl is worse than none; fed into broken-access-control (BAC) checks it manufactures false positives, so Hugin refuses to hand one downstream.

Two things carry the session through a long crawl:

  • Set-Cookie accumulation. The Crawler keeps a cookie jar seeded with your session and folds in every Set-Cookie it sees across redirects and pages, so rotating tokens and multi-step login chains keep working.
  • Require authenticated. Turn this on to force the proof even when your auth rides in a header rather than cookies.

If your auth is custom headers and most responses still come back logged-out, the Crawler warns at the end that the capture is effectively anonymous — a sign the header didn't take.

Keep the crawl in scope

Same host and subdomains

The default: the seed host plus its subdomains. Turn subdomains off to pin to the exact host.

Include / exclude patterns

Regular expressions matched against the full URL. Exclude wins over include; include overrides the seed-host rule. Exclude /logout, scope-in a CDN host, whatever the engagement calls for.

Accepted URLs only

Strict mode: follow only the seed's path subtree (and your include patterns). An authenticated app shell that links to /pricing or a www. sibling won't drag marketing pages into the crawl.

Project scope

The crawl also obeys your global scope from the Scopes tab — a URL has to pass both gates to be followed.

Add as many seed URLs as you want; each one's host and path subtree joins the scope. The Crawler won't downgrade either: if a seed was HTTPS, it won't follow an HTTP version of the same path, which would leak your session cookies in clear.

Route the traffic

Send the crawl wherever you need it:

  • Routing mode — Direct (default), through Mullvad's SOCKS5 proxy, through Hugin's own proxy (so every crawl request is captured in History too), or a custom HTTP/SOCKS proxy.
  • Custom request headers — added to every request and every redirect hop. This is how you inject Authorization: Bearer … or a fixed Cookie to carry a session the cookie field doesn't cover.
  • User-Agent rotation — fixed, round-robin, or random, drawn from your UA pool and picked per request. It defaults to real browser strings, not a crawler-identifying one.
  • Accept invalid certificates — off by default, so a man-in-the-middle (MITM) can't swap responses on the static engine, the passive sources, or the robots/sitemap fetch. Turn it on only for a target that legitimately serves a broken certificate.

HTTP/3 (QUIC) and extra method probes (OPTIONS/HEAD per URL) are toggles too.

What it pulls out of JavaScript

JavaScript analysis is on by default. It's static — regex over the script text, no execution — so it's safe against anything you've already fetched. From every inline and external script it pulls:

Secrets

12 types: AWS access keys (AKIA…) and secret keys, JWTs, Google (AIza…), GitHub (ghp_… / github_pat_…), Slack (xox…), Stripe (sk_live_… / pk_live_…), PEM private-key blocks, generic API keys, bearer / access / refresh tokens, passwords, and long high-entropy strings. Values are redacted in the report.

Endpoints and API routes

fetch, axios, XMLHttpRequest, and GraphQL calls, plus Express-style route definitions — each tagged with its HTTP method when the code reveals it (GET/POST/PUT/DELETE/PATCH), so you know how to hit it.

Build-time config

Environment-variable references (process.env, import.meta.env, NEXT_PUBLIC_, REACT_APP_, webpack DefinePlugin) and source-map files that can rebuild the original source.

Live channels and storage

WebSocket, Socket.IO, and SockJS endpoints; localStorage / sessionStorage keys; and service / web / shared worker scripts.

It also flags weaknesses in the page's CSP. Anything that looks like an endpoint goes back into the crawl with its method hint, so a route discovered in a bundle gets fetched like any other link.

Render JavaScript-heavy pages

The static engine only sees the server's pre-JavaScript HTML. For a single-page application (SPA) that builds its UI in the browser, turn on headless rendering: a real Chromium drives the page and the same extractors run over the rendered DOM.

Headless rendering is compiled in behind a build-time cdp (Chrome DevTools Protocol) feature and needs a Chromium binary present. Official builds include it; if a build doesn't, requesting it quietly falls back to the static crawl instead of failing. With it on you get:

  • DOM-XSS probing with canaries. Before the page's own scripts run, Hugin injects tagged canary values (HUGIN_TAINT_…) into every query parameter and the URL fragment (or a probe parameter if the URL has none), hooks the dangerous sinks (innerHTML, document.write, eval, …), then checks which canaries reached a sink. Source-to-sink DOM-XSS findings come out with a severity, ready to confirm.
  • SPA auto-merge. For a page that looks JavaScript-built, the rendered DOM is merged back into the result — forms (file-upload forms included), the links the JS draws, and endpoints / source maps / workers / secrets from inline scripts that don't exist in the pre-JS body.
  • Cookies, storage, and traffic. The headless run reads back the page's cookies and localStorage, records the XHR/fetch URLs and console errors it sees, and can capture a screenshot of the rendered page.

Don't get caught in a trap

A naive crawler drowns in calendars and infinite parameter permutations. The Crawler watches for and skips:

  • Calendars — date paths like /2024/01/15 or /events/2024-06-15.
  • Parameter explosion — more than 15 distinct values for one query parameter on a host (session ids, cache-busters).
  • Path-template repetition — more than 10 URLs sharing the same shape once numbers, UUIDs, hashes, dates, and slugs collapse to {param}.
  • Runaway depth — paths past 15 segments.

For deduplication it normalizes every URL first: drops the fragment and default ports, sorts query parameters, strips tracking junk (utm_*, fbclid, gclid, and friends), trims trailing slashes, and lowercases the host — so ?a=1&b=2 and ?b=2&a=1&utm_source=x count as one page.

Submit forms

Form submission is off by default — turn it on deliberately. With it on, the Crawler fills detected forms with test values, submits them, and feeds the response and any new URLs back into the crawl.

File-upload forms (multipart/form-data, or any <input type="file">) are not skipped. With submission on, the Crawler uploads a benign 1×1 GIF as the file part and captures the exact multipart request as a flow. That hands the Scanner a real upload request to aim its file-upload checks at — the Scanner sends the malicious variants; the crawl just finds the endpoint and its shape.

Destructive actions are skipped by action pattern — logout, delete, password reset, unsubscribe, payment/checkout, admin routes, and more — so an auto-submit doesn't log you out or wreck data.

Passive URL sources

Off by default: pull known URLs for the target from the Wayback Machine, Common Crawl, and AlienVault OTX without touching it. Every passive URL still has to pass scope before it joins the crawl.

Read and route the results

Discovered URLs land in a tree grouped by host; pick one to read its depth, source, status, and content type. From the detail panel you can:

  • Send to Repeater — fire the request and jump to Repeater.
  • Send to Intruder — hand the captured flow to Intruder to fuzz.
  • Send to Discover — seed Discover with the host and path prefix.
  • Add host to scope — a two-click confirm so a stray click can't widen your authorized surface.
  • Copy URL — http/https only.

Export the discovered URLs as JSON (records plus run stats), CSV (url,depth,status,content_type,source), plain text (one URL per line), or JSONL for streaming into other tools. The deeper loot — parameters, forms, and the secrets above — rides on each result and on the flows in History, where the full request and response live.

Pause, resume, and stop the crawl any time from the toolbar.

Form submission and headless crawling are active: they can create accounts, send mail, or change data, and an authenticated crawl acts as you. Scope the crawl, mind the skip-list, and think before pointing it at something stateful.

Feed the crawl into the Scanner, pair it with Discover for the paths nothing links to, and review what lands in Findings.

Last updated 2026-06-17.