docs

Discover

Find unlinked endpoints — admin panels, backups, .git, .env — with soft-404 calibration that kills false positives.

The Crawler finds what the app links to. Discover finds what it doesn't: hidden directories, forgotten admin panels, backup files, .git, .env, debug endpoints. It requests likely paths from a wordlist and reports the ones that really exist. It ships in Community — no license needed.

The Discovery view listing found paths with status and classification
Request likely paths from a wordlist and report the ones that really exist — soft-404 calibration kills the fake 200s.

Soft-404 calibration

Many apps answer every URL with 200 — a single-page-application (SPA) catch-all, a soft 404. Discover calibrates before it trusts a single hit: it fires several requests to random, non-existent paths, fingerprints the "not found" response (status plus a hash of the body, with a small length tolerance for pages that vary by a timestamp or request id), and then reports only the paths that genuinely differ. The samples have to agree on status and body before a baseline is trusted, and Discover rechecks that baseline as the run goes — roughly every couple hundred confirmed hits — so a subsection with its own custom 404 page doesn't start leaking fake 200s.

Calibration is the difference between a clean list of real findings and a thousand fake 200s. The hits Discover shows you are worth opening.

Sensitive-path heuristics

Discover knows which hits matter. Instead of an anonymous status-code row, a high-value path is classified to a Common Weakness Enumeration (CWE) id and a severity:

  • Exposed version-control directory (.git, .svn, .hg, .bzr) — CWE-527, High.
  • Exposed credential or secret file (.env, cloud creds, kubeconfig, id_rsa) — CWE-538, Critical.
  • Exposed backup or source archive (.bak, .old, .sql, .zip, .tar.gz, …) — CWE-530, High.
  • Exposed debug or management endpoint (actuator, server-status, phpinfo, /metrics, /swagger, /graphql, Jolokia) — CWE-200, Medium.
  • Administrative interface reachable (/admin, /manager, /wp-admin, /phpmyadmin) — CWE-200, Low.

A 401 or 403 still counts as present — a locked admin panel or a denied .env proves the path exists, so Discover keeps it. But a denied sensitive file is the secure configuration, not a leak: a /.env that returns 403 is downgraded to Info, while the same /.env served at 200 stays a Critical exposure. A served hit on a high-signal path (.env, .git, secrets, credentials) is auto-promoted into Findings, with the URL's auth-bearing query stripped.

Turn on Detect DOM sinks and Discover also scans each response body it already fetched — no extra request — for DOM-XSS sinks (innerHTML, eval, document.write, …) and the sources that feed them (the URL, cookies, storage), so a discovered script that builds the page from the URL is flagged for follow-up.

Pick a wordlist

Discover ships one built-in list and takes your own:

common

Roughly 290 high-signal paths, ordered most-likely-first and grouped by attack class — admin and management, API and GraphQL, auth, debug and actuator, framework paths, version-control and infrastructure files, configs, backups, CI/CD, and data stores. Selected by default.

Custom — paste or file

Pick Custom and paste a list, one path per line. Or drop wordlists into ~/.hugin/wordlists/ and they appear in the picker as custom:<name> (hit Refresh after adding one). Files stream up to 4 MB; anything over 16 MB falls back to common.

Extensions

Every extensionless entry is also probed with each extension you list (default .php,.html,.js,.txt), so admin becomes admin, admin.php, admin.html, and so on.

Scan backups

Turn it on and Discover also probes .bak, .old, .zip, .tar.gz, and ~ variants of every path — the forgotten copy of a file the live one won't give up.

For ffuf-style FUZZ-keyword fuzzing with its own engine wordlists, reach for FFuzzer; Discover is the structured content-discovery tool with sensitive-path classification.

Recurse into what you find

Turn on Recursive and every directory-like hit becomes a new scan rooted at that path, so Discover walks /api into /api/v1 into /api/v1/admin without you re-seeding. Three controls keep it from running forever:

  • Depth stops at 3 levels deep.
  • Queue cap holds at most 50 pending sub-scans; past that, the deepest subtree is dropped (logged, not silently lost).
  • No recursion on 401/403. Discover only recurses into 200/301/302 directories. An auth-walled tree answers every sub-path with the same 403 and would loop forever; to walk one, carry credentials in the headers field (see Route the traffic) so the codes turn into 200s.

A visited-path set means the same directory is never queued twice in one run.

Pace the scan

Forced browsing is loud. Tune it to the target:

  • Threads — 1 to 50 concurrent probes (default 10).
  • Delay between probes — 0 to 5000 ms on top of the concurrency cap, for a target that rate-limits or that you'd rather not hammer.
  • Adaptive backoff — a sustained run of 429 or 503 responses stretches the throttle automatically (the more in a row, the longer the pause, up to a few seconds), and only clearly-OK responses reset it, so one stray 404 mid-wall doesn't let the rate creep back up.
  • Stop after N hits — cap the run once N results land (0 = unlimited), for a quick "is anything here" pass.
  • Pause, Resume, Stop at any time. Pause holds the in-flight probes; Stop cancels them and drops the partial run.
  • Resume: skip first N entries — restart a long run past where the last one stopped instead of re-probing from the top.
  • Shuffle paths randomizes probe order; set HUGIN_DISCOVER_SHUFFLE_SEED=<number> to make that order reproducible across operators.

Filter by status

Two filters decide what reaches the table:

  • Status code filter (include) — only these codes are kept; default 200,301,302,403. Leave it empty to keep everything.
  • Exclude status — codes to drop even when the include filter would keep them, for example 404,500.

The soft-404 baseline still applies on top of both, so a host that serves 200 for everything won't flood the include list.

Route the traffic

Through Hugin's proxy (default on)

Replay through proxy is on by default. Every probe flows through Hugin's in-process proxy, so the traffic lands in History and passes your per-project scope, the intercept queue, the match-and-replace rules, and the audit log — discovery is captured like any other tool, not fired off to the side.

An upstream proxy

Point Upstream proxy at Burp or a SOCKS hop (http://, https://, socks5://, socks5h://, socks4://, socks4a://, with optional user:pass@). An explicit upstream wins over the in-process default.

A Host header or custom headers

Override the Host header for virtual-host discovery, and add custom headers (one Name: Value per line) to carry an Authorization or Cookie into every probe — the same session that turns an auth-walled tree's 403s into 200s for recursion.

Every probe is scope-checked before it's sent: a path that resolves to an out-of-scope host is refused, and switching projects mid-scan cancels the run so it can't bleed across engagements. Two environment knobs handle awkward infrastructure: HUGIN_DISCOVER_IPV4_ONLY=1 forces A-only resolution when a target's IPv6 path is broken, and HUGIN_DISCOVER_DNS=host=ip,host2=ip:port pins DNS for a run (CDN testing, internal hosts) without touching your system resolver.

Forced browsing hits a lot of paths fast and probes for files the target never meant to serve. Keep it on hosts you're authorized to test, mind the rate, and let the scope check do its job.

Read and route the results

Hits stream into a sortable table — Status, URL, Size, Content-Type, Time (response time), and Body Hash (the calibration fingerprint, so near-identical pages line up at a glance). A status-code histogram (1xx–5xx) plus a live req/s and error count ride above it, and the table caps at 500 rows with a Show all toggle so a 5000-hit run doesn't choke the view.

Expand a row for the full URL, size, content type, response time, and any DOM sinks and sources found. From a hit you can:

Send to Repeater

Open the URL as a fresh Repeater tab to start crafting an exploit request — no retyping.

Open in browser

Pop the hit in your external browser (http/https only) to eyeball it.

Scope tag

Each hit is tagged in-scope, out-of-scope, or unscoped against the active project — an out-of-scope hit is a stop-and-check signal.

Compare runs and export

A completed run that found something is saved per project, so the result table survives a restart and you build up a record of what a target looked like over time. (Cancelled runs and empty runs aren't saved, so the history stays useful.)

  • Compare with the last run — diff two scans of the same target to see which paths appeared and which vanished. New paths since last time are exactly where to look; a path that disappeared is drift worth a note.
  • Waterfall — read per-path timing to spot latency cliffs. A path far slower than its neighbours often means an expensive endpoint, a cache miss, or a throttle kicking in.
  • Export CSV or JSONL — both strip auth-bearing query strings (tokens, session ids, API keys, passwords) by default, so a shared result file doesn't leak credentials. Flip Reveal auth in export only when you deliberately need the raw URLs. The response-body hash is never exported.

Feed anything interesting into Repeater or the Scanner, and pair Discover with the Crawler to cover both the linked and the unlinked surface.

Last updated 2026-06-17.