{
“@context”: “https://schema.org”,
“@type”: “Article”,
“headline”: “Hunting Phishing Infrastructure with urlscan.io: Pivot from Screenshot to Campaign Takedown”,
“description”: “Learn a hands-on urlscan.io OSINT workflow to pivot from a single phishing screenshot to the full campaign footprint using search operators, page.domains, and certificate data.”,
“author”: {“@type”: “Organization”, “name”: “Hmmnm – Cybersecurity Tutorials”},
“publisher”: {“@type”: “Organization”, “name”: “Hmmnm”}
}
TL;DR: How urlscan.io Turns One Phishing Screenshot into a Campaign Map
A single phishing URL is never just one URL—it’s a doorway into the attacker’s entire infrastructure. urlscan.io lets you pivot from that first screenshot to the full campaign footprint by chaining search operators (hash:, page.domain:, cert.subject:), clustering results by IP, ASN, and page.domains, and building a timeline that turns scattered domains into a takedown-ready package.
What urlscan.io Gives You on a Single Scan
Traditional phishing analysis meant pasting a suspicious URL into a disposable VM, screenshotting the result, and hoping the kit didn’t fingerprint your sandbox. urlscan.io—which has scanned billions of URLs since its 2016 launch by Johannes Gilger—flips that workflow into structured, searchable intelligence.
Every scan produces:
- Screenshot — visual proof of what the victim saw, critical for brand-impersonation evidence.
- DOM snapshot — the full rendered document, including injected scripts and obfuscated HTML.
- HTTP transactions — every request, response, redirect hop, and cookie set.
- IP, ASN, and server banner — hosting fingerprints that survive domain rotation.
- TLS certificate details — subject CN, issuer, SANs, and self-signed patterns.
- Hashed page content — SHA-256 hashes of the DOM and every loaded resource, ripe for pivoting.
- Verdicts and tags — maliciousness signals from urlscan’s own detection plus community and vendor feeds.
The scan result page is your evidence locker. The search interface is your pivot engine—and that’s where campaigns get mapped.
Core Search Operators for Phishing Hunts
urlscan’s search syntax is where OSINT becomes infrastructure hunting. The operators you’ll use daily:
page.domain:example-secure-login.com— scans of the page’s registered domain.page.url:"/wp-content/verify"— scans containing a URL path pattern.ip:185.220.101.5andasn:AS200651— hosting-level clustering.cert.subject.cn:mail-login.example.comandcert.issuer.agn:Let's Encrypt— certificate pivots.hash:2a3f9c...— identical DOM or resource files across sites.filename:logo-office365.png— specific resource filenames in scan results.domain:*.secure-portal-login.top— wildcard domain matching.
Combine these with filters like verdict:malicious, tags such as phishing, and date:>2024-06-01, and you’re no longer chasing URLs—you’re querying an attacker’s build system.
Step 1 — Triage the Initial Phishing URL Safely
Submit the URL through the web UI or the POST /api/v1/scan/ endpoint. urlscan’s fetchers run in controlled infrastructure—they see the page so you don’t have to, eliminating credential-capture risk from your own endpoints. Two rules:
- Never submit URLs containing session tokens or authenticated paths publicly. A scan of
https://portal.example.com/reset?token=abc123burns the token and leaks your brand into a public database. - Use private scans (
"visibility": "private"in the API or the unlisted option in the UI) when handling targeted lures against your organization or clients.
This aligns with CISA’s phishing guidance—analyze suspicious messages in isolation, never from your production environment. CISA’s #StopRansomware and phishing advisories at cisa.gov consistently stress isolated analysis as baseline hygiene.
Step 2 — Pivoting on the Phishing Kit’s Fingerprints
Phishing kits are recycled aggressively. The same actor deploying a Microsoft 365 credential harvester across forty domains ships identical artifacts every time—and identical artifacts produce identical hashes.
Your pivot chain:
- Hash of the main DOM: Run
hash:<dom_hash>. Sibling sites rendering byte-identical pages are running the same kit build. - Resource hashes: The scan lists every loaded asset with its hash. Pivoting on a shared
background.min.jsor the kit’s config JSON catches variants where the HTML differs but the payload doesn’t. - Resource paths and filenames:
filename:office-icon.pngor distinctive paths like/auth/verify/step2.htmlreveal kit templates that get re-hosted with cosmetic HTML changes. - Favicon hashes and brand assets: Attackers hotlinking the same stolen logo image give you a free clustering signal.
This is the same asset-fingerprinting logic OWASP documents in its application security testing guide (owasp.org): predictable artifacts enable enumeration. You’re simply pointing that principle outward at the attacker.
Step 3 — Mapping Hosting Infrastructure with page.domains, IPs, and ASNs
With sibling domains identified, cluster their hosting. ip: queries reveal shared servers; asn: queries reveal shared networks.
Read the results critically:
- Dedicated attacker infrastructure: Multiple phishing domains sharing a single IP or a small IP range on an obscure ASN signals dedicated hosting—your highest-value takedown target.
- Bulletproof hosting: Clusters on ASNs with repeated abuse history (you’ll recognize the same providers recurring across campaigns) tell you registrar/host escalation will be slow and CERT involvement may be necessary.
- CDNs and shared platforms: Domains behind Cloudflare or similar fronting will cluster on CDN IPs—pivot on
page.domainsinstead. This operator captures every domain observed in a page’s resources, links, and redirects, exposing backend call-home domains and redirect targets that the primary domain alone would never reveal.
The distinction between page.domain: (the scanned page’s registered domain) and page.domains: (all domains observed in the page) is the difference between querying one site and querying everything a site touches. For campaign mapping, page.domains is your workhorse.
Step 4 — Using Certificate Data to Link Domains
Attackers automate certificate issuance, and automation leaves fingerprints.
- cert.subject CN patterns: Search
cert.subject.cn:*.secure-login-*.topto find domains sharing naming conventions baked into a kit’s provisioning script. - Issuer reuse:
cert.issuer.agn:combined with date windows catches domains provisioned through the same ACME account in the same burst. Let’s Encrypt’s certificate transparency logs—already indexed by urlscan via CT—mean provisioning runs of 20+ domains in one hour are trivially visible. - Self-signed certificates: Kits deployed without automation often reuse a single self-signed cert;
cert.subjectorcert.fingerprinthashes match across the entire cluster.
Certificate pivots are particularly powerful against fast-flux phishing because domains rotate but issuance behavior doesn’t.
Step 5 — Building the Campaign Timeline and Actor Profile
Sort your accumulated results by scan date. A campaign timeline emerges:
- Registration-to-first-scan gaps reveal how quickly domains are weaponized after purchase.
- Kit version changes over time (visible via hash evolution) show the actor maintaining and updating their tooling—an indicator of a persistent operator, not opportunistic one-shot abuse.
- Brand rotation—Microsoft 365 lures shifting to DHL to payroll portals—maps the actor’s target portfolio.
- Redirect chains across scans show infrastructure reuse: today’s dead phishing domain was last month’s TDS gate.
This is the material that turns a domain list into a threat actor profile suitable for sharing via MISP or an ISAC.
Automating the Workflow with the urlscan.io API
Manual pivots don’t scale across a brand-protection program. urlscan’s API exposes everything. Search:
curl -H "API-Key: $URLSCAN_KEY"
"https://urlscan.io/api/v1/search/?q=page.domain%3A%22secure-login-verify.top%22&size=100"
Submit a private scan:
curl -X POST https://urlscan.io/api/v1/scan/
-H "API-Key: $URLSCAN_KEY"
-H "Content-Type: application/json"
-d '{"url":"https://suspicious-domain.top/login","visibility":"private"}'
A minimal Python loop for campaign tracking:
import requests
HEADERS = {"API-Key": "YOUR_KEY"}
q = 'filename:verify-office.html AND verdict:malicious'
r = requests.get("https://urlscan.io/api/v1/search/",
headers=HEADERS, params={"q": q, "size": 100})
for res in r.json()["results"]:
print(res["page"]["domain"], res["page"]["ip"], res["task"]["time"])
Rate limits are tier-based—free accounts get modest search quotas, paid tiers unlock bulk and live scanning. Poll the scan result endpoint (/api/v1/result/<uuid>/) after submission; scans complete in seconds to a minute. Chain the search API into your SIEM or MISP instance and new kit deployments alert themselves.
From Intelligence to Takedown: Reporting and Escalation
Intelligence that doesn’t reach a registrar is a hobby. Package your findings:
- IOC list: domains, IPs, ASNs, certificate fingerprints, file hashes, and first/last-seen timestamps.
- Evidence: urlscan screenshots, DOM snapshots, and permanent scan URLs—these carry timestamps and survive the domain’s eventual deletion.
- Escalation path: registrar abuse contacts first (via WHOIS/RDAP), hosting provider abuse second, and national CERTs or brand-protection vendors for bulletproof-hosted clusters. APWG (apwg.org) and CISA’s reporting channels accept phishing reports; for .gov and critical-infrastructure impersonation, CISA’s Central Reporting System is the correct destination.
Cite specific scan URLs in your abuse reports. Abuse desks act faster on evidence with screenshots and HTTP transaction logs than on bare domain lists.
Operational Hygiene and OPSEC Tips
- Default to private scans for anything referencing your organization, executives, or clients—public scans advertise that you’re investigating.
- Never scan authenticated or tokenized URLs publicly; you’ll leak sessions and burn detection cues.
- Assume attackers monitor urlscan. Sophisticated actors watch for new scans of their domains and rotate infrastructure when probed. Stagger your submissions and use the API’s private option consistently.
- Archive early: record scan UUIDs and download evidence before results age out of retention on lower tiers.
Practice Lab: CTF-Style Exercise
Reproduce the workflow in a sandboxed exercise:
- Start with
https://microsoft-verify-account.[TLD]/login.html(use a known-dead training domain or a historical scan from urlscan’s public database—don’t fabricate or host phishing content). - Locate an existing public scan of the pattern; extract the DOM hash and the IP.
- Pivot:
hash:<dom_hash>→ note sibling domains →ip:<ip>→asn:<asn>→page.domains:on one sibling to find call-home infrastructure. - Run
cert.subject.cn:wildcard matching the naming pattern and compare issuance dates. - Deliverable: a one-page campaign report with an IOC table, timeline sorted by scan date, and a drafted registrar abuse email citing three permanent scan URLs.
If your pivot chain produces five or more linked domains from one starting hash, you’ve internalized the methodology.
Frequently Asked Questions
Is urlscan.io safe to use on live phishing URLs?
Yes, for public URLs. urlscan.io sandboxes the fetching in its own infrastructure, so your systems never touch the site. Use private scans for sensitive or targeted lures to avoid exposing your investigation in public search results.
Which urlscan.io operator finds sites sharing the same phishing kit?
Pivot on hash: of the DOM or individual resource files—identical kit builds produce identical hashes. Combine with page.domains: to capture related infrastructure and cert.subject lookalike patterns for domains the kit redeploys with cosmetic changes.
Do I need an API key to automate urlscan.io searches?
Public search via the web UI has limits, but the API requires a key. A free API key unlocks /api/v1/search/ with rate limits based on your account tier; paid tiers add higher quotas and bulk capabilities.
Can urlscan.io results be used for takedown evidence?
Yes. Screenshots, timestamps, HTTP transaction logs, and infrastructure data constitute solid supporting evidence for abuse reports to registrars, hosting providers, and CERTs. Reference permanent scan URLs directly in your reports.
What’s the difference between page.domain and page.domains searches?
page.domain: matches the registered domain of the scanned page itself. page.domains: captures all domains observed in the page’s resources, links, and redirects—making it essential for uncovering backend and call-home infrastructure during campaign mapping.
Related reading
- Build a Docker Socket Escape Lab: Why Mounting /var/run/docker.sock Is a Root Kill Chain
- Build an AiTM Phishing Lab with Evilginx: Understanding Session Cookie Theft and Detection
