Hunting Phishing Infrastructure with urlscan.io: Pivot from Screenshot to Campaign Takedown

Hunting Phishing Infrastructure with urlscan.io: Pivot from Screenshot to Campaign Takedown

📋 Key Takeaways
  • A single phishing URL is never just one URL—it's a doorway into the attacker's entire infrastructure. urlscan.io lets you pivot from that first screenshot to the full campaign footprint by chaining search operators (hash:, page.domain:, cert.subject:), clustering results by IP, ASN, and page.domains, and building a timeline that turns scattered domains into a takedown-ready package.
  • Traditional phishing analysis meant pasting a suspicious URL into a disposable VM, screenshotting the result, and hoping the kit didn't fingerprint your sandbox. urlscan.io—which has scanned billions of URLs since its 2016 launch by Johannes Gilger—flips that workflow into structured, searchable intelligence.
  • urlscan's search syntax is where OSINT becomes infrastructure hunting.
  • Submit the URL through the web UI or the POST /api/v1/scan/ endpoint. urlscan's fetchers run in controlled infrastructure—they see the page so you don't have to, eliminating credential-capture risk from your own endpoints.
10 min read · 1,881 words
Educational & Ethical Use Only — This article is provided for educational and ethical cybersecurity research purposes only. The techniques described should only be used on systems you own or have explicit permission to test. Always follow responsible disclosure and the laws applicable to you. Mitigations are included so engineers can harden real systems.
Security· 10 min read

{
“@context”: “https://schema.org”,
“@type”: “Article”,
“headline”: “Hunting Phishing Infrastructure with urlscan.io: Pivot from Screenshot to Campaign Takedown”,
“description”: “Learn a hands-on urlscan.io OSINT workflow to pivot from a single phishing screenshot to the full campaign footprint using search operators, page.domains, and certificate data.”,
“author”: {“@type”: “Organization”, “name”: “Hmmnm – Cybersecurity Tutorials”},
“publisher”: {“@type”: “Organization”, “name”: “Hmmnm”}
}

TL;DR: How urlscan.io Turns One Phishing Screenshot into a Campaign Map

A single phishing URL is never just one URL—it’s a doorway into the attacker’s entire infrastructure. urlscan.io lets you pivot from that first screenshot to the full campaign footprint by chaining search operators (hash:, page.domain:, cert.subject:), clustering results by IP, ASN, and page.domains, and building a timeline that turns scattered domains into a takedown-ready package.

What urlscan.io Gives You on a Single Scan

Traditional phishing analysis meant pasting a suspicious URL into a disposable VM, screenshotting the result, and hoping the kit didn’t fingerprint your sandbox. urlscan.io—which has scanned billions of URLs since its 2016 launch by Johannes Gilger—flips that workflow into structured, searchable intelligence.

Every scan produces:

  • Screenshot — visual proof of what the victim saw, critical for brand-impersonation evidence.
  • DOM snapshot — the full rendered document, including injected scripts and obfuscated HTML.
  • HTTP transactions — every request, response, redirect hop, and cookie set.
  • IP, ASN, and server banner — hosting fingerprints that survive domain rotation.
  • TLS certificate details — subject CN, issuer, SANs, and self-signed patterns.
  • Hashed page content — SHA-256 hashes of the DOM and every loaded resource, ripe for pivoting.
  • Verdicts and tags — maliciousness signals from urlscan’s own detection plus community and vendor feeds.

The scan result page is your evidence locker. The search interface is your pivot engine—and that’s where campaigns get mapped.

Core Search Operators for Phishing Hunts

urlscan’s search syntax is where OSINT becomes infrastructure hunting. The operators you’ll use daily:

  • page.domain:example-secure-login.com — scans of the page’s registered domain.
  • page.url:"/wp-content/verify" — scans containing a URL path pattern.
  • ip:185.220.101.5 and asn:AS200651 — hosting-level clustering.
  • cert.subject.cn:mail-login.example.com and cert.issuer.agn:Let's Encrypt — certificate pivots.
  • hash:2a3f9c... — identical DOM or resource files across sites.
  • filename:logo-office365.png — specific resource filenames in scan results.
  • domain:*.secure-portal-login.top — wildcard domain matching.

Combine these with filters like verdict:malicious, tags such as phishing, and date:>2024-06-01, and you’re no longer chasing URLs—you’re querying an attacker’s build system.

Step 1 — Triage the Initial Phishing URL Safely

Submit the URL through the web UI or the POST /api/v1/scan/ endpoint. urlscan’s fetchers run in controlled infrastructure—they see the page so you don’t have to, eliminating credential-capture risk from your own endpoints. Two rules:

  • Never submit URLs containing session tokens or authenticated paths publicly. A scan of https://portal.example.com/reset?token=abc123 burns the token and leaks your brand into a public database.
  • Use private scans ("visibility": "private" in the API or the unlisted option in the UI) when handling targeted lures against your organization or clients.

This aligns with CISA’s phishing guidance—analyze suspicious messages in isolation, never from your production environment. CISA’s #StopRansomware and phishing advisories at cisa.gov consistently stress isolated analysis as baseline hygiene.

Step 2 — Pivoting on the Phishing Kit’s Fingerprints

Phishing kits are recycled aggressively. The same actor deploying a Microsoft 365 credential harvester across forty domains ships identical artifacts every time—and identical artifacts produce identical hashes.

Your pivot chain:

  1. Hash of the main DOM: Run hash:<dom_hash>. Sibling sites rendering byte-identical pages are running the same kit build.
  2. Resource hashes: The scan lists every loaded asset with its hash. Pivoting on a shared background.min.js or the kit’s config JSON catches variants where the HTML differs but the payload doesn’t.
  3. Resource paths and filenames: filename:office-icon.png or distinctive paths like /auth/verify/step2.html reveal kit templates that get re-hosted with cosmetic HTML changes.
  4. Favicon hashes and brand assets: Attackers hotlinking the same stolen logo image give you a free clustering signal.

This is the same asset-fingerprinting logic OWASP documents in its application security testing guide (owasp.org): predictable artifacts enable enumeration. You’re simply pointing that principle outward at the attacker.

Step 3 — Mapping Hosting Infrastructure with page.domains, IPs, and ASNs

With sibling domains identified, cluster their hosting. ip: queries reveal shared servers; asn: queries reveal shared networks.

Read the results critically:

  • Dedicated attacker infrastructure: Multiple phishing domains sharing a single IP or a small IP range on an obscure ASN signals dedicated hosting—your highest-value takedown target.
  • Bulletproof hosting: Clusters on ASNs with repeated abuse history (you’ll recognize the same providers recurring across campaigns) tell you registrar/host escalation will be slow and CERT involvement may be necessary.
  • CDNs and shared platforms: Domains behind Cloudflare or similar fronting will cluster on CDN IPs—pivot on page.domains instead. This operator captures every domain observed in a page’s resources, links, and redirects, exposing backend call-home domains and redirect targets that the primary domain alone would never reveal.

The distinction between page.domain: (the scanned page’s registered domain) and page.domains: (all domains observed in the page) is the difference between querying one site and querying everything a site touches. For campaign mapping, page.domains is your workhorse.

Attackers automate certificate issuance, and automation leaves fingerprints.

  • cert.subject CN patterns: Search cert.subject.cn:*.secure-login-*.top to find domains sharing naming conventions baked into a kit’s provisioning script.
  • Issuer reuse: cert.issuer.agn: combined with date windows catches domains provisioned through the same ACME account in the same burst. Let’s Encrypt’s certificate transparency logs—already indexed by urlscan via CT—mean provisioning runs of 20+ domains in one hour are trivially visible.
  • Self-signed certificates: Kits deployed without automation often reuse a single self-signed cert; cert.subject or cert.fingerprint hashes match across the entire cluster.

Certificate pivots are particularly powerful against fast-flux phishing because domains rotate but issuance behavior doesn’t.

Step 5 — Building the Campaign Timeline and Actor Profile

Sort your accumulated results by scan date. A campaign timeline emerges:

  • Registration-to-first-scan gaps reveal how quickly domains are weaponized after purchase.
  • Kit version changes over time (visible via hash evolution) show the actor maintaining and updating their tooling—an indicator of a persistent operator, not opportunistic one-shot abuse.
  • Brand rotation—Microsoft 365 lures shifting to DHL to payroll portals—maps the actor’s target portfolio.
  • Redirect chains across scans show infrastructure reuse: today’s dead phishing domain was last month’s TDS gate.

This is the material that turns a domain list into a threat actor profile suitable for sharing via MISP or an ISAC.

Automating the Workflow with the urlscan.io API

Manual pivots don’t scale across a brand-protection program. urlscan’s API exposes everything. Search:

curl -H "API-Key: $URLSCAN_KEY" 
  "https://urlscan.io/api/v1/search/?q=page.domain%3A%22secure-login-verify.top%22&size=100"

Submit a private scan:

curl -X POST https://urlscan.io/api/v1/scan/ 
  -H "API-Key: $URLSCAN_KEY" 
  -H "Content-Type: application/json" 
  -d '{"url":"https://suspicious-domain.top/login","visibility":"private"}'

A minimal Python loop for campaign tracking:

import requests

HEADERS = {"API-Key": "YOUR_KEY"}
q = 'filename:verify-office.html AND verdict:malicious'
r = requests.get("https://urlscan.io/api/v1/search/",
                 headers=HEADERS, params={"q": q, "size": 100})
for res in r.json()["results"]:
    print(res["page"]["domain"], res["page"]["ip"], res["task"]["time"])

Rate limits are tier-based—free accounts get modest search quotas, paid tiers unlock bulk and live scanning. Poll the scan result endpoint (/api/v1/result/<uuid>/) after submission; scans complete in seconds to a minute. Chain the search API into your SIEM or MISP instance and new kit deployments alert themselves.

From Intelligence to Takedown: Reporting and Escalation

Intelligence that doesn’t reach a registrar is a hobby. Package your findings:

  • IOC list: domains, IPs, ASNs, certificate fingerprints, file hashes, and first/last-seen timestamps.
  • Evidence: urlscan screenshots, DOM snapshots, and permanent scan URLs—these carry timestamps and survive the domain’s eventual deletion.
  • Escalation path: registrar abuse contacts first (via WHOIS/RDAP), hosting provider abuse second, and national CERTs or brand-protection vendors for bulletproof-hosted clusters. APWG (apwg.org) and CISA’s reporting channels accept phishing reports; for .gov and critical-infrastructure impersonation, CISA’s Central Reporting System is the correct destination.

Cite specific scan URLs in your abuse reports. Abuse desks act faster on evidence with screenshots and HTTP transaction logs than on bare domain lists.

Operational Hygiene and OPSEC Tips

  • Default to private scans for anything referencing your organization, executives, or clients—public scans advertise that you’re investigating.
  • Never scan authenticated or tokenized URLs publicly; you’ll leak sessions and burn detection cues.
  • Assume attackers monitor urlscan. Sophisticated actors watch for new scans of their domains and rotate infrastructure when probed. Stagger your submissions and use the API’s private option consistently.
  • Archive early: record scan UUIDs and download evidence before results age out of retention on lower tiers.

Practice Lab: CTF-Style Exercise

Reproduce the workflow in a sandboxed exercise:

  1. Start with https://microsoft-verify-account.[TLD]/login.html (use a known-dead training domain or a historical scan from urlscan’s public database—don’t fabricate or host phishing content).
  2. Locate an existing public scan of the pattern; extract the DOM hash and the IP.
  3. Pivot: hash:<dom_hash> → note sibling domains → ip:<ip> → asn:<asn> → page.domains: on one sibling to find call-home infrastructure.
  4. Run cert.subject.cn: wildcard matching the naming pattern and compare issuance dates.
  5. Deliverable: a one-page campaign report with an IOC table, timeline sorted by scan date, and a drafted registrar abuse email citing three permanent scan URLs.

If your pivot chain produces five or more linked domains from one starting hash, you’ve internalized the methodology.

Frequently Asked Questions

Is urlscan.io safe to use on live phishing URLs?

Yes, for public URLs. urlscan.io sandboxes the fetching in its own infrastructure, so your systems never touch the site. Use private scans for sensitive or targeted lures to avoid exposing your investigation in public search results.

Which urlscan.io operator finds sites sharing the same phishing kit?

Pivot on hash: of the DOM or individual resource files—identical kit builds produce identical hashes. Combine with page.domains: to capture related infrastructure and cert.subject lookalike patterns for domains the kit redeploys with cosmetic changes.

Do I need an API key to automate urlscan.io searches?

Public search via the web UI has limits, but the API requires a key. A free API key unlocks /api/v1/search/ with rate limits based on your account tier; paid tiers add higher quotas and bulk capabilities.

Can urlscan.io results be used for takedown evidence?

Yes. Screenshots, timestamps, HTTP transaction logs, and infrastructure data constitute solid supporting evidence for abuse reports to registrars, hosting providers, and CERTs. Reference permanent scan URLs directly in your reports.

What’s the difference between page.domain and page.domains searches?

page.domain: matches the registered domain of the scanned page itself. page.domains: captures all domains observed in the page’s resources, links, and redirects—making it essential for uncovering backend and call-home infrastructure during campaign mapping.

Hmmnm
Published by Hmmnm

Hands-on cybersecurity tutorials, CVE breakdowns, and guided learning paths — written and lab-tested by the Hmmnm team.

🛡️ Hmmnm also delivers this expertise as a service — security testing, assessment & training.
Keep going — the structured way
This post is one step. The learning paths chain the next ones for you, with progress tracking and no account needed.
Follow a learning path →

Prabhu Kalyan Samal

Application Security Consultant at TCS. Certifications: CompTIA SecurityX, Burp Suite Certified Practitioner, Azure Security Engineer, Azure AI Engineer, Certified Red Team Operator, eWPTX v3, LPT, CompTIA PenTest+, Professional Cloud Security Engineer, SC-900, SC-200, PSPO I, CEH, Oracle Java SE 8, ISP, Six Sigma Green Belt, DELF, AutoCAD. Writing about ethical hacking, security tutorials, and tech education at Hmmnm.