In Open Source Intelligence (OSINT), the current state of a website tells only a fraction of the story.

Websites are constantly updated, redesigned, defaced or deleted — but every change leaves a digital footprint that can be tracked down by an expert.

/

By leveraging web archives, search engine caches, staging subdomains, and advanced API queries, an investigator can recover deleted content, track a target's digital evolution, and uncover sensitive data that no longer exists on the live site.

This guide takes you from a complete beginner (clicking through the Wayback Machine) to an advanced practitioner (automating CDX API pulls, hunting staging leaks, and analyzing archived JavaScript).

By the end, you will have a self-contained toolkit that eliminates the need to consult any other resource.

1. Web Archives — Where the Past Lives

The Wayback Machine is the most comprehensive public archive, but it is far from the only one. A complete OSINT workflow uses multiple archives in sequence because each one captures different content, at different times, with different crawler behaviour.

1a. Wayback Machine (Internet Archive) - The Primary Source

Key Points:

  • Calendar view: Paste any URL and a visual timeline shows exactly when snapshots were captured. Dense dot clusters = frequent crawling; gaps = content was likely added/removed during that window.

  • Quick URL patterns (use these directly in your browser):

  • Compare tool: The Internet Archive has a built-in Compare feature that highlights differences between two snapshots side-by-side. Extremely useful for tracking policy changes, retractions, or altered press releases.

  • Outlinks: On any archived page, the top-right panel lists URLs that were linked from that page at the time of capture — including URLs that no longer exist.

  • Save Page Now: Use this to archive a target's current page as a baseline before it changes.

  • Limitations: JavaScript-heavy SPAs (React, Angular, Vue) often capture only a blank shell. Sites that block archive.org via robots.txt will have gaps.

1b. Archive.today (archive.ph) — The Gap-Filler

Archive.today is an independent, on-demand archive that often has snapshots the Wayback Machine missed, and vice versa.

Key Points:

  • Critical UI tip: The site has two input boxes. The dark grey box is for searching existing snapshots. The red box is for submitting a new page to archive. Beginners frequently use the wrong one and waste time.

  • Search syntax:

  • When to use: When the Wayback Machine returns no results or has a gap in the exact date range you need. Also useful for pages that block the archive.org crawler but were manually submitted to Archive.today by someone else.

  • Caveat: Archive.today is slower and less reliable than the Wayback Machine for bulk queries. It's a supplement, not a replacement.

1c. Memento Time Travel — The Multi-Archive Aggregator

Memento Time Travel (timetravel.mementoweb.org) is not an archive itself — it is a search aggregator that queries multiple web archives simultaneously, including the Wayback Machine, Archive.today, national archives, and library collections.

Key Points:

  • Enter a single URL and it returns results from every connected archive in one view.

  • This is the fastest way to find a snapshot that exists in some archive but not the Wayback Machine.

  • When to use: As your second check after the Wayback Machine returns nothing. One lookup covers sources you would never think to check individually.

1d. urlscan.io — Historical Scan Database

urlscan.io is not a traditional archive, but its public scan database is a goldmine for OSINT. Every time anyone submits a URL to urlscan.io, the full DOM, all subresources, and all network requests are captured and stored.

Key Points:

  • Search syntax:

  • Why it matters: Historical scans reveal old endpoints, subdomains, internal URLs, and API paths that are no longer linked on the live site.

  • Passive by nature: You are reading other people's scan data — you never interact with the target infrastructure.

  • Free tier: 100 scans/day without an API key.

Perma.cc creates permanent, tamper-evident links to web pages. It is built for legal proceedings, academic research, and journalism where a page must remain retrievable for years.

Key Points:

  • Free for academic institutions, courts, and libraries. Individual free tier allows 10 permanent links.

  • If you need to cite an archived page in a report or legal document, Perma.cc provides a stable URL that will not break even if the original page is deleted.

  • When to use: Not for discovery — for preservation and citation of content you've already found in other archives.

1f. National & Institutional Archives

Several governments and institutions maintain their own web archives:

Archive

Coverage

URL

UK Web Archive (British Library)

1B+ resources from .uk domains since 2013

Library of Congress Web Archives

US government & cultural sites since 2000

Bibliothèque nationale de France

French web content (legal deposit)

Stanford Web Archive

Stanford University Libraries collections

These are niche but can contain snapshots of government, academic, or regional sites that the Wayback Machine never crawled

1g. Common Crawl — The Raw Data Layer

Common Crawl publishes massive, open datasets of web crawls (billions of URLs) on a monthly basis. It is not a browseable archive — you query it via S3 buckets or the Common Crawl Index.

Key Points:

  • Use the CC-MAIN index to search for URLs across all crawls.

  • Best for large-scale, programmatic analysis (e.g., "find all PDFs ever published on target.com").

  • Overkill for single-page lookups — use the Wayback Machine or Archive.today for those.

2. Search Engine Caches — The Quick-Check Layer

Search engines crawl and store cached copies of pages that can differ from both the live site and the Wayback Machine.

Key Points:

  • Bing is the most reliable cache source today. Type site:target.com in Bing, right-click a result, and select "Cached". This often shows a version newer than the latest Wayback snapshot.

  • Google has progressively reduced direct cache access (the classic cache: operator is largely deprecated as of 2024), but cached content may still surface through third-party tools.

  • When to use: When you suspect a page was modified between two Wayback crawls. The cache often fills that gap.

3. Deleted Pages & 404 Hunting - The "Ghost URL" Technique

This is one of the highest-value OSINT moves: a URL that returns 404 on the live site may still be fully archived.

Key Points:

  • Manual approach: Guess or enumerate likely URLs (/blog/2022/security-update, /old-portfolio, /admin-notes). Check each in the Wayback Machine.

  • Link-based enumeration: Use the Wayback Machine's "Outlinks" feature (visible in the top-right of any archived page) to discover URLs that existed in the past but are now gone.

  • Brute-force wordlists: For advanced users, run wordlists (e.g., SecLists' Discovery/Web-Content) against the Wayback Machine's CDX API (see Section 4) to find archived paths that no longer resolve live.

  • Why it matters: Deleted blog posts, old team pages, removed PDFs, and abandoned admin panels are goldmines.
    </aside>

4. The CDX API — Precision Archaeology for Advanced Users

The CDX (Chronological Data eXchange) API is the Wayback Machine's backend query interface. It lets you filter snapshots by date range, URL pattern, status code, and MIME type — turning a visual search into a programmatic one.

Key Points:

  • Basic syntax:

  • Date filtering:

  • Status code filtering (find only successful captures):

  • MIME type filtering (find archived PDFs, images, JS files):

  • Wildcard and prefix matching: url=target.com/blog/* narrows to the blog section only.

  • Automation: Pipe CDX output into jq (for JSON) or grep/awk (for text) to filter, sort, and bulk-download snapshots with wget or Python's requests.

  • Rate limits: The API is free but rate-limited (~1 request per 200ms). Add delays in scripts to avoid being throttled.

5. Staging & Development Subdomains - The "Left-Behind" Version

Previous versions of a website are frequently not in any archive at all — they're still live on forgotten subdomains, or they existed only as archived URLs that no tool has surfaced yet. Subdomain discovery is the bridge between "what the live site shows" and "what the target has ever exposed."

The workflow below moves from passive (zero interaction with target) to active (DNS queries) to correlation (cross-referencing sources)

5a. Passive Sources (Zero Interaction)

These sources reveal subdomains without sending a single packet to the target. This is your first and safest layer.

Source

What It Reveals

How to Query

Certificate Transparency Logs

Every SSL cert ever issued for the domain, including subdomains

crt.sh → search %.target.com

All subdomains ever scanned by anyone

domain:*.target.com in urlscan search

SecurityTrails

Historical DNS + subdomain database

securitytrails.com → domain → "Subdomains" tab

VirusTotal

Subdomains linked via shared IP, malware reports, or community data

virustotal.com → domain → "Relations" tab

AlienVault OTX

Threat intelligence community data

otx.alienvault.com → domain search

Wayback Machine / CDX API

Subdomains that appeared in any archived URL

url=*.target.com in CDX API

Common Crawl

Subdomains in raw crawl data

Common Crawl Index API

BuiltWith

Subdomains linked via shared tech stack, analytics tags

builtwith.com → domain → "Relationships"

5b. Tool-Based Passive Enumeration

Instead of manually checking each source, use tools that aggregate all passive sources in one command:

# Subfinder (ProjectDiscovery) — fastest, 40+ passive sources
subfinder -d target.com -all -o subfinder.txt

# Amass (OWASP) — deeper, 87+ sources, slower
amass enum -passive -d target.com -o amass.txt

# Assetfinder (Tomnomnom) — lightweight, good for quick checks
assetfinder --subs-only target.com > assetfinder.txt

# theHarvester — also pulls emails, hosts, DNS records
theHarvester -d target.com -b all -f results.html

# BBoT (BlackLanternSecurity) — newer-gen, correlates infrastructure
bbot -d target.com -o bbot.json

Merge and deduplicate:

cat subfinder.txt amass.txt assetfinder.txt | sort -u > all_passive.txt

5c. Active DNS Enumeration

Once you have a passive list, you can go one step further with active DNS queries (these do interact with the target's DNS, but do not touch any web service):

# DNS brute-force with a wordlist
puredns bruteforce wordlist.txt target.com -r resolvers.txt -w brute.txt

# Or with Amass
amass enum -active -d target.com -o amass_active.txt

# Or with Gobuster
gobuster dns -d target.com -w wordlist.txt -o gobuster.txt

Permutation-based discovery (generates subdomain variations from known ones):

# dnsgen — generates permutations from a seed list
dnsgen -d target.com -w known_subs.txt -o permutations.txt

# Resolve the permutations
dnsx -l permutations.txt -o live_permutations.txt

Reverse DNS (PTR) mapping — if you know the target's IP range:

dnsx -l ip_range.txt -reverse -o ptr_results.txt

5d. Wayback Machine as a Subdomain Source

This is a highly underused technique. The Wayback Machine has archived URLs for subdomains that no longer resolve in DNS. Tools like waybackurls and gau (GetAllURLs) extract these:

# waybackurls — pulls all URLs the Wayback Machine knows about
echo "target.com" | waybackurls > wayback_urls.txt

# gau — pulls from Wayback Machine + Common Crawl + AlienVault OTX + URLScan
gau --subs target.com > gau_urls.txt

# Extract unique subdomains from the URLs
cat wayback_urls.txt gau_urls.txt | sort -u | \
  awk -F/ '{print $3}' | sort -u > archived_subdomains.txt

This often reveals subdomains that are no longer in DNS — dead, but their archived content is still accessible via the Wayback Machine.

5e. JavaScript & DOM Recon

Modern web apps embed subdomains in JavaScript files, API calls, and DOM elements that are invisible to passive DNS enumeration.

Key Points:

  • Use Katana (ProjectDiscovery) or Gau to crawl the live site and extract subdomains from JS:

  • Burp Suite passive crawling also collects subdomains from HTTP responses and DOM.

  • Look for patterns in JS: api., cdn., assets., auth., ws. (WebSocket), grpc.

5f. ASN-Based Discovery

If the target owns their own IP range (ASN), you can enumerate all domains pointing to that range:

# Find the ASN for the target
whois target.com | grep -i "origin\|AS"

# Enumerate all domains on that ASN (via SecurityTrails, Censys, or Shodan)

This is extremely powerful for large organizations (banks, governments, universities) that own /16 or /8 blocks.

5g. The Full Automated Workflow

Here's a production-ready pipeline that chains everything together:

# 1. Passive discovery (all sources)
subfinder -d target.com -all -o 01_passive.txt
amass enum -passive -d target.com -o 02_amass.txt
echo "target.com" | waybackurls > 03_wayback.txt
gau --subs target.com >> 03_wayback.txt

# 2. Extract subdomains from URLs
cat 01_passive.txt 02_amass.txt 03_wayback.txt | \
  awk -F/ '{print $3}' | sort -u > 04_all_subs.txt

# 3. Active brute-force
puredns bruteforce wordlist.txt target.com -r resolvers.txt -w 05_brute.txt

# 4. Merge everything
cat 04_all_subs.txt 05_brute.txt | sort -u > 06_merged.txt

# 5. DNS resolve (filter dead subdomains)
dnsx -l 06_merged.txt -o 07_live.txt

# 6. HTTP probe (find which live subdomains actually respond)
httpx -l 07_live.txt -title -tech-detect -o 08_final.txt

# 7. Screenshot everything (optional, for visual triage)
gowitness file -f 08_final.txt
  • Passive enumeration (crt.sh, urlscan, SecurityTrails, Wayback): Always safe. You are reading public data.

  • Active DNS queries (dig, puredns, gobuster): Generally safe. You are querying public DNS.

  • HTTP probing (httpx, gowitness): Grey area. You are sending requests to the target. For OSINT purposes, a single GET request to check if a subdomain is live is widely accepted. Do not run vulnerability scanners, fuzzers, or brute-force login forms.

  • Staging/development sites: Passive observation only. Do not attempt logins, form submissions, or active scanning. This can constitute unauthorized access under computer misuse laws (IT Act 2000 in India, CFAA in the US, Computer Misuse Act 1990 in the UK).

Real-World & Imaginary Scenarios

Scenario 1: The Deleted Security Post

An OSINT analyst is researching a fintech company that recently rebranded. During due diligence, they recall reading a blog post two years ago where the company's CTO discussed a specific vulnerability in their payment gateway. The live site no longer has the post — the entire /blog section was removed during the redesign.

The investigation:

  1. The analyst goes to the Wayback Machine and searches fintechcompany.com/blog/*.

  2. The calendar view shows heavy crawling in March 2024 — right after the post was published.

  3. They click the 2024-03-15 snapshot and recover the full post, including the CTO's description of the vulnerability and a reference to an internal ticket number.

  4. Using the CDX API, they query url=fintechcompany.com/blog/*&from=2024&to=2024&mime=text/html and discover three other deleted posts, one of which mentions a partnership that was later quietly terminated.

  5. Result: The analyst now has a documented history of the company's security posture and business relationships that are no longer publicly visible — critical intelligence for a risk assessment report.

Scenario 2: The Staging Site That Was Never Takedown

A threat intelligence team is tracking a cybercrime group that operates through a series of scam websites. The group's current site is a polished clone of a legitimate bank. The team needs to understand the group's infrastructure history to attribute the operation.

The investigation:

  1. Passive reconnaissance with crt.sh reveals a certificate for test.scamdomain.com issued 8 months before the live site went up.

  2. The team checks the Wayback Machine for test.scamdomain.com — no snapshots exist (the group likely blocked crawlers).

  3. They directly visit test.scamdomain.com in a disposable browser profile (no personal IP, no login). The site loads — it's an older, less polished version of the scam page.

  4. In the page source, they find:

    • An old WordPress version number (4.9.8, with known CVEs)

    • A hardcoded Telegram bot API token (used for victim data exfiltration)

    • An internal comment: <!-- TODO: fix payment gateway before launch -->

  5. Result: The Telegram token, even if rotated, provides a link to the group's communication channel and the WordPress version pinpoints the timeline of the operation's development. This is attribution-grade intelligence that no live-site analysis could have provided.

Resources & Tools

Tool / Resource

Link

Purpose

Access

Wayback Machine

Primary archive for viewing historical snapshots

Free

Save Page Now

Archive a target's current page as a baseline

Free

CDX API

Advanced querying by date, URL, MIME, status

Free, rate-limited

CDX API Docs

Full parameter reference

Free

On-demand + search archive; gap-filler for Wayback

Free

Memento Time Travel

Multi-archive aggregator search

Free

Permanent, tamper-evident links for legal/academic citation

Free tier + Paid

Historical scan database; subdomain + endpoint discovery

Free (100 scans/day)

Common Crawl

Large-scale raw web crawl datasets

Free (S3)

UK Web Archive

UK national web archive

Free

Library of Congress

US government & cultural web archives

Free

Certificate transparency log search

Free

SecurityTrails

Historical DNS, IP, WHOIS & subdomain data

Free tier + Paid

VirusTotal

Domain relations, subdomains, malware flags

Free

AlienVault OTX

Threat intel community data

Free

BuiltWith

Tech stack + linked domain discovery

Free tier + Paid

Subfinder

Passive subdomain enumeration (40+ sources)

Free, open-source

Amass

Deep subdomain enumeration (87+ sources)

Free, open-source

Assetfinder

Lightweight passive subdomain discovery

Free, open-source

theHarvester

Emails, subdomains, hosts from public sources

Free, open-source

BBoT

Newer-gen recon framework with correlation

Free, open-source

waybackurls

Extract all URLs from Wayback Machine for a domain

Free, open-source

gau (GetAllURLs)

Aggregates URLs from Wayback + Common Crawl + OTX + URLScan

Free, open-source

dnsx

Fast DNS resolver for subdomain validation

Free, open-source

httpx

HTTP probing for live subdomain detection

Free, open-source

puredns

High-speed DNS brute-forcer

Free, open-source

dnsgen

Subdomain permutation generator

Free, open-source

Katana

Web crawler for JS/DOM subdomain extraction

Free, open-source

Gowitness

Screenshot tool for visual triage

Free, open-source

SecLists

Wordlists for brute-forcing

Free, open-source

jq

JSON parsing for CDX API output

Free, open-source

Browser DevTools

Built into Chrome/Firefox/Edge (F12)

Inspecting page source, JS, HTML comments

Free

Quick install (Linux/Kali):

# Core subdomain tools
go install -v github.com/projectdiscovery/subfinder/v2/cmd/subfinder@latest
go install -v github.com/owasp-amass/amass/v4@latest
go install -v github.com/projectdiscovery/dnsx/cmd/dnsx@latest
go install -v github.com/projectdiscovery/httpx/cmd/httpx@latest
go install -v github.com/dnhc/puredns/cmd/puredns@latest
go install -v github.com/tomnomnom/waybackurls@latest
go install -v github.com/lc/gau/v2/cmd/gau@latest
go install -v github.com/projectdiscovery/katana/cmd/katana@latest

# Supporting tools
sudo apt install jq
git clone <https://github.com/danielmiessler/SecLists> /opt/SecLists

Limitations

  • JavaScript-heavy sites (SPAs built with React, Angular, Vue) often capture poorly in archives. The Wayback Machine may store only a blank shell. Mitigate by checking if the site has a server-rendered fallback or by querying the CDX API for associated API endpoints.

  • robots.txt exclusions can create gaps in the archive. If a site blocks archive.org, you may need to rely on search engine caches, Archive.today, or staging subdomains instead.

  • Legal considerations: Accessing publicly available archived content is generally low-risk, but storing, redistributing, or using personal data found in archives may be subject to privacy laws (GDPR, India's DPDP Act 2023, etc.). Always assess the legal context of your jurisdiction before acting on findings.

  • Staging sites: Passive observation only. Do not attempt logins, form submissions, or active scanning of staging environments. This can constitute unauthorized access under computer misuse laws in most jurisdictions.

Keep reading