In Open Source Intelligence (OSINT), the current state of a website tells only a fraction of the story.
Websites are constantly updated, redesigned, defaced or deleted — but every change leaves a digital footprint that can be tracked down by an expert.
/
By leveraging web archives, search engine caches, staging subdomains, and advanced API queries, an investigator can recover deleted content, track a target's digital evolution, and uncover sensitive data that no longer exists on the live site.
This guide takes you from a complete beginner (clicking through the Wayback Machine) to an advanced practitioner (automating CDX API pulls, hunting staging leaks, and analyzing archived JavaScript).
By the end, you will have a self-contained toolkit that eliminates the need to consult any other resource.
1. Web Archives — Where the Past Lives
The Wayback Machine is the most comprehensive public archive, but it is far from the only one. A complete OSINT workflow uses multiple archives in sequence because each one captures different content, at different times, with different crawler behaviour.
1a. Wayback Machine (Internet Archive) - The Primary Source
Key Points:
Calendar view: Paste any URL and a visual timeline shows exactly when snapshots were captured. Dense dot clusters = frequent crawling; gaps = content was likely added/removed during that window.
Quick URL patterns (use these directly in your browser):
Compare tool: The Internet Archive has a built-in Compare feature that highlights differences between two snapshots side-by-side. Extremely useful for tracking policy changes, retractions, or altered press releases.
Outlinks: On any archived page, the top-right panel lists URLs that were linked from that page at the time of capture — including URLs that no longer exist.
Save Page Now: Use this to archive a target's current page as a baseline before it changes.
Limitations: JavaScript-heavy SPAs (React, Angular, Vue) often capture only a blank shell. Sites that block
archive.orgviarobots.txtwill have gaps.
1b. Archive.today (archive.ph) — The Gap-Filler
Archive.today is an independent, on-demand archive that often has snapshots the Wayback Machine missed, and vice versa.
Key Points:
Critical UI tip: The site has two input boxes. The dark grey box is for searching existing snapshots. The red box is for submitting a new page to archive. Beginners frequently use the wrong one and waste time.
Search syntax:
When to use: When the Wayback Machine returns no results or has a gap in the exact date range you need. Also useful for pages that block the
archive.orgcrawler but were manually submitted to Archive.today by someone else.Caveat: Archive.today is slower and less reliable than the Wayback Machine for bulk queries. It's a supplement, not a replacement.
1c. Memento Time Travel — The Multi-Archive Aggregator
Memento Time Travel (timetravel.mementoweb.org) is not an archive itself — it is a search aggregator that queries multiple web archives simultaneously, including the Wayback Machine, Archive.today, national archives, and library collections.
Key Points:
Enter a single URL and it returns results from every connected archive in one view.
This is the fastest way to find a snapshot that exists in some archive but not the Wayback Machine.
When to use: As your second check after the Wayback Machine returns nothing. One lookup covers sources you would never think to check individually.
1d. urlscan.io — Historical Scan Database
urlscan.io is not a traditional archive, but its public scan database is a goldmine for OSINT. Every time anyone submits a URL to urlscan.io, the full DOM, all subresources, and all network requests are captured and stored.
Key Points:
Search syntax:
Why it matters: Historical scans reveal old endpoints, subdomains, internal URLs, and API paths that are no longer linked on the live site.
Passive by nature: You are reading other people's scan data — you never interact with the target infrastructure.
Free tier: 100 scans/day without an API key.
1e. Perma.cc — For Legal & Academic Citations
Perma.cc creates permanent, tamper-evident links to web pages. It is built for legal proceedings, academic research, and journalism where a page must remain retrievable for years.
Key Points:
Free for academic institutions, courts, and libraries. Individual free tier allows 10 permanent links.
If you need to cite an archived page in a report or legal document, Perma.cc provides a stable URL that will not break even if the original page is deleted.
When to use: Not for discovery — for preservation and citation of content you've already found in other archives.
1f. National & Institutional Archives
Several governments and institutions maintain their own web archives:
Archive | Coverage | URL |
|---|---|---|
UK Web Archive (British Library) | 1B+ resources from .uk domains since 2013 | |
Library of Congress Web Archives | US government & cultural sites since 2000 | |
Bibliothèque nationale de France | French web content (legal deposit) | |
Stanford Web Archive | Stanford University Libraries collections |
These are niche but can contain snapshots of government, academic, or regional sites that the Wayback Machine never crawled
1g. Common Crawl — The Raw Data Layer
Common Crawl publishes massive, open datasets of web crawls (billions of URLs) on a monthly basis. It is not a browseable archive — you query it via S3 buckets or the Common Crawl Index.
Key Points:
Use the CC-MAIN index to search for URLs across all crawls.
Best for large-scale, programmatic analysis (e.g., "find all PDFs ever published on target.com").
Overkill for single-page lookups — use the Wayback Machine or Archive.today for those.
2. Search Engine Caches — The Quick-Check Layer
Search engines crawl and store cached copies of pages that can differ from both the live site and the Wayback Machine.
Key Points:
Bing is the most reliable cache source today. Type
site:target.comin Bing, right-click a result, and select "Cached". This often shows a version newer than the latest Wayback snapshot.Google has progressively reduced direct cache access (the classic
cache:operator is largely deprecated as of 2024), but cached content may still surface through third-party tools.When to use: When you suspect a page was modified between two Wayback crawls. The cache often fills that gap.
3. Deleted Pages & 404 Hunting - The "Ghost URL" Technique
This is one of the highest-value OSINT moves: a URL that returns 404 on the live site may still be fully archived.
Key Points:
Manual approach: Guess or enumerate likely URLs (
/blog/2022/security-update,/old-portfolio,/admin-notes). Check each in the Wayback Machine.Link-based enumeration: Use the Wayback Machine's "Outlinks" feature (visible in the top-right of any archived page) to discover URLs that existed in the past but are now gone.
Brute-force wordlists: For advanced users, run wordlists (e.g., SecLists'
Discovery/Web-Content) against the Wayback Machine's CDX API (see Section 4) to find archived paths that no longer resolve live.Why it matters: Deleted blog posts, old team pages, removed PDFs, and abandoned admin panels are goldmines.
</aside>
4. The CDX API — Precision Archaeology for Advanced Users
The CDX (Chronological Data eXchange) API is the Wayback Machine's backend query interface. It lets you filter snapshots by date range, URL pattern, status code, and MIME type — turning a visual search into a programmatic one.
Key Points:
Basic syntax:
Date filtering:
Status code filtering (find only successful captures):
MIME type filtering (find archived PDFs, images, JS files):
Wildcard and prefix matching:
url=target.com/blog/*narrows to the blog section only.Automation: Pipe CDX output into
jq(for JSON) orgrep/awk(for text) to filter, sort, and bulk-download snapshots withwgetor Python'srequests.Rate limits: The API is free but rate-limited (~1 request per 200ms). Add delays in scripts to avoid being throttled.
5. Staging & Development Subdomains - The "Left-Behind" Version
Previous versions of a website are frequently not in any archive at all — they're still live on forgotten subdomains, or they existed only as archived URLs that no tool has surfaced yet. Subdomain discovery is the bridge between "what the live site shows" and "what the target has ever exposed."
The workflow below moves from passive (zero interaction with target) to active (DNS queries) to correlation (cross-referencing sources)
5a. Passive Sources (Zero Interaction)
These sources reveal subdomains without sending a single packet to the target. This is your first and safest layer.
Source | What It Reveals | How to Query |
Certificate Transparency Logs | Every SSL cert ever issued for the domain, including subdomains | crt.sh → search |
All subdomains ever scanned by anyone |
| |
SecurityTrails | Historical DNS + subdomain database | securitytrails.com → domain → "Subdomains" tab |
VirusTotal | Subdomains linked via shared IP, malware reports, or community data | virustotal.com → domain → "Relations" tab |
AlienVault OTX | Threat intelligence community data | otx.alienvault.com → domain search |
Wayback Machine / CDX API | Subdomains that appeared in any archived URL |
|
Common Crawl | Subdomains in raw crawl data | Common Crawl Index API |
BuiltWith | Subdomains linked via shared tech stack, analytics tags | builtwith.com → domain → "Relationships" |
5b. Tool-Based Passive Enumeration
Instead of manually checking each source, use tools that aggregate all passive sources in one command:
# Subfinder (ProjectDiscovery) — fastest, 40+ passive sources
subfinder -d target.com -all -o subfinder.txt
# Amass (OWASP) — deeper, 87+ sources, slower
amass enum -passive -d target.com -o amass.txt
# Assetfinder (Tomnomnom) — lightweight, good for quick checks
assetfinder --subs-only target.com > assetfinder.txt
# theHarvester — also pulls emails, hosts, DNS records
theHarvester -d target.com -b all -f results.html
# BBoT (BlackLanternSecurity) — newer-gen, correlates infrastructure
bbot -d target.com -o bbot.json
Merge and deduplicate:
cat subfinder.txt amass.txt assetfinder.txt | sort -u > all_passive.txt
5c. Active DNS Enumeration
Once you have a passive list, you can go one step further with active DNS queries (these do interact with the target's DNS, but do not touch any web service):
# DNS brute-force with a wordlist
puredns bruteforce wordlist.txt target.com -r resolvers.txt -w brute.txt
# Or with Amass
amass enum -active -d target.com -o amass_active.txt
# Or with Gobuster
gobuster dns -d target.com -w wordlist.txt -o gobuster.txt
Permutation-based discovery (generates subdomain variations from known ones):
# dnsgen — generates permutations from a seed list
dnsgen -d target.com -w known_subs.txt -o permutations.txt
# Resolve the permutations
dnsx -l permutations.txt -o live_permutations.txt
Reverse DNS (PTR) mapping — if you know the target's IP range:
dnsx -l ip_range.txt -reverse -o ptr_results.txt
5d. Wayback Machine as a Subdomain Source
This is a highly underused technique. The Wayback Machine has archived URLs for subdomains that no longer resolve in DNS. Tools like waybackurls and gau (GetAllURLs) extract these:
# waybackurls — pulls all URLs the Wayback Machine knows about
echo "target.com" | waybackurls > wayback_urls.txt
# gau — pulls from Wayback Machine + Common Crawl + AlienVault OTX + URLScan
gau --subs target.com > gau_urls.txt
# Extract unique subdomains from the URLs
cat wayback_urls.txt gau_urls.txt | sort -u | \
awk -F/ '{print $3}' | sort -u > archived_subdomains.txt
This often reveals subdomains that are no longer in DNS — dead, but their archived content is still accessible via the Wayback Machine.
5e. JavaScript & DOM Recon
Modern web apps embed subdomains in JavaScript files, API calls, and DOM elements that are invisible to passive DNS enumeration.
Key Points:
Use Katana (ProjectDiscovery) or Gau to crawl the live site and extract subdomains from JS:
Burp Suite passive crawling also collects subdomains from HTTP responses and DOM.
Look for patterns in JS:
api.,cdn.,assets.,auth.,ws.(WebSocket),grpc.
5f. ASN-Based Discovery
If the target owns their own IP range (ASN), you can enumerate all domains pointing to that range:
# Find the ASN for the target
whois target.com | grep -i "origin\|AS"
# Enumerate all domains on that ASN (via SecurityTrails, Censys, or Shodan)
This is extremely powerful for large organizations (banks, governments, universities) that own /16 or /8 blocks.
5g. The Full Automated Workflow
Here's a production-ready pipeline that chains everything together:
# 1. Passive discovery (all sources)
subfinder -d target.com -all -o 01_passive.txt
amass enum -passive -d target.com -o 02_amass.txt
echo "target.com" | waybackurls > 03_wayback.txt
gau --subs target.com >> 03_wayback.txt
# 2. Extract subdomains from URLs
cat 01_passive.txt 02_amass.txt 03_wayback.txt | \
awk -F/ '{print $3}' | sort -u > 04_all_subs.txt
# 3. Active brute-force
puredns bruteforce wordlist.txt target.com -r resolvers.txt -w 05_brute.txt
# 4. Merge everything
cat 04_all_subs.txt 05_brute.txt | sort -u > 06_merged.txt
# 5. DNS resolve (filter dead subdomains)
dnsx -l 06_merged.txt -o 07_live.txt
# 6. HTTP probe (find which live subdomains actually respond)
httpx -l 07_live.txt -title -tech-detect -o 08_final.txt
# 7. Screenshot everything (optional, for visual triage)
gowitness file -f 08_final.txt
5h. Ethical & Legal Boundary
Passive enumeration (crt.sh, urlscan, SecurityTrails, Wayback): Always safe. You are reading public data.
Active DNS queries (dig, puredns, gobuster): Generally safe. You are querying public DNS.
HTTP probing (httpx, gowitness): Grey area. You are sending requests to the target. For OSINT purposes, a single GET request to check if a subdomain is live is widely accepted. Do not run vulnerability scanners, fuzzers, or brute-force login forms.
Staging/development sites: Passive observation only. Do not attempt logins, form submissions, or active scanning. This can constitute unauthorized access under computer misuse laws (IT Act 2000 in India, CFAA in the US, Computer Misuse Act 1990 in the UK).
Real-World & Imaginary Scenarios
Scenario 1: The Deleted Security Post
An OSINT analyst is researching a fintech company that recently rebranded. During due diligence, they recall reading a blog post two years ago where the company's CTO discussed a specific vulnerability in their payment gateway. The live site no longer has the post — the entire /blog section was removed during the redesign.
The investigation:
The analyst goes to the Wayback Machine and searches
fintechcompany.com/blog/*.The calendar view shows heavy crawling in March 2024 — right after the post was published.
They click the 2024-03-15 snapshot and recover the full post, including the CTO's description of the vulnerability and a reference to an internal ticket number.
Using the CDX API, they query
url=fintechcompany.com/blog/*&from=2024&to=2024&mime=text/htmland discover three other deleted posts, one of which mentions a partnership that was later quietly terminated.Result: The analyst now has a documented history of the company's security posture and business relationships that are no longer publicly visible — critical intelligence for a risk assessment report.
Scenario 2: The Staging Site That Was Never Takedown
A threat intelligence team is tracking a cybercrime group that operates through a series of scam websites. The group's current site is a polished clone of a legitimate bank. The team needs to understand the group's infrastructure history to attribute the operation.
The investigation:
Passive reconnaissance with crt.sh reveals a certificate for
test.scamdomain.comissued 8 months before the live site went up.The team checks the Wayback Machine for
test.scamdomain.com— no snapshots exist (the group likely blocked crawlers).They directly visit
test.scamdomain.comin a disposable browser profile (no personal IP, no login). The site loads — it's an older, less polished version of the scam page.In the page source, they find:
An old WordPress version number (4.9.8, with known CVEs)
A hardcoded Telegram bot API token (used for victim data exfiltration)
An internal comment:
<!-- TODO: fix payment gateway before launch -->
Result: The Telegram token, even if rotated, provides a link to the group's communication channel and the WordPress version pinpoints the timeline of the operation's development. This is attribution-grade intelligence that no live-site analysis could have provided.
Resources & Tools
Tool / Resource | Link | Purpose | Access |
Wayback Machine | Primary archive for viewing historical snapshots | Free | |
Save Page Now | Archive a target's current page as a baseline | Free | |
CDX API | Advanced querying by date, URL, MIME, status | Free, rate-limited | |
CDX API Docs | Full parameter reference | Free | |
On-demand + search archive; gap-filler for Wayback | Free | ||
Memento Time Travel | Multi-archive aggregator search | Free | |
Permanent, tamper-evident links for legal/academic citation | Free tier + Paid | ||
Historical scan database; subdomain + endpoint discovery | Free (100 scans/day) | ||
Common Crawl | Large-scale raw web crawl datasets | Free (S3) | |
UK Web Archive | UK national web archive | Free | |
Library of Congress | US government & cultural web archives | Free | |
Certificate transparency log search | Free | ||
SecurityTrails | Historical DNS, IP, WHOIS & subdomain data | Free tier + Paid | |
VirusTotal | Domain relations, subdomains, malware flags | Free | |
AlienVault OTX | Threat intel community data | Free | |
BuiltWith | Tech stack + linked domain discovery | Free tier + Paid | |
Subfinder | Passive subdomain enumeration (40+ sources) | Free, open-source | |
Amass | Deep subdomain enumeration (87+ sources) | Free, open-source | |
Assetfinder | Lightweight passive subdomain discovery | Free, open-source | |
theHarvester | Emails, subdomains, hosts from public sources | Free, open-source | |
BBoT | Newer-gen recon framework with correlation | Free, open-source | |
waybackurls | Extract all URLs from Wayback Machine for a domain | Free, open-source | |
gau (GetAllURLs) | Aggregates URLs from Wayback + Common Crawl + OTX + URLScan | Free, open-source | |
dnsx | Fast DNS resolver for subdomain validation | Free, open-source | |
httpx | HTTP probing for live subdomain detection | Free, open-source | |
puredns | High-speed DNS brute-forcer | Free, open-source | |
dnsgen | Subdomain permutation generator | Free, open-source | |
Katana | Web crawler for JS/DOM subdomain extraction | Free, open-source | |
Gowitness | Screenshot tool for visual triage | Free, open-source | |
SecLists | Wordlists for brute-forcing | Free, open-source | |
jq | JSON parsing for CDX API output | Free, open-source | |
Browser DevTools | Built into Chrome/Firefox/Edge ( | Inspecting page source, JS, HTML comments | Free |
Quick install (Linux/Kali):
# Core subdomain tools
go install -v github.com/projectdiscovery/subfinder/v2/cmd/subfinder@latest
go install -v github.com/owasp-amass/amass/v4@latest
go install -v github.com/projectdiscovery/dnsx/cmd/dnsx@latest
go install -v github.com/projectdiscovery/httpx/cmd/httpx@latest
go install -v github.com/dnhc/puredns/cmd/puredns@latest
go install -v github.com/tomnomnom/waybackurls@latest
go install -v github.com/lc/gau/v2/cmd/gau@latest
go install -v github.com/projectdiscovery/katana/cmd/katana@latest
# Supporting tools
sudo apt install jq
git clone <https://github.com/danielmiessler/SecLists> /opt/SecLists
Limitations
JavaScript-heavy sites (SPAs built with React, Angular, Vue) often capture poorly in archives. The Wayback Machine may store only a blank shell. Mitigate by checking if the site has a server-rendered fallback or by querying the CDX API for associated API endpoints.
robots.txt exclusions can create gaps in the archive. If a site blocks
archive.org, you may need to rely on search engine caches, Archive.today, or staging subdomains instead.Legal considerations: Accessing publicly available archived content is generally low-risk, but storing, redistributing, or using personal data found in archives may be subject to privacy laws (GDPR, India's DPDP Act 2023, etc.). Always assess the legal context of your jurisdiction before acting on findings.
Staging sites: Passive observation only. Do not attempt logins, form submissions, or active scanning of staging environments. This can constitute unauthorized access under computer misuse laws in most jurisdictions.

