CIA 2010 covert communication websites / Wayback Machine CDX scanning

The Wayback Machine has an endpoint to query cralwed pages called the CDX server. It is documented at: github.com/internetarchive/wayback/blob/master/wayback-cdx-server/README.md.

This allows to filter down 10 thousands of possible domains in a few hours. But 100s of thousands would be too much. This is because you have to query exactly one URL at a time, and they possibly rate limit IPs. But no IP blacklisting so far after several hours, so it's not that bad.

Once you have a heuristic to narrow down some domains, you can use this helper: ../cia-2010-covert-communication-websites/cdx.sh to drill them down from 10s of thousands down to hundreds or thousands.

We then post process the results of cdx.sh with ../cia-2010-covert-communication-websites/cdx-post.sh to drill them down from from thousands to dozens, and manually inspect everything.

From then on, you can just manually inspect for hist on your browser.

Table of contents
- Wayback Machine CDX scanning with Tor parallelization Wayback Machine CDX scanning
- JS CDX scanning Wayback Machine CDX scanning

Wayback Machine CDX scanning with Tor parallelization

 0  0

Dire times require dire methods: ../cia-2010-covert-communication-websites/cdx-tor.sh.

First we must start the tor servers with the tor-army command from: stackoverflow.com/questions/14321214/how-to-run-multiple-tor-processes-at-once-with-different-exit-ips/76749983#76749983

tor-army 100

and then use it on a newline separated domain name list to check;

./cdx-tor.sh infile.txt

This creates a directory infile.txt.cdx/ containing:

infile.txt.cdx/out00, out01, etc.: the suspected CDX lines from domains from each tor instance based on the simple criteria that the CDX can handle directly. We split the input domains into 100 piles, and give one selected pile per tor instance.
infile.txt.cdx/out: the final combined CDX output of out00, out01, ...
infile.txt.cdx/out.post: the final output containing only domain names that match further CLI criteria that cannot be easily encoded on the CDX query. This is the cleanest domain name list you should look into at the end basically.

Since archive is so abysmal in its data access, e.g. a Google BigQuery would solve our issues in seconds, we have to come up with creative ways of getting around their IP throttling.

The CIA doesn't play fair. They're actually the exact opposite of fair. So neither shall we.

Distilled into an answer at: stackoverflow.com/questions/14321214/how-to-run-multiple-tor-processes-at-once-with-different-exit-ips/76749983#76749983

This should allow a full sweep of the 4.5M records in 2013 DNS Census virtual host cleanup in a reasonable amount of time. After JAR/SWF/CGI filtering we obtained 5.8k domains, so a reduction factor of about 1 million with likely very few losses. Not bad.

5.8k is still a bit annoying to fully go over however, so we can also try to count CDX hits to the domains and remove anything with too many hits, since the CIA websites basically have very few archives:

cd 2013-dns-census-a-novirt-domains.txt.cdx
./cdx-tor.sh -d out.post domain-list.txt
cd out.post.cdx
cut -d' ' -f1 out | uniq -c | sort -k1 -n | awk 'match($2, /([^,]+),([^)]+)/, a) {printf("%s.%s %d\n", a[2], a[1], $1)}' > out.count

This gives us something like:

12654montana.com 1
aeronet-news.com 1
atohms.com 1
av3net.com 1
beechstreetas400.com 1

sorted by increasing hit counts, so we can go down as far as patience allows for!

New results from a full CDX scan of 2013-dns-census-a-novirt.csv:

219.90.61.123 journeystravelled.com

JS CDX scanning

 0  0

JAR, SWF and CGI-bin scanning by path only is fine, since there are relatively few of those. But .js scanning by path only is too broad.

One option would be to filter out by size, an information that is contained on the CDX. Let's check typical ones:

grep -f <(jq -r '.[]|select(select(.comms)|.comms|test("\\.js"))|.host' ../media/cia-2010-covert-communication-websites/hits.json) out | out.jshits.cdx
sort -n -k7 out.jshits.cdx

Ignoring some obvious unrelated non-comms files visually we get a range of about 2732 to 3632:

net,hollywoodscreen)/current.js 20110106082232 http://hollywoodscreen.net/current.js text/javascript 200 XY5NHVW7UMFS3WSKPXLOQ5DJA34POXMV 2732
com,amishkanews)/amishkanewss.js 20110208032713 http://amishkanews.com/amishkanewss.js text/javascript 200 S5ZWJ53JFSLUSJVXBBA3NBJXNYLNCI4E 3632

This ignores the obviously atypical JavaScript with SHAs from iranfootballsource, and the particularly small old menu.js from cutabovenews.com, which we embed into ../cia-2010-covert-communication-websites/cdx-post-js.sh.

The size helps a bit, but it's not insanely good unfortunately, only about 3x, these are some common JS sizes right there!

CIA 2010 covert communication websites / Wayback Machine CDX scanning

Wayback Machine CDX scanning with Tor parallelization

JS CDX scanning

 Ancestors (15)

 Incoming links (7)

 Discussion (0)

 Articles by others on the same topic (0)

CIA 2010 covert communication websites / Wayback Machine CDX scanning

Wayback Machine CDX scanning with Tor parallelization

JS CDX scanning

 Ancestors (15)

 Incoming links (7)

 Discussion (0)  Subscribe (1)

 Articles by others on the same topic (0)

 Discussion (0)