Sunday, November 12, 2006

Heritrix Activity Report

Heritrix isn't being adopted at the same rapid pace as Nutch is, but it keeps showing up from more and more places.

Here's the list of sightings, but the one that gives me the biggest giggle is the first, which claims to be "google.com" that came from Mannheim University in Germany.

134.155.241.9 "Mozilla/5.0 (compatible; heritrix/1.10.0 +http://google.com)"

137.82.84.97 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.worio.com/)"

137.82.84.97 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.worio.com/)"

152.163.214.140 "Mozilla/5.0 (compatible; heritrix/1.8.0
+http://wiki.office.aol.com/wiki/SEO)"

152.163.214.141 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://wiki.office.aol.com/wiki/SEO)"

152.163.214.144 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://wiki.office.aol.com/wiki/SEO)"

193.40.192.35 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://erika.nlib.ee)"

195.39.35.118 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.researcher.cz)"

198.162.51.70 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.worio.com/)"

207.241.233.35 "Mozilla/5.0 (compatible;archive.org_bot/heritrix-1.9.0-200608171144 +http://pandora.nla.gov.au/crawl.html)"

209.128.119.17 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://innovationblog.com)"

209.128.119.46 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://innovationblog.com)"

216.182.228.85 "Mozilla/5.0 (compatible; heritrix/1.4.0 +http://www.hanzoweb.com)"

217.91.71.203 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.schluetersche.de)"

24.8.197.68 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://crawlerx51.com)"

67.162.138.161 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://crawlerx51.com)"

71.229.152.72 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://crawlerx51.com)"

71.56.215.150 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://crawlerx51.com)"

72.20.99.46 "Mozilla/5.0 (compatible; heritrix/1.8.0 +http://www.accelobot.com)"

87.98.198.194 "Mozilla/5.0 (compatible; heritrix/1.4.0 +http://www.hanzoweb.com)"
The other one I found amusing was the Accelobot which claims to "help automate market research" and I wonder if their research showed them I wasn't interested in their help?

Not nearly as popular as other tools, but picking up a little steam unfortunately.

We'll keep an eye on this and let people know when it hits epidemic proportions.

Tracking HTTrack Website Downloader

I'm just curious why over 100 people in the last few months thought they could just download my whole website (not this blog) with HTTrack?

What were these dumb fucks going to do with it once they got it anyway?

  • Run a scraper script on the results?
  • Blatantly republish the content with their own template?
  • Run some data mining scripts on it?
  • Keep a copy just for shits and giggles?
Who the fuck knows, who the fuck cares, they aren't downloading 40K pages so they got NADA!

Here's a list of attempts to download the site:
12.218.132.246 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
151.196.39.206 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
151.44.39.130 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
151.57.203.117 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
157.150.112.6 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
160.75.107.93 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
166.102.234.113 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
168.209.97.34 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
193.194.84.227 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
193.253.222.153 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
194.138.39.53 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
194.51.93.106 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
194.57.91.165 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
195.115.20.132 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
195.229.242.53 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
195.246.48.241 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
196.1.179.77 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
196.30.245.149 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
196.31.142.11 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
200.170.96.119 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
201.0.55.48 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
202.147.168.130 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
202.58.205.163 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
202.65.119.252 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
202.83.173.59 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
202.90.87.7 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
203.189.231.13 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
203.87.188.194 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
206.223.8.30 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
208.102.27.19 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
208.255.142.57 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
208.255.142.57 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.117.81.45 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.129.60.250 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.200.201.102 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.200.201.214 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.200.203.48 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.251.8.5 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.81.218.82 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
212.93.224.35 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
213.136.106.252 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
213.216.199.2 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
213.228.0.86 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
213.23.124.2 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
216.108.210.225 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
216.76.80.93 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
220.247.221.131 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
24.205.6.210 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
61.90.220.86 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
62.210.102.125 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
62.57.32.142 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
64.222.233.72 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
68.220.248.94 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
69.22.0.123 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
69.88.8.6 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
69.88.8.6 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
70.71.114.43 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
71.227.195.118 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
71.70.233.219 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
72.255.6.100 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
74.132.128.2 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.103.33.75 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.144.203.67 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.144.234.32 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.170.26.10 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.170.39.87 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.191.116.41 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
80.53.155.234 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.208.36.91 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.245.178.4 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.246.203.43 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.250.148.63 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.29.232.56 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.50.176.143 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.56.85.53 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
81.90.175.201 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.16.147.149 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.225.167.110 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.228.167.150 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.239.139.105 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.242.65.70 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.245.61.27 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.248.45.214 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.65.0.229 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.66.135.81 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.83.202.247 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.93.27.229 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
82.93.27.229 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.135.199.34 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.135.224.26 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.16.51.174 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.179.163.75 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.93.133.158 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
83.93.133.158 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.162.79.29 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.245.166.176 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.4.209.62 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.6.122.9 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.72.193.77 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
84.90.2.1 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
86.195.214.61 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
86.68.132.131 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
87.218.59.4 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
87.81.178.38 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
87.89.114.228 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
88.139.139.203 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
88.73.106.129 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
89.54.130.12 "Mozilla/4.5 (compatible; HTTrack 3.0x; Windows 98)"
The best part is, trying to download my site from my server gets them all automatically banned.

Greed will get you nowhere, not on my site anyway!

Here a Nutch, There a Nutch, Everywhere a Nutch Nutch

Nutch usage seems to be breeding faster than cousins in Kentucky so I figured it was time to post a sequel to the original How Much Nutch is Too Much Nutch.

Here's a complete breakdown on every IP that I've seen using Nutch with the actual word Nutch in the user agent for a grand total of 190 IP's crawling to date. Several of them like Cazoodle, MQBOT, and a few .EDU's are crawling from a block of IPs but the majority seem to be scattered all over the place.

Here's the list of all the creepy crawling Nutches:

124.32.246.36 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

124.32.246.45 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

128.208.3.173 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; raphael@unterreuth.de)

128.208.6.125 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.200 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.207 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.226 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.227 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.232 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.75 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.208.6.77 NutchCVS/0.8-dev (Nutch running at UW; http://www.nutch.org/docs/en/bot.html; sycrawl@cs.washington.edu)

128.95.1.189 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

128.97.88.68 ilial/Nutch-0.9-dev

128.97.88.70 ilial/Nutch-0.9-dev

129.242.19.138 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

129.34.20.19 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

129.78.64.106 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.137.86 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.202 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.205 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.206 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.211 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.212 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

13.1.139.213 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

131.112.125.102 asked/Nutch-0.8 (web crawler; http://asked.jp; epicurus at gmail dot com)

131.112.125.103 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html;
nutch-agent@lucene.apache.org)

131.112.125.104 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

131.112.125.106 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

131.112.16.220 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

131.211.84.21 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

140.247.62.79 blogsearch/Nutch-0.9-dev

140.247.62.80 blogsearch/Nutch-0.9-dev

147.202.90.2 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

159.226.5.82 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

164.67.195.201 ilial/Nutch-0.9-dev

164.67.195.245 ilial/Nutch-0.9-dev

164.67.195.26 ilial/Nutch-0.9-dev

164.67.195.27 ilial/Nutch-0.9-dev

164.67.195.67 ilial/Nutch-0.9-dev

164.67.195.68 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

164.67.195.86 ilial/Nutch-0.9-dev

166.214.93.76 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

192.17.240.19 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.20 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.41 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.43 MQBOT/Nutch-0.9-dev (MQBOT Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.44 MQBOT/Nutch-0.9-dev (MQBOT Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.46 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.47 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.48 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.52 MQBOT/Nutch-0.9-dev (MQBOT Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.56 MQBOT/Nutch-0.9-dev (MQBOT Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.57 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu;
mqbot@cs.uiuc.edu)

192.17.240.58 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.60 MQBOT/Nutch-0.9-dev (MQBOT Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.71 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.74 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

192.17.240.76 MQBOT/Nutch-0.9-dev (MQBOT Nutch Crawler; http://falcon.cs.uiuc.edu; mqbot@cs.uiuc.edu)

193.145.45.68 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.203.240.117 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.203.240.118 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.203.240.119 HouxouCrawler/0.8-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)

193.203.240.120 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.203.240.121 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.203.240.122 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.252.148.51 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

193.42.229.3 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

195.72.131.70 HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)

195.72.131.72 HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)

195.72.131.73 HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)

195.72.131.80 HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)

203.113.130.205 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

203.147.0.44 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

203.199.83.162 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

203.244.218.1 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

207.176.224.241 Nutch/Nutch-0.8.1

207.176.224.245 Nutch/Nutch-0.8.1

207.214.93.42 MyNutch/V 0.3 (JP's Nutch Test Search Engine; jpnutch at yahoo dot com)

208.64.57.65 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html;
nutch-agent@lucene.apache.org)

210.174.3.130 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

210.196.73.193 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

210.245.31.15 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

210.245.31.18 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

211.152.34.34 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

212.101.97.63 test/Nutch-0.8.1 (test; www.apache.org; test@apache.org)

212.12.114.238 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

212.137.33.140 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

212.156.230.210 BilgiBetaBot/0.8-dev (bilgi.com (Beta) ; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

212.58.116.72 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

213.132.175.101 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

213.157.204.141 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

213.251.133.12 Misterbot-Nutch/0.7.1 (Misterbot-Nutch; http://www.misterbot.fr; nutch at misterbot.fr)

216.182.225.186 NutchEC2Test/Nutch-0.9-dev (Testing Nutch on Amazon EC2.; http://lucene.apache.org/nutch/bot.html; ec2test at lucene.com)

216.182.236.46 NutchEC2Test/Nutch-0.9-dev (Testing Nutch on Amazon EC2.; http://lucene.apache.org/nutch/bot.html; ec2test at lucene.com)

216.182.237.45 NutchEC2Test/Nutch-0.9-dev (Testing Nutch on Amazon EC2.; http://lucene.apache.org/nutch/bot.html; ec2test at lucene.com)

216.93.185.12 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

217.153.59.26 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html;
nutch-agent@lucene.apache.org)

217.31.51.128 Megatext/Nutch-0.8.1 (Beta; http://www.megatext.cz/; microton@microton.cz)

218.25.39.81 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

220.130.191.231 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.232 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.233 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.234 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.235 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.236 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.237 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.238 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.239 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

220.130.191.240 Cazoodle/Nutch-0.9-dev (Cazoodle Nutch Crawler; http://www.cazoodle.com; mqbot@cazoodle.com)

221.114.253.210 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

221.116.237.114 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

221.221.237.35 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

222.173.249.33 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

222.173.249.33 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

24.222.153.250 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

24.6.168.184 test/Nutch-0.8.1 (Test robot; http://test.com; info at test.com>)

58.186.61.164 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

58.187.12.236 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

58.215.74.242 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

58.215.75.2 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

58.87.139.90 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

59.160.240.115 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

59.160.240.116 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

59.160.240.183 Nutch-test/Nutch-0.9-dev

59.160.240.184 Nutch-test/Nutch-0.9-dev

59.160.240.185 Nutch-test/Nutch-0.9-dev

59.176.10.136 NutchCVS/0.01-beta (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

60.248.9.114 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

61.135.151.175 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

62.129.132.47 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

62.168.188.151 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

62.40.33.173 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

62.40.36.87 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

63.133.162.98 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

63.246.7.209 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

64.105.36.210 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html;
nutch-agent@lists.sourceforge.net)

64.241.242.18 NutchCVS/0.05 (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

64.242.88.10 NutchCVS/0.05 (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

64.242.88.60 NutchCVS/0.05 (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

64.34.172.78 BurstFind Crawler 1.0/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; crawler@burstfind.com)

64.34.180.167 Nokia6620/2.0 (4.22.1) SymbianOS/7.0s Series60/2.1 Profile/MIDP-2.0 Configuration/CLDC-1.0/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

64.38.10.26 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

64.71.164.125 Krugle/Krugle,Nutch/0.8+ (Krugle web crawler; http://www.krugle.com/crawler/info.html; webcrawler@krugle.com)

65.220.67.9 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

65.92.160.39 JLA/Nutch-0.8.1 (beta; http://dynamic.com/index.htm; info at test.com)

66.132.240.180 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

66.132.249.23 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

66.15.68.234 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

66.207.120.226 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

66.243.31.34 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

67.111.28.139 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

67.184.246.61 Nutch/Nutch-0.8 (Nutch Test; none; none)

67.52.101.242 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

68.178.171.109 test/Nutch-0.8.1 (Test robot; http://test.com; info at test.com>)

68.178.202.79 test/Nutch-0.8.1 (Test robot; http://test.com; info at test.com>)

68.205.124.164 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

68.205.127.94 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

68.97.222.117 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

69.248.26.83 Comrite/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

69.36.233.8 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

69.55.233.28 Argus/1.1 (Nutch; http://www.simpy.com/bot.html; feedback at simpy dot com)

70.143.79.234 JPNutchTest/Nutch-0.9-dev-JP-0.1 (JP Nutch Test; jpnutch at yahoo dot com)

70.197.81.79 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

70.56.66.216 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

70.90.188.18 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

70.96.99.254 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

71.216.0.210 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

71.217.33.149 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

71.241.153.125 NutchCVS/0.7 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

71.35.163.79 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

72.0.207.162 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

72.2.25.66 abcxyz/Nutch-0.8 (nutchtesting; nutch; abc@xyz.com)

72.2.25.67 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

72.2.25.71 Nutch/Nutch-0.8

72.5.173.22 sdcresearchlabs-testbot/Nutch-0.9-dev (www.shopping.com/bot.html; researchbot@shopping.com)

72.51.37.148 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html;
nutch-agent@lucene.apache.org)

72.84.30.230 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

75.44.225.44 NutchCVS/0.06-dev (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)

81.173.148.94 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

81.173.155.210 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

81.203.142.109 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

81.93.168.211 TRankBot/Nutch-0.8.1 (T-Rank AS; http://www.trank.no/; robot at trank dot no)

83.246.79.28 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

84.191.111.92 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

84.231.72.32 agent/Nutch-0.8 (http://lucene.apache.org/nutch/bot.html)

84.231.74.47 nutch/Nutch-0.8.1

85.117.62.114 NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

85.18.14.22 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

87.139.106.60 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)

88.191.23.109 NutchCVS/0.8-dev (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)
Wasn't that fascinating reading?

This is some crazy shit that's almost like a DoS attack of non-stop web crawlers and I suspect it will get even worse as more people try to mine the Internet for free money.

Load up the firewall and your .htaccess filters with protection and brace for impact.

Thursday, November 09, 2006

JAP Anonymization Protects Scrapers Privacy

Isn't this nice, the JAP anonymization service is so busy trying to protect people's privacy that they don't give a shit that people will use their technology the assault web servers. Their slogan proclaims "ANONYMITY ISN'T A CRIME" but aiding and abetting an assault on a server could be considered a crime, questionable ethics at a minimum.

I got hit by someone utilizing their bullshit yesterday:

141.76.45.35 [proxy2.anon-online.org.] "Mozilla/4.0 (compatible; MSIE 5.0; Windows NT 4.0)"
141.76.45.34 [proxy1.anon-online.org.] "Mozilla/4.0 (compatible; MSIE 5.0; Windows NT 4.0)"
What these dipshits don't know is the few pages of data they downloaded, before the bot blocker kicked in and stopped the assault, has all been injected with hidden tags using CSS. Humans don't see these tags but the scraper, when stripping the HTML to get my text, will expose these tags to the search engine, and then I'll be able to hunt them down like the dogs they are.

Anonymous doesn't mean anonymous for scrapers anymore because even if you hide where you crawl from if this data shows up on the web it will expose where you live so be careful what you do with that data you sneaky little bastards.

Wednesday, November 08, 2006

Bot Blocking Obsession for Men

Today I realized that my bot blocking has become such an obsession that I'm almost worse than a cultist running around spreading the word of Bot.

What started as a simple effort to save my own website from virtually daily DoS attacks from Asia turned into a hobby as it was kind of fun looking for the next big thing hiding out there.

Then that turned into a product idea as it evolved and I realized the tools I needed didn't exist which is why I resorted to building them in the first place.

Then my universe turned upside down when I realized how much crap was going on that people weren't even aware of lurking under the cover of stealth on the net, THEN it became an obsession to build a product to restore privacy and control back to the web.

The upside is, obsession is a good quality for people trying to launch a new product but it severely impacts your social skills when your mind can't get off the topic as you're now burning all of your processing power 24/7 dwelling on the topic to come up with new insights and innovations to stopping crawlers on a daily basis.

Some days I wonder if I should call a priest so he can splash me with holy water and watch my head spin and spit green pea soup across the room like Linda Blair, but that's being POSSESSED, and I'm only OBSESSED.

The good news is the embroidery company sent my new Polo shirts and the logo translated to thread very nicely, I'm happy about it so now I have my uniform for Vegas next week ;)

Even better news is the graphics designer is working on the new layout for the bot blocker control panel and it should be HOT with CSS clickable bar charts with hover and shit, I can hardly wait to see them and bolt them into the software this weekend.

The crawler apocalypse is on the horizon, there's a light at the end of the tunnel, thanks to all you people for being patient as I wanted this bot blocker thing to be a real solution, not just rushed out the door, and taking my time has truly paid off in what my bot blocker is capable of doing in the real world.

The best is yet to come, I had polo shirts made, you can tell I'm serious ;)

Saturday, November 04, 2006

Hunting PicScout, the Copyright Crawler Getty Uses

Everyone knows about PicScout used by Getty Images but nobody seems to know anything about PicScout's crawler, no user agent information, no IP's where they crawl from, nothing. When someone asked me if I knew anything about them I did a little research and nothing related could be found ANYWHERE, not even anything initially obvious in my bot blocker log files. Based on my initial observations PicScout actually seemed to be hiding better than all the other corporate crawlers I've researched to date, but maybe we can shed some light on this.

Not that I advocate copyright violation, as a matter of fact, I'm a staunch copyright defender.

However, attempting to crawl under the radar, refusal to honor robots.txt files, or identify your bot in any fashion and bypass website security measures gets under my skin more than anything so I picked up the gauntlet and tried to find signs of PicScout activity.

After the usual simple research methods failed, I decided to start by seeing where they were hosted.

host picscout.com
picscout.com has address 82.80.254.37

host 82.80.254.37
37.254.80.82.in-addr.arpa domain name pointer bzq-80-254-37.dcenter.bezeqint.net.
Ah ha!

I remember a rash of activity I shut down from bezeqint.net a while back so I looked a little deeper into this angle.
inetnum: 82.80.248.0 - 82.80.255.255
netname: BEZEQINT-HOSTING
descr: BEZEQINT-HOSTING
country: IL
Ah yes, they're the guys from Israel that were hammering one of my servers.

I found a high volume of crawling from these IP's that was trapped by the bot blocker automatically and never answered the challenges, so it was definitely bot traffic.
82.80.249.195
82.80.249.196
82.80.249.197
82.80.249.201
82.80.249.202
82.80.249.203
82.80.249.204
82.80.252.130
These IPs have only been spotted using the two following user agents:
Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; (R1 1.1); .NET CLR 1.1.4322)
My theory is that this is PicScount attempting to crawl under the radar.

Check your logs people, see if you have any activity in this range, I think it's them.

I would just block this range out of principle at this point as those IPs crawling aren't honoring any internet standards, and if it is PicScout, blocking them could possibly save you a massive chunk of money if some web designer used stolen images building your website.

UPDATE:

After posting this the fine people from PicScout visited the blog and revealed more information about their facilities.

The log showed this visit:
Host Name mail.picscout.com
IP Address 62.0.8.2
Country Israel
ISP Nv-picscout
The information I found from that, including another IP block is here:
inetnum: 62.0.8.0 - 62.0.8.255
netname: NV-PICSCOUT
descr: NV-PICSCOUT
country: IL
admin-c: OG570-RIPE
tech-c: NN105-RIPE
status: ASSIGNED PA
mnt-by: NV-MNT-RIPE
mnt-lower: NV-MNT-RIPE
source: RIPE # Filtered
So, there's a few more IPs you might want to block, but I doubt they're scanning from the office.

UPDATE: Caught Getty keeping an eye on everyone today.

My blog log showed this:
Time: 12th June 200712:24:53 PM
Host Name outbound.gettyimages.com
IP Address 206.28.72.1
Country United States
Region Washington
City Seattle
ISP Getty Images
Referrer: http://www.webproworld.com/graphics-design-discussion-forum/56384-invoiced-getty-images-unlawful-use-images.html

It appears they were snooping on WebProWorld and followed the link here. The user agent claimed to be MSIE 6.0 but it's possibly an automated crawler, hard to say.

Anyway, we're watching you watch us, it works both ways.

Monday, October 30, 2006

Net::Trackback Rocks D-Block

Why is it every time someone puts some code out on the net like Net::Trackback that some asshole will download it and then aim their new creation at my server?

This is where they attempted to hammer my server this morning:

209.9.169.66 [209-9-169-66.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.78 [209-9-169-78.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.67 [209-9-169-67.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.70 [209-9-169-70.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.71 [209-9-169-71.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.69 [209-9-169-69.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.68 [209-9-169-68.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.72 [209-9-169-72.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.73 [209-9-169-73.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.75 [209-9-169-75.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.74 [209-9-169-74.sdsl.cais.net.] "Net::Trackback/1.01"
Of course they got nothing but error message for their troubles, but this is still.... BULLSHIT!

Can't even research the source as ARIN.NET's website won't load at this moment and CAIS.NET never responds to WHOIS inquiries and just hangs like this:
[Querying whois.arin.net]
[Redirected to rwhois.cais.net:4321]
[Querying rwhois.cais.net]
Never got a response...

Bunch of BULLSHIT, that's what this is!

Sunday, October 29, 2006

Hand Spammers Waving the White Flag?

Ever since I implemented techniques to automatically moderate hand spammers (aka Indian SEO's) they seem to have noticed they aren't getting through and have gone away. The first couple of weeks it didn't seem like they were slowing down at all, but they were moderated at least so nobody else saw them. Then I made some other changes in how I'm handling spammers that still did it by hand and suddenly they are just gone.

Before the last few changes I easily had about 10 hand spams getting trapped as moderated posts a day, then suddenly nothing moderated has shown up for over a week now.

Did they just give up?

We shall see, but this is promising!

Saturday, October 28, 2006

Ignoring my adoring fans, both of them!

I feel like I've been ignoring all of you lately but it's not true. I've just been so damn busy programming my little ass off, updating massive databases, writing web pages and PowerPoints, ordering custom embroidered polo shirts, so on and so forth, it's just crazy.

Sadly, I feel like the blog has recently become the red headed stepchild that nobody is playing with, not even the dog even if you hung a pork chop around the poor kids neck.

Seriously though, trying to get a massive update completed on an old site and roll a new product out the door while at the same time is crazy stuff. Top if off with getting ready for speaking at PubCon Vegas in November and SES Chicago in December is making me burn the candle at both ends but that hot wax feels OH SO GOOD on my nipple, but that's a different post.

It's like I've become some demoniacally possessed worker bee or some shit and I just can't get enough. I'm spinning out of control and probably heading for a serious burnout but it will be worth it as there are press releases and shit that are going to hit the web on a couple of fronts before the end of the year and I'm stoked.

Hell, I haven't even been going out to see movies, gamble or haunt strip clubs in almost 2 months but I'm sure a week in Vegas at PubCon will solve the gambling and hookers, um strippers issue.

Don't get too upset though, I'm still making it out at least twice a week for a nice 2-3 hour lunch with a friend and a LOT of beer.

Gotta treat myself right a little ;)

Thursday, October 26, 2006

Google, Yahoo and MSN Like Indexing Pure Garbage Sites

The other day I was working on a link checking filter so I could comb many thousands of linked sites and eliminate all sites from my index automatically that no longer contain valuable content.

What I did was make a filter that checked the profile of information on the page looking for signals that detected any sites that have reverted to default registrar pages, default hosting pages, or have become part of domain parks or scraper sites.

After successfully detecting and filtering out many sites that had fallen by the wayside, I started to wonder if the search engines actually indexed all of this crap.

Sure enough, a quick check of Google, Yahoo and MSN confirmed that the search engines eat these shit sites like candy although they can be easily detected and eliminated either by profiling the page content or checking the whois information, or a combination of both.

What purpose does indexing these millions of garbage web sites serve for any search engine?

I mean seriously, the scraper spam sites are one thing, but these are so easily detected there's no ryhme or reason they show up as results to any search being they are 100% crap.

Anyone from one of the major search engines mind dropping a note to explain why hundreds of thousands of cloned garbage sites are being indexed?

We'd really love to hear from you on this topic, please feel free to post a comment :)

Sunday, October 22, 2006

Scrapers Abandoning My Site?

In a rather unusual turn of events it appears my bot blocking efforts have went way better than expected to the point they might be backfiring. It appears that not only have some of the more serious scrapers stopped including my content in their sites, as they're being burned in the search engines (thanks guys) and new appearances of directly stolen content has went down drastically.

However, what I'm noticing is a new trend in sites that used to scrape my server now appear to be just scraping snippets off of other websites which I mentioned recently. Where this became most apparent was when I recently launched a boatload of new pages that had some breadcrumbs cloaked into all the pages not being served to the search engines. After a few weeks after releasing the new pages I went searching for references to these pages in Google, Yahoo, etc. and sure enough found some but they didn't contain my bread crumbs.

Doing a bit of quick investigation showed that the snippets indexed in Yahoo and Google actually came from the search engines themselves. This means that my site is being bypassed completely and the search engines are now the target for what little content the scrapers can get to use from my site.

Remember, I installed NOARCHIVE, NOCACHE and did a bunch of other things to minimize my exposure to the scrapers via the search engines many months ago yet they're still scrambling for the last few scraps that can get.

Just goes to show you how desperate these assholes are for any little scrap of information that ranks high.

Kind of sad that the search engines can't tell they're eating their own dog food though...

Spam Free Accomplishment Zone

Just thought I'd post a follow up after my latest anti-spam measures were put into place that it's been so blissfully quiet that I haven't even bothered to rant about these idiots or much of anything else lately because I've actually been doing more productive work.

Sure, the spammers keep knocking at the doors, banging on the walls, tapping on the windows, but other than falling silently into my spam log just to keep track of what kind of trash was at the door and silently swept away, I see none of it.

The last few weeks were absolutely amazing as I forgot just how much work a person could get done when you aren't constantly trying to clean up after those fucking spammers trying to shit all over every web form they can find.

Sorry spamming assholes, your days are numbered and I'm loving every minute without you.

Thursday, October 19, 2006

Nutch used to advertise Houxou?

I keep seeing this crawler for Houxou:

195.72.131.72 "HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)"
When you go to their link http://www.houxou.com/crawler it doesn't say anything about the crawler, it just shows you their homepage. I'm not sure what special interest on-line communities you can possible be serving when you can't even post the page your user agent links claim to be on your website.

Before I gave up altogether, I decided to see what I could come up with in Google and found some interesting results but the site appears to be down.
Nutch: search results
help. Hits 1-9 (out of about 9 total matching pages): WHOIS - 193.203.240.120 ... 20030922 source: RIPE person: Monu Ogbe address: 15 Penman Close, ...
nutch6.houxou.com:8080/search.jsp?query=ogbe&hitsPerPage=10 - 10k - Supplemental Result - Cached - Similar pages

Nutch: 搜索帮助 - [ Translate this page ]
搜索英文单词不区分大小写, 因此搜索NuTcH 等同于搜索nUtCh. ... 评分详解)显示Nutch如何给该网页打分. (anchors)显示指向该网页而被Nutch索引的anchor文本. ...
nutch6.houxou.com:8080/zh/help.html - 7k - Supplemental Result - Cached - Similar pages

So what's the deal?

Why is Houxou crawling with a link to a missing page about bots?

Is this just a ploy to get webmasters trying to figure out what the Houxou crawler is to look at their hosting services?

Who knows, guess we'll just have to wait and see but it smells fishy to me.

Smiley Face User Agent

This should be filed under "What the fuck is wrong with people".

Here's the user agent with a hyperlinked smiley face to some bullshit website in The Netherlands:

80.126.0.125 - "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; <A HREF=http://jult.net>;-)</A>; .NET CLR 1.1.4322; InfoPath.1)"
Looks like some asshole might've done this to his browser as the request did have a Google referrer so it's probably a real human that landed on my site.

Well pal, you got an error message when you hit my site didn't you?

Bet you're not so fucking smiley faced now.

Stupid shit.

Sunday, October 15, 2006

Netsweeper Caught Using Multiple Brooms

In the badly behaving corporate bots dept. we offer Netsweeper as our newest entry from Canada. They run one of those content filtering companies that thinks they should be allowed to crawl your site no matter what just to protect their clients.

Sorry, but we happen to disagree with all these content filtering spiders that feel the need to crawl without any regard for robots.txt and we really don't need a whole buttload of content filtering companies scanning the fucking web.

Yes, I threw in the word fucking just so your asshole spider will flag this post as bad content so none of your goddamn customers can read this so blow that out your ass.

Let's see what Netsweeper runs:

66.207.120.226 "webcollage/1.127"

66.207.120.226 "NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)"

66.207.120.227 Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.5) Gecko/20041107 Firefox/1.0
These IP addresses have the following host names:
66.207.120.226 -> firewall.net-sweeper.com

66.207.120.227 -> host227.net-sweeper.com
Let's just cut thru the chase and here's the information to block their ass:
CustName: Netsweeper
Address: 4-512 Woolwich Street
City: Guelph
StateProv: ON
PostalCode: N1H-3X7
Country: CA
RegDate: 2003-04-08
Updated: 2003-04-08

NetRange: 66.207.120.224 - 66.207.120.239
Ta ta Netsweeper, you've been blocked and swept under my rug.

Cya!

Saturday, October 14, 2006

No Referrer and Visitors vs Spambots

I've been struggling recently on better ways to handle requests to form pages, like a comments page, that are more secure yet more visitor friendly than just slapping up a "403 Forbidden" which runs off humans as well as bots.

After contemplating the issue and taking bookmarks and disabled referrers into account, I decided to simply redirect these potential bad hits to my home page instead of the old 403 error. This way a valid visitor that bookmarked the page could just navigate back, referrer intact and post as usual. So far I've seen a few humans that were redirected off the page for whatever reason navigate back to where they wanted to go, so it doesn't appear to be stopping determined people that aren't just there for malicious purposes.

Additionally, I only do this redirect after verifying the request isn't coming from the search engines as I obviously don't want to confuse the SEs by redirecting them to the home page.

The fun part is it seems to be confusing the shit out of the spambots, they're bouncing all over the place, quite hysterical to see them freak out.

For even more fun, verify that the referrer to your form page isn't a direct hit from a search engine to that page as many of the hand spammers from places like the Ukraine and India seem to like to use Google to find pages to post their clients' listings. For those reasons, now I'm also redirecting any direct hits to the posting page that comes from Google, Yahoo, MSN and ASK. Redirecting the hand spammers (aka SEOs) back to the home page seems to stop the hand spams as well since I've just made the job a little harder they just seem to move along instead of spending more time to find the page

One last trick, if you have the means to track these things like I do, is reject or auto-moderate anything with just a single page view to your site, which is just that form page. Some figured out I was looking for a specific referrer and plugged it in so the post looks 100% legit. Well bummer dudes, you need to view more than one page to make a submission so cleaning up the data to add a valid referrer was a nice try but you still have a bad bot profile by having no previous page views.

Gotta love all the fun and games with both sides escalating but so far I'm still spam free and winning this war.

Friday, October 13, 2006

Bot Busting or Spam Hunting?

Since I made a few posts about spam lately, mostly web spam, it seems I have a few spammers all bent and other people thinking that I'm a spam hunter.

For starters, I'm a bot buster and not a spam hunter. Whether it's a scraper, a stealth crawler, or a spambot, they're all bots so anything that malicious bots do to my sites that need to be stopped will be of interest to me. Defeating a bot that scrapes or spams each has it's own challenges and it's not been too terribly hard to stop them either way, so far.

Spam Hunter?

Get real, who has to hunt spam?

I just sit here minding my own business and spam comes from every angle via email, online submission forms, blog comments, or anything else you can put online and they'll spam it. When I see a trend emerging in my own spam logs, like the recent wiki/tiki and phpBB abuse, I just look to see how deep the problem is which isn't hunting spam exacly, it's analyzing the trends and patterns of the software vulnerabilities that are being exploited.

Maybe someone will notice and close some of those exploits, wouldn't that be nice, unless you use those exploits...

Just because a few of the latest topics revolve around spam, it's still bots in action and bots I've automatically blocked and logged, just spambots is all, but still bots all the same.

Enough of this noise, back to busting bots, scraper, spam or otherwise.

Wednesday, October 11, 2006

Web spammers abuse GuestCity's hospitality

While researching the depth of the wiki/tiki spam abuse problem there was one particular redirect link that caught my eye to some site called GuestCity.



When you see the smoke from web spammers in a search engine there's usually a web spammed fire somewhere close so I decided to look deeper to see what GuestCity was all about and assess the damage.

There on the home page was an encouraging anti-spam symbol on the lower left of the screen!



I clicked on the anti-spam symbol and read their get tough policy on spam, cool.




If these guys are really tough on spam, shouldn't find much web spam over there, right?

Sorry, took about 2 seconds to spot sites overflowing with crap like phentermine, viagra, cialis, and on and on.

If you want to see the funniest shit ever, click on their DEMO link right off the home page that is spammed upside down and inside out with a couple of years worth of garbage.

The most priceless quote is this one from 2004:
4249. Old demo book was removed due to lot of spam messages. Welcome to new one!
2004-10-21 08:33:59, Webmaster,
I think I fell of my chair laughing hysterically about then.

Sure wouldn't take more than a few days to write some code that would stop the spammers and eradicate all the splogs over there, hope they're up for the challenge.

Zone Communications sends SEO Spam

Just when I thought some of the SEO's were getting smarter, since I haven't been spammed by one for a while, here comes a nice juicy one from our friends at Zone Communications in southern California.

The spam came from this IP:

71.128.4.233
ppp-71-128-4-233.dsl.irvnca.pacbell.net.
Here's the lovely spam:
Your website can be at the top of the first page on all major search engines. Zone Communications has a great service that is very low cost and is billed month to month with a full refund if you are not satisfied. With this service you�re your ranking on Google, Yahoo and MSN within certain cities will be at the top of the page.

The pricing is simple. We charge $59 per month for your first listing and $40.00 per month for each additional listing. (There is also a small one time set up fee.)

If you are not satisfied with your position within 30 days, we will give you a 100% refund.

[spammers name]
Zone Communications
800-xxx-xxx
714-xxx-xxx (fax)
[spammers name]@zonecominc.net
OK, if they even took a look at the site they spammed they would know I'm all over the top of the 4 top search engines for keywords, cities and just about everything in my niche short of Mom's Apple Pie and Kitchen Sinks.

For people that spam me, my pricing is simple. We out spammers for FREE, and there is no per month fee to continue being outted. (There is also a small one time fee called you've been BUSTED for sending me spam.).

Now stay off my damn websites, you weren't invited in the first place.

Saturday, October 07, 2006

Vulnerable Tikis Ruthlessly Spammed and Google Indexed

The other day I posted about how VT.EDU's tiki was overflowing with spam so today I went thru my spam filter log to just see how many attempted spams there were last week using tiki redirect pages.

Here's a short list of the most recent attempted spams linking to tikis that hit my server:

http://www.lug-viersen.de/tiki-directory_redirect.php?siteId=136#viagra
http://ipvs.informatik.uni-stuttgart.de/BV/swarmrobot/tikiwiki-1.9.2/tiki-directory_redirect.php?siteId=474#viagra
http://i60p4.ira.uka.de/tiki/tiki-directory_redirect.php?siteId=24#viagra
http://www.xsl-rp.de/tiki-directory_redirect.php?siteId=1018#cialis
http://www.neurotransmitter.net/wiki/tiki-directory_redirect.php?siteId=243#viagra
http://research.cs.vt.edu/advance/tiki/tiki-directory_redirect.php?siteId=3284#viagra
http://meverhagen.nl/tikiwiki/tiki-directory_redirect.php?siteId=19#viagra
http://www.namurantifasciste.be/tiki-directory_redirect.php?siteId=996#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=10#viagra
http://www.xsl-rp.de/tiki-directory_redirect.php?siteId=1015#viagra
http://www.railfuture.org.uk/tiki-directory_redirect.php?siteId=61#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=9#viagra
http://herenaforge.org/tiki-directory_redirect.php?siteId=38#phentermine
http://herenaforge.org/tiki-directory_redirect.php?siteId=51#viagra
http://www.derrychineseschool.org/DCS/tiki-directory_redirect.php?siteId=7#viagra
http://openg.org/tiki/tiki-directory_redirect.php?siteId=54#viagra
http://www.prospace.org/tiki-directory_redirect.php?siteId=2385#viagra
http://www.milwaukeelug.org/tiki/tiki-directory_redirect.php?siteId=1349#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=18#viagra
http://dev.librehwdb.tuxfamily.org/tiki-directory_redirect.php?siteId=18#viagra
What's distressing is that Google and the other SE's really love these spammed pages too, just gobble them up, and it's probably unwittingly passing PR from all these spammed tiki sites on such terms as viagra, cialis, levitra and a whole lot more.

So Google gives spammers a 2-for-1 special by giving them SEO value for their spamming activities, it's just a crying shame, it really is.

What's pathetic is this problem could be stopped on both sides of the coin. The tiki/wiki software developers could get off their lazy asses and implement some tools to allow webmasters to stop this rampant spamming of their software, it's easily doable. Additionally, the search engines like Google can easily identify and stop indexing spammed web pages to eliminate the value they give to the spammer.

Remember, I'm reporting about ATTEMPTED spams, all those links and a shitload more were automatically dumped, it's not rocket science, it's barely programming above a rudimentary level to identify and filter that shit out.

Why does this continue when the solutions are so simple for all involved?

Amazing that it's allowed to continue, simply amazing.

Thursday, October 05, 2006

Podomatic Vulnerability Enables Spammer Redirects

Here's another instance in a rash of reported vulnerabilities in member registration pages being spammed. Never heard of Podomatic before but it appears the spammers sure have and some nitwit registered as a member called Valium to do his spamming.

The link to the member's site is:

http://www.podomatic.com/profile/member/valium
The javascript redirect code appears to be this shit embedded in the memberpage:
<script>
var mbht872 = 'on=';
var bikmr354 = 'qiqyi199';
var zlh171 ='ment';
var k97='.lo';
var ydxglyjedai737='ti';
var bmmp211='docu';
var mzcra833='http://drsearch.net/search.php?aff=15313&q=';
var ertmj632='valium';
var qiqyi199 = 'ca';
var lflx482='"';
if(bikmr354 = 'qiqyi199')eval(bmmp211+zlh171+k97+qiqyi199+ydxglyjedai737+mbht872+lflx482+mzcra833+ertmj632+lflx482);
</script>
Just goes to show you that if you don't secure your sites some spammer will abuse it but people just don't listen.

Wednesday, October 04, 2006

Automatic Detection of Spam Hand Jobs

Sometimes certain anti-spam ideas just hit you upside the head when you least expect them and seem so obvious you wonder what took you so long to figure it out.

I've already blogged about the fact that I've stopped all automated spam dead in it's tracks on my sites, but people manually posting can of course correct all of the errors detected and continue to make an unwanted garbage post.

I have an extensive junk detection filter that rejects anything with the usual suspects like viagra, cialis, gambling, poker, etc. which stops the nastiest of these posts. However, some little pain in the ass SEO aka spammer might slip thru with a hand job posting about his store in India selling magic beetle dung or something that you would never imagine putting in your junk filter in the first place.

A few days ago I decided to review the last 30 days of legitimate submissions and compare them to the few off topic hand jobs that slipped through the cracks and see if I could come up with anything that would allow me to stop the hand jobs of absolutely random and crazy things outside the realm of the typical common auto-spam posts.

Then, like a lightning bolt it suddently hit me, that with these random off topic hand spams it's not what's IN the posts it's what's NOT in the posts that makes them easily identifiable. The concept is to scan for a list of words that SHOULD be in the post, like quotes from anything in the thread or certain keywords related to the topic and automatically set everything to MODERATE that doesn't fit the usual posting patterns.

Basically it's a 'lack of content filtering' technique and off topic posts, like spam, stand out like a sore thumb.

Using this blog as as example for a topic, you would expect most comments to contain words like bot, spam, IP, host, crawl, firewall, htaccess, apache, etc. or a set of keywords derived from the original post title and text. The absence of any of these words is a clue that the post just might be SPAM or otherwise off topic and should be placed on moderation for the admin to review.

Since I've started using this new 'lack of content filtering' technique it's snared the few hand submissions to my other site that were completely off topic, those that I would've deleted immediately. The beauty is I can continue to leave the posting wide open for humans, not moderate everything, with only those posts that don't match the topic getting instantly set to moderate.

I expect a few false positives but so far 'lack of content filtering' is doing exactly what I expected it do and set a couple of crap submissions last night for shit like "zanaflex information", apparently some pill I've never heard of and "News, Stores, People, Careers at Finditt", some wannabe search engine, to moderate automatically while letting 20 on topic things thru without a hitch.

Another automated weapon in the war on spam!

GoodBidWords.com Scrapes LookSmart

Noticed at hit from one of my scraper probes in GoodBidWords.com which contained the IP address of the original crawler.

Looked up the IP address and guess where it came from:

"Mozilla/4.0 compatible ZyBorg/1.0 (wn-14.zyborg@looksmart.net; http://www.WISEnutbot.com)"
Isn't this precious that GoodBidWords got caught because of all the places to scrape they decided to scrape a search engine that I don't permit to crawl my site!

What a hoot, second-hand scraper busting, this rocks!

Tuesday, October 03, 2006

phpBB Membership Spamming for Authority

We first reported about phpBB spamming the other day when we stumbled upon this "DISY registration spamming script" and since then have had a little time to examine what spammers are doing with phpBB trying to gain authority.

Let's just check a few of these spammers in Google:

pimpdomain.net
thewestgategazette.com
ritalin-pharmacy.com
Hell, just try any of the domains listed in my Technorati Loves Spam post and search for the domain name and phpBB and see what shows up.

Just amazing what these assholes do with this shit cluttering up the net with spam.

Technorati Loves Tasty Cloaked Blog Spam

I've noticed that Technorati has been happily eating up scraped and cloaked blog spam for ringtone sites, among other things, like it's fucking candy.

Let's use a search on my blog name as an example:



Click on those links and it's always to the same spammy page name like these hosted on theplanet.com of course:
http://artinexis.net/#comment-341
http://themetrogiant.com/#comment-329
http://pimpdomain.net/#comment-341
One server is 70.87.88.121 or 79.58.5746.static.theplanet.com with these domains all spewing ringtone ads:
about-levitra.net
acvfa.net
artinexis.net
cariculture.net
catsfive.net
citadel1.net
cloudsite.net
eightonefive.net
rennenmotorsports.net
t3linkcom.net
Another annoying server is 70.87.88.108 better known as 6c.58.5746.static.theplanet.com which has these goddamn domains:
talonpro.com
tempuspercussion.com
terminal34.com
the-god-poll.com
theincrediblesuckingspongies.com
themetrogiant.com
thepulse2000.com
thespinet.org
thewestgategazette.com
thoweu.com
tlc-express.com
Or this fucking spam filled server host 70.87.88.106 hosted by our fucking friends 6a.58.5746.static.theplanet.com:
perseidslive.com
pimpdomain.net
poemnet.net
posses1consent.com
projhind.com
ptcsucks.com
r1g4t2you.com
rbigkitty.com
rep1icas.com
ricohtour.com
rising7.com
ritalin-pharmacy.com

Here's the same shit about ringtones they all show:


Who are the fucking idiots buying all these goddamn ringtones anyway?

How about you just set the phone on buzz, stick it in your pocket, and you'll never miss a call or be confused it's someone else's phone ringing, and best of all you can do it without lining the pockets of the cell companies or perpetuating this spam. Better yet, just shove that phone up your ass as most people that feel the need to never miss a call by using goddamn custom ringtones are probably talking out of their ass anyway. While you're at it, shove some custom phone face plates and a nice blue tooth headset up your ass too, but I digress.

WAIT A FUCKING MINUTE...

I think I see a pattern here 70.87.88.106, 70.87.88.108, 70.87.88.121...

Let's try 70.87.88.120 and see what we find:
asmort.net
bevirusproof.net
conlajusticiaysociedad.net
fabionne.net
friendshipmotorinn.net
macoszone.net
palick.net
phila-ibiz.net
themikecam.net
wesmn.net
More ringtone spam spam spam....

Or let's try 70.87.88.115:
audio-wire.net
buy-cheap-2u.com
chabadofbuffalo.com
cheap-online-buy-free.com
el-condor-pasa.net
ellemtel.net
fairy-wings.net
free-top-sex.net
gotobiz.net
healthcybeline.com
javabooks.net
jemison-nealon.net
lolitasexlinks.net
macromediaseminars.com
netnetn.net
remax-powell-m-corpus-christi.com
stopsundiata.com
wangmatongli.com
xinyifang.net

Yes!

More spam spam spam spam spam!

OK, this is obviously a big operation with lot's of shit domains serving up spam on lots of IP's, I'm bored with this already, if you want to help fill in more blanks with this ringtone spammer go to Domain Tools Reverse-IP page and type in the IP address in that range and see what's on the servers.

Maybe they should change their name to Spamorati as they seem to love these fake blogs reposting old posts.

BTW, if you need help with automating the identification of spam over at Technorati just drop me a line as I'd be more than happy to show you how to automate the process for a small fee!

The things I could teach them on ways to clean up their listings and improve their service would boggle their minds.

Saturday, September 30, 2006

ShoeMoney's Blog Spam Stopping Primer

The day after my battle cry to Rally the Anti-Spammers here comes ShoeMoney with some great suggestions for stopping blog spam. Everything ShoeMoney posted is very solid advice but some spammers have already been evolving past some of those patches which is why I use my draconian anti-spam methods. Basically, ShoeMoney's advice will stop the majority of your garden variety spammers, but not all as they are constantly adapting, so as you improve your defenses they improve their ability to bypass those defenses.

Remember, security is built in layers and the more layers you pile on, the more the spammers will chip away at your security so building the better spamtrap just results in smarter spammers and they're already here which I'll address with examples below.

Let's examine ShoeMoney's anti-spam advice, see what some state of the art spammers are already doing, and add a few more tricks here and there for even better security.

Starting with the first item he listed:
5) Deny Access to No Referrer Requests

The approach does work on most spammers but I had about 10 requests today where it would've failed. Not that you shouldn't implement this, it's a good trick to stop a lot of spam, just be aware it won't stop everything.

Example:

My bounced spam log shows the following:

IP: 84.110.248.226
User Agent: "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Subject: "Viagra"
URL: http://anol.webhosting.gs/viagrageneric.html#viagra
Take a look at what's in my server log:
84.110.248.226 - "POST /formsubmit.html HTTP/1.0" 200 11918 "http://www.mysite.com/formsubmit.html" "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Yup, that's right, a referrer, and I had about 10 of those and they were all from spambots.

Stopping the poorly coded spambots is easy, but they won't be vulnerable for long as the patch to add the domain name being spammed into the referrer is trivial so I expect this anti-spam advantage to be short-lived but I use it too, you should still do this.

Now, let's tackle the next item, which is VERY good advice:
4) Kill tor anonymous proxies

I block many proxies on my servers, which does stop a lot of spam, but don't think that all spammers use known proxies. This is the reason I also block dedicated server hosting facilities because a series of $2 webhosting accounts can be used to effectively spam and bypass the proxy lists.

Example of 4 sample spams (out of many) today that all had referrers mentioned above and came from some ISP/Host called bezeqint.net:
09/29/2006 84.110.248.226
"Viagra" http://anol.webhosting.gs/viagrageneric.html#viagra

09/29/2006 84.110.244.240
"Viagra" http://gerda.forospace.com/#viagra

09/29/2006 84.110.243.107
"Cialis" http://borea.forospace.com/#cialis

09/29/2006 84.110.241.163
"Cialis" http://kaizer.webhosting.gs/cialisbuy.html#cialis
Use this with caution:
2) Blacklist Repeat Offenders:

First off, blacklist on the FIRST offense so there is no second time. However, you really need to know what you're doing and lookup who the IP address belongs to so you aren't blocking IP addresses from places like the AOL IP pool (reused every 15 minutes or so) or any other shared proxy dial-up IP pools as those IP assignments are very temporary and the next access is probably a different visitor, not a spammer, so be very careful with this.

This is a gem and we can make it better:
1) Rename your comment file

Excellent advice as I've done that on some websites but don't be shocked when it's short-lived as spammers also have crawlers looking for these comment pages and the fact that you're still linking it under the keyword "comments" is a dead giveaway.

If you're going to change the file name, also change the word that links to the file name to "discussion", "verbal intercourse", or "rants", anything but "comments" to throw them off.

Additionally, move the actual FORM into obfuscated javascript document writes. How this works is the spambot scanning your website can't even find the webform to submit comments as most bots don't use javascript, so only an actual visitor would see an actual webform written into the web page via javascript.

Don't forget the CAPTCHA!

Now, the one thing ShoeMoney didn't mention which works wonders is a simple CAPTCHA and it's keeping a few of my sites spam free without ANY other work involved. Yes, there are ways to bypass a captcha but it's not easy for the spammer. So far most captcha protected sites are safe with such simple protection, but I expect that situation to escalate soon.

Kudos to ShoeMoney for spreading the word, we need more anti-spam information spreading and more people jumping on the anti-spam bandwagon so we can rid the 'net of this scourge as soon as possible and move on to more productive activity.

Thursday, September 28, 2006

Virginia Tech's Computer Science: Wiki Spam 101

My website stops spam posts cold, and logs them, so that eventually I can glance over the list of bounced spams now and then just to see what was caught and this one was priceless:

09/28/2006
200.88.223.98
"Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Subject: "Viagra"
URL: http://research.cs.vt.edu/advance/tiki/
tiki-directory_redirect.php?siteId=3284#viagra

I looked and thought, "Viagra spam linking to VT.EDU? Could their server be hacked like SpamHuntress is posting about?" So I click the link and of course it uses VT.EDU's server to redirect me to some viagra sales site just like the URL would make you think it would, no surpise there.

So I trimmed the URL to see what in the heck this site was and it's ADVANCE, FOR THE ADANCEMENT OF WOMEN IN ACADEMIC SCIENCE AND ENGINEERING CAREERS and it's full of advances for MEN such as viagra, cialis and levitra spam plus a whole bunch more.

Well, the IT dept. and professors in charge of the VT computer science program should probably start quaking in their boots as I would be VERY UNHAPPY if I was the Dean.

This is completely unacceptable when the IT guys and CS Profs aren't using even rudimentary anti-spam technology like, oh, maybe a simple CAPTCHA to stop this shit.

I want my tuition refunded.

BTW, whoever these spammers are, they've been VERY BUSY little beavers.

Time to Rally the Anti-Spammers

After the demise of Blue Security and this recent meaningless default judgement against SpamHaus, the spammers are getting braver and bolder by the day. Now, one of the most vocal anti-spammers around, SpamHuntress, has recently come under attack after exposing a few people that really didn't want to be exposed.

Even one self-professed blackhat SEO web spammer has the audacity to tell SpamHuntress to "get a life" because she must be cutting into his livelihood. Maybe I'm just too lazy, but who would've ever thought of registering for a bunch of forums and never posting as an SEO tactic? Using his DISY registration spamming script probably sped it up and he's busy making friends [scroll to bottom] as well.

OK, so now the phpBB people will need to be alerted to add NOFOLLOW to all those links in the registration page to stop this SEO vulnerability, but I digress, will rant about that later.

Unlike email spam, which is a real pain in the ass to stop, there is absolutely no reason we have blog, forum or guestbook spam whatsoever except for shitty programmers writing the stuff and people using it that either:

  • have abandoned their websites or forgotten that old guestbook or blog now littered with junk
  • aren't aware there is a problem as many spambots post on older threads
  • don't know there are solutions to these problems
  • aren't capable of installing the patches even if they are aware of the solutions
I've posted before how I stop spam on all my pages that have forms for submitting a variety of things, WITHOUT the use of captcha's, and although it's a pretty draconian approach to the problem it's also highly effective. My solution was to simply reject any posts with embedded HTML and URLs, just bounce them with an error about the content, and it works 100% against real spam. Maybe it's a tad extreme but when this type of spam is dead maybe I'll open up my sites again to more robust content posts, you never know.

However, for those that like to continue to do things the hard way, here's a list of software you can install to stop the spammers:
I'm going to ask that people reading here help the cause and start educating everyone you run across with a blog or forum being overrun by spam.

Please point them to a resource to solve the problem or offer to help them add the plug-ins pro-bono or for a nominal fee if they don't understand how, or if all else fails alert the host to help sites overflowing with spam and see if they'll be of any assistance.

Don't forget, the purpose of these spammers is to drive direct traffic and also get results in Google so when you stumble upon these sites in Google, make sure you file a Google Spam Report while you're there to get them whacked from the search results.

We can stop this in the next year or two, as long as people quit being complacent and just install the upgrades, patches, captchas and other anti-spam tools.

Spread the word, let's just get this done so we can stop talking about it already!

Wednesday, September 27, 2006

MySpace: Porn Networking Spam Machine?

The other day I signed up for MySpace while researching the members with "Click The Ads" on their pages encouraging others to commit click fraud to fund their various lame causes.

Unfortunately, signing up for MySpace immediately resulted in a couple of porn spams sent to my Inbox which really pissed me off.

So I get some shit that looks like this:

FROM: MySpace Events
SUBJECT: .. has invited you to: I seen you online

Hi ,

.. has invited you to an event on MySpace:

Click the link below to view the event details:
http://events.myspace.com/index.cfm?fuseaction=PORNSPAM

Now below this, there is some bullshit message from MySpace:
At MySpace we care about your privacy. We have sent you this notification to facilitate your use as a member of the MySpace.com service. If you don't want to receive emails like this to your external email account in the future, change your Account Settings to "Do not send me notification emails."
Really, you care so much about my privacy you let goddamn porn spammers send me fucking email?

I'm touched, a tear comes to my eye ...

... yes a tear, because I realize can't reach out and smack whoever let this shit happen upside the head!

Anyway, here's the website on MySpace linked from the spam:



Here's the first site's link to a girl with a webcam:




And here's the second spam site's girl with a webcam:





I'm wondering if people under 18 get these spams too?

I think I'll just cancel the account because MySpace is no place I was to be associated with.

Monday, September 25, 2006

MySpace: A Click Fraud Social Network?

Maybe it's just Web 2.0, or Web Welfare 2.0, but it appears that stealing from advertisers is now something that is accepted in social networks. Let's look at what we find on sites like MySpace and others which are a good place to build up a nice list of friends to click your ads, especially the Google ads, because we all know that friends click friends ads, especially if you want your friends banned from AdSense.

Even on YouTube where people can't put up their own ads they beg people to come to their website and click the ads to support them putting up more videos!



The most shocking is Blogger, which is owned by Google, the creator of AdWords and AdSense, which hosts sites that encourage people to "Click the Ads" to defraud the very advertisers they rely on for their massive income.

How difficult would it be to have a single employee out of the entire Googleplex devoted just to keeping click fraud off their own property?

You know the answer, I know the answer, yet a simple search reveals that it's not being done, or not done adequately at any rate or there would be no sites returning results from Blogger on this topic if they were on top of the problem.

The technology for these sites to deploy an automated process to locate pages within their sites that contain calls to "click the ads" or "click Google ads" or any combination and eliminate this fraud on a daily basis is so trivial and rudimentary that beginning programmers could do it.

Bottom line is there's absolutely no excuse for this type of call for advertiser click fraud to be allowed unchecked on these sites, not in MySpace, YouTube, Blogger, Google, Yahoo, MSN or anywhere else and why Click Fraud 2.0 continues to perpetuate on the web when it's so easy to thwart frankly boggles the mind.

Flickr Member Requests Click Frauding Advertisers for the Children

Well, I've seen all sorts of excuses to advocate click fraud but the plea on Flickr to commit a crime for the children is a new one and more despicable than any I've seen before. Think about the precedent that this sets in impressionable young minds that it's "OK TO STEAL FOR A CAUSE" when crime is never OK. Sadly, all of the good this person has possibly done for these children was wiped away with one call to arms to defraud people for a cause.

If you want to save the children, set up a Paypal account and teach the children than they can be helped by the generosity of others, not by others commiting FRAUD!

Here's the screen shot from Flicker:



And the site it lands on in Blogger:



Come on buddy, just ask for donations and keep it legal as we all love the children but this is over the top.