Monday, June 12, 2006

Bad Karma is a potential DDoS threat

Some guy has something he's put out as freeware called the Referrer Karma which gets the referring page and checks to see if it actually has a link to the site referred. If no link to your site exists on the referring page it slams the door on the visitor assuming it's a referrer spammer.

Two problems with that approach:

  1. Links that pass thru redirect pages from directories directory sites will fail this test every time as the referrer is the redirect page itself, not a web page with links on it.
  2. Sites that block bots, like mine, toss out error pagess when stupid user agents appear and VOILA! the visitor from my site gets bounced off by this stupid script.
Here's the info:
65.98.116.226
cp5.secserverpros.com.
"Referrer Karma/2.0"
Next, let's explore my concern with potential vulernabilities with Referrer Karma.

If you think about the implementation of Referrer Karma for a minute you'll realize it would allow one kiddie script to potentially pull off a DDoS attack. This could be accomplished by issuing thousands of requests to a bunch of sites running this Referrer Karma and each request containing a faked referrer to the target site you're attacking.

You wouldn't need to wait for the page request to complete, just send out a ton of requests to a bunch of servers and terminate the socket when the websever respondes with data is ready. No need to download the resulting page as Referrer Karma has already done your dirty deed for you by hitting the other site asking for the requested page.

Ask for a few thousand pages in a few seconds from a a bunch of sites using Referrer Karma and step back and watch the fun as the target server melts.

Lack of Intelligence Competence Crawler

Well here's yet another site called the Intelligence Competence Center trying to crawl the web looking for things they can sell they to various industries.

Here's the crawler details:

212.227.103.133
s15208971.onlinehome-server.info.
"iCCrawler (http://www.iccenter.net/bot.htm)"

also...

82.165.39.218
p15197600.pureserver.info
"iCCrawler (http://www.iccenter.net/bot.htm)"

all IPs it's used with my site...

212.227.103.133
212.227.93.221
82.165.39.218
I'm really getting sick and tired of these fucking corporate leeches that keep crawling [pun intended] out of the woodwork.

Sunday, June 11, 2006

HK Creepy Crawler

Been seeing this same bot "Java/1.4.1_04" asking for the same pages from the following IPs multiple times:

210.177.215.25
210.177.215.28
210.177.215.29

Don't know if it's a hosting account or ISP, but it's worth keeping an eye on these IPs.

Friday, June 09, 2006

My sites doesn't need UPDATED!

How many comparison shopping sites can one planet possibly need?

Well, whatever your answer was, it's wrong, there's one more called Updated crawling around.

The user agent info:

38.119.96.110 "updated/0.1-alpha (updated crawler; http://www.updated.com; crawler@updated.com)"
Block it, don't block it, I don't care...

Thursday, June 08, 2006

Google Dance Makes You Shit Your Pants

What a day.

Got up this morning and my main sites traffic was wonky so I took a look around and Google traffic seemed to be askew. Checked a few data centers and my position was all over the map. Up and down, round and round, Google's chewing me up and spitting me out in all sorts of wacky places.

My default datacenter now shows me back in the top 10 but only God and Google knows where it'll be tomorrow and neither of them are talking.

What a mess, Google better have a new liver standing by for me...

Anonymous Media, just can't make this shit up

Here comes another scaper with a mission called Anonymous Media and they want your shit.

63.133.162.98 - "GET /robots.txt HTTP/1.0" 200 146 "-" "Anonymous/0.0 (Anonymous; http://www.anonymous.com; noreply@anonymous.com)"

Right on their website it says they:

  • SPY "anonymous watches what consumers watch"
  • SCRAPE "anonymous compiles data from multiple sources"
  • PROFIT "anonymous generates market reports and analysis"
  • TARGETS YOU "anonymous conducts custom market research"
OK, is it just me or are we all getting SICK OF THIS SPYBOT SHIT?

Maybe we should just install 301 redirects to other spybot companies when spybots come crawling it points to a competitor and just let them crawl each other's sites to death.

RED ALERT UPDATE - SnapBot and the Linux Firefox Revelation

Finally, after weeks of these idiots bombarding my servers with Linux Firefox for unknown reasons I finally connected all the dots as SnapBot is apparently running from 3 data centers that I know about and all 3 have run both the SnapBot crawler and Firefox from the same IPs.

65.38.102.0/24 "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1"

38.98.19.0/24 "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1"

66.234.139.0/24 "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1"

You know what SnapBot was doing from these 3 locations?

Making goddamn screen shots of every fucking webpage, or trying to anyway, all 40K+ of them too!

Fuckers.

I know this for a fact as I found a picture of my error message about their site being banned with their IP number embedded in it from the 65.38.102.0/24 block and the others are already known Snap IPs.

Listen assholes, take a picture of the home page and leave it at that, you aren't getting 40K screenshots and anyone that thinks they are can blow that shit out their ass.

Don't Worio, just another search engine

Something new hit my site claiming to be Worio and planning to go live in 2006 according to their website. Another site using Heritrix to crawl with, oh joy, almost becoming as annoying as the new-nutch-site-of-the-week club.

BAD_AGENT: 198.162.51.70 [worbo2.cs.ubc.ca.] requested 1 pages as "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://www.worio.com/)"

What makes Worio special is it's supposed to be created for computer scientists and programmers.

OK, there goes my lunch.

Don't Worio, Be Happy.

Wednesday, June 07, 2006

REJECT SPAM, DO NOT BOUNCE!

I've been setting my email pref's on all my servers to REJECT unknown addresses instead of BOUNCE which opens your server up to all sorts of spam problems as your email queue fills up with undeliverable garbage and your server starts to choke.

Thought I was pretty clean but I imported some domains from another server and suddently started seeing 150+ emails piled up ever day. They were all failure notices just sitting there chewing up my box trying to deliver all day long.

Got sick of that shit last weekend and did a complete email audit and now ALL DOMAINS are set to REJECT unknown addresses.

Guess how many outbound emails are pending now?

ZERO!

Been zero since I did it and holding strong so spammers using randomname@mydomain.com are no longer bothering me whatsoever.

Just a reminder, check your email accounts and make sure you're locked up tight!

Block them PRONTO

Roll out the RedCarpet and let them scrape all they want.... NOT!

We found this little beast bothering us:

BAD_AGENT: 66.45.38.59 [unknown] as "RedCarpet/1.3 (http://www.pronto.com/robots.html)"
On their web site it says:
We use our patent-pending software [crawler] to scour the web [scrape] and identify the most merchant sites and products possible for you.
What a bunch of happy horseshit, web crawlers are now patent-pending software?

Go take a long walk of a short pier and jump into lake go-fuck-yourself.

PlanetLab comes knocking

Here we go again with what looks like someone testing my PlanetLab barriers to see which proxies do or don't work and this just started a few minutes ago and will probably get more interesting in an hour or two.

Note the same browser from each location:

BANNED: 128.4.36.12 [planetlab2.pc.cis.udel.edu.] requested 1 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"
BANNED: 129.10.120.111 [planetlabone.ccs.neu.edu.] requested 1 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"
BANNED: 155.98.35.2 [planetlab1.flux.utah.edu.] requested 1 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"
This PlanetLab bullshit needs the plug pulled on it ASAP.

Google Gags on Drop Down Menu

This was an amusing find in my log file:

66.249.72.171 - "GET /%20Select%20An%20Item%20 HTTP/1.1" 404 "-" "Mediapartners-Google/2.1"
Wouldn't you think all those high-priced Google employees would know better?

RED ALERT #2 UPDATE - It's SNAP snapping images

Recently I issued a RED ALERT on BBCOM and it appears it's Snap making screen shots.

I'm positive it's Snap as one of those screen shots has my firewall message displaying the IP address of 65.38.102.146 and the user agent "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1"

Too fucking bad Snap, I'm still blocking your dumb ass until you put your ID on the User Agent from those IPs.

No wonder people think their conversion rates are declining when fucking search engines are even skewing the numbers will bullshit like this unprofessional horseshit.

Tuesday, June 06, 2006

Charlotte, what a web you weave

In the BIG SHOCKER of the day department, when I installed my prototype bot blocker in another of my own websites today to expand my data stream wouldn't you know it but it caught an active spider that was crawling many hundreds of pages the second it was enabled.

This crazy assed spider it still going strong, or thinks it is, well over an hour after stopping it from getting real data.

The spider is Charlotte, with these specs:

209.249.86.4 "Mozilla/5.0 (compatible; Charlotte/1.0b; charlotte@betaspider.com)"
Never heard of Charlotte before so I took a peek to see what others might know about it and stumbled into the most hysterical thread I've ever seen in the OsCommerce Forums about all these store owners frantically chasing down spider names and dropping spider names it into something called their spiders.txt file or some shit.

Dudes, blocking by user agent is so 1990's, you'll stress out and become incontinent doing it that way. What a waste of time, wish I was ready to help you all already, but such is life and quality takes time.

Monday, June 05, 2006

Blog Spam is not a problem

That's right, blog spam isn't the problem whatsoever and technically it shouldn't even exist if it wasn't for the sloppy work of the idiot programmers that write blog software. Anyone leaving the comment forms wide open so that any script kiddie could abuse it should have their programming license revoked.

Now there are even solutions springing up to monitor and stamp out blog spam and it's fucking ridiculous or would hysterical be the proper word?

For example, their feature claims:

AntiSpam Deluxe.
Get rid of referrer and comment spam thanks to a community maintained spam database.
Community maintained database, you must be kidding?

Can you people say CAPTCHA?

I know you can, put a captcha on the comments and registration page to stop automation.

The part that just kills me is whoever was naive enough to put "registration" as the only means of defense without captchas in blogs thinking that making people register to post would stop spam because bots don't use cookies or register. I've noticed a few people resorting to registration alone as an anti-spam technique lately and that simply won't cut it. Wrong again, bots use cookies, so you better put a captcha on your registration page otherwise the bots just register as some random name and post anyway.

Now address manual spamming as follows:
  • Filter out all posts that contain embedded HTML and URLs from all first time posters, just bounce any post that looks like spam.
  • Don't hyper-link names of posters to websites until they become a trusted poster or register.
I do these exact techniques on some of my websites (not this website, except the captcha) and the spammers simply went away. There was nothing they could do of value unless they wanted to just vandalize out of spite, and most are interested in making money so once I removed their incentive they stopped coming back.

Can't post a link to your website?

Oh boo-fucking-hoo, then your spam won't drive visitors, so go the fuck away.

Wow, that was easy wasn't it.
  • No community database
  • No anti-spam service
  • No bullshit
  • No spam
Now you're probably going to start yelling that my concept breaks the premise of blogs in that it's all about the community and the linking and all that other happy horseshit.

Nope, you can still link out, just not on your first damn post, or maybe your first 10 posts. The easiest way to build up "post count" would be using automation which is blocked again, that's where the captcha's come back into play.

If you really want to piss off the spammers, require both registration AND a captcha.

Now go fix your crappy blog software, download a captcha and install it, and stop whining about spam ya bunch of cry babies.

Saturday, June 03, 2006

RED ALERT #5 - Everyones Scraping Internet

There seems to be both legitimate and questionable activity from Everyones Internet so it's been scrutinized a bunch before deciding to issue a warning about this mess.

Basically, to see what was coming out of ev1.net there were a couple of filters installed to see what was coming from them and we got several possible legit bots, some proxy servers, and some wacky crap.

You'll note Chitika listed which is definitely using part of that block:

Everyones Internet EVRY-BLK-15 (NET-67-15-0-0-1)
67.15.0.0 - 67.15.255.255
Chitika, Inc EVRY-398 (NET-67-15-219-0-1)
67.15.219.0 - 67.15.219.63
The linksmanager.com looks possibly legit based on reverse dns (linkchecker02.linksmanager.com), but I think picsearch.com (217.212.245.198) is potentially bogus.

On linksmanager website it says:
LinksManager.com runs an automated Reciprocal Link Checker and Dead Link Checker (User-agent: linksmanager and User-agent: linksmanager_bot ) for all LinksManager customer's web sites. If you are linking with a website that is powered by LinksManager.com, you might see LinksManager.com/linkchecker.html listed in your server log reports.
However, in the log file I see this:
67.15.16.30 Mozilla/5.0 (compatible; LinksManager.com_bot +http://linksmanager.com/linkchecker.html)
So is linksmanager being spoofed, guilty of sloppy outdated documentation or all of the above?

Here's a sample of what I'm seeing:
67.15.0.24
67.15.0.89 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
67.15.119.25 User-Agent: User-Agent: Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0)
67.15.126.25 Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)
67.15.136.199 psbot/0.1 (+http://www.picsearch.com/bot.html)
67.15.138.14 PyQuery / 0.1
67.15.14.5
67.15.143.22 Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)
67.15.16.30 Mozilla/5.0 (compatible; LinksManager.com_bot +http://linksmanager.com/linkchecker.html)
67.15.182.4 Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
67.15.184.3 HTTP/1.0
67.15.184.41 HTTP/1.0
67.15.189.16
67.15.191.19
67.15.2.67 WordPress/1.5.2 PHP/4.4.1
67.15.219.10 Chitika ContentHit 1.0
67.15.219.11 Chitika ContentHit 1.0
67.15.219.12 Chitika ContentHit 1.0
67.15.219.14 Chitika ContentHit 1.0
67.15.219.15 Chitika ContentHit 1.0
67.15.219.16 Chitika ContentHit 1.0
67.15.219.17 Chitika ContentHit 1.0
67.15.219.18 Chitika ContentHit 1.0
67.15.219.3 Chitika ContentHit 1.0
67.15.219.9 Chitika ContentHit 1.0
67.15.221.2 ia_archiver
67.15.221.26 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
67.15.232.3 User-Agent: Mozilla/5.0 (; U;; en-US; rv:1.7.10) Gecko/20050716 Firefox/1.0.6
67.15.35.26 HTTP/1.0
67.15.38.27 Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)
67.15.56.4 Mozilla/5.0 (compatible; LinksManager.com_bot +http://linksmanager.com/linkchecker.html)
67.15.6.64 Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)
67.15.76.148 Mozilla/5.0 (Macintosh; U; PPC Mac OS X; en) AppleWebKit/412.6.2 (KHTML, like Gecko) Safari/412.2.2
67.15.77.119 Mozilla/5.0 (Macintosh; U; PPC Mac OS X; en) AppleWebKit/412.6.2 (KHTML, like Gecko) Safari/412.2.2
67.15.77.223 Mozilla/5.0 (Macintosh; U; PPC Mac OS X; en) AppleWebKit/412.6.2 (KHTML, like Gecko) Safari/412.2.2
67.15.78.93
67.15.8.2 Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)
Did you note the various Netscape/7.1's in the list?

That was yours truly using an outdated Netscape browser very few use to spot proxy servers when I'm testing lists of proxies. Some of the smarter ones mask the browser which makes it's a little more difficult, but then I just check a special page that has never been indexed that only I know about so there is NO hiding from my prying eye.

OK, now you know a new trick, are you happy yet?

Anyway, we cranked this small list thru the reverse DNS meat grinder and here's the results:
24.0.15.67.in-addr.arpa name = hu-tethys.com.
89.0.15.67.in-addr.arpa name = cpanel.masgrafx.com.
25.119.15.67.in-addr.arpa name = ev1s-67-15-119-25.ev1servers.net.
25.126.15.67.in-addr.arpa name = ev1s-67-15-126-25.ev1servers.net.
199.136.15.67.in-addr.arpa name = ev1s-67-15-136-199.ev1servers.net.
14.138.15.67.in-addr.arpa name = ev1s-67-15-138-14.ev1servers.net.
5.14.15.67.in-addr.arpa name = ev1s-67-15-14-5.ev1servers.net.
22.143.15.67.in-addr.arpa name = ev1s-67-15-143-22.ev1servers.net.
30.16.15.67.in-addr.arpa name = linkchecker01.linksmanager.com.
4.182.15.67.in-addr.arpa name = ev1s-67-15-182-4.ev1servers.net.
3.184.15.67.in-addr.arpa canonical name = 3.184.15.67.in-addr.ev1.opticaljungle.com.
3.184.15.67.in-addr.ev1.opticaljungle.com name = lhb-us-b-1.mailhostingserver.com.
41.184.15.67.in-addr.arpa canonical name = 41.184.15.67.in-addr.ev1.opticaljungle.com.
41.184.15.67.in-addr.ev1.opticaljungle.com name = lhb-us-b-2.mailhostingserver.com.
16.189.15.67.in-addr.arpa name = jnchost.net.
19.191.15.67.in-addr.arpa name = ev1s-67-15-191-19.ev1servers.net.
67.2.15.67.in-addr.arpa name = ev1s-67-15-2-67.ev1servers.net.
10.219.15.67.in-addr.arpa name = ev1s-67-15-219-10.ev1servers.net.
11.219.15.67.in-addr.arpa name = ev1s-67-15-219-11.ev1servers.net.
12.219.15.67.in-addr.arpa name = ev1s-67-15-219-12.ev1servers.net.
14.219.15.67.in-addr.arpa name = ev1s-67-15-219-14.ev1servers.net.
15.219.15.67.in-addr.arpa name = ev1s-67-15-219-15.ev1servers.net.
16.219.15.67.in-addr.arpa name = ev1s-67-15-219-16.ev1servers.net.
17.219.15.67.in-addr.arpa name = ev1s-67-15-219-17.ev1servers.net.
18.219.15.67.in-addr.arpa name = ev1s-67-15-219-18.ev1servers.net.
3.219.15.67.in-addr.arpa name = ev1s-67-15-219-3.ev1servers.net.
9.219.15.67.in-addr.arpa name = ev1s-67-15-219-9.ev1servers.net.
2.221.15.67.in-addr.arpa name = arcadehub.com.
26.221.15.67.in-addr.arpa name = ev1s-67-15-221-26.ev1servers.net.
3.232.15.67.in-addr.arpa name = assista.com.
26.35.15.67.in-addr.arpa canonical name = 26.35.15.67.in-addr.ev1.opticaljungle.com.
26.35.15.67.in-addr.ev1.opticaljungle.com name = 67-15-35-26.opticaljungle.com.
27.38.15.67.in-addr.arpa name = web.ir.cx.
4.56.15.67.in-addr.arpa name = linkchecker02.linksmanager.com.
64.6.15.67.in-addr.arpa name = ev1s-67-15-6-64.ev1servers.net.
148.76.15.67.in-addr.arpa name = 67.15.76.148.
119.77.15.67.in-addr.arpa name = 67.15.77.119.
223.77.15.67.in-addr.arpa name = 67.15.77.223.
93.78.15.67.in-addr.arpa name = mail.aucoffre.com.
2.8.15.67.in-addr.arpa name = mail.caromhosting.com.
Since it's a mixed bag and I haven't really decided what to do with this mess yet I'm just blocking anything that responds to reverse DNS as ".ev1servers.net" and taking all the other accesses in this range on a case by case basis.

What a freak'n mess, oh my freak'n head...


Friday, June 02, 2006

Jeteye should Jetpak it up, not ready for primetime

Here we go again with another new web service JetEye that allows you to pack up your links and things you find in some nonsense called a Jetpak and share it with people.

71.5.15.254 [firewall.jeteye.com.] "jeteyebot/0.1; http://www.jeteye.com/bot.html"
I'll give them this much credit, at least they looked in robots.txt before hitting the website.

But I think I'm giving them too much credit for robots.txt, you'll find out at the bottom...

I don't even know what to say about Jeteye as they appear just to be passing around a bunch of links to stuff, except somewhere down the road when a page is moved or changed there will be a flood of 404's with outdated Jetpaks, oh joy.

Who am I kidding, it won't be down the road, I got some 404s with their current Jetpaks, let the bullshit begin!

We give Jeteye a try!

Maybe because it was late Friday afternoon and I was bored shitless, who knows, but I downloaded and installed the damn thing just for shits and giggles.

When I first open Jeteye and click the "sign up here" link the first thing it does is overwrite the current tab in Firefox with their stupid sign-up page.

Opening a new tab must've been too hard for them but I digress...

I fill out the form like the wannabe bleeding-edge netizen that I am and click SUBMIT and it kicks back telling me my verification code failed. OK, would it fucking kill you to tell us in the text above the verification box that the code is case sensitve? Some captchas are sensitive, some aren't, but it's fucking nice to know which is which.

Submit with the new verification code and it kicks back again with something about defining a fucking password. Listen assholes, I typed in a password the FIRST time I filled out the form and you discarded it when you bounced my verification code because I didn't capitalize the fucking "Q". You're starting to get me pissed off already with this bullshit form but I'm trying to remain calm and open minded.

Luckily for them the email validation was swift and went off without a hitch or I'd be going off on a rant about now, so we save the ranting for later.

Jeteye is starting to get on my nerves as clicking on anything in Jeteye keeps zapping my current window. Almost 20 tabs open and which tab is currently being viewed gets blasted when I click on something in Jeteye. They better work on that as I'm about to scream with that behavior, they need some options or some shit, maybe just look and reuse the tab already opened labelled "Jeteye" but that would make too much sense wouldn't it?

Figured this mess out, usability can be a bit challenging and I have a few more complaints but I'm bored writing about this as it's just not quite ready for primetime yet and they aren't paying me to QA the goddamn product. However, I've managed to build my first little Jetpak while watching the request for robots.txt from their site hit my server at almost every action which is pissing me off immensely. Cache the robots.txt for a while, geez, give me a break.

Now the fun stuff:

Just for fun I drop them in the robots.txt file like they say to do to see what happens.

User-agent: jeteyebot
Disallow: /

Well, it did FUCKING NOTHING!

Every link I drop into a Jetpak from my site still shows up in the Jetpak so I assume it means they won't crawl my site but I'm still forced to be an unwilling participant in these goddamn Jetpaks. They still hit the server reading the robots.txt file for every stinking link and it looks like maybe the page I linked as well. It certainly doesn't tell the end user "piss off, this site doesn't want to play with Jetpak" - it just keeps on accessing my site like nothing changed.

Let get this straight, it's MY FUCKING WEBSITE and if I don't want MY LINKS to be in JETPAKS you better give me a way to STOP THIS SHIT!

Guess what else?

It saves direct links to GOOGLESYNDICATION!

Holy Shit! You can save AdSense links in a Jetpak?!?

What a concept, load up a Jetpak full of CPC and Affiliate links and start your click ring Jetpak!

You can drop in some images from any old page that show up stand alone in a Jetpak without attribution to the author. Oh sure, it shows the page of origin but we're not naive and we know that even with copyright notices on a page staring people in the face they steal shit anyway. However, with images in a Jetpak taken out of context they are more prone to get the right mouse "save as..." without even a second thought, especially without any warning about the image being copyrighted or anything.

Taking a little liberty with a phrase from some movie critics I'm giving Jeteye "2 thumbs up"

.... "2 thumbs up the ass" that is, as this shit sucks!

Time to whip up a couple of rewrite rules to block their IP 71.5.15.254 and referer "http://www.jeteye.com/jetpak/" in .htaccess before this gets out of control.

Locating and Blocking Proxy Servers

Since some of my readers want to know how I'm doing it, here's a few tips on how you too can eliminate the anonymous proxies from your site. Probably won't get them all and you might get a few false positives as well but it's better to have some defense against this menace than none at all.

A large number of these proxy servers are on .EDU domains because of all the bleeding heart crap about making information free for all without censorship and being able to surf without fear of retribution. That's a very noble and altruistic motive but you open the doors for competitve spying, scrapers, phishing theives and a lot more so don't take this the wrong way when I don't appreciate what you're doing with our tax and college dollars and send out a big "FUCK YOU" to establishments of higher learning that permit this bullshit. If people in other countries don't like being censored, let them overthrow their fucking government, it's not our problem and my server and copyrighted content shouldn't be vulnerable to attack because of the gaping holes opened up by your bleeding heart asses, but I'm off on a tangent.

The other groups of asshole proxies are the many web-based CGI and PHP proxy servers (like eatmoreblueberries) being used to bypass restricted internet access imposed on corporate, library and school networks. Well I'm sorry but you're supposed to be WORKING or STUDYING so let me give you a big "FUCK YOU" as well. Not only do they download your pages, they strip out YOUR ads and insert their OWN ads, assholes. So for all you slackers using those proxies, zip it up, close the porn sites, go back to work, and get a life you little fuckers as MySpace isn't it.

So, with a bit of ranting aside, back to blocking proxies...

New proxy servers pop up every 5 seconds so my method requires multiple techniques:

  1. Import lists of known proxies and block them
  2. Look for proxy environment variables
  3. Test the IP for typical proxy ports and see if it works
  4. Check for a port number being appended to your domain done by lame proxies
  5. Monitor for proxy crawl thru of known services
1. Import Lists

This step is pretty obvious and can be automated by downloading the lists from a few well known proxy list sites, or if you're lazy you can subscribe to a service or two already doing that.

Probably doesn't hurt to validate these proxies, which can be done automatically, otherwise your list will grow infinitely as they appear and disappear very quicky

2. Proxy Environment Variables


You can check for the following:
HTTP_VIA
HTTP_X_FORWARDED_FOR
HTTP_PROXY_CONNECTION

Yes, those will tell you a proxy made the request but remember that AOL and many others are also a proxy so then it becomes more complicated as you have to evolve a list of known good proxies vs. all the rest and do further processing on those you don't know.

FYI, the really good anonymous proxies don't send that information so you'll never know it's a proxy.

3. Test for Proxy Ports

It will look simple but it's way more complicated to get right.

in PHP you can check to see if you can open port 80 on the incoming IP to see if it's an open proxy like this:

$fp = @fsockopen($theIP, 80, $errno, $errstr, 5);
if ($fp) {
// OPEN PORT
}

But that's very simplistic as most don't use :80, they use port :8080 and other weird #s like :3128, to avoid what the admins are currently blocking.

Not to mention, some proxies are very slow so you want to do exhaustive testing on post-page processing so you don't slow down the user experience on the front end of the page. You only have to do this once per IP, but someone could think your website is down if the process takes to long and worse case you get a positive answer that it's a proxy they've only accessed one page and you block the next page.

Once you detect the proxy add it to your proxy lists built in step 1 above and you'll never have to worry about this one again.

Remember, you may end up blocking IPs from colleges and universities but remember our alma mater, good old FU.

4. Port Numbers Appended to Domain

The dumbest of the dumb append a port number to your domain name which is easy to test in the HTTP_HOST variable. The only exceptions I've have to make to this rule so far is for the poor dumb bastards still using prodigy.net.mx which astonished me that prodigy still existed even as a name on a block of IPs!

5. Proxy Crawl Thru

What some of these dumb fuck proxy operators do is set up a cloaked directory, probably a clone of DMOZ or some shit, and cloak this directory to the search engines.

When you see things like Googlebot, Mediabot, Msnbot, etc. hitting your servers outside of their known range of IP's it means only 1 of 2 possibilities.
  1. Someone is trying to spoof the user agent to get onto your server
  2. The crawler is coming thru a proxy port
The best defense is to throw up an error or something in this case as I've had some pages hijacked by this nonsense so you really don't want to serve up real pages in this event as the search engines simply aren't that smart.

BTW, before serving up an error message, it's wise to do a reverse DNS lookup to make sure that Googlebot really isn't on a new block of IP's owned by google.com.

Summary

Probably not as simple as you had hoped but a couple of techniques are very straight forward and stop some level of the proxy nonsense without fear of blocking innocents.

Good luck trying this and may all your proxy requests bounce off your server like a rock skipping across a pond.

AJAX vs JSON, does anyone care?

Apparently someone cares as the Silicon Valley WebGuild is going to host a program about Yahoo! Web Services Using JSON by Douglas Crockford on June 14th and it'll be hosted at Google.

I might wander down there just to see what Douglas has to say on this topic not to mention they usually put out some appetizers or something at these events, don't know about Google, but when Microsoft used to host it the food was sometimes the best part of the evening!

Not that I condone abandoning XML, but I'll listen to his reasons.

Anyone else in the Bay Area interested in going?

Thursday, June 01, 2006

Stupid Spammers Snared

OK, calling this spammer stupid gives stupid spammers a bad name.

How do I define stupid?

All of the attempted submissions below from "195.225.177.6" don't even have active domains so THAT's fucking STUPID!

I've never really had a problem with spammers before but last week these idiots started using automated submission to pump junk into my site and a couple of hours of programming later they're all blocked.

Here's the first few that hit today:

85.140.17.80 "Mozilla/4.0 (AllSubmitter)" "Business" http://business.alti.ru/
210.214.192.42 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; FunWebProducts)" "sports book for free" http://sportsbookusa.us
210.214.192.42 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; FunWebProducts)" "sports book for free" http://sportsbookusa.us
200.208.239.2 "Mozilla/4.0 (compatible; MSIE 5.01; MSNIA; Windows 98)" "Protonix order" http://stvincent.uzhgorod.ua/anal-masturbation-tip.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://loversweekend.org/~ahmad_3139/files/hornyhousewives.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://lovingvacation.org/~ahmad_3139/files/amateursex.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://realromanticlove.org/~ahmad_3139/files/wifelovers.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://romanticdirect.org/~ahmad_3139/files/amatuersex.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://thefranticromantic.org/~ahmad_3139/files/amatuersex.html
195.225.177.6 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET" "" http://theromanticwoman.org/~ahmad_3139/files/amateurbus.html

If you didn't catch it, block "AllSubmitter" from allowed user agents and stop this fucknut tool from being used on your site.

So sad, too bad, no spam for you.

Monday, May 29, 2006

Linux Rulez, Scraper Droolz

Just couldn't help it, this user agent cracked me up:

195.70.35.179 [palatinus.sanomabp.hu.] requested 3 pages as "KummHttp/1.1 (compatible; KummClient; Linux rulez)"
What exactly is the KummClient?

A porn scraper?

Is it looking for bukkake sites?

Anonymous Proxies Out Cheap Hosts

Now that I've figured out how to accurately detect and block most CGI and PHP proxy servers I'm just sitting here going down the list trying them all and they're leaving my presents.

Yes, they're giving me.... PRESENTS!

What kind of presents you might ask?

Well, how about a list of cheap web hosts that these sites use, also used by scrapers, so I'm picking up a ton of valuable intelligence on these operation in very little time and will be massively expanding the list of blocked hosts to monitor those locations moving forward.

ServePath to more IPs

Just after I blocked the last batch of IPs from ServePath someone popped up on yet a new location ranging from 69.59.128.0 to 69.59.191.255 . All of their reverse DNS starts with "customer-reverse-entry." so I think I'll just zap them by that host address phrase and save some trouble here.

Proxy List Connecting PlanetLab Dots

It's hard to imagine that the same servers hosting PlanetLab experiments are also hosting anonymous proxies, but it appears to be true so now it's hard to say whether it was the PlanetLab network or just the list of proxies that attacked the server last week.

The only clue that it was really the PlanetLab network in that original attack was the entry "pli1-pa-3.hpl.hp.com." as HP has a special hardware deal for PlanetLab members.

Discounted Hardware
We are pleased to announce that HP is now providing a special discount price on PlanetLab machines. PIs will be able to see configuration and pricing information when they log in.
Mostly what has been learned from this proxy list is that institutions of higher education appear to be giving the internet bottom feeders the ability to steal.

Here's the reverse DNS on most of the proxy list that I processed as it's interesting to see who's hosting these things:
1.2.103.142.in-addr.arpa name = planetlab1.cs.ubc.ca.
10.202.108.129.in-addr.arpa name = planetlab1.utep.edu.
10.23.72.132.in-addr.arpa name = planetlab1.bgu.ac.il.
101.65.151.128.in-addr.arpa name = planet1.cs.rochester.edu.
102.65.151.128.in-addr.arpa name = planet2.cs.rochester.edu.
105.150.22.129.in-addr.arpa name = planetlab-2.EECS.CWRU.Edu.
106.139.112.128.in-addr.arpa name = planetlab-10.CS.Princeton.EDU.
108.139.112.128.in-addr.arpa name = planetlab-9.CS.Princeton.EDU.
109.118.83.147.in-addr.arpa name = planetlab2.upc.es.
109.146.68.207.in-addr.arpa name = sasch1031210.phx.gbl.
11.1.31.128.in-addr.arpa name = planetlab1.csail.mit.edu.
11.23.72.132.in-addr.arpa name = planetlab2.bgu.ac.il.
11.36.4.128.in-addr.arpa name = planetlab1.pc.cis.udel.edu.
110.139.112.128.in-addr.arpa name = planetlab-11.CS.Princeton.EDU.
111.120.10.129.in-addr.arpa name = planetlabone.ccs.neu.edu.
111.126.8.128.in-addr.arpa name = salt.planetlab.cs.umd.edu.
111.139.112.128.in-addr.arpa name = planetlab-13.CS.Princeton.EDU.
112.120.10.129.in-addr.arpa name = planetlabtwo.ccs.neu.edu.
112.126.8.128.in-addr.arpa name = pepper.planetlab.cs.umd.edu.
12.36.4.128.in-addr.arpa name = planetlab2.pc.cis.udel.edu.
122.100.44.219.in-addr.arpa name = softbank219044100122.bbtec.net.
123.17.42.194.in-addr.arpa name = planetlab-1.cs.ucy.ac.cy.
124.17.42.194.in-addr.arpa name = planetlab-2.cs.ucy.ac.cy.
125.120.148.210.in-addr.arpa name = catv120-125.lan-do.ne.jp.
126.60.247.140.in-addr.arpa name = righthand.eecs.harvard.edu.
129.20.195.216.cpe.townisp.com name = dhcp-0-11-11-8e-ee-34.cpe.townisp.com.
129.20.195.216.in-addr.arpa canonical name = 129.20.195.216.cpe.townisp.com.
130.182.167.193.in-addr.arpa name = pl-1.hip.fi.
130.208.76.64.in-addr.arpa name = c6476112-130.impsat.com.co.
133.204.23.138.in-addr.arpa name = planet-lab1.cs.ucr.edu.
137.143.128.in-addr.arpa nameserver = athena.cs.Virginia.EDU.
137.143.128.in-addr.arpa nameserver = athena.cs.Virginia.EDU.
138.228.240.129.in-addr.arpa name = planetlab2.simula.no.
14.1.31.128.in-addr.arpa name = planetlab4.csail.mit.edu.
14.63.114.128.in-addr.arpa name = planetslug1.cse.ucsc.edu.
143.109.38.192.in-addr.arpa canonical name = 143.128-191.109.38.192.in-addr.arpa.
143.128-191.109.38.192.in-addr.arpa name = amigos13.distlab.diku.dk.
144.109.38.192.in-addr.arpa canonical name = 144.128-191.109.38.192.in-addr.arpa.
144.128-191.109.38.192.in-addr.arpa name = amigos14.distlab.diku.dk.
145.236.16.220.in-addr.arpa name = softbank220016236145.bbtec.net.
148.12.100.138.in-addr.arpa name = planetlab1.ls.fi.upm.es.
149.11.135.128.in-addr.arpa name = planetlab1.cs.uchicago.edu.
149.12.100.138.in-addr.arpa name = planetlab2.ls.fi.upm.es.
15.1.31.128.in-addr.arpa name = planetlab5.csail.mit.edu.
15.63.114.128.in-addr.arpa name = planetslug2.cse.ucsc.edu.
150.145.245.130.in-addr.arpa name = planetlab1.mnl.cs.sunysb.edu.
152.11.135.128.in-addr.arpa name = planetlab3.cs.uchicago.edu.
152.145.245.130.in-addr.arpa name = planetlab3.mnl.cs.sunysb.edu.
154.127.233.220.in-addr.arpa name = 154.127.233.220.exetel.com.au.
154.159.210.216.in-addr.arpa name = 216-210-159-154.atgi.net.
154.247.194.61.in-addr.arpa canonical name = 154.SUB152.247.194.61.in-addr.arpa.
154.40.161.130.in-addr.arpa name = planetlab2.ewi.tudelft.nl.
154.SUB152.247.194.61.in-addr.arpa name = ns.m-mbc.co.jp.
156.192.6.128.in-addr.arpa name = orbpl2.rutgers.edu.
157.112.20.221.in-addr.arpa name = softbank221020112157.bbtec.net.
157.12.27.220.in-addr.arpa name = softbank220027012157.bbtec.net.
159.19.200.in-addr.arpa nameserver = curau.pop-mg.rnp.br.
159.19.200.in-addr.arpa nameserver = quindim.pop-mg.rnp.br.
16.1.31.128.in-addr.arpa name = planetlab6.csail.mit.edu.
16.63.114.128.in-addr.arpa name = planetslug3.cse.ucsc.edu.
161.33.24.141.in-addr.arpa name = planet1.prakinf.tu-ilmenau.de.
161.76.186.219.in-addr.arpa name = softbank219186076161.bbtec.net.
168.128.38.202.in-addr.arpa name = vw.ihep.ac.cn.
17.1.31.128.in-addr.arpa name = planetlab7.csail.mit.edu.
176.160.37.220.in-addr.arpa name = softbank220037160176.bbtec.net.
178.234.59.200.in-addr.arpa name = inalambrico178-234-regina.neunet.com.ar.
18.75.63.193.in-addr.arpa name = planetlab-1.ic.ac.uk.
181.152.67.219.in-addr.arpa canonical name = 181.176.152.67.219.in-addr.arpa.
181.176.152.67.219.in-addr.arpa name = vcs.tokyuhotel.co.jp.
184.139.in-addr.arpa nameserver = dns0.cs.bham.ac.uk.
184.139.in-addr.arpa nameserver = ns1.susx.ac.uk.
184.139.in-addr.arpa nameserver = ns2.susx.ac.uk.
19.119.216.208.in-addr.arpa name = planetlab1.gti-dsl.nodes.planet-lab.org.
19.75.63.193.in-addr.arpa name = planetlab-2.ic.ac.uk.
190.193.6.207.in-addr.arpa name = d207-6-193-190.bchsia.telus.net.
190.68.30.220.in-addr.arpa name = softbank220030068190.bbtec.net.
191.214.170.129.in-addr.arpa name = planetlab1.cs.dartmouth.edu.
192.214.170.129.in-addr.arpa name = planetlab2.cs.dartmouth.edu.
193.128.77.199.in-addr.arpa name = planet1.cc.gt.atl.ga.us.
193.152.252.132.in-addr.arpa name = planetlab1.exp-math.uni-essen.de.
193.152.252.132.in-addr.arpa name = planetlab1.iem.uni-due.de.
193.152.252.132.in-addr.arpa name = planetlab1.iem.uni-duisburg-essen.de.
194.115.145.136.in-addr.arpa name = planetlab-01.ece.uprm.edu.
194.128.77.199.in-addr.arpa name = planet.cc.gt.atl.ga.us.
196.19.242.129.in-addr.arpa name = planetlab1.cs.uit.no.
197.19.242.129.in-addr.arpa name = planetlab2.cs.uit.no.
197.4.208.128.in-addr.arpa name = planetlab01.cs.washington.edu.
198.4.16.221.in-addr.arpa name = softbank221016004198.bbtec.net.
198.4.208.128.in-addr.arpa name = planetlab02.cs.washington.edu.
199.160.83.130.in-addr.arpa name = planetlab2.rbg.informatik.tu-darmstadt.de.
199.4.208.128.in-addr.arpa name = planetlab03.cs.washington.edu.
199.69.127.216.in-addr.arpa name = gilletts.com.au.
2.142.19.139.in-addr.arpa name = swsat1501.mpi-sws.mpg.de.
2.142.19.139.in-addr.arpa name = planetlab02.mpi-sws.mpg.de.
2.2.103.142.in-addr.arpa name = planetlab2.cs.ubc.ca.
2.60.116.195.in-addr.arpa name = planetlab2.warsaw.rd.tp.pl.
20.102.204.132.in-addr.arpa name = crt1.PLANETLAB.UMontreal.CA.
20.19.252.128.in-addr.arpa name = vn1.cse.wustl.edu.
200.160.83.130.in-addr.arpa name = planetlab3.rbg.informatik.tu-darmstadt.de.
201.4.213.141.in-addr.arpa name = planetlab1.eecs.umich.edu.
201.67.59.128.in-addr.arpa name = planetlab2.comet.columbia.edu.
202.215.117.219.in-addr.arpa name = 219.117.215.202.user.rb.il24.net.
202.4.213.141.in-addr.arpa name = planetlab2.eecs.umich.edu.
202.67.59.128.in-addr.arpa name = planetlab3.comet.columbia.edu.
203.103.232.128.in-addr.arpa name = planetlab3.xeno.cl.cam.ac.uk.
209.218.149.141.in-addr.arpa name = planetlab4-dsl.cs.cornell.edu.
21.19.252.128.in-addr.arpa name = vn2.cse.wustl.edu.
21.254.136.130.in-addr.arpa name = planetlab1.CS.UniBO.IT.
210.130.82.206.in-addr.arpa name = quize.onatel.bf.
210.91.100.12.in-addr.arpa name = 210.mula.mlwk.chcgil24.dsl.att.net.
217.101.192.128.in-addr.arpa name = itchy.cs.uga.edu.
218.101.192.128.in-addr.arpa name = scratchy.cs.uga.edu.
219.135.41.192.in-addr.arpa canonical name = 219.deleg-192.135.41.192.in-addr.arpa.
219.deleg-192.135.41.192.in-addr.arpa name = planetlab2.csg.unizh.ch.
22.102.204.132.in-addr.arpa name = crt3.PLANETLAB.UMontreal.CA.
22.19.252.128.in-addr.arpa name = vn3.cse.wustl.edu.
22.254.136.130.in-addr.arpa name = planetlab2.CS.UniBO.IT.
224.35.228.194.in-addr.arpa name = zakskola.nosovice.indos.cz.
225.17.239.132.in-addr.arpa name = planetlab2.ucsd.edu.
227.202.2.134.in-addr.arpa name = peace.ri.uni-tuebingen.de.
228.202.2.134.in-addr.arpa name = freedom.ri.uni-tuebingen.de.
230.144.90.85.in-addr.arpa name = 85-90-144-230.DSL.ycn.com.
230.152.163.198.in-addr.arpa name = planetlab2.win.trlabs.ca.
231.79.114.140.in-addr.arpa name = pads21.cs.nthu.edu.tw.
233.79.114.140.in-addr.arpa name = pads23.cs.nthu.edu.tw.
235.226.113.128.in-addr.arpa name = planet1.ecse.rpi.edu.
236.229.225.143.in-addr.arpa name = planetlab01.dis.unina.it.
238.229.225.143.in-addr.arpa name = planetlab02.dis.unina.it.
238.75.97.129.in-addr.arpa name = blast.cs.uwaterloo.ca.
242.38.80.194.in-addr.arpa name = planetlab1.cs-ipv6.lancs.ac.uk.
243.198.37.130.in-addr.arpa name = planetlab1.cs.vu.nl.
243.38.80.194.in-addr.arpa name = planetlab2.cs-ipv6.lancs.ac.uk.
244.198.37.130.in-addr.arpa name = planetlab2.cs.vu.nl.
244.47.240.60.in-addr.arpa name = infinitem.com.
246.3.150.142.in-addr.arpa name = planetlab01.erin.utoronto.ca.
247.3.150.142.in-addr.arpa name = planetlab02.erin.utoronto.ca.
249.137.143.128.in-addr.arpa name = planetlab1.cs.Virginia.EDU.
249.99.246.138.in-addr.arpa name = planetlab1.lkn.ei.tum.de.
25.18.246.64.in-addr.arpa name = ev1s-64-246-18-25.ev1servers.net.
25.191.136.193.in-addr.arpa name = planetlab-1.iscte.pt.
250.137.143.128.in-addr.arpa name = planetlab2.cs.Virginia.EDU.
251.70.92.130.in-addr.arpa name = planetlab01.cnds.unibe.ch.
252.70.92.130.in-addr.arpa name = planetlab02.cnds.unibe.ch.
253.253.137.129.in-addr.arpa name = planetlab1.uc.edu.
26.106.199.203.in-addr.arpa name = 203.199.106.26.static.vsnl.net.in.
26.191.136.193.in-addr.arpa name = planetlab-2.iscte.pt.
26.203.88.130.in-addr.arpa name = planet1.manchester.ac.uk.
26.27.9.35.in-addr.arpa name = planetlab1.cse.msu.edu.
27.203.88.130.in-addr.arpa name = planet2.manchester.ac.uk.
28.247.220.128.in-addr.arpa name = planetlab1.isi.jhu.edu.
29.247.220.128.in-addr.arpa name = planetlab2.isi.jhu.edu.
3.142.19.139.in-addr.arpa name = swsat1502.mpi-sws.mpg.de.
3.142.19.139.in-addr.arpa name = planetlab03.mpi-sws.mpg.de.
34.248.207.206.in-addr.arpa name = planetlab1.arizona-gigapop.net.
34.60.116.195.in-addr.arpa name = planetlab2.olsztyn.rd.tp.pl.
35.159.19.200.in-addr.arpa name = planetlab2.pop-mg.rnp.br.
35.248.207.206.in-addr.arpa name = planetlab2.arizona-gigapop.net.
4.20.6.193.in-addr.arpa name = planet1.colbud.hu.
40.127.203.130.in-addr.arpa name = planetlab00.cse.psu.edu.
40.221.49.130.in-addr.arpa name = planetlab1.cs.pitt.edu.
41.127.203.130.in-addr.arpa name = planetlab01.cse.psu.edu.
41.221.49.130.in-addr.arpa name = planetlab2.cs.pitt.edu.
42.188.161.205.in-addr.arpa name = cache.gua.net.
44.146.68.207.in-addr.arpa name = sasch1031305.phx.gbl.
49.60.116.195.in-addr.arpa name = planetlab1.swidnik.rd.tp.pl.
5.142.19.139.in-addr.arpa name = planetlab05.mpi-sws.mpg.de.
5.142.19.139.in-addr.arpa name = swsat1504.mpi-sws.mpg.de.
5.20.6.193.in-addr.arpa name = planet2.colbud.hu.
50.48.217.141.in-addr.arpa name = planetlab1.cs.wayne.edu.
51.249.107.210.in-addr.arpa name = planetlab2.icu.ac.kr.
51.48.217.141.in-addr.arpa name = planetlab2.cs.wayne.edu.
52.19.10.128.in-addr.arpa name = planetlab1.cs.purdue.edu.
52.2.216.144.in-addr.arpa name = planetlab-1.unk.edu.
53.19.10.128.in-addr.arpa name = planetlab2.cs.purdue.edu.
53.2.216.144.in-addr.arpa name = planetlab-2.unk.edu.
55.48.184.139.in-addr.arpa name = planetlab1.rn.informatics.scitech.susx.ac.uk.
56.218.37.80.in-addr.arpa name = 56.Red-80-37-218.staticIP.rima-tde.net.
56.240.11.133.in-addr.arpa name = planetlab1.iii.u-tokyo.ac.jp.
56.68.93.129.in-addr.arpa name = planetlab1.unl.edu.
57.240.11.133.in-addr.arpa name = planetlab2.iii.u-tokyo.ac.jp.
6.142.19.139.in-addr.arpa name = planetlab06.mpi-sws.mpg.de.
6.142.19.139.in-addr.arpa name = swsat1505.mpi-sws.mpg.de.
61.52.111.128.in-addr.arpa name = planet1.cs.ucsb.edu.
62.52.111.128.in-addr.arpa name = planet2.cs.ucsb.edu.
65.60.116.195.in-addr.arpa name = planetlab1.piotrkow.rd.tp.pl.
65.88.238.128.in-addr.arpa name = planetlab2.poly.edu.
66.169.49.62.in-addr.arpa name = no-dns-yet.demon.co.uk.
69.126.8.128.in-addr.arpa name = planetlab2.cs.umd.edu.
69.202.221.201.in-addr.arpa name = 201-221-202-69.bk11-dsl.surnet.cl.
70.0.132.200.in-addr.arpa name = planetlab2.pop-rs.rnp.br.
70.112.179.131.in-addr.arpa name = Planetlab1.CS.UCLA.EDU.
70.255.159.200.in-addr.arpa name = planetlab1.pop-rj.rnp.br.
71.112.179.131.in-addr.arpa name = Planetlab2.CS.UCLA.EDU.
71.139.112.128.in-addr.arpa name = planetlab-1.CS.Princeton.EDU.
71.70.91.139.in-addr.arpa name = planet2.ics.forth.gr.
72.139.112.128.in-addr.arpa name = planetlab-2.CS.Princeton.EDU.
73.139.112.128.in-addr.arpa name = planetlab-3.CS.Princeton.EDU.
74.139.112.128.in-addr.arpa name = planetlab-6.CS.Princeton.EDU.
74.3.12.129.in-addr.arpa name = planetlab1.ukc.ac.uk.
74.44.201.212.in-addr.arpa canonical name = 74.72/29.44.201.212.in-addr.arpa.
74.72/29.44.201.212.in-addr.arpa name = planetlab2.eecs.iu-bremen.de.
75.3.12.129.in-addr.arpa name = planetlab2.ukc.ac.uk.
77.35.248.60.in-addr.arpa name = 60-248-35-77.HINET-IP.hinet.net.
79.109.165.216.in-addr.arpa name = planetx.scs.cs.nyu.edu.
80.139.112.128.in-addr.arpa name = alice.CS.Princeton.EDU.
81.108.54.202.in-addr.arpa name = delhi-202.54.108-81.vsnl.net.in.
81.109.165.216.in-addr.arpa name = planet1.scs.cs.nyu.edu.
82.109.165.216.in-addr.arpa name = planet2.scs.cs.nyu.edu.
82.139.112.128.in-addr.arpa name = planetlab-8.CS.Princeton.EDU.
82.56.227.128.in-addr.arpa name = planetlab2.acis.ufl.edu.
82.60.116.195.in-addr.arpa name = planetlab1.krakow.rd.tp.pl.
83.60.116.195.in-addr.arpa name = planetlab2.krakow.rd.tp.pl.
87.11.99.137.in-addr.arpa name = planetlab2.engr.uconn.edu.
87.2.17.66.in-addr.arpa name = 66-17-2-87.biz.bkfd.arrival.net.
90.150.22.129.in-addr.arpa name = planetlab-1.EECS.CWRU.Edu.
91.112.214.128.in-addr.arpa name = planetlab1.hiit.fi.
92.112.214.128.in-addr.arpa name = planetlab2.hiit.fi.
96.139.112.128.in-addr.arpa name = planetlab-4.CS.Princeton.EDU.
97.139.112.128.in-addr.arpa name = planetlab-5.CS.Princeton.EDU.

Someone else has also published a list showing PlanetLab proxies all over the place.

Saturday, May 27, 2006

Scrapers Impacting Conversion Rates?

Was just reading on Threadwatch about the Shop.org report on the decline in conversion rates for online stores and suddenly had an epihany that scrapers may be involved in this equation.

Let's assume that these conversion rate facts and figures include many of the non-human stealth crawlers that I'm blocking on a daily basis. There's no way your average online retailer is probably aware of this situation and you know they're being scraped just like the rest of us, maybe even scraped MORE than the rest of us, who knows.

Using one of my websites as an example, it averages 13,500 visitors a day and 50-200 stealth crawlers are being blocked which accounts for .5% - 1.5% of my daily traffic, which would definitely impact the conversion rate for any store with similar traffic.

Perhaps scrapers on a very large website getting a million visitors a day wouldn't have much impact unless the site attracts a lot more scrapers than my site. However, a smaller online retailer with similar traffic to the site I'm protecting would obviously notice a difference in their conversion rate, a HUGE difference, just by adjusting their stats to include pages downloaded by stealth crawlers.

Just another example of how scraping and stealth crawling is BAD FOR THE WEB and needs to be stopped.

ServePath to being banned

Found a bunch of random stuff coming from a hosting company called ServePath today while running historical analysis on a batch of IPs.

Now these are the visible crawlers that came from ServePath:

64.151.75.252 PEAR HTTP_Request class ( http://pear.php.net/ )
64.151.64.212 "Jakarta Commons-HttpClient/3.0"
64.151.65.12 "Jakarta Commons-HttpClient/3.0"
64.151.111.116 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0)
64.151.112.44 NutchCVS/0.7.1 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)
Here's the whole range:
OrgName: ServePath, LLC
NetRange: 64.151.64.0 - 64.151.127.255
CIDR: 64.151.64.0/18
I'm going to block the whole thing and see if there are any stealth crawlers operating out of that location that haven't tripped any alarms yet and see what happens.

Amazon's A9 Amateur Hour

Guess what boys and girls?

We've all been forced to OPT-IN to yet another non-standard "web tool" that Amazon's A9 has thrust upon us. A9's blog said they introduced this crap last July but it's obviously been so low key compared to everything else hitting my server that I overlooked this small slice of idiocy.

This has been showing up in my logs for a while now"
207.171.167.25 - “GET /siteinfo.xml HTTP/1.1″ 404 1639 “-” “Java/1.5.0_04″
The only reason I noticed it today was the amount of times it hit the server escalated and they're racking up a bunch of 404 errors requesting this file I've never heard about which is idiotic and stupid.

Ever hear of any internet standard such as ROBOTS.TXT to see if I even want you looking for this stupid file on my server?

Apparently not as the only file being hit is "siteinfo.xml".

Had to resort to a reverse DNS lookup just to find out it was iad-fw-global.amazon.com who was doing this stupid crap. Didn't the vaudeville programmers that wrote this joke ever hear of setting the USER AGENT to identify who and what this is instead of Java/1.5?

Amazon, if you happen to read this pay very close attention to the fact that many web applications bombard my server with the user agent of "Java/1.whatever" on a daily basis which are all BLOCKED so you will never ever get access to siteinfo.xml until you properly identify yourself.

Here's a sample "siteinfo.xml" file that you can install in your root web directory:
<?xml version="1.0"?>
<siteinfo xmlns="http://a9.com/-/spec/siteinfo/1.0/">
<webmenu>
<name>Amazon SiteInfo Sucks</name>
<menu>
<item>
<text>Doesn't use standards</text>
<url>http://www.robotstxt.org/</url>
</item>
<item>
<text>Doesn't identify itself</text>
<url>http://www.mozilla.org/build/revised-user-agent-strings.html</url>
</item>
</menu>
</webmenu>
</siteinfo>
I commented about their lack of professionalism and standards being used in this implementation on their blog but it's awaiting moderation and I doubt they'll let my less than happy comments be published, but we shall see.

Thursday, May 25, 2006

PlanetLabs Bombards Server - Abused or Compromised?

Well here's a new one that was uncovered this week when a tipster wishing to remain anonymous sent me a very suspicious looking log file snippet with a bunch of identical accesses from over 130 IP addresses ranging over a couple of hours.

After doing a little bit of research it looks like this "attack" came from a consortium of computers called PlanetLab located in various universities and research institutions around the world and this appears to be only a portion of the network that was aimed at our tipsters server. We don't know at this point if this was an isolated demonstration of their network, whether they were being abused by a member or if a hacker has breeched the protocol, but the potential for damage here is huge.

Their website claims the following stats:

PlanetLab currently consists of 668 machines, hosted by 325 sites, spanning over 25 countries. Most of the machines are hosted by research institutions, although some are located in co-location and routing centers (e.g., on Internet2's Abilene backbone). All of the machines are connected to the Internet. The goal is for PlanetLab to grow to 1,000 widely distributed nodes that peer with the majority of the Internet's regional and long-haul backbones.

Below are sample of the log files, IPs involved, and the reverse DNS of all the IPs which is what we used to figure out this was probably PlanetLab. There were other files accessed as well, but browsers don't typically look at robots.txt so that's all we needed to suspect something was wrong with this situation and treated it as a potential attack.

If this was an actual PlanetLab project aimed at crawling the web undetected and aggregate tons of data, then it failed miserably. Now that we know who you are and where you are, our servers will be watching to see if you strike again.

If this was an unauthorized test then PlanetLab better beef up security as this network is one big DDoS attack just waiting to happen under control of the wrong person.

Here's a sample snippet of the log file:
216.165.109.81 - - [11/May/2006:08:45:18 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
216.165.109.82 - - [11/May/2006:08:45:18 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
216.165.109.79 - - [11/May/2006:08:45:18 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
141.161.20.32 - - [11/May/2006:08:45:59 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
138.100.12.149 - - [11/May/2006:08:46:03 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
138.100.12.148 - - [11/May/2006:08:46:03 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
199.77.128.193 - - [11/May/2006:08:46:38 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
199.77.128.194 - - [11/May/2006:08:46:39 -0400] "GET /robots.txt HTTP/1.1" 200 452 "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.10) Gecko/20050720 Fedora/1.0.6-1.1.fc3 Firefox/1.0.6" "-"
Here's the complete list of IP's involved:
216.165.109.81
216.165.109.82
216.165.109.79
141.161.20.32
138.100.12.149
138.100.12.148
199.77.128.193
199.77.128.194
128.31.1.11
128.31.1.15
128.31.1.14
128.31.1.16
138.23.204.232
138.23.204.133
128.31.1.12
128.31.1.13
130.245.145.152
128.83.122.181
128.83.122.180
152.3.138.2
152.3.138.3
128.220.247.28
12.46.129.21
169.229.50.13
169.229.50.10
169.229.50.17
12.46.129.22
169.229.50.8
12.46.129.23
169.229.50.9
169.229.50.12
169.229.50.16
198.133.224.145
129.115.248.225
194.36.10.154
194.36.10.156
128.208.4.199
171.66.3.181
205.189.33.178
193.6.20.4
193.6.20.5
129.170.214.191
129.170.214.192
130.37.198.243
128.143.137.250
128.111.52.62
194.80.38.242
194.80.38.243
131.234.66.161
131.234.66.160
132.227.74.40
169.229.50.18
169.229.50.7
129.186.205.77
155.98.35.4
155.98.35.3
130.104.72.201
130.104.72.200
128.227.56.82
132.252.152.193
147.83.118.123
147.83.118.124
147.83.118.109
147.83.118.125
130.88.203.26
130.88.203.27
193.10.64.36
142.103.2.2
142.103.2.1
193.10.133.128
193.1.201.26
212.201.44.74
133.11.240.56
133.11.240.57
193.167.182.130
132.72.23.11
132.72.23.10
195.116.60.82
195.116.60.83
132.204.102.20
132.204.102.22
130.203.127.40
130.203.127.41
129.242.19.196
169.229.50.11
128.252.19.21
129.242.19.197
129.22.150.105
138.251.214.18
138.251.214.19
128.151.65.101
128.151.65.102
193.63.75.19
134.76.81.241
134.76.81.242
128.59.67.202
130.136.254.22
210.125.84.16
210.125.84.15
140.109.17.181
200.132.0.70
195.116.60.65
204.123.28.53
131.188.44.101
128.8.126.112
128.8.126.69
128.8.126.111
131.246.19.202
163.221.11.73
163.221.11.71
163.221.11.72
193.144.21.130
193.144.21.131
165.230.49.114
165.230.49.115
143.248.139.170
128.232.103.201
128.232.103.203
202.249.37.212
192.33.210.16
193.136.191.26
193.136.191.25
130.49.221.41
192.17.239.251
192.17.239.250
192.41.135.218
192.41.135.219
142.150.3.247
142.150.3.246
200.159.255.70
128.232.103.202
134.226.52.34
134.226.52.35
To make sense of all this mess, I crunched them all thru NSLOOKUP to see if any patterns emerged and what was a common theme was .EDU and PLANETLAB all over the place.

Here's the reverse DNS on all the IPs for your viewing pleasure:
35.52.226.134.in-addr.arpa name = planetlab02.cs.tcd.ie.
81.109.165.216.in-addr.arpa name = planet1.scs.cs.nyu.edu.
82.109.165.216.in-addr.arpa name = planet2.scs.cs.nyu.edu.
79.109.165.216.in-addr.arpa name = planetx.scs.cs.nyu.edu.
32.20.161.141.in-addr.arpa name = planetlab1.georgetown.edu.
149.12.100.138.in-addr.arpa name = planetlab2.ls.fi.upm.es.
148.12.100.138.in-addr.arpa name = planetlab1.ls.fi.upm.es.
193.128.77.199.in-addr.arpa name = planet1.cc.gt.atl.ga.us.
194.128.77.199.in-addr.arpa name = planet.cc.gt.atl.ga.us.
11.1.31.128.in-addr.arpa name = planetlab1.csail.mit.edu.
15.1.31.128.in-addr.arpa name = planetlab5.csail.mit.edu.
14.1.31.128.in-addr.arpa name = planetlab4.csail.mit.edu.
16.1.31.128.in-addr.arpa name = planetlab6.csail.mit.edu.
232.204.23.138.in-addr.arpa name = planet-lab2.cs.ucr.edu.
133.204.23.138.in-addr.arpa name = planet-lab1.cs.ucr.edu.
12.1.31.128.in-addr.arpa name = planetlab2.csail.mit.edu.
13.1.31.128.in-addr.arpa name = planetlab3.csail.mit.edu.
152.145.245.130.in-addr.arpa name = planetlab3.mnl.cs.sunysb.edu.
181.122.83.128.in-addr.arpa name = planetlab3.csres.utexas.edu.
180.122.83.128.in-addr.arpa name = planetlab2.csres.utexas.edu.
2.138.3.152.in-addr.arpa name = planetlab2.cs.duke.edu.
3.138.3.152.in-addr.arpa name = planetlab3.cs.duke.edu.
28.247.220.128.in-addr.arpa name = planetlab1.isi.jhu.edu.
21.129.46.12.in-addr.arpa canonical name = 21.0/25.129.46.12.in-addr.arpa.
21.0/25.129.46.12.in-addr.arpa name = planet1.berkeley.intel-research.net.
13.50.229.169.in-addr.arpa name = planetlab11.Millennium.Berkeley.EDU.
10.50.229.169.in-addr.arpa name = planetlab8.Millennium.Berkeley.EDU.
17.50.229.169.in-addr.arpa name = planetlab15.Millennium.Berkeley.EDU.
22.129.46.12.in-addr.arpa canonical name = 22.0/25.129.46.12.in-addr.arpa.
22.0/25.129.46.12.in-addr.arpa name = planet2.berkeley.intel-research.net.
8.50.229.169.in-addr.arpa name = planetlab6.Millennium.Berkeley.EDU.
23.129.46.12.in-addr.arpa canonical name = 23.0/25.129.46.12.in-addr.arpa.
23.0/25.129.46.12.in-addr.arpa name = planet3.berkeley.intel-research.net.
9.50.229.169.in-addr.arpa name = planetlab7.Millennium.Berkeley.EDU.
12.50.229.169.in-addr.arpa name = planetlab10.Millennium.Berkeley.EDU.
16.50.229.169.in-addr.arpa name = planetlab14.Millennium.Berkeley.EDU.
145.224.133.198.in-addr.arpa name = planetlab1.cs.wisc.edu.
225.248.115.129.in-addr.arpa name = pl1a.pl.utsa.edu.
154.10.36.194.in-addr.arpa name = planetlab1.nrl.dcs.qmul.ac.uk.
156.10.36.194.in-addr.arpa name = planetlab2.nrl.dcs.qmul.ac.uk.
199.4.208.128.in-addr.arpa name = planetlab03.cs.washington.edu.
181.3.66.171.in-addr.arpa name = planet1.scs.stanford.edu.
178.33.189.205.in-addr.arpa name = planet1.ottawa.canet4.nodes.planet-lab.org.
4.20.6.193.in-addr.arpa name = planet1.colbud.hu.
5.20.6.193.in-addr.arpa name = planet2.colbud.hu.
191.214.170.129.in-addr.arpa name = planetlab1.cs.dartmouth.edu.
192.214.170.129.in-addr.arpa name = planetlab2.cs.dartmouth.edu.
243.198.37.130.in-addr.arpa name = planetlab1.cs.vu.nl.
250.137.143.128.in-addr.arpa name = planetlab2.cs.Virginia.EDU.
62.52.111.128.in-addr.arpa name = planet2.cs.ucsb.edu.
242.38.80.194.in-addr.arpa name = planetlab1.cs-ipv6.lancs.ac.uk.
243.38.80.194.in-addr.arpa name = planetlab2.cs-ipv6.lancs.ac.uk.
161.66.234.131.in-addr.arpa name = planetlab-2.cs.upb.de.
160.66.234.131.in-addr.arpa name = planetlab-1.cs.upb.de.
40.74.227.132.in-addr.arpa name = planetlab-01.ipv6.lip6.fr.
18.50.229.169.in-addr.arpa name = planetlab16.Millennium.Berkeley.EDU.
7.50.229.169.in-addr.arpa name = planetlab5.Millennium.Berkeley.EDU.
77.205.186.129.in-addr.arpa name = planetlab-4.ece.iastate.edu.
4.35.98.155.in-addr.arpa name = planetlab3.flux.utah.edu.
3.35.98.155.in-addr.arpa name = planetlab2.flux.utah.edu.
201.72.104.130.in-addr.arpa name = planetlab2.info.ucl.ac.be.
200.72.104.130.in-addr.arpa name = planetlab1.info.ucl.ac.be.
82.56.227.128.in-addr.arpa name = planetlab2.acis.ufl.edu.
193.152.252.132.in-addr.arpa name = planetlab1.iem.uni-duisburg-essen.de.
193.152.252.132.in-addr.arpa name = planetlab1.exp-math.uni-essen.de.
193.152.252.132.in-addr.arpa name = planetlab1.iem.uni-due.de.
123.118.83.147.in-addr.arpa name = planetlab3.upc.es.
124.118.83.147.in-addr.arpa name = planetlab4.upc.es.
109.118.83.147.in-addr.arpa name = planetlab2.upc.es.
125.118.83.147.in-addr.arpa name = planetlab5.upc.es.
26.203.88.130.in-addr.arpa name = planet1.manchester.ac.uk.
27.203.88.130.in-addr.arpa name = planet2.manchester.ac.uk.
36.64.10.193.in-addr.arpa name = planetlab2.sics.se.
2.2.103.142.in-addr.arpa name = planetlab2.cs.ubc.ca.
1.2.103.142.in-addr.arpa name = planetlab1.cs.ubc.ca.
128.133.10.193.in-addr.arpa name = planetlab-1.it.uu.se.
26.201.1.193.in-addr.arpa name = planetlab-1.tssg.org.
74.44.201.212.in-addr.arpa canonical name = 74.72/29.44.201.212.in-addr.arpa.
74.72/29.44.201.212.in-addr.arpa name = planetlab2.eecs.iu-bremen.de.
56.240.11.133.in-addr.arpa name = planetlab1.iii.u-tokyo.ac.jp.
57.240.11.133.in-addr.arpa name = planetlab2.iii.u-tokyo.ac.jp.
130.182.167.193.in-addr.arpa name = pl-1.hip.fi.
11.23.72.132.in-addr.arpa name = planetlab2.bgu.ac.il.
10.23.72.132.in-addr.arpa name = planetlab1.bgu.ac.il.
82.60.116.195.in-addr.arpa name = planetlab1.krakow.rd.tp.pl.
83.60.116.195.in-addr.arpa name = planetlab2.krakow.rd.tp.pl.
20.102.204.132.in-addr.arpa name = crt1.PLANETLAB.UMontreal.CA.
22.102.204.132.in-addr.arpa name = crt3.PLANETLAB.UMontreal.CA.
40.127.203.130.in-addr.arpa name = planetlab00.cse.psu.edu.
41.127.203.130.in-addr.arpa name = planetlab01.cse.psu.edu.
196.19.242.129.in-addr.arpa name = planetlab1.cs.uit.no.
11.50.229.169.in-addr.arpa name = planetlab9.Millennium.Berkeley.EDU.
21.19.252.128.in-addr.arpa name = vn2.cse.wustl.edu.
197.19.242.129.in-addr.arpa name = planetlab2.cs.uit.no.
105.150.22.129.in-addr.arpa name = planetlab-2.EECS.CWRU.Edu.
18.214.251.138.in-addr.arpa name = planetlab1.dcs.st-and.ac.uk.
19.214.251.138.in-addr.arpa name = planetlab2.dcs.st-and.ac.uk.
101.65.151.128.in-addr.arpa name = planet1.cs.rochester.edu.
102.65.151.128.in-addr.arpa name = planet2.cs.rochester.edu.
19.75.63.193.in-addr.arpa name = planetlab-2.ic.ac.uk.
241.81.76.134.in-addr.arpa name = planetlab1.informatik.uni-goettingen.de.
242.81.76.134.in-addr.arpa name = planetlab2.informatik.uni-goettingen.de.
202.67.59.128.in-addr.arpa name = planetlab3.comet.columbia.edu.
22.254.136.130.in-addr.arpa name = planetlab2.CS.UniBO.IT.
** server can't find 16.84.125.210.in-addr.arpa: NXDOMAIN
** server can't find 15.84.125.210.in-addr.arpa: NXDOMAIN
181.17.109.140.in-addr.arpa name = planetlab2.iis.sinica.edu.tw.
70.0.132.200.in-addr.arpa name = planetlab2.pop-rs.rnp.br.
65.60.116.195.in-addr.arpa name = planetlab1.piotrkow.rd.tp.pl.
53.28.123.204.in-addr.arpa name = pli1-pa-3.hpl.hp.com.
101.44.188.131.in-addr.arpa name = planetlab2.informatik.uni-erlangen.de.
112.126.8.128.in-addr.arpa name = pepper.planetlab.cs.umd.edu.
69.126.8.128.in-addr.arpa name = planetlab2.cs.umd.edu.
111.126.8.128.in-addr.arpa name = salt.planetlab.cs.umd.edu.
202.19.246.131.in-addr.arpa name = planetlab2.informatik.uni-kl.de.
73.11.221.163.in-addr.arpa name = planetlab-03.naist.jp.
71.11.221.163.in-addr.arpa name = planetlab-01.naist.jp.
72.11.221.163.in-addr.arpa name = planetlab-02.naist.jp.
130.21.144.193.in-addr.arpa name = planetlab.urv.net.
131.21.144.193.in-addr.arpa name = planetlab2.urv.net.
114.49.230.165.in-addr.arpa name = planetlab1.rutgers.edu.
115.49.230.165.in-addr.arpa name = planetlab2.rutgers.edu.
170.139.248.143.in-addr.arpa name = csplanetlab3.kaist.ac.kr.
201.103.232.128.in-addr.arpa name = planetlab1.xeno.cl.cam.ac.uk.
203.103.232.128.in-addr.arpa name = planetlab3.xeno.cl.cam.ac.uk.
212.37.249.202.in-addr.arpa name = planetlab2.koganei.wide.ad.jp.
16.210.33.192.in-addr.arpa name = lsirextpc01.epfl.ch.
26.191.136.193.in-addr.arpa name = planetlab-2.iscte.pt.
25.191.136.193.in-addr.arpa name = planetlab-1.iscte.pt.
41.221.49.130.in-addr.arpa name = planetlab2.cs.pitt.edu.
251.239.17.192.in-addr.arpa name = planetlab2.cs.uiuc.edu.
250.239.17.192.in-addr.arpa name = planetlab1.cs.uiuc.edu.
218.135.41.192.in-addr.arpa canonical name = 218.deleg-192.135.41.192.in-addr.arpa.
218.deleg-192.135.41.192.in-addr.arpa name = planetlab1.csg.unizh.ch.
219.135.41.192.in-addr.arpa canonical name = 219.deleg-192.135.41.192.in-addr.arpa.
219.deleg-192.135.41.192.in-addr.arpa name = planetlab2.csg.unizh.ch.
247.3.150.142.in-addr.arpa name = planetlab02.erin.utoronto.ca.
246.3.150.142.in-addr.arpa name = planetlab01.erin.utoronto.ca.
70.255.159.200.in-addr.arpa name = planetlab1.pop-rj.rnp.br.
202.103.232.128.in-addr.arpa name = planetlab2.xeno.cl.cam.ac.uk.
34.52.226.134.in-addr.arpa name = planetlab01.cs.tcd.ie.
35.52.226.134.in-addr.arpa name = planetlab02.cs.tcd.ie.
Best we can tell it was definitely PlanetLab involved with this and I'm very upset that an organization like this would aim a large section of their network at a single server at the same time without permission.

This is abuse, pure and simple, without proper user agent attribution or anything, and I welcome them to come here and let us know what really happened.

While we're waiting on PlanetLab to respond, and I wouldn't hold my breath, I'm going to block the IPs listed above and probably ban anything with "planetlab" or "planet-lab" in the reverse DNS location name until further notice.

RED ALERT #4 - NiceBot Neighborhood

Found another distributed IP batch sitting in a hosting farm claiming to be "nicebot".

Nicebot my ass...

Here's the range of IP's spotted with user agent nicebot:

69.60.120.165 - nicebot
69.60.120.166 - nicebot
69.60.120.167 - nicebot
69.60.120.168 - nicebot
69.60.120.169 - nicebot
69.60.120.172 - nicebot
69.60.120.173 - nicebot
69.60.120.174 - nicebot
69.60.120.176 - nicebot
NSLOOKUP claims they belong to ServerPronto.
nslookup 69.60.120.169
Server: 64.34.160.92
Address: 64.34.160.92#53

Non-authoritative answer:
169.120.60.69.in-addr.arpa name = 169-120-60-69.serverpronto.com.
So I think I'm going to just block this range from ServerPronto as it's a hosting farm:
Serverpronto INMM-69-60-114-0 (NET-69-60-114-0-1)
69.60.114.0 - 69.60.125.255
Some of you might naively think that you can just block "nicebot" with rewrite rules and solve your problem. However, my research has shown that many of these bots eventually change names when they get blocked by too many sites. You're best off blocking the source permanently so they don't slip thru the cracks next week crawling as something like ""Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Q312461; BTOW V9.0; SV1)" which you can't detect.

Remember, they're desperate when you cut off their source of revenue and they'll attempt to adapt so use the best prevention up front which is lock them out by location and don't waste your time fighting changing user agent names.

Bots gone WILD!

This is just a follow up on a couple of the bots using distributed IP's I've highlighted recently which just won't take NO for an answer. Ever since their little cluster of scraping IPs has been uncovered and blocked it's still been a non-stop daily request for hundreds of pages per scraper.

These bots are very nasty so if you weren't paying attention the first time, go back and block THIS, THIS and THIS as they are some hungry-assed bots that need to be stopped.

Wednesday, May 24, 2006

BEZEQINT-HOSTING has a scraper

Coming from the lovely land of Israel is a scraper from Bezeq International, and I can't tell if this is a hosting IP or a DSL connection, but I'm guessing it's hosting but could just be DHCP.

Like I can read their website, feh!

Anyway, the bot always claims to be:

Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; (R1 1.1); .NET CLR 1.1.4322)
Comes from the following addresses:
82.80.249.201
82.80.249.203
82.80.249.204
So block those at a minimum and 82.80.249.0/24 if you want to be safe.

Ta ta, no scrape for you!

Monday, May 22, 2006

Odd traffic from Hong Kong, Middle-East and Africa

Has anyone noticed any huge spikes in traffic from Hong Kong, Saudi Arabia, South Africa or Dubai lately?

They are setting off alarms all over the place with my bot blocker and it's all coming from shared networks so I can't tell yet if it's just a lot of people using a few IPs or a few crawlers going crazy in a scraper haven.

I'm thinking about just setting the whole bunch of them to "CAPTCHA-mode" which is the equivalent of forcing them to login before accessing my site. This will quickly determine the number source of the activity based on the number of unanswered CAPTCHA's vs. a valid response from a human.

Let's see what happens next, I'll keep you posted ;)

RED ALERT #3 - GoDaddy hosting distributed scraper

This one may have just moved to a new location as I've been watching similar activity before which stopped. These new antics have been going on at this location for a week now and I waited just to make sure it was really coming from a common location which appears to be a block of IPs on some GoDaddy hosting farm secureserver.net.

This creepy crawler doesn't use any user agent string whatsoever and keeps asking for pages like "/#top" and other stupid stuff. Below is the range of IPs and the number of pages asked for just today. You'll note it was a slow day for them asking for only 75 pages, but the day isn't over yet.

68.178.242.111 [ip-68-178-242-111.ip.secureserver.net.] requested 30 pages as ""
68.178.242.126 [ip-68-178-242-126.ip.secureserver.net.] requested 15 pages as ""
68.178.242.128 [ip-68-178-242-128.ip.secureserver.net.] requested 15 pages as ""
68.178.242.127 [ip-68-178-242-127.ip.secureserver.net.] requested 15 pages as ""
Performed an nslookup and got this:
nslookup 68.178.242.111
Server: 64.34.160.92
Address: 64.34.160.92#53

Non-authoritative answer:
111.242.178.68.in-addr.arpa name = ip-68-178-242-111.ip.secureserver.net.

When I did a whois on the IP there came the surprise:
[Querying whois.arin.net]
[whois.arin.net]

OrgName: Go Daddy Software, Inc.
OrgID: GDS-31
Address: 14455 N Hayden Road
Address: Suite 226
City: Scottsdale
StateProv: AZ
PostalCode: 85260
Country: US
178.128.0 - 178.255.255

Now do a whois on secureserver.net:
NetRange: Registrant:
Special Domain Services, Inc.
14455 N Hayden Rd
Scottsdale, Arizona 85260
United States

Registered through: WWDomains.com
Domain Name: SECURESERVER.NET
Created on: 30-Mar-98
Expires on: 29-Mar-12
Last Updated on: 07-Feb-06

Not sure it makes sense to block the entire GoDaddy IP range, so for now 68.178.242.0/24 is all I'm blocking unless I see more rogue activity in their network.

BTW, anyone notice how many sneaky crawler networks I'm busting now that I have proximity alarms in place to spot organized activity?

This proximity alarm is great as it doesn't care if the crawlers ask for 1 page or 100 pages, the minute it detects multiple IP addresses in a similar range doing these things it pops up on my radar. The best thing is that the distributed crawler doesn't even have to use more than one IP address per day as long as they break one of my "bad bot rules" on each visit so the IP is flagged and archived. The proximity report of archived bad bot activity will then expose those archived bots operating from a single location.

Pretty tricky, eh?

You stupid bots better wise up quick, you can't hide behind a bank of IPs, your days are numbered!

Sunday, May 21, 2006

Publicly Available Website

That's the current buzzword most often used when you confront someone crawling your site, especially a corporation, that it's a "Publicly Available Website".

Well just because something is publicly available doesn't mean you have the right to do whatever you like with it. It's publicly available for the PUBLIC, meaning visitors, to read individual pages and it's also available to the 6 search engines that I permit to crawl my site. Other than that, just like any other publicly available business, I have the RIGHT TO RESTRICT ACCESS to anyone else that I so desire.

For instance many brick and mortar businesses say "No Shoes, No Shirt, No Service".

Well my website has similar rules "No Humans, No Permission, No Service".

If I even get a whiff off a robot on the site, permission denied.

You corporate and private scrapers just better get over your loser mantra as putting a website online, even on a public network, does NOT give everyone complete access to do whatever they feel like with your site. There are terms of service on that site which distinctly prohibit the use of unauthorized tools to crawl that site, and if you have to ask what's authorized then you don't have permission in the first place so go away.

The site doesn't have a "GNU Free Documentation License", instead it has one of those funny things called a "copyright" which means I own it, not YOU. Additionally, I pay for the server, not YOU. Which means, it's up to ME what is and isn't allowed, even when it's a "Publicly Available Website", NOT YOU!

Let's make it so simple even a 2 year old can understand it:

The website is MINE! MINE! MINE! ALL MINE! and NOT YOURS!

Is that language clear enough for the mental midgets scraping the web to comprehend?

Saturday, May 20, 2006

RED ALERT #2 - Distributed IP Scraper on BBCOM

This one is virtually identical to scraper spotted on Vericenter, same profile to the letter.

Claims to be the same exact browser:

Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1
This spider was seen operating from these IP addresses:
66.234.139.195
66.234.139.200
66.234.139.203
66.234.139.204
66.234.139.206
66.234.139.207
66.234.139.211
66.234.139.213
66.234.139.216
66.234.139.217
This IP range belongs to BBCOM:
OrgName: Backbone Communications, Inc.
OrgID: BBCM
Address: 515 South Flower Street
Address: Suite 4350
City: Los Angeles
StateProv: CA
PostalCode: 90071
Country: US

NetRange: 66.234.128.0 - 66.234.159.255
I'm blocking 66.234.139.0/24 for the moment, but keeping an eye on the rest of their network in case these scrapers switch to a new block.

Another Proximity Alert on PCCW

Getting multiple hits from the same IPs over and over asking for the exact same web page, no images, nothing else, just the same web page day after day. One day the user agent was blank and another day simply"Mozilla/5.0".

This is the range of IPs that just keep repeating the same request.

210.87.251.111
210.87.251.107
210.87.251.41
210.87.251.106
Appears to be coming from HK, can't tell if it's a shared DHCP situation or what, but I'm blocking 210.87.251.0/24 just to be safe.

Another DUMBASS Crawler

This one hails from South Africa, tried to crawl "/#top" as "GET /%23top"

You poor dumb fucker, BUSTED!

RED ALERT - Distributed IP Scraper Hosted on Vericenter

Here's a real sneaky scraper using distributed IPs that is using a bot that almost appears designed to fly under my bot blockers radar. No single IP address accessed enough pages or did anything obnoxious enough to set off any triggers but the collective accesses set off a proximity alarm and they got nailed anyway.

The scraper is pretending to be Firefox for Linux:

http://www.Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.8.0.1) Gecko/20060124 Firefox/1.5.0.1/
The range of IP's noticed in this scrape attack are as follows:
65.38.102.138 ip-65-38-102-138.hou.vericenter.com
65.38.102.139 ip-65-38-102-139.hou.vericenter.com
65.38.102.141 ip-65-38-102-141.hou.vericenter.com
65.38.102.143 ip-65-38-102-143.hou.vericenter.com
65.38.102.145 ip-65-38-102-145.hou.vericenter.com
65.38.102.146 ip-65-38-102-146.hou.vericenter.com
65.38.102.147 ip-65-38-102-147.hou.vericenter.com
65.38.102.148 ip-65-38-102-148.hou.vericenter.com
65.38.102.150 ip-65-38-102-150.hou.vericenter.com
65.38.102.153 ip-65-38-102-153.hou.vericenter.com
65.38.102.155 ip-65-38-102-155.hou.vericenter.com
65.38.102.157 ip-65-38-102-157.hou.vericenter.com
The host information is as follows:
OrgName: VeriCenter, Inc.
OrgID: VRCT
Address: 757 N Eldridge Parkway
City: Houston
StateProv: TX
PostalCode: 77079
Country: US
NetRange: 65.38.96.0 - 65.38.111.255
Athough the attack seems to be centered on the 65.38.102.0/24 block at the Houston datacenter of Vericenter, I think I'm going to completely block Vericenter as it doesn't appear to have any ISP facilities [ie. NO HUMANS] and see if anything else bounces off the bot blocker from their facilities.

Thursday, May 18, 2006

Another Web 2.0 Scraper Company

Don't roll your eyes and think that Bill's just making a fuss as this company claims they scrape:

Real-time Data Collection - Technologies for crawling, monitoring and scraping newly posted web content including the content from the “deep web."
See, I'm not making this shit up, honest to god admitted scrapers!

Not only that, they proudly display the number of sources they scrape updated constantly on their home page.

They appear to crawl without looking at robots.txt best I can tell, don't identify the source of the crawler other than it's the "Jakarta Commons-HttpClient", and their primary interest in my site seems to be attempting to crawl content referenced from my XML feed.

I'm not sure what information they could possibly think is on my site that could help "Institutional Investors leverage the latest technology and data to make better investment decisions" but they'll just have to be in the dark and use the Magic 8-ball from now own.

They have been seen using these IPs:
206.188.0.11 "Jakarta Commons-HttpClient/3.0"
206.188.0.22 "Jakarta Commons-HttpClient/3.0"
206.188.0.10 "Jakarta Commons-HttpClient/3.0"
206.188.0.23 "Jakarta Commons-HttpClient/3.0"
They are all part of the Geometric Group:
Geometric Group DP-206-188-0-0 (NET-206-188-0-0-2)
206.188.0.0 - 206.188.0.63
I'd just block the whole range and be done with it and hope we don't cause the market to crash.

UPDATE: They switched to Java in 2007!

01/22/2007 206.188.0.22 "Java/1.5.0_03"
01/22/2007 206.188.0.23 "Java/1.5.0_06"

Then mysteriously, stopped pinging my server on 03/15/2007 after a year of being fed garbage.

Think someone finally realized they were getting bounced?

Wednesday, May 17, 2006

Blue Frog Legs and Spam for Dinner

Normally I don't comment on the news and such but Blue Security shutting down their anti-spam Blue Frog operation is stunning. Obviously the spammers were feeling the pressure and it was working or they wouldn't have attacked, the plan was working. Then in this shocking turn of events the anti-spamming Generals leading the attack not only retreated, they resigned from the effort!

So what if the spammers brought down a few servers and services here and there?

Isn't that the whole idea to get those spamming bastards out in the open so people can track them and block their asses once and for all, put them out of business?

When you start a war you certainly don't pack up and go home the minute you get a bloody nose so I suspect that there's a lot more to this story as people just don't fold so easily from a purely technological war. With all the money at stake in spamming, I'm suspecting the threats got a bit more personal which resulted in the sudden shut down, but that's purely speculating on my part.

If this escalated into seriously dangerous territory, where it seemed to be heading, the big service providers, government and everyone else would've gotten involved and put a permanent stop to those responsible.

Unfortunately you Blue Pussies gave up before we ever got a chance for it to get really interesting.

Maybe someone with balls will step up and continue where you left off.

Tuesday, May 16, 2006

Port 80 Proxies Expose Themselves

Don't know whoever the fucking morons are writing those stupid fucking proxy servers, but a shitload of them were just blocked today when we noticed they were appending ":80" to the URL and it shows up in the HTTP_HOST parameter.

Normally HTTP_HOST just has something like "www.domain.com" but when the connection is initiated from a certain cluster of proxy servers it shows up as "www.domain.com:80" which is trivial to block.

Saved me a shitload of work tracking down the IPs to block.

Thank you VERY MUCH you stupid fucking assholes!

Search Engine Harvesting

Probably never would've noticed this but the referral string changing set off my referral spam trap. Turned out it wasn't a referral spammer at all but someone crawling my site trying to mine specific topic search engine results from both Google and Yahoo.

Very strange behavior too.

Cloaked as MSIE, downloaded all the images as well to remain hidden, yet crawled so fast it also set off a speed trap.

Stupid stupid stupid.

Dragonfly Crawler

No clue what this thing is and nobody seems to have any real information about it, but it's been seen visiting from 5 different IP addresses.

72.29.233.182 "dragonfly(ebingbong@playstarmusic.com)"
72.29.233.183 "dragonfly(ebingbong@playstarmusic.com)"
72.29.233.185 "dragonfly(ebingbong@playstarmusic.com)"
72.29.233.186 "dragonfly(ebingbong@playstarmusic.com)"
72.29.233.188 "dragonfly(ebingbong@playstarmusic.com)"
Claims to be associated with www.playstarmusic.com (72.29.233.167) and the IP address is so close it's possible.

If anyone has any additional information it would be helpful.

When Innovations Collide

Some web crawler has hit my site a few times called Heritrix which appears to be written mostly by the team at Archive.org, the same team that created ia_archiver for those of you that haven't had your coffee yet.

Yes, it supports robots.txt, but if you didn't know this damn thing existed you wouldn't bother blocking it now would you?

People writing crawlers wonder why webmasters get pissed tracking and opt-ing out all this nuisance crawling on their websites, but I digress, that's an old rant.

The real amusement is that Heritrix claims their technology is designed to "collect the digital artifacts of our culture and preserve them for the benefit of future researchers and generations" which is a bunch of pretty language to try to sidestep downloading a website without permission, especially when the webmaster probably isn't aware of your crawler, doesn't matter how you try to candy coat it.

Now comes the fun part,
let's see who was using it and why!


Today's attempted crawl was HUGE so it's safe to assume this thing has been on my site in the past and apparently the crawler was even banned on a previous bayarea.net IP address:

209.128.119.46 "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://innovationblog.com)
Today the crawler used a different bayarea.net IP, could be DHCP, could be on purpose to sidestep the previous ban. Who knows, but it only got a couple of pages before the doors were automatically slammed by the bot blocker:
209.128.119.17 - "Mozilla/5.0 (compatible; heritrix/1.6.0 +http://innovationblog.com)"
With my curiosity in overdrive, it was time to research innovationblog.com and see why they were crawling my site. Not a clue as there's nothing but a "This Web Site Coming Soon" site under construction page, but the WHOIS for the site was very revealing.
Registrant:
Michael Osofsky
1758 Shoreline Blvd. Suite B
Mountain View, California 94043
United States

Registered through: GoDaddy.com
Domain Name: INNOVATIONBLOG.COM
Created on: 09-Mar-05
Expires on: 09-Mar-07
Last Updated on: 06-Mar-06

Administrative Contact:
Osofsky, Michael mosofsky@accelovation.com
1758 Shoreline Blvd. Suite B
Mountain View, California 94043
United States
(650) 968-4741 Fax --

Technical Contact:
Osofsky, Michael mosofsky@accelovation.com
1758 Shoreline Blvd. Suite B
Mountain View, California 94043
United States
(650) 968-4741 Fax --

Domain servers in listed order:
WSC1.JOMAX.NET
WSC2.JOMAX.NET
This Michael seems to be involved with a company called accelovation.com and he seems to be big in the innovation circles having founded the MIT Innovation Club.

According to the Accelovation website:
Accelovation is the first and only Market Discovery System (MDS) that allows innovators to mine the online world for insights into unmet needs, trends, innovations and market activity.
Sound familiar?

We crawl you and use your information without permission to make a profit.

Where have we heard this before?

I'll bet they'll be surprised at my attitudes about this but they should try reading some webmaster forums and find out what they're doing probably isn't welcome without permission, some clue posted about what the benefits are to the webmaster to allow his site to be crawled, yada yada yada we've been down this path a few times before, it's getting old.

Sorry, but your innovation collided with my innovation called a bot blocker.

Your crawl is denied, and thanks for playing.