Tuesday, February 28, 2006

Comment Wars in Rantville

Normally I wouldn't just blog about a stream of comments but this article about Referer Spammer Revenge is almost a month old and the shit just hit the fan yesterday. I'm thinking the person I mentioned figured out it was me posting about them from my comment at Spam Huntress and went off the deep end.

We're talking some serious flaming here, cracked me up, but you be the judge.

Let this be a lesson to any children that may be reading this shit as it proves you shouldn't stick your finger too far up your nose as you can impair your thinking if you accidentally stab your brain and you'll end up a referer spammer too.

FIRST SIGHTING - New Bot Discovered

It's very rare I run across a bot you can't find any details about in Google but last night something calling itself "gnootBot" came from 66.39.177.8 which just has a default Apache web page.

No information about this beast except it started asking for pages in the middle of my website.

Very bizarre.

Monday, February 27, 2006

Peeling the Scraper Onion with Reverse DNS

Stopping scrapers appears to be like peeling an onion in that when you peel away one layer of bad bot activity you unearth yet another whole new layer that you couldn't see before. You start with user agent filtering then put up speed bumps and honeypots to stop others and even profile other behavior to stop them and they still keep coming. Now we're several layers into this scraper onion all sorts of new things are showing up that require more sophisticated methods to detect and block as they're hiding as browsers, running low key crawls, but still obvious to anyone that it's not a human if you look at the pattern of access.

To help thwart more of this nonsense the latest tool added to the bot blocking arsensal is reverse DNS lookups to see where the IP originates from and blocking or challenging bad sources from the start. There have been a few trends to become very obvious in that many scraper IPs that were auto-blocked didn't resolve to a domain name whatsoever or come from some suspicious hosting farms, the two most notable and persistent ones have been in Taiwan and the UK.

Now they'll need to find yet another way to get around the bot blocker as the newly installed steel door now has a chain, 2 deadbolts and a mean as piss rottweiler waiting on the other side just in case.

Stay tuned for more on the next episode of As the Onion Peels.

Saturday, February 25, 2006

Real Impression Tracking in Bot Buster

There was an unexpected benefit to all this bot busting and information recording in that now I can track on a daily basis all pages served to legitimate bots, blocked bots and actual people. This has allowed me to come up with some simple stats that shows a breakdown of where pages are served and I'm getting real close to almost matching exactly what you see in Google AdSense for page impressions.

This side effect alone could be an enormous benefit to people that need actual page impressions vs bot page impressions for selling advertising on their website and I've just made it a heck of a lot simpler to get that information since every single page passes thru my code in real time.

Coolio!

Cleverly Masked Bots Evolving

It would appear that the war over my content has been cranked up a level as bots masking as browsers and modifying their behavior to appear like people seems to be escalating. There are still some tell-tale signs that are easy to spot when you look at the server log but a couple of them that the bot blocker didn't catch are finding ways to game the system.

I didn't want to make the site more difficult for visitors but the only way to stop these guys would appear to be tossing in more random challenges like captchas and such after a pre-determined number of pages. To stop the typical captcha blow-thrus the challenges are very random and nobody could program a way to bypass them all as you don't know what they all are and I can add new ones daily if I wanted.

There's also something I noticed which isn't earth shattering but only humans seem to use my javascript menu which is a HUGE tell. Robots navigate the text links only but humans love those drop down lists and that's a clear sign that differentiates the two of them most of the time.

At the end of the day, it's just like trying to secure money in a bank, no matter how hard you try someone is going to rob you eventually but the best you can hope for is to make the number of times you get robbed as minimal as possible without pissing off all your customers in the process.

Thursday, February 23, 2006

Fuck Your Intellectual Property

Some asshats claiming to "defend your brand" sent their little AIPBOT to crawl my pages looking for anything of their clients on my site. Listen up fucknuts, you can use my SEARCH tool and look for something being on my site but you can kiss my ass when it comes to a 40K page crawl just to see if I'm violating someone's precious brand name.

This sense of entitlement of everyone to crawl the web is really starting to piss me off.

Take a hike assholes.

Link Me Or Else!

Here we go again with a persistent raging linkaholic badgering the shit out of me to link to his crappy little directory website so he can get a few AdSense clicks.

Hi YouBigWebStudYou,

I sent you a link request for bullshit-directory.com to see
if you would be interested in exchanging links.

I realize you are probably but wanted to let you know that I will be
removing your link next Wednesday if I don't hear back from you.

You can verify your link is by going to:
http://realfuckingannoying.com

with the following details
Title- Link To Me Please
URL-www.imbeggingyou.com
Description-I'm the biggest pain in the ass link-to-my-site whining spammer you've ever seen so link to me now before I beg and plead more.

If you do add my site please use the below information and let me know
the location you added it so I don't remove the link to yours
unknowingly.

I hope to hear from you before 2nd March 2006 but if not then I'm bound
to remove the link from my site.
Well fucking remove me already and stop sending me this shit!

Didn't the dead silence after your first spam give you a fucking clue?

If I could get my hands on you they'd have a new opening homicide scene for CSI next week so just keep it up, your luck is about to run out.

Oh yeah, it's a good thing you're in India so CAN-SPAM can't be used against you and you aren't registered with GoDaddy so they can't blackmail you to get your domain back. However, if there is a god one of those nasty little bugs in your water will give you atomic diahrea and you'll shit out a vital organ and die.

Wednesday, February 22, 2006

Search Engines Let Scrapers Bypass Spider Traps!

Just when you thought you've seen it all the actual search engines themselves can be used by scrapers to bypass spider traps. How this is accomplished is the scrapers find all of the indexed page names from your site in Google or Yahoo and then download pages the using known page names from your site thus side-stepping spider traps as they aren't actually spidering your site at all.

Therefore, just eliminating your pages from being CACHED in the search engines doesn't stop scrapers from still using the remaining data to their advantage.

Some days it just doesn't pay to get out of bed.

Tuesday, February 21, 2006

Another Plug-n-Scrape Component

Yet another toolkit letting armchair programmers attempt to grab my web pages.

Yawn.

This one ID's itself as:

IP*Works! V5 HTTP/S Component - by /n software - www.nsoftware.com
And their web site claims:
The HTTP component can be used to retrieve documents from the World Wide Web.
Might want to revise that to "used to be able to retrieve documents" as it went splat against my brick wall but I found it's calling card in my auto-blocked bot log.

Chitty Content

Must be the new wave of affiliate bots as I also got hit today by the Chitika ContentHit crawler or whatever the heck it is.

No way to verify this bot as reverse DNS for this IP address just claimed to be from Charter Communications.

71.10.233.52 Chitika ContentHit 1.0
This new bot of the day chit's just getting old.

Cell Phones and PDAs Can Piss Off

All these damn cell phones and PDAs all have unique user agent strings and for the last few months the handful that hit my website are all being told to piss off.

You people making cell phones and PDAs better wake up and smell the coffee as I'll be damned if I whitelist a bazillion user agents just to let your pissy products see 5 lines of my web site.

You all better come up with some better ideas for cell phone user agents as this unique name per phone shit isn't gonna fly.

CJ Quality Bot

Well here's a new one that I've never seen before from our friends at Commission Junction.

216.34.209.23 CJNetworkQuality; http://www.cj.com/networkquality
Unforunately they bounced off the walls, think I should let them in?

They might delist my site if I don't but based on the revenues I earned with them last month it's kind of a why bother IMO.

Fine, time to whitelist CJ, sigh.

Saturday, February 18, 2006

Robots.txt gives Bad Bots clues to access

In a rather lengthy debate with the owner of Majectic-12 on WebmasterWorld the issue of robots.txt came up over and over and I finally revealed that robots.txt is arcane and a real problem in the world of scrapers as it gives them clues to accessing your content.

Not only does robots.txt reveal which user agents may be blocked in the .htaccess file but it also reveals which agents are allowed into your server. Any roque bot not getting access to your site can simply examine the robots.txt file and use any allowed user agent name to get past the barracades.

My recommendation is to use a generic robots.txt file such as follows:

User-agent: *
Disallow: /stayout.html
Disallow: /keepaway.html
Disallow: /cgi-bin/
Disallow: /someotherdirectory/

Then allow and disallow robots privately in your .htaccess file only to keep that information away from prying eyes and give the lamest of the scrapers fewer clues how to penetrate your site's defences.


UPDATE: the Majectic-12 conversation at WMW is on hold pending review now so maybe I spilled to many beans on site security issues. It was a great debate, hope it comes back only slightly altered.

Cornucopia of Random User Agent Strings

When my bot buster first started operating I noticed a few gibberish user agent strings now and then as I'm sure the theory behind this is if a website is blocking known user agents then you can skirt past that technique with a string of gibberish.

The problem is that they've noticed nothing is getting thru and random user agent string usage against my site is escalating to the point it's hysterical to witness them thrashing.

Small sampling of thousands the other day:

66.148.68.37 2uigq2oecesvv2nwso rwiakBsBue Bobgw2nuB
202.125.44.200 efeSthqvkr11ticgo1iovjjrdwakbbd
66.148.68.34 emwx4cxnd pedafhfpac
66.148.68.34 ymdexin7xpebtulwnxew
202.125.44.199 pepgfu wjdjqrxckulhwiflmrdsmkc mjvldn
84.180.94.183 mairwthe Ifirpl8tiwotwyi lsu
84.180.94.183 r9Hreiynmkxmpjh ioHmmknpdmid
66.148.68.34 ewoqaohlcegoD emkdywx
66.148.68.34 obtDrqhxogxsewDfcDktb
209.190.21.100 bedmdFjkFhc4a noFjajakffieapvngdtpwxk
209.190.21.100 gdouk6Ss6nnykg66hvojc6txjsecuu
209.190.21.100 aphErvbtijj vulgctlslo
209.190.21.100 jgbhwntsdlprxcwogijI8orrw b8
209.190.21.101 DrbspcgyubxrpeikfiihxD mh
209.190.21.101 jvAhnviAjwwud8gymvewtcqhehgbAcytyqdxq
209.190.21.101 cvwkvl6kfujhqlujqblFl dffrepmrxdspmdFjq
209.190.21.101 obmJJkjtslbqreh6pwx6epruhptrpJbk
This is why I keep preaching that blocking by user agent only works for the legit crawlers that want to allow you to block them but the scrapers aren't playing by any of the old rules and I'm shocked these idiots just didn't use a browser string which has a better chance at least getting a handfull of pages of my web site until they get too greedy.

Sorry webmasters but the rules have changed in how this game is being played and you really need to block all non-browser agents and allow legit crawlers like Google, Yahoo and MSN by IP only. Any other method is just wasting your time as random user agents cannot be stopped by your old traditional techniques.

The blacklist is out, the whitelist is in.

Wednesday, February 15, 2006

Fuckers For Sale!

Found this little gem while playing around with AdSense testing relevant ads on dynamic search pages by using a GET vs a POST and just for shits and giggles added "?q=fuckers" to the end of the URL.

Well, much to my surprise one of the ads was extremely relevant and not only that, it appears you can compare places that sell fuckers on shopping.com and get a fucker of your choice for the best price!

I don't make this shit up, see it for yourself:



So much for family friendly AdSense ads!

Tuesday, February 14, 2006

Some days it seems so obvious

Today I was looking in the search engines for more signs of scrapers showing up in the SEs with my error messages and found a few more of these idiots. Then it crossed my mind that it would sure be nice to know which IP addresses that got blocked from scraping were assoicated with which web sites.

Well duh.

It's obvious I need to embed a bug in my text that ties the scraper to his web site so today starts the next wave of linking the abusers to their websites and opens up opportunities for automatic abuse reports that tie them all together.

Now this has me giddy, I can hardly wait!

Crawl Delayed RSS

Sometimes interesting solutions to problems just present themselves out of the blue as a discussion about Aaron Pratt's SEOBUZZBOX being supplemental results instead of the authoritative source gave me an idea.

What if you simply delayed updating your RSS feed until Google, Yahoo, etc. had already crawled your new content pages?

Theoretically this would improve your chances of being the first place that the new content was discovered and the news aggregators would then all become secondary sources for your information.

Would Google still make SEOBUZZBOX a supplemental result because of it's lower page rank or would fresh original content float to the top based on a first come first served basis and make the aggregator take the hit on duplicate content?

This is an experiment well worth trying, Aaron, you listening?

Anyone got any guesses what would happen?

Carnival of Scraper Sites

These circular scraper sites must be the worst sites I've encountered so far as all the scraped content is linked internally and no matter what you click on you just keep going round and round inside the scrapers web site until you click on an advertisement to escape.

These come in a couple of flavors such as the Moron Loop-de-Loop, the Loser Landing Strip and the Doorway Pages to Hell.

The Moron Loop-de-Loop sites contains snippets of your site and something that looks like it's a link to your web site but instead it links deeper into itself showing yet another page that is more thematic based on what you clicked. You can click and click and go round and round in this site because the only exits are AdSense exits. This is actually quite ingenious as you get more specific ads related to your area of interest if you keep clicking on more thematic links that seem to narrow your focus of interest but never provide anything interesting, except the ads.

The Loser Landing Strip is an interesting cloaking variation as this supposedly single page web site appears to have an infinite amount of cloaked pages behind it but only Google's IP addresses see that cloaked content, you only seel the silly landing page. No matter what snippet of page content you see in the search engine, no matter what nav links you click on in the web site, it's always that same landing page. All paths lead round and round to the same page so just click the ads already and give them what they want!

Door Way Pages to Hell is an interesting variant as it is part Moron Loop-de-Loop and Loser Landing Strip in that the first site you hit looks like a blog with links to actual web sites but the blog is actually the starting point on a trip to hell. Each link in the blog that says it's taking you to a useful web site actually has a link to one of their other scraped content sites instead, which of course link to yet other scraped content sites. It's a freaking AdSense pyramid scheme disquised like a domain park, it's absolutely frightening that someone would build such a large network just to keep a surfer going in circles until they catch a click.

Definitely a new low in scraping as the original content owner gets no value, not even a link, nothing but being used to get keywords and phrases to divert people to their sites.

Not sure it can get much worse than this but I've been wrong before!

Monday, February 13, 2006

Bullshit Gourmet

Sometimes when the little frozen food meals are on sale at the store I'll snap them up to get the 10 for $20 sale price and have a nice quick microwave lunch every now and then. Well, this weekend the fridge was bare and the usual brands weren't on sale at the market so what-the-hell let's see what these Budget Gourmet's are like since they're on sale.

Today is the moment of truth.

Opened one of these little boxes up and it has all the food, what little there is, on one side and all the sauce on the other.

Wait a fucking minute, half of this itty bitty box is reserved for a thin layer of sauce?

You must be fucking joking.

Completely unsatisfied ripped open yet another box.

Same bullshit, small pile of crap on one side, sauce on the other side.

BULLSHIT!

FUCKING BULLSHIT!

I'M STILL HUNGRY AND I'VE JUST BEEN FUCKED IN THE ASS BY A FUCKING FROZEN FOOD COMPANY AND NOW I'M FUCKING PISSED!

McDonald's anyone?

Sunday, February 12, 2006

Cloaking Scrapers Busted

Yesterday was a sad day for one Russian scraper that just had a large number of websites busted for cloaking in Google and they were just reported to Yahoo today as well.

How they were busted is simple, they scraped my site!

Since my scraper stopper was deployed it's been giving all the scrapers a very unique message which they were happily scraping and including with their additional scraped web content. Wait a few weeks for the search engines to continue to crawl and update and VOILA! this message starts to appear on web sites. Clicking on the website in question and the message didn't appear but clicking on the CACHE copy of the page in Google and there it is, cloaked in all it's glory.

Report these sites and POOF! they are gone.

It's like shooting fish in a barrel and more fun than allowed by law.

Come on you cloaking scrapers, take my pages, I dare you...