Wednesday, May 03, 2006

Google Exact Phrase Search is BUSTED ALL TO HELL!

Have you done any exact phrase searches in Google lately?

It's serving up all sorts of off topic shit, I mean, it's really busted.

Couple of examples of topics from this blog:

Do a search on "Lava Pockets" and you get crap liks this in listing #13:

Take Five - The Boston Globe(Correction: Because of a photographer's error, the names of two members of the Click Five were transposed in some edition's of Sunday's Arts ...
www.boston.com/ae/music/articles/2005/08/07/take_five/ - 41k - Supplemental Result - Cached - Similar pages
Neither the word LAVA or POCKET are even on that page

Even b
etter is "Florida Drinks Cat Piss" with 99% of the results complete shit!

Let's examine the #2 result:
eg Forums -> Pretentious bottle--ordinary wineOutline · [ Standard ] · Linear+. Pretentious bottle - - ordinary wine. Track this topic | Email this topic | Print this topic ...
forums.egullet.org/index.php?showtopic=72214 - 71k - Supplemental Result - Cached - Similar pages

That's right, nothing on that page about Florida whatsoever but the irony of wine coming up doing a search for cat piss certainly didn't pass me by.

Hey Google, you even know about SQA?

That's Software Quality Assurance for those that don't know, and I can recommend some very good people if you need them down at Google.

Some freaky shit

OK, today when I was taking a shower a very loud noise happened in the bathroom, nearly jumped out of my skin. Looked out and the door was still shut, so the cat hadn't been in the room and the wife was at work.

When I got out of the shower it turned out a tall perfume bottle, one that has sat in the same spot for ages, just randomly fell over and bounced on the counter top.

What the fuck?

I immediately thought perhaps it was an earthquake jolt but the dining room chandelier wasn't swinging, and it swings for the mildest of jolts, so I'm not sure what's going on at this point.

Just what I need, ghosts in the crapper, nothing the aftermath of a good mexican lunch won't scare away.

Friday, April 28, 2006

Jews for sale!

Who'll start the bidding at 10 shekels, do I hear 15?

That's right, some dumb fuck didn't filter out the plural of kike on Google Adwords and some dumber fuck is now trying to sell them on eBay.



Oy vey!

Wednesday, April 26, 2006

Microsoft Promotes Web Theft

Microsoft is promoting a whole page full of offline browsers just to help people steal shit.

They gleefully promote it too:

Quickly and efficiently download entire sites or sections of a site for viewing and browsing when you are not connected to the Internet.
Just what we need is a major corporation sticking some shit like this in front of people's faces that may not even be aware such tools existed in the first place. That's right, America's largest software company has become the Pied Piper of scraping websites.

Any words of caution about copyright, that your IP may get banned for abusive practices or anything else?

Hell no.

Hey Microsoft, you gonna pay for my web hosting bill when the assholes that download all these offline browsers blow my fucking server out of the water with ridiculously high bandwidth abuse?

I didn't think so.

Well, fuck you Microsoft, just fuck you.

Tuesday, April 25, 2006

India Link Whores Just Don't Understand

It's the spring link season once again and all the India link whores are out doing their ritualistic website link mating dance. Unfortunately, the shit they want me to link to would be like trying to mate a monkey with a goat, wrong species pal, nothing going to come from this attempt to defy nature.

Here's a slightly modified version of the most recent link blackmail to conceal the identity of the clueless asshole:

Hi webmaster,

We have already added your site to our resource page
http://www.OFFTOPICBULLSHIT.com/ long back.

However, we just visited your site and we still do not find the reciprocal link to our site. If we do not found our reciprocal link we will remove your site link from our resouce page. Please confirm the url where the link is posted as soon as possible to allow us keep your link.

Title :- Spamming link whores at nominal rates.

Description :- We don't fucking understand that people don't like off topic links, link to us or we'll add your link to Jihad.com and then you'll be screwed

URL :- http://www.REALANNOYINGFUCKERS.com

We look forward to receiving this information promptly.

Thank
Stoopidja Dumbfukja
webmaster@IMFUCKINGCLUELESS.com
Not only was the site completely off topic, the whole fucking thing was ORANGE, even a significant amount of the text!

My eyes started bleeding the minute I saw that shit and I couldn't tell if my link was on their site or not, but it didn't really matter as I told them to take their link and shove it up their ass sideways.

Monday, April 24, 2006

Crawl my site, go to jail, it's the law.

Who the hell put open season on my website?

Not talking about this blog, but the website that pays my bills and shit.

It's NOT a free articles site, plainly posted copyright notices, technology is in place to stop assholes dead in their tracks and they just keep coming.

Did someone post a sign on the internet that I've never seen that says:
"HEY! FREE SHIT TO SCRAPE HERE! COME AND FUCKING GET IT WHILE IT'S HOT!"

The other day I wrote some dry ass bullshit about data mining and then followed it up with some boring assed statistics I've been collecting but suddenly something changed and the amount of thieves hitting the site is suddenly going off the charts compared to a month ago.

Let's put this into perspective:

If my website were a grocery store and these assholes were looters or shoplifters the shelves would be bare, the clerks would be guarding the doors with shotguns, the police would have the place surrounded and the parking lot blocked off and a riot squad would be firing rubber bullets and cracking skulls of people trying to get away with th loot.

The local jails, needless to say, would be overflowing.

I realize there is no physical theft involved like the grocery store example above, and I also realize putting something out on a public network invites a certain amount of risk, and lunatics from the fringe, but when you install barriers and roadblocks to stop that activity and they just keep coming it's beyond and above what falls under 'normal access' and is well into some serious realms of abuse and harassment.

But this is just 'copyright infringement' and 'bandwidth theft', right?

Well, maybe it is when you identify yourself as robot and behave in an acceptable way that allows me, the webmaster, to stop you with reasonable efforts. However, when you mask to conceal the nature of your visit, change the identity of the crawler, and then attempt to crawl undetected and bypass mechanisms such as firewalls put in place to stop your activity, then TECHNICALLY this becomes hacking.

Isn't hacking to gain unwanted access a FUCKING CRIME?

I'm thinking someone needs to make a test case on this and just bypass copyright altogether and try filing a criminal complaint against them for hacking and see what happens.

I have a simple case I could use to test this theory already, as initially I put simple roadbloacks in place for humans to get past just waiting to see when the bandits would code around it, and sure enough they did a few months later. Then of course I made it harder and I'm waiting to see if they'll take a shot at bypassing this roadblock as well. Don't worry, there are deadbolts I can install to keep them out if needed, but my current cat and mouse game is more fun and useful to learn from as they expose the level of their technology.

OK, since I put a lock on the door that stopped the unwelcome technology that was hitting my site and someone deliberately programmed to bypass it, isn't that technically hacking?

I'm thinking it's just about time to drop some money on a lawyer to see where this idea stands within the definition of the U.S. laws on hacking as a crime as it would be nice just to scare the shit out of the local scrapers.

Not to mention it would sure be a cool to put a logo on my site that basically explains to these assholes:

"Crawl my site, go to jail, it's the law."

Who am I kidding?

Some asshole in some foreign country would just start a whole new business selling scraping services or pre-scraped websites if they aren't doing it already.

Then we're right back to fighting copyright infringement.

PicGrabAss Image Theft Technology

Some fuckers on a website far, far away have a product called PicGrabber that let's you quickly and easily crawl a website and steal every picture, movie, and any other goddamn thing that isn't bolted to the walls and nailed down to the fucking floor.

Their slogan is:

PICgrabber - The software that finds millions of free movies and free pictures for you!
And catch this "feature" list:
  • Scans the web with Keywords
  • Scans any given start URL
  • Downloads all found Images / Movies / MP3 automatically'
Shows up in your log file as:
Mozilla/PICgrabber - (http://www.movies-free.net/)
Luckily their little grand theft tool got nothing but bitch slapped with error messages.

Hey PicGrabber, try my slogan: FUCK YOU!

Sunday, April 23, 2006

Live Servers Update

The more IPs I block from this group, they just seem to move to a new block of IPs. All are being hosted by live-servers.net and they just using block after block and keep on coming.

The current ranges being blocked are as follows:

88.208.192.
88.208.193.
88.208.194.
88.208.195.
88.208.200.
88.208.202.
These ranges never send any traffic with the exception of multiple crawlers from the same block at the same time and all IPs appear to be pointing back to a server farm so I'm recommending you just block 'em.

More IP updates on this group as they become available.

Friday, April 21, 2006

Where are all the kick ass PHP programmers?

Now that I'm ready to actually build my bot blocker for commercial purposes it seems like it's next to near impossible to find any quality PHP programmers, at least none that aren't already booked thru 2010, to help build the damn thing.

For those thinking it's going to be some simple piddly-assed fire-and-forget script think again as we're talking a product of much larger scope, including a centralized server component. The scope is more that just the "bot blocking" that I rant about here as it will be a complete crawler management and protection system. It will allow webmasters to manage access to their site without having to deal with the technical aspects, nor keep up with the latest crawlers. Additionally, there is a security component that will harden web sites against a variety of exploits.

I was prepared to just pay someone to convert it to PHP, or even be an equity partner if they wanted some of the back end action, but it's starting to look like I may just have to roll up my sleeves and do it all myself. That's going to throw a nasty monkey wrench into the time frame as there is still pending R&D that needs to be done.

So much to be done, so little time, and a planned summer BETA release not looking so good at the moment.

Monday, April 17, 2006

Blocked Spiders DO NOT Go Away

There are a few bold, albeit naive, statements by other so-called "bot blockers" that scrapers just go away after you deny them a few pages which is complete and utter BULLSHIT!

Some of the scrapers being blocked on my server have been set to BANNED for months now, haven't gotten a single page of value, yet they just keep coming over and over, attempting to get pages they remember regardless of the outcome.

Most bot blockers I've reviewed just set speed traps or page limits and then throw a captcha in their face to make them go away for a brief period of time, maybe a few hours, maybe a day or two, but many of them will come back over and over and get another chunk of pages when they return. The stakes are high and the scrapers want your content badly so putting silly little bandages on your website for short term solutions do not cure the long term problems.

The only way to truly stop them is to profile their behavior over time as my bot blocker throws all first-time suspected bot IP's into QUARANTINE. Once an IP makes it into quarantine they are immediately suspended for 24 hours and then challenged immediately when they return to the web site after 24 hours. This stops repeat offenders from getting any pages whatsoever when they return and also protects against permanently blocking a DHCP address by accident that is used to scrape only once. After a couple of repeated scrape attempts without breaking thru the challenge, which a human can easily do, the site is escalated from quarantine to BANNED which no longer presents challenges and just gives error messages on repeat visits.

Not rocket science but it has a lot more finesse than some of the more simplistic methods others employ and better hardens the site against repeated attempts at scraping.

Sunday, April 16, 2006

Gosh Goes Wild

I blocked one IP address and they came back with a new IP address 72.51.37.210 doing their sneaky masked crawl bullshit.

Go away little crawler, stay away little crawler, or this WILL get ugly when I start cloaking 'YO MAMMA' snaps into your fucking search engine for major keywords like: "Pool Cleaning Services - Yo Mamma so stupid she saw shit in the pool and thought the anal porn queen was trying to teach her baby to swim!"

BTW, fix your dumb fucking crawler so it can properly parse a goddamn web page as hitting my server for "GET / <font" is about as bright as asking for 30 pages in a minute while claiming to be "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"

Stupid fucking shit really gets on my nerves.

Saturday, April 15, 2006

Alexa Hiding in the Shadows

What the hell's going on with Alexa trying to crawl my site with NO user agent string!

Had an attempted anonymous crawl from 209.237.238.224 which is:

whois 209.237.238.224
Alexa Internet ALEXA-INTERNET (NET-209-237-237-0-1)
209.237.237.0 - 209.237.238.255
They still read the robots.txt file, but no clue who it was:
209.237.238.224 - - "GET /robots.txt HTTP/1.0" 200 111 "-" ""
What the hell's going on Alexa?

Someone break the crawler or you just trying to sneak under the radar after everyone blocked your ass?

You claim to crawl as "ia_archiver" but I don't see any ID here whatsoever!

Please explain, inquiring minds want to know!

Friday, April 14, 2006

Cloaked Bots Snared, Charted and Graphed

Since improving the profiling algorithm of my bot blocker so that cloaked bots could be trapped more accurately the statistical data has been piling up and now that data has been compiled and graphed for your viewing pleasure.

Hopefully this will explain to those that think I'm trying to stop beneficial spiders that those spiders aren't what I've been going on about whatsoever.

For those new to the blog we'll quickly define a couple of terms.

  • Cloaked Bot - a web crawler that uses the user agent string of a browser like Internet Explorer, Firefox or Opera
  • Blocked Agents - those crawlers that plainly identify themselves, such as Googlebot, but are unwanted and typically blocked via robots.txt, .htaccess, or other methods.
The chart below shows the amount of requested pages by cloaked bots pretending to be browsers (in purple) compared to blocked agents that have a plainly identifiable user agent string. One thing immediately obvious is that page requests by the cloaked bots far exceed the number of page requests made by all the other blocked agents combined.



The trend analysis of the cloaked bots reveal that, other than a couple of recent spikes, their page requests are on a slow but steady decline over the period charted. This could be a direct result of blocking them and sending them pages of error messages since they aren't getting what they need. Note that in the same time period the normal blocked agents keep attempting to request pages at a steady pace, regardless of the fact they're getting nothing but error pages, with no obvious trend evident at this time.

While some of this may look underwhelming at first, realize that many crawls would've been for hundreds or thousands of pages that are now being stopped before they ever start. Much of the attempted crawling appears to be repeat offenders that already have a complete listing of all the pages of the web site from prior visits. Some cloaked bots do manage to get the site map before being stopped and try to crawl all the links they know about, albeit fruitlessly.

At a minimum this information proved my theory, for my website anyway, that it's not the bots you know and can see that are the real problem, it's the bots you can't see.

With all this in mind, do you think you really know how many actual visitors and pages views your site gets?

Thursday, April 13, 2006

Spy vs. Blog Cache

OK, this is getting ridiculous, spy on someone that's stupid and falls for your tricks.

I reported on blog spying the other day and the direct spying stopped showing up in my activity log but now I'm seeing information spies watching my site via Google cache.

Let me be the first to tell you sneaky snoopers that you've met your match this time, the gloves are off so keep fucking with me so I can add every trick in your book to my bot blocker product. Just keep showing off trying to spy on me as every time you play a new trump card you can never use that trump card against me, or anyone that will install my bot blocker in the future.

Talk about playing with fire, sheesh.

Did you think that you could stop hitting my blog and hit my cache and not be noticed?

Look at this listing from Google cache that was scrubbed of identifiable data to mask these people:

http://72.14.203.104/search?q=cache:incredibill.blogspot.com/ %22[THE IP OF THEIR WEB SITE]%22

Very clever, scanning Google for the domain name of my blog, and narrowing the search for LINKS to your site from my blog, by your IP address instead of your domain name hoping it would go unnoticed.

Maybe other people are stupid and fall for this shit but I noticed that IP address, did an NSLOOKUP, and VOILA!, it was the same block of IPs and spybot company name that I've blocked on my main website and discussed in this blog before and now this continuing surveillance is bordering on, bordering hell, it's harassment.

Just go away, we're done already.

UPDATE: After lunch I set the blog to NOARCHIVE so as soon as Google and all the rest update and kill the cache we'll see what they try next.

Tuesday, April 11, 2006

Beware Blog Spies

This isn't the first time I've noticed this but it's definitely the quickest that my blog has attracted spying operations that monitor information about specific topics.

The first time it was overtly obvious what was happening when I posted about corporate crawlers that spy on websites and the very next day my blog was under daily surveillance from these corporate clowns.

The second most obvious time started yesterday when I posted the blurb "Dubious from Dubai" which immediately started getting hits from what appeared to be automated tools performing automated queries for Dubai on search.blogger.com and one of them was from 217.164.235.140 which according to whois is Dubai: "P.O. Box 1150, Dubai, UAE". The other hits were from several places in New York which make me wonder who's camping on any information about these Arab countries and for what purpose?

I'm not a cloak and dagger kind of guy but sometimes you just see things that raise your eyebrows.

Monday, April 10, 2006

Yugoscrapia My Ass

The "Get a Fucking Clue" award for today goes to whoever the hell is sitting behind 212.102.136.25 hiding as Mozilla/5.0 (compatible; MSIE 5.0) from Yugoslavia that has spent over 8 hours now downloading 1,000+ error messages from my site and is currently still going strong.

Well pal, whoever you are, my automation is better than your automation because mine has a built-in BULLSHIT DETECTOR that stopped yours from doing it's job.

Dubious from Dubai

Just got hit for a couple hundred pages each from this triumvirate of IPs from Dubai:

195.229.241.181
195.229.241.180
195.229.241.187
Reverse lookup said NXDOMAIN and Whois didn't get much either so I'm just gonna block 195.229.241.* and be done with it.

Data Mining Kills the User Agent String

FACT: Your website is raw material for the Internet data mining industrial machine.

Much like the California Gold Rush there's a lot of free money on the table and everyone is scrambling to grab their share on the Internet. This time instead of sifting thru rocks looking for gold nuggets they're using bots instead of shovels and crawling websites instead of standing in a creek. The purpose of all this web crawling is to sift out information from a variety of websites looking for gold nuggets of content that will help get free money in the form of internet advertising. Everyone wants a share of the free money and your website could contain just the right couple of nuggets of gold that the Internet claim jumpers need to succeed.

Simplistic filtering of the User Agent string to block these claim jumping bots has definitely become obsolete because most undesirable bots already don't identify themselves as anything unique and try to hide their presence as the prize is too big to let a webmaster stop them. Don't think this behavior is limited to simple content thieves trying to capitalize on your hard work with AdSense as there are several corporations that I've caught in my snare and probably a bunch more lurking behind IPs that don't expose them with a simple reverse DNS lookup.

What kind of data mining happens on your site?

  • Search Engines
  • Data Aggregators
  • Web Copiers/Offline Readers
  • Copyright Compliance
  • Branding Compliance
  • Corporate Security Monitoring
  • Media Monitoring (mp3, mpeg, etc.)
  • Link Checkers
  • Privacy Checkers
  • Content Scrapers (pure theft)
  • so on and so forth

Other than search engines which provide a valuable service bringing you traffic, many of these so-called services are just one-way bandwidth hogs that not only earn money off your back but you get to pay for the privelege!

Not all of the aforementioned services try to hide who they are and the more legit ones still check robots.txt and present a user agent string so you can opt-out (don't get me started) of their service. However, as the free money flows on the internet so does the desire not to get caught and stopped such as the spy services and scrapers.

More and more crawlers daily are pretending to be users than admit what they truly are to permit the webmaster to stop them, and that trend seems to be growing rapidly as the stakes are higher.

Use robots.txt and .htaccess while you can but you're only stopping the good guys as everything else has gone underground and there doesn't appear to be any reversal of that trend anytime soon.

Saturday, April 08, 2006

ShitMoreBlueberries

Here we go again with another anonymous proxy site called EatMoreBlueberries.

On the site it goes blah blah blah about circumventing censorship then has the balls to strip my ads and shit, which censors MY ability to earn from my pages, might even be considered a tortious intereference with business, then slaps their ads on top of content that THEY DON'T FUCKING OWN!.

They appear to have a block of IP's so it looks like you can just block 69.49.99.* for starters and they're gone.

Well guess what fuckwads?

I've blocked your ass so you've gotten the ultimate censorship which is a big FUCK YOU from ME to YOU!

SeekOn Elsewhere

It appears that the people at SeekOn think that if "someone" submits your site to their service, just anyone according to them, they have the right to crawl your site. Better yet, they don't even bother to look for ROBOTS.TXT, oh fuck no, they expect you to use META TAGS on all your pages to control robots instead of that one handy little robots.txt file.

Just block "SeekOn Spider" in your .htaccess file and they can kiss your ass at that point.

JetBrains Damaged

The product claims to be an RSS Reader and Newsgroup Aggregator but I caught something claiming to be JetBrains Omea Reader 2.1.4 looking at pages, well, attempting to look at pages, that were neither RSS or Newsgroup.

Never seen this before and I'm not sure if it's a fake or this thing can be used for other purposes but you might want to keep an eye on this thing just in case.

Friday, April 07, 2006

Scraping Gold Medal

Most scrapers only try to get a few hundred pages at a shot and some of the more aggressive ones attempt to get a couple of thousand pages but this scraper gets the gold by attempting to extract 12,935 pages in one shot.

What he got was 12, 935 pages of error messages.

Way to go dumbass!

Wednesday, April 05, 2006

Brazilian Bot Busted, Oh GoshME!

Well some AdSense monetized bullshit search engine or something called GoshME masking as a BROWSER of all things has been very slowly extracting 56 pages for a total of 58,054 seconds or 16 hours from the address of 64.34.172.77.

BTW, what kind of happy horseshit is a Brazilian website that's all in English?

What's best is this fucking thing is POSTing to my web site, not using GETs like a normal crawler, so God only knows what these assholes are stealing and he's not telling.

FWIW my website shows up as #1 on their site for my category but I don't give a flying fuck.

MASKING AS A BROWSER?

FUCK YOU!!!

Le Tub and Le Return From Florida

Back from Florida so no more vacation, birthdays or weddings in my near future, done with all that shit and not a moment too soon.

While in Hollywood, that's Hollywood FLORIDA you morons, I decided to try what GQ magazine claimed was the #1 choice of "The 20 Hamburgers You Must Eat Before You Die" and gave Le Tub a try.

For starters, it should be called Le Dump as it makes every dive I've ever been in look like a sparkling palace by comparison. Everything including the tables and benches, if you can charitably call them that, look like they were all made of the same rotting aged dock lumber. The place is cluttered with bath tubs, sinks and toilets used as planters and stuff. Some would call this shit "decor" but overall it looks like you're dining in a fucking junkyard. How this place isn't condemned is fucking amazing.

Don't worry, we read several reviews of this place before going over there and knew what we were getting ourselves into but reality cramps are still a bitch.

We walk in and head toward what looks like a half empty bar and the cook shouts from the cloud of smoke in his bird cage sized kitchen "Those seats are taken and it will be an hour to hour and a half for a burger!" so we head outside and plop down on a table where my bench looks like it's about to collapse but surprisingly didn't wobble at all.

That's correct, you're reading this properly, 1 to 1 1/2 hours for a burger, which we also expected going in the door as it's the Disneyland of hamburgers and it will take that long to get your turn at the hamburger rollercoaster.

Why does it take so long?

All of the burgers are 8 ounces (1/2 pound) which makes them huge fucking piles of beef to start with and takes a long time to cook even to medium. The grill is tiny too, about a 3x3 grill so it doesn't hold very many of these gut bombs in the first place.

The menus are printed on crappy paper and of course mine had mustard stains all over it and might've been slightly soggy but there was only one item on the menu of any interest and that was the hamburger. It's only served one way, and that's the 8 ounce charbroiled way for $10 with the optional cheese costing $0.50 more. Comes with lettuce and tomato but you have to pay an extra $3.50 for small fries or $5 for a basket to share and the fries are well done, nice and crispy the way we like 'em.

Whoever wrote the review claiming Le Tub had a "good if not stellar selection of beer" wouldn't know a good beer if the bottle was smashed upside their head. Good beer is not named Bud, Bud Lite, Coors, Coors Lite, Heineken, Amstel Lite and Becks for fuck's sake!

Our skinny-assed waitress, and I mean skinny as my dick is bigger than her thigh, was pretty attentive and after taking our order it only took 4 Beck's before the burger arrived. While we were waiting we noticed so was everyone else with every table full of hungry zombies just staring at each other wondering when or if their food would ever arrive. Once the orders started flying out of the kitchen it was pandemonia at all tables with squeeze bottles belching out ketchup and mustard in a chorus culminating with the starved gobbling up a generous chunk of scorched cow on a bun.

The burgers had a nice charbroiled flavor and since I've now eaten the #1 hamburger that I must eat before I die, according to GQ, I can now peacefully pass knowing that Le Tub probably contributed to the clogged artery that will most likely kill me.

If you're ever in the neighborhood of Hollywood, FL you simply must try Le Tub as you can't even imagine this place until you see it first hand as there are simply no words that can truly describe this place.

Bon Appetit.




Friday, March 31, 2006

Slow Blogging

Sorry the blogging is so slow but I've been stuck in Florida all week waiting for this big wedding that's going to happen tomorrow. Been trying to avoid all the cat piss on tap they call beer down here and starting to notice that although they mostly serve Bud, Miller and Coors there appears to he Amber Bock in most places for reasons I can't quite explain. Even ran into some Killian's Red the other night among the kegs of cat piss, I was stunned.

Anyway, I'll be back to California sometime next week and raising hell as usual.

In the mean time John Andrews went of the deep end, enjoy.

Wednesday, March 29, 2006

Bot Blocker PLUS Security Features

Been playing around with adding security features to the bot blocker as it became obvious there were additional things I could do such as punting pages with SQL injection, cross site scripting (XSS), certain overflow attempts and a whole bunch of other things when I was evaluating the page requests.

Now if I can add some anti-password hacking, anti-password sharing and anti-blog spamming then I'll have a complete set of tools to combat wide spread abuse without each website having to code it's own protection.

What a concept, stopping bandwidth abuse, content theft and some hack attempts all rolled into one!

Tuesday, March 28, 2006

Almost Ten Percent of All Pages Blocked

The total number of pages being blocked from non-humans masking as browsers on my site is about 10% of all pages displaying daily and this rate has been holding strong since I started blocking them.

That's right, for every 48K pages displayed daily between 4K-5K are going to scrapers, crap search engines that send no visitors, and other useless wastes of bandwidth. That's a heck of a lot of pages that are being scraped and the purposes for all of this are still only slowly unfolding as each week something new shows up somewhere on the net.

Simply amazing.

It will be interesting to get more sites profiled moving forward and see if this is a common trend or not as 10% of all site traffic being wasted, not to mention server resources, could take a decent load off many servers starting to feel the pinch.

Will get some more detailed stats together next week hopefully and put them up for everyone to take a gander at as it's blowing my mind for sure.

Friday, March 24, 2006

WebCorp Crawls Why?

Saw this entry on my blocked log a couple of days in a row now:

03/24/2006 193.60.130.67 "WebCorp/1.0"
So I looked it up and sure enough it's WebCorp
webcorp.uce.ac.uk.
Went to their website and they have some mental mind fuck linquistic mashup running they call SEARCH and it's the slowest thing I've ever seen since I used a Commodore-64.

However, it claims to cull results from Google, Altavista, Metacrawler and AllTheWeb so why in the hell is this thing crawling attempting to crawl my website if my content isn't even being taken into consideration for their results?

Sorry, you PhDs in linquistics are just too smart for me so maybe you have some higher purpose for attempting to crawl that just escapes us bot blocking neanderthals.

Here's a phrase you cunning linquists might know - "fuck off"

Thursday, March 23, 2006

How Clever Yet so STOOOOOPID!

These assholes that crawl my site must think I'm looking for browser-like behavior such as image loading or something, or they're just using an API to drive Firefox in an effort to completely mask their tracks running complete browser operation including loading ads.

Now let's follow the fun antics of this crawler:

63.197.247.13 - - "GET /robots.txt HTTP/1.1" 200 111 "-" "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8.0.1) Gecko/20060111 Firefox/1.5.0.1"

Reading robots.txt with a browser?

I should've stopped you here but I didn't just because I know you won't get far, 10-20 pages tops, and since I'm a bit on the sadistic side, I've been letting them continue on past robots.txt lately just to see how good the rest of my traps are working.

STRIKE #1

Now you access a page the robots.txt told you not to use?

More importantly, you can only see this page name if you're looking in robots.txt or finding it hidden in my HTML, normal visitors don't see this page.

63.197.247.13 - - "GET /dont_click_this_page.html HTTP/1.1" 200 12895 "-" Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8.0.1) Gecko/20060111 Firefox/1.5.0.1"

STRIKE #2

Then your stupid ass program continues to attempt to load 60 more pages while not noticing you're getting a captcha after stepping into the spider trap and after a few pages of that, getting error messages telling you that YOU'VE BEEN BUSTED.

STRIKE #3 - YOU'RE OUTTA THERE!

Which just goes to show that even the ones that do things that would bypass my 3 page robot stopping technique still get stopped in record time. Not only that, they could've been stopped before the first page as reading robots.txt is a cardinal sin for a browser.

However, stopping them too fast at this point and I wouldn't have had the fun of getting more profile information from this type of crawler.

The harder they try, the harder they fall, as someone obviously went to some extremes to pull this off and still ultimately failed in the end.

Cyveillance Keeping An Eye on My Blog

You tell the world Cyveillance is snooping on your website and sure enough here they come snooping on my blog too.

Referring Link http://blogsearch.google.com/blogsearch?hl=en&q=cyveillance&amp;amp;ie=UTF-8&scoring=d
IP Address 65.213.208.155
Country United States
Region Virginia
City Arlington
ISP Cyveillance
We know you're watching us watching you spy on us.

FYI, they also seem to have this range of IPs:
CYVEILLANCE UU-65-213-208-128-D4 (NET-65-213-208-128-1)
65.213.208.128 - 65.213.208.159

Wednesday, March 22, 2006

Stop that bot, 3 pages or LESS

Some days you just wake up and BINGO! you have the best damn idea you ever had in a long time. That's right, something so simple and as plain as the nose on my face just slapped me upside the head today and I'm positive I can shut down bots masking as humans in 3 pages or less into a crawl.

However, it was real a bonus day for me as I came up with not one but TWO new techniques to add to my arsenal of bot blocking weapons. The only downside is both of these tricks require changes in the web pages in order to make it work, but it's well worth the trouble if it can stop bots dead in less than 3 pages.

We'll be doing some testing for a week or so to see if it's really as effective as I think it is and verify it's not snaring humans and let you all know how it works but I'm REALLY excited as this KICKS ASS so far!

Anyone know a good patent attorney that's reasonably priced? ;)

More Fun With Stupid Bots

The website I'm protecting from all these idiots uses javascript navigation for some pull-down menus and it appears that a couple of the scrapers appear to be attempting to scrape something out of the javascript looking for embedded URLs in the script itself.

Unfortunately for them, but lucky for me, their code is mildly brain damaged and doesn't bother parsing the extracted information to see if it's a valid URL and these morons are trying to access a page name from the server like "/getURLfromMenu".

Now that I've noticed this little tidbit I went back and checked my bot blockers error logs and it's already stopped this from about 40 different IPs in the last couple of weeks. A couple of the others attempting this had invalid user agents, meaning they didn't scrape it lately as they haven't been getting into the site for many months, so this has been going on a LONG time before I caught them.

At least I have another bit of criteria to add to my instant block list.

Stealing a line from Seinfeld with apologies to the Soup Nazi:
NO SCRAPE FOR YOU!

Monday, March 20, 2006

Clone Wars

Who are all these wannabe assholes?

I've been the original IncrediBILL online since I first got my hands on a 300 bps modem and if you don't believe ME, then ask my wife FRANtastic!

Most of these fuckers were probably hopping from ball to ball just to keep from landing in their Dad's tissue or being the glue sticking the pages of the Playboy together, or maybe shitting yellow in a diaper at the time.

Look at this shit, IncrediBill's to the left, IncrediBill's to the right, the fuckers are crawling out of the goddamn woodwork!

For the love of god make up your own fucking names.

Sunday, March 19, 2006

Educating the Public About Scraping

After talking to a lot of people lately, many webmasters and Silicon Valley internet savvy types, it has become obvious that they simply are oblivious to the entire problem with rogue bots and scrapers. Most people I've been discussing this with are aware of crawlers and they're aware of things like robots.txt, but completely in the dark about what goes on bypassing so-called standards. Ultimately, they leave the conversation with a new level of fear about the security of their online content and run to the nearest console and start searching for unauthorized usage of their content which, as we all know, they typically find without too much trouble.

The real eye-opener for most that aren't building sites that thrive off Webmaster Welfare™ (aka AdSense) seems to be the entire AdSense economy that fuels the bottom-feeding scraper sub-culture that Google has unwittingly created. Once they understand the motivation not only is it clear why scrapers scrape to anyone, but many wonder why they didn't think of it first! Then it's obvious that the low hanging fruit has universal appeal and everything on the net is fair game for the unethical types that pluck that fruit at any cost.

So the question remains, after this small sampling of industry savvy folks, is how wide is the blissful ignorance to this pandemic?

Wonder how many people learn something about this for the first time hitting this web site and just think I'm a paranoid loon with a tinfoil hat dancing with a flute celebrating the summer solstice?

I'm suspecting the depth of the problem is not known by most, even by webmasters fighting one-off copyright infringement, those that even have a hint think it's being overblown and from what I'm seeing in the last week in my banned log files, it will get a lot worse before it gets better.

Friday, March 17, 2006

Spyveillance, Block 'em if you got 'em

OK, this must be a clue that my bot blocker has graduated to the head of the class as I've snared 2 coporations bypassing security measures within 24 hours pretending to be browsers.

Remember what I said about bot blocking being an onion that you keep peeling layer by layer?

The next one in our list of sneaky snoopers is Cyveillance, which apparently has been around for a while but went silently unnoticed until I cranked up the level of bot profiling on my site just a bit to see if I was missing anyone and BINGO! got 2 big fish in a day looking at the next layer of the onion.

According to what I've been reading at linuXgod's site, these boys spy for the RIAA, government and god knows who else or for what purposes. He's been trying to get them to stop crawling his site via a small back and forth of emails and they don't seem to be interested in complying.

My favorite quote is where they justify ignoring internet standards like robots.txt and mask the user agent string as a browser ""Mozilla/4.0 (compatible; MSIE 6.1; Windows XP)".

Because many sites use redirection pages to route robots to special "indexing" pages, we identify our web crawler as an IE browser to ensure it receives the same content as the majority of web surfers on the internet and to allow our programmers to concentrate on a single interpretation of thehtml standard.
Well hell, doesn't that logic just make it fucking OK to ignore whether I want your robot on my server in the first place?

So you're justified in bypassing my security to stop browsers just to concentrate on a single html standard?

Well guess what, NO, YOU'RE NOT JUSTIFIED!

Here you go people, the range of IPs so block them as we're not being given any other means to detect this crawler:
whois 63.148.99.239

Cyveillance QWEST-63-148-99-224 (NET-63-148-99-224-1)
63.148.99.224 - 63.148.99.255

and...

CYVEILLANCE UU-65-213-208-128-D4 (NET-65-213-208-128-1)
65.213.208.128 - 65.213.208.159
Wish I had the bot blocker commercialized now to go mainstream and nail this nonsense.

Corporate Crawler Masking as MSIE

Well, color me stunned shocked and appalled as I ran into an actual real live corporation with a legitimate product that is deploying a crawler that sets the user agent as MSIE ""Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; ....)".

Yeah, that's right, forget robots.txt, forget letting you block them by normal user agent filtering means, they're getting into your website whether you like it or not because they have MANIFEST DESTINY!

They are ENTITLED TO YOUR CONTENT!

Not.

These lovely sneaky snoopers that boldly bypass your firewalling efforts are Lightspeed Technologies and they appear to be operating from this IP range 66.17.15.128 - 66.17.15.191.

Just block them now as this is about the lowest I've seen a corporate crawler get and they should be blocked on principle alone by not honoring internet standards.

Thursday, March 16, 2006

Terms of Service vs. Fair Use

Here's my next thought about how to combat ill behaved spiders that include snippets from your website and claim fair use. Include something in your TERMS OF SERVICE or LEGAL page on your website that prohibits unauthorized robots.

Therefore, even if they are within their rights of fair use they've violated your terms of service and you possibly have an actionable item on your hands.

Thinking about running this one past a lawyer as we need some boilerplate text like the GNU license that can be distributed and used everywhere as leverage against scrapers.

Film @ 11

Tuesday, March 14, 2006

Related Movie Bullshit

If my last rant just moments ago about sucky movies didn't get the point across, I was just reading some entertainment news and J Lo may star in a movie adaptation of Dallas and Ice Cube may star in the big screen version of Welcome Back Kotter.

Dallas?

Welcome Back Kotter?

You must be shitting me!

IT'S OBVIOUS WHY PEOPLE DON'T GO TO THE FUCKING MOVIES YOU MORONS!

Looks like I'll need to take up golf or some shit since I obviously won't be at the movies anymore.

Hollywood Blames Piracy Instead of SUCK MOVIES

This rant is rated TV-MA for fucking language.

Pay attention Hollywood, just sit down and listen the fuck up, it's not piracy that's stopping movie goers from watching films and buying DVD's, it's the long stream on non-stop shit you've been cranking out this year stopping people from wasting their money. In case you missed it the first time, listen up you shithead demographic chanting morons, it's your BULLSHIT MOVIES keeping my money in my pocket, not piracy, not video rentals, not On Demand, not cable TV. Don't put the blame on anything but your garbage product as it's ALL YOU, nothing else, causing your decline.

My wife and I used to go see movies 1 or 2 times a week and in the last year we've been having a real hard time finding anything worth wasting money on 1 time a month so obviously we're stealing shit instead of paying to watch shit according to your theory just because you couldn't make a movie worth 2 thumbs up your ass most of the year.

Other than Walk the Line, Matchpoint, Syriana, and Good Night, And Good Luck plus a few others I can't remember at the moment the choices have been real fucking slim this year.

Nothing fun stood out like that anticipated sequel to the First Wives Club those bastards shelved, give me Goldie, Bette and Diane before I get pissed or one of them has a stroke! Nor did they show anything outstanding like American Beauty or being a Nicholson fan we could use more films like Something's Gotta Give, About Schmidt and As Good As It Gets and I could care less if Jack stars in them either, just good quality movies you can watch!

Not to mention the massive vacuum of any real stand-out superhero flicks in a while but the new Superman is on the way, supposedly. Nor have we seen any good SciFi / space epics in a loooooong while but TV's Friday night SciFi line-up is kicking their asses anyway so if you decide to make a new space movie it better kick ass like the original Alien or be fun like Starship Troopers, something special because people running around a rusty bucket ship shooting each other over some fucking conspiracy in a low budget space movie is BOOOORING.

I like to go out, I like a night at the movies, I'd go twice a week easily, but I refuse to ruin my night watching any old shit you think I'll pay for because you're WRONG WRONG FUCKING WRONG so make a good movie or just shut the fuck up and go broke already.

BTW, while I have your attention, spread out the goddamn movie times you assholes. People are still getting off work and want to have dinner when all the movies are starting at 7pm and most people move on and do other things before the next showing at 10pm. It's stupid, it's always been stupid, it will continue to be stupid and you can lose more business until you stagger showtimes a bit more so working slobs can see movies at 8pm and 9pm which is more reasonable.

So BLAME PIRACY when you continue to MAKE SHITTY MOVIES and continue with showtimes on CRAPPY SCHEDULES so.... FUCK YOU, FUCK YOUR PROFITS and STOP FUCKING WHINING YOU RICH HOLLYWOOD GOLD-PLATED GATED MANSION LIVING MOTHER FUCKERS just FUCKING FUCK YOU!

Just make me some good movies and we won't have a talk again, ok?

Monday, March 13, 2006

This Means War

While I was out to lunch this afternoon some nitwits that I've banned over and over trotted out a new IP address and tried to scrape 1,000 pages when nobody was watching.

Sorry pals, someone WAS watching, it was my little silent sentinel buddy that I wrote myself that blocked your ass after about 20 pages and sent you a nice whopping 900+ pages of error messages.

I think I've had enough of your shit though and perhaps it's time we see what ABUSE@SCRAPERHOSTING.COM has to say about your repeated attacks on my server.

Hope you get shut down or at a minimum find yourself in bed with Lorena Bobbitt and wake up with a Frankendick.

Sunday, March 12, 2006

Knuckle Scraping Neanderthal

When a scraper reads your robots.txt file don't you think they would avoid the disallowed pages and directories?

Then would you believe the scraper reads your robots.txt file a SECOND time after just downloading a few pages and immediately opens the page that it's told to leave alone and WHAMMO! gets stopped.

How FUCKING STUPID can you be to write such brain damaged code?

Chitika ContentHit IPs

Chitika took another shot at my server with the user agemt "Chitika ContentHit 1.0" but this time tried a whole bunch of IPs on a single web page, one that Chitika doesn't even appear on which was most amusing.

  • 67.15.219.3
  • 67.15.219.11
  • 67.15.219.18
  • 67.15.219.14
  • 67.15.219.15
  • 67.15.219.10
  • 67.15.219.9
  • 67.15.219.16
  • 67.15.219.17
So there you have it, let 'em in block their ass at your leisure.

Saturday, March 11, 2006

Looksmart using Nutch?

When I was looking thru the blocked bot log today I ran across a single nutch hit that caught my attention which upon closer inspection appears to Looksmart playing with Nutch, not even identifying themselves as Looksmart.

VERY ODD.

This was the entry:

03/11/2006 07:07:16 BAD_AGENT 64.241.242.18 "NutchCVS/0.05 (Nutch; http://www.nutch.org/docs/en/bot.html; nutch-agent@lists.sourceforge.net)" "index.html"
So I looked it up and there they were:
nslookup 64.241.242.18
Server: 64.34.160.76
Address: 64.34.160.76#53 Non-authoritative answer:
18.242.241.64.in-addr.arpa name = sv-fw.looksmart.com.

Open Source replacing jobs in failing companies?

Think someone got fired in the seach dept. down there, if you can call it that.

Stupid Bots Can't Parse

The other day I posted about the idiots trying to locate "/#" and "/#top" but I didn't even notice a couple of nasty bots that aren''t handling SGML or HTML properly and are leaving things like "&amp;" in the URIs.

Well, that just made my life a WHOLE BUNCH EASIER as a couple of the worst offenders are really sloppy like that so at the moment I've got at least 5 snares running just on their idiocy alone.

BTW, if I haven't mentioned it lately, that little bastard gnootBot is still attempting to crawl my site all these days later. Persistent little fucker, I'll give it that much.

Saturday is Scrapeday™?

Did I miss a memo?

Was Saturday designated Scrapeday™ and nobody told me?

Today's attackers should've been filmed and released onto video as "Bots Gone Wild" as there were as many as 5 at a time hitting me that were masking as a browser, well masking the user agent only, not their bad behavior.

The boys over a live-servers.net must've noticed they were being stopped and tried to send a new can of whoop ass from the following IPs:

  • 88.208.193.65
  • 88.208.193.67
  • 88.208.200.198
  • 88.208.200.236
  • 88.208.200.220
I'm so done with them now that anything coming from live-servers.net is just going to get the big FUCK YOU from first contact.

The other thing that caught my attention today were several scrapers that used multiple IPs to avoid detection, big shock, but was being tracked by the cookie again which stopped the morons as they hit several landmines. It's just priceless to think someone goes to all the trouble to use what most likely was a series of proxies to avoid detection but don't dump the cookie when the IP switched to a different network.

God I love idiots, they just make my work easier.

Unipeak Privacy Scraper

Don't you just love all these so-called privacy proxy sites like these monkey spankers over at Unipeak that claim to "Filter out unwanted advertising such as banners and popups" when in fact they insert their own unwanted ads on the page AFTER TAKING MINE OFF!

They appear to be downloading your pages from 207.234.209.125 so just block their asses.

Friday, March 10, 2006

CALL TO BOT BUSTERS - ACCESS LOGS NEEDED

We're moving into the next phase of bot busting and need your help!

Anyone out there willing to open their access logs to let the Mad Scientist take a peek by letting my scraper analyzer comb a months worth of data to see what kind of shit is hitting your fan?

I would be interested in other sites that generate traffic from 100K-400K visitors a month on this first pass but won't turn down 1M+ visitors either. Your site and traffic will be kept 100% confidential with only vague statistical data comparisons used from the final analysis.

The purpose of this experiment is to see if you're being abused by the same IPs and agents, to see if there is commonality in the spider traps being used, etc. and if not, WHO is abusing you and would my current scripts stop them dead in their tracks.

Basically, this is a quest for help to build the bot blocker Rosetta Stone.

In return you'll get a report of all potentially bad activity happening on your server that you could take action against immediately.

If you wish to contact me privately and anonymously, not in blog comments, then I suggest private mail to IncrediBILL on SearchEngineWatch, WebmasterWorld and ThreadWatch.

When we have enough sample subjects we'll post a comment here that we're closed for submissions.

Be a part of internet history that's about to unfold, submit your logs today!

Thursday, March 09, 2006

Matt Cutts Confirms People Are Stupid

Rarely do I write about something posted on someone else's blog but Matt's post about "How to sign up for WebmasterWorld" is absolutely priceless. According to Matt the problem arises when he references people to WebmasterWorld and then people write to him asking if there is a way to get into WebmasterWorld for free because the login screen implies you have to pay to join.

Let's give Brett Tabke kudo's for such brilliant marketing to underplay the free registration to access WebmasterWorld because there is a lot of valuable information over there and after all, he's running a business and not a charity. He once told me how much he pays a month for his servers and bandwidth and it's a shitload. I can't really say I blame Brett for making people think they need to open their wallets to get inside.

After a small amount of begging and pleading with Matt to kill the thread, I realize he's just too nice a guy to the people he's trying to help and isn't concerned that he's foisting idiots onto WebmasterWorld that can't even figure out a simple IQ test to get inside.

According to Matt:

if I wanted to post a pointer to a WMW thread, I didn’t want to get “how do I do it?” questions. And I’ve seen that from some people with high IQ.
Well just how freaking smart can they be Matt?

If you put cheese in a maze even a rat can find it eventually if they're hungry enough so why coddle these whiners that can't even help themselves to a free registration?

Not that I mind helping people, but I draw the line at holding their hands and wiping their asses when all they need to do is read and click, it's nothing mind bending.

If you want coddling then Matt's your guy, and a nice one too, no question about it.

I'll continue challenging people to think for themselves as I'm a firm believer in that teach a man to fish theory.

Googlenoia [warning, major rant]

What the hell is wrong with all you people?

Don't you have a fucking life that doesn't revolve around Google?

The last few days it's nothing but AdSense is crashing, BigDaddy is broken, Google stock is tumbling, blah blah get a fucking life blah.

First, losing some traction in the search engines thanks for some PhDs having brain farts at the 'plex doesn't mean your AdSense is broken or any of the conspiracy theories you can concoct to add to the mystique of AdSense. People using AdSense dance around it like the monkeys banging on the monolith in the opening scenes of 2001 a Space Odyssey. Come on fucknuts, unwrap the tightly wound tinfoil hats and use that steaming heap of gray shit called brains and realize that trends change, advertisers and budgets wax and wane, Google tests things now and then and on top of the list SHIT HAPPENS.

Funny, once upon a time Yahoo stock was skyrocketing and like all things that go up it came down but for some reason Google is different and held to different standards. Well too fucking bad you psuedo-religious Google freaks, it's not different. Google made a lot of people a lot of money, some would call them filthy fucking rich, and if you were too stupid to buy early or sell high then FUCK YOU for being a moron so stop whining and move on. For what it's worth you should probably sell at a huge loss before you lose everything and you're homeless living in your goddamn car because it's not going to rebound, the wild ride is over.

Last but not least, those crybabies that come out in droves every time every search engine changes, especially Google, and you or your customers go up and down from #1 to #5 or heaven forbid your sorry ass slipped to PAGE TWO and you're not in the top 10 anymore.

Well for those of you whining about your SERPs I have 2 words:
FUCK YOU!

With hundreds of thousands of pages competing for these top keywords you were lucky you got there in the first place and the fact that you couldn't survive a new way of indexing is just too damn bad. Nobody OWES you that position in the search engine and you don't have the right to demand getting it back, adapt and deal with it or just fuck off as I'm sick of hearing all your shit.

I'll bet conversely someone else is happy as shit they moved up and are dancing in the aisles getting all that free traffic that your whiny ass is bitching about as one man's tragedy is another man's blessing when it comes to search engines.

This Googlenoia is just getting so fucking old, can't we find something new to talk about?

Tell me about your latest project, found any good blogs lately?

JUST SHUT THE FUCK UP ABOUT GOOGLE ALREADY!

Enough.

Wednesday, March 08, 2006

Dumbest of the Dumb

Ever see a crawler look for "/#top" on your web site?

I'm still laughing over whoever wrote that shit.

UPDATE:
The dumb fucker came back today 3 times and tried to get "/#top" every time.

Appears to be some stupid fucker from The Netherlands using DHCP as each access was from something like dipshits.too.stupid.to.live.CHELLO.NL

Funny Funny Scraper Shit

Well, my scraper challenge page contains a URL in the sticky challenge loop that can vary per page that let's a human get past right away but keeps spiders looping until they try to index that particular new link, which would let them pass if they indexed it quickly, which of course is too late by the time they actually index it and are locked out for a while.

Ok, now the funny shit, these idiot spiders are now coming back looking for this link as an actual page so if you're not already in the sticky spider loop and ask for the page directly, WHAMMO!, you go DIRECTLY TO JAIL, DO NOT SCRAPE PAST GO, DO NOT SCRAPE 200 PAGES!

It's like a high-tech comedy show at times and you just sit back sipping bourbon waiting for the first asshole to set foot in the latest snare.

Ah, my side is killing me from laughing so hard at these fucking idiots.

Anonz Azz booted to infinity and beyond!

That's right, someone was on my site that came from anonz.azz.ru and I'll bet you're shocked that it has something to do with their website www.Anonymizer.ru aren't you?

Well a BOOT TO HEAD for both of them!

And one for Jenny and the wimp...

Maxwell's Silver Hammer

Most of you probably think I'm very brash and run amok implementing things on my server all willy-nilly with hardly a concern for the damage that I might be inflicting on my visitors but that's further from the truth than you can imagine. I'm actually very cautious and do a lot of testing with each new approach I phase into my bot buster by first executing the rules and giving me a preview of what would happen for a day or so before I make the rules live.

That means to date all the IPs I've been banning are being banned in software so that I could monitor their returns and activity to verify they're really a permanent source of abuse or a one-shot attack from a dynamic IP.

Well, enough of them are returning on a regular basis that I've decided it's time to start the next phase of the project which I call Maxwell's Silver Hammer where it will decide automatically that the source of abuse is bad enough and just drop them in the .htaccess file so they simply bounce off the server and don't even tie up my scripts keeping an eye on them anymore.

So here we go sweating bullets that this code won't accidentally crash some night and leave the .htaccess file all banged up and bring the site down to it's knees.

Progress, gotta love it.

Monday, March 06, 2006

Pathologically Extreme

Yep, that's what my bot busting obsession was called today in private email.

Now that I'm "pathologically extreme" I must thank the person for his bluntness as it did bring up the point that there's a lot more to this bot busting issue that someone sitting on the sidelines only casually familiar with my quest and this blog may know.

In all fairness, if you told me I'd be on this quest to abolish unauthorized access to my site 12 months ago I would've laughed in your face and said "what harm does a little crawling do anyway?" and yes, I used to hold those tightly wound content control freaks in low regard as misguided time wasting fools.

However, then I decided to get out of the consulting game and focus more attention just on my own web sites which bring in a decent revenue stream without all of the whining and hassles of customers.

That's when all hell broke loose as suddenly both the spammers and scrapers started hammering my old server so hard it was going down all the time. Not physically crashed mind you, but it was just so busy serving the needs of spammers and scrapers that my income needs weren't being met whatsoever. We're talking DOS attacks because of the sheer speed and volume of this nonsense and the server just didn't respond for 5, 10, 15 and the worst was 90 minutes at a shot. It got SO BAD at one point I had to completely get rid of server side spam filtering as that tool itself could use up all the CPU when some spammer came along doing a pump and dump of spam.

These shameless greedy bastards were impacting my site, my SERPs, my wallet and really pissing me off - the shit had to stop.

First was the easy part which was just getting rid of the spam. I blocked email coming from most of Asia and Russia which eliminated the majority of the high speed spam dumps and gave me some breathing room to work. Then I made the only way to contact me a form on the web sites, eliminating all email addresses but 2, and literally set the server not to BOUNCE emails but REJECT emails. Why I did this is bounce emails still come into your server and attempt to send a response back but most spam has a bogus reply address and thousands of bounce emails quickly fill up the queue and your email system grinds to a screaming halt processing bounce deliveries all day long. Trust me on this, just REJECT those undeliverable emails, no bandwidth or CPU wasted at all as they just bounce off your server harmlessly never to be seen again.

Guess what?

Asia and Russia are no longer blocked as REJECTing their emails stopped them from being a threat.

At this point there are never more than 10 emails sitting in my mail queue at any time and the spam that gets thru is literally a handful of emails a day, blissfully under control, I love it.

However, after solving this problem the old server was still going down like it was under a spam attack and after a while I came to the conclusion my site was probably just too busy to handle the load and my older slower server just couldn't deal with the demands of all the visitors, search engines, etc. and set out to upgrade.

Now, with a big fast shiny dual Xeon box it's back up and running faster than ever.

Two weeks later some fuckers took it offline for 90 minutes in the middle of the night and I lost my shit, that was it, the straw that broke the camels back, no more Mr. Nice Guy.

.... this was war....

Then the whole process kind of evolved into a huge eye opening adventure at this point and being a naturally curious guy and a programmer with a huge ego [yes, I am IncrediBill and I can stop these bastards] it kind of took on a life of it's own.

First, stopping the high speed scrapers was easy, totally childs play.

Next, the sheer volume of scraping became apparent once I was monitoring real-time site activity while squashing the high speed scrapers and looking for other unauthorized resource wasting bots.

Evolution just kept happening as one thing led to another, stopping more scrapers unearthed even more scrapers, that the errors I fed scrapers unveiled tons of sites with MY SHIT on them, and that many people had apparently built AdSense-incentivized businesses based on bottom feeding off my business and in the process were diluting keywords I was earning money from using my own content against me.

OK, now THAT pissed me off even more.

So while some of you may call the depths and extremes I'm taking to protect my shit as "pathologically extreme" my side of the story is self-defense for my very survival and I'll be damned if some bottom-feeding leeches are going to take me down without a good fucking fight.

Yes, that's it people, as far as I'm concerned at this point it's a fight to the death, theirs and not mine if I have anything to say about it. So far I've spent a ton of time and money addressing the issues and at this point it's paying off but for how long only time will tell but it's definitely going to be a death match for one of us.

Now, go buy my CD's and t-shirts in the lobby to help fight the cause and invite me as a motivational speaker at your next nerd conference to spread the word!

Nah, that would be WAY too extreme!

UK Scraperz in da Hood

There really aren't too many truly persistent jerks out there and I hate to rat out sites or IPs unless I'm sure they are truly rotten but this wins the award for most annoying IP address of the year. I blocked 88.208.194.241 a long time ago and it just never stops, it's relentlessly attempting to crawl, day after day, asking for thousands of pages and doesn't fucking take NO! for an answer.

A quick reverse DNS on this idiot shows it's apparently on some UK server farm, possibly run by fasthosts.co.uk, that IDs itself as server88-208-194-241.live-servers.net.

Out of curiousity I went back and ran a reverse DNS on all my blocked IPs and big shock, this server farm has a few hits in my list.

server88-208-192-204.live-servers.net
server213-171-220-120.live-servers.net
server88-208-194-241.live-servers.net
server88-208-194-252.live-servers.net
server88-208-195-4.live-servers.net

Matter of fact, I just reviewed my logs for today and all of the 88.208.19* IPs listed above hit the server so it appears to be distributed scraping.

So I'm thinking it's probably not such a bad idea just to block their whole damn neighborhood based on the non-stop abuse I'm getting from a couple of these IPs.

I would block at a minimum:

88.208.192.0/24
88.208.194.0/24
88.208.195.0/24
213.171.220.0/24

I'd keep an eye out on anything from this range as well:

inetnum: 213.171.220.0 - 213.171.223.255
netname: FASTHOSTS-UK-NETWORK

They claim to resell dial-up and broadband otherwise I'd suggest blocking the whole damn network but at this point I'm not sure 100% if we're looking just at servers or surfers but so far blocking what I've blocked doesn't appear to have any negative impact on my site except to stop their stupid bots from downloading my content.

Let me know if anyone else is seeing activity from these guys and what IPs you're seeing as this is about as bad as I've seen and they need to be stopped.

Alexa Bowling for Matt Cutts

Don't know who did it or how they did it but it was a stroke of genius to load Alexa up with a bunch of related sites for Matt Cutts that are obviously spam.

Wonder how hard this was, just write a script using the MSIE toolkit with Alexa's toolbar installed to just referer spam the crap out of his site?

Is it possible from a single IP to game Alexa or would someone need to enlist a legion of anonymous proxy servers to pull this off?

Or, could it be so simple that someone has dissected the Alexa toolbar API and just fed it a load of junk about Matt?

We may never know unless someone confesses but I'm thinking it might be worth a giggle to try and see what it takes to make something like that happen.

Good thing I'm not bored today!

Sunday, March 05, 2006

Best Bot Ever - Almost!

Tonight I saw the future of sneaky bots and it did everything to look like a human so I'm thinking they used a developer toolkit to drive the crawl via MSIE. This thing downloaded images, banner ads from 3rd party servers, ran javascript and even accessed AdSense ads so it was as convincing as anything you can imagine.

Spooky and amazing in how well it cloaked itself.

I couldn't tell by looking at the log files either, very impressive.

Then after patiently crawling for a whopping 11,097 seconds or more than 3 hours for those of you that can't divide 11,097 / 3600 in your head, it exceeded my max page count which is set fairly high.

Then it proceeded to very slowly and stealthily ask for 20 more pages after being told it had exceeded it's daily limit of pages.

BUSTED!

However, the point being if it had been just a bit smarter to realize it was getting a repeated error page I'd have never known it wasn't a human.

Not a good turn of events whatsoever!

Nasty.

Every Webmasters Worst Nightmare

Here I sit working on my website late at night, we're talking LATE at 3am, so I can test some new bot blocking technology in the middle of the night when the traffic is low so impact will be minimal if I screw up and knock everyone offline.

Ok, upload some code, bug, quick patch, bug, something weird going on so hop onto SSH and look at server and something is chewing up CPU like crazy and website isn't responding properly. Whew, it's an automated nightly update that only lasts a couple of minutes and everything is back to normal.

Test some code, find another minor bug, upload some new code, run the page, and it goes and goes and goes. What the heck, did I just blow something big time? Check SSH again, staggers a little and stops, then generates a socket error and closes.

Holy Fuck! Server Down!

So I click here and there in my browser, EVERYTHING is down!

YAY! I didn't crash the server, the cable modem is offline.

30 harrowing minutes later the damn cable modem comes back online and everything on the server is running just fine except I left a debugging message that's displaying on all the pages.

Remove that message, test it again, everything seems to be OK, time for bed.

Whew!

Bracing for nightmares tonight....

Bandwidth sucking Arachmo

Some speedy little crawler from Japan named Arachmo that I've never seen before set off my speed trap tonight so you may want to just block this pest before it hits your server.

60.237.36.157 "Mozilla/4.0 (compatible; Arachmo)"
At one point this little bastard was asking for more than 10 pages a second.

Maybe Godzilla will just step on the damn server while fighting with Mothra some day!

Saturday, March 04, 2006

Competition claims WE'RE NUMBER TWO!

Just about chuckled my ass off when I was snooping some of my competitors sites, not the blogs silly, my money making site, the site that keeps me in expensive booze and wide screen TVs, not this bullshit.

Back to the story as I'm DigressJacking™ it already.

Anyway, this site we've discussed before as they undersell ad prices, send out newsletters begging for advertisers to fund new projects, etc. but now it gets even more amusing. Somehow this guy has mysteriously moved up the ranks in Alexa which is now listing him as #2 in his category right below me, oh whoop de doo, I've been ranked #1 in that slot in Alexa since they first set up shop, so you bumped someone to get to #2, big fucking deal. Funny he doesn't mention anything about Google's Directory by Rank that has me still at #1 for years and he's way down the list over there, not a peep.

Then he calls out one other place where he's #2 below me for some bullshit meaningless 2 keyword search term in Google that doesn't even drive traffic.

Excuse me?

Is that smoke I feel blowing up my ass?

Hate to burst your bubble buddy boy but I'm in the top 10 for 2 keywords that are tops in the field, ONE WORD, NOT TWO, yes a SINGLE KEYWORD that drives thousands of visitors daily and it's not a BULLSHIT TERM, people actually actively fucking search this term!

So then he goes on lamenting about how his ads are a better value at his rock bottom bargain basement prices than "some other more expensive sites".

Give me a fucking break, I send people more traffic from my ads in a day than you sometimes send in a month. He knows it's true too because he used to have a page showing the traffic for all his ads and took it down, good thing too, it was embarassing.

I've been toying with making any mention to his references but so far all I did was put up a nice bold line of text on my ad page that says "We may cost more but you get what you pay for - performance" and left it at that.

I'm thinking at this point the best way to deal with this putz is just ignore him and let him dig a deeper hole until Google determines he's supplemental like they just did to an even lesser competitor and relegate them to the search engine dungheap.

Wrong Number, Hang Up Already

Well, someone with their Nokia phone just didn't take WRONG NUMBER for an answer and tried to beat the shit outta my server looking for a place that would take their call.

209.191.82.253 Nokia6600/1.0 (4.09.1) SymbianOS/7.0s Series60/2.0 Profile/MIDP-2.0 Configuration/CLDC-1.0

Actually, it looks like Yahoo might've done this on their behalf as the IP address resolves to msfp02.search.mud.yahoo.com so it might not have been the phone that was so insistent.

03/04/2006 05:30:53 "/"
03/04/2006 05:30:53 "/mob"
03/04/2006 05:30:53 "/index.wml"
03/04/2006 05:30:53 "/index.xhtml"
03/04/2006 05:30:54 "/default.wml"
03/04/2006 05:30:54 "/default.xhtml"
03/04/2006 05:30:54 "/home.wml"
03/04/2006 05:30:54 "/home.xhtml"
03/04/2006 05:30:55 "/mobile"
03/04/2006 05:30:55 "/mobile/index.wml"
03/04/2006 05:30:55 "/mobile/index.xhtml"
03/04/2006 05:30:55 "/mobile/default.wml"
03/04/2006 05:30:55 "/mobile/default.xhtml"
03/04/2006 05:30:56 "/mobile/home.wml"
03/04/2006 05:30:56 "/mobile/home.xhtml"
03/04/2006 05:30:56 "/mob"
03/04/2006 05:30:56 "/mob/index.wml"
03/04/2006 05:30:56 "/mob/index.xhtml"
03/04/2006 05:30:57 "/mob/default.wml"
03/04/2006 05:30:57 "/mob/default.xhtml"
03/04/2006 05:30:57 "/mob/home.wml"
03/04/2006 05:30:57 "/mob/home.xhtml"
03/04/2006 05:30:57 "/wml/index.wml"
03/04/2006 05:30:58 "/wml/default.wml"
03/04/2006 05:30:58 "/wml/home.wml"
03/04/2006 05:30:58 "/xhtml/index.xhtml"
03/04/2006 05:30:58 "/xhtml/default.xhtml"
03/04/2006 05:30:58 "/xhtml/home.xhtml"
03/04/2006 05:30:59 "/wap/index.wml"
03/04/2006 05:30:59 "/wap/index.xhtml"
03/04/2006 05:30:59 "/wap/default.wml"
03/04/2006 05:30:59 "/wap/default.xhtml"
03/04/2006 05:31:00 "/wap/home.wml"
03/04/2006 05:31:00 "/wap/home.xhtml"
Sorry, if you'ld like to dial my website, tough shit.

Keep your hands on the fucking steering wheel.

gnootBot still going and going....

As reported the other day we seem to be the first ones reporting on gnootBot and it has never downloaded a single real page from our site but seems to somehow knows every page that's on my server and keeps slowly asking for pages day after day.

Wonder where the pages names from?

Possibly crawled me before installing the bot blocker or downloaded a list of pages from a search engine or some shit.

Doesn't matter as it's still going and there's obviously nobody at the wheel as they are getting nothing but error messages.

Hope it's worth your time when it's over asshole.

FIRST SIGHTING: Sproose Goose got Plucked

Yet another Silicon Valley startup search engine called Sproose came crawling this morning tagged as sproose/0.1-alpha using Nutch. Well, in their site it claims they have seed funding from VC's, also reported elsewhere, but you can't do any searches yet as they are currently building their Knowledge Rank™ which sure sounds a lot like PageRank, huh?

The first knowledge they got when they hit my site was that they didn't rank high enough to crawl my content and got the door automatically slammed in their faces by being an unauthorized bot. Sorry boys, robots.txt is so 90's, we use razor wire around the compound to keep people out these days.

You may have seed capitol, but I require being wined and dined before you may crawl my 40K pages, or just email mail me with a PLEASE as this entitlement mentality to crawl every site and run up our costs online just because you have been funded is BULLSHIT.

BTW, the people that wrote NUTCH should be hauled out in front of a firing squad and shot as I'm seeing more and more crawling from their little engine that couldn't constantly bouncing off my site.

Friday, March 03, 2006

POP QUIZ: Your Site Already Has a Spider Trap?

Most of you web site owners already have a spider trap on your web site and you don't even know it. There are about 3 pages that humans almost NEVER READ and spiders gobble up daily so all you have to know is which pages these are and then grep for them in your access logs and VOILA! you see a list of mostly spiders hitting your web site that can be blocked at will.

Once you get a list of who's been looking at youre spider trap pages, simply take each IP in the list and then grep for all activity for that in your access log. When you see hits to all pages and no images loaded it's a clincher you got a spider but just dont get carried away and block Google/MSN/Yahoo.

Even if an entry in your access log says Googlebot as the user agent it may not be Google, so check out where the IP resolves and make sure it's Google.com in the domain with a reverse DNS lookup, which you can do on DNS Stuff if you don't have other tools available.

Now, anyone want to guess which 3 pages or files on a web site are spider traps?

I know I give you all a lot of information but anyone should be able to figure this out by staring at any typical web site and see which links you would never click.

If nobody can figure it out MAYBE I'll tell you on Monday, if I'm in the mood and if I remember.

Come on people, POP QUIZ! post your guesses, don't be shy!

Wednesday, March 01, 2006

Scrapevertising

Well here's a new combination that really pissed me off with a scraper and referer spammer all rolled into one. Saw this spunk monkey crawling my site, the bot blocker was already stopping his ass, but something caught my attention in that every page crawled had the same referer as the origin. Sure enough, when I went to look at the site attempting to get notoreity from my access log it was a new directory web site that was using scrapings to get attention to the site for people submitting listings.

Holy shit, you mean to tell me people are stupid enough to click "submit url" when they can plainly see all the listings are junk that don't even have valid links out?

Apparently they are that stupid as the bottom of each so-called directory page appears to be actual submitted listings opposed to the scraped crap content without outbound links at the top trying to snare search engine love.

Now I'm steamed, this is NOT a pretty turn of events.

New AdSense Setup Wizard!

Apparently as more less technical types have joined AdSense in droves Google decided to dumb down the AdSense setup interface to reduce the tech support load and dependence on certain browser features.

The tabs across the top of the AdSense website for AdSense for Content, AdSense for Search and Referrals have all been replaced with a new tab AdSense Setup which now contains all the options for the previous tabs.

When you click on the new link for AdSense for Content you're in a 3-step wizard that drives you thru the process.

  1. Choose Ad Type (ad unit/link unit)
  2. Choose Ad Format and Colors
  3. Get Ad Code
My suspicion is many people confused people were accidentally changing options and getting the wrong code, or forgot to get the code, etc. and this will simplify the process for that overloaded page. It also makes it more clear which ad type you're getting as some complaining noobs in forums weren't understanding why they only got links and not text ads, so you know those poor people in AdSense support were crying in their beer every night.

Problem is, anyone with a serious amount of channels trying to get the code will be suicidal with this back and forth nonsense between those two pages. Anyone want to join a class action for Google inducing carpal tunnel with this new page? Just kidding, well, maybe.

I can see why some of these changes were needed for them but it sure would be nice to have an "expert mode" for us AdSense old-timers that actually are power users of the page they just ripped apart.

The only thing I can't believe they didn't fix in this update is making the CHANNEL page accept multiple URLs in a textarea instead of one URL at a time. Try setting up 200 URLs with this stupid page opposed to a quick cut-n-paste from your list of page names on your web log analyzer and you'll see what I mean, it's brutal and stupid, yet it remains unchanged.

Now for the million dollar question:
"What are they about to add to AdSense that required this overhaul to make more space?"

Let's keep an eye out and see what happens.

Tuesday, February 28, 2006

Comment Wars in Rantville

Normally I wouldn't just blog about a stream of comments but this article about Referer Spammer Revenge is almost a month old and the shit just hit the fan yesterday. I'm thinking the person I mentioned figured out it was me posting about them from my comment at Spam Huntress and went off the deep end.

We're talking some serious flaming here, cracked me up, but you be the judge.

Let this be a lesson to any children that may be reading this shit as it proves you shouldn't stick your finger too far up your nose as you can impair your thinking if you accidentally stab your brain and you'll end up a referer spammer too.

FIRST SIGHTING - New Bot Discovered

It's very rare I run across a bot you can't find any details about in Google but last night something calling itself "gnootBot" came from 66.39.177.8 which just has a default Apache web page.

No information about this beast except it started asking for pages in the middle of my website.

Very bizarre.

Monday, February 27, 2006

Peeling the Scraper Onion with Reverse DNS

Stopping scrapers appears to be like peeling an onion in that when you peel away one layer of bad bot activity you unearth yet another whole new layer that you couldn't see before. You start with user agent filtering then put up speed bumps and honeypots to stop others and even profile other behavior to stop them and they still keep coming. Now we're several layers into this scraper onion all sorts of new things are showing up that require more sophisticated methods to detect and block as they're hiding as browsers, running low key crawls, but still obvious to anyone that it's not a human if you look at the pattern of access.

To help thwart more of this nonsense the latest tool added to the bot blocking arsensal is reverse DNS lookups to see where the IP originates from and blocking or challenging bad sources from the start. There have been a few trends to become very obvious in that many scraper IPs that were auto-blocked didn't resolve to a domain name whatsoever or come from some suspicious hosting farms, the two most notable and persistent ones have been in Taiwan and the UK.

Now they'll need to find yet another way to get around the bot blocker as the newly installed steel door now has a chain, 2 deadbolts and a mean as piss rottweiler waiting on the other side just in case.

Stay tuned for more on the next episode of As the Onion Peels.

Saturday, February 25, 2006

Real Impression Tracking in Bot Buster

There was an unexpected benefit to all this bot busting and information recording in that now I can track on a daily basis all pages served to legitimate bots, blocked bots and actual people. This has allowed me to come up with some simple stats that shows a breakdown of where pages are served and I'm getting real close to almost matching exactly what you see in Google AdSense for page impressions.

This side effect alone could be an enormous benefit to people that need actual page impressions vs bot page impressions for selling advertising on their website and I've just made it a heck of a lot simpler to get that information since every single page passes thru my code in real time.

Coolio!

Cleverly Masked Bots Evolving

It would appear that the war over my content has been cranked up a level as bots masking as browsers and modifying their behavior to appear like people seems to be escalating. There are still some tell-tale signs that are easy to spot when you look at the server log but a couple of them that the bot blocker didn't catch are finding ways to game the system.

I didn't want to make the site more difficult for visitors but the only way to stop these guys would appear to be tossing in more random challenges like captchas and such after a pre-determined number of pages. To stop the typical captcha blow-thrus the challenges are very random and nobody could program a way to bypass them all as you don't know what they all are and I can add new ones daily if I wanted.

There's also something I noticed which isn't earth shattering but only humans seem to use my javascript menu which is a HUGE tell. Robots navigate the text links only but humans love those drop down lists and that's a clear sign that differentiates the two of them most of the time.

At the end of the day, it's just like trying to secure money in a bank, no matter how hard you try someone is going to rob you eventually but the best you can hope for is to make the number of times you get robbed as minimal as possible without pissing off all your customers in the process.

Thursday, February 23, 2006

Fuck Your Intellectual Property

Some asshats claiming to "defend your brand" sent their little AIPBOT to crawl my pages looking for anything of their clients on my site. Listen up fucknuts, you can use my SEARCH tool and look for something being on my site but you can kiss my ass when it comes to a 40K page crawl just to see if I'm violating someone's precious brand name.

This sense of entitlement of everyone to crawl the web is really starting to piss me off.

Take a hike assholes.

Link Me Or Else!

Here we go again with a persistent raging linkaholic badgering the shit out of me to link to his crappy little directory website so he can get a few AdSense clicks.

Hi YouBigWebStudYou,

I sent you a link request for bullshit-directory.com to see
if you would be interested in exchanging links.

I realize you are probably but wanted to let you know that I will be
removing your link next Wednesday if I don't hear back from you.

You can verify your link is by going to:
http://realfuckingannoying.com

with the following details
Title- Link To Me Please
URL-www.imbeggingyou.com
Description-I'm the biggest pain in the ass link-to-my-site whining spammer you've ever seen so link to me now before I beg and plead more.

If you do add my site please use the below information and let me know
the location you added it so I don't remove the link to yours
unknowingly.

I hope to hear from you before 2nd March 2006 but if not then I'm bound
to remove the link from my site.
Well fucking remove me already and stop sending me this shit!

Didn't the dead silence after your first spam give you a fucking clue?

If I could get my hands on you they'd have a new opening homicide scene for CSI next week so just keep it up, your luck is about to run out.

Oh yeah, it's a good thing you're in India so CAN-SPAM can't be used against you and you aren't registered with GoDaddy so they can't blackmail you to get your domain back. However, if there is a god one of those nasty little bugs in your water will give you atomic diahrea and you'll shit out a vital organ and die.

Wednesday, February 22, 2006

Search Engines Let Scrapers Bypass Spider Traps!

Just when you thought you've seen it all the actual search engines themselves can be used by scrapers to bypass spider traps. How this is accomplished is the scrapers find all of the indexed page names from your site in Google or Yahoo and then download pages the using known page names from your site thus side-stepping spider traps as they aren't actually spidering your site at all.

Therefore, just eliminating your pages from being CACHED in the search engines doesn't stop scrapers from still using the remaining data to their advantage.

Some days it just doesn't pay to get out of bed.

Tuesday, February 21, 2006

Another Plug-n-Scrape Component

Yet another toolkit letting armchair programmers attempt to grab my web pages.

Yawn.

This one ID's itself as:

IP*Works! V5 HTTP/S Component - by /n software - www.nsoftware.com
And their web site claims:
The HTTP component can be used to retrieve documents from the World Wide Web.
Might want to revise that to "used to be able to retrieve documents" as it went splat against my brick wall but I found it's calling card in my auto-blocked bot log.

Chitty Content

Must be the new wave of affiliate bots as I also got hit today by the Chitika ContentHit crawler or whatever the heck it is.

No way to verify this bot as reverse DNS for this IP address just claimed to be from Charter Communications.

71.10.233.52 Chitika ContentHit 1.0
This new bot of the day chit's just getting old.

Cell Phones and PDAs Can Piss Off

All these damn cell phones and PDAs all have unique user agent strings and for the last few months the handful that hit my website are all being told to piss off.

You people making cell phones and PDAs better wake up and smell the coffee as I'll be damned if I whitelist a bazillion user agents just to let your pissy products see 5 lines of my web site.

You all better come up with some better ideas for cell phone user agents as this unique name per phone shit isn't gonna fly.

CJ Quality Bot

Well here's a new one that I've never seen before from our friends at Commission Junction.

216.34.209.23 CJNetworkQuality; http://www.cj.com/networkquality
Unforunately they bounced off the walls, think I should let them in?

They might delist my site if I don't but based on the revenues I earned with them last month it's kind of a why bother IMO.

Fine, time to whitelist CJ, sigh.