Saturday, February 11, 2006

Scrape Me Up Scotty

Finally I've had an out-of-this-world scraping experience when someone tried to offload about 500 pages via a satellite. Thanks to the fine people over at DirecPC for hosting this little bandit my scraper stopping efforts have now entered the realm of aerospace.

I can see the headlines now:
Earthly Bot Stopper Blocks E.T. Scrapers

It's safe to hazard a guess Ray Bradbury never, in his wildest imagination, thought people would use satellites to steal websites.

Friday, February 10, 2006

Scraper Sites are GOOD?

Usually I have a good deal of respect for Martinibuster as he's a cool dude and has some good insights into a lot of SEO and webmastering topics. However, when I read his article "Scraper Sites are Good for You - Surrender Your Content" my first thought was to print it on toilet paper so I could give it the proper respect it deserved.

Come on Martini, you can't be serious about letting people waste your bandwidth, steal your content, and spam the SERPs with your own junk just to give you links?

Martini claims "Scrapers: There is Nothing You Can do About it" which is totally wrong! I'm stopping those bots, AlexK's script snares them and BotBuster and Bandwidth Protector claim they will also stop them. Can't comment or endorse the latter products as I've never used them but they do exist as well as some others so webmasters not stopping scrapers just means they're lazy or cheap at this point.

Then Martini makes me spit soda across my keyboard with "Surrender to the Scrapers... It is Better for You". What a load of crap my friend. Just ask Aaron Pratt of SEOBUZZBOX how good surrendering to the scraper is as he's being lumped into supplemental results as the scrapers aggregators are being indexed first and Google thinks Aaron is the duplicate content which is just wrong.

So Martini says "scrapers do more good to you than harm" which gets the bullet list:

  • Scrapers can put your content in supplemental results
  • Scrapers can rank above you in the SERPs and get the first shot at AdSense and affiliate income using your content
  • High speed scrapes of 100s pages/second are like DDOS attacks and knock servers offline until they stop scraping keeping visitors from clicking your ads
  • When servers can't respond due to scrape attacks Google, Yahoo and MSN get time outs on pages and SERPs drop
  • Rampant scraping can run up your bandwidth charges and you pay for their excess
Nope, no harm no foul, nothing wrong with scraping.

What I can state with certaintly is that since I've started blocking scrapers my SERPs and REVENUE are both up substantially and that's about the only major change I've made to my site recently.

Try reading one of Martini's other articles that made sense instead, he's really a nice guy, just slightly misguided on the topic of scrapers is all.

P.S. Doesn't scraper and bot stopping sound like a great session topic for PubCon Vegas this year?

Thursday, February 09, 2006

Referer Spammer Revenge!

Today I caught some referer spammer bombarding my web server by hitting the same page over and over and over with only the referer changing.

To put it mildy, this pissed me off.

When I started looking up each domain I noticed something fascinating in that they all had the same AdSense account and were all registered at GoDaddy.

Recent topics on ThreadWatch about GoDaddy locking abusers domains provided a true inspiration today. Instead of wasting my time putting this asshole in my banned list of IPs and domains to keep him out as would be my normal routine, I reported him to both AdSense and GoDaddy abuse and will be waiting and watching to see if either of them take action.

If GoDaddy shuts his entire array of websites down this will be the best defensive action to take against referer spamming yet, I'll post the results if any when I see the domains go offline.

BTW, I also added referer spamming detection to my bot blocker today so I'll be snagging more of these idiots on a regular basis if this is a huge problem. I'm considering making the bot blocker automatically perform a whois on the domains when this is detected and send automatic abuse letters to the proper parties.

This could be fun ;)

Wednesday, February 08, 2006

Polish Robot You Can't Pronounce

Only a polish robot would be looking for polish websites on my server in Texas - sigh.

Anyway, we found Szukacz trying to snoop around but alas, it slammed into the great wall protecting my site.

Looks like it supports robots.txt according to the web page but who knows.

Burf Barf Puke

Someone in jolly old England unleashed Norbert the Spider upon my site which appears to be sent on behalf of the fine people at Burf that claim "BURF - Alternative Search Engine and Entertainment Portal"

Alternative to what, finding what I'm actually looking for?

Then I dediced to click on the "Your Ad Here" link just to see what one gets for the money to advertise with Burf and it's placement on a truckload of search sites I've never even heard about.

Well, it's cheap advertising but then again my URL would probably get about as much exposure putting it on the bottom of my shoe.

Yacy my Assy

Filed under "Who Needs Another Crappy Search Engine" we find something called Yacy that claims to be a P2P distributed search engine whatever the fuck that means. I'm translating that back into English as a free-for-all scrapefest bandwidth waster.

They claim it's "Easy to install!"

I claim it's "Easy to block!"

Toodles.

Monday, February 06, 2006

Anti-Social Bookmarking

Well here comes the new leech-of-the-week Susie came crawling and went head first into an error page. Sych2It claims that "Dead and out of date links are automatically reported" but I'm not sure how they would know for sure as they got a slap in the face instead of a web page.

Fine German Engineering? HA!

The AnonyMouse proxy slammed head first into my bot blocker.

Come on people, if you're gonna waste your fucking time writing an anonymous proxy server the least you can do is attempt to fake being a browser instead of making your user agent string your domain name.

Give me a goddamn break.

Friday, February 03, 2006

Big Decisions Time for Bot Blocker

Getting real serious about converting the bot blocker prototype into a product and all the agony that goes along with developing, launching and supporting a new product brings back fond nightmares, um, memories of products launched in the past.

Trying to determine if I should write it myself or hire someone to write it to speed it along, look for equity partners from the beginning, whether to even sell it as a product, open source it, or whatever the hell to do with it and the whole process is just maddening.

Talking to a few people the last few days, we'll see how it goes.

I knew there was a reason I've been claiming to be retired the last few years!

SuperBot Found My Kryptonite

Sorry you're not so SuperBot after all as you couldn't even leap over my index page without tripping and falling.

The author claims:

Unlike other offline browsing tools, SuperBot is fast AND powerful AND and easy to use...
Whoops!

Should now read "Just like all offline browsing tools it was stopped dead in it's tracks when it hit IncrediBILL's Bot Blocker"

Better hope your customer that tried to download my website doesn't come looking for a refund!

EUREKA! Proxy Detection Thanks to Idiots

Thanks to some sloppy work by some rank amateurs running proxy servers they pointed out a flaw that many anonymous proxy servers share that are now allowing me to automatically detect proxy usage and block the damn things in real-time.

Of course this doesn't work on all proxy servers but it caught 10 of them just today.

I should've spotted this happening weeks ago but at the time I came up with a different hypothesis for the data that presented itself which today, with additional clues, makes it obvious many of these hits are via proxy servers.

Wow, it's amazing how such simple revelations can rock your world so easily.

This is cool - blog ya later as I need to work on proxy busting now ;)

Scientific Search Halted

Here's another crawler that lost it's way called Scirus that claims to be a scientific search engine yet was trying to roam around my non-scientific web site. They claim to be powered by Fast but they were stopped dead in their tracks by my bot blocker that said not-so-Fast, heh.

Sorry boys, you need to find a new lab rat to play with.

Rufus is a Dufus

When it just couldn't get any sillier along comes something claiming to be RufusBot which claims to be a good little bot but others claim it's personal scrapeware.

I don't care either way as it's not scraping here and don't let the door hit you on the ass on your way out Rufus.

Proxy Mouse Ate the Poison Cheese

Too funny, my spider trap killed a mouse, ProxyMouse as a matter of fact.

These leeches are just like the other proxy service that strips all of my ads off the page by default and slaps their own Yahoo ads on the top of the page but NOT ANYMORE!

What's best is how I caught them thanks to MSN which somehow attempted to crawl my site thru their proxy. When MSN was crawling their proxy server just passed whatever user agent string, in thie case msnbot, thru to my server. Suddenly my server sees msnbot on an IP address that doesn't belong to Microsoft and SNAP! they are busted.

No more Yahoo income for you off my back, your filtering asses are done.

Buh bye, see ya, wouldn't want to be ya!

Thursday, February 02, 2006

Yahoo Blogs Doing a Content Crawl

Not sure what good old Yahoo is up to exactly as this spider hasn't shown up in my spider trap before but Yahoo Blogs is attempting to crawl the content from my RSS Feeds directly. Not sure if they intend to display the content directly to the user or still redirect to my site so I'm not sure I'm letting them step past the RSS feed just yet.

209.191.83.13 Yahoo-Blogs/v3.9 (compatible; Mozilla 4.0; MSIE 5.5; http://help.yahoo.com/help/us/ysearch/crawling/crawling-02.html )

Official Name: crc4.opn.search.mud.yahoo.com
IP address: 209.191.83.13

Maybe it's harmless and I'll let it pass, something to contemplate over a beer.

Why is WaveFire crawling?

Some Canadian consulting company called WaveFire has a bot trying to crawl for reasons not divulged on their website. Would've been nice if they posted something about what purpose they had in attempting to crawl sites.

64.141.15.109 Wavefire/0.8-dev (Wavefire; http://www.wavefire.com; info@wavefire.com)

Official Name: search-d-02.internal.wavefire.ca
IP address: 64.141.15.109
Sorry, but your fire was put out and your spider was splattered.

Wednesday, February 01, 2006

Expanding Bot Blocker to More Sites

This week I'm going to install my bot blocker on 2 additional websites and see what level of abuse these sites are taking just to see what kind of an impact this technology could have on smaller less active web sites.

Don't get your tits in a twist just yet as it's still not converted to PHP and is still just a protoype no where near being a product for release.

On the more amusing side the message my site pops up when scrapers hit telling them they've been stopped has started showing up in the search engines as these automated scraping idiots haven't realized they're showing the world just how stupid they are.

What's more telling is which search engines show which sites with these messages vs. others that don't as I'm getting some insights into what some search engines are blocking as spam sites.

Blasphemers use IncrediBILL in VAIN!

Someone is using my name in vain to describe someone else in some big embittered bullshit battle about some religious horseshit.

Just what I need, now a bunch of idiots [you know who you are] will think I'm that person.

Fucking lovely.

Bend Over When Upgrading WebCeo

Finally decided to upgrade to the latest WebCeo and it did pop up a message about "losing reports" in that all existing reports should be printed before upgrading etc.

OK, big whoop, didn't need the old reports, could care less.

What that little message DIDN'T SAY was you would lose all projects, all profiles, and all that stuff and didn't even bother trying to automatically upgrade the data from the previous version leaving me to think all I was going to lose was the ranking information in the reports.

The new WebCeo version is now installed and pops up completely blank.

FUCK!

WebCeo needs to make those warnings a little more specific:

  • You'll lose all reports
  • You'll lose all projects
  • You'll lose all profiles
  • You'll lose EVERY FUCKING THING
Not that this was a complete crisis as I maintained a list of all the keywords for the ranking reports in a set of separate text files as WebCeo doesn't [didn't] have any easy way of just exporting my keyword list so I maintained it externally which worked out in the end as I just cut and paste the whole list of keywords back into the projects as I recreated them.

Back in business but the next time they release a major upgrade I may be tempted to try a different product instead as this was not amusing.

Excrutiating Back Pain

Pulled a muscle the other day which resulted in the brief hiatus from the blog as sitting in my desk chair long enough to write a rant was just too painful and I was in too much pain to use the laptop even as looking down pulled the muscle and hurt so I caught up on television watching instead.

BTW, when I'm in pain the level of cursing escalates to a new plateau so the blog may become TV-MA rated over the next few days until my back gets better as I need to release all this pent up hostility somewhere.

You've been warned so brace for impact.