Saturday, November 04, 2006

Hunting PicScout, the Copyright Crawler Getty Uses

Everyone knows about PicScout used by Getty Images but nobody seems to know anything about PicScout's crawler, no user agent information, no IP's where they crawl from, nothing. When someone asked me if I knew anything about them I did a little research and nothing related could be found ANYWHERE, not even anything initially obvious in my bot blocker log files. Based on my initial observations PicScout actually seemed to be hiding better than all the other corporate crawlers I've researched to date, but maybe we can shed some light on this.

Not that I advocate copyright violation, as a matter of fact, I'm a staunch copyright defender.

However, attempting to crawl under the radar, refusal to honor robots.txt files, or identify your bot in any fashion and bypass website security measures gets under my skin more than anything so I picked up the gauntlet and tried to find signs of PicScout activity.

After the usual simple research methods failed, I decided to start by seeing where they were hosted.

host picscout.com
picscout.com has address 82.80.254.37

host 82.80.254.37
37.254.80.82.in-addr.arpa domain name pointer bzq-80-254-37.dcenter.bezeqint.net.
Ah ha!

I remember a rash of activity I shut down from bezeqint.net a while back so I looked a little deeper into this angle.
inetnum: 82.80.248.0 - 82.80.255.255
netname: BEZEQINT-HOSTING
descr: BEZEQINT-HOSTING
country: IL
Ah yes, they're the guys from Israel that were hammering one of my servers.

I found a high volume of crawling from these IP's that was trapped by the bot blocker automatically and never answered the challenges, so it was definitely bot traffic.
82.80.249.195
82.80.249.196
82.80.249.197
82.80.249.201
82.80.249.202
82.80.249.203
82.80.249.204
82.80.252.130
These IPs have only been spotted using the two following user agents:
Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.0; (R1 1.1); .NET CLR 1.1.4322)
My theory is that this is PicScount attempting to crawl under the radar.

Check your logs people, see if you have any activity in this range, I think it's them.

I would just block this range out of principle at this point as those IPs crawling aren't honoring any internet standards, and if it is PicScout, blocking them could possibly save you a massive chunk of money if some web designer used stolen images building your website.

UPDATE:

After posting this the fine people from PicScout visited the blog and revealed more information about their facilities.

The log showed this visit:
Host Name mail.picscout.com
IP Address 62.0.8.2
Country Israel
ISP Nv-picscout
The information I found from that, including another IP block is here:
inetnum: 62.0.8.0 - 62.0.8.255
netname: NV-PICSCOUT
descr: NV-PICSCOUT
country: IL
admin-c: OG570-RIPE
tech-c: NN105-RIPE
status: ASSIGNED PA
mnt-by: NV-MNT-RIPE
mnt-lower: NV-MNT-RIPE
source: RIPE # Filtered
So, there's a few more IPs you might want to block, but I doubt they're scanning from the office.

UPDATE: Caught Getty keeping an eye on everyone today.

My blog log showed this:
Time: 12th June 200712:24:53 PM
Host Name outbound.gettyimages.com
IP Address 206.28.72.1
Country United States
Region Washington
City Seattle
ISP Getty Images
Referrer: http://www.webproworld.com/graphics-design-discussion-forum/56384-invoiced-getty-images-unlawful-use-images.html

It appears they were snooping on WebProWorld and followed the link here. The user agent claimed to be MSIE 6.0 but it's possibly an automated crawler, hard to say.

Anyway, we're watching you watch us, it works both ways.

Monday, October 30, 2006

Net::Trackback Rocks D-Block

Why is it every time someone puts some code out on the net like Net::Trackback that some asshole will download it and then aim their new creation at my server?

This is where they attempted to hammer my server this morning:

209.9.169.66 [209-9-169-66.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.78 [209-9-169-78.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.67 [209-9-169-67.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.70 [209-9-169-70.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.71 [209-9-169-71.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.69 [209-9-169-69.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.68 [209-9-169-68.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.72 [209-9-169-72.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.73 [209-9-169-73.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.75 [209-9-169-75.sdsl.cais.net.] "Net::Trackback/1.01"
209.9.169.74 [209-9-169-74.sdsl.cais.net.] "Net::Trackback/1.01"
Of course they got nothing but error message for their troubles, but this is still.... BULLSHIT!

Can't even research the source as ARIN.NET's website won't load at this moment and CAIS.NET never responds to WHOIS inquiries and just hangs like this:
[Querying whois.arin.net]
[Redirected to rwhois.cais.net:4321]
[Querying rwhois.cais.net]
Never got a response...

Bunch of BULLSHIT, that's what this is!

Sunday, October 29, 2006

Hand Spammers Waving the White Flag?

Ever since I implemented techniques to automatically moderate hand spammers (aka Indian SEO's) they seem to have noticed they aren't getting through and have gone away. The first couple of weeks it didn't seem like they were slowing down at all, but they were moderated at least so nobody else saw them. Then I made some other changes in how I'm handling spammers that still did it by hand and suddenly they are just gone.

Before the last few changes I easily had about 10 hand spams getting trapped as moderated posts a day, then suddenly nothing moderated has shown up for over a week now.

Did they just give up?

We shall see, but this is promising!

Saturday, October 28, 2006

Ignoring my adoring fans, both of them!

I feel like I've been ignoring all of you lately but it's not true. I've just been so damn busy programming my little ass off, updating massive databases, writing web pages and PowerPoints, ordering custom embroidered polo shirts, so on and so forth, it's just crazy.

Sadly, I feel like the blog has recently become the red headed stepchild that nobody is playing with, not even the dog even if you hung a pork chop around the poor kids neck.

Seriously though, trying to get a massive update completed on an old site and roll a new product out the door while at the same time is crazy stuff. Top if off with getting ready for speaking at PubCon Vegas in November and SES Chicago in December is making me burn the candle at both ends but that hot wax feels OH SO GOOD on my nipple, but that's a different post.

It's like I've become some demoniacally possessed worker bee or some shit and I just can't get enough. I'm spinning out of control and probably heading for a serious burnout but it will be worth it as there are press releases and shit that are going to hit the web on a couple of fronts before the end of the year and I'm stoked.

Hell, I haven't even been going out to see movies, gamble or haunt strip clubs in almost 2 months but I'm sure a week in Vegas at PubCon will solve the gambling and hookers, um strippers issue.

Don't get too upset though, I'm still making it out at least twice a week for a nice 2-3 hour lunch with a friend and a LOT of beer.

Gotta treat myself right a little ;)

Thursday, October 26, 2006

Google, Yahoo and MSN Like Indexing Pure Garbage Sites

The other day I was working on a link checking filter so I could comb many thousands of linked sites and eliminate all sites from my index automatically that no longer contain valuable content.

What I did was make a filter that checked the profile of information on the page looking for signals that detected any sites that have reverted to default registrar pages, default hosting pages, or have become part of domain parks or scraper sites.

After successfully detecting and filtering out many sites that had fallen by the wayside, I started to wonder if the search engines actually indexed all of this crap.

Sure enough, a quick check of Google, Yahoo and MSN confirmed that the search engines eat these shit sites like candy although they can be easily detected and eliminated either by profiling the page content or checking the whois information, or a combination of both.

What purpose does indexing these millions of garbage web sites serve for any search engine?

I mean seriously, the scraper spam sites are one thing, but these are so easily detected there's no ryhme or reason they show up as results to any search being they are 100% crap.

Anyone from one of the major search engines mind dropping a note to explain why hundreds of thousands of cloned garbage sites are being indexed?

We'd really love to hear from you on this topic, please feel free to post a comment :)

Sunday, October 22, 2006

Scrapers Abandoning My Site?

In a rather unusual turn of events it appears my bot blocking efforts have went way better than expected to the point they might be backfiring. It appears that not only have some of the more serious scrapers stopped including my content in their sites, as they're being burned in the search engines (thanks guys) and new appearances of directly stolen content has went down drastically.

However, what I'm noticing is a new trend in sites that used to scrape my server now appear to be just scraping snippets off of other websites which I mentioned recently. Where this became most apparent was when I recently launched a boatload of new pages that had some breadcrumbs cloaked into all the pages not being served to the search engines. After a few weeks after releasing the new pages I went searching for references to these pages in Google, Yahoo, etc. and sure enough found some but they didn't contain my bread crumbs.

Doing a bit of quick investigation showed that the snippets indexed in Yahoo and Google actually came from the search engines themselves. This means that my site is being bypassed completely and the search engines are now the target for what little content the scrapers can get to use from my site.

Remember, I installed NOARCHIVE, NOCACHE and did a bunch of other things to minimize my exposure to the scrapers via the search engines many months ago yet they're still scrambling for the last few scraps that can get.

Just goes to show you how desperate these assholes are for any little scrap of information that ranks high.

Kind of sad that the search engines can't tell they're eating their own dog food though...

Spam Free Accomplishment Zone

Just thought I'd post a follow up after my latest anti-spam measures were put into place that it's been so blissfully quiet that I haven't even bothered to rant about these idiots or much of anything else lately because I've actually been doing more productive work.

Sure, the spammers keep knocking at the doors, banging on the walls, tapping on the windows, but other than falling silently into my spam log just to keep track of what kind of trash was at the door and silently swept away, I see none of it.

The last few weeks were absolutely amazing as I forgot just how much work a person could get done when you aren't constantly trying to clean up after those fucking spammers trying to shit all over every web form they can find.

Sorry spamming assholes, your days are numbered and I'm loving every minute without you.

Thursday, October 19, 2006

Nutch used to advertise Houxou?

I keep seeing this crawler for Houxou:

195.72.131.72 "HouxouCrawler/Nutch-0.8.2-dev (houxou.com's nutch-based crawler which serves special interest on-line communities; http://www.houxou.com/crawler; crawler at houxou dot com)"
When you go to their link http://www.houxou.com/crawler it doesn't say anything about the crawler, it just shows you their homepage. I'm not sure what special interest on-line communities you can possible be serving when you can't even post the page your user agent links claim to be on your website.

Before I gave up altogether, I decided to see what I could come up with in Google and found some interesting results but the site appears to be down.
Nutch: search results
help. Hits 1-9 (out of about 9 total matching pages): WHOIS - 193.203.240.120 ... 20030922 source: RIPE person: Monu Ogbe address: 15 Penman Close, ...
nutch6.houxou.com:8080/search.jsp?query=ogbe&hitsPerPage=10 - 10k - Supplemental Result - Cached - Similar pages

Nutch: 搜索帮助 - [ Translate this page ]
搜索英文单词不区分大小写, 因此搜索NuTcH 等同于搜索nUtCh. ... 评分详解)显示Nutch如何给该网页打分. (anchors)显示指向该网页而被Nutch索引的anchor文本. ...
nutch6.houxou.com:8080/zh/help.html - 7k - Supplemental Result - Cached - Similar pages

So what's the deal?

Why is Houxou crawling with a link to a missing page about bots?

Is this just a ploy to get webmasters trying to figure out what the Houxou crawler is to look at their hosting services?

Who knows, guess we'll just have to wait and see but it smells fishy to me.

Smiley Face User Agent

This should be filed under "What the fuck is wrong with people".

Here's the user agent with a hyperlinked smiley face to some bullshit website in The Netherlands:

80.126.0.125 - "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; <A HREF=http://jult.net>;-)</A>; .NET CLR 1.1.4322; InfoPath.1)"
Looks like some asshole might've done this to his browser as the request did have a Google referrer so it's probably a real human that landed on my site.

Well pal, you got an error message when you hit my site didn't you?

Bet you're not so fucking smiley faced now.

Stupid shit.

Sunday, October 15, 2006

Netsweeper Caught Using Multiple Brooms

In the badly behaving corporate bots dept. we offer Netsweeper as our newest entry from Canada. They run one of those content filtering companies that thinks they should be allowed to crawl your site no matter what just to protect their clients.

Sorry, but we happen to disagree with all these content filtering spiders that feel the need to crawl without any regard for robots.txt and we really don't need a whole buttload of content filtering companies scanning the fucking web.

Yes, I threw in the word fucking just so your asshole spider will flag this post as bad content so none of your goddamn customers can read this so blow that out your ass.

Let's see what Netsweeper runs:

66.207.120.226 "webcollage/1.127"

66.207.120.226 "NutchCVS/0.7.2 (Nutch; http://lucene.apache.org/nutch/bot.html; nutch-agent@lucene.apache.org)"

66.207.120.227 Mozilla/5.0 (X11; U; Linux i686; en-US; rv:1.7.5) Gecko/20041107 Firefox/1.0
These IP addresses have the following host names:
66.207.120.226 -> firewall.net-sweeper.com

66.207.120.227 -> host227.net-sweeper.com
Let's just cut thru the chase and here's the information to block their ass:
CustName: Netsweeper
Address: 4-512 Woolwich Street
City: Guelph
StateProv: ON
PostalCode: N1H-3X7
Country: CA
RegDate: 2003-04-08
Updated: 2003-04-08

NetRange: 66.207.120.224 - 66.207.120.239
Ta ta Netsweeper, you've been blocked and swept under my rug.

Cya!

Saturday, October 14, 2006

No Referrer and Visitors vs Spambots

I've been struggling recently on better ways to handle requests to form pages, like a comments page, that are more secure yet more visitor friendly than just slapping up a "403 Forbidden" which runs off humans as well as bots.

After contemplating the issue and taking bookmarks and disabled referrers into account, I decided to simply redirect these potential bad hits to my home page instead of the old 403 error. This way a valid visitor that bookmarked the page could just navigate back, referrer intact and post as usual. So far I've seen a few humans that were redirected off the page for whatever reason navigate back to where they wanted to go, so it doesn't appear to be stopping determined people that aren't just there for malicious purposes.

Additionally, I only do this redirect after verifying the request isn't coming from the search engines as I obviously don't want to confuse the SEs by redirecting them to the home page.

The fun part is it seems to be confusing the shit out of the spambots, they're bouncing all over the place, quite hysterical to see them freak out.

For even more fun, verify that the referrer to your form page isn't a direct hit from a search engine to that page as many of the hand spammers from places like the Ukraine and India seem to like to use Google to find pages to post their clients' listings. For those reasons, now I'm also redirecting any direct hits to the posting page that comes from Google, Yahoo, MSN and ASK. Redirecting the hand spammers (aka SEOs) back to the home page seems to stop the hand spams as well since I've just made the job a little harder they just seem to move along instead of spending more time to find the page

One last trick, if you have the means to track these things like I do, is reject or auto-moderate anything with just a single page view to your site, which is just that form page. Some figured out I was looking for a specific referrer and plugged it in so the post looks 100% legit. Well bummer dudes, you need to view more than one page to make a submission so cleaning up the data to add a valid referrer was a nice try but you still have a bad bot profile by having no previous page views.

Gotta love all the fun and games with both sides escalating but so far I'm still spam free and winning this war.

Friday, October 13, 2006

Bot Busting or Spam Hunting?

Since I made a few posts about spam lately, mostly web spam, it seems I have a few spammers all bent and other people thinking that I'm a spam hunter.

For starters, I'm a bot buster and not a spam hunter. Whether it's a scraper, a stealth crawler, or a spambot, they're all bots so anything that malicious bots do to my sites that need to be stopped will be of interest to me. Defeating a bot that scrapes or spams each has it's own challenges and it's not been too terribly hard to stop them either way, so far.

Spam Hunter?

Get real, who has to hunt spam?

I just sit here minding my own business and spam comes from every angle via email, online submission forms, blog comments, or anything else you can put online and they'll spam it. When I see a trend emerging in my own spam logs, like the recent wiki/tiki and phpBB abuse, I just look to see how deep the problem is which isn't hunting spam exacly, it's analyzing the trends and patterns of the software vulnerabilities that are being exploited.

Maybe someone will notice and close some of those exploits, wouldn't that be nice, unless you use those exploits...

Just because a few of the latest topics revolve around spam, it's still bots in action and bots I've automatically blocked and logged, just spambots is all, but still bots all the same.

Enough of this noise, back to busting bots, scraper, spam or otherwise.

Wednesday, October 11, 2006

Web spammers abuse GuestCity's hospitality

While researching the depth of the wiki/tiki spam abuse problem there was one particular redirect link that caught my eye to some site called GuestCity.



When you see the smoke from web spammers in a search engine there's usually a web spammed fire somewhere close so I decided to look deeper to see what GuestCity was all about and assess the damage.

There on the home page was an encouraging anti-spam symbol on the lower left of the screen!



I clicked on the anti-spam symbol and read their get tough policy on spam, cool.




If these guys are really tough on spam, shouldn't find much web spam over there, right?

Sorry, took about 2 seconds to spot sites overflowing with crap like phentermine, viagra, cialis, and on and on.

If you want to see the funniest shit ever, click on their DEMO link right off the home page that is spammed upside down and inside out with a couple of years worth of garbage.

The most priceless quote is this one from 2004:
4249. Old demo book was removed due to lot of spam messages. Welcome to new one!
2004-10-21 08:33:59, Webmaster,
I think I fell of my chair laughing hysterically about then.

Sure wouldn't take more than a few days to write some code that would stop the spammers and eradicate all the splogs over there, hope they're up for the challenge.

Zone Communications sends SEO Spam

Just when I thought some of the SEO's were getting smarter, since I haven't been spammed by one for a while, here comes a nice juicy one from our friends at Zone Communications in southern California.

The spam came from this IP:

71.128.4.233
ppp-71-128-4-233.dsl.irvnca.pacbell.net.
Here's the lovely spam:
Your website can be at the top of the first page on all major search engines. Zone Communications has a great service that is very low cost and is billed month to month with a full refund if you are not satisfied. With this service you�re your ranking on Google, Yahoo and MSN within certain cities will be at the top of the page.

The pricing is simple. We charge $59 per month for your first listing and $40.00 per month for each additional listing. (There is also a small one time set up fee.)

If you are not satisfied with your position within 30 days, we will give you a 100% refund.

[spammers name]
Zone Communications
800-xxx-xxx
714-xxx-xxx (fax)
[spammers name]@zonecominc.net
OK, if they even took a look at the site they spammed they would know I'm all over the top of the 4 top search engines for keywords, cities and just about everything in my niche short of Mom's Apple Pie and Kitchen Sinks.

For people that spam me, my pricing is simple. We out spammers for FREE, and there is no per month fee to continue being outted. (There is also a small one time fee called you've been BUSTED for sending me spam.).

Now stay off my damn websites, you weren't invited in the first place.

Saturday, October 07, 2006

Vulnerable Tikis Ruthlessly Spammed and Google Indexed

The other day I posted about how VT.EDU's tiki was overflowing with spam so today I went thru my spam filter log to just see how many attempted spams there were last week using tiki redirect pages.

Here's a short list of the most recent attempted spams linking to tikis that hit my server:

http://www.lug-viersen.de/tiki-directory_redirect.php?siteId=136#viagra
http://ipvs.informatik.uni-stuttgart.de/BV/swarmrobot/tikiwiki-1.9.2/tiki-directory_redirect.php?siteId=474#viagra
http://i60p4.ira.uka.de/tiki/tiki-directory_redirect.php?siteId=24#viagra
http://www.xsl-rp.de/tiki-directory_redirect.php?siteId=1018#cialis
http://www.neurotransmitter.net/wiki/tiki-directory_redirect.php?siteId=243#viagra
http://research.cs.vt.edu/advance/tiki/tiki-directory_redirect.php?siteId=3284#viagra
http://meverhagen.nl/tikiwiki/tiki-directory_redirect.php?siteId=19#viagra
http://www.namurantifasciste.be/tiki-directory_redirect.php?siteId=996#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=10#viagra
http://www.xsl-rp.de/tiki-directory_redirect.php?siteId=1015#viagra
http://www.railfuture.org.uk/tiki-directory_redirect.php?siteId=61#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=9#viagra
http://herenaforge.org/tiki-directory_redirect.php?siteId=38#phentermine
http://herenaforge.org/tiki-directory_redirect.php?siteId=51#viagra
http://www.derrychineseschool.org/DCS/tiki-directory_redirect.php?siteId=7#viagra
http://openg.org/tiki/tiki-directory_redirect.php?siteId=54#viagra
http://www.prospace.org/tiki-directory_redirect.php?siteId=2385#viagra
http://www.milwaukeelug.org/tiki/tiki-directory_redirect.php?siteId=1349#viagra
http://www.ee.aston.ac.uk/intranet/tiki-directory_redirect.php?siteId=18#viagra
http://dev.librehwdb.tuxfamily.org/tiki-directory_redirect.php?siteId=18#viagra
What's distressing is that Google and the other SE's really love these spammed pages too, just gobble them up, and it's probably unwittingly passing PR from all these spammed tiki sites on such terms as viagra, cialis, levitra and a whole lot more.

So Google gives spammers a 2-for-1 special by giving them SEO value for their spamming activities, it's just a crying shame, it really is.

What's pathetic is this problem could be stopped on both sides of the coin. The tiki/wiki software developers could get off their lazy asses and implement some tools to allow webmasters to stop this rampant spamming of their software, it's easily doable. Additionally, the search engines like Google can easily identify and stop indexing spammed web pages to eliminate the value they give to the spammer.

Remember, I'm reporting about ATTEMPTED spams, all those links and a shitload more were automatically dumped, it's not rocket science, it's barely programming above a rudimentary level to identify and filter that shit out.

Why does this continue when the solutions are so simple for all involved?

Amazing that it's allowed to continue, simply amazing.

Thursday, October 05, 2006

Podomatic Vulnerability Enables Spammer Redirects

Here's another instance in a rash of reported vulnerabilities in member registration pages being spammed. Never heard of Podomatic before but it appears the spammers sure have and some nitwit registered as a member called Valium to do his spamming.

The link to the member's site is:

http://www.podomatic.com/profile/member/valium
The javascript redirect code appears to be this shit embedded in the memberpage:
<script>
var mbht872 = 'on=';
var bikmr354 = 'qiqyi199';
var zlh171 ='ment';
var k97='.lo';
var ydxglyjedai737='ti';
var bmmp211='docu';
var mzcra833='http://drsearch.net/search.php?aff=15313&q=';
var ertmj632='valium';
var qiqyi199 = 'ca';
var lflx482='"';
if(bikmr354 = 'qiqyi199')eval(bmmp211+zlh171+k97+qiqyi199+ydxglyjedai737+mbht872+lflx482+mzcra833+ertmj632+lflx482);
</script>
Just goes to show you that if you don't secure your sites some spammer will abuse it but people just don't listen.

Wednesday, October 04, 2006

Automatic Detection of Spam Hand Jobs

Sometimes certain anti-spam ideas just hit you upside the head when you least expect them and seem so obvious you wonder what took you so long to figure it out.

I've already blogged about the fact that I've stopped all automated spam dead in it's tracks on my sites, but people manually posting can of course correct all of the errors detected and continue to make an unwanted garbage post.

I have an extensive junk detection filter that rejects anything with the usual suspects like viagra, cialis, gambling, poker, etc. which stops the nastiest of these posts. However, some little pain in the ass SEO aka spammer might slip thru with a hand job posting about his store in India selling magic beetle dung or something that you would never imagine putting in your junk filter in the first place.

A few days ago I decided to review the last 30 days of legitimate submissions and compare them to the few off topic hand jobs that slipped through the cracks and see if I could come up with anything that would allow me to stop the hand jobs of absolutely random and crazy things outside the realm of the typical common auto-spam posts.

Then, like a lightning bolt it suddently hit me, that with these random off topic hand spams it's not what's IN the posts it's what's NOT in the posts that makes them easily identifiable. The concept is to scan for a list of words that SHOULD be in the post, like quotes from anything in the thread or certain keywords related to the topic and automatically set everything to MODERATE that doesn't fit the usual posting patterns.

Basically it's a 'lack of content filtering' technique and off topic posts, like spam, stand out like a sore thumb.

Using this blog as as example for a topic, you would expect most comments to contain words like bot, spam, IP, host, crawl, firewall, htaccess, apache, etc. or a set of keywords derived from the original post title and text. The absence of any of these words is a clue that the post just might be SPAM or otherwise off topic and should be placed on moderation for the admin to review.

Since I've started using this new 'lack of content filtering' technique it's snared the few hand submissions to my other site that were completely off topic, those that I would've deleted immediately. The beauty is I can continue to leave the posting wide open for humans, not moderate everything, with only those posts that don't match the topic getting instantly set to moderate.

I expect a few false positives but so far 'lack of content filtering' is doing exactly what I expected it do and set a couple of crap submissions last night for shit like "zanaflex information", apparently some pill I've never heard of and "News, Stores, People, Careers at Finditt", some wannabe search engine, to moderate automatically while letting 20 on topic things thru without a hitch.

Another automated weapon in the war on spam!

GoodBidWords.com Scrapes LookSmart

Noticed at hit from one of my scraper probes in GoodBidWords.com which contained the IP address of the original crawler.

Looked up the IP address and guess where it came from:

"Mozilla/4.0 compatible ZyBorg/1.0 (wn-14.zyborg@looksmart.net; http://www.WISEnutbot.com)"
Isn't this precious that GoodBidWords got caught because of all the places to scrape they decided to scrape a search engine that I don't permit to crawl my site!

What a hoot, second-hand scraper busting, this rocks!

Tuesday, October 03, 2006

phpBB Membership Spamming for Authority

We first reported about phpBB spamming the other day when we stumbled upon this "DISY registration spamming script" and since then have had a little time to examine what spammers are doing with phpBB trying to gain authority.

Let's just check a few of these spammers in Google:

pimpdomain.net
thewestgategazette.com
ritalin-pharmacy.com
Hell, just try any of the domains listed in my Technorati Loves Spam post and search for the domain name and phpBB and see what shows up.

Just amazing what these assholes do with this shit cluttering up the net with spam.

Technorati Loves Tasty Cloaked Blog Spam

I've noticed that Technorati has been happily eating up scraped and cloaked blog spam for ringtone sites, among other things, like it's fucking candy.

Let's use a search on my blog name as an example:



Click on those links and it's always to the same spammy page name like these hosted on theplanet.com of course:
http://artinexis.net/#comment-341
http://themetrogiant.com/#comment-329
http://pimpdomain.net/#comment-341
One server is 70.87.88.121 or 79.58.5746.static.theplanet.com with these domains all spewing ringtone ads:
about-levitra.net
acvfa.net
artinexis.net
cariculture.net
catsfive.net
citadel1.net
cloudsite.net
eightonefive.net
rennenmotorsports.net
t3linkcom.net
Another annoying server is 70.87.88.108 better known as 6c.58.5746.static.theplanet.com which has these goddamn domains:
talonpro.com
tempuspercussion.com
terminal34.com
the-god-poll.com
theincrediblesuckingspongies.com
themetrogiant.com
thepulse2000.com
thespinet.org
thewestgategazette.com
thoweu.com
tlc-express.com
Or this fucking spam filled server host 70.87.88.106 hosted by our fucking friends 6a.58.5746.static.theplanet.com:
perseidslive.com
pimpdomain.net
poemnet.net
posses1consent.com
projhind.com
ptcsucks.com
r1g4t2you.com
rbigkitty.com
rep1icas.com
ricohtour.com
rising7.com
ritalin-pharmacy.com

Here's the same shit about ringtones they all show:


Who are the fucking idiots buying all these goddamn ringtones anyway?

How about you just set the phone on buzz, stick it in your pocket, and you'll never miss a call or be confused it's someone else's phone ringing, and best of all you can do it without lining the pockets of the cell companies or perpetuating this spam. Better yet, just shove that phone up your ass as most people that feel the need to never miss a call by using goddamn custom ringtones are probably talking out of their ass anyway. While you're at it, shove some custom phone face plates and a nice blue tooth headset up your ass too, but I digress.

WAIT A FUCKING MINUTE...

I think I see a pattern here 70.87.88.106, 70.87.88.108, 70.87.88.121...

Let's try 70.87.88.120 and see what we find:
asmort.net
bevirusproof.net
conlajusticiaysociedad.net
fabionne.net
friendshipmotorinn.net
macoszone.net
palick.net
phila-ibiz.net
themikecam.net
wesmn.net
More ringtone spam spam spam....

Or let's try 70.87.88.115:
audio-wire.net
buy-cheap-2u.com
chabadofbuffalo.com
cheap-online-buy-free.com
el-condor-pasa.net
ellemtel.net
fairy-wings.net
free-top-sex.net
gotobiz.net
healthcybeline.com
javabooks.net
jemison-nealon.net
lolitasexlinks.net
macromediaseminars.com
netnetn.net
remax-powell-m-corpus-christi.com
stopsundiata.com
wangmatongli.com
xinyifang.net

Yes!

More spam spam spam spam spam!

OK, this is obviously a big operation with lot's of shit domains serving up spam on lots of IP's, I'm bored with this already, if you want to help fill in more blanks with this ringtone spammer go to Domain Tools Reverse-IP page and type in the IP address in that range and see what's on the servers.

Maybe they should change their name to Spamorati as they seem to love these fake blogs reposting old posts.

BTW, if you need help with automating the identification of spam over at Technorati just drop me a line as I'd be more than happy to show you how to automate the process for a small fee!

The things I could teach them on ways to clean up their listings and improve their service would boggle their minds.

Saturday, September 30, 2006

ShoeMoney's Blog Spam Stopping Primer

The day after my battle cry to Rally the Anti-Spammers here comes ShoeMoney with some great suggestions for stopping blog spam. Everything ShoeMoney posted is very solid advice but some spammers have already been evolving past some of those patches which is why I use my draconian anti-spam methods. Basically, ShoeMoney's advice will stop the majority of your garden variety spammers, but not all as they are constantly adapting, so as you improve your defenses they improve their ability to bypass those defenses.

Remember, security is built in layers and the more layers you pile on, the more the spammers will chip away at your security so building the better spamtrap just results in smarter spammers and they're already here which I'll address with examples below.

Let's examine ShoeMoney's anti-spam advice, see what some state of the art spammers are already doing, and add a few more tricks here and there for even better security.

Starting with the first item he listed:
5) Deny Access to No Referrer Requests

The approach does work on most spammers but I had about 10 requests today where it would've failed. Not that you shouldn't implement this, it's a good trick to stop a lot of spam, just be aware it won't stop everything.

Example:

My bounced spam log shows the following:

IP: 84.110.248.226
User Agent: "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Subject: "Viagra"
URL: http://anol.webhosting.gs/viagrageneric.html#viagra
Take a look at what's in my server log:
84.110.248.226 - "POST /formsubmit.html HTTP/1.0" 200 11918 "http://www.mysite.com/formsubmit.html" "Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Yup, that's right, a referrer, and I had about 10 of those and they were all from spambots.

Stopping the poorly coded spambots is easy, but they won't be vulnerable for long as the patch to add the domain name being spammed into the referrer is trivial so I expect this anti-spam advantage to be short-lived but I use it too, you should still do this.

Now, let's tackle the next item, which is VERY good advice:
4) Kill tor anonymous proxies

I block many proxies on my servers, which does stop a lot of spam, but don't think that all spammers use known proxies. This is the reason I also block dedicated server hosting facilities because a series of $2 webhosting accounts can be used to effectively spam and bypass the proxy lists.

Example of 4 sample spams (out of many) today that all had referrers mentioned above and came from some ISP/Host called bezeqint.net:
09/29/2006 84.110.248.226
"Viagra" http://anol.webhosting.gs/viagrageneric.html#viagra

09/29/2006 84.110.244.240
"Viagra" http://gerda.forospace.com/#viagra

09/29/2006 84.110.243.107
"Cialis" http://borea.forospace.com/#cialis

09/29/2006 84.110.241.163
"Cialis" http://kaizer.webhosting.gs/cialisbuy.html#cialis
Use this with caution:
2) Blacklist Repeat Offenders:

First off, blacklist on the FIRST offense so there is no second time. However, you really need to know what you're doing and lookup who the IP address belongs to so you aren't blocking IP addresses from places like the AOL IP pool (reused every 15 minutes or so) or any other shared proxy dial-up IP pools as those IP assignments are very temporary and the next access is probably a different visitor, not a spammer, so be very careful with this.

This is a gem and we can make it better:
1) Rename your comment file

Excellent advice as I've done that on some websites but don't be shocked when it's short-lived as spammers also have crawlers looking for these comment pages and the fact that you're still linking it under the keyword "comments" is a dead giveaway.

If you're going to change the file name, also change the word that links to the file name to "discussion", "verbal intercourse", or "rants", anything but "comments" to throw them off.

Additionally, move the actual FORM into obfuscated javascript document writes. How this works is the spambot scanning your website can't even find the webform to submit comments as most bots don't use javascript, so only an actual visitor would see an actual webform written into the web page via javascript.

Don't forget the CAPTCHA!

Now, the one thing ShoeMoney didn't mention which works wonders is a simple CAPTCHA and it's keeping a few of my sites spam free without ANY other work involved. Yes, there are ways to bypass a captcha but it's not easy for the spammer. So far most captcha protected sites are safe with such simple protection, but I expect that situation to escalate soon.

Kudos to ShoeMoney for spreading the word, we need more anti-spam information spreading and more people jumping on the anti-spam bandwagon so we can rid the 'net of this scourge as soon as possible and move on to more productive activity.

Thursday, September 28, 2006

Virginia Tech's Computer Science: Wiki Spam 101

My website stops spam posts cold, and logs them, so that eventually I can glance over the list of bounced spams now and then just to see what was caught and this one was priceless:

09/28/2006
200.88.223.98
"Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Subject: "Viagra"
URL: http://research.cs.vt.edu/advance/tiki/
tiki-directory_redirect.php?siteId=3284#viagra

I looked and thought, "Viagra spam linking to VT.EDU? Could their server be hacked like SpamHuntress is posting about?" So I click the link and of course it uses VT.EDU's server to redirect me to some viagra sales site just like the URL would make you think it would, no surpise there.

So I trimmed the URL to see what in the heck this site was and it's ADVANCE, FOR THE ADANCEMENT OF WOMEN IN ACADEMIC SCIENCE AND ENGINEERING CAREERS and it's full of advances for MEN such as viagra, cialis and levitra spam plus a whole bunch more.

Well, the IT dept. and professors in charge of the VT computer science program should probably start quaking in their boots as I would be VERY UNHAPPY if I was the Dean.

This is completely unacceptable when the IT guys and CS Profs aren't using even rudimentary anti-spam technology like, oh, maybe a simple CAPTCHA to stop this shit.

I want my tuition refunded.

BTW, whoever these spammers are, they've been VERY BUSY little beavers.

Time to Rally the Anti-Spammers

After the demise of Blue Security and this recent meaningless default judgement against SpamHaus, the spammers are getting braver and bolder by the day. Now, one of the most vocal anti-spammers around, SpamHuntress, has recently come under attack after exposing a few people that really didn't want to be exposed.

Even one self-professed blackhat SEO web spammer has the audacity to tell SpamHuntress to "get a life" because she must be cutting into his livelihood. Maybe I'm just too lazy, but who would've ever thought of registering for a bunch of forums and never posting as an SEO tactic? Using his DISY registration spamming script probably sped it up and he's busy making friends [scroll to bottom] as well.

OK, so now the phpBB people will need to be alerted to add NOFOLLOW to all those links in the registration page to stop this SEO vulnerability, but I digress, will rant about that later.

Unlike email spam, which is a real pain in the ass to stop, there is absolutely no reason we have blog, forum or guestbook spam whatsoever except for shitty programmers writing the stuff and people using it that either:

  • have abandoned their websites or forgotten that old guestbook or blog now littered with junk
  • aren't aware there is a problem as many spambots post on older threads
  • don't know there are solutions to these problems
  • aren't capable of installing the patches even if they are aware of the solutions
I've posted before how I stop spam on all my pages that have forms for submitting a variety of things, WITHOUT the use of captcha's, and although it's a pretty draconian approach to the problem it's also highly effective. My solution was to simply reject any posts with embedded HTML and URLs, just bounce them with an error about the content, and it works 100% against real spam. Maybe it's a tad extreme but when this type of spam is dead maybe I'll open up my sites again to more robust content posts, you never know.

However, for those that like to continue to do things the hard way, here's a list of software you can install to stop the spammers:
I'm going to ask that people reading here help the cause and start educating everyone you run across with a blog or forum being overrun by spam.

Please point them to a resource to solve the problem or offer to help them add the plug-ins pro-bono or for a nominal fee if they don't understand how, or if all else fails alert the host to help sites overflowing with spam and see if they'll be of any assistance.

Don't forget, the purpose of these spammers is to drive direct traffic and also get results in Google so when you stumble upon these sites in Google, make sure you file a Google Spam Report while you're there to get them whacked from the search results.

We can stop this in the next year or two, as long as people quit being complacent and just install the upgrades, patches, captchas and other anti-spam tools.

Spread the word, let's just get this done so we can stop talking about it already!

Wednesday, September 27, 2006

MySpace: Porn Networking Spam Machine?

The other day I signed up for MySpace while researching the members with "Click The Ads" on their pages encouraging others to commit click fraud to fund their various lame causes.

Unfortunately, signing up for MySpace immediately resulted in a couple of porn spams sent to my Inbox which really pissed me off.

So I get some shit that looks like this:

FROM: MySpace Events
SUBJECT: .. has invited you to: I seen you online

Hi ,

.. has invited you to an event on MySpace:

Click the link below to view the event details:
http://events.myspace.com/index.cfm?fuseaction=PORNSPAM

Now below this, there is some bullshit message from MySpace:
At MySpace we care about your privacy. We have sent you this notification to facilitate your use as a member of the MySpace.com service. If you don't want to receive emails like this to your external email account in the future, change your Account Settings to "Do not send me notification emails."
Really, you care so much about my privacy you let goddamn porn spammers send me fucking email?

I'm touched, a tear comes to my eye ...

... yes a tear, because I realize can't reach out and smack whoever let this shit happen upside the head!

Anyway, here's the website on MySpace linked from the spam:



Here's the first site's link to a girl with a webcam:




And here's the second spam site's girl with a webcam:





I'm wondering if people under 18 get these spams too?

I think I'll just cancel the account because MySpace is no place I was to be associated with.

Monday, September 25, 2006

MySpace: A Click Fraud Social Network?

Maybe it's just Web 2.0, or Web Welfare 2.0, but it appears that stealing from advertisers is now something that is accepted in social networks. Let's look at what we find on sites like MySpace and others which are a good place to build up a nice list of friends to click your ads, especially the Google ads, because we all know that friends click friends ads, especially if you want your friends banned from AdSense.

Even on YouTube where people can't put up their own ads they beg people to come to their website and click the ads to support them putting up more videos!



The most shocking is Blogger, which is owned by Google, the creator of AdWords and AdSense, which hosts sites that encourage people to "Click the Ads" to defraud the very advertisers they rely on for their massive income.

How difficult would it be to have a single employee out of the entire Googleplex devoted just to keeping click fraud off their own property?

You know the answer, I know the answer, yet a simple search reveals that it's not being done, or not done adequately at any rate or there would be no sites returning results from Blogger on this topic if they were on top of the problem.

The technology for these sites to deploy an automated process to locate pages within their sites that contain calls to "click the ads" or "click Google ads" or any combination and eliminate this fraud on a daily basis is so trivial and rudimentary that beginning programmers could do it.

Bottom line is there's absolutely no excuse for this type of call for advertiser click fraud to be allowed unchecked on these sites, not in MySpace, YouTube, Blogger, Google, Yahoo, MSN or anywhere else and why Click Fraud 2.0 continues to perpetuate on the web when it's so easy to thwart frankly boggles the mind.

Flickr Member Requests Click Frauding Advertisers for the Children

Well, I've seen all sorts of excuses to advocate click fraud but the plea on Flickr to commit a crime for the children is a new one and more despicable than any I've seen before. Think about the precedent that this sets in impressionable young minds that it's "OK TO STEAL FOR A CAUSE" when crime is never OK. Sadly, all of the good this person has possibly done for these children was wiped away with one call to arms to defraud people for a cause.

If you want to save the children, set up a Paypal account and teach the children than they can be helped by the generosity of others, not by others commiting FRAUD!

Here's the screen shot from Flicker:



And the site it lands on in Blogger:



Come on buddy, just ask for donations and keep it legal as we all love the children but this is over the top.

Saturday, September 23, 2006

Search Engine Spammers Extraordinaire

OK, these idiots made the classic mistake of scraping one of MY pages so they're about to get outted in a massive way. Unfortunately, in this case I didn't get an IP address and my content was already missing from their site thanks to the slow crawl and index of MSN, but a little research proved this was a HUGE operation of mind blowing proportions.

I got bored checking all the domains as some are hosted in the same place, some aren't, too many to look at but it's all spam. Perhaps the same person, or perhaps a bunch of idiots running some automatic website generating tools.

The sites tend to come in 3 flavors, AdSense monetized articles, AdSense monetized scrapers sites (scroll WAY down) and AdSense + Shareasale sites.

Just search for the phrase "When we had a difficult think about this project" in Google, Yahoo or MSN and you'll see a shitload of pages from these search engine spammers.

Also, try a search for the phrase "Foraging for the best file on" in Google, Yahoo and MSN and see more shitloads of pages.

You can see all sorts of key phrases these sites repeat and bust more and more of them like this "Everyones path is incomparable and everyone" one on Google or Yahoo.

And even more shit like "If you've worked with a portal" on Yahoo.

Someone noticed their terms were hijacked in these bullshit pages and blogged about their suspicion on what's going on.

Seriously though, I bet I could write a script to identify and locate all the bullshit spammers using this data with all their common phrases as it's so easy to spot once you have a data sample like these to analyze.

Spam, spam, fucking spam, and not so smart fucking spammers.

Whitelist OPT-IN htaccess file

People are always asking me how to build an OPT-IN .htaccess file, which I advocate, opposed to the traditional blacklist methods.

The problem with OPT-IN is it's VERY unforgiving and you really need to check your visitor stats and make sure you're letting in all the crawlers that are sending you traffic.

Belows is a bare bones sample of how it works and anything not in the list gets a 403 Forbidden error so you'll probably need to add more items and refine this for your particular website.

Sample .htaccess file for Apache 2.0:

#allow just search engines we like, we're OPT-IN only

#a catch-all for Google
BrowserMatchNoCase Googlebot good_pass
BrowserMatchNoCase Mediapartners-Google good_pass

#a couple for Yahoo
BrowserMatchNoCase Slurp good_pass
BrowserMatchNoCase Yahoo-MMCrawler good_pass

#looks like all MSN starts with MSN or Sand
BrowserMatchNoCase ^msnbot good_pass
BrowserMatchNoCase SandCrawler good_pass

#don't forget ASK/Teoma
BrowserMatchNoCase Teoma good_pass
BrowserMatchNoCase Jeeves good_pass

#allow Firefox, MSIE, Opera etc., will punt Lynx, cell phones and PDAs, don't care
BrowserMatchNoCase ^Mozilla good_pass
BrowserMatchNoCase ^Opera good_pass

#Let just the good guys in, punt everyone else to the curb
#which includes blank user agents as well

<Limit GET POST PUT HEAD>
order deny,allow
deny from all
allow from env=good_pass
</Limit>

Just save the above as a file named ".htaccess" in your httpdocs or root web folder in your hosting account and all the crazy bots abusing your site will get bounced from now on.

Remember, anything not listed will no longer have access so be careful and make sure everything your site needs allowed is in the list.

Enjoy.

Googlebot Validation

Google has finally completed a DNS project that will allow us to use a simple reverse and forward DNS check to verify it's really, truly, honestly Googlebot and not a cheap immitation, or Google crawling thru a proxy, or anything else you can imagine.

I'm so sick of explaining why you might need this and what it solves you'll just have to follow a few links and read the threads at these various places.

Here's the official How To Verify Googlebot post on Google's blog.

Then you can check out what's been said about How To Verify Googlebot on Matt's blog.

Then a couple of threads on WMW about Verifying Googlebot that should answer any other questions on this topic.

Thanks again to Matt for getting this project finished!

Tuesday, September 19, 2006

How Important Are Plurals

Many people ignore plurals when they optimize their website and miss a lot of opportunity for additional search engine traffic.

Here's a few trend examples:

Take a look at plumbing, plumber and plumbers and you'll note that the plural is just as often the search term as the singular plumber.

How about teaching, teacher and teachers where all 3 run very close and teachers appears to dominate the search trend by a thin margin.

Last but not least, something closer to home with blogger and bloggers, where blogger clearly stands out as the dominate term but bloggers is statistically significant enough to merit ranking for the plural.

So don't forget to rank for your keyword plurals or someone else will rank there instead of you and they will KICK YOUR S!

Request from India

Just when I thought it was going to be a boring day I got a link-exchange spam from one of those wonderful Indian SEO's that wouldn't know how to promote a website to save his own life.

I'm actually shocked this email didn't include the usual threat that "you have 24 hours to confirm a reciprocal link before we remove yours from our site".

Boy, doesn't this shit look familiar:

Dear Webmaster
Greetings from India

Happened to visit your Webpages : [FILL IN BLANK OF SPAM RECIPIENT HERE] & liked it very much.
Would like to request you to have a look on our site :: [FILL IN BLANK OF SITE BEING SPAMMED HERE]

Hope you'll like this site. We are trying our best to spam the shit out of everyone in the name of India, You can help us by just adding our link on your wonderful website. And these exchanging link with good quality websites is beneficial for both the site to get a good ranking in search engines & that will help both of us in driving Traffic.

So We request you to add our link at your Website

Here is the Link Information of our Sites ::

URL : [LINK TO OFF TOPIC SHIT GOES HERE]
Link Text : We Spamma U Ass
Desc. : That's Right, This is Spam, its no more a dream!!

Just do let us know if this acceptable for you.
Hope to have quick & positive response.
Thanks in Advance

Best Regards
Sendjay Sumspam
Spamming-Our-Ass-Off.com

BTW, if you're the Indian fuckhead sending this shit, FUCK NO I WON'T LINK TO YOUR SITE you goddamn moron.

Just a lovely way to start the day.

Say it with spam.

Thursday, September 07, 2006

Counting Scrapers on your Abacus

Had a couple of persistent little fuckers hosting with Abacus that just keep trying and trying to download a boatload of pages that I've been monitoring for months now.

The specific IPs of these boxes are:

206.225.82.155 "Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)"

206.225.91.164 "Mozilla/4.0 (compatible; MSIE 5.0; Windows NT; DigExt)"

206.225.83.179 "Evaal/0.7.2 (Evaal search engine; http://evaal.coml; bot@evaal.com)"

216.55.161.38 "Java/1.4.1_04"

216.55.142.118 "Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)"

216.55.162.3 "PEAR HTTP_Request class ( http://pear.php.net/ )"

216.55.147.80 "sna-0.0.1 mikeelliott@hotmail.com"
Toss in a couple of proxies:
206.225.85.127
206.225.86.86
And some other miscellaneous bullshit not worth mentioning.

Here's what to block:
OrgName: Abacus America Inc.
OrgID: ABAC
NetRange: 206.225.80.0 - 206.225.95.255

OrgName: Abacus America Inc.
OrgID: ABAC
NetRange: 216.55.128.0 - 216.55.191.255
Now you've been COMPLETELY BLOCKED so count THAT on your Abacus!

More Evolving Scrapers

Like I've been reporting, they're all going stealth.

I keep seeing user agent change from this:

62.163.33.234 "Java/1.4.1_04"
To this:
62.163.33.234 "Mozilla/4.0 (compatible; MSIE 6.0; Windows 98)"
Soon the usual blocking methods won't work whatsoever.

Wake up and smell the COPY before it's too late!

Block the Bots Tonight

Time for a little lunacy break for people feeling blue battling the bad bots.

Sing along boys and girls...

Sung to the tune of "Rock Around the Clock"
with apologies to Bill Haley and the Comets.

One, two, three bots, four bots, blocked.
Five, six, seven bots, eight bots, blocked,
Nine, ten, eleven bots, twelve bots, blocked,
We're gonna block all the bots tonight.

Put your firewall on and lock em out,
We'll have some fun when they scream and shout,
We're gonna block all the bots tonight,
We're gonna block, block, block, their scraping blight.
We're gonna block, gonna block, all the bots tonight.

When the block strikes two, three and four,
If the scrapers slow down we'll yell for more,
We're gonna block all the bots tonight,
We're gonna block, block, block, their scraping blight.
We're gonna block, gonna block, all the bots tonight.

When the server dings five, six and seven,
We'll be right in bot blocker heaven.
We're gonna block all the bots tonight,
We're gonna block, block, block, their scraping blight.
We're gonna block, gonna block, all the bots tonight.

When it's eight, nine, ten, eleven too,
I'll be blocking bots and so will you.
We're gonna block all the bots tonight,
We're gonna block, block, block, their scraping blight.
We're gonna block, gonna block, all the bots tonight.

When the counts hit twelve, we'll laugh and yell,
As a dozen bad bots have just went to hell!
We're gonna block all the bots tonight,
We're gonna block, block, block, their scraping blight.
We're gonna block, gonna block, all the bots tonight.

University of Toronto Goes Bat Shit for VPI

Something coming from the University of Toronto keeps making periodic pitstops at my server and only request _vpi.xml like I give a shit about this file.

142.150.4.114 [kahuna.erin.utoronto.ca.] "Firebat 2.5.22" "/_vpi.xml"
Looks like a bunch of bullshit to me as I tried to weed through the ramblings about Jabber groupchat protocol since I've never had anything remotely related on my server whichs brings up the million dollar questions, why is this little fucker looking for it?

Dunno what the motives are but they didn't get far, back to class asshole.

Tuesday, September 05, 2006

Firefox Memory Leaks

Leaving Firefox 1.5 up and running too long without ever closing it for days always seems to eventually cause issues like the swap drive running non-stop or something.

Anyway, I decided to keep the Windows Task Manager up and running all the time so I can monitor Firefox performance and it appears there are some serious memory leaks and issues with closed or stopped downloads that may not be stopping the thread reading the data in the background.

A couple of easily reproduced problems involves stopping a very large page downloading, we're talking thousands and thousands of lines of text, but it appears to keep loading into memory even after it's no longer visible, pushing the memory footprint up to 200MB+ with only a couple of tabs open.

Sure hope they do some better testing on the 2.0 code as I may switch back the IE 7 if it's substantially better as one thing Microsoft does know how to do is keep their code from leaking memory and not leaving zombie threads running in the background.

Sunday, September 03, 2006

Scrapers4U.de

Today I noticed another hit from this same server farm in Germany with something pretending to be a Windows browser:

62.75.218.82 [elbe016.server4you.de.] requested 16 pages as "Mozilla/4.0 (compatible; MSIE 6.0; Windows 98; Win 9x4.90)"
So I checked my archives and sure enough it's been here a time or two before attempting to get inside and there was some hit's from other assocated IP's in their range.

Who hosts this mess appears to be intergenia.de:
netnum: 62.75.128.0 - 62.75.255.255
org: ORG-iGCK1-RIPE
netname: DE-INTERGENIA-20010727
descr: intergenia AG
Which also owns plusserver.de, server4you.de, server4you.com, netfabrik.de, and some end user services who's IP's may be a part of intergenia.de's range, no clue.

The plusserver.de, server4you.de and netfabrik.de both appear to use this range:
inetnum: 217.172.167.0 - 217.172.169.255
netname: PLUSSERVER-1
descr: PlusServer - Dedicated Premium Serverhosting
descr: http://www.plusserver.de
The server4you.com seems to have this block:
OrgName: Server4You Inc.
NetRange: 69.64.32.0 - 69.64.63.255
Comment: http://www.server4you.com
Which means the crawler that started this search still can't be pinned down to a specific hosting block for server4you other than the reverse DNS claims it's server4you.de. I poked around doing a few nslookups in that range and they return either return static-ip-62-75-*-*.inaddr.intergenia.de or someserver.server4you.de so I'm a little hesitant just to block the whole intergenia.de range.

So it looks like I'll block the obvious hosting ranges by IP and server4you.de by reverse DNS for now.

Bots from ServerDeli at Mediopia

Something came crawling from ServerDeli hosted at Mediopia, and it was the typical bot with an invalid user agent if you notice the space between "compatible" and ";" and nevers asks for robots.txt, just pages.

Here's the crawler info:

209.125.47.35 [win1.serverdeli.com.] requested 26 pages as "Mozilla/4.0 (compatible ; MSIE 6.0; Windows NT 5.1)"
Sorry, but my site isn't a deli snack for whatever bullshit you're running.

It always always gives me pause when you see a webhosting company using HOTMAIL addresses for their contact information:
OrgName: MEDIOPIA TECHNOLOGIES (IMA'D W/ 69998
OrgID: MTIW6
Address: 9507 34TH AVE
City: JACKSON HEIGHTS
StateProv: NY
PostalCode: 11372
Country: US

NetRange: 209.125.47.0 - 209.125.47.255
CIDR: 209.125.47.0/24
NetName: ATWORK-65024-55156
NetHandle: NET-209-125-47-0-1
Parent: NET-209-125-0-0-1
NetType: Reassigned
Comment:
RegDate: 2005-05-02
Updated: 2005-05-02

OrgTechHandle: ACH48-ARIN
OrgTechName: CHICO, ALFREDO
OrgTechPhone: +1-718-476-0313
OrgTechEmail: MYMEDIOPIA@hotmail.com
So it looks like blocking 209.125.47.* wouldn't hurt anything.

Core-Project Hijacks an IP

Saw these idiots again today looking for FrontPage on my server:

207.226.161.69 - "POST /_vti_bin/_vti_aut/author.dll HTTP/1.1" 404 1176 "-" "core-project/1.0"
207.226.161.69 -"HEAD / HTTP/1.0" 200 - "-" "-"
207.226.161.69 - "POST /_vti_bin/_vti_aut/author.dll HTTP/1.1" 404 1176 "-" "core-project/1.0"
The IP appears to be dedicated to a single customer hosted on Rackco.com:
cigar-review.com
cigarreview.com
Sadly, Rackco has shared and dedicated hosting so I was unable to easily pin down if this was a compromised server or some little script monkey running in a different account on a shared server.

I guess the only thing I'm amused with is how would some random script in shared hosting, if that is indeed the case, crawl out using a different IP than the server default.

Traceroute have a few clues:
ge6-14.colo02.ash01.pccwbtn.net (206.223.115.48)
ge13-1.br01.ash01.pccwbtn.net (63.218.44.125)
209-8-237-222.rackco.net (209.8.161.222)
mike.rackco.com (209.8.238.194)
cigar-review.com (207.226.161.69)
Still nothing pointing out more than one IP to block.

Ah well, either way, can't seem to narrow down the IP range assigned to Rackco because rwhois.cais.net isn't responding and ARIN just shows the major block assigned to PCCW formerly "Beyond The Network America, Inc.".

Wel'll keep an eye on this one.

Monday, August 28, 2006

Google Utilized in Phishing Exploits

Maybe the title is a little bit of link bait but it's also accurate as I received a WellsFargo phishing email today with a redirect link through Google.

Some of you may remember how I've complained a time or two about being abused via various Google proxy servers and sure enough they have something else that's vulnerable to being used by abusers.

The link to the phishing site used Google to redirect victims:

http://www.google.com/url?sa=t&ct=res&cd=7
&url=http%3Awebtracpro.valleyvistamortgage.com/wellsfargo/Update.html
How's that for Google's war on anti-phishing?

Yes, I know that's a cheap shot but they really need to fix some vulnerabilities over there and maybe after enough cheap shots someone will pay attention, who knows.

Onward with our phishing expedition!

Here's a screenshot of the email sent by the Wells Fargo "Safehaebor Department" which is amusing that they didn't even bother spell checking their phish but most people are illiterate and wouldn't notice such details.



Here's a screenshot of the actual "Update Sistem" (typo in the title) phishing page itself on the compromised server:



And the form sends the data to some place in The Czech Republic:
http://mailform.cz/
The only amazing part is that I notified the people with the compromised server a couple of hours ago and the phish site is still live as I write this, supposedly after their IT dept. was going to handle it ASAP.

So there you have it, another exciting episode of Gone Phishing.

Until next time...

Thursday, August 24, 2006

Inhoster Spammer Hits My Unprotected Contact Form

To allow visitors to let me know that my bot blocker MIGHT be making a mistake, which has happened now and then as it evolved, I had to leave one email contact form unprotected and wide open to potential bot abuse.

This has never been a problem for a long time and suddenly some jerk hosted on Inhoster started fucking with me which has actually been quite interesting.

85.255.117.253 [85.255.117.253-xbox.dedi.inhoster.com]
"User-Agent: Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1)"
Of course my page requires a POST method and isn't abused by the simple GETs, and for my own reasons I didn't think a CAPTCHA was appropriate on this page as I wanted feedback without making it too hard for people.

I was breaking my own anti-spam rules on this page just because I didn't want to reject any legit posts by accident as I was trying to collect all the information I could, but now I'm implementing a few of the filters.

This first thing I did after the spambot started messing with the form was to simply start rejecting all posts with specific HTML tags. To further filter the spam, I'm rejecting any post that is nothing more than a pile of links as they were dumping a bunch of links per post, but still allowing people to send me a link or two as long as it falls within my framework of what legit content looks like.

This seems to be bouncing them at the moment and I'm not sure what the purpose would be for them to continue to spam my form if I don't allow them to dump links, but we'll see what happens.

One added benefit discovered when I was testing was it even bounced a couple of those spammy "link request" emails because they have too many links in them.

Sweet.

Try the javascript trick...

A really cute trick to play on spammers is to make the form submit activate javascript that includes additional data fields that wouldn't be submitted unless they run the javascript as another way to verify human vs. bot without using a CAPTCHA.

The only drawback to this trick, which is inconsequential IMO, is that the Google and Yahoo translation proxies bust this all to hell as they replace all of your links with links back to their translation proxy, which of course doesn't send the data through the proxy properly.

SCRAPER BUSTED #3 - UPDATE Cloaker Surfaces on Netfirms

The same cloaking bullshit artist I wrote about before has surfaced on Netfirms server.

Details:

IP Address: 80.77.80.103
User Agent: "" [blank]
Where scraping content and redirect appear:
rbmusicartist.netfirms.com/artistic-family-portrait.html
Which redirects to some Ukranian or Russian bullshit artist's site:
Domain Name: DEVAMATRI.COM

Registrant:
Oleg Povaljaev
Oleg Povaljaev (anandasat@narod.ru)
Tereshkovoj
Odessa
null,65072
UA
Tel. +380.482648166
Guess what?

They host it on ThePlanet.com, you could knock me over with a feather, I'm so surprised.

DEVAMATRI.COM (70.87.136.118)
OrgName: ThePlanet.com Internet Services, Inc.
OrgID: TPCM
Address: 1333 North Stemmons Freeway
Address: Suite 110
City: Dallas
StateProv: TX
PostalCode: 75207
Country: US
NetRange: 70.84.0.0 - 70.87.255.255
Guess we should drop Netfirms in our blocked list too just to be safe:

rbmusicartist.netfirms.com (64.34.66.18)
Netfirms Inc PEER1-NETFIRMS-02 (NET-64-34-66-0-1)
64.34.66.0 - 64.34.66.255
Well, it's not much, but a little blocking each day will keep the scrapers away.

Now, here comes the real fun...

I was curious what else was on the server with DEVAMATRI.COM (70.87.136.118) and found a shitload of cloaking spam sites:
derrdek1234.info
devamatri.com
fred00med.info
fredodermok2.info
goramon.com
greddertrniko.info
koljazzza.info
nikkasder4ee.info
nikkrongz.info
niko0lwerty.info
nikolannsw12.info
nikolansedd.info
nikolas1qqq4.info
nikolas1qwe.info
nikolazqwii.info
nikolfdsaz.info
ringvvv.info
vvvorgs.org
vwwvcom.info
wvvver54.info
xkoljazzzao.info
Note: The sites are indexed in both Yahoo and MSN but they aren't in Google.

Probably not the last of the sites from this slimeball, most likely the tip of the iceberg, but it's definitely a start to unearthing his network of crap.

SCRAPER BUSTED #11- Inhoster Scraper Indexed by Yahoo

Couple of weeks back I posted about blocking Inhoster which was oozing with spambots with one scraper in their midst and that scraper has finally surfaced.

The scraper's ID is:

IP Address: 85.255.116.178
User Agent: Snoopy v1.2
Which showed up on a page buried on this domain:
index-se.com (85.255.116.182)
What a concept, 2 IPs in Inhoster for one scraper.

Now let's dig for some dirt!

A reverse-IP lookup reveals the scraping IP address 85.255.116.178 is also the IP for FINDALLBEST.COM which looks just like index-se.com.

85.255.116.178: FINDALLBEST.COM
Domain Name: FINDALLBEST.COM
Registrant:
N/A
Nekto (nekto@utopia.com)
Jamaica 17
Cuba
null,12476
CU
Tel. +543.56576767
The info for index-se.com claims to be from the US:
Domain Name: INDEX-SE.COM
Registrant:
Index SE
Index SE (admin@index-se.com)
67 Mt. Auburn St.
Cambridge
,02138
US
Tel. +617.4959659
85.255.116.182: SEARCHADULTSEX.COM:
Domain Name: SEARCHADULTSEX.COM
Registrant:
N/A
Nekto (nekto@utopia.com)
Jamaica 17
Cuba
null,12476
CU
Tel. +543.56576767
So I got curious what else was between 85.255.116.178 - 182 and it was all the same crap:

85.255.116.179: right-pharmacy.com

Different registrant but domain redirects to buy-soma-online.findallbest.com, there's a shock:
Registrant:
N/A
Alexei Aniskevich (alex@coolsearch.biz)
Sopruse pst 15
Tallinn
Harjumsa,50707
EE
Tel. +372.715713
85.255.116.180: wagemax.com

This one is just a Plesk domain placeholder page at this time and another registrant.
Domain Name: WAGEMAX.COM
Registrant:
N/A
Alexei Aniskevich (alex@coolsearch.biz)
Sopruse pst 15
Tallinn
Harjumsa,50707
EE
Tel. +372.715713
85.255.116.180: search-paga.com

Yes, same registrant and site looks like all the rest of the crap.
Domain Name: SEARCH-PAGA.COM
Registrant:
N/A
Alexei Aniskevich (alex@coolsearch.biz)
Sopruse pst 15
Tallinn
Harjumsa,50707
EE
Tel. +372.715713
85.255.116.181: coolsearch.biz

Pay dirt! We found the domain linked to the other domains on 85.255.116.180
Domain Name: COOLSEARCH.BIZ
Domain ID: D6614592-BIZ
Sponsoring Registrar: ESTDOMAINS INC
Sponsoring Registrar IANA ID: 832
Domain Status: ok
Registrant ID: DI_2271261
Registrant Name: Alexei Aniskevich
Registrant Organization: N/A
Registrant Address1: Moisavahe 64-1
Registrant City: Tartu
Registrant State/Province: Tartumsa
Registrant Postal Code: 50707
Registrant Country: Estonia
Registrant Country Code: EE
Registrant Phone Number: +372.715713
Registrant Email: alex@coolsearch.biz
When you go to coolsearch.biz it automatically takes you to: www.gigasearch.biz
Domain Name: GIGASEARCH.BIZ
Domain ID: D7182275-BIZ
Sponsoring Registrar: ESTDOMAINS INC
Sponsoring Registrar IANA ID: 832
Domain Status: clientTransferProhibited
Registrant ID: DI_2191316
Registrant Name: Alexei Aniskevich
Registrant Organization: N/A
Registrant Address1: Sopruse pst 15
Registrant City: Tallinn
Registrant State/Province: Harjumsa
Registrant Postal Code: 50707
Registrant Country: Estonia
Registrant Country Code: EE
Registrant Phone Number: +372.715713
Registrant Email: alex@coolsearch.biz
85.255.116.181: your-searcher.com
Domain Name: YOUR-SEARCHER.COM

Registrant:
N/A
Alexei Aniskevich (alex@coolsearch.biz)
Sopruse pst 15
Tallinn
Harjumsa,50707
EE
Tel. +372.715713
Let us continue with more of this puzzle...

Let's explore gigasearch.biz a bit more:

69.50.163.9: gigasearch.biz

We did find some similar scraping in this range:
69.50.190.242 "Snoopy v1.2"
Actually, the range 69.50.*.* has a ton of scraping so seeing a link to this scraper and the Snoopy user again yet again was no surprise.

GigaSearch.biz is hosted on our old friends Intercage which hosted Scraper #4 and Scraper #6 which I think may be all the same scraper as everything just keeps linking them together from host to host, some similar IP ranges and the same user agent. Nothing concrete but all the circumstantial evidence is overwhelming that they may be somehow related.

Most amusing is all the links on gigasearch.biz redirect to find.fm, and this relationship could be interesting but I'm getting sick of chasing this scraper / spammer at this point.

The host of our busted scraping pals #4, #6 and #11:
OrgName: InterCage, Inc.
OrgID: INTER-359
Address: 1955 Monument Blvd.
Address: #236
City: Concord
StateProv: CA
PostalCode: 94520
Country: US
NetRange: 69.50.160.0 - 69.50.191.255
Let's see what else is on the Gigasearch.biz server:

69.50.163.9: blanksearch.biz

This domain is NSFW with raw porn all over it.
Domain Name: BLANKSEARCH.BIZ
Domain ID: D6761115-BIZ
Sponsoring Registrar: ESTDOMAINS INC
Sponsoring Registrar IANA ID: 832
Domain Status: ok
Registrant ID: DI_3009123
Registrant Name: Ivars Kaupers
Registrant Organization: No
Registrant Address1: Skirgailos 15
Registrant City: Kaunas
Registrant Postal Code: 75128
Registrant Country: Lithuania
Registrant Country Code: LT
Registrant Phone Number: +370.571689
Registrant Email: ivars@blanksearch.biz

69.50.163.9: tgp-porno.net

This site brings up another of the same old porn links again.
Domain Name: TGP-PORNO.NET
Registrant:
N/A
Alexei Aniskevich (alex@coolsearch.biz)
Moisavahe 64-1
Tartu
Tartumsa,50707
EE
Tel. +372.715713
Last but not least, the server with find.fm hosts a few other garbage domains with the same links about pills and porn on them all, with "find.fm" on the bottom of the page which was a big shocker as well:

Domains on 64.111.196.119 (Find.fm)

adultwebfind.com
carwebsearch.com
cashwebsearch.com
dmns4sale.com
gamblingwebsearch.com
pharmacywebsearch.com
travelwebsearch.com
your-needs.info

Well, that's all for now.

Needless to say, they can't hide for long as they leave a slimey trail that can be followed.

Scrape me again assholes, let's unravel the rest of your bullshit sites.

Wednesday, August 23, 2006

Slow Blog Week

Sorry if you aren't getting your daily dose of bad bots and the usual run-on ranting sentences packed full of expletives but I've been busy the last few days catching up on some accounting and doing some work on software and websites.

If you really need your 'fix' you can catch up on the latest of the Nutchies that are still harassing the shit out of me from a link from Doug Cutting's site.

As you can tell in the final comment posted, and a few previous comments, that I'm losing patience with this bunch of crawling-the-web-is-our-right cultists.