Minasan, Watashiwa Wawan Desu...

NurCell Movies
Showing posts with label Content. Show all posts
Showing posts with label Content. Show all posts

Friday, February 18, 2011

How to Make a WordPress Blog Duplicate Content Safe

Supplementary indexIn one of my recent posts I wrote about the duplicate content issue. This topic is especially important to me since my blog uses the WordPress content management system which, when used with the default configuration, is not duplicate content proof. In fact this CMS is capable to render almost 100% of your content duplicate. As usual the fault of the system has roots in its advantages. WordPress has many features facilitating blogging and linking, such as RSS feeds to posts and comments, trackback URLs, monthly archives and so on. In the same time this variety of URLs returning similar or identical pages represents a clear case of duplicate content.

The first evidences of duplicate content produced by your WordPress CMS can be found in your sidebar. They are category pages and monthly/daily archives. Category pages store your articles posted under the same topic – a category. Such pages have no unique content; they are just a collection of your previous posts. Monthly and daily archives also simply group your previous articles by the date of posting. Sometimes when you have only one post in a given day, the archive page for the date and your post are totally identical.

The next case of duplicate content is even more prominent. It can be your home page itself. If it contains not excerpts but the full text of your posts, then it duplicates your post pages. This also applies to the ‘next/previous entries’ pages – those accessible via /page/2, /3, /4 etc.

Feeds. Search engine spiders crawl all the content they can reach and of course this includes RSS feeds too. The additional problem with them is that Google may choose to display your RSS URL in the search results over the link to the original post. In this case the user who clicks this result will see an XML formatted page which is not ‘human-friendly’.

Trackback URLs. Many WordPress templates add trackback links after posts. This links enable authors to track who links to their posts. Usually, if your post URL looks like ‘www.yoursite.com/2006-11-30/yourpost/’ its trackback URL will be ‘www.yoursite.com/2006-11-30/yourpost/trackback/’.

Identical meta-description. By default WordPress doesn’t provide a tool to add unique meta description tags to your posts, and they either have none or share a single site-wide description. Having no meta description at all is a disadvantage, as a properly written one can make your snippet stand out in a SERP. Having an identical description for all your pages is a threat, as Google might get them filtered out as too similar. (see a thread here)

Because of the duplicate content Google search can return less desired URLs (such as feeds or archives instead of original posts); your pages can be moved out of their index, or placed into the supplemental results, which are rarely displayed to users.

What can you do to avoid this problem? You can tell the search engines what URL to index by using ‘noindex, follow’ meta tag, robots.txt exclusions or 301 redirects. Let’s say you want Google to index your front page, posts, single pages and category pages and forbid the spiders from crawling the content of archives, feeds and ‘next entries’ pages – page/2, /3, … To do this you have to add to your header.php the following code:

if((is_home() && ($paged < 2 )) || is_single() || is_page() || is_category()){echo '';} else {echo '';}

For those not familiar with editing templates in WordPress: in your dashboard click Presentation menu item and after the new page is opened – click Theme Editor. In the Theme Editor choose ‘header.php’ and then paste the above code into the editor form. This code has to be inserted anywhere between head tags .

Here the tag is added to the home page but not the ‘next entries’ page (is_home() and ($paged<2)), to your posts (is_single()); to solo pages, like ‘About me’, if you created any (is_page()); and to category pages (is_category()). If you don’t want your categories to be indexed just delete || is_category(). All the other pages will get . They will not be indexed, but this will not prevent crawlers from following their outgoing links.

For this purpose I use Head Meta Description plugin. This plugin can be configured to use an excerpt of your post as a meta description – this is especially useful if you have to add this tag to hundreds of existing pages. Or you can add your own manually as a custom field, which is my personal preference.

By using this tag you tell WordPress to display only the first few lines of your post. This greatly reduces the similarity of home page and your articles. If you have too many existing posts to edit, you can use an ‘excerpt’ plugin, such as this one from Semiologic

You should edit your .htaccess file to perform 301 redirects. Non-www addresses like yoursite.com should be redirected to www.yoursite.com. URL without trailing slashes like www.yoursite.com/category should be rewritten to include it: www.yoursite.com/category/ This can be done by inserting the following code into your .htaccess file:


RewriteEngine On
RewriteCond %{HTTP_HOST} !^www\.yoursite\.com$ [NC]
RewriteRule ^(.*)$ http://www.yoursite.com/$1 [R,L]
RewriteBase /
RewriteCond %{REQUEST_FILENAME} !-f
RewriteCond %{REQUEST_FILENAME} !-d
RewriteRule . /index.php [L]

For more details I advise you to read this: the process or rewriting the URL layout.

For this purpose you should edit your robots.txt file by inserting the following code

User-agent: *
Disallow: /wp-
Disallow: /search
Disallow: /feed
Disallow: /comments/feed
Disallow: /feed/$
Disallow: /*/feed/$
Disallow: /*/feed/rss/$
Disallow: /*/trackback/$
Disallow: /*/*/feed/$
Disallow: /*/*/feed/rss/$
Disallow: /*/*/trackback/$
Disallow: /*/*/*/feed/$
Disallow: /*/*/*/feed/rss/$
Disallow: /*/*/*/trackback/$

Some people find it useful to restrict the number of posts displayed in your home page to 4-5, as less posts are duplicated.

A great article on customizing the more tag in Wordpress.

To avoid the duplicate content issue in WordPress include you should do:Add ‘noindex, follow’ meta tag to your monthly/weekly/daily archives, ‘next entries’, and if necessary, category pagesEnsure that all your pages have unique meta-description tagsSet up 301 redirects for your non-www URL and URLs without trailing slashesRestrict search engine crawlers from indexing your feeds and trackbacksUse more tag to show excerpts in your home page instead of full postsRestrict the number of posts displayed in your home pagereddit_url='http://www.seoresearcher.com/how-to-make-your-wordpress-blog-duplicate-content-safe.htm'

View the original article here

Monday, February 14, 2011

Duplicate content sin #2: Default page linking


Last week I wrote about duplicate content sin #1 - screwy pagination. Today I'm going to explain a much simpler, but bigger problem: The inconsistent default page link.


When I say 'default page', I mean whatever page you'd first see if you navigated to a folder on a web site.


So the default page for Conversation Marketing (the whole site) can be found at www.conversationmarketing.com/. That's the root folder - the main folder housing my whole site.


The default page for all of this month's posts can be found at http://www.conversationmarketing.com/2010/10/. That's the sub-sub folder /10/, in the sub-folder 2010, in the root folder for www.conversationmarketing.com:


cm-folder-structure.gif


You can also find the default page for Conversation Marketing at http://www.conversationmarketing.com/index.htm. And you can find the default page for this month's posts at http://www.conversationmarketing.com/2010/10/index.htm.


Web servers automatically deliver these default pages when a visitor requests the folder - that's why you don't have to add 'index.htm' to these addresses.


The problems arise when a developer or designer links to default pages using different link styles at different times. For example, if your site has a 'home' link that points at '/index.htm' or 'default.aspx' or whatever your default page is, you've created duplication:

Search engines and most people see your home page as www.yoursite.com. Most other sites link to you there, too.But search engines crawling your site also see the link to www.yoursite.com/index.htm, and follow that link.To a search engine, the '/index.htm' page and the www.yoursite.com page are two unique pages with the exact same content.

Voila. Duplication.


The same thing happens if you inconsistently link to subfolders in your site.


I won't even waste time explaining what this does to your link profile. It's bad.


The problem here is duplication. And, as we know, duplicate content sucks.


If you want to avoid this kind of problem, apply Ian's Rule of Simplicity: Always use the shortest version of any default page's address. That version should typically be:


www.yourdomain.com + folders


No filenames.


Do that, and you'll eliminate one huge duplication problem. Best part is, most of your default page links will be in your navigation. If your site was built by a relatively sane person, you can make one change to your site template and fix a site-wide duplication issue. Woo hoo!


By the way, this is also considered a canonicalization problem. I'll never stop ranting about canonicalization - you know that, right?

I've been writing up a storm this week, so no fancy conclusions or funny animal pictures. Bye.



View the original article here

Duplicate content sin #1: Pagination


Earlier this week I wrote about why duplicate content sucks in SEO. I'm going to start mixing in tutorials/explanations of common ways folks end up duplicating content on their sites, too.


Today's topic: Pagination. It's oh-so-easy to generate duplicates with those little 1 2 3 4 >> at the bottom of the page.


Say you've got a site called www.blah.com. You've written an article that's 12 pages long, and added pagination at the bottom, like this:


Typical pagination


It's purty, and it works. When Google or Bing land on the page www.blah.com/articleaboutx/, they see the pagination and the page URL, and they get it. This page is page 1 of your article:



Nice.


Now, Googlebot crawls to page 2 of the article. That page is located at www.blah.com/articleaboutx/p2. Also no problem.


But when it attempts to crawl the '1' link, it sees a new URL: www.blah.com/articleaboutx/p1


That page has the same content as the first article page we saw at www.blah.com/articleaboutx/, because it is the first article page. But it's got a different URL.


google-gets-confused-pagination.gif


Two URLs, same page? Uh-oh. That's a duplication problem of the canonicalization variety.


If you have a large publication with, oh, 2000 articles, and all of those articles are paginated the way I described above, you've created 2000 duplicate pages on your site. And they happen to be the first page of every article - the most important page you've got.


Bloggers will link to the '/' or the '/p1' version randomly, depending on which URL they're viewing when they cut and paste.


Your caching software will have to cache both URLs.


And search engine crawlers will waste their time crawling all of those duplicates.


Blech. Luckily, this is an easy one to avoid.


This one's magical... it's tricky... wait for it...


Link the '1' in your pagination to the original URL for the first page of your article.


So, if your article's first page was at www.blah.com/articleaboutx/, make the '1' link point there, too. Don't point it at /p1.


Wow.


This sounds silly, I bet, but I have yet to see a publisher site, a designer blog, or any other site that paginates get it right the first time. If it's right, it's because a cranky SEO whined about it.


If you don't like the sound of me whining, go ahead and fix it now.


There you have it: One duplicate content problem fixed.



View the original article here

Thursday, February 3, 2011

Forget “Content Production”, Think “Idea Exploration”

Since I wrote about creativity, I’ve been considering the issue of getting inspiration for blog posts.

Some who commented on that article said they never have trouble finding post ideas. But others revealed that they really struggle with getting inspiration. Sometimes, they’re struck by it; other times, they have to go out and track it down bodily.

I’ve found that my approach has a lot to do with how many post ideas I have. I wanted to share my approach here, and see if you felt the same way, or take a different approach.

The burden of having to “produce” content can be overwhelming to the point where it stops production altogether. Feeling that you need to “produce” to a schedule, or on a regular basis, can make you feel a bit like a machine, and make your content seem like an “output”.

If I take this approach, my writing can become mechanical, my posts formulaic, and my points vague and unfocussed. The last thing I want to be is a content sausage-factory, but if I take a “production” philosophy, that’s how I wind up feeling—and it shows in my content.

The “production” approach doesn’t work for me, but the “exploration” approach does.

I find writing posts is a great way to explore the ideas that are on my mind. You probably started your blog because you have an interest in your blog’s topic. What aspects of that topic are on your mind? What elements are you curious about? What areas within that field annoy you, and why?

These kinds of considerations are precisely what inspire me to write. I look at what others are doing and saying and creating, and I reflect on that—maybe not immediately, with a pen in my hand and notebook open, but over time. I let these ideas, motivations, and questions filter, settle, and develop in my mind. They’re always there—we’re always thinking, right? Then, when something really starts to stick in my conscious, I write about it.

Your blog is the ideal place to explore those ideas that are rattling around your mind. Rather than “producing posts”, you might find it helpful to think of your blogging as an opportunity to:

formulate disparate thoughts into coherent conceptsadvance your own theoriessuggest alternative viewpoints or approaches that have occurred to yousee if your readers agree with a hunch you’ve gotinvite readers to help shape your perception, idea, or viewpoint

This post is an example of exactly that. As I mentioned at the outset, this idea—of post writing being a way to explore and develop thinking on a topic—has been sifting through my mind ever since I wrote that post on creativity. There are many half-formed, embryonic ideas in my mind, as I’m sure there are in yours, and this one has finally surfaced as something that I wanted to get a second opinion on.

So here I am, posting about it, in the hopes that you’ll share your thoughts on using your blog to explore and develop ideas in the comments. I’d love to hear them!


View the original article here