Showing posts with label homepage. Show all posts
Showing posts with label homepage. Show all posts

Sunday, August 21, 2011

TagCloud and SEO in Plone: problem solved and lesson learned

This article is an analysis of how a minor bug in a Plone product we are maintaining give us some problems on a production site, how fixing it was not enough to revert problems, and what lesson I learned.

The environment
The Plone site I'm talking about is an old Plone 3.3 installation with an old hardware. It worked without problems for years then suddenly it started to be slow (sometimes really slow).
In front of this Plone installation there's also a Varnish installation that cache also HTML for anonymous users (a standard for us).

So what's can be turned wrong?!

After started checking the problem I also find that it wasn't a memory consumption problem (a lot of free memory, thanks to the use of plone.app.blob also on this Plone 3.3 installation), but only a CPU usage (unluckily we have only a single core for the Plone site).

Next step was checking what the site was doing to keep the CPU so busy. Before installing products like zope.healthwatcher, I spent some minutes simply checking the HTTP log and I found that the site is... "really popular". I mean it is very often visited by web crawlers, mainly Googlebot. This is not strange: the site is the main site of a well know public agency, and update really often.

The Vaporisation problem found
Is this enough to make a Plone site "slow"? Obviously no! The problem was not Plone but the TagCloud portlet.

One of the best features of collective.vaporisation (maybe the main feature that convinced us to takeover the project and maintain it) is the joint navigation.
When activating it, clicking on one of the links inside the cloud will not simply display a search result, but a search result where you can also navigate through found items using related additional terms (something like a faceted navigation).

The first problem was that customer site use a lot of different keyword, using them widely in a lot of site contents.
In this way the joint navigation became complex. Imagine a web crawler that starts to scan the tag cloud results page: it's able to follow a lot of additional link that create a great permutation of different search results (let call this big number N).

The second problem (the real one) was a bug in the way Vaporisation was creating links from the cloud to the result page (note that version 1.2.0 of the TagCloud fixed this problem so recent releases will not give you such behavior): older releases were calling the cloud_search template on the context.
This mean that when the visitor was checking the cloud links from the home page, it will call Plone as http://thesite.com/cloud_search, but when visiting another page, the URL will became http://thesite.com/path/to/document/cloud_search.

This is a disaster for the cache that Varnish is trying to produce for our site: the two URL are different from the cache point of view, but in facts on Plone they are generating the same result. This is really bad (also for page rank).

This also raise the number of possible cloud search pages from N to NxM (where M is the number of documents in the site). Terrible!

The fix was simple: change the way the URL to the cloud is created, always give to users the http://thesite.com/cloud_search version.

Google: the elephant
When I released this fix I make another mistake (not very lucky...).
Always remember that in Plone the context can be important. If you created a site view that must be called onto the site root is better to define it callable only on this context. An common error can be define this view callable everywhere (like old CMF skins templates).

So even if all URL were always generated in the right way, the view was defined still as follow:

<browser:page
	 for="*"
	 name="cloud_search"
	 class=".search.CloudSearch"
	 template="search.pt"
	 permission="zope2.View"
	 />

In this way calling (manually) again the http://thesite.com/path/to/document/clud_search will still be valid.

Why this is still a problem? Haven't you fixed all possible wrong links?
Because of Google long memory.

Even if we removed all links to the useless cloud_search call, Google already indexed them so it continue to call all those URLs and make the Varnish cache useless.
Maybe with the time those kind of call could decrease (as the search engines will find no links to this page anymore), but after applied the fix we get no benefits.

Version 1.2.2 fixed this, defining also the right context for the view:

<browser:page
	 for="Products.CMFPlone.interfaces.siteroot.IPloneSiteRoot"
	 name="cloud_search"
	 class=".search.CloudSearch"
	 template="search.pt"
	 permission="zope2.View"
	 />

After this change all calls to something that isn't http://thesite.com/cloud_search will simply generate a NotFound error page. This is good: Plone is very fast to generate this page, and this kind of pages (if no link lead to them) will rapidly disappear from the index.

Some SEO enhancement can help
Version 1.2.2 helped a lot, but in the meantime I was also reading a book about SEO and talked with coworkers (real SEO expert) of those kind of problems.

Must Google be aware of the cloud_search page? Is a good thing that this page will be indexed?Also, the site still use the joint navigation, so a major number of pages are still called, but in any case are those pages giving some good feedback to Web users that performs searches?

I think that there isn't an universal answer (I've doubt also related to simple Plone SERP... must be indexed by Google? Maybe Plone itself must think about this) but in this case no: let search engines index results page of search performed by tag cloud is useless.

Version 1.2.3 make the Vaporisation product more "SEO friendly".

First of all we can suggest search engines to not follow a link when scanning our site putting a rel='nofollow' attribute on it.

The new version of Vaporisation will put this rel nofollow value on every tag cloud link. I read recently that this attribute is only a suggestion, but can help.

The second and last change of this version is the use of a tag meta in the cloud SERP:

<meta name="robots" content="noindex, nofollow" />

This says to the search engine to not index this page (noindex value) but also don't follow any links from this page (nofollow value, less important there).

In this way Google quickly stopped to index also the big number of possible joint navigation results pages.

Damn Google Reader!
After every fix I applied, the site quickly became faster.
After all those fixes I still saw an heavy Google access to the site on a lot of different pages: this time to the search_rss page (used by Plone to generate RSS feeds from search result), commonly called with an URL like http://thesite.com/search_rss.
The site now was not so slow, but I liked to continue the investigation.

The problem was still the first one I analyzed above: the search_rss URL was available from the cloud_search page and so, when using older releases, was created on the context and not on the portal root.
So while a version lower than 1.2.0 was active, Google indexed a lot of cloud_search versions but for each one also a useless search_rss page.

Again, have fixed the problem for the "master" cloud_search page will not stop Google from indexing all the RSS calls.

This time we can't rely onto a well-done meta tag, as the RSS isn't HTML! Where I can put the tag?!

We have two alternative ways.

The quicker one is to block every web crawler access to the search_rss page, using simply a robots.txt file on the site root (but this will works with query parameters?):

User-agent: *
    Disallow: /search_rss

The other way, that also leave the search_rss page indexable by services like Google Reader, is to make also the search_rss template callable only from the site root.

This need some simple Plone customization (outside the Vaporisation product itself). Maybe Plone 5 (or 6?) will not use anymore CMF templates, but right now we have a lot of them all around Plone.
One of the problem of old-style template is that they can be called on every possible context. And obviously search_rss template is a very old template...

So the fix is not elegant as the cloud_search ones. We can:
  • configure Apache to disallow any search_rss call outside the site root
  • manually check if the site context is the site root; if not, we manually raise a NotFound error.

Conclusion
I found the resolution of this someway "funny".

Other Plone sites that installed directly the 1.2.0 version of Vaporisation didn't suffer all problems that are here described because the cloud_search page was always called in the "most correct way": on the site root.
So in this case Google did not indexed all other useless alternative way to call the same page (this mean also no need to fix the Plone basic search_rss)

Last thing: if you create a Plone template planned to be called on a context be sure to register it only on this context. If this template is an old CMF ones keep an eye on how you create links to this template.

Saturday, April 2, 2011

Plone portlets are not enemies... just rude friends

I want to put in words all I like (and don't like) in Plone portlets, and all the problems and limit that everyday a Plone user can find using them.

First main reason for liking portlets
Portlet are quite simple to be developed, and using formlib you can generate your add/edit form in a very straightforward way. You have no limit to the number of configuration option you can add to your form. Then, after that, you can use them all to develop you target portlet template.
All of those configuration are persistent, and are saved in your Plone site.

Second main reason for liking portlets
If the first reason is only a "developer toy", the real good thing of using portlet in Plone is that you are now forced to keep them in the classic two-columns.

Adding a new area where storing portlet (a new portlets manager) is very simple and well documented, for example see "Adding Portlets Manager" guide.

This is not a so common task however, just because two great product did this for you.

ContentWellPortlets
The first one I want to introduce is Products.ContentWellPortlets. The idea behind this is very simple: Plone gives to you two areas where managing you portlets? So, ContentWellPortlets add two additional ones: one above the content main area, another under it. New portlets area works exactly as the standard ones!

This simple general new behavior can resolve a lot of additional theme requests. Can you think about some other Plone area where you need to put portlets?!

collective.portletpage
...yes. There's at least one additional place where you can like portlets: inside the content area that your visitor is reading. Maybe not in every content (for this maybe is better to wait for Deco super drag&drop feature), but to be honest Plone today didn't offer much out-of-the-box for creating homepages.

A simple document, or a collection (even if with a custom view) is not always enough. A lot of Plone websites offer as main homepage (but also but main page for site subsections) a set of blocks, one with a static text, another that display last news, another with the box to register users to a web service... and so on.

As I said, is difficult to do this with Plone out-of-the-box... so what? We must simulate dynamic content with a static use of TinyMCE? We must create some Plone Site root specific view?
Obviously no! In the great star system of Plone add-ons, you can find a product for creating your homepages. Our choice is 90% of times collective.portletpage. We found it perfect for making homepages:
  • simple to be customized for your theme
  • simple for the power user that manage the homepage to be customized
  • simple to add new features (AKA: new portlets)
Using the static text portlet, the collection portlet (distributed directly with Plone), and a good skinner in the team, you can satisfy maybe more the 50% of requests. For the other half, rarely you need to develop new portlets (you can find a lot of theme as Plone add-ons if the initial set is not enough).

NB: a lot of other Plone users choose another well know project for making homepages. I'm talking of Collage.

So: what's bad in portlets?
There nothing really bad in portlets in Plone: it's a great technology, but you need to know some thing to live with them without headache.

Lesson 1: remember to uninstall
This is a good thing to be done with all Plone products but some of theme didn't leave problems behind if you remove theme without uninstall. Sometimes they left some garbage, sometimes nothing...

But not portlets.

With portlets you must remember to uninstall the product that gave you the portlet before remove it from your buildout.
If you don't follow this simple rules sooner or later you will try to access the "manager portlet" section; then you'll get the standard not found error that says "We’re sorry, but that page doesn’t exist…".

What is missing? To see this, let say that we installed a product named example.portlet.foo, we installed it, then we removed it from our buildout.

On Plone 3, you will see a NotFound error when accessing the "manage portlets" views; to see the real error in console you need to enable the logging of NotFound errors at /error_log in the ZMI.

The only know way to restore proper behavior seems to add the product again, then uninstall it, finally remove it from buildout for the last time.

Luckily latest Plone 4 releases has removed this problem; someway the exception is now swallowed. You can remove the product then don't see any problem in the "Manager portlets" section... but please: don't do this! Uninstall is good!

Lesson 2: you can't remove broken portlets
What is a broken portlet?
When you uninstalled the Plone product that give you a portlet, you can also forget to delete some portlet objects. Let explain this, but be warned that this time we have a deep difference between Plone 3 and Plone 4.

Plone 3 has the best behavior: if you remove the product you can continue to use your Plone site, and also the portlet management section.

Broken portlet in Plone 3Obviously you can't continue to see, use and manage the portlet you left there, but the site still works.
What you see is only a portlet that says "this object from the unknown project is broken" and the common "x" icon for deleting it.

But if you try to delete the portlet in this way you'll get no result, and another error in the log:
Traceback (innermost last):
Module ZPublisher.Publish, line 127, in publish
Module ZPublisher.mapply, line 77, in mapply
Module ZPublisher.Publish, line 47, in call_object
Module plone.app.portlets.browser.kss, line 65, in delete_portlet
Module zope.container.ordered, line 243, in __delitem__
KeyError: 'broken'
So, seems that there isn't a way to remove the portlet if you removed (and uninstalled) the product.

My personal opinion about this is different from the bad attitude of didn't uninstalling products we see at "Lesson 1": forget to uninstall is bad, but this is different!
Maybe I've used the portlet in 100 different places I don't remember (or my users did it)... how can I scan the whole site to know where I added locally portlet, or to what group or user's dashboard it had been installed?

The best behavior there is make possible for the user to remove portlet, alway, also if is broken. No Plone versions right now give this: be warned.

Let's talk of Plone 4.

Plone 4 is more problematic. If you remove the product then forget a portlet you will have a real problem, an immediate error:
TypeError: ('Could not adapt', ... 
I was not aware of this problem with Plone 4 until the time I've started this article, so is possible that this is a bug.

Lesson 3: know how to extend portlets
The foo product I named as example is there, but is splitted in many versions:
http://svn.plone.org/svn/collective/example.portlet.foo/

For the rest of the article I will rely heavily on this example project, take a look at the source when needed.

The problem I'm introducing is about errors you will find when you add fields to you portlet (maybe because a new releases is ready), but this code will be used in sites where old instances of the portlet still lives.
There's a know method of writing portlet code that make you don't fall into this. You can still find this error when playing with old products, or relying onto old tutorials, or because you are using older version of ZopeSkel (because... you are using ZopeSkel, right?).

The error came from the portlet Assignment class, or better from how you make it.
The first sub example will introduce you the early version on the portlet (that is a stupid piece of code, that show a message you put into it). The portlet has a single, string, field. Very simple!

The code is there:
http://svn.plone.org/svn/collective/example.portlet.foo/version1/

You need only to note for now (but will become clear later) the Assignment code:
class Assignment(base.Assignment):
"""Portlet assignment.

This is what is actually managed through the portlets UI and associated
with columns.
"""

implements(IFooPortlet)

def __init__(self, foo_field1=u""):
self.foo_field1 = foo_field1
After that, a quick view to the main expression in the portlet template:
tal:content="view/data/foo_field1"
The question is: how if later we add a new field to the portlet? How Plone and portlets react?

So, let's go the version 2 of the example:
http://svn.plone.org/svn/collective/example.portlet.foo/version2

We have a second silly field that the user can fill with data,
Starting from example 1, this new version will be simply like this:
class Assignment(base.Assignment):
"""Portlet assignment.

This is what is actually managed through the portlets UI and associated
with columns.
"""

implements(IFooPortlet)

def __init__(self, foo_field1=u"", foo_field2=u""):
self.foo_field1 = foo_field1
self.foo_field2 = foo_field2
Trivial... But this new version will badly works with portlet created with previous code.

The first problem is related to the portlet template:
tal:content="view/data/foo_field2"
This expression will raise a simple error. Why? Because the old portlet assignment (that is persistent) has no foo_field2 data at all. We will get:
NotFound: foo_field2
The guy who develop the portlet need to think to his previous version, so you can think that he must provide a template that didn't break if a parameter is missing.
However this is not really needed. Let's go on for now.

Now the worst behavior. If you go to the manage portlet, then you'll try to edit the old portlet with the version 2 of the code, you'll get:
Traceback (innermost last):

* Module ZPublisher.Publish, line 127, in publish
* Module ZPublisher.mapply, line 77, in mapply
* Module ZPublisher.Publish, line 47, in call_object
* Module plone.app.portlets.browser.formhelper, line 123, in __call__
* Module zope.formlib.form, line 782, in __call__
* Module five.formlib.formbase, line 50, in update
* Module zope.formlib.form, line 745, in update
* Module zope.formlib.form, line 820, in setUpWidgets
* Module zope.formlib.form, line 408, in setUpEditWidgets
* Module zope.schema._bootstrapfields, line 171, in get

AttributeError: foo_field2
I can't edit my old portlet anymore! I need to delete it, the create it again?
What is the problem here? The Assignment class don't provide a value for the foo_field2 field and this break the edit form.

There's a simple way to fix this, writing you Assignment class in a different way:
class Assignment(base.Assignment):
"""Portlet assignment.

This is what is actually managed through the portlets UI and associated
with columns.
"""

implements(IFooPortlet)

foo_field1=u""
foo_field2=u"foo"

def __init__(self, foo_field1=u"", foo_field2=u""):
self.foo_field1 = foo_field1
self.foo_field2 = foo_field2
This Assignment class (and the final version of the example) is there:
http://svn.plone.org/svn/collective/example.portlet.foo/version3/

This final way of making Assignment, where you have default values fallback taken from the class instead of the instance, solve both problems:
  • the template will always find a foo_field2 value (to make this more clear, the default of the field is not an empty string
  • the edit form can be used normally. New values put in the foo_field2 field are then saved.
Conclusions
  • Portlets are great for compose your layout
  • Plone has a lot of great portlet products
  • You must remember to uninstall portlet you don't need anymore
  • You must delete all portlets before removing the product
  • Create new portlet, but use class level attribute for our assignment
  • Use ZopeSkel templates