Wednesday 19 November 2008
SpeciesIndex - a new taxonomy?
The split between the notion of a name and a taxon is now widely understood. It is no good knowing the name of an organism if this information isn't accompanied by a description of some kind. The use of a name alone relies heavily on a shared understanding of the meaning of that name. If you remain to be convinced of this take a few minutes to view Rich Pyle's excellent presentation at TDWG 2008 - Names, Concepts, Codes and Lots of Confusion
Moving to a world where biodiversity data is manipulated by machines will require a greater level of precision than is given by the use of names alone. The shared understanding that provides the names meaning is ripped away along with the context in which the name is used. I simply isn't possible to mix-and-match data from different sources on the basis of the names having the same meaning across all of them.
The initial response to this situation was to encourage the use of sensu or sec whenever a name was cited so that a name was alway accompanied not only by the name of the author who first published it but also by the name of the author who published the prefered taxon description for this name. This works well in the written world but still doesn't help machines much. There are many ways to cite the sec authors and it is non-trivial to try and match up different spellings automatically. Far better for each taxon circumscription to carry a globally unique id (GUID) such as a URL, LSID or even a UUID. We can then cite a name and qualify it with an ID that points to the precise meaning for that name - in this context.
Now there has been an ongoing debate about which is the best technology to use for GUIDs and what should be returned when a GUID is called but for species descriptions it seems clear that this data should at least come in the form of a human readable document. The majority of existing descriptions in the literature are text documents with illustrations and their counterparts on line are increasingly refered to a "Species Pages".
A species page (from here) is a single web page describing a named species taxon. It contains descriptive information about the species such as morphology, ecology, geography, uses etc. It may include text and/or still and moving images and/or sound files. The information may be free standing or may be interpreted in context of other pages in which case it is obvious from the presentation that this is the case. Analogs to species pages in the physical world are entries for species in encyclopedias, floras, faunas and guide books. Questions that SpeciesPages help Users answer are: Given a name what kind of organism is this? What does it look like? How does it live? Where can I find it? Is it endangered? Given a specimen how can I confirm its identity? Does this match the descriptive data given in the page? Species pages contain more than one piece of information about the species - they are compiled works. A single photograph is not therefore a species page but a photograph with notes pointing out why the specimen in the photograph should be consider a particular species would count as a species page. Generally species pages are only existed for taxa that are considered accepted by the Publisher. If they exist for non-accepted taxa (e.g. synonymous taxa) then they are clearly linked to the accepted taxa.
There are hundreds of thousands of Species Pages already available. Wikipedia has over 90,000 pages that contain Taxoboxes and could be considered species pages. USDA Plants has over 40,000 plant profiles. There are many other such website, large and small, that contain species pages.
This is a tremendous resource. We have taxonomic descriptions tagged with GUIDs (in the form of URLs) in a standard format (HTML). This isn't totally ideal. Perhaps the GUIDs should be DOIs or LSIDs and perhaps the descriptions should be marked up in some standard XML or tagged with RDF but