04 March 2009

The trouble with names is they belong to people

I recently read Dorothea Salo's latest article, 'Name authority control in institutional repositories', which will appear in the April issue of Cataloging and Classification Quarterly. (You can find the preprint here).

As a repository manager, Salo is aware that name authority problems have a significant impact, not just for librarians responsible for content management in repositories, but also on repository users and the discoverability of our content. She believes that one of the reasons for the problems we experience managing author names is that we never envisaged our institutional repositories as library-managed databases; they were meant to be 'do-it-yourself' (Salo 2009) author deposit mechanisms. This means we didn't plan how to control our author metadata in the first instance.

But even if we had, how would we have controlled it?

Traditional cataloguing standards like AACR2 are designed by librarians for librarians, and for library systems frankly more concerned with stock inventory than resource discovery. Authors have no input in the way their works are represented in a library catalogue; cataloguing standards treat them as just another piece of descriptive metadata.

Whether we populate our repositories through self-deposit or librarians recruiting content themselves, there's no doubt that authors are much more to IRs than just another metadata element.

For starters, without authors institutional repositories have no content, and without content, they don't exist. And the location of authors at the time they create a work is the sole basis for their inclusion in an institutional repository's collection.

Salo's viewpoint is that the problems with consistency in repository content are tied to software. But a quick glance at institutional repositories using a variety of software solutions shows that name variant problems affect them all. No single repository, regardless of architecture, can escape this issue, because it's not a software but a human element. And people are always much trickier than technology.

Salo believes that eventually institutional repository software will improve, and that '[i]n the meantime, institutional-repository managers can only plan to plow large amounts of staff time into managing names' (Salo 2009). But the truth is that it's not that easy. We've already spent inordinate amounts of time trying to find a way to manage author names in Swinburne Research Bank, and we've drawn a blank.

Salo notes that EPrints software (unlike DSpace and Fedora) has an autocomplete function, which allows depositors to select from names in the repository's existing vocabulary when they create author metadata. But this is not a long-term solution. While it might help with cases where authors use their initials on some papers and their full names on others (assuming we're comfortable with overwriting these differences---and that's a big assumption), it's just not appropriate when authors associate a different identity with a particular name variant (eg a married name, legal change of name, etc).

Names are not just about software. They're about people.
And this is where NicNames comes in.

--Rebecca Parker, NicNames Subject Matter Expert


UPDATE: The article has now been published, and is available here.

30 January 2009

Draft specification for NicNames application

The following document is a rough attempt to describe the way that the NicNames application might work, from my perspective.

From a high level, it describes:
  • Data model
  • User interface
  • Query service
  • Bulk import or 'harvesting'
Some questions are as yet unanswered. All comments are welcome.

>> Nicnames spec 0.1.1 20090130 (PDF, 248KB)

23 January 2009

The logic of persistent identifiers

“Authority control is the process of grouping multiple terms for the same entity into a single record for the purposes of disambiguation and collocation”1 and has a long history in the library world. But, because of that long history, some practices have accumulated which are not appropriate in a digital context.

In particular, the authorised (form of name) heading concept is an artefact of card catalogues, which was used as a mechanism to collocate entries for all works (or more precisely FRBR group 1 entities: Work, Expression, Manifestation, Item) by a named entity (more precisely a FRBR group 2 entity: a person or a corporate body), including those created under variant forms of name. See and See also entries (tracings) were then used to refer to the main entry authorised form.

The authorised form of name used in this way also, confusingly, concatenates a particular name form with collocation.

In a digital environment, we don’t need an authorised form of name because any form of name can be used to link to all works by the named entity. But, we do need some form of persistent identifier (PID) to identify the group 2 entities to which the variant names and group 1 entities can be linked.

That PID could be in the form of a URI which links to information about the group 2 entity, but it should be noted that that again concatenates two logically distinct functions; that is, (a) providing a linking function between group 2 (named) entities, their names and works (group 1 entities) and (b) providing information about the group two entity.

In a local system, the PID could be as simple as any non-meaningful (that is, not linked to or derived from any data in the record) (most likely numeric) string. As long as suitable policies2, such as those developed by the PILIN project, are in place and resources provided to implement the policies, then such PIDs will work for local purposes.

However, in a situation where there is a need to identify a group 2 entity beyond the local system, as is the case with the NicNames Project, a higher level PID is required. This is because we are now trying to link namedEntityA@Swin with namedEntityA@UNSW with namedEntityA@UNew. That is, an Australian researcher may have works deposited at any of a number of Australian research repositories and we want to be able to identify both the works and any authority data not held locally.

This is where an educational or national name identification service, such as the National Library of Australia's People Australia service, could play an important role.

If the first repository to generate authority data for a researcher submits it to People Australia, a PID could be assigned for that researcher which other repositories could then use when incorporating the authority data into their own systems. If works (group 1 entities) were also linked to the authority data, then, in principle, it should be possible to easily find all works by that researcher, in whatever repository they happen to reside.

The implications of this logic are that each repository creates authority data for new researchers as they deposit work into the repository. That authority data, including any attached works and any relevant entity attributes, is submitted to People Australia, who assign a PID which is later added to the local record.

When the researcher changes institution and deposits material in that institution’s repository, the authority data is retrieved from People Australia and incorporated into the local system complete with the already assigned PID. The new work and any further attributes, such as the new affiliation, is then added to the authority data and resubmitted to People Australia.

It should then be possible, in principle, to incorporate a metasearching component into repository searches which will query People Australia to retrieve all works by a given researcher.

References

  1. Norrish, Jamie (2007). EATS: an entity authority tool set. http://researcharchive.vuw.ac.nz/handle/10063/220
  2. Nicholas, Nick, Ward, Nigel and Blinco, Kerry (2009). A policy checklist for enabling persistence of identifiers. D-Lib Magazine. 15 (1/2). http://www.dlib.org/dlib/january09/nicholas/01nicholas.html

09 January 2009

Progress Report January 2009

Happy New Year!

Most of the team is back at work this week after a break over the Christmas New Year period and pressing on with the project. Our Business Analyst, Damien Ingle, started just before Christmas and has spent some time with both Swinburne and Newcastle staff gathering information to feed into stakeholder requirements and institutional analyses.

Rebecca Parker, our Subject Matter Expert (because she actually works with the Swinburne Research Bank repository) is putting together a set of researcher personal name use cases and our Programmer, Tom Rutter, has begun development of tools to work with personal names.

Now that we have a better idea of costs, it seems as though there may be enough funding to take the project beyond the original March deadline or to increase the resources devoted to it over a shorter time frame. So, while we are still not being overly ambitious, it may be that we can do a little more than we had originally thought. Fingers crossed.

05 December 2008

“Names touch everything…” here too!

Derek Whitehead drew my attention to this entry in the hangingtogether blog. The idea of a Cooperative Identities Hub as a more broadly based name authority file suitable for use by a wide range of data custodians (libraries, archives, museums, repositories, aggregators, publishers) certainly fits well with what this project hopes to do.

It has occurred to us that, because fresh new researchers frequently publish for the first time as a co-author of a paper while still a graduate student, university research repositories will often be the first to see the researcher's name and, as a consequence, be the ones to do the original authority work (and will also be in the best position to gather researcher persona attribute data).

So, if there is data which might not be of immediate interest to a repository manager, but is nevertheless easily accessible and likely to be of use to other institutions later, then we probably should gather it and pass it on.

Progress Report December 2008

It has taken a while to make appointments, but the project is finally underway, albeit in a somewhat cart before horse fashion – the stakeholder requirements analysis will now be done in parallel with at least schema design and some preliminary investigation of name matching and distinguishing algorithms, all of which will be happening through December and January.

We are currently looking at how well EAC-CPF (Encoded Archival Context – Corporate bodies, Persons and Families) might meet our needs after doing a rough comparison of FRAD, RDA, DC, MADS, EAC, FOAF and VCARD against a set of possible attributes and relationships that might be readily available to ARROW repository managers. Rough because most of these are in a state of flux and because our learning time is limited.

EAC-CPF is attractive because it is a rich namespace structured to represent relationships as well as entities and because People Australia is proposing to use it. Once we learn how to code EAC, the next step will be to try to test it by generating some use cases and attempting to render them in EAC.

At this stage, we are not intending to go to the next step of defining an application profile and wrapping our EAC and whatever other vocabulary elements we might need into an RDF structure. It would be a desirable outcome, but we will probably not have time to get that far.

On the application side, we are going to have a look at how the BibApp application might fit into what we are doing – it does seem to have some effective mechanisms for disambiguating and distinguishing names that seem to overlap with what we are doing.

07 November 2008

Comments

We have had a few comments, just not through the blog! They fall into a few categories as follows;

1. this is a project whose time has come;

This has come from a number of quarters; from the library world and from people interested in learning object repositories as well as those running research repositories. Authority control has been around for a long time, but it seems the new context of digital repositories has led to the issue bubbling up for a rethink.

2. the timeline is very short;

Yes, indeed. Particularly with Christmas and the New Year in the middle, we recognise that we may have to cut our cloth to fit the timeline (sorry for the mixed metaphor). There may be a some flexibility with the March deadline, but we will see.

We are also focusing solely on personal names as an attempt to keep it as simple as possible.

3. will there be an operational relationship with People Australia?

We have already had a discussion about how this project might interact with People Australia and, without wanting to prejudice the outcome of the project, it does seem only sensible to build and use People Australia as the authority file for Australian researchers.

4. identifiers and vocabularies;

We have had mention of URL/URIs, People Australia persistent identifiers, ISNI, ISADN, OpenID and various commercial researcher numbers as identifiers. There are also many developments to do with schemas, DTDs, vocabularies, etc and sorting something reasonable out of all that will be a core part of the project.