It has been a little quiet over here lately. At the moment, I'm writing a revised literature review on names. The JISC landscape review was a great summary of the names environment in June 2008, but it has been a busy year in our area and we'd like to share some of the more interesting new literature with you as well.
The JISC Names Project released its Phase One final report in July. This partnership between the University of Manchester and the British Library is building a national authority file for the whole of the UK. It's an ambitious task, and we salute them for it. They've already released a prototype of their web service; you can have a play here (I did).
Also in July, Peter Sefton from the CAIRSS Project wrote a blog post about how a NicNames web service might interact with People Australia (I particularly liked the picture of the happy repository manager and hope that will be me soon ...)
The scholarly literature is also reflecting some very interesting developments. I summarised Dorothea Salo's paper on the absence of name authority control in institutional repositories in an earlier post. It's exciting to see that the big journals are starting to weigh in on the action, too. If 2008 will be remembered as the year The Lancet published an article about two clinical researchers who had decided to become numbers, 2009 was the year Science started to care about names. Both articles discussed the merits of the ResearcherID product from Thomson Reuters, which they described as 'ready and available now'. (I'm not so sure about that ...)
And finally, a few weeks ago, Ernesto Ruelas Inzunza from Dartmouth published what looks like a very interesting paper, 'Writing and citing 'international' names'. As soon as I can get my hands on a copy, I'll let you know all about it.
Interested in more literature about names? Feel free to contact Rebecca.
08 October 2009
18 September 2009
NicNames Project Plan
The draft project plan has been finalised to reflect any changes to the project outcomes as the project has progressed and the requirements have been refined, and to reflect the new completion dates of the project. This has been released as The ARROW NicNames Project Project Plan Version 1.1.
The draft project plan has been finalised to reflect any changes to the project outcomes as the project has progressed and the requirements have been refined, and to reflect the new completion dates of the project. This has been released as The ARROW NicNames Project Project Plan Version 1.1.
10 August 2009
Of Beagles and men: a cautionary tale of Charles Darwins
Are you one of the people who finds it difficult to see the problem we're trying to address with NicNames? C'mon, don't be shy ... I know you're out there.
Sure, researchers at our institutions publish work under a variety of name variants. But we know who they (really) are, so why not slap the same name on all their papers so they're easier for our users to find? It's how libraries do it.
Susan Stone from Intergalactic University might produce research as Sue Stone, S. Stone, S. G. Stone, S. Gilligan Stone and write horror novels as Susan Sly Stallone, but we could ignore all that untidiness and bring everything together under a single authoritative name. It would be so much neater.
Except that it's not. And I'm going to prove it.
Today, while I was adding some new records to Swinburne Research Bank, I noticed a familiar author name come up on a paper. Let's say the name was Charles Darwin so I don't have to give away any real names. And the artist to be known as Charles Darwin is familiar to me because he works at Swinburne's humanities faculty and has contributed lots of his work already (under the name Charles Darwin).
But the Charles Darwin I saw today was not Charles Darwin, Swinburne atheist. It was Charles Darwin from MIT, co-author of Emily Pankhurst, Swinburne professor of innovation.
Oh dear. So now what?
Here's how Swinburne Research Bank's author browse handles two completely different Charles Darwins spelled exactly the same way.

Would your repository be any different? I doubt it.
If all we have is their names, two Charles Darwins will always appear as the same person, both in search results and browse menus. And in a display like this, there's no way for our users tell the difference between them.
We need to be able to record and present defining features---such as fields of research and institutional affiliations---to be able to make sense of these names. It's all about context. And this is where NicNames will come in handy.
In the meantime, we at Swinburne Research Bank have a problem, and we won't be alone. I bet there's more than one Susan Smith out there. How are you going to answer when she knocks at your door?
--Rebecca Parker, NicNames Subject Matter Expert
Note: There are a few Easter eggs in the browse table. Be sure to let me know in the Comments if you find them.
Sure, researchers at our institutions publish work under a variety of name variants. But we know who they (really) are, so why not slap the same name on all their papers so they're easier for our users to find? It's how libraries do it.
Susan Stone from Intergalactic University might produce research as Sue Stone, S. Stone, S. G. Stone, S. Gilligan Stone and write horror novels as Susan Sly Stallone, but we could ignore all that untidiness and bring everything together under a single authoritative name. It would be so much neater.
Except that it's not. And I'm going to prove it.
Today, while I was adding some new records to Swinburne Research Bank, I noticed a familiar author name come up on a paper. Let's say the name was Charles Darwin so I don't have to give away any real names. And the artist to be known as Charles Darwin is familiar to me because he works at Swinburne's humanities faculty and has contributed lots of his work already (under the name Charles Darwin).
But the Charles Darwin I saw today was not Charles Darwin, Swinburne atheist. It was Charles Darwin from MIT, co-author of Emily Pankhurst, Swinburne professor of innovation.
Oh dear. So now what?
Here's how Swinburne Research Bank's author browse handles two completely different Charles Darwins spelled exactly the same way.
Would your repository be any different? I doubt it.
If all we have is their names, two Charles Darwins will always appear as the same person, both in search results and browse menus. And in a display like this, there's no way for our users tell the difference between them.
We need to be able to record and present defining features---such as fields of research and institutional affiliations---to be able to make sense of these names. It's all about context. And this is where NicNames will come in handy.
In the meantime, we at Swinburne Research Bank have a problem, and we won't be alone. I bet there's more than one Susan Smith out there. How are you going to answer when she knocks at your door?
--Rebecca Parker, NicNames Subject Matter Expert
Note: There are a few Easter eggs in the browse table. Be sure to let me know in the Comments if you find them.
12 June 2009
NicNames as a webservice
The current perspective is of NicNames as a webservice and as such supplying an API set allowing for submission of names and the extraction of names and associated metadata. The standard set of DB maintenance methods (Add, Edit, Reports etc) is supplied together with extensions that enable the tie of the service to a web application (e.g. Valet) supplying resolution of names and metadata usable to populate application fields. As such there are two forms of access, via direct access to NicNames or via calls to NicNames through such as Valet or repository management applications (E.g. VITAL) - the latter requiring customisation to integrate NicNames with the application.
Since the Valet environment covers self-submission of data to the repository there are (a) some restrictions on access to NicName methods and (b) requirements for repository staff to later validate name entries when input from a Valet environment (The X-Files element - trust no one!). The security also covers harvesting attempts where NicNames data can be extracted (OAI-PMH format) covering the name(s) and a defined set/subset of existing data - the definition as set down by the associated repository manager(s) and so limiting access to such as staff IDs etc that could be exploited as part of identity theft etc. but are essential for disambiguation methods.
As a webservice NicNames is dominantly passive; population of the DB with existing names from the repository done in the form of repository staff extracting data into an XML file and submission of that file to NicNames. Once data has been added, any additional data defined when setting up the NicNames schema is required to be added to the system. The amount of data is determined by the repository manager or else one can accept the default schema that is comprehensive in its coverage of data usable to aid in the disambiguation process.
The simplicity of the approach i.e. a webservice that enables the 'transcending' of current authority control data (as MARC format etc), hides the complexity in use of that data to disambiguate names where the essential feature of NicNames is in the speed and precision achieved in the disambiguation focus. The additional benefits include access to additional metadata beyond their use in resolving ambiguities.
Since the Valet environment covers self-submission of data to the repository there are (a) some restrictions on access to NicName methods and (b) requirements for repository staff to later validate name entries when input from a Valet environment (The X-Files element - trust no one!). The security also covers harvesting attempts where NicNames data can be extracted (OAI-PMH format) covering the name(s) and a defined set/subset of existing data - the definition as set down by the associated repository manager(s) and so limiting access to such as staff IDs etc that could be exploited as part of identity theft etc. but are essential for disambiguation methods.
As a webservice NicNames is dominantly passive; population of the DB with existing names from the repository done in the form of repository staff extracting data into an XML file and submission of that file to NicNames. Once data has been added, any additional data defined when setting up the NicNames schema is required to be added to the system. The amount of data is determined by the repository manager or else one can accept the default schema that is comprehensive in its coverage of data usable to aid in the disambiguation process.
The simplicity of the approach i.e. a webservice that enables the 'transcending' of current authority control data (as MARC format etc), hides the complexity in use of that data to disambiguate names where the essential feature of NicNames is in the speed and precision achieved in the disambiguation focus. The additional benefits include access to additional metadata beyond their use in resolving ambiguities.
NicNames & Disambiguation - moving into higher dimensions
In considering disambiguation issues -
"It is a lot like the difference between solids, where the atoms are locked into place, and fluids, where the atoms tumble over one another at random. But right in between the two extremes, at a kind of abstract phase transition called the edge of chaos, you also find complexity: a class of behaviors in which the components of the system never quite lock into place, yet never quite dissolve into turbulence, either. These are the systems that are both stable enough to store information, and yet evanescent enough to transmit it. These are the systems that can be organized to perform complex computations, to react to the world, to be spontaneous, adaptive, and alive." M. Mitchell Waldrop, from Complexity [p. 293]
We are dealing with an area of mathematics called 'hinge theory':
Plastic Hinge Theory covers http://en.wikipedia.org/wiki/Plastic_hinge
The emphasis is on the "plastic rotation [deformation] of an otherwise rigid column connection" - for us the 'rigid column connection' is the key, the identifier, for people in the form of a list of names. As such we are focused on the static/dynamic, the solid/fluid border of identity.
The use of Baysian probabilities introduces a partials perspective as we try to identify the 'whole' but is still focused on a one-dimensional POV and this issue is under consideration whilst at the same time being focused on the more practical implementation of a refined one-dimensional POV methodology; refinement in the form of the metadata schema of NicNames allowing for extended analysis of name associations and so extending current authority control material used in the disambiguation process.
"It is a lot like the difference between solids, where the atoms are locked into place, and fluids, where the atoms tumble over one another at random. But right in between the two extremes, at a kind of abstract phase transition called the edge of chaos, you also find complexity: a class of behaviors in which the components of the system never quite lock into place, yet never quite dissolve into turbulence, either. These are the systems that are both stable enough to store information, and yet evanescent enough to transmit it. These are the systems that can be organized to perform complex computations, to react to the world, to be spontaneous, adaptive, and alive." M. Mitchell Waldrop, from Complexity [p. 293]
We are dealing with an area of mathematics called 'hinge theory':
Plastic Hinge Theory covers http://en.wikipedia.org/wiki/Plastic_hinge
The emphasis is on the "plastic rotation [deformation] of an otherwise rigid column connection" - for us the 'rigid column connection' is the key, the identifier, for people in the form of a list of names. As such we are focused on the static/dynamic, the solid/fluid border of identity.
The use of Baysian probabilities introduces a partials perspective as we try to identify the 'whole' but is still focused on a one-dimensional POV and this issue is under consideration whilst at the same time being focused on the more practical implementation of a refined one-dimensional POV methodology; refinement in the form of the metadata schema of NicNames allowing for extended analysis of name associations and so extending current authority control material used in the disambiguation process.
29 May 2009
‘Monthly’ Progress Report for May 2009
I see we haven’t posted an entry here since early March and no progress report since January(!). The excuse is that we have actually been getting on with it.
Because of the earlier delays, the NicNames project has now been extended until the end of October. This will give us time to complete the tasks we have set for ourselves and, hopefully, produce a set of applications and documentation which will not only describe the problem, but provide Institutional Repository managers with a way to deal with it.
We have a Business Requirements Specification and will soon have systems and application specifications. By the end of June we should have a working application and by about the end of August a completed usability study and a guidelines toolkit.
These products will then be implemented and tested in the partner institutions, Swinburne University of Technology, University of Newcastle and University of New South Wales, during September and October and once we are happy with the way it works, it will all be released to the wider IR community.
Stay tuned.
Because of the earlier delays, the NicNames project has now been extended until the end of October. This will give us time to complete the tasks we have set for ourselves and, hopefully, produce a set of applications and documentation which will not only describe the problem, but provide Institutional Repository managers with a way to deal with it.
We have a Business Requirements Specification and will soon have systems and application specifications. By the end of June we should have a working application and by about the end of August a completed usability study and a guidelines toolkit.
These products will then be implemented and tested in the partner institutions, Swinburne University of Technology, University of Newcastle and University of New South Wales, during September and October and once we are happy with the way it works, it will all be released to the wider IR community.
Stay tuned.
Subscribe to:
Posts (Atom)