Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

If only they had a decent Catalogue…

File lables the accession slips of a collection of snuff spoons in the African Collection at the British Museum.

The story which broke in late August, about over 1500 objects going missing from the collection at the British Museum, has had a lot of press coverage. In the UK media at least, it has prompted all sorts of hand-wringing over how something so unexpected could happen at one of the world’s largest, oldest and most illustrious institutions, which supposedly set the benchmark for security. Heads have rolled – one Keeper has been dismissed, and perhaps more significantly, the director, Hartwig Fischer, resigned. And a few months after the losses were made public, the museum announced a raft of security measures, including the planned digitisation of their entire collection (which will probably cost over 12 million pounds)

But while these kinds of stories often make a big splash in the media, to many people who work in museums, and those of us who think of ourselves as ‘British Museum observers’, the only real question is why it took so long for these losses to be discovered and made public.

Digitisation in large museums like the BM is often the best (and sometimes only) way for curators and other staff to get a handle on what their holdings actually are. This is especially true in museums where most of the collection is stored offsite. During my PhD when I was studying exactly how the process took place at the BM, I spoke to a retired keeper who had been closely involved in the digitisation processes. He was hilariously frank about the state of the collection in the 1970s – explaining that several collections were a mess, “real basket cases” as he termed them, and the order from on high was “Sort yourselves out, or else!”. Items were mislabelled, lost or miscatalogued. Record keeping was haphazard at best, having evolved over a couple of centuries from hand-written ledgers and access books, which even then were not guaranteed to be accurate. Cross-referencing items in collections that large was impossible, and even with the best will in the world, it was almost guaranteed that stuff would go missing because it was so hard to keep track of. (Two colleagues and I wrote about this in a paper on how memory functioned at the BM, if you’re interested)

Photo of a page old Accession Book from the British Museum African collection, showing notes by various generations of curators, as they tried to catalogue and organise the materials.

30 years later, when I was observing catalogue digitisation, they still weren’t finished with the process, but some order had been brought to the collections. Many of the records, particularly in the Prints and Drawings collections were extraordinary pieces of research in themselves, containing essay-long notes on the objects. The Museums’s Collection Online tool was being continually updated and improved, and there were promises of better technical backends to create links between items. There were also grand plans for Research Space, the museums linked data knowledge graph, which would provide new ways to access the collection, using semantic technology to create linkages and associations between collections and across the institutions, generating new insights into the materials.

Over the years this vision has been eroded away to almost nothing. Accessing the collection via the SPARQL endpoint was almost impossible from the start. Staff working on the database and other digital experiments were made redundant. Collection online presents wonderful high-resolution images and some catalogue data for about 2 million extraordinary objects, but this is less than half of the collection, and the research uses are limited. Research Space is a hermetically sealed digital sandbox, which is far beyond the reach of the average user. And at the end of the day, items are still being lost, misplaced and as we now know, stolen, because their records are patchy, incomplete or non-existent.

Good museum cataloging practice isn’t just an exercise in knowledge management. It’s a security issue, and any institution which chooses not to treat collection management systems in such a way, risks having to make expensive corrections in the future.

TSLIM

TSLIM is the research project blog for a 3-year project at the University of Vienna, funded by the REWIRE programme – a Marie Sklodowska Curie Actions COFUND project funded by the European Commission. TSLIM is a multidisciplinary project, based in the Digital Humanities group in the Department if History. It deals with topics which fall broadly under the themes of digital humanities, museum studies, data science, critical data studies and digital heritage studies. The objectives of the project are to examine the use of linked data in museums, and in ethnographic museums in particular. The digitisation of cultural heritage data is creating previously unimagined possibilities for humanities researchers, making more museum, archive and library materials available online than ever. Scholars are able to overcome many of the logistical challenges of using multiple primary sources, such as the fragmentation and geographic dispersal of collections. Many cultural heritage collections have embraced the semantic web as a mechanism for opening up their databases, allowing collections to become interconnected, and creating networks of highly complex, rich and heterogenous data. However, the highly contextualised nature of humanities sources and the need to make allowance for complexity and ambiguity when interpreting them, mean that the rigid data model of the semantic web does not always provide enough contextual data for the complex evidentiary reasoning that takes place in humanities research. At the same time, researchers using this data have to grapple with data of inconsistent quality and depth. One of the objectives of TSLIM is to develop an evaluation framework with a specific focus on the sustained linking of ethnographic materials, which will allow scholar to evaluate the completeness of the data they are using. At the same time, the research will also contribute to increasingly urgent discussions about how to manage this vast and growing body of data, by conducting a nuanced and rigorous study of both LOD creation and exploitation. In a context where searching for digital content is increasingly taken to mean ‘just Google it’, it is crucial for data producers and data consumers be able to find what they need. To do this, it is essential to devise new approaches to how the data is stored, categorised and made available. The blog will be used to track the progress of the project, announce any workshops, seminars and other collaborative activities which take place under the auspices of the project, and provide a space for regular, informal but considered writing and reflection. It is expected to be of interest to digital humanities scholars, data scientists, museum professionals and other humanities scholars.