File lables the accession slips of a collection of snuff spoons in the African Collection at the British Museum.
The story which broke in late August, about over 1500 objects going missing from the collection at the British Museum, has had a lot of press coverage. In the UK media at least, it has prompted all sorts of hand-wringing over how something so unexpected could happen at one of the world’s largest, oldest and most illustrious institutions, which supposedly set the benchmark for security. Heads have rolled – one Keeper has been dismissed, and perhaps more significantly, the director, Hartwig Fischer, resigned. And a few months after the losses were made public, the museum announced a raft of security measures, including the planned digitisation of their entire collection (which will probably cost over 12 million pounds)
But while these kinds of stories often make a big splash in the media, to many people who work in museums, and those of us who think of ourselves as ‘British Museum observers’, the only real question is why it took so long for these losses to be discovered and made public.
Digitisation in large museums like the BM is often the best (and sometimes only) way for curators and other staff to get a handle on what their holdings actually are. This is especially true in museums where most of the collection is stored offsite. During my PhD when I was studying exactly how the process took place at the BM, I spoke to a retired keeper who had been closely involved in the digitisation processes. He was hilariously frank about the state of the collection in the 1970s – explaining that several collections were a mess, “real basket cases” as he termed them, and the order from on high was “Sort yourselves out, or else!”. Items were mislabelled, lost or miscatalogued. Record keeping was haphazard at best, having evolved over a couple of centuries from hand-written ledgers and access books, which even then were not guaranteed to be accurate. Cross-referencing items in collections that large was impossible, and even with the best will in the world, it was almost guaranteed that stuff would go missing because it was so hard to keep track of. (Two colleagues and I wrote about this in a paper on how memory functioned at the BM, if you’re interested)
Photo of a page old Accession Book from the British Museum African collection, showing notes by various generations of curators, as they tried to catalogue and organise the materials.
30 years later, when I was observing catalogue digitisation, they still weren’t finished with the process, but some order had been brought to the collections. Many of the records, particularly in the Prints and Drawings collections were extraordinary pieces of research in themselves, containing essay-long notes on the objects. The Museums’s Collection Online tool was being continually updated and improved, and there were promises of better technical backends to create links between items. There were also grand plans for Research Space, the museums linked data knowledge graph, which would provide new ways to access the collection, using semantic technology to create linkages and associations between collections and across the institutions, generating new insights into the materials.
Over the years this vision has been eroded away to almost nothing. Accessing the collection via the SPARQL endpoint was almost impossible from the start. Staff working on the database and other digital experiments were made redundant. Collection online presents wonderful high-resolution images and some catalogue data for about 2 million extraordinary objects, but this is less than half of the collection, and the research uses are limited. Research Space is a hermetically sealed digital sandbox, which is far beyond the reach of the average user. And at the end of the day, items are still being lost, misplaced and as we now know, stolen, because their records are patchy, incomplete or non-existent.
Good museum cataloging practice isn’t just an exercise in knowledge management. It’s a security issue, and any institution which chooses not to treat collection management systems in such a way, risks having to make expensive corrections in the future.
Sometimes the process of just trying to get organised ends up illustrating the objectives of research as clearly as the actual research itself. In this case, my initial intention to get this blog up and running, ended up reminding me that managing large collections of digital assets in cultural heritage collections can be complicated, and that sometimes things end up mis-classified, and in pretty odd places. Because, essentially, computers are not infallible. Who knew??
Getting started…
Scholarly blogs are one of my favourite formats for reading about research, and one which I am trying to get better at writing myself. I like the slightly less formal tone I can adopt as a writer, and the way in which they enable both researchers and readers to share an idea, an experience or a step in a research process, in what feels like a lower-stakes context than a journal article or conference paper. That’s not to say that I don’t think scholarly blogs should be rigorous or that they should maintain the same standards of citation that we expect of other scholarly communications. But as someone who spent 10 years as a journalist before coming back to academia, I still feel a lot more comfortable drafting and publishing a blog post than I do working my way through the multiple stages of preparing an article for publication. So a few weeks ago, I decided to pay some attention to this blog, beginning with making it look good.
Finding the right image to use as a main illustration was important. I wanted something that would convey the themes of the project – something that encapsulated the practices of classification, arranging and ordering which are required in museum documentation, and that also spoke to the problem of abundance – of objects and of records. Abundance (or volume) is one of the major areas of enquiry in TSLIM; one of the questions I am looking at in my research is how to manage culturally and historically difficult data (such as provenance information) in museum documentation, at scale. Computers are very good at managing volume, but they can’t make the same judgements as humans about whether it is appropriate to make some data public or not. However, we are increasingly relying on automated systems to manage the huge amounts of heritage data out there, and sometimes that can mean that things wind up in places where they should not be.
Vague search
I enjoy searching for images because image metadata is a different to metadata which describes texts. The logic is different – titles are often missing, and the metadata which describes the content of an image may be more subjective, the vocabulary less controlled. In my experience, it is always better to approach an image search with no clear idea of what you want, because you are likely to be disappointed. Rather, the serendipitous approach is often more fruitful – having a dig, seeing what you come up with often yields more fruitful results. My search for a good header image sent me to my usual sources for historically interesting, openly licensed images from various heritage collections online. I usually find what I need (and more) by searching through repositories and aggregations like Europeana, various museum websites, and the Flickr Commons. This last one is a particular favourite of mine – I think the navigation and search interface are really good and clear, I love their approach to crowd annotation and the range of cultural heritage institutions who have added material to the Commons is great. It went quiet for long time, and I was very happy to hear earlier this year that George Oates, a museum and digital heritage professional whose work I have admired for a long time, has been asked to help revitalise it and is reporting on progress.
This time I ended up at the Wellcome Collection – another excellent source of images. The Wellcome is a museum and library in London devoted to the subject of health. Their digital collections of images, books and museum objects are incredibly rich, and they have been innovating in digital asset management for a long time. When I was an MA student almost 10 years ago, we used the Wellcome’s digitisation work as a case study, and visited their digital teams onsite to learn about how they managed their metadata and other digital assets. After searching with rather vague keywords, such as ‘collection’ and ‘landscape layout’ and using their ‘search for similar images’ feature, which uses machine learning to find other images which have similar shapes and structural features based on various visual similarities across all images in the collection, I found myself on this page in the screenshot below:
Screenshot of search result from Wellcome Collection, captured 13 July 2021
The third image from the left, in the top row looked like something that might suit my needs. Collections of many small things, a range of different types of objects, organised by some kind of principle, and set up for display – it seemed perfect. Free to reuse, thanks to the Wellcome’s use of Creative Commons licenses, all I needed to do was find out a bit more about the image, so that I could attribute it correctly, and link back to the original.
A larger version of the image, out of context of the collection it was grouped in.
Unexpected Results
But when I clicked through to find out what the source was, I was taken to the digitised version of a book called Isagoge breves prelucide ac uberime in anatomiam humani corporis. A communi medicorum academia usitatam, written by an Italian surgeon named Jacopo Berengario da Carpi in 1522. And while it was immediately clear that the other images in the search result were indeed all images from da Carpi’s book (and by the way, he sounds like he was quite a guy – notorious for treating Romans who had syphilis with doses of mercury and charging them a lot of money for doing so)1 the image of the pendants seemed like a fairly clear mistake. Too recent, not an illustration, and thematically off. But where did this image come from, and how could I find it’s correct attribition?
I should point out here that this mix-up is not because the Wellcome’s digital teams are not good at their job. It is just one of the catches of working with large collections of digital materials. Sometimes, things wind up in the wrong place, and it can be very difficult to find the mistakes, or to automatically correct them. Many times, these kinds of small glitches slip under the radar until a user, like me, who is usually looking for something else, finds it.
Next, I decided to do a reverse Google image search. This is a fairly blunt search tool, but several art historians I know use it for their research, as a starting point for more detailed and careful searches, and I figured that it might yield some results. I wasn’t disappointed – I got several hits, and they only served to deepen the mystery:
Screenshot of the Google Image Search results for the amulet image, captured 28 September 2021.
I found several results, also from the Wellcome, for the image. As you can see from the screenshot above, it has been classified with a range for different terms, including “Dissection, 16th Century”, “Uterus” and “Anatomy”. It has also found its way into Wikimedia Commons, with the title “Uterus, Berengarius, 1523”. Using that as a search term produced other illustrations from Carpi’s book, as well as the photo of the amulets, which leads me to think that somehow that image has been associated with the book, which is why it keeps turning up in the search results. This probably means that not only does that association need to be ‘broken’ but also that the image’s original metadata also needs to be located, and a new ‘association’ created, so that it doesn’t get totally lost in the system.
A mystery, as yet unsolved
I have mailed the digital team at the Wellcome, to highlight what I found, and ask them if they could tell me how these kinds of glitches happen, and how they manage them when they do discover them. I’ll update this post as and when I hear back from them. To be honest, I am a bit sad that this odd anomaly might get totally lost. I’ve made a bunch of screenshots, and created some permanent links in the Internet Archive, so that I can call up the mistakes again, if I need to. This strange case is just too good to lose, if and when the correction is made.
As for the header image – I found one. Also from the Wellcome. And guess what? If you search for this image, these are the ‘visually similar images’ you’ll be offered by the ML algorithm:
There are our friends the amulets, second from the left…
TSLIM is the research project blog for a 3-year project at the University of Vienna, funded by the REWIRE programme – a Marie Sklodowska Curie Actions COFUND project funded by the European Commission. TSLIM is a multidisciplinary project, based in the Digital Humanities group in the Department if History. It deals with topics which fall broadly under the themes of digital humanities, museum studies, data science, critical data studies and digital heritage studies. The objectives of the project are to examine the use of linked data in museums, and in ethnographic museums in particular. The digitisation of cultural heritage data is creating previously unimagined possibilities for humanities researchers, making more museum, archive and library materials available online than ever. Scholars are able to overcome many of the logistical challenges of using multiple primary sources, such as the fragmentation and geographic dispersal of collections. Many cultural heritage collections have embraced the semantic web as a mechanism for opening up their databases, allowing collections to become interconnected, and creating networks of highly complex, rich and heterogenous data. However, the highly contextualised nature of humanities sources and the need to make allowance for complexity and ambiguity when interpreting them, mean that the rigid data model of the semantic web does not always provide enough contextual data for the complex evidentiary reasoning that takes place in humanities research. At the same time, researchers using this data have to grapple with data of inconsistent quality and depth. One of the objectives of TSLIM is to develop an evaluation framework with a specific focus on the sustained linking of ethnographic materials, which will allow scholar to evaluate the completeness of the data they are using. At the same time, the research will also contribute to increasingly urgent discussions about how to manage this vast and growing body of data, by conducting a nuanced and rigorous study of both LOD creation and exploitation. In a context where searching for digital content is increasingly taken to mean ‘just Google it’, it is crucial for data producers and data consumers be able to find what they need. To do this, it is essential to devise new approaches to how the data is stored, categorised and made available. The blog will be used to track the progress of the project, announce any workshops, seminars and other collaborative activities which take place under the auspices of the project, and provide a space for regular, informal but considered writing and reflection. It is expected to be of interest to digital humanities scholars, data scientists, museum professionals and other humanities scholars.