As a last event of the academic year, Altergraphy attended the five days summer school organized by DISTAM, one of the labelled consortiums of the research infrastructure (IR) Huma-Num in the gorgeous convent of La Tourette by Le Corbusier. Since 2022, DISTAM identifies challenges in the digital transition for research in area studies and provides technical advice for scholars and projects dealing with corpora in non-latin scripts. Apart from the several plenary sessions where common issues and new research projects were discussed, the summer school was tripartite, with workshops on text encoding, web archive and image annotation.


Plenaries
A first roundtable gathering a few area specialists in the social sciences and humanities suggested that for researchers with our profiles, opting for descriptive languages and data structures readable and sharable by machines is one way to overcome discplinary weaknesses and language boundaries.
The plenary by Emmanuelle Morlock centered on the sustainability of data, the relevance of stored data (for their scientific or patrimonial value), good practices in sharing source codes and metadata schemes (in open protocoles, readable formats, controlled vocabularies), examples of solid data management plans and forms of publication in the process of structuring corpora in line with these standards. After a few months of scratching our heads within Altergraphy and in dialogue with other projects having shared their metadata schemes or conceptual maps with us, it was good to learn how this form of engagement requiring the invisible time and effort to train oneself and a team, can be valorized. Emmanuelle also evoked the case of retrospective ordering of previous research project or archiving a lifetime’s corpora for retiring scholars, two configurations which are represented in our project, with the thousands of images by Lia Wei and Zhang Qiang since 2009, and the 1500 images recently scanned by Sakata Gensho spanning the years 1980-2010.
A plenary on identifiers and the nature of metadata by Michel Jacobson, was followed by a talk on interoperability by Benjamin Guichard. The latter, scientific director at BULAC, introduced the notion of Knowledge Organization System (KOS), with three steps: from controlled vocabularies, to the thesaurus (a hierarchized list of interconnected concepts) and the ontology (thesaurus of thesaurus). He gave the interesting example of PACTOLS, which thesaurus is also available on Opentheso, the environment we are about to list our current lists of descriptors in. PACTOLS was initiated by archaeologists, so that one finds entities such as “mason’s mark” (an immaterial entity>symbolic object>mark). Benjamin also provided references for ontologies, higher up in the hierarchy of concepts, that provide guidelines for the creation of thesauri, as well as platforms for sharing terminologies accross disciplines and languages in the French-speaking environment, and their UK counterparts, and tools for aligning metadata such as Mix’n’Match. As an actor from the library world, Benjamin also stressed the importance of existing structures in France such as IdRef, BnF data, Calame and Sudoc.
Finally, a last plenary on the concept of “annotation” was given by Emmanuelle Morlock. Tracing back to the first consensus on annotating conventions between epigraphists – indeed the perfect example of unmoveable text, difficult to share, and that must be enabled to travel and be exported into other formats! – she located in the 1960s the separation between textual form and content, and the development of generic systems for the exchange of information about text. She defined the marked up text has having been analyzed in its visual aspects (logical and physical structures such as titles and colophons), indexable elements (people, places, dates) and analytical categories proper to the research question (phenomena of interest). She numbered the possible tools for text annotation and markup from manual to semi-automatic or form-based interfaces.
Image annotation workshop
We were looking forward to learn more about how to embed comments, notes, citations and links to authorities within images of the landscape, landmarks, infrastructures, inscriptions or rubbings, down to the single character. We have been drifting away from the idea of inserting all the contextual information in the encoded texts of the inscriptions, since this would need a heavy “physical description header” for only a few characters’ text, or in over-detailed Excel sheets, and we were wondering how to deal with the different nested scales of visual representation. Mapping also requires reconsideration, since the subjective spatial narratives and representations we have gathered from the late imperial period to the 1980s, together with current brainstorming about alternative ways to document our own subjective experience of the sites: the “mapbox + physiographic or historical map” association now prooves unsatisfactory. When subscribing to this workshop, we were hoping to become able to annotate photographs and images directly, without necessarily converting the visual/material into textual or spatial information.
In preparation to the workshop, we were encouraged to think about the composition of our image corpus, its provenances and involved rights. In our case, these are mostly field photographs collected between 2009 and 2019 by Lia Wei and Zhang Qiang, at different seasons and times of the day, havinng served as illustrations to some publications. They document the evolution of the sites in an unsystematic way. Photographs by the same duo since 2019 have become more systematic, and are classified by stele or site, with drone views of the landscape and character-by-character shots, almost systematically with a LED light. However, uneven formats, colorimetry, and blurred photographs still represent major flaws of this photographic record. To this personal record that still needs to go through a selection and the last addition of our fieldwork in October 2024, since the beginning of the project in October 2023, the Altergraphy corpus has grown with the addition of Gensho Sakata’s photographic archive, which he recently scanned. Some photographs taken by Taneya between the years 1982 and 2004 could be added to this corpus. Photographs from local archives may be shared by the county level offices of Pingdu and Laizhou during our next survey. Photographs of rubbings in personal, private and public collections constitute the other half of our image corpus.
A presentation by Vincent Thérouin, member of the Callfront project, introduced us to the way the project approached calligraphy, from data collection in tabular form to content display through Omeka, a free content management system developed by the Roy Rosenzweig Center for History for managing the back-end construction of a database and its front-end online interface. Contributors to Callfront for example, are invited to add examples from their corpora illustrating forms of writing that have been previously listed, structured into a thesaurus and approved by a community of researchers, with the help of a calligrpahy practitioner. The writing samples rather than the whole manuscript are first listed in a tabular form, so that neither the text, nor the line of text or even the letter (letters are the second level of classification in the proposed thesaurus, stroke morphology coming first) is the unit of analysis, but rather certain types of strokes. These writing habits are described in geometric terms (based on basic geometric shapes and angles, the direction of the movement, etc) in an attempt to establish a traceology of Arabic calligraphy. While the terms are not necessarily related with an emic terminology from treatises etc, they are sometimes derived from the practioners’ vocabulary, related to the human or animal body (ex: head, body and tail of character). The home-made thesaurus by Callfront (about 50 additions to Dublin Core) can be hosted as customized vocabulary in Omeka S (a list called customvocab), and is stored in a hierarchized form, with full definitions and more elaborate descriptors in Opentheso (for example, 21 descriptors for “ink”), with accompanying images. In Omeka S, the geographic areas covered by Callfront will appear as collections, stressing the “geopolitical” dimension of Arabic calligraphy in the investigated areas.
In our case, the entry points to the Altergraphy corpus could be the two “calligraphers” or stylistic paradigms, ZDZ and SADY, with the attached calligraphic markers as points of entry. If classified by types of data, the four entries could be: (1) the landscapes or portraits of the mountains (scale of the site, visitors’ comments and heritagization); (2) medieval epigraphy (at the scale of the character, with variants, and transfer technique from calligraphy to engravings); (3) rubbings (and their circulation, versions or reproductions innto other media); (4) (history of) calligraphy (SADY and ZDZ as paradigms, commentaries from treatises and reference works for comparison).
Chahan Vidal-Gorène, digital paleographer currently working on damaged armenian inscriptions and responsible for the masters programme at the Ecole des Chartes, introduced us to two opensource tools for image annotation: Labelstudio and Yolo. Computer-assisted vision aims at analyzing large corpora, identify relevant views and quantifying the extracted information, but it can also augment our peception of an artifact with the addition of embedded notes, or automatize a repetitive task. Today, a large number of data is not necessary anymore to create a model specialized in a task (fine tuning), a sample of about 50 images is enough. The sample should be quite homogenous, of good quality, with homogenous and constant annotations, with all various forms represented and all categories represented in sufficient amount (if not, the categories should be less refined). Chahan then detailed the 4+1 existing approaches in computer vision : (1) classification+localization (adding a label); (2) object detection (target object is isolated in a box like in Transkribus); (3) semantic segmentation (terms attributed to the scale of the pixel like in eScriptorium) ;(4) instance segmentation (mixte like in Calfa); and (5) keypoint detection.
Michèle Galdemar, documentalist in the dh service of INHA, whom we previously met in a first very inspiring and helpful contact with project Callfront, introduced us to the use of Tropy, an open source software for the local organisation of images, but that can be moved from one machine to another together with it source folder containing all images, and based on Dublin core fields that can be then edited to one’s purposes. Tropy allows one to create collections through tagging, and to re-name files in batches (still to do in our case with a consistent system of image IDs, to avoid duplicate names). We will opt for this software for the post-fieldwork phase of our project.
Bruno Morandière the father of Persée, then proposed us a bridge from Tropy to Nakala – the repository maintained by Huma-Num – and Omeka to test image annotation locally in Tropy, import the annotations as metadata into Nakala and try it out on Omeka, by ourselves. There is in fact a OmekaS plug-in for Tropy, but this chain was proposed to avoid duplicating the annotated zones of the images as supplementary images in Omeka, and mixing the raw and edited versions of the images in Omeka, without an additional permanent storage location. Tropy, Nakala and OmekaS are all three structured in Dublin Core and adopt the IIIF manifest.
This part of the workshop was probably the most directly useful for the Altergraphy project, where two interns, Francesca and Paula, are in fact already using Tropy for their own research. While we will classify the photographs from our survey, from the Sakata archiven and from the rubbing collections in Tropy, annotating landmarks and markers of (calli)graphic variation, we will also continue to work with the google sheets for landmarks, inscriptions, rubbings and reception. Another important encounter during this workshop was with Mercedes Volait, director of the CNRS laboratory InVisu, which hosts visual corpora that wish to propose an alternative kind of scientific publications, image-based, where the reader would be travelling “into the image” with tags, controlled vocabularies and storylines… right what we have been dreaming about!
The last intervention was by Dominique Roux, father of Métopes, who for 20 years worked on harmonizing the French world of edition to adopt TEI publishing standards with his team of four at the University of Caen. It was incredible to listen to him, explaining the evolution of the editor’s craft from metal plates to Indesign and away from it, with stages of impovershment where the visual logic of layout became an obstacle to the structural description of text. Métopes invests a lot of time thinking solutions for very specific publication projects, and ways for authors and editors to work with XML without necessarily coding directly, by producing publishable content from an adapted version of word that is then translated into TEI or more specific forms of TEI such as Epidoc and helps putting order. They are currently editing an epigraphy catalogue for various publication formats, always with the images in Nakala and, while images are for the time being still an external resource, they are working towards the possibility of integrating annotated images in the editorial workflow.
OpenEdition suggests that you cite this post as follows:
Lia Wei (July 25, 2024). Image annotation workshop with DISTAM (8-12 July 2024). ALTERGRAPHY. Retrieved September 12, 2026 from https://doi.org/10.58079/12enx