Tag Archives: text

Transcribing, annotating and translating our textual corpus in Tacteo

Since the beginning of the project, we have tried different softwares and platforms to share and work collaboratively on our notes, transcriptions and translations, but we relied quite heavily on google docs and a shared google drive. After about a year of activity, and when getting familiar with Sharedocs, a platform provided by the Huma-Num research infrastructure in France, that we also used to sort and comment on images, we wished to turn to safer and more open environments.

When looking at the Huma-Num services, we discovered a plain text editing tool called Stylo. The inscriptions’ texts transcriptions that we had found in the existing publications, mostly by Wang Sili and Lai Fei (compiled in Lai Fei 2023), Sakata Gensho (Sakata 1984) and Robert Harrist (Harrist 2008), were saved as pdf documents, but Stylo could help us get rid of the formatting issues that we encountered in google docs or word. Stylo, however, requires to manage the notes and bibliography in separate documents, and for straightforward transcription we needed a tool that could both comprise image and text, allowed to share and edit versions collaboratively and could be exported as plain text or xml file.

Fortunately… such a tool exists: Tacteo, also proposed by the Huma-Num research infrastructure, allows a team to open a project, upload or import image files , and display the image next to a text editing box with different options (transcribe/visualize/comments/logs). One can redact a manuel with instructions for transcribers, define one’s own annotation scheme (based on TEI, but also with the possibility to add specific tags), and the transcriber can submit the text for one or several proofreadering steps.

We opted for a simplified workflow and annotation scheme in line with our research questions and objectives and the skills and composition of the team.

(1) the texts previously transcribed were copied and pasted into Tacteo, with the corresponding rubbings gathered by Anna Le Menach (when our own rubbings were not the best available version we based ourselves on versions available in public repositories such as the Berkeley library, which possesses a good number of rubbings of epigraphy attributed to Zheng Daozhao, from donated Japanese collections).

(2) a first character-to-character verification of the transcription is accompanied by the tagging of variants (the existing <g> tag in TEI) by Francesca Berdin. The variants are listed in a table in our relational database in the grist environment, and attributed an ID bearing the inscription ID, column and line number (e.g. TZS03_02_06 for the sixth character of the second column of the third inscription in Mount Tianzhu). In the grist table, a first classification of variants is proposed by the project, resulting from the various discussions and workshops led in the previous years with our partners in Heidelberg and Ghent. The attributes of the <g> tag are thus not shown in Tacteo, only the variant ID is. During this step, line breaks <lb> and <supplied> characters are also tagged (also extant in TEI, the latter being further developed in Epidoc), to account for the layout and state of preservation of the inscriptions.

(3) a second round of proofreading based on the translations is led by Zhang Rui, our research assistant who is responsible for the translation of the inscriptions, who adds to Tacteo a punctuated version of the text with no line breaks in a second <div>.

(4) In a third and last step of proofreading, this version is also tagged for authorities of person, date and place. This last step of “conceptual” tagging feeds the authorities tables of the relational database in grist and allow us to build stronger relationships between our entities.

Tacteo allows us to work with three text divisions (the translation counts as the third) and a two-step proofreading process, and a minimal set of tags. We encountered some technical issues in the day-to-day use of Tacteo, such as difficulties in uploading images, or issues when using the <div> tag, which seemed to duplicate some of the embedded tags.

Unfortunately, Omeka S can only display our transcriptions as image caption/metadata, which means that our annotation work will not be exploited in that environment. the text dsiplayed in the website interface will thus not be dynamic. However, the relational database tables built in grist and the xml files exported from Tacteo will be stored and accessible in Nakala, so our partners and fellow epigraphists can still exploit them for different needs (either for variants or authorities), and we can base ourselves on those trnascribed/translated versions (with missing/damaged/supplied characters and punctuated versions) for future publication and editing purposes.

We are looking forward to share this experience with the organizers and participants of the 10th Epigraphy.info conference and workshop planned in Graz on the 25th of March 2026, where we will also present recent progress on two other section of our corpus, the rubbings and variants.