Automatic recognition of configurations in plays – Introduction

Sure, there are some quantitative, computer-based methods applicable to dramatic texts which do not require any data mining, as for example stylometry. Others, like topic modelling, require a basic linguistic tagging which can be done with existing tools like tree tagger. But today I want to present my small experiments to extract the scenic presence of the characters of a play which I believe is a crucial structural aspect of dramatic texts. This is fundamental for network analysis but also for matrix information. [1]

The idea of charting the scenic presence in a configuration matrix goes back to Solomon Marcus who proposed to make a matrix for every single play which [2] holds one row for every character and one column for every scene. If a character appears on stage, it is registered in the corresponding cell. Based on these matrices Marcus computes for instance, the co-presence of characters, their lack of encounters or their scenic presence over long periods as well as the overall configuration density of a play.[3] The latter is the result of dividing the number of cells holding a 1 by the total number of cells (i.e., the number of characters multiplied by the number of scenes). In other words the configuration density indicates how many of the potential character entrances have actually been realized. It can therefore be designated as a measure of a play’s ‘population density’.[4]

This a simple task if the texts are marked up like the ones from the Folger Shakespare library. http://www.folgerdigitaltexts.org In this data characters in stage directions are tagged with a genuine who-tag, scenes are marked as <scene>, so that it is easy to extract information about who is present on stage in every scene. As far as I can see the markup of character-names in the stage directions is done by hand. For the case of German drama in the 17th and 18th century the markup of the Texts in Textgrid https://textgridrep.de  or the Gutenberg Collection do not not provide this information but only tags for scenes, acts, stage directions and speakers. In some cases there are no scene divisions at all.[5] In others there are, but often characters are entering and leaving the stage during a scene so that it is not possible to assume the scenic presence of all characters just for the <scene>-elements.

A research agenda in this area may look as follows.

  1. Classify stages by hand and train an algorithm to identify action information in stage directions.
  2. Train an algorithm to identify names of characters
  • Bring the information of I and II together
  1. Write a script which reads the results of III and checks if there are contradictions with speaker-information. For example: Show cases where a character identified as dead by a stage directions appears as a speaker afterwards. These cases might be wrong or might denotate a ghost’s appearance.

During the next weeks I will present the results of a small project with Pouyan Azari, student of computer science and Johnathan Gaede, student of digital humanities, working as graduate student researchers, funded by the Bavarian Academy of Sciences and Humanities.[6]


 

[1] Peer Trilcke: Social Network Analysis (SNA) als Methode einer textempirischen Literaturwissenschaft. In: Philip Ajouri, Katja Mellmann u. Christoph Rauen (Hg.): Empirie in der Literaturwissenschaft, Münster 2013, S. 201-247; Frank Fischer, Dario Kampkaspar, Mario Göbel, Peer Trilcke: Digitale Literaturwissenschaftliche Netzwerkanalyse trilcke.de/digitale-lina/; Wilhelm, T., Burghardt, M., & Wolff, C. (2013). “To See or Not to See” – An Interactive Tool for the Visualization and Analysis of Shakespeare Plays. In R. Franken-Wendelstorf, E. Lindinger, & J. Sieck (Eds.), Kultur und Informatik: Visual Worlds & Interactive Spaces (pp. 175–185). Glückstadt: Verlag Werner Hülsbusch. Retrieved from http://epub.uni-regensburg.de/28417/; Katrin Dennerlein “Measuring the average population densities of plays. A case study of Andreas Gryphius, Christian Weise and Gotthold Ephraim Lessing”. In: Semicerchio. Rivista di poesia comparata LIII (2015), S. 80-88.

[2]   Solomon Marcus, Mathematische Poetik (Frankfurt am M.: Athenäum 1973 [first published Bucuresti: Editura Academiei 1970]), p. 290.

[3]   See Marcus, Mathematische Poetik, 292–301.

[4]    Katrin Dennerlein “Measuring the average population densities of plays. A case study of Andreas Gryphius, Christian Weise and Gotthold Ephraim Lessing” Semicerchio. Rivista di poesia comparata LIII (2015).

[5]   In cases where just the tags of the scenes are missing, it is possible to identify scenes automatically via XPath because Acts don’t have a text child.

[6]     http://www.badw.de (08.12.2015).

Katrin Dennerlein
Katrin Dennerlein

OpenEdition suggests that you cite this post as follows:
Katrin Dennerlein (May 2, 2016). Automatic recognition of configurations in plays – Introduction. History of Comedy. Retrieved July 26, 2026 from https://comedy.hypotheses.org/61


Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.