Have a look at the Introduction to this series of blogposts on Automatic recognition of configurations in plays
I Train an algorithm to identify action information in stage directions (enter-, exit-, dead- and aside)
In a first step stage directions of 15 dramatic texts from Textgrid (www.textgridrep.de) were classified by hand. (https://github.com/pouyana/teireader/tree/ner/drama/hand_classified). This was done in large part by undergraduates participating in the course “Quantitative Dramenanalyse” taught in winter term 2013/2014 together with Dr. Christof Schöch in the Digital Humanities Bachelor-Program at the University of Würzburg. I am very grateful for their work. They were enriched with one of the following type-Attributes in the tag <stage>: enter, exit, aside, dead. This resulted in a total number of 1986 (https://github.com/pouyana/teireader/tree/ner/drama/corpora) stages. Secondly the stages were extracted from the documents and linguistically classified with the help of the Natural Language Toolkit (http://www.nltk.org/). NLTK consits of several text-processing libraries for python scripts to classify and tokenize utterances as well as adding part of speech tags. and analyzing them semantically.
| Other | 1504 |
| Enter | 176 |
| Exit | 230 |
| Aside | 68 |
| Dead | 8 |
Table 1: numbers of the different stage types in the manually classified corpus
In a second step Random Forest was trained with all these types of stages in / with http://scikit-learn.org with all types of satges (exit, enter, aside, dead). This was done with NLTK-Trainer (http://nltk-trainer.readthedocs.org/en/latest/; http://www.nltk.org/howto/data.html; https://github.com/japerk/nltk-trainer)The training corpus consisted of the following plays:
| Lessing, Gotthold Ephraim | Emilia Galotti | 1772 |
| Hafner, Philipp | Mägera, die förchterliche Hexe | 1761 |
| Gryphius, Andreas | Horribilicribrifax | 1663 |
| Lessing, Gotthold Ephraim | Der Freigeist | 1749 |
| Lessing, Gotthold Ephraim | Der Misogyn | 1748 |
| Lessing, Gotthold Ephraim | Die Juden | 1749 |
| Lessing, Gotthold Ephraim | Die alte Jungfer | 1748/49 |
| Lessing, Gotthold Ephraim | Miss Sara Sampson | 1755 |
| Lessing, Gotthold Ephraim | Nathan der Weise | 1779 |
| Lessing, Gotthold Ephraim | Der junge Gelehrte | 1747 |
| Möser, Justus | Die Tugend auf der Schaubühne oder Harlekins Heirath | 1798 |
| Schiller, Friedrich | Die Räuber | 1781 |
| Schiller, Friedrich | Don Carlos | 1787 |
| Schiller, Friedrich | Kabale und Liebe | 1784 |
| Weise, Christian | Masaniello | 1682 |
Table 2: Dramatic texts with hand-classified stages
Then a first automatic classification of new material was done for the following plays, complementing the corpus to 29 dramatic texts, representing a variety of dramatic subgenres in 17th and 18th century :
| Bostel, Lukas von | Der hochmüthige, gestürzte und wieder erhabene Crösus | 1684 |
| Goethe, Johann Wolfgang | Stella | 1775 |
| Gryphius, Andreas | Absurda Comica oder Herr Peter Squentz | 1658 |
| Gryphius, Andreas | Verlibtes Gespenst, geliebte Dornrose | 1660 |
| Kurz, Joseph Felix von | Prinzessin Pumphia | 1767 |
| Lessing, Gotthold Ephraim | Der Schatz | 1748 |
| Lessing, Gotthold Ephraim | Damon | 1748 |
| Lessing, Gotthold Ephraim | Minna von Barnhelm | 1767 |
| Lessing, Gotthold Ephraim | Philotas | 1774 |
| Schiller, Friedrich | Don Carlos | 1787 |
| Stranitzky, Joseph Anton | Türckisch bestraffter Hochmuth | 1725 |
| Weise, Christian | Bäurischer Machiavellus | 1681 |
| Weise, Christian | Der niederländische Bauer | 1700 |
| Weise, Christian | König Wentzel | 1700 |
Table 3: Dramatic texts in the test corpus for automatic classification
The automatically classified action information of the stages (see Table 4: before) then was revised by hand (see Table 4: after):
| Before | After | Count |
| enter | other | 76 |
| aside | enter | 1 |
| aside | dead | 1 |
| aside | other | 4 |
| other | dead | 2 |
| enter | exit | 3 |
| exit | other | 12 |
| other | aside | 21 |
| other | enter | 274 |
| enter | aside | 4 |
| other | exit | 80 |
| exit | dead | 2 |
Table 4: Correction of the automatically classified action information
To improve the performance of that algorithm for an intended use in a bigger context we did some further training with the help of cross-validation. We used 75% of the corpus with the extracted stages for training and validated the results on the remaining 25%. Having classified the whole corpus beforehand it was thus possible to see if the algorithm gave the same results as classified. This procedure is done four times, always with another quarter as a testing corpus.
Here are the results (for an explanation of ‚precision‘, ‚recall‘ and ,f-measure’ see: https://en.wikipedia.org/wiki/Precision_and_recall)
| aside precision | 0.878788 |
| aside recall | 0.878788 |
| aside f-measure | 0.878788 |
| enter precision | 0.936842 |
| enter recall | 0.631206 |
| enter f-measure | 0.754237 |
| exit precision | 0.820000 |
| exit recall | 0.694915 |
| exit f-measure | 0.752294 |
| other precision | 0.866667 |
| other recall | 0.972222 |
| other f-measure | 0.916415 |
Table 5: Results of cross-validation
Cross-validation lead to excellent results for those stages, containing an “aside”. For “enter” and “exit” theres were high precsion values (meaning that those entities which were found trully prooved to be “enter” or “exit”). But there were also quite low values for recall, showing there were not few cases which were not found but classified as “other”.
To improve the results we thought of doing named entity recognition for two purposes:
- To eliminate such cases were stages only consist of names (as they are “enter”)
- To train random forest again with stages containig the information “Person Name” instead of various names.
These effort will be the topic of my next blogpost.
OpenEdition suggests that you cite this post as follows:
Katrin Dennerlein (June 26, 2016). Automatic recognition of configurations in plays – Training a Machine Learning Algorithm. History of Comedy. Retrieved July 25, 2026 from https://comedy.hypotheses.org/75