Automatic recognition of configurations in plays – Training a Machine Learning Algorithm

Have a look at the Introduction to this series of blogposts on Automatic recognition of configurations in plays

I Train an algorithm to identify action information in stage directions (enter-, exit-, dead- and aside)

In a first step stage directions of 15 dramatic texts from Textgrid (www.textgridrep.de) were classified by hand. (https://github.com/pouyana/teireader/tree/ner/drama/hand_classified). This was done in large part by undergraduates participating in the course “Quantitative Dramenanalyse” taught in winter term 2013/2014 together with Dr. Christof Schöch in the Digital Humanities Bachelor-Program at the University of Würzburg. I am very grateful for their work. They were enriched with one of the following type-Attributes in the tag <stage>: enter, exit, aside, dead. This resulted in a total number of 1986 (https://github.com/pouyana/teireader/tree/ner/drama/corpora) stages. Secondly the stages were extracted from the documents and linguistically classified with the help of the Natural Language Toolkit (http://www.nltk.org/). NLTK consits of several text-processing libraries for python scripts to classify and tokenize utterances as well as adding part of speech tags. and analyzing them semantically.

Other 1504
Enter 176
Exit 230
Aside 68
Dead 8

Table 1: numbers of the different stage types in the manually classified corpus

In a second step Random Forest was trained with all these types of stages in / with http://scikit-learn.org with all types of satges (exit, enter, aside, dead). This was done with NLTK-Trainer (http://nltk-trainer.readthedocs.org/en/latest/; http://www.nltk.org/howto/data.html; https://github.com/japerk/nltk-trainer)The training corpus consisted of the following plays:

Lessing, Gotthold Ephraim Emilia Galotti 1772
Hafner, Philipp Mägera, die förchterliche Hexe 1761
Gryphius, Andreas Horribilicribrifax 1663
Lessing, Gotthold Ephraim Der Freigeist 1749
Lessing, Gotthold Ephraim Der Misogyn 1748
Lessing, Gotthold Ephraim Die Juden 1749
Lessing, Gotthold Ephraim Die alte Jungfer 1748/49
Lessing, Gotthold Ephraim Miss Sara Sampson 1755
Lessing, Gotthold Ephraim Nathan der Weise 1779
Lessing, Gotthold Ephraim Der junge Gelehrte 1747
Möser, Justus Die Tugend auf der Schaubühne oder Harlekins Heirath 1798
Schiller, Friedrich Die Räuber 1781
Schiller, Friedrich Don Carlos 1787
Schiller, Friedrich Kabale und Liebe 1784
Weise, Christian Masaniello 1682

Table 2: Dramatic texts with hand-classified stages

Then a first automatic classification of new material was done for the following plays, complementing the corpus to 29 dramatic texts, representing a variety of dramatic  subgenres in 17th and 18th century :

Bostel, Lukas von Der hochmüthige, gestürzte und wieder erhabene Crösus 1684
Goethe, Johann Wolfgang Stella 1775
Gryphius, Andreas Absurda Comica oder Herr Peter Squentz 1658
Gryphius, Andreas Verlibtes Gespenst, geliebte Dornrose 1660
Kurz, Joseph Felix von Prinzessin Pumphia 1767
Lessing, Gotthold Ephraim Der Schatz 1748
Lessing, Gotthold Ephraim Damon 1748
Lessing, Gotthold Ephraim Minna von Barnhelm 1767
Lessing, Gotthold Ephraim Philotas 1774
Schiller, Friedrich Don Carlos 1787
Stranitzky, Joseph Anton Türckisch bestraffter Hochmuth 1725
Weise, Christian Bäurischer Machiavellus 1681
Weise, Christian Der niederländische Bauer 1700
Weise, Christian König Wentzel 1700

Table 3: Dramatic texts in the test corpus for automatic classification

The automatically classified action information of the stages (see Table 4: before) then was revised by hand (see Table 4: after):

Before After Count
enter other 76
aside enter 1
aside dead 1
aside other 4
other dead 2
enter exit 3
exit other 12
other aside 21
other enter 274
enter aside 4
other exit 80
exit dead 2

Table 4: Correction of the automatically classified action information

To improve the performance of that algorithm for an intended use in a bigger context we did some further training with the help of cross-validation. We used 75% of the corpus with the extracted stages for training and validated the results on the remaining 25%. Having classified the whole corpus beforehand it was thus possible to see if the algorithm gave the same results as classified. This procedure is done four times, always with another quarter as a testing corpus.

Here are the results (for an explanation of ‚precision‘, ‚recall‘ and ,f-measure’ see: https://en.wikipedia.org/wiki/Precision_and_recall)

aside precision 0.878788
aside recall 0.878788
aside f-measure 0.878788
enter precision 0.936842
enter recall 0.631206
enter f-measure 0.754237
exit precision 0.820000
exit recall 0.694915
exit f-measure 0.752294
other precision 0.866667
other recall 0.972222
other f-measure 0.916415

Table 5: Results of cross-validation

Cross-validation lead to excellent results for those stages, containing an “aside”. For “enter” and “exit” theres were high precsion values (meaning that those entities which were found trully prooved to be “enter” or “exit”). But there were also quite low values for recall, showing there were not few cases which were not found but classified as “other”.

To improve the results we thought of doing named entity recognition for two purposes:

  • To eliminate such cases were stages only consist of names (as they are “enter”)
  • To train random forest again with stages containig the information “Person Name” instead of various names.

These effort will be the topic of my next blogpost.

 

 

 

 

Katrin Dennerlein
Katrin Dennerlein

OpenEdition suggests that you cite this post as follows:
Katrin Dennerlein (June 26, 2016). Automatic recognition of configurations in plays – Training a Machine Learning Algorithm. History of Comedy. Retrieved July 25, 2026 from https://comedy.hypotheses.org/75


Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.