Full text
Artjoms Šeļa, Ben Nagy, Joanna Byszuk, Laura Hernández-Lorenzo, Botond Szemes, and Maciej Eder From Stage to Page: Stylistic Variation in Fictional Speech Abstract:Stylometryismostlyappliedtoauthorialstyle.Morerecently,researchers have begun investigating the style of characters, finding that although there is detectable stylistic variation, the variation remains within authorial bounds. In this article, we address the stylistic distinctiveness of characters in drama. Our primary contribution is methodological; we introduce and evaluate two nonparametric methods to produce a summary statistic for character distinctiveness that can be usefully applied and compared across languages and times. This is a significant advance – previous approaches have either been based on pairwise similarities (which cannot be easily compared) or indirect methods that attempt to infer distinctiveness using classification accuracy. Our first method is based on bootstrap distances between 3-gram probability distributions, the second (reminiscent of ‘unmasking’ techniques) on word keyness curves. Both methods are validated and explored by applying them to a reasonably large corpus (a subset of DraCor): we analyze 3301 characters drawn from 2324 works, covering five centuries and four languages (French, German, Russian, and the works of Shakespeare). Both methods appear useful; the 3-gram method is statistically more powerful, but the word keyness method offers rich interpretability. Both methods are able to capture phonological differences such as accent or dialect, as well as broad differences in topic and lexical richness. Based on exploratory analysis, we find that smaller characters tend to be more distinctive and that women are cross-linguistically more distinctive than men, with this latter finding carefully interrogated using multiple regression. This greater distinctiveness stems from a historical tendency for female characters to be restricted to an ‘internal narrative domain’ covering mainly direct discourse and family/romantic themes. It is hoped that direct, comparable statistical measures will form a basis for more sophisticated future studies, and advances in theory. Artjoms Šeļa, Ben Nagy, Joanna Byszuk, Maciej Eder, Polish Academy of Sciences Laura Hernández-Lorenzo, University of Seville Botond Szemes, Institute for Literary Studies Budapest Artjoms Šeļa, University of Tartu Open Access. ©2024 the authors, published by De Gruyter. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. https://doi.org/10.1515/9783111071824-008
150 | Šeļa et al. 1 Introduction Since Vladimir Propp’s work, structural narratology has approached fictional characters mainly through their role or function – by what they do or whatis done to them (Eder et al. 2010). This character typology relied on recurring functions in the narrative (lover, villain, victim, detective, etc.) and the same perspective was often adopted in computational research, where characters in novels were modeled on the basis of narrative passages rather than dialogue (Bamman et al. 2014; Bonch-Osmolovskaya and Skorinkin 2017; Underwood et al. 2018; Stammbach et al. 2022). In dramatic texts, however, the dominant device for characterization is an utterance. While the script usually contains some stage directions, the specifics of characterization and style of performance are not determined by the text itself, but developed by a specific theater, director or a troupe. Over the course of history, many plays were written for specific theater stages, and it was common practice to write characters for specific actors (Fischer-Lichte 2002). Of course, this kind of ‘outsourced characterization’ was supported by dramatic conventions and formulas. Viewers’ expectations could be shaped without a single word being uttered on stage, just by a character wearing a costume, operating a puppet, or changing a dell’arte stock mask. At the same time, the things characters say and how they say them are the main textual source of information about them. It is reasonable to assume that dramatists make significant efforts to create linguistic distinctions between princes and paupers, lovers and schemers, aristocrats and merchants. Tragic monologue is written differently from a comedic exchange between servants. Some previous computational works treat linguistic distinctiveness of characters from the perspective of this stylistic continuum (Vishnubhotla et al. 2019), noting that it can be influenced by genre, character gender, or their social and professional dispositions. A parallel narratological tradition, tied to Bakhtin’s ideas of heteroglossia, focuses not on abstract character roles, but on the words characters say (Sternberg 1982; Culpeper 2001; Bronwen 2012). The modern novelistic space of dialogic exchange, ‘educated conversation’ (Moretti 2013, p. 20) and the clash of styles in reported discourse become central here. Available stylometric research on fictional speech and micro-stylistic variation suggests that characters within a text are often distinguishable by their local linguistic patterns without obscuring the global authorial trace (Burrows 1987; Hoover 2017). As put by Burrows and Craig: “Characters speak in measurably different ways, but the authorial contrasts transcend this differentiation. The diversity of styles within an author always remains within bounds” (Burrows and Craig 2012, pp. 307–308).
From Stage to Page | 151 Conceptuallyandmethodologically,themajorityofpreviousworksexamined notthedistinctiveness ofcharacters,buttheir(pairwise) similarity.Similaritymeasures are meaningful in pairwise contexts but cannot be analyzed and compared asindividual summarystatistics.SinceBurrows’seminalstudyof speechpatterns in Jane Austen’s characters (Burrows 1987), these approaches focused on calculating similarity within a collection of characters: how different is character X from characterY,and eachof themfromcharacterZ. Burrows measured the correlation between characters’ usage of 30 most frequent words (technically, he fit a linear regressionfortwosetsof log-frequencies);later,similaritywasmostofteninferred throughclusteringbasedonpairwisedistancecalculations(Reeve2015;Craigand Greatley-Hirsch2017;Hoover 2017).Sometimeslinguistic similarity servedas a basisforarguingfunctionalsimilarityaswell.ArecentstudythatlinkedBakhtin’sdialogism and the stylistic diversity of characters’ speech (Vishnubhotla et al. 2019) proposed the analysis of distinctiveness rather than similarity using supervised classification. Instead of using a network of pairwise relationships, the authors asked how well a classifier can recognize character X as being written by author A. Classification accuracy in this scenario becomes an explicit summary statistic for distinctiveness that can be assigned to a character (or, in an aggregated manner, to a play or an author). However, the supervised approach, proposed by Vishnubhotla et al., is data hungry: it suffers from extreme class imbalance, an abundance of short samples (most characters speak only a little), and is dependent on language-specific feature construction procedures. By contrast, this paper will present a simple, non-parametric measure of character distinctiveness that is based on bootstrapped probability distributions representing a character and all others present in a given play: an approach largely informed by authorship verification techniques. This measure is languageindependent and relies only on the context of a single work, which, in turn, minimizes problems of language variation, authorial signal and chronological change in a comparative setting. Individual distinctiveness scores can then be tested against other measures and metadata categories in a hypothesis-driven manner, not only across languages, but also across genres (e.g., novel vs. drama). Do comedies tend to employ more distinct characters? Does distinctiveness increase (authors get better), or decrease (social and linguistic homogenization occurs) over time? Is there a difference between the distinctiveness of fictional women and men? If so is it the direct result of perceived gender differences, or is it constructed by imagined differences in social and professional status? Lackinggooddescriptivemetadataonthedramaticcharacters,thispaperwill not answer the above-mentioned questions in any satisfying way. Instead, we focus on presenting and justifying the measure of distinctiveness and exploring sev-
152 | Šeļa et al. Tab. 1: A summary of the corpus. All word and 3-gram counts are for the filtered corpus (characters that speak at least 2000 words) only. Total Characters Unique Unique Total Total Corpus Characters Analyzed 3-grams Words 3-grams Words French 15462 1744 9896 79994 29.79 m 5.47 m German 14010 1182 14341 150956 24.80 m 4.31 m Russian 3707 248 12542 71217 4.05 m 0.72 m Shakespeare 1431 127 5921 19595 2.16 m 0.43 m eral factors that might shape the final scores (like the year of composition, character gender and characters’ sample size). 2 Materials As the beginning of our exploration of cross-linguistic variation, we examined fourdramatic corpora fromDraCor (Fischeretal.2019): Shakespeare,French,German, and Russian. DraCor is a project that gathers dramatic corpora in various languages, primarily European, encoded in TEI-XML. With 15 corpora available so far, including the Shakespeare corpus available both in English and German, DraCor facilitates large scale analysis of dramatic conventions across language traditions and offers a wide variety of useful metadata at the level of both plays and characters. While the analysis of all DraCor corpora would be possible with the methods we developed, for the purpose of this preliminary study we focused onthelanguagesand dramatictraditions well-knowntothemembers of our team, eventually selecting the full corpora for Shakespeare, French, German, and Russian: a total of 2324 texts, the majority of which come from French and German. The corpus is summarized in Table 1. 3 Methods 3.1 General Approach and Definitions Our understanding of character distinctiveness is largely informed by ‘authorship verification’ approaches, which center around verifying that a text is written by a target author. This problem is more general than ‘authorship attribution,’ which tries to identify the nearest stylistic neighbor for a text (Halvani et al. 2019). In-
From Stage to Page | 153 stead, authorship verification asks about the relative magnitude of similarity: is a targettextmoresimilartosame-authorsamples or different-authorsamples?With this in mind, we define a character’s ‘distinctiveness’ as the degree to which the style of their speech differs from that of other characters. We understand ‘style’ here instrumentally, as a deviation from an unobserved average language (Herrmann et al. 2015), and do not introduce aggressive feature filtering, allowing both ‘grammatical’ and ‘thematic’ signals to contribute to the final measures. We anchor our distinctiveness measure in the context of the specific text in which a character appears. In theory, the frame of reference could be all plays from one author, or all playsfromthesameperiod,orevensomeexternalcorpus– however, all of these would greatly complicate any comparative study. 3.2 Bootstrap 3-Gram Distinctiveness Basedonourdefinition ofdistinctivenessabove,weconsidereda character’sstyle to be an idiolect sampled from a frequency distribution of character 3-grams. As a natural language distribution, this was expected to be generally Zipfian, a family of heavy-tailed distributions, so non-parametric methods were seen to be important. 3-grams were preferred to words for a number of reasons: first, they capture sub-word information, which means they will reflect general sonic preferences (so they can capture things like accent) and, particularly in inflected languages, also reflect some grammatical style; second, as a practical matter, they effectively expandthe sample data, sincea string of text producesapproximately one3-gram per character. This increased sample size should reduce the variance of the statistics. Finally, the number of unique 3-grams in a language is considerably smaller than the number of words, so the frequency data is less sparse, which again is expected to increase robustness. To now operationalize the distinctiveness, as defined, we used standard bootstrap methods to measure the median energy distance (Székely and Rizzo 2013) with bootstrap confidence intervals between the two distributions (character 3-gram frequencies vs. ‘other’ 3-gram frequencies). The energy distance is one of a family of related metrics that are commonly used to measure the difference between probability distributions. Some limitations and choices were required. As mentioned, we measured distinctiveness only within the context of a single work (even for authors with multiple works). To expand beyond single works would produce very mismatched sample sizes, since some authors were prolific and some produced just one play; even with non-parametric methods, hugely mismatched sample sizes are problematic. Furthermore, the plays span four languages and roughly five centuries, making the ‘distant’ context seem ridiculous. As well as the selected distinctiveness statis-
154 | Šeļa et al. tic (median energy distance) we also recorded a ‘baseline’ distinctiveness, this being each character’s distance from themselves. The theoretical baseline is, of course, zero, but the sample baselines will not be, meaning that this gives us an idea of the inherent variance of the samples. Finally, when selecting characters to examine, we chose a minimum size of 2000 words. Sample sizes are somewhat arbitrary, and are matters of debate (Eder 2015, 2017), but this seemed to be a reasonable, or perhaps even slightly aggressive, lower bound. 3.3 Area under Keywords Our second, supplementary approach was informed by ‘unmasking’ techniques often employed in stylometric research (Koppel and Schler 2004; Kestemont et al. 2016; Plecháč and Šeļa 2021). Unmasking refers to a range of methods that share one goal: to measure and compare the depth of the differences between two sets of texts. For example, an author might write both high fantasy fiction and historical novels: a classifier would have little difficulty distinguishing one genre from another by simply using superficial features (e.g., ‘dragons,’ ‘magic,’ ‘elves’). However,byassumption, ifthesemostdistinctivefeaturesareremoved,theclassifierwillhavemoretroubledetermining whichtextcamefromwhichpool,because the texts share one deep similarity – a common authorial style. Conversely, if we compare books by two different fiction writers, these texts will also have superficial differences. However, while removing more and more distinctive features, the classifier should remain confident in distinguishing the authors from each other, because the texts do not share an authorial style that is deeply rooted in common linguistic elements and distributed over many features. By comparing the speed with which the rates of accuracy decay, we can approach authorship verification problems, i.e., how plausible is it that this text belongs to author A? We applied the same thinking to fictional characters, as opposed to authors: the distinctiveness of a character may rely on a small number of catch-phrases (‘Gadzooks!’ or ‘Cowabunga!’) or it may be driven by non-stylistic, referential factors(Mary,speakingtoJohn,isnotlikely tousetheword‘Mary,’but is likely touse the word ‘John,’ and vice-versa). On the other hand, there are characters whose speech systematically differs from the neutral language: such as when the author imitates dialects, slang, regionalism, speech and phonetic idiosyncrasies. In the former case, an imaginary classifier should quickly lose accuracy (since John and Mary speak quite similarly), but in the latter case the removal of a small number of features would not be enough to disrupt classification. In our case, it was impractical to use ‘standard,’ supervised (i.e., classifierbased) unmasking because individual characters, as samples, were simply too
From Stage to Page | 155 small. Instead we used word keyness – a character’s relativepreferencefor a word in the context of a given drama – to calculate an alternative distinctiveness score together with a bag of easily interpretable features per character. First, we use weighted log-odds (Monroe et al. 2008) to calculate keywords for a character relative to the speech pool of the rest of the cast; second, we represented each character by their 100 words with the highest keyness, arranged by rank; finally, we measured the area under this curve, which we interpret as distinctiveness – characters with just a few key words will exhibit less area under the keyness curve. By comparing these final areas, we can compare the relative difference between each character and rest of the speech in the play. In a similar manner to the bootstrapped approach, we upsample each character’s word pool to match the size of the rest of the words in the play to minimize the effect of the sample size as much as possible. 4 Results Overall, the distinctiveness energy statistic appears useful. The baseline (character vs. self) is quite stable cross-linguistically, although it is slightly higher for characters with a very large share of dialogue (Figure 1). Note also that the distinctiveness statistic appears roughly Gaussian (see Appendix B for more discussion) and its range is relatively consistent between languages (peaking at roughly 0.20), although this consistency does not apply at the level of authors. The obvious issue is that there is a strong negative correlation between character size and distinctiveness, but this is not only a limitation of the method – lead characters naturally set the dominant style of a text (and possibly inherit more of the ‘true’ authorial voice). Importantly, distinctiveness does not increase with the number of speakers in a play. The method works best when there are reasonable sample sizes for both the examined character and the ‘other’ class. This is illustrated by the ‘U’ curve visible in the French corpus in Figure 1 as the examined characters’ dialogue sharepasses50%. Ashoped,the energy-distancemethoddoes appearto capture characters who are written with distinctive idiolects, representing things like foreign accents or social class. For a discussion of this see Section 5. As seen in Figure 2, there is no clear correlation between the date of composition and character distinctiveness, which suggests that language change does not disturb the measure. The finding that seems clear is that women are written differently from men. Female characters are generally more distinctive in all corpora (Figure 3b), although this is not visible using the keyness AUC measure – leading us to conclude that the keyness measure has lower power. This difference in the
156 | Šeļa et al. Fig. 1: Character distinctiveness, per corpus, versus % Dialogue. Women are shown smaller and in orange, men (and undefined) larger and in blue. GAM (Generalized Additive Model) trendlines are superimposed in the same colors. Baseline data (GAM trend for distinctiveness of character vs. self) is shown as a dashed line. distinctiveness of female characters can partly be explained by the fact that they tend to have smaller parts (Fig 3a), and smaller characters in general are more distinctive (Figure 1), but that is not the whole story. Female parts have more restricted 3-gram vocabularies (Figure 3c), suggesting that they are also restricted in their semantic fields. This becomes clearer when the relative frequencies of their (word) vocabularies are examined. As well as the stereotypical tendencies (women say ‘love,’ men say ‘sword’), the female characters, cross-linguistically, seem to be less likely to reference the ‘external world’ of the drama. As seen in Appendix A, relatively more frequent words for women are dominated by personal
From Stage to Page | 157 Fig. 2: Character distinctiveness, per corpus, versus year composed (DraCor data). Women are shown smaller and in orange, men (and undefined) larger and in blue. GAM (Generalized Additive Model) trendlines are superimposed in the same colors. Baseline data (GAM trend for distinctiveness of character vs self) is shown as a dashed line. pronouns representing ‘I,’ ‘me,’ ‘you’ etc. or words relating to family. The male lists are dominated by indicative articles and political terms (‘law,’ ‘noble,’ ‘king’ etc.). The higher distinctiveness of female characters is further supported by a formal linear model: we fit a Bayesian multiple regression where distinctiveness was conditioned on both gender and size (characters’ percentage of total dialogue). A direct gender effect is present in all corpora, as expected from Figure 3a, but when we account for variation among authors, the effect may be less pronounced than it appears (for analysis and more detailed discussed, the posterior estimates are described in Appendix B). Our finding interlocks with the observation by Under-
164 | Šeļa et al. B Bayesian Regression Models: Effect of Gender on Distinctiveness Fig. 4: Character distinctiveness, predicted from posterior, estimate of grand mean (no grouplevel effects), 6000 draws. Predictions are made for a counterfactual “median” character role, who has 20.9% of dialogue share. Predictions are presented at natural scale. Istheperceivedgendereffect‘real’?Intechnicalterms,whatisthedirectinfluence of character gender (G) on distinctiveness scores (D) across traditions (T), conditioned on the share of dialogue they have (S)? To answer this, we fit a Bayesian multilevel multiple regression with group-level estimates for individual plays (P). We chose to model at the level of plays both because our D statistic is tied to the context of a single play, and because character features coming from the same
From Stage to Page | 165 Fig. 5: Posterior predictions for gender, marginal of individual plays. Errorbars show .95 CI. Empirical data is plotted in colour, 5 extreme cases (>0.3) are filtered out. Predictions are presented at natural scale. play are not independent (e.g. there cannot be two characters with 60% of the dialogue). Modelling this way also significantly improved predictions. Gender is allowed to interact by corpus, yielding a single, cross-linguistic model that makes compatible predictions for different traditions. In brms formula syntax: log(𝐷) ∼ 𝐺 ∗ 𝑇 + 𝑇 ∗ (𝑆 + 𝐼(𝑆^2)) + (1|𝑃 ) (1) Based on sample observation, we used a Gaussian prior for log-transformed D scores. We could have also fitted the original values, but D scores have extreme outliers that extend the tail: the model has much easier time with sampling and chain convergence on a log-transformed domain. We chose a quadratic term for S,
166 | Šeļa et al. because the relationship between D and S is U-shaped. Importantly, ‘unknown’ gender entities are filtered, because often (but not always) this is not data that is missing, but entries that are incompatible with a binary classification:¹primarily collective or compound entities (people, choirs, soldiers). It would have been possible to use standard strategies, like imputation, to ‘repair’ the data, but that approach would be incorrect. Posteriorestimatesfordistinctivenessbygenderare shown in Figure 4.Based on the figure, we can be most confident about the difference in German and least confident in Shakespeare (few characters and, specifically, few women with large dialogue shares). The differences in means, however, appear consistent. As calculated from the posterior: in French, female characters are more distinctive by only .009 (±.003); in German, by .017 (±.003); in Russian by .023 (±.009, the widest CI); and in Shakespeare by .012 (±.008). To understand the full extent of variation across different plays, it is useful to look at the marginal posterior means of the plays (Fig. 5). Here, the difference in distinctiveness between genders remains visible, but there is a better estimation of the global uncertainty and variation across different texts. Note that the confidence intervals in Fig. 5 are asymmetric (wider on the upper arm), having been transformed from symmetric intervals on a log domain. 1In modern terms, it is vexing to be forced to reduce characters to a gender binary, but since gender non-conforming characters are virtually unrepresented in this predominantly historical corpus, the point is moot.