207 41 7MB
English Pages 185 [182] Year 2024

Challenges in Corpus Linguistics Rethinking corpus compilation and analysis edited by Mark Kaunisto Marco Schilk
Studies in Corpus Linguistics
118 JOHN BENJAMINS PUBLISHING COMPANY
Challenges in Corpus Linguistics
Studies in Corpus Linguistics (SCL) issn 1388-0373
SCL focuses on the use of corpora throughout language study, the development of a quantitative approach to linguistics, the design and use of new tools for processing language texts, and the theoretical implications of a data-rich discipline. For an overview of all books published in this series, please see benjamins.com/catalog/scl
General Editor
Founding Editor
Ute Römer-Barron
Elena Tognini-Bonelli
Georgia State University
The Tuscan Word Centre/University of Siena
Advisory Board Laurence Anthony
Susan Hunston
Antti Arppe
Michaela Mahlberg
Michael Barlow
Anna Mauranen
Monika Bednarek
Andrea Sand
Tony Berber Sardinha
Benedikt Szmrecsanyi
Douglas Biber
Elena Tognini-Bonelli
Marina Bondi
Yukio Tono
Jonathan Culpeper
Martin Warren
Sylviane Granger
Stefanie Wulff
Waseda University
University of Alberta University of Auckland University of Sydney Catholic University of São Paulo Northern Arizona University University of Modena and Reggio Emilia Lancaster University University of Louvain
University of Birmingham University of Birmingham University of Helsinki University of Trier Catholic University of Leuven The Tuscan Word Centre/University of Siena Tokyo University of Foreign Studies The Hong Kong Polytechnic University University of Florida
Stefan Th. Gries
University of California, Santa Barbara
Volume 118 Challenges in Corpus Linguistics. Rethinking corpus compilation and analysis Edited by Mark Kaunisto and Marco Schilk
Challenges in Corpus Linguistics Rethinking corpus compilation and analysis Edited by
Mark Kaunisto Tampere University
Marco Schilk University of Hildesheim
John Benjamins Publishing Company Amsterdam / Philadelphia
∞
TM
The paper used in this publication meets the minimum requirements of the American National Standard for Information Sciences – Permanence of Paper for Printed Library Materials, ansi z39.48-1984.
Cover design: Françoise Berserik Cover illustration from original painting Random Order by Lorenzo Pezzatini, Florence, 1996.
doi 10.1075/scl.118 Cataloging-in-Publication Data available from Library of Congress: lccn 2024029608 (print) / 2024029609 (e-book) isbn 978 90 272 1588 8 (Hb) isbn 978 90 272 4653 0 (e-book)
© 2024 – John Benjamins B.V. No part of this book may be reproduced in any form, by print, photoprint, microfilm, or any other means, without written permission from the publisher. John Benjamins Publishing Company · https://benjamins.com
Table of contents Acknowledgements From fallacies and pitfalls to solutions and future directions: Navigating the evolving terrain of corpus linguistics Mark Kaunisto Engaging with bad (meta)data in historical corpus linguistics Turo Vartiainen & Tanja Säily Named entities as potentially problematic items in corpora Mark Kaunisto
vii 1 9 35
Challenges in the compilation, annotation and analysis of learner corpus data Marcus Callies
55
Early newspapers as data for corpus linguistics (and Digital Humanities): Issues in using the British Library Newspapers database as a corpus Turo Hiltunen
68
Open Corpus Linguistics – Or how to overcome common problems in dealing with corpus data by adopting open research practices Stefan Hartmann
89
Text length and short texts: An overview of the problem Aatu Liimatta
106
Corpus genre categories: Issues at the intersection of linguistics and literature Daniel Ocic Ihrmark
126
Modeling fine-grained sociolinguistic variation: The promises and pitfalls of Twitter corpora and neural word embeddings Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
142
Subject index
171
Acknowledgements Most of the chapters of this volume have their origins in the pre-conference workshop of the 42nd ICAME conference held at TU Dortmund University in August 2021. The workshop, organised by Mark Kaunisto, Marco Schilk, and Jukka Tyrkkö, was titled “Corpus pitfalls: dealing with messy data (and other traps for the unwary)”, and it brought together linguists interested in addressing some of the not-always-so-smooth aspects of corpus linguistic study. The convenors and participants acknowledged the pioneering work of corpus linguistic scholars, such as Matti Rissanen and Charles Fillmore, who drew attention to some of the core challenges in corpus design and interpretation. Their earlier work served as an inspiration for a continued watchful treatment of classic and newer corpus data. It was felt important to share the concerns, demonstrate the practical nature of the problems, and contemplate solutions to work around the challenges, a sentiment that culminated in the present volume. The editors would like to thank the organisers of the ICAME conference, as well as Jukka Tyrkkö for his insights and contributions in the early stages of planning the present volume. A number of people were called upon to act as external peer reviewers for the chapters, and we are grateful for the expert work done by Tine Breban, Katrin Menzel, Dong Nguyen, Arja Nurmi, Alexandr Rosen, Christina Sanchez-Stockhammer, Fabian Vetter, Valentin Werner, and Martin Weisser. The resulting chapters have greatly benefitted from the review process. We are also indebted to John Benjamins for including the volume as part of the Studies in Corpus Linguistics series. It has been a pleasure to work with the people at Benjamins through various stages of the project. We thank the series editor Ute Römer-Barron and Kees Vaes for their cordiality and support.
From fallacies and pitfalls to solutions and future directions Navigating the evolving terrain of corpus linguistics Mark Kaunisto
Tampere University
In a short but important paper published thirty-five years ago in the ICAME Journal, Rissanen (1989) identified and discussed three potential problems that the use of corpora may present to unwitting scholars not necessarily closely familiar with the contents, structure, and overall character of the corpora that they set out to examine. Rissanen referred to the problems with the terms “The philologist’s dilemma”, “God’s truth fallacy”, and “The mystery of vanishing reliability”, and although the names themselves were arguably distracting in their likeness with titles of detective novels, the identification of these issues was very perceptive. The first problem related to the possibility of increased distance between scholars and the actual texts or other sources of language that a corpus includes, in other words, Rissanen was worried that the use of corpora, which enable quick searches of words and patterns from large amounts of language data, may become less familiar with the true nature of the sources themselves. As regards the second problem, the concern with the use of corpora has to do with a user trusting the corpus and the findings based on it without due consideration to its representativeness or limitations. The third problem related to the danger of relatively small corpora being inadequate for the study of sociolinguistic variation, the reliable analysis of which would require relatively large amounts of data representing a broad range of different types of variables. The insightfulness of Rissanen’s article has been generally recognized by corpus linguists (see, e.g., McEnery & Wilson 2001: 124–125; Adamson & AyresBennett 2011: 203–204; Rütten 2017: 8; Hoffmann et al. 2018: 7), and it has continued to inspire scholars, even though it could be seen as having a cautionary, or even a pessimistic tone. Examples of works building upon Rissanen’s article include Bennett et al. (2013), and the special issue “Focus on the Philologist’s Dilemma” of Anglistik (Vol. 28, 2017, edited by Tanja Rütten), and on the whole, https://doi.org/10.1075/scl.118.01kau © 2024 John Benjamins Publishing Company
2
Mark Kaunisto
Rissanen’s paper is still frequently cited in connection with discussions on the need to balance between quantitative and qualitative (or philological) approaches into language study (e.g., Dollinger 2016; Mindt 2017; Moskowich & Crespo 2023). As solutions to the issues raised, Rissanen (1989) promoted the compilation of larger and more representative corpora, as well as getting acquainted with the data itself. However, what has been seen in practice is that the increase in the size of the corpus is not necessarily a straightforward solution to the problems. Today we have access to corpora which are massive, compared to the ones available thirty years ago, which undoubtedly has enhanced the possibilities of conducting a broader range of studies. At the same time, paradoxically, the increases in corpus size make it considerably less likely that corpus users will have a clear understanding of the characteristics of many of the texts (or text types) included in the corpora. In other words, the sound advice of “knowing your data” becomes increasingly hard to follow when working with corpora that consist of billions of words from a huge number of different sources. It may therefore be argued that larger corpora have also brought about new kinds of challenges (see, e.g., Hiltunen, McVeigh & Säily 2017; Rütten 2017: 9; Mindt 2017: 71, as well as Vartiainen & Säily in the present volume). Weisser (2016: 255) also openly questions the factuality of the notion that having larger corpora is always desirable, stating that “the idea that ‘bigger is better’ […] may not always be fully justified if the quantity of data isn’t equally matched by quality”. In general, critical self-assessment is unquestionably an essential part in the evolution of any new methodological approach into language study, and corpus linguists have consistently recognized the theoretical and practical challenges inherent in the compilation, design, organization, and analysis of data. Over the years, scholars have noted and discussed the occasionally elusive nature of the challenges and suggested solutions to the problems, tackling broader and abstract notions such as representativeness and balance of corpora, but also more concrete questions like sampling, various levels of annotation, precision and recall of corpus queries, dispersion of search hits across a corpus, statistical evaluation, among other points. The attitudes of corpus linguists seem to include both a healthy sense of humility as regards what is possible to achieve, as well as an understanding of the need to continuously remind the scholarly community of the core concerns in the field. As illustrations of the former idea, we can consider the often-quoted statements that “[t]he results are only as good as the corpus” (Sinclair 1991: 13) and “[a]bsolute representativeness is an unattainable aim” (Mukherjee 2004: 114), as well as the observation that “[n]o resource is going to meet the needs of all possible users, but the point is to create as useful of tools as is possible for the largest number of researchers” (Davies 2010: 413). As regards the necessity to maintain a
From fallacies and pitfalls to solutions
critical mindset at every turn, the assessments of several interviewed corpus linguists in Viana et al. 2011 are a good indication of the recognition of limitations that the use of corpora still has. Having noted that it is only sensible to be aware of the impossibility to create perfect corpora and analytical methods which are entirely free of pitfalls, it would feel wrong not to take note of the increase of the levels of analytical sophistication and efficiency in the field. The toolbox that a corpus linguist has today is immensely superior to a corresponding one twenty or thirty years ago. However, some problems persist, and it is safe to assume that the number of pitfalls that a corpus linguist needs to try to avoid has not necessarily decreased, quite the contrary. But it is not difficult to imagine that for novice users of corpora, the numerous functions readily available in concordancers and online corpus interfaces have created another fallacy that might be called a “fallacy of sophisticated technology”: with such state-of-the-art systems, what could go wrong? In a way this would be related to Rissanen’s “God’s truth fallacy”, but instead of a false sense of security arising from the seemingly representative datasets, the sheer impressiveness of the software and tools of analysis may suggest that pitfalls do not exist. Improvements in some areas do not mean that all of the old problems have been solved. In many ways we can then say that the challenges do not only relate to corpora and their use themselves, but also to how we educate new corpus users and inform them about the fundamental concepts and concerns. There are undoubtedly many types of persistent problems, common frustrations, and messiness in corpus data that seasoned scholars have encountered and know about, but which are seldom specifically addressed. Yet beginning corpus users might benefit from learning about what may be regarded as tacit knowledge in corpus linguistics, and even the more advanced scholar may encounter issues new to them that have been addressed earlier. This is also the main motivating factor behind the present volume. Comprising eight chapters, the book discusses the nature of these problems and seek solutions to the perplexities faced by themselves and other scholars. While the chapters mostly concentrate on English corpora, many of the themes are relevant from a general corpus-linguistic point of view. The topics range from addressing issues relating to grammatical annotation (Vartiainen & Säily; Kaunisto, Callies), attempts to use contents of more or less unorganized databases as resources for corpus linguistic studies (Vartiainen & Säily; Hiltunen), problems of replicability of studies based on data with limited availability (Hartmann), issues in text-analytic studies caused by differences between text length (Liimatta), practices of genre categorizations of fictional texts in corpora (Ihrmark), as well as the study of fine-grained sociolinguistic phenomena (Miletić, Przewozny-Desriaux & Tanguy).
3
4
Mark Kaunisto
It deserves to be mentioned that Rissanen’s concerns on the use of corpora were explicitly presented from the point of view of historical linguistics, and while many of the issues he raises are also relevant to the study of contemporary forms of languages, historical corpus linguistics faces significant challenges of its own. The chapter by Turo Vartiainen and Tanja Säily (Chapter 2) raises a number of practical points which still pose problems in the analysis of historical corpora. The authors examine a set of case studies, highlighting issues like part-of-speech annotation, miscategorization challenges, and problems related to metadata and digitized texts in historical databases, such as the 10.5-billion-word Eighteenth Century Collections Online, a database which was originally not compiled for linguistic purposes. Reporting on a number of instances where initial findings from corpora turned out to be misleading or inconclusive upon closer inspection of the data, the authors advocate for methodological improvements, user feedback channels, and balancing subgenres. They stress the significance of “knowing one’s data” for researchers dealing with vast corpora, emphasizing collaboration with historians, and systematic text examination for qualitative analyses and data interpretation. The impact of insufficient metadata is also the topic of Chapter 3 by Mark Kaunisto, who examines the problems seen in the annotation of named entities in corpora, delving into the problems that arise in interpreting corpus data. Despite the long-recognized importance of considering proper nouns and names in corpus annotation schemes, many contemporary linguistic corpora lack the capability to exclude these items effectively when setting up queries. The presence of names in corpora introduces a layer of complexity, potentially including items that may not reflect the active linguistic choices of the represented writers or speakers. Through small-scale case studies utilizing prominent English language corpora like the British National Corpus and the Corpus of Contemporary American English, the chapter underscores the necessity for careful post-processing of search results. The occurrence of searched items within named entities poses challenges in analyzing word frequencies and collocational behavior. The chapter advocates for more detailed annotation of named entities in large linguistic corpora that are already available. Notable improvements could also be gained by the increased collaboration between traditional corpus linguists and scholars in the fields of computational linguistics and NLP, of which there are already encouraging examples. Chapter 4 by Marcus Callies discusses challenges related to the special characteristics of learner corpus data, emphasizing the importance of valid data representation for studying L2 production and development. Learner corpora, particularly those containing academic texts, often include instances of multilingual practices, code-switching, and borrowed content, presenting difficulties for annotation and analysis. Specific attention is called for in tagging elements like
From fallacies and pitfalls to solutions
expert terminology, metalinguistic language use, and quoted passages to ensure the accuracy of word counts and concordance analyses. Compilers and users of learner corpora also grapple with the issue of unwanted lexical bias, stemming from task topics, writing prompts, or other input materials. Identifying and addressing lexical bias is crucial, as it can impact the validity of research findings, especially given the consideration of lexical variation as a proxy for L2 proficiency. Methods to tackle lexical bias involve treating biased words as stopwords or excluding L2 structures likely induced by bias. Additionally, task and prompt materials may influence the recurrent use of specific grammatical constructions, highlighting the need for careful analysis. Callies concludes his chapter by addressing potential bias in annotation methods within Learner Corpus Research (LCR) and related disciplines. Overall, the chapter emphasizes the need to move beyond a monolingual native-speaker norm in assessing learner data, acknowledging the significance of diverse linguistic expressions. Callies also raises an important point about the challenges that the use of AI or other writing tools may pose in the compilation of learner corpora in the future – a point that will also be relevant to the compilation of many different types of corpora – as one needs to make sure that the samples compiled truly reflect the writers’ own language choices. In Chapter 5, the question of the useability of databases for corpus linguistic purposes is addressed by Turo Hiltunen, who discusses issues relating to the use of materials in the British Library Newspapers database. Perspectives on the matter vary between scholarly approaches: the criteria for determining what is “good data” – for example, considerations on the accessibility and processing of the data – can be quite different in Digital Humanities and corpus linguistics. For a corpus linguist, it is important to have detailed knowledge of how the data has been compiled, what editorial or reformatting practices have been applied to the data, and with what tools the data can be examined. Hiltunen highlights ways in which the useability of databases can be assessed, also noting the different conceptualizations of notions such as register in different scholarly disciplines as posing a challenge. As a way forward in trying to find solutions to improve the adaptability of databases for linguistic research, Hiltunen calls for more interdisciplinary discussions and collaboration between corpus linguistics, Digital Humanities, and related disciplines, which are still too far from each other because of the disparities in their research agendas. The subject of accessibility of data and its repercussions in corpus study is also covered by Stefan Hartmann (Chapter 6), who brings up the problem of replicability of corpus studies, often resulting from the limited accessibility of the corpora studied, as some corpora are only available behind a paywall. Seeing connections between these problems and those outlined by Rissanen, Hartmann advocates for open research practices in corpus linguistics, including the
5
6
Mark Kaunisto
open availability of entire corpora and sharing of concordances, annotations, and analysis scripts. Although there are practical challenges, such as legal uncertainties with copyright-protected content, it is argued that Open Corpus Linguistics not only addresses concerns about transparency and non-reproducibility, but also helps to contextualize research findings and to overcome Rissanen’s problems. With the rise of social media and web data, the question of how texts of distinctly different shapes and forms, compiled together into a corpus, can reliably be examined becomes increasingly pertinent. Aatu Liimatta’s chapter (Chapter 7) focusses on the challenges posed by variation in text length and the specific issue of short texts, an issue which so far has not received much attention from quantitative corpus linguists. Largely the reason for this is how previously short texts have not been regarded as being a major problem, but with the advent of new forms in contemporary digital communication, the ‘problem of text length’, as Liimatta calls it, needs to be addressed. Although solutions have been proposed, they often appear to be limited as regards their suitability in different kinds of studies. Liimatta proposes potential avenues for improvement and new method development, including the exploration of resampling methods for estimating distribution and approaches utilizing the large size of datasets. While perfect solutions may not yet exist, Liimatta offers useful insights to the study of linguistic questions within the context of varying text lengths. In Chapter 8, Daniel Ocic Ihrmark explores the challenges of categorizing fiction genres in corpus compilation, especially when catering to both linguistic and literary research fields. General-purpose linguistic corpora have traditionally aimed to include works of fiction; however, as Ihrmark notes, the practices of categorizing (and subcategorizing) literary genres in the disciplines tend to differ, which may result in difficulties in trying to make use of corpora in literary studies. The differences between how literary genres are categorized largely reflect the different viewpoints and goals between the fields, with corpus linguistic categorizations aiming to focus on the different communicative purposes of texts, whereas the descriptions of genres by literature scholars are based more on content and stylistic concerns. Ihrmark examines various methods employed in corpus stylistics, such as keyword analysis and n-gram searches, and their reliance on genre categorization for comparative studies. Overall, Ihrmark suggests adopting broader genre categorizations at higher levels for wider applicability, while allowing for additional granularity as needed. As observed in other chapters in the volume, corpus linguistics stands to benefit from the innovative ideas, methods, and expanded possibilities developed in related fields such as natural language processing and computational linguistics, enriching its analytical toolkit and enhancing its capacity for nuanced linguistic analysis. In the concluding chapter of the volume, Filip Miletić, Anne
From fallacies and pitfalls to solutions
Przewozny-Desriaux and Ludovic Tanguy examine the potential of natural language processing techniques and social media data in studying language variation, particularly contact-induced semantic shifts in Quebec English. By analyzing English tweets from Montreal, Toronto, and Vancouver, the authors investigate instances where English words take on meanings influenced by phonologically similar French words. They employ BERT (Bidirectional Encoder Representations from Transformers), a model based on a neural network, to identify clusters of tweets sharing similar meanings of target words. The findings reveal that contact-related meanings tend to be more prevalent in tweets from Montreal, indicating a connection to the city’s large French-speaking population. However, the study also highlights challenges such as false positives, cultural factors, and idiolectal preferences, which can skew the analysis. Despite these limitations, the research demonstrates the potential of large corpora and statistical models in advancing sociolinguistic understanding but underscores the importance of careful data collection and interpretation. In conclusion, it is worth noting again that corpus linguists have acknowledged the continuous nature of the struggle to find solutions to the problems that we are facing. Although Labov’s (1994: 11) often-quoted notion of “making the best use of bad data” originally referred to historical linguistics, to a degree the idea can be extended to almost all corpora. It may seem frustrating to recognize that the work in trying to find solutions to the perplexes in corpus linguistics will never be complete, and while doing so, scholars will always be a few crucial steps behind, as new types of corpora, datasets, analysis methods, computational tools are being developed and become available. Correspondingly, although it might otherwise seem counterintuitive to say that the editors of this volume would wish its contents to become out-of-date or obsolete, considering the broader goals of the present volume, this would of course be desirable, and the sooner this happens, the better. At the same time, one must resign to the fact that this is not likely to happen, at least not in its entirety, as some of the issues dealt with here have to do with fundamental issues of corpus linguistics, and because of the evolving nature of corpus linguistics as a scholarly field. Books with similar goals as those of the present volume will undoubtedly still be needed in the future.
References Adamson, Sylvia & Ayres-Bennett, Wendy. 2011. Linguistics and philology in the Twenty-first Century: Introduction. Transactions of the Philological Society 109(3): 201–206. Bennett, Paul, Durrell, Martin, Scheible, Silke & Whitt, Richard J. (eds). 2013. New Methods in Historical Corpora. Tübingen: Gunter Narr.
7
8
Mark Kaunisto
Davies, Mark. 2010. More than a peephole: Using large and diverse online corpora. International Journal of Corpus Linguistics 15(3): 412–418. Dollinger, Stefan. 2016. On the regrettable dichotomy between philology and linguistics: Historical lexicography and historical linguistics as test cases. In Studies in the History of the English Language, VII: Generalizing vs. Particularizing Methodologies in Historical Linguistic Analysis, Don Chapman, Colette Moore & Miranda Wilcox (eds), 61–89. Berlin: Mouton de Gruyter. Hiltunen, Turo, McVeigh, Joe & Säily, Tanja. 2017. How to turn linguistic data into evidence? In Big and Rich Data in English Corpus Linguistics: Methods and Explorations [Studies in Variation, Contacts and Change in English 19], Turo Hiltunen, Joe McVeigh & Tanja Säily (eds). Helsinki: eVarieng. (19 May 2024). Hoffmann, Sebastian, Kytö, Merja, Nevalainen, Terttu & Taavitsainen, Irma. 2018. A tribute to Matti Rissanen. ICAME Journal 42: 5–8. Labov, William. 1994. Principles of Linguistic Change, Vol. 1: Internal Factors. Oxford: Blackwell. McEnery, Tony & Wilson, Andrew. 2001[1996]. Corpus Linguistics: An Introduction. 2nd edn. Edinburgh: EUP. Mindt, Ilka. 2017. Analyzing corpus data from within. Anglistik 28(1): 57–73. Moskowich, Isabel & Crespo, Begoña. 2023. Stance in the Corpus of English Life Sciences Texts: A vindication of text-reading. Language Value, 16(1): 1–22. Mukherjee, Joybrato. 2004. The state of the art in corpus linguistics: Three book-length perspectives. English Language and Linguistics 8(1): 103–119. Rissanen, Matti. 1989. Three problems connected with the use of diachronic corpora. ICAME Journal 13: 16–19. Rütten, Tanja. 2017. Introduction. Anglistik 28(1): 7–12. Sinclair, John. 1991. Corpus Concordance Collocation. Oxford: OUP. Viana, Vander, Zyngier, Sonia & Barnbrook, Geoff (eds). 2011. Perspectives on Corpus Linguistics [Studies in Corpus Linguistics 48]. Amsterdam: John Benjamins. Weisser, Martin. 2016. Practical Corpus Linguistics: An Introduction to Corpus-Based Language Analysis. New York NY: John Wiley & Sons.
Engaging with bad (meta)data in historical corpus linguistics Turo Vartiainen & Tanja Säily University of Helsinki
In this chapter, we discuss some common pitfalls related to historical data and its use in linguistic analysis. We argue that the “philologist’s dilemma”, as originally proposed by Rissanen (1989), should be reconceptualized to meet the needs of the fast-evolving field of corpus linguistics, where scholars make increasing use of big-data resources and sophisticated statistical modelling. By providing examples of errors and uncertainties related to, for example, corpus metadata, sampling, balance, and OCR accuracy, we argue that corpus linguists should pay increasingly close attention to the sampling and annotation principles employed in the compilation of historical corpora as well as to the quality of the linguistic data. We propose that the principle of “knowing one’s corpus” in terms of its compilation principles has become all the more important in the age of bigdata corpora, where it is not feasible for individual researchers, or corpus compilers, to validate their data manually. Keywords: historical corpus linguistics, metadata, part-of-speech annotation, big data, corpus compilation, sampling
1.
Introduction
When Professor Matti Rissanen wrote a short article entitled “Three problems connected with the use of diachronic corpora” for the ICAME Journal in 1989, the landscape of corpus linguistics looked quite different from today. At the time, Rissanen was leading a group of scholars responsible for the compilation of the first diachronic corpus of English, the Helsinki Corpus of English Texts. The Helsinki Corpus, which was to be published in 1991, would revolutionize the historical research of English for decades to come, and its importance for historical corpus linguistics cannot be overstated. Against this background, it is interesting to note that Rissanen felt compelled to discuss some potential pitfalls related to historical corpus linguistics two years prior to the publication of the Helsinki Corpus. https://doi.org/10.1075/scl.118.02var © 2024 John Benjamins Publishing Company
10
Turo Vartiainen & Tanja Säily
From a present-day perspective, some of the issues raised by Rissanen have been addressed, while others remain a source for concern. Furthermore, there are some new trends in corpus-based research that Rissanen could not have anticipated in his article, and these introduce challenges of their own. In this paper, we revisit Rissanen’s ideas in order to identify some of the present-day pitfalls in historical corpus linguistics. We pay particular attention to what Rissanen (1989: 16) calls “the philologist’s dilemma” and how this dilemma might be recast from a present-day perspective. Originally, the philologist’s dilemma referred to a situation where the linguist’s work becomes increasingly detached from the language of the texts under study, as well as from the history, culture, and customs of the communities in which the texts were produced. In other words, the danger according to Rissanen was that the study of linguistic variation and change, with all its complexities, would be reduced to the reading of concordance lines on a computer screen. Such a narrow focus would have obvious negative effects both on the researcher’s mastery of the older forms of the language, and thus on their analysis of the data, as well as on their ability to contextualize their results with respect to the author, genre, or society, etc. When Rissanen discussed the philologist’s dilemma, it was impossible to anticipate the scope of the changes in corpus-based historical linguistics that would take place in the coming decades. For example, the advent of historical mega-corpora (on the scale of hundreds of millions or even billions of words), and the increased interest in statistical approaches to language variation and change within the research community, have in some ways shifted the focus of research even further away from the original texts and the culture in which they were produced. Today, many quantitatively-oriented historical linguists work primarily with frequencies and probabilities but might not always make enough room for detailed linguistic analysis and interpretation (cf. Larsson, Egbert & Biber 2022). Furthermore, the sheer size of the historical mega-corpora and databases makes it impossible for the compilers of these resources to collect the texts manually, which may raise questions about their representativeness, balance, and sampling accuracy. However, it is not our intention to advise against the use of big-data corpora or sophisticated statistical methods in historical corpus linguistics. Having larger corpora at our disposal is arguably one of the most important advances of the recent decades, and the more statistically-oriented research projects have permitted the analysis of highly complex questions pertaining to language change with increased quantitative rigour (Szmrecsanyi et al. 2014; Petré & Anthonissen 2020). Instead, we wish to draw the reader’s attention to the fact that the introduction of modern resources and applications has introduced new problems. Some of these can be resolved during the research process, but others may be harder to tackle. By dis-
Bad (meta)data in historical corpus linguistics
cussing examples of specific methodological challenges and their potential resolutions below, we aim to illustrate that the philologist’s dilemma is no less topical today than it was in the late 1980s. However, we also argue that the researcher’s detachment from the texts under study has become slightly different in nature. In particular, we argue that it is not possible for linguists who make use of megacorpora in their research to know their data as well as those using smaller, “tidier” corpora (see also Hundt & Leech 2012). Consequently, it is typically the case that present-day corpus linguists working with very large corpora need to accept a degree of data-related uncertainty that would have been unacceptable for scholars like Matti Rissanen in the early days of historical corpus linguistics. Accordingly, we propose that the changing focus from detailed analyses of small datasets to more abstract and statistically-oriented analyses of big data requires a reconceptualization of the philologist’s dilemma. While the original formulation of the philologist’s dilemma is no less relevant for scholars engaged in, say, historical sociolinguistics, we argue that in more grammatically oriented studies that make use of big data, increasing attention should be paid to the corpus itself: to its sampling principles and layers of annotation. Indeed, while the handling of linguistic big data always includes a degree of uncertainty, there is no excuse for ignoring the potential problems related to these resources. We acknowledge that some of the issues that we discuss in this paper may not be specific to historical corpora (see, e.g, Sinclair 2004: 191), but we feel that they are particularly common in historical research, and we therefore restrict our discussion to historical corpora and databases. After this short introduction, we proceed directly to our examples of presentday pitfalls in historical corpus-based research. Each of the following sections is intended to illustrate one general problem associated with historical corpus linguistics by discussing individual studies that we have carried out previously. In Section 2, the focus is on part-of-speech (POS) annotation, which can be regarded as one of the most useful kinds of linguistic annotation in contemporary research, but which, as we shall see, does not come without problems. Section 3 focuses on some pitfalls related to big corpora, and our examples in this section concern both the reliability of the semi-automated sampling of such resources and the comparability of the research results when new genres are introduced to the corpus. In Section 4, the focus is on linguistic databases, and on Eighteenth Century Collections Online, in particular. Here, we discuss issues related to the reliability of the database with respect to OCR (Optical Character Recognition) accuracy as well as problems pertaining to the balance of the database and the metadata associated with it. Section 5 brings the chapter to a close with a discussion of the main topics and some final conclusions.
11
12
Turo Vartiainen & Tanja Säily
2.
POS annotation in diachronic datasets
Part-of-speech (POS) annotation arguably provides one of the most useful layers of linguistic annotation to assist the researcher in the retrieval of relevant constructions from corpus data. Having each word in the corpus tagged according to its part of speech (i.e., word class) permits, for instance, the study of partiallyfilled constructions (e.g., N-PROP of N-PROP; “the Einstein of Italy”) and of lexical items of a specific part of speech (e.g., love_V; “I love you”).1 The words in a corpus are typically also annotated at the level of POS subcategories, so that nouns can be categorized into common and proper nouns, as above, while verbs can be classified into present and past tense forms, for example. Most of the widely-used historical and present-day corpora of English make use of an automated tagging software from the CLAWS (Constituent Likelihood Automatic Word-tagging System) family, first developed at the University of Lancaster in the early 1980s. Many versions of the CLAWS tagset have been introduced over the years, with the most recent edition, C8, published in 2001. Although C8 contains over 160 different tags, the coding scheme has been kept relatively simple and uncontroversial from a theoretical perspective: categories such as “noun” and “verb”, “singular” and “plural”, or “present tense” and “past tense” are not likely to pose substantial problems for researchers in data recall – or so one would think. However, applying software developed for Present-day English to historical data is somewhat problematic (Rayson et al. 2007). First, the tagger may only recognize modern spellings, which means that the high degree of spelling variation in historical data (as well as optical character recognition errors; see Section 4.2 below) may lead to reduced accuracy. Second, language change often leads to changes in the functions of individual words, and this may have repercussions on the way they are categorized into word classes. As a consequence, the same word should be assigned a different POS tag in the corpus at different periods, and this poses a potential methodological problem if the researcher is not familiar with the way the corpus is annotated.
2.1 Accounting for category change One of the areas of research where POS annotation might fail is category change, that is, a process where a word of one word class gradually begins to be used like a word from another class (see, e.g., Denison 2001; Aarts 2007). Two common examples of category change are key and fun, which are originally nouns, but which have gradually become more adjective-like in recent English (e.g., Denison 1. Here N-PROP stands for ‘proper noun’, whereas V stands for ‘verb’.
Bad (meta)data in historical corpus linguistics
2013: 160–164). For instance, the adjectival uses of the words in Examples (1) and (2), taken from the Corpus of Historical American English (COHA; Davies 2010), seem perfectly normal to many speakers of Present-day English, even though the words appear in contexts where adjectives are typically found: they are used in attributive and predicative functions, respectively, and they are modified in degree with very and extremely. (1) I think that the role of the Ambassador in the Soviet Union is a very key one. (COHA, News, 1964) (2) It’s extremely fun to imagine kids in sneakers running through the halls where corsets and petticoats used to be worn, or playing hide and go-seek in the palace’s gardens. (COHA, Mag, 2019)
Both key and fun in the above examples are correctly tagged as adjectives in COHA. Whether or not the tag has been probabilistically assigned by the tagger or manually inserted as a rule is not obvious, but at this stage we can note that the tagger can correctly identify at least some usages of these items as adjectival. Figure 1 provides a visualization of data extracted from COHA with the queries key_j and fun_j, which target all forms of key and fun that have been assigned an adjective tag by the POS tagger. The figure suggests that the adjectival uses of these items have become more frequent in American English in the course of the twentieth century, and the trendlines suggest that their popularity as descriptive items has not yet peaked.
Figure 1. Frequency of key and fun as adjectives in COHA (normalized to 1,000,000 words)
13
14
Turo Vartiainen & Tanja Säily
However, on a closer look it soon becomes clear that as convincing as the trends in Figure 1 look, something is not right. First, contrary to what has been established in previous research (De Smet 2012: 622–624), the data in Figure 1 suggest that the adjectivization process of key and fun started to pick up pace in the 1930s and the 1940s, which is some decades earlier than observed in, for example, De Smet (2012: 624). Second, the normalized frequencies are roughly ten times higher than reported in De Smet (2012), which raises suspicion that the POS tagger has not identified the adjectival uses of these two items correctly. Indeed, when taking a closer look at the results, it becomes obvious that while the tagger has correctly managed to identify some forms that can be labelled “adjectival”, such as Examples (3) and (4), most instances represent types where key and fun are used referentially, and should therefore be tagged as nouns, as in (5) and (6). (3) This sets up a situation where, in fact, nutrition is key.
(COHA, Mag, 2001)
(4) The whole tent was very fun, very lively.
(COHA, Mag, 2005)
(5) With a key ring in his hand, he went to the door of 7D.
(COHA, Fic, 1966)
(6) We’re going to have fun with this craft before we leave it!
(COHA, Fic, 1911)
In short, the tagger is ill-prepared to deal with the kind of gradual change that has taken place with key and fun, and consequently, the POS tags are of little use to the researcher when collecting the data from the corpus. The situation becomes even more complicated when the research focuses on items that are attested with a lower frequency, or on items whose adjectival use has not yet become conventionalized to the same degree as key and fun. In (7) to (11), taken from COHA’s sister corpus, the Corpus of Contemporary American English (COCA), all the “nouns” in boldface are used in a descriptive and non-referential function (i.e., similarly to key and fun in Examples (1) to (4) above), and yet none of them has been annotated as an adjective in the corpus. (7) Heard you gave a killer sermon last weekend. (8) It’s been a rubbish week all round, all in all.
(COCA, Fic, 2011) (COCA, Blog, 2012)
(9) This was some very dynamite, for lack of a better adjective, information, if it were true. (COCA, News, 2003) (10) Everybody was doing their job. It was textbook.
(COCA, Spoken, 2001)
(11) The construction project was mammoth by the standards of the day. (COCA, Mag, 2006)
Bad (meta)data in historical corpus linguistics
Examples (12) and (13) below show that this problem does not only affect nounto-adjective category change, but category change generally. Here, killer and textbook are used in functions where degree adverbs (e.g., very and really) are typically found, and yet such items are consistently tagged as nouns in the corpus. (12) He wasn’t just killer good-looking. He was to die for.
(COCA, Fic, 2008)
(13) Business meeting: Claudia Lester was textbook perfect.
(COCA, Fic, 2002)
To summarize, the pitfall related to POS annotation in the case of category change is that even though word classes are treated as static entities in corpus annotation, they are in fact dynamic; word classes exhibit both category-internal and category-external gradience, and this gradience may be a result of ongoing language change. This problem is of course not limited to historical corpora but also affects present-day corpora when the process of change has not reached its conclusion. While software like CLAWS also provide probabilistic information about word class, in practice this information is typically not available to the end users of the corpus. A possible solution to these kinds of problems is to make use of queries that not only combine lexical items with POS tags (e.g., key_j, fun_j) but also target the surrounding context of the item under study.
2.2 Theoretical choices in the design of the annotation scheme There are also cases where it may be impossible to make use of POS annotation because of the idiosyncratic tagging of the corpus. For example, in Säily, Vartiainen and Siirtola (2017) we were interested in establishing whether POS annotation can be used as a tool for studying sociolinguistic variation and genre evolution in the Parsed Corpus of Early English Correspondence (PCEEC). We focused particularly on the question of whether the genre of personal correspondence showed signs of colloquialization in the Early Modern period by measuring the incidence of different parts of speech that have been argued to be particularly frequent in informal speech, such as personal pronouns and lexical verbs, as well as the frequency of parts of speech associated with phrasal complexity and informational orientation, such as nouns and prepositions (e.g., Biber & Finegan 1997). We hypothesized that if personal correspondence had become increasingly speech-like in the period under study (c.1410 to 1681), we would see an increase in the proportions of personal pronouns and lexical verbs and a decrease in the proportions of nouns and prepositions out of all POS categories in the corpus. However, our approach was partly unsuccessful because of the theoretical choices made in the tagging of the corpus (see Taylor & Santorini 2006). Most significantly, the annotation of prepositions in the PCEEC was based on Huddleston
15
16
Turo Vartiainen & Tanja Säily
and Pullum’s (2002) analysis, according to which subordinating conjunctions should be categorized as prepositions with clausal complements. Normally, it makes sense to include this information in the POS tag, in other words, to annotate words as either prepositions or subordinating conjunctions depending on whether the complement is a noun phrase or a clause, respectively. However, because the PCEEC is syntactically parsed, the corpus compilers were able to include this information in the parsing trees (see also López-Couso & Méndez-Naya 2020 for problems related to the retrieval of “minor declarative complementizers”, i.e., adverbial connectives that are infrequently used as complementizers). Whether or not Huddleston and Pullum’s analysis is ideal from the perspective of word class theory is not relevant in this case; the important thing is that the PCEEC is annotated in a way that makes it impossible to study colloquialization by measuring the frequency of prepositions in different time periods. As pointed out in Biber and Gray (2011), prepositional phrases (in the traditional sense) are connected with NP complexity and informational orientation in text production, while subordinating conjunctions are related to a speech-like style. When the annotation conflates these two functions, colloquialization can only be studied from the perspective of preposition usage by resorting to lexical queries and going through thousands of concordance lines manually, or by retagging the entire corpus.
2.3 Annotation tailored to specific research questions Our third example of a potential POS-related pitfall pertains to a case where the corpus compilers were themselves working on the long diachrony of English and used a conservative annotation scheme to facilitate comparability over time. The PCEEC is one of the parsed corpora of historical English produced in collaboration between the universities of Penn, Helsinki, and York, among others. These corpora are intended to cover all stages of the history of English, and as such, the annotation scheme has been designed to be backwards compatible all the way to Old English. This is of course a laudable aim and makes the corpora invaluable for research on historical syntax. However, it severely complicates the study of certain types of research questions that rely on POS annotation, as will become apparent below. In our case, we were interested in examining changes in the POS frequencies in the PCEEC over time. In our first explorations, we focused on the frequency of nouns and personal pronouns in the corpus (Säily, Nevalainen & Siirtola 2011). These categories have been associated with an informational orientation and a personally involved style, respectively (e.g., Biber & Finegan 1997). Our goal was to assess the stability of the correspondence genre in this respect and to see if any changes observed might be influenced by social factors like gender, as earlier stud-
Bad (meta)data in historical corpus linguistics
ies on Present-day English had observed gender differences in the use of nouns and pronouns (e.g., Rayson, Leech & Hodges 1997). The annotation scheme of these POS categories in the PCEEC is simple, at least in principle: nouns receive tags beginning with N (divided into common and proper, singular and plural, and possessive), while pronouns are tagged with PRO (personal) or PRO$ (possessive). In addition, there are a number of compound tags with a noun tag as the first or the last component, and this is where our troubles began. We initially assumed that any compound tag ending in a noun tag could be counted as a noun. Sometimes this was indeed the case, as in sixpence_NUM+N. At other times, however, these represented other parts of speech that had grammaticalized from nouns (see, e.g., Peitsara 1997), such as adverbs (otherwise_OTHER+N), ‘quantifiers’ (anything_Q+N), subordinating conjunctions (because_P+N) and reflexive pronouns (himself_PRO+N). In the long diachrony, it makes perfect sense to annotate these words like this to be able to retrieve them with the same tags throughout, but for our purposes and the time periods we were analysing, it would have been misleading to count these instances as nouns, as the grammaticalization process had already been completed (or mostly completed; see Säily, Nevalainen & Siirtola 2011: 185). To resolve these issues, we ended up retagging the entire corpus in order to move the grammaticalized items from nouns to their current parts of speech. This was highly labour-intensive, especially as some of the compound tags were ambiguous in that they included both genuine instances of nouns, as in gentleman_ADJ+N, as well as adverbs, as in likewise_ADJ+N; there were 1,199 instances of ADJ+N alone. Moreover, because the corpus contains private letters from a period before English was standardized, there is a great deal of spelling variation in the corpus. Consequently, it was not enough to simply list all adverbs etc. that could have been tagged as nouns; instead, we had to identify all their spelling variants as well, including those that had been written as two words (e.g., like_ADJ wise_N). These tokens were combined and reclassified, so the retagging involved changes in tokenization, too. Time-consuming though this process was, it enabled us to gain interesting and reliable results, which we refined and augmented in our study discussed above (Säily, Vartiainen & Siirtola 2017).
17
18
Turo Vartiainen & Tanja Säily
3.
Large corpora
3.1 Inaccuracies in text sampling In this section, we give an example of a potential pitfall related to the semi-automated sampling of corpus texts. Here, it is not our intention to argue that the texts included in big-data corpora should be sampled manually; indeed, when the corpus includes samples from tens or even hundreds of thousands of individual texts, it is obviously not feasible to check all the data sources individually. Consequently, big-data corpora are likely to contain errors pertaining to the dating of some of the texts, their genre, and even the language variety, which the researcher must be aware of. Needless to say, even though some amount of data-related noise may be tolerable in the analysis, sampling errors like these can sometimes lead to disastrous results if one does not exercise sufficient care in data collection and analysis. As an example of a pitfall related to text sampling, we discuss a particular usage of the complex adverb as well, which caught our attention some time ago. While as well is commonly used in a clause-final position in cases like (14) and (15), we were interested in studying what seemed to be an emerging sentenceinitial use of as well in American English, exemplified in (16). Prior to performing our corpus queries, we carried out a preliminary literature survey of English connectives, but were unable to find mentions of this usage in some of the major works on the topic (e.g., Lenker & Meurman-Solin 2007; Lenker 2010). Therefore, a more detailed examination seemed warranted. (14) Not only are we stranded on this dreadful planet, but we are starving as well. (COHA, TV/Mov, 1968) (15) This interaction with readers continues not only through the mail but on-line as well. (COHA, NF/Acad, 1995) (16) These communities have significant African-American and HispanicAmerican populations, among others. As well, they are largely low-income communities. (COHA, NF/Acad, 1994)
We consequently decided to investigate the frequency of sentence-initial as well in COHA, a corpus that provided us with enough material to examine its diachronic development in recent American English. According to our corpus queries, the frequency of the usage was generally very low in the corpus, but it was clearly becoming more common in the late twentieth century (Figure 2).2 2. We carried out our queries on an older version of COHA in 2015. The search string included all sentence-final punctuation marks (. ! ? ; :) followed by as well and a comma.
Bad (meta)data in historical corpus linguistics
Figure 2. As well as a sentence-initial connective in COHA, 1830–2009 (normalized to 1,000,000 words)
We also obtained data on sentence-medial and sentence-final uses of as well, noting a substantial increase in both positions: the frequency of the sentencemedial use had increased from 0.32 tokens per million words (pmw) in 1830–1859 to 2.00 tokens pmw in 1980–2009, while the frequency of the sentence-final use had gone up from 2.87 pmw to 50.86 pmw during the same period. An exploration of the Corpus of Contemporary American English, on the other hand, showed that as well was particularly associated with spoken language in recent American English. To top it all, we also found evidence of two sentence-initial uses of as well that might have functioned as bridging contexts in the development of the sentenceinitial connective use. In (17), the distance between we well know and as well we know is long enough to possibly allow for a connective reading (‘moreover’) instead of a qualitative reading (‘we know equally well’). In (18), on the other hand, as well is short for might as well. From a semantic perspective, the usage in (18) is perhaps not as convincing an example of a bridging context as (17), but the formal similarity may have been enough to support reanalysis. Importantly for our preliminary analysis, the frequency of these potential bridging contexts and the frequency of the sentence-initial connective uses converged in an interesting way in the corpus: we see the frequency of the bridging contexts decrease as the frequency of the sentence-initial connective uses increases (Figure 3).
19
20
Turo Vartiainen & Tanja Säily
(17) We well know, that the Colonists are charged by many persons in Great Britain, with attempting to obtain such an exclusion of any power of parliament over these Colonies and a total independence on her. As well we know the accusation to be utterly false. (COHA, Mag, 1844) (18) An attempt is made to bind up and fetter this country’s expanding energies and prospects with Old World theories and methods. As well attempt to put baby-clothes upon the statue of Hercules. (COHA, Mag, 1874)
Figure 3. The frequency of sentence-initial connective uses and possible bridging contexts in COHA, 1830–2009 (normalized to 1,000,000 words)
Considering all this evidence, we were optimistic that the phenomenon was indeed real and that we were studying a usage that had thus far been ignored in the literature. We were only concerned about the low frequency of the form: in the last period studied (1980–2009), there were 32 tokens of the sentence-initial usage in COHA, which meant that our results might have been tarnished by errors made in the sampling of the corpus. Unfortunately, this is exactly what had happened. Indeed, our initial enthusiasm was quickly dampened by the observation that the connective use of as well was dispersed very unevenly in the corpus. On closer inspection, we were able to establish that 13 of the 32 tokens, all from the 1990s and the 2000s, were in fact produced by Canadian authors. Six additional tokens in the 1980s sample were produced by an author from Portland, Oregon, and one from Seattle, Washington, which raised strong suspicion of Canadian influence. Finally, the author of one sentence-initial usage turned out to be Irish. When these data are excluded from the results, the frequency of the form is so negligible that
Bad (meta)data in historical corpus linguistics
the argument for an incoming connective use of as well in American English can no longer be maintained. Now that we knew that the connective use was probably due to sampling errors in COHA, we conducted a new literature survey with a focus on Canadian English (CanE). We promptly found brief mentions of the sentence-initial usage in Denison (1998: 242) and Brinton & Fee (2001: 432), who simply noted that as well is commonly used in a sentence-initial position in CanE. To confirm beyond any doubt that the form is particularly associated with Canadian English, we compared its frequency across COCA (AmE), the British National Corpus, and the Strathy Corpus (CanE). Figure 4 leaves little room for doubt: the result that we had acquired from COHA was nothing more than a mirage caused by inaccuracies in the automated sampling procedure.
Figure 4. Frequency of sentence-initial as well in the British National Corpus (BNC; BrE), COCA (AmE), and the Strathy Corpus (CanE) (normalized to 1,000,000 words)
While this small-scale project ended in failure, we were grateful for the lesson it taught us: you should always try to gain an understanding of the texts behind the concordance lines. In our case, we were fortunate to be studying a lowfrequency item because the situation would have been far worse if we had examined an item with a token frequency measured in the thousands; in such a case, we would probably have drawn false conclusions, as it might not have been feasible to carry out a detailed investigation of the form’s dispersion.
21
22
Turo Vartiainen & Tanja Säily
3.2 Changes in the balance of subgenres Our second example pertains to genre balance, and in particular to the effect of including new subgenres in a corpus over time. In Säily and Vartiainen (forthcoming), we analysed a relatively recent change in the modification patterns of -ed participles from the more verbal much -ed to the more adjectival very -ed (e.g., Denison 1998; Vartiainen 2021), as in Examples (19) and (20). (19) He has been much interested in your movements – quite anxious about your return. (COHA, Fic, 1846) (20) We are very pleased with the court’s ruling.
(COHA, News, 2017)
Because we were interested in the potential impact of gender on the change, we chose as our dataset the fiction section of COHA, which contains named authors for whom gender metadata has been generated by Öhman, Säily and Laitinen (2019). We treated much/very -ed as a linguistic variable and calculated the proportion of different participle types representing the incoming variant, very -ed, out of all very -ed and much -ed types. We used types rather than tokens to safeguard against the possibility that the data could be skewed by a few frequent participles, such as interested. Figure 5 shows our most intriguing result: in the final decades of the twentieth century, men seem to start lagging behind in the change. However, when we began to look for potential reasons for this lag, we noticed that the internal balance of the fiction subcorpus changes over time. While most of the corpus consists of novels, the most recent periods in the corpus include increasing amounts of drama, short stories and movie scripts, for example. When we restricted our corpus queries to novels alone, which is possible if one uses the downloadable version of COHA rather than the online interface, the gender difference disappeared. This example serves as another illustration of the pitfalls of mega-corpora that are not as carefully sampled and balanced as smaller corpora. Of course, the issue of genre balance also affects smaller diachronic corpora like the Helsinki Corpus, because some genres may disappear and new ones appear over time, and compromises must be made in the compilation process in terms of representativeness vs. diachronic comparability. In the same vein, it is completely justified to include movie scripts in the fiction section of COHA as movies begin to be made in the twentieth century. On the other hand, the history of many fiction genres, including drama and short stories, is certainly longer than the coverage of COHA, so their proportions could have been more optimally balanced in the corpus; considering the automated sampling of COHA, their better representation in the most recent periods probably reflects their increased availability in digital form.
Bad (meta)data in historical corpus linguistics
Figure 5. Proportion of very -ed types out of very -ed and much -ed types in the fiction section of COHA over time (solid line = men; shaded regions = typical proportion in each period in 10,000,000 random subcorpora of the same size). Reproduced from Säily and Vartiainen (forthcoming: Figure 5)
4.
Historical databases
4.1 Issues with balance and metadata In recent years, corpus linguists have been increasingly interested in making use of massive historical databases, such as the Eighteenth Century Collections Online (ECCO), in their research. ECCO consists of more than half of all known British publications in the eighteenth century, or about 200,000 texts and 10.5 billion running words (Tolonen et al. 2021; Hill & Hengchen 2019). The huge size of ECCO offers many possibilities for linguistic research, but because it was not originally compiled as a corpus, there are some significant drawbacks related to its use. One of the biggest issues with ECCO is the lack of available information on the source texts: the metadata provided by Gale, the company that owns and distributes the database, is minimal, and we have little idea about the balance and representativeness of ECCO in terms of genres, for example. Fortunately, there is ongoing research that aims to shed more light on what exactly the database contains. For example, Tolonen et al. (2021) compared ECCO
23
24
Turo Vartiainen & Tanja Säily
with the English Short Title Catalogue (ESTC), which provides comprehensive bibliographic metadata on eighteenth-century British publications. Even though ECCO allegedly represents “every significant English-language and foreign-language title printed in the United Kingdom between the years 1701 and 1800”,3 we should not fall victim to the “God’s truth fallacy” (Rissanen 1989: 17) and assume that it completely accurately reflects the reality of eighteenth-century publishing. Moreover, up to 31% of the texts are reprints or new editions of earlier, often wellknown works, which implies highly problematic biases and complicates the study of real-time language change (Tolonen et al. 2021). With the help of the ESTC metadata, harmonized and augmented by the Helsinki Computational History Group (COMHIS), the ECCO dataset can be narrowed down to first editions only, which greatly facilitates linguistic research. The metadata also provides opportunities for generating principled subcorpora of ECCO, such as economic literature (Liimatta et al. 2023a); this subsetting makes it feasible to get to know the data in more detail, and scholars can take samples to go through exhaustively in order to better understand their results. As a further mitigation of the “God’s truth fallacy”, Rissanen’s (1989: 17) suggestion of keeping the corpus open-ended for future improvements has in a sense been taken up by Gale, although the sampling of each part has not followed corpus-linguistic practices. Part I of ECCO was published in 2002, Part II (50,000 titles focusing on literature, social science, and religion) following later in 2009, and there are plans to publish a third part (90,000 titles) in the future (Tolonen et al. 2021).
4.2 OCR errors Crucially, the digitized texts in ECCO do not even represent “God’s truth” about the original publications on which they are based, owing to severe errors in optical character recognition (OCR) in the digitization process. These errors, the rate of which varies between Parts I and II (Tolonen et al. 2021), are so prevalent that it is often impossible for a human to make sense of the digitized text. Nevertheless, recent research has shown that many quantitative methods common in corpus linguistics and the digital humanities can still be applied to the database at a reasonable level of accuracy (Hill & Hengchen 2019; Liimatta et al. 2023b). Some areas of linguistic inquiry may prove more difficult to pursue than others, however.
3. 〈https://www.gale.com/primary-sources/eighteenth-century-collections-online〉
Bad (meta)data in historical corpus linguistics
4.2.1 Hapax legomena To give a concrete example, many studies of lexical and morphological productivity rely on accurate type counts, particularly of hapax legomena, that is, words that only occur once in the corpus. Here, the small subset of ECCO that has been keyed in manually by the Text Creation Partnership (ECCO-TCP), and which is hence free from OCR errors, provides an illustrative example. Comparing the texts shared by ECCO and ECCO-TCP, we find that more than 90% of the hapax legomena in the ECCO sample (henceforth ECCO-OCR) are not found in ECCO-TCP. This means that they are in fact spurious OCR errors and that the precision of queries based on rarity is abysmal (Figure 6). On the other hand, more than half of the hapax legomena in ECCO-TCP do not occur in ECCOOCR, meaning that recall is also poor and that our query is able to retrieve less than half of the actual instances (Figure 7).
Figure 6. Proportion of types in ECCO-OCR that are not found in ECCO-TCP, by frequency of occurrence
While this means that studying the entire ECCO based on hapax legomena is unfeasible both because the results of the queries are unreliable and because it is humanly impossible to clean them, taking a small enough subset that could be gone through manually would resolve the problems related to precision. To reduce the amount of manual labour required, and hence to increase the potential size of the subset, it might be possible to group the search results based on, for example, edit distance (roughly, by how many characters a word differs from another), so that spurious hapaxes that are simply longer words with an OCR error could be automatically mapped to the correct word. Recall would in this
25
26
Turo Vartiainen & Tanja Säily
Figure 7. Proportion of types in ECCO-TCP that are not found in ECCO-OCR, by frequency of occurrence
case remain poor, but if the errors were distributed equally enough, what is caught in the net might be sufficient for research. To improve recall, the researcher could expand the query from hapaxes to types occurring, for example, 1 to 10 times, which would also be theoretically justified in studies of productivity (cf. Baayen 1993: 195–196). In addition to these solutions, subsetting by the measure of OCR quality provided by Gale would enable the researcher to zoom in on cleaner data (Hill & Hengchen 2019), from which a sensible sample could then be taken based on other metadata. 4.2.2 Historical lexis A recent study of economic vocabulary in ECCO (Liimatta et al., 2023a) further illustrates the challenges posed by OCR. Firstly, normalized frequencies are the historical corpus linguist’s bread and butter, but owing to OCR errors, the tokenization in ECCO is so poor that the number of running words in the texts cannot be reliably estimated, which means that calculating normalized frequencies based on the number of running words is quite problematic. We resolved this issue by normalizing the data by character count rather than word count. This admittedly raises problems of its own because words vary in length, but it still provides more reliable estimates. The second issue relates to the fact that some characters are more likely to result in OCR errors than others: in particular, the long “s” and ligatures are more likely to be incorrectly identified. For instance, we aimed to study the spelling variants of economy to analyse the diachrony of the transition from the Latinate
Bad (meta)data in historical corpus linguistics
variants, “oeconomy” and “œconomy”, to “economy”. However, by taking a small sample of the instances and going back to the document images, we discovered that about a third of the instances of “economy” were in fact OCR errors for “œconomy”. This made accurate timing impossible, unless we were to take a more representative sample and check all the document images manually, which would have been a very labour-intensive task. Moreover, in our multi-dimensional analysis of economic vocabulary where we mapped the words in a subset of ECCO to the “trade and finance” section of the Historical Thesaurus of the Oxford English Dictionary, none of the words that emerged as significant had an “s” in them, suggesting that these might have been missed owing to the problem with the long “s”. Here one solution would be fuzzy searching, or at least adding variants with “f ” or “l” for “s” to the queries. The texts in ECCO were digitized and OCR’d in the 1990s and 2000s. Since then, OCR methods have vastly improved, so it is to be hoped that the document images – poorly scanned as they may sometimes be – could be re-OCR’d in the future, with much better results. These methods rely on deep learning and neural networks, which are becoming increasingly common in other linguistic applications as well, for instance word embeddings, which are used for lexical semantics. Deep learning is an area of computer science that is highly resource intensive, which means that the restrictions of hardware and software, already mentioned by Rissanen (1989: 18), again become an issue at this scale. Utilizing neural networks may call for access to a supercomputer, and high-performance computing of this kind requires special expertise. Here, historical corpus linguists may benefit from collaboration with data scientists and experts in natural language processing.
5.
Discussion and conclusion
In this paper, we have discussed pitfalls related to some of the widely used historical corpora and databases in the context of current corpus-linguistic research. Our main argument was that the recent trends and advances in the field, such as the publication of big-data corpora, the increased reliance on statistical approaches to linguistic data, and the exploitation of various kinds of metadata, require that the linguist should have a detailed understanding of the structure of the corpus, its sampling procedure, and the principles followed in the construction of the metadata. In other words, the increased use of big-data resources in historical research has resulted in fundamental changes in research methods, data analysis and research questions, and while these changes have provided researchers with exciting new possibilities, they have also presented new challenges.
27
28
Turo Vartiainen & Tanja Säily
We also suggest that these challenges are ultimately not so different from the pitfalls discussed by Rissanen in his “philologist’s dilemma”: just like in the early days of historical corpus linguistics, today’s challenges pertain to the linguist’s knowledge and understanding of their data. However, because the focus of research has in many cases shifted from small, carefully compiled corpora to bigdata resources, the nature of the “data” and methods with which historical corpus linguists typically work has changed. Consequently, we suggest that it would be prudent to think about the principle of “knowing one’s data” from a new perspective: as the sheer size of the modern mega-corpora prevents scholars from engaging with either all or the majority of the original source texts in as much detail as in the early days of historical corpus linguistics, they should strive for an intimate understanding of the historical corpora as mediators of the original texts. Admittedly, this perspective signifies a trade-off between two types of knowledge. On the one hand, the linguist must accept that while sociocultural contextualization remains as important as ever, it is not always possible to check every corpus text or concordance line manually. On the other hand, the linguist can compensate for this to some degree by acquiring a thorough understanding of the corpus and its description. This knowledge is more abstract in nature, and it is partly a reflection of the more abstract kinds of research questions explored in corpus-based research today, as well as of a shifting focus towards an increasingly statistical orientation in research design. This shift also means that researchers must be willing to accept a certain degree of data-related uncertainty or “noise”, which cannot always be controlled as well as when one works with smaller corpora. In the following, we discuss these issues and potential solutions to them in light of our case studies. Our first examples concerned problems related to part-of-speech annotation, and these challenges are hard to overcome entirely. Even when the coding schemes are intended to be theoretically as neutral as possible, the POS tags always add an analytical layer to the corpus, which reflects a particular theoretical stance towards word class categorization. Because of this analytical interference, Tognini-Bonelli (2001: 74) has criticized the use of POS annotation in corpus-based research, arguing that the reason for why POS-based research provides intuitively satisfactory results is that it is in line with the linguist’s preconceptions and categorizations of the data, and as such, can only offer a partial picture of language use. While we think that this criticism is probably too harsh, and we agree with Leech (1991: 15) that the benefits of corpus annotation far outweigh the disadvantages, there are also problematic cases that particularly pertain to language change, as shown by our examples on category change and the grammaticalized elements in the PCEEC. In such cases, the practical thing to do is first to read the corpus manuals carefully and then to experi-
Bad (meta)data in historical corpus linguistics
ment with different queries in order to gain a general understanding of how the corpus has been tagged from the perspective of the research question at hand. In our case study on as well in Section 3.1, the confounding influence of miscategorized Canadian sources could be controlled because of the low frequency of the form, but if the frequency of the form had been higher, the likelihood of our spotting the miscategorized texts would have been lower. However, based on the proportion of the Canadian texts that yielded a number of hits of as well in our query out of all the texts sampled in COHA, the effect of CanE on most research questions studied with COHA is likely to be negligible. Indeed, we wish to emphasize that resources like COHA are invaluable for diachronic linguistic research, and that these resources are even more useful when one knows exactly how they are constructed and annotated. Furthermore, the biases can be combatted methodologically, for example, by using “dispersion-aware” methods (Säily 2014: 46) that will alert the researcher if the instances are poorly dispersed, that is, concentrated in a very limited number of sources. The use of such methods, however, requires methodological expertise as well as access to text-level information on the numbers of running words and instances, which is not always readily available. In the case of changes in the balance of subgenres (Section 3.2), one solution is for corpus compilers to provide detailed, or at least basic metadata, that enables users to zoom in on maximally comparable subcorpora or to take balanced samples of their own. As also suggested by Rissanen (1989: 17), keeping the corpus open-ended makes it possible to improve the balance and representativeness of the corpus as new material becomes available, even though this will negatively impact comparability with previous research performed with the older versions of the corpus. When the corpus metadata turns out to be incorrect, as in our Canadian English case, corpus compilers could provide a channel for user feedback that would be taken into account in future updates of the corpus. In the case of COHA, both the corpus and its metadata were improved in an update that was made available in 2021 (Alatrash et al. 2020). In addition to problems with metadata, the noise present in the digitized texts in historical databases (Section 4) may lead to compromises in accuracy in corpus-linguistic research. While the precision of queries can always be increased by going through the hits manually or semi-automatically (taking a smaller sample of the full dataset if needed), striving for perfect recall in data afflicted by OCR errors may prove to be a doomed endeavour. As noted in Section 4.2.1, however, if the OCR errors are distributed relatively equally across the corpus data, settling for lower recall may be justified, as there is so much data that missing some of it does not significantly impact the results. Whether or not lower recall is acceptable ultimately depends on the research question.
29
30
Turo Vartiainen & Tanja Säily
The principle of knowing one’s data and being “on really intimate terms” with the data (Rissanen 1989: 16) becomes increasingly difficult to follow if the material consists of 10.5 billion running words, as in ECCO. On the other hand, researchers who have access to this dataset also have access to images of the original publications via Gale, which means that they have the “original editions” recommended by Rissanen (1989: 17) at their fingertips. The examination of the source texts can be time-consuming, but some time can be saved by computational means. For instance, a new interface based on the Octavo API, developed by Eetu Mäkelä at COMHIS, provides simultaneous access to the plain texts and document images, which allows the researcher to gain a better understanding of the data without expending too much time on individual texts. To conclude, we argue that all linguistic research conducted on historical corpora and text databases benefits from an understanding of not only the texts and the historical language variety but also from the sociocultural contexts in which the texts were produced. From the corpus compiler’s perspective, one way of contextualizing the texts is to provide relevant metadata to the end users (e.g., Menzel, Knappen & Teich 2021). Ideally, the metadata would be standardized across corpora to facilitate comparability, and sociocultural metadata would be developed in collaboration with historical (socio-)linguists and historians, which would make it maximally usable in the digital humanities.4 The end users can gain further knowledge by reading historical research on the topic as well as by collaborating with historians, which is becoming increasingly common among historical corpus linguists in general and is the modus operandi in the digital humanities (e.g., McEnery & Baker 2019; Hill & Hengchen 2019). Furthermore, while going through all of the search results may be unfeasible in the case of mega-corpora and while a degree of uncertainty is acceptable in more statisticallyoriented research projects in particular, we argue that it is still important for historical corpus linguists to base their qualitative analyses and interpretations of the data on a systematic examination of the actual texts.
Funding We gratefully acknowledge the financial support of the Research Council of Finland (grant numbers: 333944, 323390, and 333717).
4. We thank an anonymous reviewer for drawing our attention to these points.
Bad (meta)data in historical corpus linguistics
Acknowledgements We thank the anonymous reviewers and the editors of this volume for their helpful comments and suggestions.
References Aarts, Bas. 2007. Syntactic Gradience: The Nature of Grammatical Indeterminacy. Oxford: OUP. Alatrash, Reem, Schlechtweg, Dominik, Kuhn, Jonas & Schulte im Walde, Sabine. 2020. CCOHA: Clean Corpus of Historical American English. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk & Stelios Piperidis (eds), 6958–6966. Marseille: European Language Resources Association. Baayen, R. Harald. 1993. On frequency, transparency and productivity. In Yearbook of Morphology 1992, Geert Booij & Jaap van Marle (eds), 181–208. Dordrecht: Kluwer. Biber, Douglas & Finegan, Edward. 1997. Diachronic relations among speech-based and written registers in English. In To Explain the Present: Studies in the Changing English Language in Honour of Matti Rissanen [Mémoires de la Société Néophilologique de Helsinki 52], Terttu Nevalainen & Leena Kahlas-Tarkka (eds), 253–275. Helsinki: Société Néophilologique. Biber, Douglas & Gray, Bethany. 2011. The historical shift of scientific academic prose in English towards less explicit styles of expression: Writing without verbs. In Researching Specialized Kanguages [Studies in Corpus Linguistics 47], Vijay Bhatia, Purificación Sánchez Hernández & Pascual Pérez-Paredes (eds), 11–24. Amsterdam: John Benjamins. Brinton, Laurel & Fee, Margery. 2001. Canadian English. In The Cambridge History of the English Language, Vol. 6: English in North America, John Algeo (ed.), 422–440. Cambridge: CUP. COHA = Davies, Mark. 2010–. The Corpus of Historical American English: 400 million words, 1810–2009. (10 May 2024). Denison, David. 1998. Syntax. In The Cambridge History of the English Language, Vol. 4: 1776–1997, Suzanne Romaine (ed.), 92–329. Cambridge: CUP. Denison, David. 2001. Gradience and linguistic change. In Historical Linguistics 1999: Selected Papers from the 14th International Conference on Historical Linguistics, Vancouver, August 1999 [Current Issues in Linguistic Theory 215], Laurel Brinton (ed.), 119–144. Amsterdam: John Benjamins. Denison, David. 2013. Parts of speech: Solid citizens or slippery customers? Journal of the British Academy 1: 151–185. De Smet, Hendrik. 2012. The course of actualization. Language 88(3): 601–633.
31
32
Turo Vartiainen & Tanja Säily
ECCO = Eighteenth Century Collections Online. Gale. (19 May 2024). Hill, Mark & Hengchen, Simon. 2019. Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case study. Digital Scholarship in the Humanities 34(4): 825–843. Huddleston, Rodney & Pullum, Geoffrey K. (eds). 2002. The Cambridge Grammar of the English Language. Cambridge: CUP. Hundt, Marianne & Leech, Geoffrey. 2012. “Small is beautiful”: On the value of standard reference corpora for observing recent grammatical change . In The Oxford Handbook of the History of English [Oxford Handbooks in Linguistics], Terttu Nevalainen & Elizabeth Closs Traugott (eds), 175–188. Oxford: OUP. Larsson, Tove, Egbert, Jesse & Biber, Douglas. 2022. On the status of statistical reporting versus linguistic description in corpus linguistics: A ten-year perspective. Corpora 17(1): 137–157. Leech, Geoffrey. 1991. The state of the art in corpus linguistics. In Karin Aijmer & Bengt Altenberg (eds), English Corpus Linguistics. Studies in Honour of Jan Svartvik. London: Longman. Lenker, Ursula. 2010. Argument and Rhetoric: Adverbial Connectors in the History of English. Berlin: Mouton de Gruyter. Lenker, Ursula & Meurman-Solin, Anneli (eds). 2007. Connectives in the History of English [Current Issues in Linguistic Theory 283]. Amsterdam: John Benjamins. Liimatta, Aatu, Marjanen, Jani, Tahko, Tuuli, Tolonen, Mikko & Säily, Tanja. 2023a. Dimensions of incoming economic vocabulary in eighteenth-century Britain. Linguistica 63(1–2): 353–374. Special issue Sociocultural Change and the Development of Vernacular Languages in Early Modern Europe, Oliver Currie (ed.). Liimatta, Aatu, Ryan, Yann, Säily, Tanja & Tolonen, Mikko. 2023b. Results from rough data? The large-scale study of early modern historiography with multi-dimensional register analysis. In Proceedings of the Digital Humanities in the Nordic Countries 7th Conference, Oslo – Stavanger – Bergen, Norway, March 8–10, 2023 [DHNB Publications 5:1], Annika Rockenberger, Sofie Gilbert, Juliane Tiemann & Elisa Pierfederici (eds), 297–312. Oslo: University of Oslo. López-Couso, María José & Méndez-Naya, Belén. 2020. Masked by annotation: Minor declarative complementizers in parsed corpora of historical English. Research in Corpus Linguistics 8(2): 133–158. McEnery, Anthony & Baker, Helen. 2019. Language surrounding poverty in early modern England: A corpus-based investigation of how people living in the seventeenth century perceived the criminalized poor. In From Data to Evidence in English Language Research, Carla Suhr, Terttu Nevalainen & Irma Taavitsainen (eds), 225–257. Leiden: Brill. Menzel, Katrin, Knappen, Jörg & Teich, Elke. 2021. Generating linguistically relevant metadata for the Royal Society Corpus. Research in Corpus Linguistics 9(1): 1–18. Special issue Challenges of Combining Structured and Unstructured Data in Corpus Development, Tanja Säily & Jukka Tyrkkö (eds).
Bad (meta)data in historical corpus linguistics
Öhman, Emily, Säily, Tanja & Laitinen, Mikko. 2019. Towards the inevitable demise of everybody? A multifactorial analysis of one/-body/-man variation in indefinite pronouns in historical American English. Presentation at the 40th Annual Conference of the International Computer Archive of Modern and Medieval English (ICAME 40), Neuchâtel, Switzerland, June. (19 May 2024). PCEEC = Parsed Corpus of Early English Correspondence, tagged version. 2006. Annotated by Arja Nurmi, Ann Taylor, Anthony Warner, Susan Pintzuk & Terttu Nevalainen. Compiled by the CEEC Project Team. York: University of York and Helsinki: University of Helsinki. Distributed through the Oxford Text Archive. Peitsara, Kirsti. 1997. The development of reflexive strategies in English. In Grammaticalization at Work: Studies of Long-term Developments in English [Topics in English Linguistics 24], Matti Rissanen, Merja Kytö & Kirsi Heikkonen (eds), 277–370. Berlin: De Gruyter Mouton. Petré, Peter & Anthonissen, Lynn. 2020. Individuality in complex systems: A constructionist approach. Cognitive Linguistics 31(2): 185–212. Rayson, Paul, Archer, Dawn, Baron, Alistair, Culpeper, Jonathan & Smith, Nicholas. 2007. Tagging the Bard: Evaluating the accuracy of a modern POS tagger on Early Modern English corpora. In Proceedings of Corpus Linguistics 2007, 27–30 July, University of Birmingham, UK, Matthew Davies, Paul Rayson, Susan Hunston & Pernilla Danielsson (eds), article #192. (19 May 2024). Rayson, Paul, Leech, Geoffrey & Hodges, Mary. 1997. Social differentiation in the use of English vocabulary: Some analyses of the conversational component of the British National Corpus. International Journal of Corpus Linguistics 2(1): 133–152. Rissanen, Matti. 1989. Three problems connected with the use of diachronic corpora. ICAME Journal 13: 16–19. Säily, Tanja. 2014. Sociolinguistic Variation in English Derivational Productivity: Studies and Methods in Diachronic Corpus Linguistics [Mémoires de la Société Néophilologique de Helsinki XCIV]. Helsinki: Société Néophilologique. Säily, Tanja, Nevalainen, Terttu & Siirtola, Harri. 2011. Variation in noun and pronoun frequencies in a sociohistorical corpus of English. Literary and Linguistic Computing 26(2): 167–188. Säily, Tanja, Vartiainen, Turo & Siirtola, Harri. 2017. Exploring part-of-speech frequencies in a sociohistorical corpus of English. In Exploring Future Paths of Historical Sociolinguistics [Advances in Historical Sociolinguistics 7], Tanja Säily, Arja Nurmi, Minna Palander-Collin & Anita Auer (eds), 23–52. Amsterdam: John Benjamins. Säily, Tanja & Vartiainen, Turo. Forthcoming. Historical linguistics. In The Bloomsbury Handbook of Corpus Linguistics, Gavin Brookes & Michaela Mahlberg (eds). London: Bloomsbury. Sinclair, John. 2004. Trust the Text. Language, Corpus and Discourse. New York NY: Routledge. Szmrecsanyi, Benedikt, Rosenbach, Anette, Bresnan, Joan & Wolk, Christoph. 2014. Culturally conditioned language change? A multi-variate analysis of genitive constructions in ARCHER. In Late Modern English Syntax, Marianne Hundt (ed), 133–152. Cambridge: CUP.
33
34
Turo Vartiainen & Tanja Säily
Taylor, Ann & Santorini, Beatrice. 2006. The Parsed Corpus of Early English Correspondence. York: University of York. (19 May 2024). Tognini-Bonelli, Elena. 2001. Corpus Linguistics at Work [Studies in Corpus Linguistics 6]. Amsterdam: John Benjamins. Tolonen, Mikko, Mäkelä, Eetu, Ijaz, Ali & Lahti, Leo. 2021. Corpus linguistics and Eighteenth Century Collections Online (ECCO). Research in Corpus Linguistics 9(1): 19–34. Vartiainen, Turo. 2021. Trends and recent change in the syntactic distribution of degree modifiers: Implications for a usage-based theory of word classes. Journal of English Linguistics 49(2): 228–251.
Named entities as potentially problematic items in corpora Mark Kaunisto
Tampere University
This chapter discusses problems in the interpretation of corpus data arising from the insufficiencies in the annotation of named entities. Many corpora nowadays still do not adequately enable corpus users to set up queries that would exclude items appearing in names when needed to improve precision of the searches. Through an examination of case studies in major English language corpora, the chapter highlights the need to carefully post-process the search results, as irrelevant occurrences of named entities may pose challenges in the analyses of word frequencies and their collocational behaviour. The chapter calls for more detailed annotation of named entities in already available large linguistic corpora and reminds of the importance of close inspection of the search hits. Keywords: named entities, proper names, annotation, corpus linguistics
1.
Introduction
As we have reached the 2020s, corpus linguistics as a branch of linguistic study has progressed in many ways from its early days, nowadays offering a multitude of possibilities of examining various facets of language use. Not only have new corpora become available that are massive in size compared to those compiled in the 1980s and 1990s, but search engines or online interfaces have also become technically more and more sophisticated, giving users the opportunity to observe statistically organised search results that go far beyond the presentation of mere concordance lines. In concrete terms, the changes are noticeable on the level of educating students about the basics of corpus linguistics: with the evolving nature of the field, there is a constant need to update course materials to provide novices with an overview of both the possibilities and challenges relevant to corpus study. With new corpora containing billions of words and allowing complex searches by making use of various search functions and different types of linguistic annohttps://doi.org/10.1075/scl.118.03kau © 2024 John Benjamins Publishing Company
36
Mark Kaunisto
tation, novice corpus users may find it difficult to question the relevance of their search results adequately and carefully. In line with the “God’s truth fallacy” noted by Rissanen (1989), automatically produced frequency lists or rankings of statistically most significant collocates may be accepted without further scrutiny, and with relevant tokens numbering in the thousands, there may be a growing temptation not to check the concordance lines, as the corpus user feels secure about the reliability and representativeness of the corpus. There are many types of items commonly found in corpora which, while they are perfectly representative of language use, may present problems when investigating patterns of language use. For example, Rissanen (1992) paid attention to the role of quotations in historical sermons and observed that biblical quotations contained relatively higher proportions of pronouns compared to other parts of sermons. In a similar vein, Kaunisto (2017) examined the overall number of words from quotations as well as foreign language passages in Samuel Taylor Coleridge’s Biographia Literaria, and found that approximately 12 per cent of the entire word count of the book was made up of items which do not exactly represent Coleridge’s own choices of linguistic patterns; in some instances, the quotations represented language written hundreds of years before Coleridge’s time. In other words, texts may sometimes feature items which have the potential of skewing the results. When we use idioms, common proverbs, or quote someone, those sequences of words are more frozen, that is, we are bound to follow the patterns and choices of words made by other people. It is therefore of interest to examine more closely the characteristics of such items and to assess the seriousness of such effects on different types of linguistic analyses. There are also items which, because of their high frequency, pose other, quite serious challenges in corpus analyses, and the ubiquitous references to named entities can be identified as one such issue. Corpus users are probably well aware that items found in proper names – for example, words found in titles of books, articles, films, etc. – may be regarded as frozen items and would normally be excluded from further analysis because the choice of the exact words in such instances was usually made by someone else than the authors or speakers themselves. The identification of multi-word units functioning as proper names is a task that has received a great deal of attention from scholars to solve problems relating to different purposes – for example, data mining, automatic translation, named entity recognition – generally in the field of Natural Language Processing (NLP). Much work has gone into developing efficient annotation strategies for differentiating items belonging to named entities from those representing regular uses of the words. The task is obviously not straightforward, as names differ greatly as regards their structural complexity, and in some cases, they can be rather long, as in Example (1) from a review of a music album.
Named entities as potentially problematic items
(1) Originally titled ‘The Southern Harmony And Musical Companion Featuring A Choice Collection Of Tomes, Songs, Odes And Anthems From The Most Eminent Authors In The United States’, this is the record that was struggling to be heard inside the more studio-constricted ‘Money Maker’. (BNC, CHA 4244)
In this example, we can see that the title of the album is a lengthy one, and it can be observed that although the mention of the title in practice constitutes a single reference to only one entity, the 25 words of the name are similar to a quote in that none of the words and their combinations were originally produced by the author of the review. Proper names and multi-word proper names in particular pose challenges to the study of language use, and many currently available linguistic corpora are lacking in this kind of annotation. In automated collocational analyses, for instance, it may not always be possible to make use of part-of-speech tags to distinguish between instances of words where the word has been a part of a proper name or not. The situation is somewhat easier if we need to distinguish between proper nouns and common nouns, as far as we can rely on the accuracy of the tagging. But as regards, for example, the occurrences of adjectives within proper names – as in magical in Magical Mystery Tour – no annotation is necessarily available for us to separate such instances from regular uses of the adjective, and close manual inspection is therefore needed to exclude such items from the analysis. In this chapter, I will examine through different types of examples how the frequencies of English proper name uses can distort studies focussing on word frequency and collocational behaviour, and how the occurrences of proper name use may show different degrees of prominence of words in different genres and regional varieties. Section 2 provides the broader background in relation to English proper nouns, proper names, and the strategies of annotating names. Section 3 presents small-scale studies of corpus searches involving words that may be frequently used as named entities, or parts of named entities. Section 4 provides further discussion and concluding remarks. Overall, the main argument is that due caution must be given to such items and sufficient manual inspection of concordance lines is needed to avoid the possibility of misinterpreting initial findings from corpus data.
37
38
Mark Kaunisto
2.
Background
2.1 The concepts of proper nouns and proper names Many studies on the annotation of named entities start off with descriptions on how the concept of “name”, or “proper name” has been described and defined, going back to the theoretical postulations by philosophers such as John Stuart Mill, Bertrand Russell, and Ludwig Wittgenstein (see, e.g., Anderson 2004, 2007; Pierini 2008; Simon 2019). As noted by Pierini (2008: 43), proper names have been the subject of much study from the perspective of philosophy, logic, psychology, and anthropology. The semantic notions of sense, reference, and denotation are central in the understanding and description of what a name is, but in simplest terms – and for purposes most directly relevant here – we can start with the distinction between proper nouns and common nouns. Aarts et al. (2014, s.vv. proper noun; common noun) define proper nouns as nouns “referring to a unique person, place, animal, etc.”, for example, Tom, London, Dumbo, while common nouns are ones that are “not the name of any particular person, place, thing, etc.”, but which refer to a class of entities, for instance, janitor, street, reason. In addition to their semantic difference and the fact that proper nouns are usually spelled with a capital initial, proper nouns and common nouns can typically be distinguished by their grammatical behaviour. Unlike proper nouns, common nouns show contrasts in number and definiteness, but as noted by Biber et al. (2021: 243–244), there are also borderline cases as regards the two types of nouns and their grammatical behaviour (e.g., proper nouns can in some instances be used with possessive determiners, as in they want our Harry for the job, even though this feature is more typical of common nouns), and there are also instances where proper nouns can be used as common nouns (e.g., he’s our new Shakespeare; we have several Napoleons in our ward). Some grammarians (e.g., Huddleston & Pullum 2002; Aarts 2011) prefer to distinguish between the concepts proper noun, reserving it for single-word items (e.g., Bob, London, Japan), and proper names. The latter term covers proper nouns as well as names comprised of noun phrases with the definite article the and a proper noun (e.g., the Thames), the + a common noun (e.g., the Queen, the Plough), combinations of common and proper nouns (with or without the definite article, e.g., Lake Erie, Windsor Castle, the Guggenheim Museum), adjectives and common nouns (e.g., Great Britain, the United Nations, the British Isles), or longer, more complex noun phrases (e.g., the United States of America, The 7th Annual Global Congress of Knowledge Economy). Proper names may also include phrases other than noun phrases: for example, the film title Almost Famous is an adjective phrase consisting of a modifying adverb and an adjective as head, and the book title In Cold Blood
Named entities as potentially problematic items
is a prepositional phrase consisting of a preposition with a noun phrase complement. Commenting on the wide variety of grammatical structures that can stand as proper names, Huddleston and Pullum (2002: 517) also give examples of clauses functioning as names (e.g., White Men Can’t Jump, How to Marry a Millionaire). Nevertheless, such titles are proper names, and when used in sentences, they perform the same functions as noun phrases. Because the prototypical proper name is a proper noun, we tend to associate names with the word class noun (for a detailed discussion on the grammatical categorization of names, see, e.g., Anderson 2007, Chapter 6, and Colman 2008). In the domain of NLP, the term ‘named entities’ is often used as a broader umbrella term, including temporal and numerical expressions in addition to conceptually more straightforward ‘proper nouns’ (see, e.g., Ševčiková 2007). The terms ‘proper name’ and ‘named entity’ are used interchangeably in this chapter.
2.2 Annotation of named entities As regards the task of annotating corpus data, the identification and annotation of named entities has been an area of interest from the very beginning, both in terms of manual and automated annotation strategies. There have been different ways of annotating named entities in corpus data, and the strategies tend to reflect both the various interests (e.g., translation and information retrieval, alongside general language study) as well as the practical possibilities of doing so. The guidelines used for the earliest English corpora compiled for linguistic analysis such as the Brown Corpus already observed the problems in the tagging of proper names. The annotation of multi-word proper names was also recommended in the corresponding guidelines for the Penn Treebank Project, in which Santorini (1990: 31) notes that “[i]f a series of words is capitalized as part of a name, the capitalized words should be tagged as proper nouns (NP or NPS)”, with one of the examples being A Tale of Two Cities, in which case all words expect for of was to be tagged as a proper noun. In the USAS category system (see Archer, Wilson & Rayson 2002), multi-word units such as phrasal verbs, proper names, and idioms are annotated with a similar strategy, so that in instances such as Operatic Society both words are tagged as proper nouns. Considerably more in-depth work on the development of automatic algorithms to identify and annotate named entities in digital texts has been done in the realm of NLP and Computational Linguistics, with the aim of increasing accuracy and precision related to tasks such as data mining for named entities and machine translation. Within NLP, Named Entity Recognition (NER) is now seen as one of the major tasks under the broader field of Information Extraction (IE). The challenges in such tasks include, for example, taking into consideration the
39
40
Mark Kaunisto
differences between languages in terms of how the concepts of proper nouns and proper names are understood and how such items manifest themselves on a grammatical level. The use of capital initials is not a fully reliable marker of the status of a proper noun even in English (see, e.g., Biber et al. 2021: 247), and the degrees to which capital initials are indicative of proper noun status can vary from one language to another, posing further challenges to automated processing of parallel language data (see, e.g., Fung 1995; Cucchiarelli et al. 1998; Ševčiková 2007). The design of automatic recognition, translation, and annotation of named entities has developed to such an extent that named entity taggers can perform more fine-grained semantic annotation, allowing the use of tags to distinguish between names for persons, places, organizations, titles, and so on.1 However, the results of this work have not been applied in linguistic corpora. To study annotated corpora efficiently, it is important to be aware of what kinds of things have been annotated and how, as this ultimately affects our understanding of the expected levels of precision and accuracy of our search hits. In addition, it is also true, as stated by Kübler and Zinsmeister (2015: 162) that “[w]hen we use annotated corpora, we need to be aware of the fact that these annotations will not be perfect; they will contain errors.” Kübler and Zinsmeister (2015: 163) further note that particularly with automatically annotated corpora, there may be errors which are more likely systematic in nature, such as the mistagging of sentence-initial common nouns as proper nouns. In general, they recommend that “the user of such automatically annotated corpora look for a discussion of typical errors that the annotation tool in question makes if such information is available” and that “it may make sense to look at a sample of data, to see which error types may occur” (Kübler & Zinsmeister 2015: 163–164; see also Denison 2013). Through a set of some relatively straightforward sample queries relating to items that make up or are part of named entities, the present chapter aims to emphasize and reiterate this very recommendation. The series of case studies is based on queries from corpora such as the British National Corpus (BNC; Hoffmann & Evert 1996–) and other corpora hosted at English-corpora .org. These corpora have been annotated with the CLAWS tagger, which identifies proper nouns and assigns them with tags separating them from common nouns, but it does not assign tags for multi-word proper names.
1. To give an idea of the multitude of different types of named entities identified in the field of NER, as many as 150 types of named entities were defined in Sekine, Sudo and Nobata (2002).
Named entities as potentially problematic items
3.
Case studies
This section looks into a number of lexical items whose analyses based on corpus searches are complicated by a high number of named entity instances found among the search results. What is the extent of potential noise, and why is the only option often to manually inspect the concordance lines in order to exclude irrelevant items from the analysis? Some of the items examined here are selected based on observations of individual problematic instances made over the years, while some near-synonymous word pairs chosen for closer study have one member of the pair fairly frequently occurring in different types of named entities, and the aim is to see how much such cases influence the studies of the near-synonyms. In the 100-million-word British National Corpus, (for the purpose of the searches for this chapter, I have used the BNCweb interface),2 proper nouns are assigned the grammatical tag of _NP0, and common nouns are tagged either with _NN0 (neutral for number), _NN1 (singular) or _NN2 (plural). As regards multi-word proper names, the aim has been to apply the tag _NP0 to all individual parts of the name, so that Lake Tanganyika has both Lake and Tanganyika tagged with _NP0. However, as noted in the manual to the second version of the BNC (Leech & Smith 2000: n.p.), it is acknowledged that “in practice the openendedness of naming expressions makes it difficult to capture all possible types consistently”, and that “[w]e have confined its coverage mainly to personal and geographical names”. The manual is thus very clear in acknowledging the fact that there are many multi-word named entities where many parts of the name may not be tagged with NP0. In some instances, all parts of a name may be tagged as “ordinary words”, such as World Health Organisation, with every word tagged as a singular common noun.
3.1 Common nouns used as (parts of ) proper nouns: Lifespan and samurai If one performs a frequency list search of the singular nouns in the BNCweb with the _NN1 tag and begins to browse through the most frequently occurring nouns on the list, paying attention to the structural make-up of the words, it is probably unsurprising to see that the most common singular nouns are monomorphemic ones. Nouns such as time, way, year, day, and man appear at the top of the list. Not much further down the list, we also find nouns formed with a root and an affix (e.g., government, business, and development). The most frequent singular noun consisting of two roots on the list is chairman, and other nouns of this type include football, newspaper, background, and bedroom. Among the highest2. The CQP-edition, available at 〈https://bncweb.lancs.co.uk〉
41
42
Mark Kaunisto
ranking nouns of this type is lifespan, followed by words such as database, policeman and classroom. Considering the most common two-root singular nouns, we might be surprised to find lifespan featuring as one of the most frequent ones. With a total frequency of 3,700, and a normalized frequency of 37.6 instances per million words, it can be regarded as a high-frequency word in the corpus. However, a closer look into the instances of lifespan informs us that all the 3,700 hits are found in only 83 texts, which suggests the word probably has a very uneven dispersion across the corpus.3 Indeed, a large majority of the hits of lifespan – as many as 3,562 instances (or 96 per cent) – are found in one single document in the corpus, a 200,000-word sample from a Lifespan computer manual. As in Example (2), the instances of the word refer to the named entity, typically spelled with capital letters throughout: (2) Customers who have purchased a LIFESPAN maintenance agreement are entitled to call upon the services of the LIFESPAN help desk. (BNC, HWF 15266)
The case of lifespan is an example of the problem of accurate annotation in cases where common nouns are used as proper nouns, as has been observed in previous studies by, for example, Preiss and Stevenson (2013). In the BNC, none of the instances of lifespan as a part of a name has been tagged as a proper noun (NP0) – in accordance with the tagging system. As noted by Leech and Smith (2000: n.p.) in the BNC2 manual, non-personal and non-geographical names consisting of “ordinary words (common nouns, adjectives etc.) […] receive ordinary tags (NN1, AJ0 etc.)”. When performing frequency list searches in BNCweb, case-sensitive searches are not an option, and even if we search for the word separately, the hits include some common nouns as well as proper names with a capital initial or with all capitalized spelling. In practice, the manual inspection of the concordance lines immediately suggests that lifespan clearly is not, in fact, one of the most frequent two-root singular nouns in British English. Another interesting common word that is often found in proper names is the loan word samurai. In her MA dissertation on the occurrences of Japanese borrowings in six regional varieties in the Corpus of Global Web-Based English (GloWbE; Davies 2013), Lehtonen (2021: 46) noted the high frequency of instances of samurai – as many as 40 per cent of the tokens – appearing in names (as, for example, in the film title The Seven Samurai, or Samurai Pizza Cats, the name of an animated television series), a feature especially prominent with that word compared to others included in her study. Studying the occurrence of 3. The problem of uneven dispersion has been discussed by, for instance, Leech et al. (2001); as regards statistical adjustments that account for the problem, see, for example, Gries (2008).
Named entities as potentially problematic items
common nouns in named entities in corpora such as GloWbE may be particularly interesting, as different types of names are perhaps more likely to manifest themselves not only in sections of corpora representing different genres but also different regional varieties. This has been pointed out, for example, by Stefanowitsch (2020: 365), who mentions that geographical names are more likely to be more frequent in corpora representing their corresponding language regions (i.e., references to London are likely to be more numerous in corpora of British than American English), further adding that “proper names may differ in frequency for purely cultural or for linguistic reasons; the same is true of common nouns”. Another possibility is that regional varieties differ as regards the proportions with which a common word appears in named entities; in order to assess to what degree this is the case for the present study, the instances of the word samurai were examined in five regional varieties represented in the GloWbE corpus (namely those of Great Britain, the United States, Singapore, India, and Malaysia, to include both so-called inner and outer core varieties of English). All instances of samurai in GloWbE are tagged as singular common nouns,4 except for a small number of cases where the word spelled with a capitalized initial follows a title such as Mr or Madam. Sometimes the word is spelled with an initial capital when it is used as a common noun, and the use of initial capitals is irregular, as seen in Examples (3a)–(b), requiring again a closer manual inspection of the concordance lines of the search hits to separate cases where the word appears as a part of a named entity from those where it does not. (3) a.
Inside he said he saw Mr Fisher holding a samurai sword. (GloWbE, Great Britain, General, www.thisisstaffordshire.co.uk) b. Believing he was carrying a Samurai sword, the officer called for backup and the police helicopter was scrambled […] (GloWbE, Great Britain, General, www.southwalesargus.co.uk)
Table 1 below presents frequencies of the word samurai in the five sections of GloWbE, with information given on the total numbers of search hits of the word tagged as a singular noun (_nn1), and a further breakdown of the hits into cases where the word functioned as a common noun and cases where the word is found as a part of a named entity. This serves to exemplify the potential effect of named entities in studies examining the use of a word in general; if such items are not appropriately separated from regular common noun uses, frequency data might be skewed.
4. In BNCweb, the instances of samurai are tagged as _NN0, that is, common nouns which are neutral for number. Some of the GloWbE taggings of the word as singular nouns are erroneous, as in some cases the noun is plural (as in the example The Seven Samurai).
43
44
Mark Kaunisto
Table 1. Frequencies of the word samurai in five sections in the GloWbE corpus (N = absolute frequencies of the tokens; pmw = frequencies normalized per million words) Great Britain
United States
India
Singapore
Malaysia
N
pmw
N
pmw
N
pmw
N
pmw
N
pmw
All tokens tagged _nn1
663
1.73
736
1.90
99
1.03
208
4.84
185
4.44
Common nouns
428
1.10
446
1.15
63
0.65
120
2.79
69
1.66
(Part of ) Named entities
235
0.63
290
0.75
36
0.37
88
2.05
116
2.79
percentage of named entities in all _nn1 tokens
35.4%
39.4%
36.4%
42.3%
62.7%
As we can see in Table 1, based on raw figures of tokens of samurai and their corresponding normalized frequencies in the Great Britain, United States, India, Singapore, and Malaysia sections of GloWbE, the singular form of the noun samurai appears most frequently in the Singapore and Malaysian datasets, with the normalized frequencies being more than twice those in the Great Britain and the United States sections, and the Indian section having the lowest frequencies per million words. After manually separating the instances where the word is found in a proper name from common noun uses, we find that the two types of uses do not break down in the five sections in similar proportions. As for the common noun uses, Singapore and Malaysia have the highest normalized frequencies (2.79 and 1.66, respectively), but it is worth noting that in the Malaysian section, the normalized frequency of the common nouns is proportionally lower when comparing the numbers of all tokens. In other words, the Malaysian section of the corpus has a higher proportion of samurai appearing as a part of a named entity. This is most clearly seen in the bottom row of Table 1, indicating the percentage of hits found in proper names out of all searched hits – out of all search hits of samurai in the Malaysian section, 62.7% are cases where the word is a part of a named entity (e.g., Double Burger Samurai, the name of a McDonald’s hamburger in some Asian countries, or Samurai Shodown, a video game series). The corresponding percentages in the search hits in the other sections are closer to each other, ranging between 35.4% and 42.3%. As regards the uses of samurai in named entities, the Malaysian component thus clearly stands out from the other examined varieties. The differences seen in the rates of samurai in named entities potentially also reflect some cultural differences and the interests of the writers represented in the corpus: the named entities including samurai in the GB and US sections included more references to drama films (The Seven Samurai, The Last Samurai), the Asian sections featured references to food-related items, computer software (Market Samurai, a keyword research tool), action toy figures (Samurai Predator
Named entities as potentially problematic items
AC-01). The popularity of the word in the names of video games or board games (The Way of the Samurai 4, Samurai Warriors 3, Samurai Spirit) was seen in all sections examined.5
3.2 Near-synonymous adjectives in named entities: Limited/restricted, royal/regal and fantastic/fabulous In addition to difficulties caused by common nouns used as parts of names, words representing other parts of speech may also be problematic when they are included in a name. For instance, in popular culture, the names of entities such as films, songs, books, etc. may feature a high number of adjectives, and some adjectives may appear more commonly in them than others. The tagging of some of the most well-known named entities in the BNCweb and the Corpus of Contemporary American English (COCA; Davies 2008–; which both use the CLAWS tagger), such as the White House (the official residence of the president of the United States), seems to take account of the use of capital initials, and mostly White in that name is tagged as a proper noun (_np or _NP0). Although for the most part, white in references to the White House is tagged as a proper noun, there are also some instances in BNCweb and COCA where it is tagged as an adjective; particularly in COCA, there are altogether 1,497 hits (searched August 27th, 2022) with this tagging, showing various combinations of capital or lowercase spelling of initial letters of white and house (mostly referring to the building in Washington D.C.). Interestingly, the House part in the name is usually tagged as a common noun. Lists of high-frequency names are often used when designing automatic taggers, which might account for the fact that with names of lower frequencies, the tagging of white in names appears less systematic.6 5. Considering the fashionableness of the concept of samurai in popular culture today, it is quite possible that the occurrences of the word in named entities would show changing trends in diachronic corpora, and the proportions of hits of the word in names as opposed to common noun uses vary from one period to another. 6. As noted previously, in the BNC2 manual, Leech and Smith (2000: n.p.) note that the NP0 tagging is usually assigned to the different parts of names for geographical places, whereas nonNP0 (or “ordinary”, i.e., AJ0 or NN1) tags is used with multi-word names of places not usually on maps, or with “non-personal or non-geographical” names. However, it is also stated that the tagging shows “a little arbitrariness in application”. In BNCweb, white is accordingly tagged as a proper noun in names such as White Bird Canyon, White Hart Lane, White Mountain, the White Sea, whereas adjectival tagging is found in names such as White City (district of London), the White Tower (the Tower of London), and the White Lion (name of a pub). There are also cases where both types of tagging are found with the same name, for example, the White Swan (name of a pub, with one instance of White tagged as an adjective, and four instances of White tagged as a proper noun).
45
46
Mark Kaunisto
One kind of implication of difficulties in the tagging of named entities can be seen in the corpus-based analysis of near-synonyms, and here we will take a look at some examples of adjective pairs to illustrate the types of problems that can be posed by the occurrences of adjectives in proper names. On a purely intuitive basis, we might predict that the comparison between the adjectives limited and restricted could be affected by the fact that limited appears in names of companies (e.g., Tennis Interlink Limited, County NatWest Limited, and Wirral Peninsular Care Limited). Again, when trying to assess the active use of the adjectives outside of references to named entities, names should be separated from regular uses of the words which reflect the active word choices of the writer or speaker. In the BNCweb, there are no instances of limited which are tagged as a proper noun, regardless of the use of a capital or lowercase initial. There is variation in how limited and restricted are tagged, but it concerns the identification of the words as either adjectives or participial forms of their underlying verbs. To keep the analyses simple, the words were searched in BNCweb with the tag query _AJ*. The wildcard * allows one to target instances where the words were either assigned an adjectival tag or the portmanteau tag (sometimes also referred to as an ambiguity tag) _AJ0-VVN. The order of the two tags in a portmanteau tag is significant: in this case, the automatic tagger has found insufficient evidence to determine the accurate part of speech, but that it was more likely an adjective – the first element of the tag – than a past participial form of a lexical verb – identified by the tag _VVN. In order to assess the frequencies of the adjectives limited and restricted, and then to manually inspect the frequencies with which they occur in names, searches were made in three written sections of the BNC – Academic prose, Non-academic prose and biography, and Newspapers, which also allows for comparisons of the frequencies between written domains. The frequencies of the adjectival uses of the words, starting with total numbers of hits and then providing separate figures for nonnamed entity and named entity uses, are presented in Table 2. As can be seen in Table 2, as an adjective limited outnumbers restricted in all text types, with the highest normalized frequencies for both words found in academic texts, and the lowest in newspaper texts. It is worth noting that this overall observation is largely unchanged even after excluding instances occurring in names.7 It is, however, interesting to observe how the proportions of occurrences in names among the search hits vary between the text types: while the search hits of limited in academic and non-academic prose texts have relatively low proportions of items appearing in names, a higher proportion is seen in newspaper texts.
7. The adjective restricted was rarely used in named entities; for example, the few instances of the word in newspaper texts related to the name of a horse racing category.
Named entities as potentially problematic items
Table 2. Frequencies of adjectival uses of limited and restricted in three text types in the British National Corpus (pmw = per million words) Limited_AJ*
Restricted_AJ*
1. total number of hits
2,012 (127.5 pmw)
377 (23.9 pmw)
2. frequency of regular adjectival uses
1,959 (124.2 pmw)
377 (23.9 pmw)
3. frequency in named entities
53 (3.4 pmw)
0
named entity percentage:
2.60%
0%
1. total number of hits
1,906 (78.8 pmw)
277 (11.5 pmw)
2. frequency of regular adjectival uses
1,824 (75.4 pmw)
275 (11.4 pmw)
3. frequency in named entities
82 (3.4 pmw)
2 (0.1 pmw)
named entity percentage:
4.30%
0.70%
1. total number of hits
327 (34.7 pmw)
56 (6.0 pmw)
2. frequency of regular adjectival uses
280 (29.7 pmw)
48 (5.1 pmw)
3. frequency in named entities
47 (5.0 pmw)
8 (0.9 pmw)
named entity percentage:
14.40%
14.30%
Derived text type: Academic prose
Derived text type: Non-academic prose and biography
Derived text type: Newspapers
In other words, the potential of finding unwanted or irrelevant items may in some cases vary between different sections of the corpus. Similar problems are encountered if we examine the adjective pair royal and regal. In the BNCweb, there are 56 instances where royal with a capital initial is tagged as a proper noun (NP0), as in Royal Exchange Square (as a part of an address), Palais Royal, Park Royal, and Musée Royal des Beaux-Arts, and in these cases all elements of the names are applied with the NP0 tag. However, out of the total number of hits of royal in the corpus – 14,628 hits in 1,825 texts – it appears with a capital initial as often as 10,237 times in 1,603 texts). In the vast majority of instances with an capital initial royal is tagged as an adjective, as in the Royal Academy, the Royal Shakespeare Company, the Royal Exchange Theatre, the Royal Commission on Environmental Pollution, the Royal Mile, the Royal Albert Hall, the Royal Navy, and so on.8 The adjective regal is overall considerably less frequent 8. Considering the names in which royal with a capital initial is tagged as an adjective instead of a proper noun, the types of entities in question is reflected in the tagging, as noted in the BNC2 manual; in the names of institutions or charters, royal is an adjective, in the names of locations it is treated as a proper name.
47
48
Mark Kaunisto
than royal in the BNCweb, with 304 instances found in 152 texts, all of which are tagged as adjectives, although in a number of cases the word appears in names, as in the Regal Scottish Masters (a snooker tournament), the Regal (a cinema theatre), the Regal Arms (a hotel), and Regal Trophy (a rugby league). There are altogether 194 hits of regal spelled with a capital initial, although not all instances are proper names. In other words, in the case of both royal and regal, substantial proportions of the occurrences of the words in the corpus appear in proper names, and it would be crucial to exclude such instances if we wanted to study of the use of the words as non-frozen elements. In lexical studies, near-synonyms are investigated for various purposes and interests: to reveal fine-grained distinctions and principles behind word choices, to examine variation between dialects and sociolects, or to provide assistance to translators and language learners, to name but a few. When we examine the characteristics of near-synonyms with corpus data and attempt to analyze the differences in the uses of the words, we typically examine their collocational behaviour. Many corpus interfaces and tools allow for the analysis of the strongest collocates of words based on different methods of assessment: Mutual information, log-likelihood, T-scores, raw frequency ranking, and so on. In such analyses, it might be beneficial to be able to exclude those tokens which appear as a part of a name. If the taggings of the corpus data are not helpful in this regard, one can try to perform case-sensitive searches; in the BNCweb, the searches could be restricted to all-lowercase spellings of royal and regal, for example. However, this would compromise those cases where the words appear at the beginning of a sentence. Conversely, particularly in British English, the adjective royal is sometimes spelled with a capital initial, perhaps out of respect to the monarchy, even in cases where it is not a part of a name, as in “The old Royal adage of never complain, never explain just won’t do any longer” (BNC, CH6 6584) and “Princess Diana got the star billing at a Royal banquet on board the Britannia last night” (CBF 9645). In addition, in some collocations there are variant types of spelling of royal: for example, we find instances of royal family, Royal family, and Royal Family, and the form of spelling does not necessarily make it clear whether the authors are using the phrase as a name. All such considerations potentially add to the problems of analysing of collocational strength, which in the case of royal and regal may be more complex with British English data. Some corpus interfaces such as the one at English-Corpora.org allow users to compare the collocates of two words, and the collocates which are more typically used with the compared words are listed in two separate columns in a ranked order, based on a score that takes into account the frequencies and comparative ratios of the two words in the corpus, the frequencies and ratios of different combinations of the words and their collocates, as well as the frequencies of
Named entities as potentially problematic items
the individual collocating words across the corpus. The listings are very useful in highlighting the differences between near-synonyms which are reflected in their collocational behaviour. However, we can note that analyses of collocations can also be affected by occurrences of the compared words in names; for example, we can consider the adjectives fantastic and fabulous. Adjectives sometimes differ as regards their adverbial modification, and adverbs have also been seen to differ as regards their degrees of use in different regional varieties of English (see, e.g., Biber et al. 2021: 519, 560–563). For the purpose of illustrating a case where the influence of references to a named entity can skew the analysis of adverbial modifier comparison between fantastic and fabulous, we can first perform a “Compare” search in the Great Britain section in the News on the Web Corpus (NOW; Davies 2016–) by entering the two adjectives in the Word1 and Word2 slots, ADV in the “Collocates” slot, and setting the window span to one word to the left (see Figure 1). Limiting the search further to the Great Britain section, this search provides us with lists of the adverb modifier collocates of fantastic and fabulous.
Figure 1. A screenshot of a search setup comparing the adverb collocates of fantastic and fabulous
We find adverbs such as bloody, actually, frankly, and genuinely are listed as collocating more typically with fantastic (and on the results page, these words are highlighted in green colour), whereas adverbs listed as strong collocates of fabulous include rather, totally, very, and absolutely. Of the adverbs associated more closely with fabulous, the listing of the adverb absolutely might make some corpus users wonder whether the high ranking is due to the popularity of the television comedy show Absolutely Fabulous. A look into the concordance lines shows that references to the BBC show are very numerous indeed; in fact, out of the 736
49
50
Mark Kaunisto
instances of the phrase absolutely fabulous,9 as many as 545 (i.e., 74%) are found in names, mostly indeed referencing the show or the film Absolutely Fabulous: the Movie. The modifying adverb absolutely is likewise listed as favouring fabulous compared to fantastic in the Ireland section of the NOW corpus, and in this Section 70% of the 549 instances were found in names (also including references to the flower delivery service Absolutely Fabulous Flowers). In the United States section, absolutely is not among the adverbs highlighted as ones associated more closely with fabulous, but the scores measuring the collocation ratios are tilted towards fabulous (0.5 for absolutely fantastic, and 1.8 for absolutely fabulous). In the US data, a notably lower proportion of absolutely fabulous are references to the show (133 out of 374 hits, i.e., 35.6%). Although the show appears to be popular also in the United States based on the NOW corpus data, its impact on the collocation scores is not as strong as it is in the Great Britain and Ireland sections. None of the instances of absolutely fantastic in the US, GB, and IE sections were names or parts of names. So, when we examine the adverb modifiers of adjectives, we know that the collocation scores probably exaggerate the association between the adverb absolutely and the adjective fabulous, as the scores are skewed by the frequent occurrences of the combination of those two words in names. If we tried to adjust and recalculate the scores of collocational ratios of absolutely fantastic and absolutely fabulous by removing those instances of absolutely fabulous found in names, we would also need to subtract the corresponding instances from the numbers of instances of fabulous and absolutely fabulous. The recalculated scores for collocational strength for absolutely fantastic and absolutely fabulous would then be reversed in the Great Britain and Ireland data (in the GB section, the scores for absolutely fantastic and for absolutely fabulous are 1.7 and 0.6, and in the IE section, the corresponding scores are 1.2 and 0.8), indicating a tendency for absolutely to be used more frequently as a modifier to fantastic than fabulous, although the difference between the scores would not be high enough for absolutely to be identified as collocating strongly with fantastic in comparison with fabulous (i.e., absolutely would not be highlighted in the results list with green colour). But such recalculations, of course, are not entirely accurate, as we would be assuming that the scores for all the other adverbs modifying the two adjectives would be unaffected. Subtracting the instances of absolutely fabulous might not affect the rankings of other collocations to a notable degree, but for things to be ideal, we would need information on the named entity uses of all combinations, and for that purpose manual inspection is practically impossible. Nevertheless, based on these observations, it would be highly advisable to 9. Searches on the adjectives fantastic and fabulous in the NOW corpus were made on Aug. 3, 2022.
Named entities as potentially problematic items
inspect concordance lines of the highest-ranking collocations to check for possible skewing effects by names.
4. Discussion and conclusion Considering the occurrences of names in the corpora examined, it can be argued that it would be beneficial to have the option of building a search query that would enable one to include or exclude items if they constitute a part of a name. Ultimately names also contribute to the word count of a corpus, and it might also be beneficial to exclude named entities from automated calculations in cases where names – or other types of items in a corpus which may be irrelevant to what one is trying to assess by the calculations – do not reflect the active linguistic choices of the writers or speakers represented in the corpus. Of course, the uses of words in different types of names is an interesting question in its own right, and the words found in names definitely carry meanings beyond the moment of giving an entity a name. As mentioned, the question of exclusion arises from the idea that a person referring to the entity by its conventionalized name is bound by this convention and is not at liberty to use other words than the ones in the name, expect for any understood or established variants of the name. By drawing attention to the question of words appearing in names, the present chapter has made the point that it is not always necessarily clear from the outset how much and what kind of post-processing of the search results is needed. As all users of corpora are not necessarily aware of the variety of things to watch out for, a set of examples was presented of cases where occasionally large proportions of search hits can turn out to be – depending on the point of view of the study – instances that one might want to exclude from the study, which may not always be a straightforward matter. As has been observed in instances such as the noun lifespan in the British National Corpus, and the adjectival phrase absolutely fabulous in British and Irish sections of the News on the Web corpus, sometimes the irrelevant tokens may far outnumber the relevant ones. The concerns and challenges addressed in the present chapter are by no means new ones, and many existing corpora are frequently updated and improved to make the search interfaces more user-friendly as well as to increase the reliability of the findings. As a lot of work has been done on named entity recognition in the field of computational linguistics, a key to address the problems raised in commonly used, large corpora in corpus linguistics would be to increase the collaboration between the areas, and to integrate the systems and practices developed in NLP and computational linguistics into corpora such as the BNC. An encouraging example of such collaboration is the Clean Corpus of Historical American English
51
52
Mark Kaunisto
(Alatrash et al. 2020), involving the clean-up of data in the diachronic Corpus of Historical American English to enable a broader set of possibilities of examining language change. Thinking back to the problems relating to the use of corpora as outlined by Rissanen (1989), it is perhaps useful to keep reminding corpus users of the value in scrutinizing one’s initial findings from the corpus data, checking the dispersion of the search hits, and familiarizing themselves with the annotation systems applied. As much as can be done automatically, we do not always have ways to automatically spot the outliers. The matter is important from the point of view of educating future users of corpora, as errors resulting from insufficient understanding of the corpora are frequently seen in student papers and even in published scholarly works. This sentiment was also voiced by Denison (2013: 33), who commented on the problems caused by errors in grammatical mark-up in corpora to different user groups by saying that “misclassified examples will mislead students” and that “[e]xperienced researchers can find misclassified examples if they already have suspicions, but if not, relevant examples may be missed”. Furthermore, it is possible that the multitudes of functions of corpus user interfaces that are available may influence the users’ perceptions of what kind of post-processing considerations are necessary, with the danger that the more sophisticated and impressive the system appears at the outset, the less cautious one feels one needs to be when one analyses the corpus findings. This false sense of security, related to the “God’s truth fallacy” pointed out by Rissanen, is perhaps a new kind of corpus-related fallacy.
References Aarts, Bas. 2011. Oxford Modern English Grammar. Oxford: OUP. Aarts, Bas, Chalker, Sylvia & Weiner, Edmund. 2014. The Oxford Dictionary of English Grammar, 2nd edn. Oxford: OUP. Alatrash, Reem, Schlechtweg, Dominik, Kuhn, Jonas & Schulte im Walde, Sabine. 2020. CCOHA: Clean Corpus of Historical American English. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk & Stelios Piperidis (eds), 6958–6966. Marseille: European Language Resources Association. Anderson, John M. 2004. On the grammatical status of names. Language 80(3): 435–474. Anderson, John M. 2007. The Grammar of Names. Oxford: OUP. Archer, Dawn, Wilson, Andrew & Rayson, Paul. 2002. Introduction to the USAS category system. (19 May 2024).
Named entities as potentially problematic items
Biber, Douglas, Johansson, Stig, Leech, Geoffrey N., Conrad, Susan & Finegan, Edward. 2021. Grammar of Spoken and Written English. Amsterdam: John Benjamins. BNC = Hoffmann, Sebastian & Evert, Stefan. 1996–. BNCweb (CQP-edition). Zürich: Zûrich University. COCA = Davies, Mark. 2008–. The Corpus of Contemporary American English. (27 August 2022). Colman, Fran. 2008. Names, derivational morphology, and Old English gender. Studia Anglica Posnaniensia 44: 29–52. Cucchiarelli, Alessandro, Luzi, Danilo & Velardi, Paola. 1998. Automatic semantic tagging of unknown proper names. In COLING 1998, Vol. 1: The 17th International Conference on Computational Linguistics, 286–292. Montreal: Université de Montréal. Denison, David. 2013. Grammatical mark-up: Some more demarcation disputes. In New Methods in Historical Corpora, Paul Bennett, Martin Durrell, Silke Scheible & Richard J. Whitt (eds), 17–35. Tübingen: Gunter Narr. Fung, Pascale. 1995. A pattern matching method forfinding noun and proper noun translations from Noisy Parallel Corpora. In 33 Annual Meeting of the Association for Computational Linguistics (ACL ’95), 236–243. Cambridge MA: Association for Computational Linguistics. GloWbE = Davies, Mark. 2013. Corpus of Global Web-Based English. (19 May 2024). Gries, Stefan T. 2008. Dispersions and adjusted frequencies in corpora. International Journal of Corpus Linguistics, 13(4): 403–437. Huddleston, Rodney & Pullum, Geoffrey. 2002. The Cambridge Grammar of the English Language. Cambridge: CUP. Kaunisto, Mark. 2017. Multilingualism and quotations from a corpus-linguistic perspective: A case study of Samuel Taylor Coleridge’s Biographia Literaria, in Challenging the Myth of Monolingual Corpora [Language and Computers 80], Arja Nurmi, Tanja Rütten & Päivi Pahta (eds), 220–238. Leiden: Brill. Kübler, Sandra & Zinsmeister, Heike. 2015. Corpus Linguistics and Linguistically Annotated Corpora. London: Bloomsbury. Leech, Geoffrey & Smith, Nicholas. 2000. Manual to Accompany the British National Corpus (Version 2) with Improved Word-class Tagging. UCREL. Lancaster University. (19 May 2024). Leech, Geoffrey, Rayson, Paul & Wilson, Andrew. 2001. Word Frequencies in Written and Spoken English: Based on the British National Corpus. London: Longman. Lehtonen, Sharin. 2021. Tsunami, anime, and martial arts: A corpus-based lexicological study of Japanese borrowings in a historical context and in six varieties of Present-day English. MA dissertation, Tampere University. 〉 (19 May 2024). NOW = Davies, Mark. 2016–. Corpus of News on the Web. (3 August 2022). Pierini, Patrizia. 2008. Opening a Pandora’s pox: Proper names in English phraseology. Linguistik Online, 36(4).
53
54
Mark Kaunisto
Preiss, Judita & Stevenson, Mark. 2013. Distinguishing common and proper nouns. In Second Joint Conference on Lexical and Computational Semantics (*SEM), Vol. 1: Proceedings of the Main Conference and the Shared Task, 80–84. Stroudsburg PA: Association for Computational Linguistics. Rissanen, Matti. 1989. Three problems connected with the use of diachronic corpora. ICAME Journal 13: 16–19. Rissanen, Matti. 1992. The diachronic corpus as a window to the history of English. In Directions in Corpus Linguistics. Proceedings of Nobel Symposium 82, Stockholm, 4–8 August 1991, Jan Svartvik (ed.), 185–205. Berlin: Mouton de Gruyter. Santorini, Beatrice. 1990. Part-of-Speech Tagging Guidelines for The Penn Treebank Project, 3rd revision. Philadelphia PA: University of Philadelphia. Sekine, Satoshi, Sudo, Kiyoshi & Nobata, Chikashi. 2002. Extended named entity hierarchy. In Proceedings of the 3rd International Conference on Language Resources and Evaluation (LREC 2002), Las Palmas, Canary Islands, Spain, 1818–1824. Paris: European Language Resources Association (ELRA). Ševčiková, Magda. 2007. Proper nouns in Czech corpora. In Proceedings of the Corpus Linguistics Conference (CL2007), Matthew Davis, Paul Rayson, Susan Hunston & Pernilla Danielsson (eds). University of Birmingham. 〉 (19 May 2024). Simon, Eszter. 2019. The Definition of Named Entities. In K + K = 120. Papers Dedicated to László Kálmán and András Kornai on the Occasion of their 60th Birthdays, 481–497. Budapest: Research Institute for Linguistics, Hungarian Academy of Sciences (RIL HAS). Stefanowitsch, Anatol. 2020. Corpus Linguistics: A Guide to the Methodology [Textbooks in Language Sciences 7]. Berlin: Language Science Press.
Challenges in the compilation, annotation, and analysis of learner corpus data Marcus Callies
University of Bremen
This chapter highlights and discusses the special characteristics of learner corpus data and the challenges they may present for corpus compilation, annotation, and analysis. Because learner corpus and SLA researchers use their data to study L2 production and development, it is of utmost importance that the data are valid, that is, they represent “authentic” L2 production, which means that the data must stem from the studied learners’ own language production. I discuss challenges in three areas: (1) multilingual practices and metalinguistic language use, (2) lexical and constructional bias, often brought about by the wording of task instructions or writing prompts that learners are asked to respond to, and (3) learner corpus annotation in view of the “discourse of deficit” in SLA. For each of these challenges solutions as to how they can be met are offered. Keywords: learner corpus, compilation, analysis, annotation, multilingualism, lexical bias, task instruction, writing prompt, discourse of deficit, lexical innovation
1.
Introduction and general remarks
Learner Corpus Research (LCR) is a relative newcomer to the scene of research paradigms and methodologies within applied linguistics and second language acquisition (SLA) research. LCR as a field only emerged and became visible in the course of the 1990s in the context of the popularization of corpus linguistics at large but has rapidly evolved and grown in scope and sophistication over the past four decades. Two handbooks that survey the field and discuss its links to the major neighbouring disciplines of corpus linguistics, SLA and language teaching bear witness to this rapid development (Granger, Gilquin & Meunier (eds) 2015; Tracy-Ventura & Paquot (eds) 2021).
https://doi.org/10.1075/scl.118.04cal © 2024 John Benjamins Publishing Company
56
Marcus Callies
When compared to other types of data that have traditionally been used in SLA research, learner corpora aim to provide large-scale principled collections of authentic, continuous and contextualized language use by foreign/second language (L2) learners that are stored in electronic format. They enable the systematic and (semi-)automatic extraction, visualization and analysis of large amounts of learner data in a way that was not possible before. Researchers in LCR and SLA use learner corpus data to study L2 production and development. The occurrence and contextual use of particular structures in learner data is thus taken as evidence of (the process of ) acquisition and productive use. From this follows that the data must be valid, that is, represent language actually produced by the learners under study. If this is not the case, the validity of findings derived from learner corpus data may be threatened. In this chapter I draw attention to and discuss several special characteristics of learner corpus data and the challenges these may present for corpus compilation, annotation and analysis. Examples from three areas will be used to illustrate these challenges: (1) multilingual practices and metalinguistic language use, (2) lexical and constructional bias, often brought about by the wording of task instructions or writing prompts that learners are asked to respond to, and (3) learner corpus annotation in view of the “discourse of deficit” in SLA. For each of these I will offer solutions as to how the challenges can be met.
2.
Challenges and how to respond to them
2.1 Multilingual practices and metalinguistic language use Texts contained in learner corpora have typically been produced by bi- or even multilingual individuals, and learner data are rich in multilingual practices and phenomena induced by language contact, such as code-switching, foreignizing (the morphophonological modification of an L1 form to adapt to the structure of the L2) or calquing (literal translations of expressions from the L1 into the L2). Importantly, they are indicative of interlanguage development. In SLA, the term and concept of ‘interlanguage’ refers to a systematic and independent developing learner grammar that should be studied in its own right and contains elements of the learner’s L1 and L2 but also independent structures (Selinker 1972). Multilingual practices also evidence communicative strategies used by the learners to solve communicative challenges (most often problems in lexical search) but at the same time present challenges for annotation and analysis alike (Callies & Wiemeyer 2017: 90–92). Identifying and interpreting such instances is timeconsuming and usually requires researchers to have a firm command of the learners’ L1 so that they are able to identify and interpret effects induced by crosslinguistic influence.
Challenges in learner corpus data
Moreover, learner corpora, especially those of academic texts, contain expert terminology, metalinguistic language use such as contextualised examples of language use, citations or mentioned items, and sometimes even whole abstracts or thematic summaries from other languages. Such instances usually do not represent the respective learners’ own written production as they are typically taken over or copied from secondary sources. On the one hand, from the point of view of quantitative L2 analysis, these could thus be considered unwanted items or ‘false positives’ as their inclusion in word counts and concordance analyses will affect findings on what structures learners may have acquired and are able to produce. Thus, such cases have to be treated separately so that they can be excluded from search results and word counts to not distort the data in learner corpus studies. On the other hand, they are part of the text, they cannot simply be removed, as this may reduce the context or even render the text illegible. Additionally, their textual embedding and glossing may be of interest for future analyses, for instance in studies of intertextuality and referencing in academic writing, so that it seems desirable to preserve them in their original form and syntactic function.
Response Data of the kind discussed above should be specifically annotated/tagged. From a practical point of view, the annotation of instances of multilingual practices in learner corpora facilitates their automatic search and identification through corpus software such as concordancers. It also allows for their exclusion from analysis and frequency counts if necessary. Despite the pervasiveness and importance of instances of multilingual language use in SLA, learner corpora are not commonly annotated for such features. In some corpora, however, specific types of multilingual language use have explicitly been annotated. One example is the Louvain International Database of Spoken English Interlanguage (LINDSEI; Gilquin et al. 2010), a learner corpus of spoken English interlanguage that consists of transcripts of oral interviews between EFL learners and English L1 or L2 interviewers. The interviews were transcribed orthographically following a set of guidelines1 including a specific tag for the use of so-called “foreign” words. Learners’ use of words from a language other than English is indicated by the tag before and after the respective word or expression as shown in Example (1) taken from the Spanish component of the LINDSEI. (1) I don’t know the the . the name in English but (eh) here we call it Traducción e Interpretación (LINDSEI; SP025)
1. The transcription guidelines are available from 〈https://uclouvain.be/en/research-institutes /ilc/cecl/transcription-guidelines.html〉 (15 March 2024).
57
58
Marcus Callies
A similar practice has been followed in the annotation of the Spanish Learner Language Oral Corpora (SPLLOC), a set of corpora of L2 Spanish that were transcribed using the CHAT system developed by the CHILDES project.2 In the annotation process, a number of codes were added to the CHAT system for the specific purposes of SLA research.3 Among other things, the annotation of what is referred to as “L2 adaptations” comprises the use of L1 English, that is, codeswitching at the word or phrase level and in complete utterances. The former is marked by the code @s: followed by a code indicating the part-of-speech as shown in (2a) and (2b), the latter are given in square brackets starting with the code ^eng:, as shown in (3). (2) a. b. (3)
*P63: y cómo se dice scuba@s:d diving@s:v ? *P51: no están en el sol están en shade@s:d.
*P04: [^ eng: I don’t know what that means].
Additionally, learners’ use of indeterminate forms and idiosyncratic neologisms (referred to as “invented words”, see SPLLOC Transcription Guidelines 2008: 18), apparently mostly cases of foreignizing or calquing, were marked with the code @n at the end of the word as shown in (4). (4) *P54: um ehm detrás de lo eh pictura@n eh hay [/] hay un número de turistas.
The transcription guidelines also provide for codes to mark the use of a third language different from Spanish or English (see SPLLOC Transcription Guidelines 2008: 17). From the perspective of the corpus user, the annotation of instances of multilingual strategies such as codeswitching, foreignizing, and calquing may prove useful as it makes it possible to systematically retrieve and study them in a learner corpus. Analyzing such passages can provide new insights into interlanguage development, specifically in terms of the multilingual strategies and communicative competencies of the learners. As for academic texts that contain expert terminology and instances of metalinguistic language use, Callies and Wiemeyer (2017: 90) propose using a tag that subsumes all kinds of metalinguistic elements, most of which are genre- and discipline-specific, for example, contextualised examples, quotations
2. The transcription guidelines for CHAT are available from 〈https://talkbank.org/manuals /CHAT.pdf〉 (15 March 2024). 3. These specialised transcription guidelines are available from 〈http://www.splloc.soton.ac .uk/trancon.html〉 and 〈www.splloc.soton.ac.uk/doc/SPLLOCTranscriptionGuidelines.doc〉 (15 March 2024).
Challenges in learner corpus data
or translations. Thus, mentioned items are all those linguistic items that are referred to on the meta-level (both foreign and native) in the text. This applies to single letters, phonetic symbols, emoticons, syllables, morphemes, words, phrases, clauses, and sentences.4
2.2 Task effects A further challenge that compilers and users of learner corpora have to deal with is unwanted lexical or constructional bias which results in an accidental overrepresentation of certain structures in the data. This bias can be introduced either indirectly by the topic of the task in that, for example, a large batch of essays on the topic of friendship in a corpus will lead to a disproportionally high frequencies of words and phrases that belong to the lexical field of friendship. Lexical or constructional bias can also occur when learners use words, phrases or syntactic constructions from the task description, the writing prompt or other input material. This may include patterns that are copied as unanalysed sequences by the learners to complete the task, a ‘play-it-safe’ strategy without necessarily having acquired the respective structure. O’Donnell et al. (2013) describe how this may affect the use of n-grams in argumentative writing: there are likely effects of text sampling on the recurrence of formulaic pattern, with the prompt questions driving the more common formulaic sequences in ICLE (e.g. the opium of the masses, the birth of a nation, the generation gap, ICLE French) … MICUSP (especially MICUSP-NS) reflects common formulaic sequences from reference sections (e.g. American Journal of Public Health, Hispanic Journal of Behavioral Sciences, levels of psychological well-being). (O’Donnell et al. 2013: 101)
Lexical material in the task instructions or the writing prompt may even trigger the recurrent use of a whole grammatical construction. For example, Callies (2008) notes an effect of the writing prompt on the occurrence of raising constructions with easy as in “English is an easy language to learn/teach” which is a frequently occurring essay topic in the International Corpus of Learner English. Alexopoulou et al. (2015: 110) discuss various task-effects on the use of formulaic sequences (FSs) with relative clauses (RCs) in a learner-corpus study based on the EF Cambridge Open Language Database (EFCAMDAT), an open-access database of student writing submitted to Englishtown, the online school of EF Education First. In particular, they highlight several methodological challenges 4. Such a tag, but in a narrower sense, is also used for the International Corpus of English (ICE). The Markup Manual for Written Texts (Nelson 2002: 10) specifies that it is used to “mark words which are cited as words”.
59
60
Marcus Callies
and issues arising from an analysis of big data resources by means of Natural Language Processing (NLP) procedures. They draw attention to the fact that the simple occurrence of relative clauses in the data “cannot be interpreted as direct evidence for the existence of grammatical knowledge, and productive uses of RCs must be distinguished from FSs involving those structures” and that “in an EFL context, task effects and FSs interact since FSs can be part of routines rehearsed in a lesson that learners are then asked to reproduce in writing” (2015: 109). Thus, RCs in the data can be prefabricated, that is, stored and retrieved as wholes from the input at the time of use rather than being generated subject to previous analysis by learners’ interlanguage grammar. Several task effects are distinguished: – –
– –
Explicit encouragement to include a specific structure (e.g., a gerund) or certain type of structure (e.g., a relative clause) in a specific writing task, leading to a higher frequency of use in that particular task; A task elicits a certain type of structure implicitly. A prompt may elicit a high number of occurrences of a particular structure (e.g., temporal clauses, pronouns) as a natural consequence of language required to meet functional communicative requirements of the task (e.g., past tense narrative) A task is neutral with regard to the elicitation of a specific structure. Copying directly from input prompt.
Finally, as Lozano and Mendikoetxea (2010) show, the overrepresentation of a certain lexical verb in a corpus can lead to the subsequent overrepresentation of a syntactic structure irrespective of the task prompt. They note that postverbal subjects (VS) in English occurred significantly more frequently in their learner corpus (L1 = Spanish, L2 = English) despite the fact that the constraints on full inversion are similar for both L1 and L2 users. The higher occurrence of VS in the L2 data was apparently brought about by the learners’ use of the verb exist: almost half of the learners’ VS structures contained the verb exist, as opposed to the English L1 writers (2010: 492). The corresponding Spanish verb existir would usually have a postverbal subject in Spanish (2010: 492). While this feature does not transfer directly to L2 English (there are more instances of SV with exist than inverted structures), it is the general overrepresentation of the lexical item exist by the Spanish L2 learners of English that caused an overrepresentation of inverted structures in the data.
Response It is important that researchers detect and control for the effects of lexical and/ or constructional bias but identifying it can be challenging. Essay topics, task prompts, and writing instructions should be checked carefully if related and recurrent patterns are noticed in the data. Words identified to cause lexical bias are
Challenges in learner corpus data
either treated as stopwords and thus excluded from corpus queries, and likewise, L2 structures that are likely to have been brought about by lexical or constructional bias (and thus may have been copied as unanalysed chunks by the writer) are excluded from the analysis (as in Callies 2008). Such structures are sometimes spotted accidentally in concordance searches but could also be retrieved automatically be comparing the text used in the task instructions to recurrent patterns in the learner corpus. If such bias is not discovered, its effects may challenge the validity of research findings, because lexical / lexico-grammatical variation, sophistication and complexity are often considered as proxies for L2 proficiency.
2.3 “Discourse of deficit” and learner corpus annotation Finally, I would like to discuss the potential bias of certain annotation methods used in LCR and in other disciplines. LCR has been influenced by the “discourse of deficit” (Ortega 2013) that is still prevalent in much SLA research which is linked to the persistent use of a monolingual native-speaker norm as the benchmark for the assessment of learner data. The underlying pedagogical perspective is one that considers native speakers’ language as the norm against which differences and features of learner language are evaluated and characterized as nativelike, or rather, as non-native-like. This perspective has been criticized as centering around a target-deviation perspective in which interlanguages are merely seen as more or less successful attempts to reproduce an implicit target-language norm. Learners try to do what adult native speakers do, but do so less well, thus exhibiting an imperfect and deficient imitation of the target language. A special kind of annotation in learner corpora is error annotation. This is typically carried out on the basis of an error tagset, which usually derives from error taxonomies based on structural linguistic categories (see, e.g., Lüdeling & Hirschmann 2015). In view of an idealised, monolingual native-speaker benchmark, error annotation often tends to be overly prescriptive. Lüdeling and Hirschmann (2015) note that any error annotation that uses a native ‘standard’ against which an error analysis is performed is problematic. This certainly has to be taken into account and there are many researchers that try to do error analysis in ways that avoid (or minimise) the comparative fallacy, mainly by carefully explaining and motivating any step in the analysis so as to find possible biases. (Lüdeling & Hirschmann 2015: 139)
As mentioned in Section 2.1, instances of multilingual language use such as foreign elements, foreignizing and calquing are often tagged as errors. Some error tagsets do, however, take into account that certain errors are induced by L1 transfer. For
61
62
Marcus Callies
example, the tagset used for an error-tagged sample of the Polish component of the ICLE (PICLE) contains a category for “morphological errors due to L1 transfer” (Kaszubski 2005). Examples of these include *psychics instead of psychology and *antidotum instead of antidote. Three categories allow for the categorization of what is referred to as “false friends” (Kaszubski 2005): one for incorrect dependent prepositions of nouns, one for incorrectly formed prepositional verbs, and one for phrases in which there were errors due to a transfer from the L1 Polish. For instance, one learner used the phraseological “false friend” *get back to health to mean ‘recover’, apparently due to the existence of a similar phraseme in Polish. Creative and innovative but ‘non-standard’ interlanguage forms are often contact-induced or formed on the basis of semantic or structural analogy to frequent and recurrent patterns identified in the L2 input. An often-discussed example is particle verbs like discuss about, enter into, return back (see Callies & Hehner 2021; Callies 2023). In most contexts, such structures may be considered errors from an exonormative point of view, but they actually present valuable and highly interesting data for research into SLA and nativization processes in World Englishes (see, e.g., Callies 2016).
Response As for error annotation, Lüdeling and Hirschmann (2015) outline the advantages for multi-level standoff annotation which makes it possible to define as many independent annotation layers as necessary. Hence, different annotators, who potentially use different tools, can work on the same data and their analyses can be consolidated. Importantly, different interpretations of the same data can be kept apart, which enables a transparent reconstruction of both target hypothesis and error exponent. By contrast, in flat token-tag architecture errors are annotated in some kind of inline format directly in the original text; the annotation is thus not stored separately from the corpus data. Lüdeling and Hirschmann argue that it is useful to make the target hypothesis explicit and to use a flexible corpus architecture that allows multiple target hypotheses and the addition of one or more layer(s) for the fine-grained analysis of a given error type or phenomenon (2015: 155). By contrast, instances of multilingual language use and lexical innovations should not be annotated as errors. By innovations I refer to forms that are unattested or infrequent in the main reference varieties of British and American English, but products of morphological regularity or creativity in that they are formed by either adapting L1 elements to fit L2 forms, or by recombining L2 elements (Callies 2016: 230). One solution to take a less evaluative stance and to separate annotation from (over)interpretation would be to adopt the practice followed in at least some components of the International Corpus of English (ICE). The man-
Challenges in learner corpus data
ual for written texts specifies a special mark-up with normative corrections using insertions and deletions intended to facilitate the parser that proves useful for automatically tracing potential lexical innovations.5 However, it appears that this mark-up (described in the ICE tagging manual in a section on “Normalizing the text”, Nelson 2002: 12) appears to not have been applied consistently across all ICE components.6 A related issue for annotators of learner corpora concerns foreignized words and calques. Annotating instances of these phenomena is challenging because these formations are elements neither from the L1 nor from the L2. To an annotator who is not familiar with either language or who is not an expert speaker of the target language, they may be difficult to identify as instances of foreignizing or calquing. On a formal level, they represent unattested forms in either language. From an interlanguage perspective, however, they are immensely interesting because they provide insights into the developmental stage of the learners, their employment of lexical bridges, and their (creative) application of word-formation rules. Therefore, accounting for these phenomena in an annotation scheme would afford corpus users the opportunity to study them and gain valuable insights into interlanguage development. They should thus be considered by corpus annotators, for example, by means of a tag that accounts for such hybrid interlanguage forms that result from language mixing in bi- or multilingual speakers. However, this would require linguistically highly trained annotators and would already constitute a considerable degree of intervention in the original data in that interpretations are made from the point of view of the corpus compilers and annotators.
3.
Summary and conclusion
In this chapter I have argued that in LCR and SLA the occurrence and contextual use of particular structure(s) in learner data is taken as evidence for the (process of ) acquisition and productive use, hence L2 data must be valid in that they represent language actually produced by the respective learners under study. Since texts compiled in learner corpora have been produced by multilingual individuals, learner data are rich in phenomena induced by multilingualism and language con5. In addition, ICE uses two more tags relevant to the present discussion: (for a “word or sequence of words that is foreign and non-naturalised”) and (marking words which are “non-English but are indigenous to the country in which the corpus is being compiled”). 6. In fact, as ICE components have been compiled over three decades, annotation practices seem to vary from component to component.
63
64
Marcus Callies
tact which have a hybrid nature and are challenging to classify. Learner corpora of academic writing, on the other hand, contain various kinds of metalinguistic language use and language taken over or copied from secondary sources. I have suggested that multilingual practices, metalinguistic language use and instances of intertextual reference should be identified and annotated so that they can be dealt with in or be excluded from corpus analysis. Researchers also need to pay attention to recurrent patterns in the task instructions, writing prompts and other input material (provided this has been documented!) and if these may induce lexical or constructional bias. When compiling a new corpus, the task instructions and the input material should be neutral with regard to the elicitation of specific structure(s). Finally, annotation and analysis should be separated from normative, oftentimes (over)prescriptive interpretation. Conceptual and methodological reform and innovation remain a strong desideratum in a young discipline like LCR. To provide a certain incentive for researchers to document and share their methodological practices and tools, and to give researchers engaged in advancing the methodological expertise of the field the recognition they deserve, the International Journal of Learner Corpus Research decided to introduce new publication formats (see Paquot & Callies 2020). This can be seen as another initiative in line with a “growing trend towards methodological awareness and reflection” in LCR (Paquot & Callies 2020: 122). The journal thus now invites and publishes corpus materials and methods and software reports, and, to answer calls for reflection and assessment of methodological practices, also review articles, position papers and replication studies (which receive growing importance in SLA) to reproduce findings of previous LCR studies. Issues of standardisation and best practices have frequently been addressed (see, for instance, Gilquin 2015), but when it comes to corpus compilation and annotation, such guidelines for best practice do not seem to be available yet when compared to corpus linguistics in general (see, for instance, Sinclair 2005). Such standardisation is perhaps a lot to ask of a relatively young and diverse discipline that is still in development. But perhaps the professional organisation of the field, the Learner Corpus Association, could lead the way in this endeavour. One recent initiative that aims at standardisation in corpus description concerns the collection and documentation of metadata in LCR. In view of the many variables that affect SLA, rich metadata are crucial. In LCR, however, research data management seems to have attracted less attention. A recent initiative (König et al. 2022) is aimed at the standardization of corpus description including metadata at the level of the corpus as a whole, but also metadata used to describe the individual learners and, importantly, task types and instructions. This would enhance corpus usability and comparability. König et al. (2022) propose a metadata schema
Challenges in learner corpus data
that features a set of core metadata fields considered necessary to describe learner corpora consistently and informatively. Finally, given that LCR is still biased towards written learner corpora that typically contain only the final product of the writing process, it is important to consider that more and more electronic writing tools have become available to help the learner. More recently, the field has seen some initiatives and innovative research that seeks to explore the actual writing process, for example by means of keystroke logging to examine textual revisions (Gilquin 2022), the use of electronic writing tools (Gilquin & Laporte 2021) and intertextual resources and text structuring (Wiemeyer 2022). However, new developments in the use of generative AI tools for text production provide challenging implications for learner corpus compilation. As mentioned earlier, ensuring the authenticity of learner writing is essential. When creating a corpus of student writing, there is a risk that samples generated or heavily modified by means of AI tools (if used unethically) may enter the corpus which, however, do not accurately reflect students’ own linguistic abilities. While these challenges have not yet been addressed explicitly in LCR, they are discussed for example in the assessment of writing (Hartwell & Aull 2023). One solution could be to restrict writing samples that are included in a learner corpus to those that were produced in controlled (online) contexts as in the compilation of the International Corpus Network of Asian Learners of English (ICNALE; Ishikawa 2023).
References Alexopoulou, Theodora, Geertzen, Jeroen, Korhonen, Anna & Meurers, Detmar. 2015. Exploring big educational learner corpora for SLA research: Perspectives on relative clauses. International Journal of Learner Corpus Research 1(1): 96–129. Callies, Marcus. 2008. Easy to understand but difficult to use? Raising constructions and information packaging in the advanced learner variety. In Linking Contrastive and Learner Corpus Research [Language and Computers, Studies in Practical Linguistics Series 66], Gaëtanelle Gilquin, Szilvia Papp & María Belén Díez-Bedmar (eds), 201–226. Amsterdam: Rodopi. Callies, Marcus. 2016. Towards a process-oriented approach to comparing EFL and ESL varieties: A corpus-study of lexical innovations. International Journal of Learner Corpus Research 2(2): 229–250. Callies, Marcus. 2023. Errors and innovations in L2 varieties of English: Towards resolving a contradictory practice. In Contradiction Studies – Exploring the Field, Gisela Febel, Kerstin Knopf & Martin Nonhoff (eds), 201–214. New York NY: Springer.
65
66
Marcus Callies
Callies, Marcus & Wiemeyer, Leonie. 2017. Multilingual speakers, multilingual texts: Multilingual practices in learner corpora. In Challenging the Myth of Monolingual Corpora [Language and Computers 80], Arja Nurmi, Tanja Rütten & Päivi Pahta (eds), 80–94. Leiden: Brill. Callies, Marcus & Hehner, Stefanie. 2021. Konstruktionen mit Partikelverben in Varietäten des Englischen: Zum Spannungsfeld von Präskription und Innovation an der Schnittstelle von Sprachwissenschaft, Fremdsprachendidaktik und Unterrichtspraxis. In Sprachwissenschaft und Fremdsprachendidaktik: Konstruktionen und Konstruktionslernen, Christoph Bürgel, Paul Gévaudan & Dirk Siepmann (eds), 81–93. Tübingen: Stauffenburg. Gilquin, Gaëtanelle. 2015. From design to collection of learner corpora. In Cambridge Handbook of Learner Corpus Research, Sylviane Granger, Fanny Meunier & Gaëtanelle Gilquin (eds), 9–34. Cambridge: CUP. Gilquin, Gaëtanelle. 2022. The Process Corpus of English in Education: Going beyond the written text. Research in Corpus Linguistics 10(1): 31–44. Gilquin, Gaëtanelle, De Cock, Sylvie & Granger, Sylviane (eds). 2010. The Louvain International Database of Spoken English Interlanguage. Handbook and CD-ROM. Louvain-la-Neuve: Presses universitaires de Louvain. Gilquin, Gaëtanelle & Laporte, Samantha. 2021. The use of online writing tools by learners of English: Evidence from a process corpus. International Journal of Lexicography 34(4): 472–492. Granger, Sylviane, Meunier, Fanny & Gilquin, Gaëtanelle (eds). 2015. Cambridge Handbook of Learner Corpus Research. Cambridge: CUP. Hartwell, Kelly & Aull, Laura. 2023. Editorial introduction: AI, corpora, and future directions for writing assessment. Assessing Writing 57: 1–4. Ishikiwa, Shin’ichiro. 2023. The ICNALE Guide An Introduction to a Learner Corpus Study on Asian Learners’ L2 English. New York NY: Routledge. Kaszubski, Przemysław. 2005. Typical errors of Polish advanced EFL learner writers. (currently not accessible). König, Alexander, Frey, Jennifer-Carmen, Stemle, Egon W., Glaznieks, Aivars & Paquot, Magali. 2022. Towards standardizing LCR metadata. Paper presented at the 6th International Conference for Learner Corpus Research (LCR 2022), Padova, 22–24 September. Lozano, Cristóbal & Mendikoetxea, Amaya. 2010. Interface conditions on postverbal subjects: A corpus study of L2 English. Bilingualism: Language and Cognition 13(4): 475–497. Lüdeling, Anke & Hirschmann, Hagen. 2015. Error annotation systems. In The Cambridge Handbook of Learner Corpus Research, Sylviane Granger, Gaëtanelle Gilquin & Fanny Meunier (eds), 135–158. Cambridge: CUP. Nelson, Gerald. 2002. International Corpus of English. Markup Manual for Written Texts. (15 March 2024). O’Donnell, Matthew Brook, Römer, Ute & Ellis, Nick C. 2013. The development of formulaic language in first and second language writing: Investigating effects of frequency, association, and native norm. International Journal of Corpus Linguistics 18: 83–108.
Challenges in learner corpus data
Ortega, Lourdes. 2013. SLA for the 21st century: Disciplinary progress, transdisciplinary relevance, and the bi/multilingual turn. Language Learning 63(1): 1–24. Paquot, Magali & Callies, Marcus. 2020. Promoting methodological expertise, transparency, replication, and cumulative learning: Introducing new manuscript types in the International Journal of Learner Corpus Research. International Journal of Learner Corpus Research 6(2): 121–124. Selinker, Larry. 1972. Interlanguage. International Review of Applied Linguistics 10: 209–231. Sinclair, John McH. 2005. Corpus and text – Basic principles. In Developing Linguistic Corpora: A Guide to Good Practice, Martin Wynne (ed.), 1–16. Oxford: Oxbow Books. SPLLOC Transcription Guidelines 2008. (1 August 2022). Tracy-Ventura, Nicole & Paquot, Magali (eds). 2021. The Routledge Handbook of Second Language Acquisition and Corpora. London: Routledge. Wiemeyer, Leonie. 2022. Intertextuality in Foreign-language Academic Writing in English. A Mixed-methods Study of University Students’ Writing Products and Processes in Sourcebased Disciplinary Assignments. PhD dissertation, University of Bremen.
67
Early newspapers as data for corpus linguistics (and Digital Humanities) Issues in using the British Library Newspapers database as a corpus Turo Hiltunen
University of Helsinki
The availability of large digital archives has great potential for corpus linguistic research, but their use is not without problems. These problems can often be traced to fundamentally different ideas of what might constitute “good data” in Digital Humanities and in corpus linguistics, leading to different expectations regarding how the data is made available to researchers. This chapter discusses the specific challenges involved in using the British Library Newspapers database for corpus linguistics and considers potential solutions for them. It is argued that, to take full advantage of the database, it is necessary to adopt a flexible approach enabling a critical reflection on the digital materials, how they have been collected, processed, and made available. Keywords: corpus compilation, Digital Humanities, sampling, representativeness, register
1.
Introduction
The availability of massive text archives holds great promise for corpus linguistic work, but at the same time they also present considerable methodological challenges for users (see, e.g., Hiltunen, McVeigh & Säily 2017). This statement may be obvious to corpus linguists, who for decades have been building carefully structured corpora and using them for research. Much has been written about corpus compilation and features that are generally characteristic of good linguistic corpora, such as representativeness and balance (Sinclair 2005; McEnery, Xiao & Tono 2006). A challenge with compiling such corpora has always been the large amount of work that is involved, particularly if the compilation requires source
https://doi.org/10.1075/scl.118.05hil © 2024 John Benjamins Publishing Company
Early newspapers as data for corpus linguistics
data to be transcribed and corrected manually, which is typically the case with both spoken corpora and historical corpora. The recent proliferation of large digital archives as part of Digital Humanities research projects has raised the question of whether they could be readily used as linguistic corpora, which would allow scholars to spend less time on the compilation and pre-processing of data, and more time on actual research. While this prospect is extremely attractive, anecdotal evidence suggests in many cases this is far less straightforward than what might seem to be the case at the outset (as also noted by Vartiainen & Säily in Chapter 2 in the present volume) and that there are many pitfalls that could limit the effectiveness of this approach, and even undermine the validity of the results that are obtained. The aim of this chapter is to reflect on some of the pitfalls involved in using the British Library Newspapers database as a corpus and consider some ways in which they can be avoided. I argue that the specific issues encountered with this database are due to fundamentally different ideas of what might constitute “good data” in the fields of Digital Humanities and corpus linguistics, respectively, and this in turn leads to different expectations regarding how the data is made available to researchers (e.g., Mehl 2021). To use large, digitised archives as data for corpus linguistic research, it is essential to critically think about the ways in which they have been collected and processed, as well as what kinds of software tools are available to search them. In practical terms, this may involve different processes of “remediation” (Kichuk 2007) and conversion from one data format to another, extracting samples using the available metadata, and identifying possible errors in the source data and considering what can be done about them. In addition, a frequent source of possible errors is the notion of register, which is a key concept in corpus linguistics but one that in other areas is typically conceptualised differently or even overlooked altogether. Working out suitable solutions to these issues is clearly essential if we want to capitalise on the potential of the British Library Newspapers database for corpus linguistic research (Hiltunen 2021). And as will be discussed below, conventional workflows relying on built-in interfaces and standard corpus linguistic tools may not be sufficient for identifying such issues, but better results can be obtained with a more flexible, exploratory approach using tools like Octavo (Mäkelä 2021), which allows users to extract tailorable information from the complete database and process it further. Section 2 of this chapter provides a brief outline of different approaches to digital text analysis. Section 3 identifies three pitfalls specific to the British Library Newspapers database and reviews some possible solutions to them. In Section 4 I summarise the main findings and discuss their implications to the corpus linguistic study of historical newspaper discourse.
69
70
Turo Hiltunen
2.
Digital text analysis in the humanities
There is a high degree of consensus across humanities disciplines that the emergence of digitised materials and techniques for analysing them is a major advantage for scholarship. For example, in recent Digital Humanities literature, a great deal has been written on the nature of these advantages over “traditional” ones. The argument is that owing to the availability of new sources of data, digital approaches have great potential for increased productivity, enabling scholars to automate many of the tasks that previously have taken up a lot of time. Another widely recognised advantage of the proliferation of data is the possibility to obtain new perspectives on old questions which simply would not have been feasible previously, and this is seen as conducive to significant new epistemological advances. These views are nicely captured in the two quotations by McCarty (2012) and Gale (2023) below; the first one is from a reflective essay, the second one from a website advertising the strengths of a specific text-analytic platform, Gale Digital Scholar Lab, which is one of the software tools through which the British Library Newspapers database can be accessed (Section 3.2). The phrase [a telescope for the mind?] is Margaret Masterman’s; the question mark is mine. […] She used the phrase to suggest computing’s potential to transform our conception of the human world just as in the seventeenth century the optical telescope set in motion a fundamental rethink of our relation to the physical one. The question mark denotes my own and others’ anxious interrogation of research in the digital humanities for signs that her vision, or something like it, is being realized or that demonstrable progress has been made. (McCarty 2012: 113) Gale Digital Scholar Lab provides a new lens to explore history.
(Gale 2023)
Very similar discourse is often found in corpus linguistic literature in support of a corpus-based approach. For example, in their widely used textbook Corpus Linguistics: Theory, Method and Practice, McEnery and Hardie (2011) talk about the potential of corpus linguistics to “reorient our entire approach to the study of language”: It may refine and redefine a range of theories of language. It may also enable us to use theories of language which were at best difficult to explore prior to the development of corpora of suitable size and machines of sufficient power to exploit them. (McEnery & Hardie 2011: 1)
However, to discuss possible pitfalls associated with research at the interface of Digital Humanities and corpus linguistics, it is necessary to briefly review some basic assumptions associated with these respective areas of study.
Early newspapers as data for corpus linguistics
2.1 Digital Humanities Central to the field of Digital Humanities are computers and their application to the analysis of large masses of text, but beyond that, this broad field can be defined in different ways. In a recent article, Mehl (2021) assigns the following attributes to Digital Humanities: it is about using language data to understand different humanities and social scientific questions, and a technical practice associated with digitisation, creation of digital editions; research practice in this field essentially amounts to taking unstructured datasets, finding/applying some structure, and using that for analysis. In her textbook, Drucker (2021) identifies three components as being the essential components of a Digital Humanities project and workflow: materials, processing, and presentation. For her, research in the Digital Humanities begins with materials (images, texts, sound files, etc.), which, in many cases, involves remediation (i.e., converting the information from one format into another). These materials are subjected to computational processing (data mining/statistical analysis), the outcomes of which are organised in a presentation and preserved for future research purposes. Given the breadth of the field of Digital Humanities, sub-disciplinary specialisms naturally exist within the field. Roth (2018) distinguishes between three orientations: 1. Digitised humanities (research broadly relying on digitized data) 2. Numerical humanities (which focuses specifically on numerical and, more broadly, formal models, e.g., computational social science or social informatics) 3. Humanities of the digital (i.e., the study of Computer-mediated interaction) He further argues that (1) and (2) represent different epistemic communities, between which there is not that much overlap or interaction; Digital Humanities journals predominantly contain papers representing (1), and Computational Social Science journals predominantly focus on quantitative models, and traditional Digital Humanities papers are only marginally represented. A similar situation clearly obtains for Digital Humanities and corpus linguistics: the goals and general perspectives in these fields are frequently intertwined, while simultaneously we find considerable differences between individual specialisms and research projects. As previously indicated, the role of computing in Digital Humanities is often framed with the help of a visual metaphor as being that of an instrument that enables scholars to get a new perspective on the data that would not be possible through any other means. Irrespective of the orientation, it is generally agreed that advantages of a DH approach are said to be epistemological: by making available large amounts of data, DH provides new ways of looking at data and enables scholars to ask new questions and to be more productive.
71
72
Turo Hiltunen
2.2 Corpus linguistics It is clear that many of the points made in the previous section apply equally well to corpus linguistics, where research is based on large amounts of language data which have been processed and turned into linguistic corpora. Even the same metaphors are used: As the telescope is indispensable to the astronomer, so the corpus is indispensable to the linguist. (McEnery & Hardie 2011: 53)
Such similarities are of course not surprising, given that both Digital Humanities and corpus linguistics deal with textual data. In fact, the field of corpus linguistics is often conceptualised as being part of the Digital Humanities (e.g., Roth 2018; Zottola 2020; Mehl 2021) despite the fact that it has a longer history. In the modern sense, the term “corpus linguistics” became established early in the second half of the 20th century, with such milestones as the publication of the Brown Corpus (1964) and the formation of The International Computer Archive of Modern English (ICAME) in the early 1970s (McEnery & Hardie 2013). By contrast, the rise of the term “Digital Humanities” can be dated to the first decade of the 2000s, when it took over as the main category of self-identification among scholars working in areas previously known as “humanities computing”, “humanist informatics”, and “literary and linguistic computing” (Nyhan, Terras & Vanhoutte 2010: 2). The obvious affinity between these two fields has also been noted by many writers, although Jensen (2014) notes that corpus linguistics only plays a peripheral role under the DH umbrella. However, there is one major difference that sets corpus linguistics apart from the Digital Humanities, and this relates to the definition of the term corpus. Corpus linguistics typically adopts a narrow definition: a corpus is not just any collection of texts, but a collection that has been selected to represent a language or a sublanguage following an extensional view of language (e.g., Baroni & Evert 2009). A logical consequence of this position is that any database that does not match this definition is unlikely to meet all the expectations that a corpus linguist might have. This is clearly the case with the British Library Newspapers database.
2.3 Towards a useful synergy The differences between Digital Humanities and corpus linguistics should of course not be overstated, given that the two fields share many research goals and practices. For example, the use of unstructured archives as data is not exclusive to Digital Humanities, but they are also made use of by corpus linguists as “opportunistic corpora” (McEnery & Hardie 2011: 11), which may provide large amounts
Early newspapers as data for corpus linguistics
of useful data to study linguistic questions. At the same time, there are clearly notable differences in the stance that is taken to many key questions, the form in which the data is made available to researchers, the tools and methods needed to use the data, and ultimately, the aims of research projects. Indeed, given that the term humanities has a much wider scope than linguistics, it is hardly surprising that there would be vastly different aims, interests, and also types of materials. In this sense, it is reasonable to see their activities as occupying different “intellectual territories” (Becher & Trowler 2001), and as such, representing different epistemic communities (cf. Roth 2018). The divergence between Digital Humanities and corpus linguistics may manifest itself not only intellectually, but also culturally as ways of behaving, and generally the kinds of cultural frames within which work is situated (cf. Geertz 1983: 155). Insofar as the cultural frames are different for corpus linguists and Digital Humanities scholars (including more specific denominations like ’digital historians’ or ’digital social scientists’), we can arguably talk about different disciplinary communities in both intellectual and social sense. The argument that I want to make in this chapter is that in order to take advantage of the synergy between corpus linguistics and Digital Humanities, it is often necessary to critically reflect on the digital materials, how they have been collected, processed, and made available. In particular, it is often the case that the default workflow for accessing a specific archive is not ideal to tackle corpus linguistic research questions, which may make it difficult to compare results to previous work.1 To avoid such pitfalls, adjustments are clearly needed, and this often requires carefully considering the influence of register, which is a crucial notion in most types of corpus linguistics. If this is possible – and in particular if the availability of data is not tied to a particular platform or infrastructure – we can improve the quality of opportunistic corpora and ultimately obtain results that are more accurate and are more easily situated within previous research. My aim here is to illustrate this argument by identifying and describing issues that are specific to the British Library Newspapers database, an archive that has clearly not been put together with corpus linguistic research designs in mind, presenting some possible solutions to them, and reflecting on the overall feasibility of these solutions.
1. For a fuller discussion of replicability and reproducibility specifically in the context of corpus linguistics, see Flanagan (forthcoming). Digital Humanities researchers might of course have similar reservations about research practices and data in corpus linguistics (e.g., corpus size, use of samples instead of whole texts, focus on only linguistic research questions).
73
74
Turo Hiltunen
3.
Historical newspaper prose and the British Library Newspapers database
The linguistic study of historical newspapers has become a burgeoning area of research, and plenty of corpus-based work exists on this register (e.g., Jucker 1992; Conboy 2010; Duguid 2010; Landert 2014; Rühlemann & Hilpert 2017; for an overview, see, e.g., Jucker 2009; Percy 2012; Fries 2012). Yet the problem of much of previous corpus work has been that existing corpora are relatively small and limited; for instance, the Zurich English Newspaper (ZEN) corpus (2004) consists of 349 newspaper issues from London newspapers 1661–1791, with the total word count of 1.6 million words.2 To peruse larger masses of historical newspaper language, scholars can nowadays turn their attention to the British Library Newspapers database, which is a massive digital archive of newspaper writing archive hosted by Gale. It contains around 5.5 million pages from national and regional newspapers from Britain between the 18th and the 20th centuries. To corpus linguists, however, the database also presents multiple issues, which need to be addressed to make full use of it. In the following sections, I will focus on four such issues: 1. 2. 3. 4.
Problem with available search tools Sampling, balance, representativeness Register/sub-register considerations Quality of Optical Character Recognition (OCR)
3.1 Problems with available search tools The British Library Newspapers database can be accessed through the Gale Primary Sources platform,3 together with other proprietary digital databases. The platform enables users to type search words and retrieve all the documents in the database where the search word occurs. The search results are available as facsimile images and plain text. In addition, the platform includes a “topic finder”,4 which generates visualisations of “keywords” based on article titles, subjects, and top words co-occurring with the search term (e.g., keywords obtained for war include civil, Boer, and peace). To a corpus linguist, the main shortcoming of Gale Primary Sources is the fact that the searches produce lists of documents, effectively suggesting that the next 2. See 〈https://varieng.helsinki.fi/CoRD/corpora/ZEN/〉 3. 〈https://www.gale.com/intl/primary-sources〉 4. 〈https://www.gale.com/intl//research-tools/tools-for-discovery〉
Early newspapers as data for corpus linguistics
logical step in the analysis would be to focus on individual texts and their particularities. While such an “individual-document approach” (McEnery & Hardie 2011: 232) might indeed be appropriate for some purposes, corpus linguists would usually be interested in looking at a concordance to identify patterns in the use of the search term or calculate and compare its normalised frequency across different sections of the database. With Primary Sources this is only possible if the documents are individually downloaded and analysed using some other corpus tool. To some extent, these shortcomings are addressed by Gale’s Digital Scholar Lab,5 a text analytics platform through which the British Library Newspapers database can also be accessed. It enables users to build “content sets” based on search terms and apply different Digital Humanities tools on them, including document clustering, named entity recognition, n-grams, parts-of-speech, sentiment analysis, and topic modelling. One clear advantage of Digital Scholar Lab over Primary Sources is that it allows entire content sets to be downloaded by one action. However, it should be noted that the number of documents to download is currently limited to 10,000, which for many sets is insufficient for a comprehensive analysis. The usefulness of the available text mining tools is also seriously limited by their lack of customisability. For example, the part-of-speech tool (using spaCy) only offers visualisations of POS-frequencies but does not give direct access to the tagged texts or include them in the downloadable content sets, which effectively means that the morpho-syntactic annotation provided by Digital Scholar Lab is of no use for corpus research. Fortunately, there are alternative ways of accessing the textual data. The massive full-text archive of British Library Newspapers is available to subscribing institutions as plain text XML files, and with the help of Octavo, a tool for Digital Humanities research created by Eetu Mäkelä (Mäkelä 2021), it is possible to search the indexed archive in its entirety and obtain complete lists of documents and their accompanying metadata (e.g., name of publication, date, and city) as JSON files, which can be trivially converted to other formats like XML or plain text. This workflow is preferable to Digital Scholar Lab: it enables the quick creation of tailor-made samples from the database based on specific metadata items. In this way, it is also possible to add further levels of annotation to the texts, for example part-of-speech tagging, and make use of them when doing corpus linguistic searches. To sum up, as many workflows in corpus linguistics rely on the availability of data as plain text, it is clear that the possibility to access the data in the British Library Newspapers database using Octavo addresses many of the concerns that 5. 〈https://www.gale.com/intl/primary-sources/digital-scholar-lab〉
75
76
Turo Hiltunen
a corpus linguist might have regarding the use of the British Library Newspapers database as a corpus. However, even with this approach, there are specific issues that researchers still need to consider that come from the fact that the database itself was not designed for corpus linguistic research purposes. To illustrate these issues, I will refer to the datasets used in Hiltunen (2021) to study register features of 19th-century newspaper writing.
3.2 Sampling, balance, and representativeness While a Digital Humanities research project might start by mining the entire digital archive in a bottom-up fashion, this approach is often seen as problematic in corpus linguistics, as lack of structure in the dataset makes it difficult to interpret the findings (see, e.g., Davies 2019). For this reason, corpus linguistics has traditionally valued carefully structured corpora, even if that comes at the cost of large size and comprehensiveness (e.g., Hundt & Leech 2012). To a corpus linguist interested in using the British Library Newspapers database, the main concern would thus be how to build a dataset that would not only be maximally large, but also meaningfully structured in terms of register distribution, and not ridden with errors. With the help of Octavo, we can try to answer this question empirically: using different parameters, we can create corpora and compare them in terms of their usefulness in different research tasks. This of course raises the question of how to best assess representativeness. Representativeness and balance of data are known to be tricky issues in corpus linguistics in general: while Sinclair (2005) and others have defended the use of representativeness as a guiding principle of corpus compilation, many have also noted that true corpus representativeness cannot probably be attained. For example, for McEnery, Xiao and Tono (2006: 21), there are no objective ways of assessing representativeness and balance, and that the appropriateness of a corpus ultimately depends on the specific research questions. A related issue with historical data in particular is that corpus compilation is inevitably linked to what texts are currently available, which may not match exactly the original population at the time when the source texts were produced (Nicholson 2012). Although the database includes articles from both 18th and 20th century, my focus here is exclusively on the 19th century, which is far more comprehensively represented in terms of the newspaper issues available. Even so, it is immediately clear that the distribution of texts is not even: as can be seen in Figure 1, the total number of issues increases considerably over the course of the century. In addition, the database contains in total 155 newspapers, but not all of them appear throughout the century. For example, while some, like Leeds Mercury (red) cover more or less the entire century, others only start to appear at the end of it, which
Early newspapers as data for corpus linguistics
is the case with The Portsmouth Evening News (blue). Crucially, all these factors are invisible to users accessing the database through Gale Primary Sources or Digital Scholar Lab, as the relevant information can only be obtained with the help of querying the database with Octavo and carefully exploring the results.
Figure 1. Overall distribution of 19C issues
Choosing the appropriate sampling strategy depends on the aims of the research, and the chosen method should also take into account the contents of the database itself. For example, the most basic sampling method, a simple random sampling, is not sensitive to the increase in available issues and might result in the second half of the 19th century being overrepresented in the sample. If this is deemed problematic, then a stratified sampling strategy, where the database is first divided into smaller, temporally-defined strata (e.g., periods of 1, 5, or 10 years) before extracting the samples, is likely to be a better choice. Stratified sampling is also more appropriate if the aim is to investigate differences between different newspapers and their linguistic characteristics, both synchronically and diachronically. To achieve this, it is necessary to first identify newspapers like Leeds Mercury, which are sufficiently well-represented in the database and meaningfully cover the entire century, and use them as the strata (McEnery, Xiao & Tono 2006: 20) from which the samples are extracted. An example of such a stratified sample is one of the corpora used in Hiltunen (2021), where
77
78
Turo Hiltunen
samples were taken from newspapers originating in different parts of Britain: Bath Chronicle and Weekly Gazette, Bury and Norwich Post, Morning Post (London) and York Herald.6 To sum up, relying on built-in interfaces (Primary Sources and Digital Scholar Lab) to access the materials in the British Library Newspapers database is potentially risky for corpus linguistic research: not only do they not offer direct access to all of the available linguistic data, but they also fail to highlight the uneven distribution of texts across years and newspaper titles, which is important information for building corpora. Using a tool like Octavo is therefore a necessity. Regarding sampling, the takeaway is that the choice of sampling strategy requires background research and exploration into the structure and coverage of the database. Even though this information is not directly available and accessing it requires some effort, it is crucial for obtaining samples that are meaningful and appropriate for specific corpus-linguistic research tasks. And as we will see below, these considerations related to representativeness and balance also apply to the quality of the source data and to registers, which is a topic that we will consider next.
3.3 Registers and subregisters The importance accorded to register in contemporary corpus linguistics can hardly be overstated. Register has been established as a major determinant of linguistic variation (Biber 1988; Biber & Finegan 1994), and most contemporary corpora are designed and compiled with some notion of register (or related constructs like genre) as the core structuring parameter (e.g., Biber, Finegan & Atkinson 1993; Fries & Schneider 2000; Biber & Egbert 2018). It has likewise been identified as an important factor in the study of newspaper language (Ljung 2000; Fries 2012; Percy 2012). Biber and Gray (2013: 111) have recently argued in favour of finer granularity in studies of newspaper language, as seemingly minor differences in audience, purpose, and even editorial policy often result in systematic differences in the linguistic form of texts. This being the case, meaningful sub-divisions are clearly important for making reliable generalisations based on the British Library Newspapers database. With this database, a major practical concern has to do with how the concepts of register and subregister are operationalised, and what implications this may have for the notions of representativeness and balance discussed above. The database contains potentially relevant information about text categories under two headings: Document type contains information about “the format, genre, or other characteristics of the document”, and Publication section allows 6. This corpus (referred to as Corpus A in the original study) will also be discussed below to illustrate other pitfalls.
Early newspapers as data for corpus linguistics
users to limit searches to a specific section of the publication. Of these, the latter indicates whether the text belongs to section A, B, etc., and as such is of limited use to register analysis. Document type is much more useful: it categorises each text as belonging to one of the seven types listed in Table 1. Table 1. Text categories (‘document types’) available in the British Library Newspapers database Text category Arts and entertainment Birth, death, marriage notices Business Classified ads Editorial News Sports
The possibility of using these categories is a major methodological convenience, as the alternative approach of manually classifying corpus texts becomes unwieldy with a large number of texts. The samples used in Hiltunen (2021) were extracted using these categories, as it was initially deemed that they offer a sufficiently informative and accurate classification that enables register-aware analyses of the universe of texts in the database. This assumption was also borne out by the analyses of that study, seeing as the chosen approach produced meaningful results that could be interpreted with reference to characteristics of registers. With that said, the adoption of the text categories in Table 1 clearly has its caveats. The database offers hardly any information about the definitions of these categories and the criteria for assigning individual texts to them. For example, the category Editorial includes both articles authored by the newspaper editors and letters to the editor. It is thus not clear how these categories would map onto other register and genre labels used in previous research, although correspondences can certainly be established with the categorisation in Fries (2012). Nor do we know whether any individual text is a prototypical instance of a category or a borderline case that could also be assigned to another one. But even if the existing text categorisation is accepted as a basis for corpus compilation, it is still necessary to consider the implications from the perspective of representativeness. The register categories (like newspaper titles in the previous section) are not evenly distributed across the database, and different choices will inevitably result in different representations of the text-external reality, and as before, this information is only accessible to users systematically looking into this.
79
80
Turo Hiltunen
This issue, too, can be illustrated with the help of the same sample corpus from Hiltunen (2021), Corpus A, which comprises complete issues of the four newspapers sampled at 10-year intervals. Out of the 3,475 texts, approximately 88% (3,081) represented a single text category, namely the category News; this is shown in Figure 2. A corollary of this is that any results based on this corpus are mainly representative of the register of news in the four newspapers sampled, and only marginally of the other registers.
Figure 2. Distribution of text categories in Corpus A (Hiltunen 2021)
Another potential pitfall with Corpus A emerges when we look at the lengths of samples and their chronological distribution, displayed in Figure 3. As can be seen, the samples from 1880 include several Classified ads (purple cross) whose length is between 15,000 and 30,000 words – a dramatic difference compared to the majority of other text samples in the corpus, which are considerably shorter. It is clear that these long samples are not in fact individual texts, but rather collations of several short classified ads that appear in the same newspaper issue. While grouping such short texts together may not matter for some research designs – and could even be a reasonable thing to do – it is clear that in the British Library Newspapers database, the notion of “text” is operationalised differently for Classified ads compared to other text categories. This in turn may raise further methodological issues, as comparing short and long texts is not always straightforward (e.g., Liimatta 2022, and Chapter 7 of the present volume).
Early newspapers as data for corpus linguistics
Figure 3. Distribution of article types and lengths in Corpus A (Hiltunen 2021). (jitter has been added to reduce overplotting)
Summing up the discussion of registers, it is demonstrably possible to quickly create register-specific sample corpora with the help of Octavo and the available metadata on text categories. However, these corpora are not necessarily equally representative of all subregisters of newspaper writing, given that the text categories are not evenly distributed in the database. This unevenness is in itself not surprising but rather a central trait of the register: we expect there to be a large number of News items in a single issue of a newspaper, but only a few Editorials (or just one). What is important is that this information is visible and easily accessible to a corpus linguist so that it can be taken into account in data extraction.
3.4 Optical Character Recognition (OCR) The third major pitfall in the context of the British Library Newspapers database is the poor accuracy of Optical Character Recognition (OCR). OCR quality is a well-known issue in the Digital Humanities in general, and specifically for digitised newspapers (e.g., Tanner, Munoz & Hemy Ros 2009; Hill & Hengchen 2019). It quickly becomes obvious to even a casual user of the British Library Newspapers database that OCR errors are ubiquitous. To illustrate, Figure 4 shows a snippet from an article from The Morning Post (1805).
81
82
Turo Hiltunen
Figure 4. Image from The Morning Post, 1 January 1805
The digitised text corresponding to the three paragraphs is reproduced as Examples (1) to (3): (1)
Yesterday was launched from the King’s Yard, Dfjxforr!, the Hebe frigate, a fine new ship ol” 3 > gun. Her Royal Highness the Princess of Wales MM:, present on the occasion.
(2)
Saturday afternoon were landed from the Malta, 8.1 gons, Capt. Bulier, at Plymouth, sever.il barrels, containing nearly 60,000 dollars in silver, cinsitned from merchants in Spain to their cor-tesponiients in London. They were deposited in Rlssel’s waggon warehouses previous to their being sent to London under a proper escort.
(3)
The circumstance of Mr. Addington being ■shout to be appointed Speaker of the House ©f Lou!-., lias given rise to a report of the following] changes, which we mention without vouching the ’ avi-uracy ef any part of the statement :I
It is evident that the accuracy of the text depends crucially on the quality of the image used as the basis of the digitised text. Contemporary OCR software is able to reach a high accuracy with clean, present-day English text (Prescott 2018). However, in Figure 4, most words on the left side of the column are unclear – for example, Deptford, consigned, Russel, and accuracy – and they have been incorrectly identified by the OCR software as Dfjxforr!, c-insitned, Rlsse, and avi-uracy, respectively. What is less clear is how serious this problem is specifically for corpus linguistic research, and what should be done about it. The presence of such words is clearly problematic for both recall and precision of searches, but is this problem offset by the large amount of material that is available?
Early newspapers as data for corpus linguistics
Intuitively, the usefulness of digitised texts depends crucially on the overall frequency and distribution of the OCR errors that they contain, and if this is the case, then there indeed appears to be reason for some concern: Tanner, Munoz and Hemy Ros (2009) have estimated the average OCR accuracy across the British Library Newspapers database to be 83.6% for characters and 78% for words, and suggest that if word accuracy is higher than 80%, a fuzzy search engine would nonetheless be able to reach a high search accuracy. However, achieving this across the entire database appears unrealistic, given that this level of word accuracy was only reached by a quarter of the texts in their sample, and this impression is borne out by the trial searches reported by Prescott (2018) on the Burney collection,7 which yielded a low success rate. On a more positive note, it has been shown that not all corpus linguistic research tasks are equally sensitive to OCR errors, with for example, the identification of frequent collocations providing robust results even with texts containing relatively large numbers of errors (Hill & Hengchen 2019). Similarly, it might be possible to achieve a reasonably good recall (albeit with poor precision) using fuzzy searches or wildcards. While this may be true, it is good to remember that one standard procedure in corpus research, the counting of normalised frequencies of linguistic features, does not only depend on recall but also on the token word count of the text, which becomes unreliable with texts containing lots of OCR errors. Texts with high error rates are also problematic for adding further layers of annotation, such as POS-tagging. Therefore, it seems that in general, corpus linguistic research has a comparatively low tolerance for errors. As error correction is generally unfeasible on a larger scale (Gregory et al. 2016), this implies that we need to find a way to exclude texts with high error rates. To do this, we can make use of the figures for OCR confidence (0–99.99%), which are available for each text in the British Library Newspapers database and “[represent] the OCR engine’s confidence in the accuracy of the conversion from image to text”. The OCR confidence value corresponding to each text can be conveniently accessed with Octavo, like any other metadata. While it is unclear how the values for OCR confidence are determined, Hill and Hengchen (2019: 837) nonetheless conclude that “there is evidence that one can be confident in them”. After evaluating a number of files representing different levels of OCR confidence, it was decided that texts reaching 90% were sufficiently clean to provide accurate research, and this criterion was adopted for Corpus A. This ensures the extracted sample corpus is reasonably clean and can therefore be expected to provide reliable output for standard corpus linguistic tasks. However, the disadvantage is that we also discard an enormous amount of potentially interesting lin7. See 〈https://www.bl.uk/collection-guides/burney-collection〉
83
84
Turo Hiltunen
guistic material. For example, Corpus A only represents 17% (3,475 of 21,440) of the texts that matched the original search criteria, with 83% filtered out due to not meeting the threshold.8 As we do not know how the OCR errors are distributed across newspaper titles or text categories, it is difficult to assess exactly how serious this issue is for the representativeness of individual samples. Here, too, the answer depends on the purpose and the features being studied. Based on earlier studies on word frequency distributions (e.g., Biber 1993), we can expect such samples to be in general fairly reliable for register analyses, which are based on multiple high-frequency features. Studies on less frequent features (e.g., individual grammatical constructions) are potentially riskier, as they are more susceptible to sampling variation. To sum up, removing low-quality texts is essential for accurate corpus linguistic work, and it can be accomplished by filtering out texts that do not meet a pre-determined OCR confidence value. The flipside is that this may potentially compromise the representativeness and balance of the filtered sample, and this needs to be evaluated on a case-by-case basis due to the uneven and erratic distribution of OCR errors in the database.
4. Discussion After identifying and reviewing four major pitfalls and suggesting possible ways of avoiding them, it is possible to offer some preliminary conclusions about the usefulness of the British Library Newspapers database for corpus linguistic research. Starting with the positives, an important advantage over traditional, relatively small linguistic corpora is that the database enables an exploratory data-driven approach with reasonable effort. In other words, as data extraction can easily be automated, it becomes straightforward to create multiple corpora with the search parameters and frequency thresholds and assess their suitability for different research tasks. The structure of the database also allows researchers to readily incorporate a register perspective into the study design. Finally, as it is possible to automatically discard low-quality samples, this workflow also enables the creation of reasonably tidy corpora. As a result, the British Library Newspapers database is an attractive alternative for the corpus-linguistic analysis of historical newspaper prose. However, there are also obvious caveats to consider. Compared to carefully constructed traditional corpora with often hand-picked text samples, there is obviously much less control over individual choices, and the fact that the exact 8. Including the text in Figure 4 (OCR confidence: 79.6%).
Early newspapers as data for corpus linguistics
basis of text categorisation is unclear is likewise not optimal. Yet the by far most serious issue is the presence of errors in the source material, which introduces errors to analyses, and, in the worst case, may compromise the representativeness of corpora. How problematic these issues are depends on the goal of the individual research project. As the database is large, the omission of some publications or volumes due to low OCR quality might not matter too much for the analysis of general trends in language and discourse, whereas it may effectively preclude the study of specific questions or narrower time periods. Given these caveats, corpus linguistics arguably still has a place for “small and tidy” (Mair 2006: 355) specialised corpora. At the same time, it is likely that with the increasing availability of digitised databases like the British Library Newspapers database, the role of Digital Humanities-style workflows will increase at the cost of traditional corpus compilation projects. What is needed in the compilation of high-quality corpora from large archives are flexible tools that enable researchers to access and manipulate the full-text data of the archive, and this may require interdisciplinary collaboration. Alongside this, the present chapter has shown how the notion of register, whose importance has been indisputable in traditional linguistic corpora, is equally relevant in corpora that are collected from archive data.
References Baroni, Marco & Evert, Stefan. 2009. Statistical methods for corpus exploitation. In Corpus Linguistics: An International Handbook, Vol. 2: Anke Lüdeling & Merja Kytö (eds), 777–802. Berlin: Mouton de Gruyter. Becher, Tony & Trowler, Paul. 2001. Academic Tribes and Territories: Intellectual Enquiry and the Culture of Disciplines. Buckingham: Society for Research into Higher Education & Open University Press. Biber, Douglas. 1988. Variation across Speech and Writing. Cambridge: CUP. Biber, Douglas. 1993. Representativeness in corpus design. Literary and Linguistic Computing 8(4): 243–257. Biber, Douglas & Egbert, Jesse. 2018. Register Variation Online. CUP. Biber, Douglas & Finegan, Edward. 1994. Sociolinguistic Perspectives on Register. Oxford: OUP. Biber, Douglas, Finegan, Edward & Atkinson, Dwight. 1993. ARCHER and its challenges: Compiling and exploring a representative corpus of historical English registers. In Creating and Using English Language Corpora, Udo Fries, Gunnel Tottie & Peter Schneider (eds), 1–13. Amsterdam: Rodopi. Biber, Douglas & Gray, Bethany. 2013. Being specific about historical change: The influence of sub-register. Journal of English Linguistics 41 (2): 104–134. Conboy, Martin. 2010. The Language of Newspapers: Socio-Historical Perspectives. London: Continuum.
85
86
Turo Hiltunen
Davies, Mark. 2019. Corpus-based studies of lexical and semantic variation: The importance of both corpus size and corpus design. In From Data to Evidence in English Language Research, Terttu Nevalainen, Carla Suhr & Irma Taavitsainen (eds), 66–87. Leiden: Brill. Drucker, Johanna. 2021. The Digital Humanities Coursebook: An Introduction to Digital Methods for Research and Scholarship. Abingdon, Oxon: Routledge. Duguid, Alison. 2010. Newspaper discourse informalisation: A diachronic comparison from keywords. Corpora 5(2): 109–138. Flanagan, Joseph. Forthcoming. Reproducibility, replication, robustness, and generalizability in corpus linguistics. In Reproducibility, Replication, and Robustness in Corpus Linguistics, Michael Haugh & Martin Schweinberger (eds). Special issue in International Journal of Corpus Linguistics. Fries, Udo. 2012. English and the Media: Newspapers. In English Historical Linguistics: An International Handbook, Alexander Bergs & Laurel Brinton (eds), 1063–75. Berlin: Mouton de Gruyter. Fries, Udo & Schneider, Peter. 2000. ZEN: Preparing the Zurich English Newspaper Corpus. In English Media Texts – Past and Present: Language and Textual Structure, Friedrich Ungerer (ed.), 3–24. Amsterdam: John Benjamins. Gale. 2023. Gale Digital Scholar Lab. (19 May 2024). Geertz, Clifford. 1983. Local Knowledge: Further Essays in Interpretive Anthropology. New York NY: Basic Books. Gregory, Ian Norman, Atkinson, Paul David, Hardie, Andrew, Joulain-Jay, Amelia, Kershaw, Daniel, Porter, Catherine, Rayson, Paul Edward & Rupp, Christopher John. 2016. From digital resources to historical scholarship with the British Library 19th Century Newspaper Collection. Journal of Siberian Federal University: Humanities and Social Sciences 9(4): 994–1006. Hill, Mark J. & Hengchen, Simon. 2019. Quantifying the impact of dirty OCR on historical text analysis: Eighteenth century collections online as a case study. Digital Scholarship in the Humanities 34(4): 825–843. Hiltunen, Turo. 2021. Exploring sub-register variation in Victorian newspapers: Evidence from the British Library Newspapers Database. In Corpus-Based Approaches to Register Variation [Studies in Corpus Linguistics 103], Elena Seoane & Douglas Biber (eds), 313–338. Amsterdam: John Benjamins. Hiltunen, Turo, McVeigh, Joe & Säily, Tanja. 2017. How to turn linguistic data into evidence? In Big and Rich Data in English Corpus Linguistics: Methods and Explorations [Studies in Variation, Contacts and Change in English], Turo Hiltunen, Joe McVeigh, & Tanja Säily (eds). Helsinki: Research Unit for Variation, Contacts, and Change in English. (19 May 2024). Hundt, Marianne & Leech, Geoffrey. 2012. ‘Small is beautiful’: On the value of standard reference corpora for observing recent grammatical change. In The Oxford Handbook of the History of English, Elizabeth Traugott & Terttu Nevalainen (eds), 175–188. Oxford: OUP. Jensen, Kim Ebensgaard. 2014. Linguistics in the digital humanities: (Computational) corpus linguistics. MedieKultur: Journal of Media and Communication Research 30 (57).
Early newspapers as data for corpus linguistics
Jucker, Andreas H. 1992. Social Stylistics. Syntactic Variation in British Newspapers. Berlin: Walter de Gruyter. Jucker, Andreas H. 2009. Newspapers, pamphlets and scientific news discourse in Early Modern Britain. In Early Modern English News Discourse: Newspapers, Pamphlets and Scientific News Discourse [Pragmatics & Beyond New Series 187] Andreas H. Jucker (ed.), 1–9. Amsterdam: John Benjamins. Kichuk, Diana. 2007. Metamorphosis: Remediation in early English books online (EEBO). Literary and Linguistic Computing 22(3): 291–303. https://10.1093/llc/fqm018. Landert, Daniela. 2014. Personalisation in Mass Media Communication: British Online News Between Public and Private [Pragmatics & Beyond New Series 240]. Amsterdam: John Benjamins. Liimatta, Aatu. 2022. Do registers have different functions for text length? A case study of Reddit. Register Studies 4(2): 263–287. Ljung, Magnus. 2000. Newspaper genres and newspaper English. In English Media Texts, Past and Present: Language and Textual Structure [Pragmatics & Beyond New Series 80], Friedrich Ungerer (ed.), 131–150. Amsterdam: John Benjamins. Mair, Christian. 2006. Tracking ongoing grammatical change and recent diversification in present-day standard English: The complementary role of small and large corpora. In The Changing Face of Corpus Linguistics, Andrew Kehoe & Antoinette Renouf (eds), 355–376. Leiden: Brill. Mäkelä, Eetu. 2021. Octavo. GitHub repository. (19 May 2024). McCarty, Willard. 2012. A telescope for the mind? In Debates in the Digital Humanities, Matthew K. Gold (ed.), 113–136. Minneapolis MN: University of Minnesota Press. McEnery, Tony & Hardie, Andrew. 2011. Corpus Linguistics: Method, Theory and Practice. Cambridge: CUP. McEnery, Tony & Hardie, Andrew. 2013. The history of corpus linguistics. In The Oxford Handbook of the History of Linguistics, Keith Allan (ed.), 727–745. Oxford: OUP. McEnery, Tony, Xiao, Richard & Tono, Yukio. 2006. Corpus-based Language Studies: An Advanced Resource Book. London: Routledge. Mehl, Seth. 2021. Why linguists should care about digital humanities (and epidemiology). Journal of English Linguistics 49(3): 331–337. Nicholson, Bob. 2012. Counting culture; Or, how to read Victorian newspapers from a distance. Journal of Victorian Culture 17(2): 238–246. Nyhan, Julianne, Terras, Melissa & Vanhoutte, Edward. 2010. Introduction. In Defining Digital Humanities. A Reader, Melissa Terras, Edward Vanhoutte & Julianne Nyhan (eds), 1–10. Farnham: Ashgate. Percy, Carol. 2012. Early advertising and newspapers as sources of sociolinguistic investigation. In The Handbook of Historical Sociolinguistics, Juan Manuel Hernández Campoy & Juan Camilo Silvestre Conde (eds), 191–210. Malden MA: Blackwell. Prescott, Andrew. 2018. Searching for Dr. Johnson: The digitisation of the Burney newspaper collection. In Travelling Chronicles: News and Newspapers from the Early Modern Period to the Eighteenth Century [Library of the Written Word 66], Siv Gøril Brandtzæg, Paul Goring & Christine Watson (eds), 51–71. Leiden: Brill.
87
88
Turo Hiltunen
Roth, Camille. 2018. Digital, digitized, and numerical humanities. Digital Scholarship in the Humanities 34(3): 616–632. Rühlemann, Christoph & Hilpert, Martin. 2017. Colloquialization in journalistic writing: The case of inserts with a focus on Well. Journal of Historical Pragmatics 18(1): 104–135. Sinclair, John. 2005. Corpus and text – Basic principles. In Developing Linguistic Corpora: A Guide to Good Practice, Martin Wynne (ed.), 1–16. Oxford: Oxbow Books. Tanner, Simon, Munoz, Trevor & Hemy Ros, Pich. 2009. Measuring mass text digitization quality and usefulness: Lessons learned from assessing the OCR accuracy of the British Library’s 19th century online newspaper archive. D-Lib Magazine 15(7–8). Zottola, Angela. 2020. Corpus linguistics and digital humanities. Intersecting paths. A case study from Twitter. América Crítica 4(2): 131–141.
Open Corpus Linguistics – or How to overcome common problems in dealing with corpus data by adopting open research practices Stefan Hartmann
Heinrich Heine University Düsseldorf
In recent years, many researchers have called attention to the fact that research results very often cannot be replicated – a phenomenon that has been called replication crisis. The replication crisis in linguistics is highly relevant to corpus-based research: Many corpus studies are not directly replicable as the data on which they are based are not readily available. Especially in English linguistics, the full versions of many widely used corpora are still behind paywalls, which means that they are not accessible to parts of the global research community, and even when parts of the data are freely accessible, this presents problems for state-of-the-art methods of data analysis. In this paper, I discuss the challenges that have led to this situation and address some possible solutions. In particular, I argue for using smaller but openly available corpora whenever possible and for adopting open research practices as far as possible even when using commercial corpora. Keywords: replicability, open research, accessibility, transparency, representativeness
1.
Introduction
In a seminal paper, Rissanen (1989) mentioned three pertinent problems of (diachronic) corpus linguistics, which he termed “the philologist’s dilemma”, “God’s truth fallacy”, and “the mystery of vanishing reliability”. The “philologist’s dilemma” refers to the – potentially overly pessimistic – idea that the availability of quantitative methods might distract from a detailed qualitative analysis of texts, which can be considered particularly important when dealing with earlier stages of a language. A similar case is made by Egbert et al. (2020), who remind https://doi.org/10.1075/scl.118.06har © 2024 John Benjamins Publishing Company
90
Stefan Hartmann
readers that “[l]inguistics is done by linguists, not by computers” (Egbert et al. 2020: 69). “God’s truth fallacy” refers to the assumption that a corpus “gives an accurate reflection of the entire reality of the language it is intended to represent” (Rissanen 1989: 10). As will be discussed in more detail in Section 2, this is closely connected to the issue of representativeness (see, e.g., Biber 1993) – even a large and carefully balanced corpus can hardly represent all facets of a language, and may contribute to perpetuating the potentially problematic construct of an idealized standard language. Finally, the “mystery of vanishing reliability” refers to the phenomenon that adding more parameter values, or annotation categories, to a database can make each individual data point less reliable. This means that the more annotation categories we add, the more likely it is for each individual datapoint to contain errors or uncertainties in its annotations. As Kytö and Rissanen (1988: 172) put it, “the higher the number of parameter categories, the fewer and, consequently, the less representative the items included in each category will be.” More than thirty years later, corpus linguists still struggle with some of the issues that Rissanen has identified. But in recent years, additional issues have emerged. Perhaps most importantly for the purposes of the present chapter, the “replication crisis” that has permeated various quantitatively oriented disciplines in recent years and decades has also had a significant impact on methodological discussions in (corpus) linguistics (Sönning & Werner 2021; also see the blog post by Larsson 2021). In this paper, I will argue that Rissanen’s problems and the issues identified in the discourse on replicability are more closely connected than one might think, and I will argue that adopting open research practices can provide a partial solution to some of the most important among the issues raised. The term “replication crisis” refers to the observation that many scientific findings have been found to be much less replicable than many believe they should be (Zwaan et al. 2018). The terms “replication” and “replicability/reproducibility” can mean different things, though; a relatively common stance is that the reproducibility refers to being able to duplicate the results of a previous study using the same data, while replicability refers to being able to duplicate the results of a previous study using new data (see, e.g., Goodman et al. 2016). In other typologies, replication is used as a cover term for a variety of approaches from exact replication (i.e., “reproduction” in the sense of reproducibility) to conceptual replication, which only aspires to be comparable to the original study regarding the theoretically relevant processes (Hüffmeyer et al. 2016; see also Schmidt 2009; Machery 2020). Regardless of the exact approach to replication, it has become clear that the lack of replicability is also a topic in linguistics. To mention only one prominent example from experimental linguistics, a recent multi-lab effort to replicate the seminal study by Glenberg and Kaschak (2002), which suggested that the comprehension of motor terms involves fairly concrete sensorimotor simulation, proved unsuccessful in all 18 cases (Morey et al. 2021). There are good reasons to assume
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
that the issue of non-replicability may extend to many corpus-linguistic studies (see, e.g., Winter & Grice 2021: 1264–1268), especially given that many corpuslinguistic studies include a wide array of different variables that are often closely intertwined. Consider, for example, a variable like animacy, which has been found to influence the grammars of the world’s languages in intricate ways (see, e.g., Yamamoto 1999) and which is frequently operationalized for corpus-linguistic studies (e.g., when investigating the variation between of-genitive and ’s-genitive, see Egbert et al. 2020). But it can be operationalized in many different ways, from fine-grained coding schemes like the one proposed by Zaenen et al. (2004) to simple binary or ternary schemes such as “human/animal/inanimate”. Such differences in the way the same variable can be operationalized makes conceptual replication particularly valuable for corpus linguistics. Apart from offering the possibility to compare different results, replication studies can also enable us to compare different operationalizations of the same variable(s) (see, e.g., Omidian et al. 2021; Vandeweerd et al. 2021). Sönning and Werner (2021: 1182) mention the following list of problems that have been identified as potential causes of non-replicability: – – – –
a lack of transparency in methodology and data analysis, the non-reproducibility of scholarly work, as, for example, original data and analysis procedures are not accessible, reluctance to undertake replication studies as purportedly “unoriginal” (and unprestigious) despite their potential to put previous findings in perspective, and concerns about high rates of false-positive findings in the published scientific literature.
At first glance, it might seem quite far-fetched to link Rissanen’s problems to the issues related to the “replication crisis”. In this paper, however, I will argue that there are important connections between the different issues mentioned above. And more importantly, I will argue that the measures that have been proposed to help overcome the replication crisis can also solve Rissanen’s problems – at least partly. Specifically, I will make a case for what I call Open Corpus Linguistics. This entails putting into practice principles of open research at various levels and at various (ideally, all) stages of the research process. In a best-case scenario, it involves the open availability of the entire corpus the researcher draws on, as well as sharing of concordances, annotations, and analysis scripts (if applicable). This also helps other researchers to put one’s findings into perspective, which may be seen as the main overarching issue underlying Rissanen’s problems. The remainder of this contribution is structured as follows: In Section 2, I explicate Rissanen’s problems in more detail, relating each of them to specific
91
92
Stefan Hartmann
issues raised in the replication debate. In Section 3, I discuss the main principles of Open Corpus Linguistics, taking potential challenges and pitfalls into account. Section 4 concludes the paper by bringing the two strands of the discussion together by showing how Open Corpus Linguistics can contribute to overcome several widely discussed problems in corpus linguistics. While my focus in this paper is on English corpora, many considerations brought forward here of course apply to corpora of all languages.
2.
Revisiting Rissanen’s problems
Rissanen (1989) discusses a number of challenges that every corpus linguist faces at one point or another. His problems relate to the opportunities and limitations of quantitative and qualitative approaches, the notorious question of representativeness and to basic challenges of statistical sampling in general. Not all of the problems are specific to corpus linguistics to the same degree, and not all of the problems are necessarily problematic for all corpus-linguistic approaches. For example, whether or not the full text should be taken into account depends on the research question. While it is true that ignoring the broader context can be dangerous in some cases, many research questions do not require us to take the full text into account. For example, when investigating alternation phenomena such as the dative alternation (I gave the book to her vs. I gave her the book; see, e.g., Goldberg 1995), it is definitely advisable to take the (narrower or wider) context of the attestation into account, but reading each text in which the construction occurs in full would entail a huge amount of additional effort that would hardly be justifiable in light of the research question. For other research questions, it is of course indispensable to take the full text into account. For example, a qualitative analysis of a text’s structure obviously only makes sense if the analyst has access to the full text. Consider, for instance, van Dijk’s (2005: 90–95) analysis of one New York Times article, in which he investigates which previous knowledge is assumed on the reader’s part, and which information is explicitly given: Such an analysis only really makes sense if the entire text can be taken into account. But apart from such qualitative analyses, some quantitative methods, such as collostructional analysis (Stefanowitsch & Gries 2003) also require the availability of full texts, or at least word list data derived from them. As such, it is definitely a desideratum that full corpus texts should be available, if at all possible. Thus, the actual problem captured by “the philologist’s dilemma” is that full texts are often not readily available, for example, for copyright reasons. This can lead to a lack of transparency in methodology and data analysis, which in turn can entail nonreproducibility.
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
“God’s truth fallacy” refers to the problem of representativeness. The notion of representativeness refers to the goal that a corpus should be a representative sample of a particular language variety (Baker et al. 2006: 139). This makes it necessary to determine the limits of the population that is being studied as clearly as possible (Biber 1993). As Hunston (2008: 161) points out, however, a corpus can only ever be representative in relation to a limited number of previously selected categories. This is particularly true for reference corpora, which aim to be representative of the language spoken at a particular point in time, as “it is not possible to identify a complete list of ‘categories’ that would exhaustively account for all the texts produced in a given language. No list of domains, or genres, or social groupings can ever be complete” (Hunston 2008: 161). The problem of representativeness can, to a certain extent, be resolved by tailoring one’s corpus to the specific research question at hand. When compiling reference corpora or other corpora that are not being compiled for answering one particular research question, a feasible strategy can be to “do the best that is possible in the circumstances and to be transparent about how the corpus has been designed and what is in it. This allows the degree of representativeness to be assessed by the corpus user” (Hunston 2008: 162). The German Reference Corpus, for instance, can be considered an almost prototypical example for such an approach: It is organized into “archives”, and from these archives, users can compile corpora on their own (Kupietz et al. 2010). This of course only works if sufficient amounts of data are available. More importantly, a good corpus documentation is a prerequisite for such an approach. The more metadata the users have available, the more flexible they are in compiling corpora that are balanced for specific categories that are relevant to their research questions. Thus, the actual problem is not so much that researchers tend to overestimate the representativeness of their corpus data – quite to the contrary, (most) corpus linguists tend to reflect the limits of the representativeness of their data quite thoroughly. Instead, the more relevant problem seems to be that most corpora do not offer the possibility of re-sampling/re-compilation. This may be due to several reasons. On the conceptual level, a corpus may have been compiled with a particular research question in mind, balanced for a number of categories. In this case, the corpora in question are often relatively small, and as such, it would not necessarily make sense to make even smaller samples, even if it were technically possible. On the more technical side, which is more relevant for the purposes of the present paper, even very large corpora often do not offer the possibility of creating custom subcorpora. In many cases, this is directly connected to the issues discussed above in connection to the “philologist’s dilemma”. A number of corpora can only be accessed via online interfaces that allow for working with key-word-in-context (KWIC) concordances but do not allow for accessing or exporting larger contexts, let alone full texts. Again, this can be an obstacle to
93
94
Stefan Hartmann
research transparency and reproducibility. As an example, consider the Englishcorpora.org suite compiled by Mark Davies. Comprising such widely used corpora as the Corpus of Contemporary American English (COCA) and the Corpus of Historical American English (COHA), these resources have played and continue to play an important role in English linguistics. But the default interface only allows for retrieving a limited number of results in the form of KWIC concordances or frequency lists. For all other uses, one has to purchase a commercial license. While this is understandable from an economic point of view, as developing and maintaining corpora in a sustainable way requires considerable financial resources, it is highly problematic from the perspective of research ethics. Among other things, it creates barriers to scholars in less affluent parts of the world (even if partial waivers are granted), but also to independent scholars without the financial backing of a university or another research institution. In many cases, it can therefore make sense to look for alternative corpora that can be considered equally representative for the language the researcher wants to investigate, or even to compile one’s own corpus – possibly by drawing on existing (open) corpora and using relevant subcorpora of each corpus. This can also be advantageous with regard to the “mystery of vanishing reliability”, that is, the phenomenon that each datapoint tends to become less reliable the more parameters (in corpus-linguistic terms, annotations) one adds. We can think about this in terms of a simple spreadsheet: The number of data points (rows) remains the same, the number of columns, however, increases. The problem now is of course not that more parameters are added but rather that the number of “cells” (in relation to the number of datapoints) increases and, as such, the potential for error. The obvious solution, then, is not to reduce the number of columns1 but, ideally, to increase the number of datapoints so that the individual errors weigh in less. As such, the “mystery of vanishing reliability” can be reframed in terms of a lack of extensibility: Being in control over the compilation of a corpus allows us to easily extend the database if necessary. The problem, after all, is not so much that each data point becomes less reliable if we add more annotation categories, but rather that we often do not have enough data points to obtain a truly informative picture when addressing research questions that require us to take many categories into account simultaneously. But there is another dimension to extensibility: The reliability of a particular annotation can also “vanish” because it turns out to be misguided, for whatever reason. For example, it could turn out that an annotation set used for a corpus is based on false assumptions. Thus, it is tremendously helpful if a corpus is extensible, that is, existing annotations can be amended or improved 1. Unless, of course, there is no sufficient conceptual motivation for including a specific parameter in the analysis.
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
and new ones can be added, both by the original creators and by people who reuse the data. As Garellek et al. (2020: 6) put it (albeit in a different context, discussing phonetic datasets), “access to data allows for cumulative progress”. This is why, ideally, corpus-linguistic resources should be available openly as FAIR data – a concept to which we will return in the next section, in which some obstacles on the way to this goal will also be addressed.
3.
Open Corpus Linguistics: Perspectives and challenges
In the previous section, I argued that open research practices can provide (partial) solutions to common corpus-linguistic problems. This raises the question of how exactly these open research principles can and should be put into practice, and which challenges this entails. There exist several standards and guidelines that can provide orientation. After all, the problems discussed here are not specifically corpus-linguistic ones. In terms of open data, the FAIR guiding principles (Wilkinson et al. 2016) can be considered the de-facto standard that is also required by a number of funding agencies. FAIR stands for “Findable, Accessible, Inter-operable, Reusable”. “Findability” means that it should be easily possible to retrieve the dataset(s). To this end, data should be described extensively with the help of rich metadata, and (meta)data should be assigned a persistent identifier, such as a Digital Object Identifier (DOI) (Wilkinson et al. 2016). “Accessibility” means that data should be retrievable using an open, free, and universally implementable communications protocol, protected by an authentication and authorization procedure where necessary (Wilkinson et al. 2016). File formats like the Extensible Markup Format (XML), spreadsheets in the form of comma- or tab-separated values (CSV, TSV ), or many pertinent file formats of widely-used annotation tools (e.g., CoNLLU) fulfil this criterion, as they can be opened with any text editor and do not require any specific software, and especially no commercial tools. Other file formats, for example, the native file format of the annotation software MAXQDA, are more problematic as working with them requires commercial software. The Microsoft Office file formats are somewhere in-between: they are open file formats that can also be used with other, non-commercial software. As such, there are no strong arguments against using, for example, Excel spreadsheets, but also no strong arguments against using CSV files instead. This issue is closely connected to the “I” in FAIR, which stands for “Interoperability”: “(meta)data use a formal, accessible, shared, and broadly applicable language for knowledge representation” (Wilkinson et al. 2016). This allows for integrating data and tools with minimal effort (Wilkinson et al. 2016). For example, consider a situation in which
95
96
Stefan Hartmann
data are drawn from multiple different corpora and the researcher wants to make use of part-of-speech tags: If all corpora have been tagged using the same tag set, the researcher will have a much easier time working with the full dataset. Finally, “Reusability” means that other researchers should be able to draw on and build upon the existing work. An important prerequisite for this is that “(meta)data are released with a clear and accessible data usage license” (Wilkinson et al. 2016). An explicit license, such as one of the widely used Creative Commons licenses (see https://creativecommons.org, last checked 29/10/2022), allows the researcher to clearly state what others are and are not allowed to do with the published data. The current gold standard is the CC-BY 4.0 license, which allows for virtually unlimited re-use of the data, as long as the creator is credited appropriately. However, there are also more restrictive licenses. If, for example, a researcher wants to make a dataset available for free to the scientific community but retain the possibility to receive some financial compensation if the data are used commercially, a CC-BY-NC-SA license can be used, which allows the re-use of the data only for non-commercial purposes (NC), and only if the resulting dataset is shared under the same or a comparable license (Share Alike, SA). In an ideal world, then, all linguistic corpora would be available for free in re-usable and interoperable formats under a Creative Commons license. In practice, however, there are some obstacles. One obvious problem is that corpora are usually themselves derivative works in the broadest sense, that is, they draw on existing material. And in the default case, the existing material is subject to copyright. In the case of, say, newspaper texts, the copyright holders are usually easy to find but hard to convince to make their content available for free; in the case of web data, by contrast, the copyright situation is often unclear, which can make the redistribution of data crawled from the web problematic (Schäfer & Bildhauer 2013: 4). To be on the safe side, institutions that compile reference corpora often purchase licenses from publishers. In many cases, this makes open distribution of the data impossible. Many corpora can therefore only be queried via web interfaces, often with relatively severe restrictions regarding export options, or even with a limited number of results. Examples include the above-mentioned Englishcorpora.org suite and most corpora available via SketchEngine (see https://www .sketchengine.eu/fair-use-policy/). While this is not necessarily a problem for more qualitatively oriented approaches that do not require full texts, it severely limits the number of potential quantitative approaches that can be used to work with the data. Some query systems, such as Sketch Engine, come with a variety of inbuilt analysis methods, which is generally a good thing but entails the disadvantages that if a corpus is only available via this system, the user’s choice of quantitative methods is largely limited to those offered by the software in question. This is why some corpora aim at obtaining the copyright holders’ permission to distribute full texts, which can sometimes be successful, as in the case of Schneider’s
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
(2020) German song text corpus (Songkorpus), where the annotated full texts are accessible to registered users for research purposes. Other compilers try to circumvent copyright problems by exploiting legal grey zones. For instance, Schäfer and Bildhauer’s (2012) Corpora from the Web (COW) only distribute sentence shuffles from the websites that have been crawled. The data made available via COW can therefore be regarded as quotes from the texts. As a further legal safeguard, COW has a rather restrictive access policy to make sure that they are only used for academic purposes. Mark Davies’ English-corpora.org suite pursues a very different approach: As mentioned above, the full data have to be purchased; but the full data are not entirely complete, as every twentieth word is omitted for copyright reasons. This is of course quite problematic, especially as the redacted words can differ for different users of the corpus, which limits the comparability of different studies. An even more radical solution that can still prove viable for many computationallinguistic approaches is corpus masking: Here, each word type is replaced by a specific random string (Rehm et al. 2007). Apart from copyright, personality rights can of course also be an issue. In the case of spoken corpora or child language corpora, for example, we are usually dealing with elicited data, requiring the participants’ (or their parents’) informed consent. Especially in the case of child language data, the recordings can contain sensitive information such as the child’s or the parents’ name or the place where they live. Thus, it is important to anonymize or pseudonymize the data. But data crawled from the web can also contain sensitive information. Given that the information was public at the time of crawling, one could make a case that including it in the corpus is unproblematic, but it is quite easy to imagine scenarios in which the publication of data crawled from the web can lead to legally or ethically challenging situations. In some cases, it can therefore be useful to publish a corpus in password-protected form, even though generally, the ideal should of course be maximal accessibility. How should we ideally approach the copyright problem now if we want to follow the principles of Open Corpus Linguistics? If we do not need full texts to address our research questions, the problem is quite negligible. In that case, we can work with concordances, and to the best of my knowledge, nothing usually speaks against sharing concordances via dedicated repositories like OSF 〈https://osf.io〉 (28 October 2022), as concordances can be considered collections of quotes.2 In the US, this is probably covered by the Fair Use Doctrine, which 2. It should be added, however, that this is a layperson’s perspective (hopefully an informed one, though). The legal perspective is certainly much more complex (see Collister 2022 for an insightful treatment focusing on the situation in the US), and of course the legal situation will differ between different legislations. My argument is of course less convincing if we talk about
97
98
Stefan Hartmann
allows for using a limited amount of copyright-protected material without permission for purposes such as comment, criticism, and research, although there is also some uncertainty as to which usage situations fall under “fair use” (see Lehmberg et al. 2008: 65f.).3 Under other legislations, for example, the German one, collections of quotes can be expected to be covered by laws on citation. If we need full texts, my suggestion is that we should always consider using an openly available corpus like the BNC or the Open American National Corpus first. In a broader sense, the aforementioned COW can be considered open corpora, too – for legal reasons, they are released under a relatively restrictive license, but they are freely available for academic purposes. Naturally, there will be situations in which we cannot use such open corpora because we need more or different data. In such cases, we might have to either fall back on commercial corpora or compile our own corpus. Thanks to the availability of powerful programming languages such as R or Python, this is easier than ever before; even novice users can quite easily get familiar with a tool like Barbaresi’s (2021) trafilatura, which takes as its input a list of URLs that are then crawled. What is more, the tool can also automatically clean the data and add metadata. Also, thanks to relatively permissive new legislation at least in (parts of ) the European Union,4 many data mining activities that used to take place in a legal grey zone are explicitly legal now (see, e.g., Gärtner et al. 2021, who also offer a discussion of remaining open questions). While this does not include free sharing of the data, it does allow for data to be permanently stored on servers of university libraries. At least in theory, this allows for reconciling best-practice strategies of reproducible research with copyright and other legal restrictions. In practice, however, it is still an open question how the long-term storage of mined datasets and especially the process of granting access to peers can be organized.
very large concordances that, theoretically, would allow for reconstructing the original texts. Some corpus providers therefore limit the number of concordance lines that can be shared in their terms of use (e.g., 10,000 lines in the case of COW). 3. On Fair Use and the copyright situation in the US, also see Wilkinson et al. (2005) and Lewis et al. (2006). Thanks to a reviewer for pointing out these papers to me. 4. In particular, articles 3 and 4 of the European Directive on the Digital Single Market 〈https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32019L0790&from =EN#d1e1321-92-1〉 (28 June 2022) stipulates certain exceptions from copyright for text and data mining purposes. However, as Rosati (2021) points out, by far not all EU member states have transposed the directive into national legislation.
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
4. Conclusion: Open Corpus Linguistics in practice In the preceding sections, I argued for adopting open research practices in corpus linguistics and examined a number of potential problems that such an endeavor entails. In this section, I discuss how Open Corpus Linguistics can work in practice, and how it contributes to overcoming the pertinent problems addressed in Section 2. In Section 1, I argued that adopting open research practices – and in the ideal case, using openly available corpora – helps us to overcome “Rissanen’s problems”: In the best-case scenario, we can use corpora whose full texts are readily available, which can contribute to overcoming “the philologist’s dilemma”. Such a scenario also provides some flexibility in working with pre-compiled corpora, as we are not at the mercy of the corpus creators with regard to the composition of the corpus. Instead, we can work with custom subcorpora or work with custom compilations of subcorpora from different corpora, which can help to solve the problem of “God’s truth fallacy”. And finally, such an approach ensures replicability and reproducibility, which partly solves the “mystery of vanishing reliability”. The latter is even true if we work with commercial corpora but make our concordances and analysis scripts available, as is increasingly common in corpus linguistics. For this purpose, platforms like the Open Science Framework (OSF) or the Tromsø Repository of Language and Linguistics (TroLLing) can be used (see the FAQ in Table 1 for more information). In the previous section, I addressed some potential problems, many of which are, in my view, not insurmountable. Some of them, though, call for creative solutions such as Schäfer and Bildhauer’s decision to work with sentence shuffles in the case of COW. Needless to say, it would be desirable to have more legal certainty when using copyright-protected content for linguistic purposes – the aforementioned EU directive might be seen as a promising sign that (some) political stakeholders are aware of the need to reconcile research and copyright interests. For the time being, following the ideals of Open Corpus Linguistics might in some cases require entering legal grey zones. The risk of having to tread uncertain legal ground can, however, be minimized by keeping the question of how the data should be published in mind from the earliest design phase on, and by choosing open corpora wherever possible, as discussed above. In other words, following the principles of Open Corpus Linguistics requires us to invest considerable time in Research Data Management (RDM). While this can be time-consuming, it is very likely that it will save others and ourselves much time later on. One topic that I have not addressed yet is open-access publication. As one reviewer correctly points out, open research practices and open-access publication ideally go in tandem, even though they can be treated as separate topics.
99
100
Stefan Hartmann
Many of the arguments in favor of open research practices mentioned above also apply to open-access publication, ideally in the form of “gold open access” (i.e., the final publication is available free of charge) or alternatively in the form of “green open access” (i.e., a preprint is published on a pertinent repository; see Eve 2014 for more details and discussion). While linguistic research questions and the corpus-linguistic scenarios required to address them are too diverse to provide anything like a “cookbook” for Open Corpus Linguistics, this chapter has hopefully provided some helpful guidelines, some of which are summarized in the form of Frequently Asked Questions in Table 1. Table 1. Answers to some frequently asked questions about open research practices in corpus linguistics Category
Questions and answers
Existing corpora
Which corpora should I use to follow the ideal of Open Corpus Linguistics? The choice of corpus has to be guided, first and foremost, by the research question. But in many cases, there are open alternatives to the widely-used default choices. Examples for synchronic English data include the Open American National Corpus for spoken and written American English, as well as the ENCOW corpus for World Englishes as used on the web. The BNC, which is a paradigm example of an open corpus, hardly needs to be mentioned as it is already widely used. There are also a number of multilingual corpora, for example, the COW family of corpora to which the above-mentioned ENCOW belongs, or the WaCky corpora (Baroni et al. 2009).
Corpus compilation
Which principles should I follow when compiling new corpora? If possible, try to create a corpus that can be published freely under an open license. To do so, it is very important to address legal questions at the very beginning of a project. If you cannot make the full texts freely available for copyright reasons, try to make the corpus as accessible as possible, for example, by allowing queries via flexible search engines such as NoSketchEngine or CQPweb and by publishing word and lemma lists (and ideally, n-gram lists), or by publishing it in password-protected form via a repository that allows for closed-access corpora (e.g., CLARIN). Data repositories that fit your needs can be found via 〈https://www.re3data.org/〉 (2 July 2022).
Repositories
Where can I publish my research data? There are dedicated repositories such as osf.io, zenodo.org, the TroLLing Dataverse (https://dataverse.no/dataverse/trolling) and others. I strongly recommend publishing data there, rather than on one’s own website. Even on repositories, research data may not be available forever, but we can at least be confident that they will remain available for a few decades after the researcher has retired.
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
Table 1. (continued) Category
Questions and answers Can I publish my paper (draft) along with my data? In most cases, this shouldn’t be a problem, even if you submit the paper to a commercial journal. The Sherpa-Romeo database gives a good overview of different journals’ and publishers’ open-access policies 〈https://v2.sherpa.ac.uk/romeo/〉 (25 October 2022). As a rule of thumb, you are almost always on the safe side when making your initial, not yet peer-reviewed manuscript available on your website or on a non-commercial repository. One of the go-to non-commercial preprint repositories for linguists is PsyArXiv 〈https://psyarxiv.com/〉, which also offers the convenient opportunity to link your preprint with an OSF data repository. I’m working with concordances from a commercial corpus. Am I allowed to publish them in a repository? It depends. In some cases, the terms of use of the corpus will prohibit this. In the case of smaller concordances with small context windows, I would always argue that they are collections of quotes – and quoting should in most cases be unproblematic. In case of doubt, it might be worthwhile to talk to your university’s legal department, rather than relying on my layperson’s opinion. Also, in some cases, the original data are not strictly necessary to ensure reproducibility and replicability. For example, when working with distributional-semantic methods, we only need to know the degree of co-occurrence of the target with other words, regardless of what those other words are. Instead of the original data, corpus compilers can choose to distribute “masked” data in such cases, where each word type is replaced by a specific random string (Rehm et al. 2007). Perek (2021) makes use of this fact in the published version of his distributional semantic models of English verbs and nouns based on COHA. Another option for web-based corpora is to publish link lists; in the case of Twitter data, a frequently-used possibility is to publish Tweet IDs (McCreadie et al. 2012). To what should I pay attention when publishing my research data in repositories? Make sure that everything is well-documented and self-explanatory. (This is harder than it sounds, which is why I’m not referring to any of my own repositories here as a best-practice example.) Make sure that the repository is actually public when your paper is published (OSF, for instance, offers private and public repositories). Make sure that your analysis scripts are extensively commented (formats like R Markdown or Jupyter Notebooks invite extensive comments, but plain text scripts can also be used, of course). Make use of the possibility to assign a DOI to the dataset(s). I would like to share my dataset and analysis scripts with reviewers. How can I do this without compromising the anonymity of peer-review? OSF offers “view-only” links that you can share with reviewers. Note that the nonanonymous repository can easily be retrieved from the view-only link as soon as it is public – if you want to make sure that you remain anonymous, keep it private (the view-only link will still work).
101
102
Stefan Hartmann
Table 1. (continued) Category
Questions and answers
Potential objections
If I make my data available, won’t I lose control over it, and will I not take the risk that other people will use it to do the research that I was still planning to do? This is an objection that I hear very often. And in a way, it is understandable – after all, if you have invested a lot of time and work into a dataset, releasing it might feel a bit like giving away your child. But apart from the fact that you can’t lock up your child at home in a secret chamber forever (or at least, you shouldn’t), the argument is not really convincing: Firstly, if you follow the recommendation of publishing your data with a license, you don’t lose control over it. On the contrary, you can specify quite precisely who can do what with the data. Secondly, the likelihood that someone will address exactly the same research question with exactly the same methods on the basis of your data is vanishingly low. The worst thing that can happen is that someone conducts an eerily similar study, which, on second thought, is a good thing, because you could see it as a replication by accident, and replication is good for science. I’m thinking about publishing my corpus, but I feel it’s not good enough for publication and I don’t have the resources to improve it. * This is probably quite common – not only in corpus linguistics but also in other domains, for example, when it comes to publishing programming scripts. But the thing is: Nobody expects us to be perfect. Everybody who has ever worked in corpus linguistics knows that no corpus will ever be perfect, and that the quality of a corpus depends less on the competence of the researchers involved than on the resources they were able to put into it. As such, nobody will blame you for releasing a corpus that is still more of a raw diamond. Publishing your “raw diamond” will give other people the opportunity to build on it, or to work with it while you are still continuing to develop it.
* Thanks to a reviewer for bringing this up!
To sum up, Open Corpus Linguistics can be a challenging endeavor, but given the “replication crisis”, it is a necessary one. In the long term, adopting open research practices can also help us to focus on actual linguistic research questions, rather than spending hours and hours of work on things that other people have done before, without making the results publicly available. Adopting open research practices is ethically a good choice, and it is in our own best interest, both as individual researchers and as an empirical discipline.
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
Acknowledgements I am grateful to two anonymous reviewers for providing excellent feedback on a previous version of this paper and for pointing out a number of relevant papers to me that I had overlooked.
References Baker, Paul, Hardie, Andrew & McEnery, Tony. 2006. A Glossary of Corpus Linguistics. Edinburgh: EUP. Barbaresi, Adrien. 2021. Trafilatura: A web scraping library and command-line tool for text discovery and extraction. In Proceedings of the Annual Meeting of the ACL, System Demonstrations. (25 October 2022). Baroni, Marco, Bernardini, Silvia, Ferraresi, Adriano & Zanchetta, Eros. 2009. The WaCky Wide Web: A collection of very large linguistically processed web-crawled corpora. Language Resources and Evaluation 43(3): 209–226. Biber, Douglas. 1993. Representativeness in corpus design. Literary and Linguistic Computing 8: 243–257. Collister, Lauren B. 2022. Copyright and sharing linguistic data. In The Open Handbook of Linguistic Data Management, Andrea L. Berez-Kroeker, Bradley McDonnell, Eve Koller & Lauren B. Collister (eds), 117–128. Cambridge MA: The MIT Press. Dijk, Teun A. van. 2005. Contextual knowledge management in discourse prediction: A CDA perspective. In A New Agenda in (Critical) Discourse Analysis: Theory, Methodology and Interdisciplinarity [Discourse Approaches to Politics, Society and Culture 13], Ruth Wodak & Paul A. Chilton (eds), 71–100. Amsterdam: John Benjamins. Egbert, Jesse, Larsson, Tove & Biber, Douglas. 2020. Doing Linguistics with a Corpus: Methodological Considerations for the Everyday User [Elements in Corpus Linguistics]. Cambridge: CUP. Eve, Martin Paul. 2014. Open Access and the Humanities: Con;//doi.org/texts, Controversies and the Future. Cambridge: CUP. Garellek, Marc, Simpson, Adrian, Roettger, Timo B., Recasens, Daniel, Niebuhr, Oliver, Mooshammer, Christine, Michaud, Alexis et al. 2020. Letter to the editor: Toward open data policies in phonetics: What we can gain and how we can avoid pitfalls. Journal of Speech Sciences 9: 3–16. Gärtner, Markus, Kleinkopf, Felicitas, Andresen, Melanie & Hermann, Sibylle. 2021. Corpus reusability and copyright – Challenges and opportunities. In Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-9) 2021. Limerick, 12 July 2021 (Online-Event), Harald Lüngen, Marc Kupietz, Piotr Bański, Adrien Barbaresi, Simon Clematide & Ines Pisetta (eds), 10–19. Leibniz-Institut für Deutsche Sprache. (30 October, 2022). Glenberg, Arthur M. & Kaschak, Michael P. 2002. Grounding language in action. Psychonomic Bulletin & Review 9(3): 558–565. Goldberg, Adele E. 1995. Constructions: A Construction Grammar Approach to Argument Structure. Chicago IL: The University of Chicago Press.
103
104
Stefan Hartmann
Goodman, Steven N., Fanelli, Daniele & Ioannidis, John P. A. 2016. What does research reproducibility mean? Science Translational Medicine 8(341): 341ps12. Hüffmeier, Joachim, Mazei, Jens & Schultze, Thomas. 2016. Reconceptualizing replication as a sequence of different studies: A replication typology. Journal of Experimental Social Psychology 66: 81–92. Hunston, Susan. 2008. Collection strategies and design decisions. In Corpus Linguistics: An International Handbook [HSK 29.1], Anke Lüdeling & Merja Kytö (eds), 154–168. Berlin: Walter de Gruyter. Kupietz, Marc, Belica, Cyril, Keibel, Holger & Witt, Andreas. 2010. The German Reference Corpus DeReKo: A primordial sample for linguistic research. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC 2010), Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner & Daniel Tapias (eds), 1848–1854. Valletta: European Language Resources Association. http://www.lrec-conf.org/proceedings/lrec2010/pdf/414 _Paper.pdf (19 May 2024). Kytö, Merja & Rissanen, Matti. 1988. The Helsinki Corpus of English Texts: Classifying and coding the diachronic part. In Corpus Linguistics, Hard and Soft: Proceedings of the Eighth International Conference on English Language Research on Computerized Corpora, Merja Kytö, Ossi Ihalainen & Matti Rissanen (eds), 169–179. Amsterdam: Rodopi. Larsson, Tove. 2021. Has ‘the replication crisis’ reached corpus linguistics? Blog Linguistics with a Corpus. (25 October 2022). Lehmberg, Timm, Rehm, Georg, Witt, Andreas & Zimmermann, Felix. 2008. Digital text collections, linguistic research data, and mashups: Notes on the legal situation. Library Trends 57: 52–71. Lewis, William D., Farrar, Scott & Langendoen, D. Terence. 2006. Linguistics in the Internet age: Tools and fair use. In Proceedings of the EMELD’06 Workshop on Digital Language Documentation: Tools and Standards: The State of the Art. Lansing, MI. (6 January 2023). Machery, Edouard. 2020. What is a replication? Philosophy of Science 87(4): 545–567. McCreadie, Richard, Soboroff, Ian, Lin, Jimmy, Macdonald, Craig, Ounis, Iadh & McCullough, Dean. 2012. On building a reusable Twitter corpus. In Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval – SIGIR ’12, 1113. Portland OR: ACM Press. Morey, Richard D. et al. 2021. A pre-registered, multi-lab non-replication of the actionsentence compatibility effect (ACE). Psychonomic Bulletin & Review 29: 613–626. Omidian, Taha, Balance, Oliver James & Siyanova-Chanturia, Anna. 2021. Replicating corpusbased research in English for academic purposes: Proposed replication of Cortes (2013) and Biber and Gray (2010). Language Teaching, 1–9. Perek, Florent. 2021. Distributional semantic models for English verbs and nouns. Open Science Framework. Rehm, Georg, Witt, Andreas, Zinsmeister, Heike & Dellert, Johannes. 2007. Corpus masking: Legally bypassing licensing restrictions for the free distribution of text collections. In Digital Humanities 2007, 2nd edn, Sara Schmidt, Ray Siemens, Amit Kumar & John Unsworth (eds), 166–170. Urbana-Champaign IL: University of Illinois. (19 May 2024).
Open Corpus Linguistics – How to overcome common problems in dealing with corpus data
Rissanen, Matti. 1989. Three problems connected with the use of diachronic corpora. ICAME Journal 13: 16–19. Rosati, Eleonora. 2021. The DSM Directive two years on: Do things ever get easier? IIC – International Review of Intellectual Property and Competition Law 52(9): 1139–1142. Schäfer, Roland & Bildhauer, Felix. 2012. Building large corpora from the web using a new efficient tool chain. In Proceedings of LREC 2012, Nicoletta Calzolari, Khalid Choukri, Terry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk & Stelios Piperidis (eds), 486–493. (25 October 2022). Schäfer, Roland & Bildhauer, Felix. 2013. Web Corpus Construction. San Rafael CA: Morgan & Claypool. Schmidt, Stefan. 2009. Shall we really do it again? The powerful concept of replication is neglected in the social sciences. Review of General Psychology 13(2): 90–100. Schneider, Roman. 2020. A corpus linguistic perspective on contemporary German pop lyrics with the multi-layer annotated “Songkorpus”. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk & Stelios Piperidis (eds), 842–848. Marseille: European Language Resources Association. (19 May 2024). Sönning, Lukas & Werner, Valentin. 2021. The replication crisis, scientific revolutions, and linguistics. Linguistics 59(5): 1179–1206. Stefanowitsch, Anatol & Gries, Stefan T. 2003. Collostructions: Investigating the interaction of words and constructions. International Journal of Corpus Linguistics 8(2): 209–243. Vandeweerd, Nathan, Housen, Alex & Paquot, Magali. 2021. Applying phraseological complexity measures to L2 French: A partial replication study. International Journal of Learner Corpus Research 7(2): 197–229. Wilkinson, Mark D. et al. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3(1): 160018. Winter, Bodo & Grice, Martine. 2021. Independence and generalizability in linguistics. Linguistics 59(5): 1251–1277. Yamamoto, Mutsumi. 1999. Animacy and Reference. A Cognitive Approach to Corpus Linguistics [Studies in Language Companion Series 46]. Amsterdam: John Benjamins. Zaenen, Annie, Carletta, Jean, Garretson, Gregory, Bresnan, Joan, Koontz-Garboden, Andrew, Nikitina, Tatiana, O’Connor, M. Catherine & Wasow, Tom. 2004. Animacy encoding in English: Why and how. In DiscAnnotation ’04, Bonnie Webber & Donna Byron (eds), 118–125. Stroudsburg PA: Association for Computational Linguistics. Zwaan, Rolf A., Etz, Alexander, Lucas, Richard E. & Donnellan, M. Brent. 2018. Making replication mainstream. Behavioral and Brain Sciences 41, E120.
105
Text length and short texts An overview of the problem Aatu Liimatta
University of Helsinki
Variation in text length is an unavoidable confounder in quantitative textanalytic corpus-linguistic studies. Texts can be difficult to compare across text lengths, particularly if many of them are short, due to the difficulty of calculating meaningful frequencies for the lexical items and linguistic features of interest. Traditionally, this has been less of an issue, since texts in many of the genres typically studied in linguistics have been relatively long. However, the rise of social media has brought the issue to the forefront. In this chapter, I describe the problem of text length and short texts together with a number of solutions and workarounds to this and related problems. Keywords: text length, normalization, lexical diversity, lengthwise analysis
1.
Introduction
It is a well-known fact that variation in text length is an unavoidable source of nonuniformity in corpora and can cause issues in quantitative corpus-linguistic analyses. At the very basic level, the confounding effect of variation in text length is obvious. Since a longer text by definition contains more words, there are also more opportunities for any given linguistic item or feature to appear. In other words, a longer text will, on average, contain more instances of any item or feature simply because it is longer. This is a problem particularly for text-analytic corpuslinguistic studies, which are interested in comparing how many of these items appear in different types of texts: if two texts have a different number of occurrences of the feature of interest simply because they are of different lengths, it can be very difficult to compare texts of different lengths with each other.1 1. Variationist corpus linguistics, which focuses on the proportions of variant items or constructions, is not as heavily affected by the issue. However, even variationist analyses may be affected by the distribution of text lengths in their dataset. https://doi.org/10.1075/scl.118.07lii © 2024 John Benjamins Publishing Company
Text length and short texts
Fortunately, this basic problem has a simple mathematical solution, which commonly forms the basis of typical quantitative corpus-linguistic inquiries. The number of occurrences of the feature of interest in a text can be divided by the number of words2 in the text. This process, normalization, gives us the rate of occurrence of the feature, that is, the average number of instances of the feature per each word in the text. Viewed differently, this normalized value can be understood as the probability of picking an occurrence of the item of interest if a word is picked from the text at random. Commonly, this value is further multiplied by a factor of, for example, a thousand, ten thousand, or a million, depending on the overall prevalence of the feature in the dataset, in order to bring it within a more understandable numerical range. Crucially, the normalization process harmonizes the divisors of the feature counts from texts of different lengths. For instance, 10 occurrences per 2,500 words and 20 occurrences per 3,200 words both get the same divisor after normalization, that is, 4 occurrences per 1,000 words and 6.25 occurrences per 1,000 words, respectively. This makes it possible to directly compare the texts which these rates of occurrence represent. Normalization is a tried-and-true method of comparing texts of different lengths with each other. However, while it is a working solution to a huge potential problem in a large number of cases, normalization is not without problems itself. One problem with normalization becomes particularly salient when applying the method to short texts. The problem is based on the mathematical fact that the smaller the divisor becomes, the larger the result of the division will be. In other words, the fewer words there are in a text, the larger the normalized value is. This is of course the very basis on which the normalization method is built to enable comparisons of texts of different lengths. However, when the divisor becomes very small, the effect gets magnified and the result of the calculation inflates to meaningless levels. For instance, consider a short text of only five words (such as a tweet, a postcard, or a sticky note) which contains one instance of a feature, for example, a single first-person pronoun. If we calculate the normalized frequency of first-person pronouns in this text, we get as the result 200 first-person pronouns per 1,000 words. This is the mathematical solution to the formula, but the result is quite useless in terms of comparing texts of different lengths with each other. While not every five-word text contains a first-person pronoun, it is also not that unusual if one does. But the same normalized rate of occurrence value also applies to a text of 1,000 words which contains 200 first person pronouns. While both the five-word text and the 1,000-word text have the 2. Many other bases of comparison other than words can also be used. However, the focus of the present chapter is on word count specifically, since it is the most commonly used basis in most quantitative corpus-linguistic studies.
107
108
Aatu Liimatta
same rate of occurrence of first-person pronouns, surely the 1,000-word text with 200 first-person pronouns is much more unusual than the five-word text with one first-person pronoun. Clearly, these calculated rates of occurrence are not meaningful measures for the comparison of short and longer texts in linguistics. In other words, there are two related problems caused by text length. First, texts of different lengths cannot be directly compared because they have different numbers of everything simply due to their difference in length. I call this the problem of text length. Second, extremely short texts cannot be easily compared with other texts using many typical quantitative corpus-linguistic methods, such as normalization, because the normalization results become meaningless as the length of the text becomes increasingly short. I call this the problem of short texts. In order to talk about these problems, in this chapter, I use the words “long” and “short” to refer to texts which are and are not long enough for typical quantitative corpus-linguistic analysis, respectively, though the line between the two is of course fuzzy. To combat these two problems, a number of solutions and workarounds of varying sophistication have been devised. However, many of these solutions have problems of their own. Despite the ubiquity of the problem, and the oftensuboptimal nature of its solutions, the problem is not that often discussed in much depth. In this chapter, my goal is to bring more attention to the problems of text length and short texts, and to encourage the development and application of new and improved approaches to the problem. In Section 2, I describe how the problem of text length was historically less of an issue but is coming to the forefront with the rise of research into the language of social media. I also refer to and summarize results from earlier studies, which suggest that texts of all lengths are of interest. In Section 3, I cover various methods which have been used to either solve or work around the issues caused by the problem of short texts, the problem of text length, and related problems, and discuss their upsides and downsides, as well as suggest best practices and propose potential improvements to these methods. Furthermore, I will briefly discuss the related topic of the effect of text length on measures of lexical diversity, which has been studied in more detail. Finally, in Section 4, I will conclude this chapter with some final thoughts on the problems caused by text length and their solutions.
Text length and short texts
2.
Background
2.1 Text length, corpora, and social media The problem of text length and short texts is caused by a simple mathematical relationship, and as such it has been known of since the beginning of quantitative corpus linguistics. However, historically, the practical problems it has resulted in have arguably been of relatively little actual consequence. Most genres traditionally studied in quantitative corpus linguistics tend to comprise of longer texts, compared to the extremely short text lengths of up to only a few dozen words which are overwhelmingly common on, for example, social media. Because of the comparatively long text length in such genres, the normalization method works reasonably well with them. The reasons for the focus on genres with “longer” texts have been manifold. A major reason has of course been, and still is, availability of data. Researchers have to use whatever data they have access to. Historically, this has largely been published corpora compiled by teams of researchers. But the compilers of such corpora have also been working with the data they can get access to in large enough quantities. These include genres such as newspaper articles, fiction writing, academic papers, and countless others. Many of these genres have editorial guidelines or genre conventions which place certain requirements for the length of the piece of writing. Another reason for the focus on longer texts is what has been considered important and influential enough to study. The impact of, for example, newspaper articles, academic writing, casual conversations and personal letters on people, society, and language has been evident to all, and as such it is only natural that texts in such genres are of interest to anyone studying language. But these texts also tend to be long enough for reasonable quantitative corpus-linguistic analysis. It is often easier to overlook the societal and linguistic impact of genres with mainly shorter texts. For instance, personal notes, post cards, and shopping lists are also something written and read regularly, but we might not even think to consider them and other similar genres as research subjects. The question of influence and significance also comes back to the question of availability, since it is of course more difficult to collect a large and representative corpus of post cards or shopping lists than of newspaper articles. However, it is also, to an extent, a chicken-and-egg situation. With more interest in such genres, it is possible that more such data would be collected into corpora; and with more such corpora available, there might be more interest in such genres. A similar vicious circle has also developed when it comes to the development of analysis methods which would allow us to better approach genres with short
109
110
Aatu Liimatta
texts. If genres with longer texts are easier to quantify and the available methods work better with them, it is only natural to focus on such genres; and if the focus typically is on genres with longer texts, there is no particular need to develop methods and approaches which could help analyze genres with shorter texts in more detail. However, over the past decades, many of these paradigms have been gradually but firmly upended. A central catalyst for the change has been the spread of the internet. The web and other means of computer-mediated communication have become easily accessible sources of linguistic data, which has greatly facilitated the building of new corpora and datasets to match the needs of the researcher and to enable the study of entirely new kinds of genres and registers. Of course, published corpora are still being compiled to this day, and they are a valuable tool used in a wide variety of linguistic research. But the easy access to textual data online has allowed quantitative corpus linguists to cast a wider net in terms of their research topics than ever before. At the same time, the rise of web and CMC texts has brought many of the issues with text length to the forefront. In contrast to the more “traditional” genres which make up many compiled corpora, texts from many internet genres tend to be less bound by word count limits or guidelines. While many online genres have a highly variable text length and a large proportion of shorter texts, such as blog posts or Wikipedia articles, this is particularly true for computer-mediated communication and social language use on the internet, such as postings on various social media platforms. Most social media platforms, such as Facebook or Reddit, do not limit the length of their postings to a meaningful degree. Platforms which do limit posting length, such as Twitter (now X), usually limit the maximum length, confining all of their content into the range of short texts, which is mathematically difficult to work with in quantitative corpus linguistics. Few online platforms require a minimum length for postings or even recommend postings to be of a specific length, whereas such requirements are commonplace within the publishing industry and for many of the genres included in typical published corpora, such as newspaper articles or academic writing. At the same time, online data has brought even the shortest texts to the center stage, making their societal and linguistic importance much more evident in comparison with the more traditional short genres. In other words, the free nature of internet writing has brought texts with a wide variety of lengths into the corpora of many linguists, and consequently made the problem of text length and, particularly, the problem of short texts more central than ever.
Text length and short texts
2.2 The importance of text length It is clear that variation in text length, particularly the very shortest texts, causes mathematical issues in quantitative corpus-linguistic analyses. But do we actually have to care about text length? Can’t we simply ignore the problematic cases when conducting quantitative linguistic studies? Or should we work towards finding more ways to make it possible to include texts of all lengths in our analyses? Liimatta (2022a, 2022b) provides some insight into the role text length plays in linguistic variation by analyzing functional variation between comments of different lengths on the social media platform Reddit, with the very shortest comments also included in the analysis. In order to explain how these studies work around the problems caused by text length, it is worth describing the study design in some detail. To study functional variation, Liimatta (2022a, 2022b) builds on the ideas of Multi-Dimensional Analysis (MDA; see, e.g., Biber 1988; Biber & Conrad 2009; Conrad & Biber 2001) and more widely on the so-called text-linguistic framework of register analysis (cf. Biber et al. 2020). This framework is based on the idea that linguistic features are functional, and therefore tend to be used more in texts for whose situational context and communicative purpose they are better suited, and less when the situation or function of the text does not call for their use. Consequently, it is possible to measure variation in various linguistic features and use the observed variation patterns to suggest differences in the functions of texts. The features analyzed by Liimatta (2022a, 2022b) are based on the set of functional linguistic features used by, for example, Biber (1988) in studies of register variation. In order to specifically focus on the functional variation taking place across text lengths, Liimatta (2022a, 2022b) makes use of so-called lengthwise analysis. This family of analysis methods aims to enable the analysis of linguistic variation between texts of different lengths, while at the same time circumventing many of the problems caused by variation in text length when using more typical corpuslinguistic methods. Lengthwise analysis is based on a simple insight: texts of different lengths may be difficult to compare with each other due to, for instance, mathematical reasons, as explained above, but texts which are the exact same length can typically be compared trivially. For instance, it can be difficult to say just how comparable the rates of occurrence calculated from texts of different lengths are, but it is always possible to say how different two texts of the same length are in terms of, for example, the number of occurrences of some feature. A lengthwise analysis method, then, first performs a comparison between texts of the exact same length using some suitable method, and only after this intra-length comparison compares the results between texts of different lengths or inter-length. Liimatta (2022a, 2022b) makes use of a simple but powerful lengthwise method for demonstrating the role of text length within Reddit. In the first-step intra-length
111
112
Aatu Liimatta
comparison, the frequencies of various linguistic features were pooled and averaged by text length. In the second step, the averaged frequencies were plotted across text lengths in order to show the variation in the feature frequencies across text lengths. This kind of analysis requires a very large dataset, as there need to be enough texts of every length in the dataset for meaningful results. Consequently, social media is a good source of data for this method. Reddit in particular is arguably a very fruitful source of material for quantitative linguistic analyses overall, and especially for the analysis of the effects of variation in text length. First of all, Reddit enables access to large amounts of publicly available textual data. Some other social media platforms, such as Facebook, theoretically also have a lot of data available, but in practice a large portion of it is visible only to one’s friends on the platform or those who have joined any specific discussion group. In other cases, the data may be public but difficult to access in large quantities in practice. Furthermore, since Reddit is divided into topic-based subforums called subreddits, the data is naturally subdivided into subcategories by topic and by register (see, e.g., Liimatta, 2019). This allows studies to either focus on specific topic areas or draw comparisons between multiple ones. But most importantly when it comes to analyzing the effects of variation in text length, the comment length on Reddit is not limited for most practical purposes, allowing for comments of a wide range of lengths to exist, from extremely short to reasonably long. In a sense, one weakness of social media data for typical linguistic analyses becomes its strength in the analysis of variation across text lengths. The analysis conducted by Liimatta (2022a) shows clearly that texts of different lengths play different roles on Reddit. The frequencies of various functional linguistic features vary across comment lengths: many of the features occur at different rates in Reddit comments of different lengths. This variation means that on Reddit as a whole, texts of different lengths have different functions. For instance, the results show that shorter Reddit comments include more features which are indicative of a more casual, interpersonal, and less edited style, whereas longer comments include more features which are considered more informational, edited, and narrative. The differences in the comment functions are also shown to be the greatest between the very shortest comment lengths, with more gradual differences between longer comments. Liimatta (2022b) performs a similar analysis but zooms in further to focus on a number of popular subreddits to find out whether all subreddits follow similar patterns, or if the same text length can have different functions in different subreddits. In the analysis, most subreddits analyzed are shown to follow similar patterns with each other. For example, the short comments in most subreddits contain more features which are more casual and involved, and longer comments contain more informational features. Similarly, comments of all lengths appear to
Text length and short texts
be roughly equally narrative in all of the subreddits included in the analysis. However, a handful of the analyzed subreddits, which are more focused in terms of their topic in comparison with the very relative topics of most of the included subreddits, often differ greatly from both the general patterns and from each other. For instance, in the AskReddit subreddit, longer comments are much more narrative than the shorter ones, whereas for some other subreddits the opposite is the case. Figure 1 is an example of these results for past tense verb forms, which have been associated with narrative concerns (e.g., Biber 2014). Figure 1 shows the pattern mentioned above: the frequency of past tense forms in most analyzed subreddits, shown in gray, stays fairly similar throughout the length range or even has a slightly decreasing pattern, pointing towards a relatively level spread of narrative concerns throughout the length range. However, the AskReddit subreddit in particular shows a dramatic increase in past tense verb forms as the comments get longer, implying that the longer comments in this subreddit tend to be on average increasingly narrative. These results show that not only does text length play a role in linguistic variation, but that text length can be associated with functions differently within different register categories.
Figure 1. Frequency of past tense forms across comment lengths in 20 subreddits. From Liimatta (2022b: 279)
Taken together, the results of these two studies make it clear that texts of all lengths are of interest. If we ignore any of our datasets simply based on its length, because it is problematic in terms of quantitative analysis, we run the risk of excluding some of the variation within our data from our analysis. Conse-
113
114
Aatu Liimatta
quently, the development of methods and approaches for the analysis of texts of all lengths is definitely a worthy endeavor, and potentially much more necessary than has been recognized before, particularly with the rise of sources of linguistic data including social media and other computer-mediated communication.
3.
Solutions and workarounds
As the problems of text length and short texts have been recognized, a number of solutions and workarounds for it have also been devised. In this section, I will cover some of these approaches, and some related methods. Some of these approaches help solve or work around the problems caused by variation in text length, some the problem of short texts, and some can help alleviate the effects of both. Additionally, I will describe a closely related problem, that of measures of lexical diversity. While the solutions to the problems with lexical diversity measures are not directly applicable to the problem of text length, they may still provide inspiration and starting points for new ways of approaching the problem of text length. I have divided the solutions and workarounds to the problem of text length and short texts into two main categories. In the first group of approaches, the original set of texts is manipulated in some way, after which standard methods are applied. These could also be considered the more “traditional” approaches to the problem. Their advantage is that they are simpler to implement, but this means that their downsides are often greater. Conversely, the approaches in the second group make use of various statistical and/or computational methods to see the existing data in a new light. These approaches are more complicated to implement and often only work for specific kinds of analyses, but they are much more powerful in the situations for which they are well-suited. These two groups of course overlap in practice, and methods within and between the groups can even be used together.
3.1 Manipulation of the data 3.1.1 Exclusion A commonly used workaround for the problem of short texts is to simply exclude all texts shorter than some threshold from the analysis. For instance, we might simply choose to remove all texts shorter than, say, 400 words, 500 words, or 1,000 words from our dataset. If the aim is to be able to include as much of the data in our analysis as possible, this approach is at its most reasonable when there are only a small number of outliers under the chosen length limit, as the exclusion of a
Text length and short texts
handful of outliers do not affect the overall results from a good-sized corpus very much. It could even be argued that clear outliers do not even represent the varieties of interest in the corpus particularly well, and that therefore it would even be beneficial to exclude them. However, the larger the proportion of the texts in the corpus which fall under the chosen length limit, the more problematic the exclusion method becomes. Particularly when typical texts from the shorter end of the length range start to be excluded from the analysis in addition to obvious outliers, it is clear that the dataset starts to lose some of the information potentially available within the data. Of course, it is not obvious where the line should be drawn when considering whether a text of a certain length should be considered an outlier, but from the point of view of the data, the best practice would be to have the length limit as low as possible. The optimal cutoff length when using the exclusion approach would be low enough that as many texts as possible are included in the analysis, but high enough that the desired analysis is still possible to conduct reliably. For a slightly more statistically-based approach than simply choosing some round number such as 400 or 500 as the limit, it is also possible to define the cutoff point as, for example, the 1% quantile of the length distribution, or whichever percentage gives a length limit which is workable with the chosen methods and the research questions being investigated, since this helps quantify the amount of data which has been left out. However, there are datasets for which the exclusion approach is utterly unsuitable. For instance, most social media postings are very short, and therefore would need to be excluded from the analysis under any commonly used cutoff length which would allow analysis of the data using typical analysis methods. Consequently, different solutions and workarounds to the problems of short texts and variation in text length need to be used when dealing with such data. 3.1.2 Combining In situations where discarding any data is undesirable, another workaround for the problem of short texts is available. In many studies working with, for example, social media data and other genres which have a relatively large proportion of shorter texts, texts deemed too short to comfortably conduct the intended analysis on are combined to create new “texts” which are sufficiently long for the analysis. For instance, we could decide to combine texts so that each of the combined texts is over some length limit, such as 500 or 1,000 words. As with the exclusion approach, the desirable length depends on the methods being used for the analysis and the research questions being investigated. The main upside of the combining method, when compared to the exclusion method, is that no data is completely ignored: all text available for the analysis
115
116
Aatu Liimatta
is included in the analysis. However, the downside is that by combining texts together, the texts lose their individual nature. For example, if one text is highly edited in style, and another one is highly casual, combining them together results in a loss of a lot of this information, and makes the combined text look somewhat average on both counts. In this way, the combining approach to the problem of short texts may very easily blur out some of the variation in the data. On the other hand, it is also possible that texts may end up combined in such a way that the resulting dataset overstates the importance of some feature which is actually quite rare overall, for instance, if a feature is highly frequent in a small number of texts. The combining approach also easily results in a violation of the “independence assumption” inherent in various statistical procedures, including those commonly used by corpus linguists, such as Chi-square testing (Winter & Grice 2021), and even precludes the use of some more advanced statistical methods such as the calculation of dispersion measures. The issue of blurring out or overstating variation can in some situations be mitigated by the choice of the basis of combination. If the texts which are combined are chosen in an essentially random manner, as is often the case, these potential obscuring effects cannot be reduced. However, in many cases it is possible to use a more principled basis for the combining. In the simplest case, the texts are combined based on some metadata in which we are interested in our analysis. For instance, if the analysis focuses on texts written by different sociolinguistic groups, any combining of texts needs to be done by the sociolinguistic groups in question. This, however, is done out of necessity, and it does not really help to reduce the blurring of variation taking place within the groups. For example, if we are comparing personal letters and official letters, we of course need to combine the shorter texts separately within the two categories, personal letters and official letters. But even in this case there may be variation within these two categories which gets either blurred out or overstated. In order to lessen the blurring and overstating effect, it might be useful to consider combining texts which are as similar as possible in their production circumstances, as far as reasonably possible. What texts exactly are considered “similar” is however a question which depends on the dataset and research questions. As a rule of thumb, however, the highest number of matching or similar metadata field values might be a good starting point. Raising the level of analysis might be considered a special case of the principled metadata-based combining approach. For instance, instead of studying individual classified advertisements, we might consider the entire classified advertisements section a single text for the purposes of our analysis. Or instead of focusing on individual social media comments, we might choose to focus on full comment threads. The line between raising the level of analysis and the more
Text length and short texts
general approach of principled metadata-based combining becomes blurred, however, in the case of, for example, combining Twitter tweets with their replies together to form individual texts. 3.1.3 Chunking A different, slightly less-used approach to dealing with variation in text length is the opposite of combining shorter texts together: to cut longer texts into shorter pieces of (near) equal length. For instance, Hiltunen and Tyrkkö (2019) make use of this approach when studying Wikipedia articles, which are extremely variable in length, by dividing the articles into 200-word pieces for their analysis. When using this method, the fact that all texts included in the analysis are of (roughly) the same length facilitates their comparison using feature counts or rates of occurrence, since the confounding effects of variation in text length have been diminished. At the same time, all of the textual information is included in the analysis and not discarded, even if it has been cut into smaller pieces. Texts can be split up in various ways. A straightforward approach is to simply split a text into chunks of a certain number of words. However, since sentences are a basic structural unit of language, placing chunk boundaries at sentence boundaries, making sure that every chunk includes enough words, is likely to be a more desirable solution in many cases. Another solution, which keeps the structural and discourse units of a text together even more, is to divide the text into its paragraphs, or multi-paragraph chunks. In addition to the simple chunking options above, chunking can also make use of various computational methods to create chunks which are meaningful in terms of the discourse structure. For example, Biber et al. (2004) use a computational approach to divide texts into so-called Vocabulary-Based Discourse Units (VBDUs) in an analysis of the structure of various academic registers. VBDUs are a vocabulary-based approach to segmenting a text into discourse units. The methodology behind VBDU might be useful as a chunking approach in many kinds of analyses into linguistic variation. The chunking approach may or may not help with the problem of short texts. It would be difficult to meaningfully divide the longer texts into chunks of equivalent length if the shortest texts in the dataset are very short, such as on social media. On the other hand, if the shortest texts are still of reasonable length, dividing the longer texts into chunks of similar length might actually make them more easily comparable.
117
118
Aatu Liimatta
3.2 Computational and statistical approaches 3.2.1 Lengthwise analysis In order to make feature frequencies more comparable across text lengths, Liimatta (2020) proposes a family of methods called lengthwise scaling. Closely related to the lengthwise analysis described above, this family of methods is also based on the idea that it is trivial to compare texts which are the exact same length. In lengthwise scaling methods, feature counts in each text are first compared against texts of the exact same length (“intra-length comparison”) using some suitable method of comparison. Based on the results of this comparison, each text receives a new, scaled value, which is a representation of how typical the text is in terms of the range of variation seen in all texts of the exact same length. These scaled values can then be compared between text lengths like normalized frequencies would be, such as by visual exploration of graphs or by using some further statistical or computational analysis. While the idea behind lengthwise scaling can be applied in various ways, Liimatta (2020) demonstrates the method family with two specific implementations, lengthwise rarity scaling and lengthwise quantile scaling. In lengthwise rarity scaling, when computing the scaled value for a feature count, each feature count is compared against the full set of feature counts in texts of the same length. Each feature count is then replaced with the percentage of smaller feature counts in texts of the same length. In other words, the following question is asked for each text: “what percentage of all of the texts of the same length as this text has fewer instances of this feature?” Of the two implementations, lengthwise rarity scaling is noted by Liimatta (2020) to be particularly useful for visual exploration of data. The advantages of this scaling method include the fact that it particularly highlights smaller differences in rates of occurrence within the data, making it easier to pick up on subtler differences between groups of texts, and that it scales the observed variation into a constrained range between 0% and 100%, facilitating the graphing and interpretation of the results. Figures 2 and 3 demonstrate the effect of lengthwise rarity scaling. Figure 2 is based on the typical method of calculating normalized frequency. It shows a kernel density estimation (KDE) of first-person singular pronouns in comments from three different subreddits sampled so that all three subreddits and different text lengths are represented more evenly in the figure in order to highlight the effects of the scaling method. The frequencies are clearly packed in the lower end of the frequency range, and it can be difficult to tell what is going on in this picture because of the heavy overlap. On the other hand, in Figure 3, the first-person singular pronoun counts have instead been scaled using the proposed lengthwise rarity scaling method. This scaling has created a much more clear differentiation between the use of first-person singular pronouns in the three subreddits.
Text length and short texts
Figure 2. Kernel density estimation of a sample of first-person singular pronoun frequencies across different comment lengths with contour lines for 8 bins. From Liimatta (2020: 121)
Figure 3. Kernel density estimation of a sample of lengthwise rarity scaled first-person singular pronoun counts across different comment lengths, divided into 8 bins. From Liimatta (2020: 123)
119
120
Aatu Liimatta
The second implementation of the lengthwise scaling method family proposed by Liimatta (2020), lengthwise quantile scaling, builds on commonly-used statistical properties of the dataset. In particular, the scaled scores within each text length are scaled based on the median feature count and specific quantiles of the feature count distribution. In the proposed implementation, all feature counts are scaled so that the median feature count within each text length becomes 0, and the values above and below the median are scaled separately from each other so that the .05 quantile becomes −1 and the .95 quantile becomes 1. Since lengthwise quantile scaling is based on the median and certain quantiles of the data, the −1 and 1 lines are particularly useful for the interpretation of the data in terms of recognizing texts with uncharacteristically high or low feature counts for any given text length. On the other hand, in contrast with lengthwise rarity scaling, lengthwise quantile scaling does not confine the scaled values into any particular range. It also does not highlight smaller differences as well as lengthwise rarity scaling. However, thanks to its basis on common statistical measures, lengthwise quantile scaling may be the better of the two methods to use as a preprocessing step for further statistical or computational analysis. The main downside to both lengthwise rarity scaling and lengthwise quantile scaling is that they require a very large dataset, so that there are enough texts of every individual length to make it possible to compare texts of the same length. Such datasets are most readily based on social media and other online sources. However, if the dataset mostly contains longer texts, even a slightly smaller dataset will do if texts of adjacent lengths are binned together. If the dataset is even smaller still, and/or includes shorter texts as well, the two methods may not work too well. But these two methods are only two potential implementations of the lengthwise scaling method family. In situations where the dataset is relatively small and includes a large number of shorter texts, other kinds of implementations of the lengthwise scaling method family may work better. For instance, Liimatta (2020) suggests the use of resampling methods, which could be used even with a smaller corpus, and possibly in conjunction with, for example, binning, to estimate the population parameters for different text lengths, which can then be used as the basis for the comparison. However, even this method is unlikely to work with the smallest corpora and the shortest texts, which do not have enough text for each text range to estimate the population parameters with any reliability. 3.2.2 Multiple Correspondence Analysis There also exist methods for specific purposes which can be used with shorter texts. For instance, factor analysis methods, such as those used in the multidimensional method of register analysis, rely on feature frequencies, and as such
Text length and short texts
the methodology is difficult to apply to genres which include a large proportion of short texts. In order to get around this issue in their multi-dimensional analyses of Twitter tweets, Clarke and Grieve (2017, 2019) make use of a method called Multiple Correspondence Analysis (MCA). MCA is a dimensionality reduction method which can be used to extract dimensions of variation from a set of variables. However, unlike methods such as factor analysis or principal component analysis, which rely on continuous variables such as feature frequencies, MCA extracts its dimensions from categorical variables, such as the presence or absence of a feature in a text. Consequently, MCA can be used to analyze the dimensions of variation within genres with extremely short texts, such as tweets. However, while MCA works well with genres with only short texts, it cannot be used with datasets which include longer texts. This is because the longer a text becomes, the more likely it is to include any given feature. As the texts get longer, more and more of the features of interest start appearing in every text. Due to this, the co-occurrence patterns end up saturated when analyzing longer texts, rendering the method unusable with such texts. 3.2.3 Resampling methods Resampling methods are powerful statistical methods which “make the best use of the available data” (Säily 2014: 47) and create confidence intervals and estimate the population parameters for, for example, the rate of occurrence of a linguistic feature or item. Resampling methods have been developed with the idea that the dataset is only a sample, an imperfect representation of the full population it is supposed to represent. When using resampling methods, the dataset is first divided into samples, for example, texts, groups of texts, or parts of texts, depending on the research questions and statistical assumptions being made. Then, this set of samples is repeatedly sampled randomly in order to create new artificial datasets. The exact details of how the sampling is conducted depend on the specific resampling method chosen, such as permutation or bootstrapping. The results of these different methods also have different interpretations. The new set of artificial datasets is then analyzed in order to estimate the value and its confidence intervals. Resampling methods have been used in various studies of linguistic variation. They can be used simply to estimate the rate of occurrence together with its confidence intervals, or to enable analysis in situations where using the standard method of normalization is difficult (e.g., Gries 2006, 2022; Lijffijt et al. 2016; Säily 2014). While studies making use of resampling methods often do not explicitly deal with the problem of text length specifically, the core idea of the methodology is very applicable to this problem as well.
121
122
Aatu Liimatta
3.3 A related problem: Lexical diversity While the effects of text length have generally speaking not been studied very much in corpus-linguistic research, there is a group of measures, whose relationship with text length has received some more attention: the type-token ratio and other measures of lexical diversity (or “lexical richness”). While the type-token ratio differs as a measure from the typical calculated normalized frequencies, its relationship to text length still bears discussing in this context. Like its name suggests, the type-token ratio is the ratio of the number of different words in a text (types) to the number of all words in the text (tokens). This ratio is notoriously sensitive to variation in the length of the text it is calculated for. Due to this sensitivity, for the results to be comparable, the ratio should be calculated for texts of almost the exact same length. However, since all texts in a normal-sized corpus are rarely close enough to each other in length, as a typical workaround, the ratio is calculated for a set number of words (such as 400 words) taken from the beginning of each text. While this workaround has been used for a long time to good effect, it is also not optimal, since in many cases it excludes a large majority of the text from the calculation. The solution is a lot less optimal still for datasets with a lot of variation in text length, since the 400-word sample covers a different fraction of each text, which means that every text is represented differently by the sampling. Due to these problems, and the fact that being able to measure lexical diversity in a meaningful way would be very desirable for many linguistic questions, the question of whether a method which is less sensitive to text length could be devised has received a decent amount of attention from corpus linguists and others. Hess et al. (1986) and Hess et al. (1989) test various mathematical transformations of the basic type-token ratio and conclude based on their results that no simple transformation can make the type-token ratio comparable across text lengths. Because of the inherent problems with the type-token ratio, various alternative measures of lexical diversity have been created. These include, for example, the Moving-Average Type-Token Ratio (MATTR) (Covington & McFall 2010) and the Moving Window Type-Token Ratio Distribution (MWTTRD) (Kubát & Milička 2013). Some other methods of calculating a lexical diversity score can be found in, for example, Koizumi and In’nami (2012), who compare the performance of six different measures of lexical diversity in shorter text samples between 50 and 200 words, and in Shi and Lei (2020), who more recently compare two entropy-based measures of lexical diversity. The problem of lexical diversity measures is closely related to the problem of text length and short texts in focus in the present study. While the efforts to develop a measure of lexical diversity which is less affected by text length do not
Text length and short texts
directly target the problem of text length and short texts, the implication of these efforts is clear: methods which lessen the confounding effects of variation in text length can be developed. Maybe some method created for the purpose of measuring lexical diversity could even be adapted to help with the problem of text length in feature frequencies.
4. Conclusion The present chapter has discussed two related problems, the more general problem of variation in text length and the more specific problem of short texts. While these problems have not received as much attention than they could have from quantitative corpus linguists (as evidenced by, e.g., the body of research on measures of lexical diversity), the difficulties caused by the confounding effects of text length are only going to become more central to many studies, as more and more research is being done on social media and web data. A number of solutions and workarounds to remedy the problems have been devised, all with their own advantages and disadvantages. These solutions can be used to good effect in many kinds of linguistic investigations. However, there still is no one-size-fits-all solution to the problems caused by text length and short texts in quantitative text-analytic corpus-linguistic studies. Some potential avenues for improvements and new method development have been proposed in the present chapter. Since resampling methods are very powerful for estimating the distribution based on smaller datasets, they appear as a potentially useful avenue for the development of new methods for the analysis of texts across text lengths. At the same time, larger datasets contain more information about the variation inside them, so various approaches making use of the large size of the data, such as those developed by Liimatta (2020), may also be useful in getting around many of the problems caused by text length in at least some studies. However, such approaches can naturally only be used with a limited number of datasets which are large enough. Even if a perfect all-encompassing solution does not exist yet, or is not possible at all, the solutions mentioned in this chapter can still be used to study many linguistic questions, given that one is aware of the potential implications of their use. There certainly exist many other approaches not mentioned here, particularly various more advanced statistical and computational methods, which are less affected by variation in text length. Nevertheless, there is still a lot of room left for the development of new ways to analyze datasets with a wide range of text lengths, and particularly datasets which contain extremely short texts, which are more common today than ever.
123
124
Aatu Liimatta
References Biber, Douglas. 1988. Variation across Speech and Writing. Cambridge: CUP. Biber, Douglas. 2014. Using multi-dimensional analysis to explore cross-linguistic universals of register variation. Languages in Contras, 14(1): 7–34. Biber, Douglas & Conrad, Susan. 2009. Register, Genre, and Style. Cambridge: CUP. Biber, Douglas, Csomay, Eniko, Jones, James K. & Keck, Casey. 2004. A corpus linguistic investigation of vocabulary-based discourse units in university registers. In Applied Corpus Linguistics: A Multidimensional Perspective, Ulla Connor & Thomas A. Upton (eds), 53–72. Amsterdam: Rodopi. Biber, Douglas, Egbert, Jesse & Keller, Daniel. 2020. Reconceptualizing register in a continuous situational space. Corpus Linguistics and Linguistic Theory 16(3): 581–616. Clarke, Isobelle & Grieve, Jack. 2017. Dimensions of abusive language on Twitter. In Proceedings of the First Workshop on Abusive Language Online, Zeerak Waseem, Wendy Hui Kyong Chung, Dirk Hovy & Joel Tetreault (eds), 1–10. Vancouver BC: Association for Computational Linguistics. Clarke, Isobelle & Grieve, Jack. 2019. Stylistic variation on the Donald Trump Twitter account: A linguistic analysis of tweets posted between 2009 and 2018. PLoS One 14(9): e0222062. Conrad, Susan & Biber, Douglas (eds). 2001. Variation in English: Multi-dimensional Studies. Harlow: Pearson Education. Covington, Michael A. & McFall, Joe D. 2010. Cutting the Gordian Knot: The Moving-Average Type-Token Ratio (MATTR). Journal of Quantitative Linguistics 17(2): 94–100. Gries, Stefan T. 2006. Exploring variability within and between corpora: Some methodological considerations. Corpora 1(2): 109–151. Gries, Stefan T. 2022. Toward more careful corpus statistics: uncertainty estimates for frequencies, dispersions, association measures, and more. Research Methods in Applied Linguistics 1(1). Hess, Carla W., Haug, Holly T. & Landry, Richard G. 1989. The reliability of type-token ratios for the oral language of school age children. Journal of Speech and Hearing Research 32: 536–540. Hess, Carla W., Sefton, Karen M. & Landry, Richard G. 1986. Sample size and type-token ratios for oral language of preschool children. Journal of Speech and Hearing Research 29: 129–134. Hiltunen, Turo & Tyrkkö, Jukka. 2019. Academic vocabulary in Wikipedia articles: Frequency and dispersion in uneven datasets. In From Data to Evidence in English Language Research, Carla Suhr, Terttu Nevalainen & Irma Taavitsainen (eds), 282–306. Leiden: Brill. Koizumi, Rie & In’nami, Yo. 2012. Effects of text length on lexical diversity measures: Using short texts with less than 200 tokens. System 40(4): 554–564. Kubát, Miroslav & Milička, Jiří. 2013. Vocabulary richness measure in genres. Journal of Quantitative Linguistics 20(4): 339–349. Liimatta, Aatu. 2019. Exploring register variation on Reddit: A multi-dimensional study of language use on a social media website. Register Studies 1(2): 269–295.
Text length and short texts
Liimatta, Aatu. 2020. Using lengthwise scaling to compare feature frequencies across text lengths on Reddit. In Corpus Approaches to Social Media, Sofia Rüdiger & Daria Dayter (eds), 111–130. Amsterdam: John Benjamins. Liimatta, Aatu. 2022a. Register variation across text lengths: Evidence from social media. International Journal of Corpus Linguistics 28(2): 202–231. Liimatta, Aatu. 2022b. Do registers have different functions for text length? A case study of Reddit. Register Studies 4(2): 263–287. Lijffijt, Jefrey, Nevalainen, Terttu, Säily, Tanja, Papapetrou, Panagiotis, Puolamäki, Kai & Mannila, Heikki. 2016. Significance testing of word frequencies in corpora. Digital Scholarship in the Humanities 31(2): 374–397. Shi, Yaqian & Lei, Lei. 2020. Lexical richness and text length: An entropy-based perspective. Journal of Quantitative Linguistics 29(1), 62–79. Säily, Tanja. 2014. Sociolinguistic Variation in English Derivational Productivity: Studies and Methods in Diachronic Corpus Linguistics. Helsinki: Société Néophilologique de Helsinki. Winter, Bodo & Grice, Martine. 2021. Independence and generalizability in linguistics. Linguistics 59(5): 1251–1277.
125
Corpus genre categories Issues at the intersection of linguistics and literature Daniel Ocic Ihrmark Linnaeus University
This chapter highlights genre categorizations as a pitfall at the intersection of corpus linguistics and literature and problematizes the use of the genre category from the perspectives afforded by both fields. The intention is for the paper to argue for a more explicit communication of our genre categorization practices, and by doing so suggest ways of avoiding miscommunication and confusion due to the genre term being understood differently within different disciplines and backgrounds. The conclusion is that the wider categorizations used, such as novel or short story, are likely to be the most practical, and that studies wanting to sub-categorize further using the genre term should instead apply it according to their specific needs accompanied by explicit discussion of the implementation. Keywords: genre, corpus linguistics, stylistics, literature, special corpora
1.
Introduction
Fiction is inherently messy to work with. This is not due to the material itself, but rather due to the field that surrounds it and the needs of different theoretical approaches to the reading of the materials. Since Leech and Short’s Style in Fiction (2007), the intersection between linguistics and literature has continued to develop. Mahlberg (2007) described how the relationship between corpus linguistics and literary theory can be approached, and the framework needed to successfully utilize the methodological resources within the field of literature. Likewise, Biber (2011) provides interesting examples of different ways corpora have been used in the study of literature, where methods such as keyword analysis, n-grams, and collocations have been at the center stage. However, the use of large corpora for comparative studies within literature remains problematic as the corpora were rarely constructed for this purpose. https://doi.org/10.1075/scl.118.08ihr © 2024 John Benjamins Publishing Company
Genre categories: Between linguistics and literature
The concept of special corpora, defined by Tognini-Bonelli (2010: 13) as corpora where the selection is not made to be representative of a language but of a specific use-case, such as the learner corpora available in the International Corpus of Learner English (ICLE) (Granger et al. 2020), plays an important role here. Literary corpora are often designed to be representative of an author, a period, or a genre, which makes the categorization procedures very different from the procedures used in general corpora. While categorization in large corpora may make use of broad text type categorizations in combination with temporal and spatial categorization, as in the British National Corpus (BNC), the Corpus of Contemporary American English (COCA), and the Corpus of Historical American English (COHA), special corpora often require other category tags. The temporal and spatial categorization of texts along with the text type becomes central for comparative studies of literature. An example would be when the works of a single American author of fiction active during the 1920s to the 1960s are being contrasted with American fiction written between the 1900s to 1960s (as done in Ihrmark & Nilsson 2021). As the interpretation of data begins, further questions regarding the categorizations arise, often to do with genre and style, and one must consider whether the American author active during the 1920s to 1960s was a modernist and so forth. Consequently, comparing the fiction of one author to that of any other author matching the temporal and spatial categorization becomes problematic. It becomes more of an issue when presenting to an audience mainly engaged with other aspects of the author’s work rather than the “when and where” (as attempted in Ihrmark 2018, 2019). Categorization within corpora is a well-researched topic which has produced multiple excellent methods of approach using keywords (Özgür, Özgür & Güngör 2005), named-entity recognition (Sahin et al. 2017) or machine learning (Sebastiani 2002), but these are of limited use when the desired categorizations are based on features not directly tied to the language itself. An example more closely related to the topic of this chapter would be the genre classification scheme employed in the British National Corpus (BNC) by David Lee (2001). As the intersection between linguistics and literature becomes more popular and more populated with resources, it becomes important to discuss how these uses of corpora as contrastive, or comparative resources beyond language variants and variation could look. How could this new arena influence our categorization habits, and what are the consequences of deeper categorization of fictional texts? This chapter highlights genre categorizations as a pitfall at the intersection of corpus linguistics and literature and problematizes the use of the genre category tag from the perspectives afforded by both fields. The chapter aims to contribute towards a more explicit communication of our genre categorization practices, and avoidance of miscommunication and confusion due to the genre term being understood differently within different disciplines and backgrounds.
127
128
Daniel Ocic Ihrmark
2.
Looking up from the pit
I have presented papers using reference corpora to distinguish author-specific traits on several occasions (for instance, Ihrmark 2018 and Ihrmark 2019) and regularly received questions specifically about the issue of the comparative aspects. One paper was presented to the Hemingway Society in 2018, and compared sentence lengths, noun distribution and lexical density in Ernest Hemingway’s writing, using the COHA fiction category as reference. The reference corpus was further limited to include only materials produced between 1900 and 1960 to correctly match the time period. One line of questioning after the presentation was especially interesting for the current paper: First, was my selection of reference materials appropriate? And, second, what should be considered appropriate materials for distinguishing features specific to Hemingway? Regarding the first question about whether my selection was appropriate or not, I had approached it from an admittedly simplistic perspective. My line of thought had been that the texts belonged within the same category as they were all fiction, and that they were produced during the same time period. In retrospect, the temporal aspect had likely played too large of a role in my justification, and more time should instead have been spent considering the nature of the content. While the reference corpus did offer sub-categorization of the fiction component of the corpus, aligning those sub-categories with Hemingway’s oeuvre is not a straightforward task. Partially this has to do with the sub-categorization of the reference corpus in question, but a more prominent issue is the fluctuating idea of Hemingway’s genre belonging throughout his work, such as his short stories (Donahue 2003). The genre of a story could be a matter of debate, and perspectives rooted in different assessments of the appropriate genre could result in different readings of the text when applying close reading. This makes it difficult to argue for a higher granularity in corpus categorizations of fiction being a solution, as the use-cases for the genre categories, and their influence, are different between the fields. Moving on to the second question: what should be considered appropriate materials for distinguishing features specific to Hemingway’s writing? The way we end up in the pit, to me, begins with what we consider as “specific to an author’s writing”. From a general standpoint, one could do as I did and simply use fiction in general as the reference, and asking the question of what makes this author stand out amongst a sizeable sample of other authors active during the same period. However, this leaves the comparison vulnerable for criticisms regarding the oversimplification of fiction, as distinguishing what is specific to an author from what is specific to their peers can be done on many levels. Continuing with Hemingway as an example, it could for instance be operationalized as asking what
Genre categories: Between linguistics and literature
the differences are between Hemingway and other authors belonging in the Lost Generation of post-war American expatriates, other authors who write about war, other authors sharing his journalistic background, or other authors connected to Gertrude Stein or to Sherwood Anderson. In my own practice, I decided on the first of these suggestions when working on the Lost Generation Corpus. The corpus is aimed at collecting fiction and nonfiction for Lost Generation authors and sorting the texts diachronically in order to explore language development over time within the group. The timeline is a central component to the concept of the Lost Generation, as it was initially coined by Gertrude Stein in attempt to describe the post-World War 1 generation, popularized through Hemingway, and has since been used to refer to a group of expatriate writers active in Paris during the 1920s. After the 1920s, the group scattered internationally and the timing of the texts thus tie into important information regarding the individual authors. In addition to the temporal tags, texts were also identified as being novels, short stories, essays, letters, or diaries. The categories were intended to enable researchers to look at language use in fiction as compared to personal communications and notes. Turning to the focus of this chapter, the relevant categories are novels and short stories. The initial rationale for applying these two categories was that they were easily applied with an acceptably high consensus for a majority of the authors based on other descriptions of the Lost Generation, and that they provided a clear, functional separation of texts. This allows for the corpus to be used for comparisons between the authors short stories and novels, as well as comparisons between different authors considered to belong within the same group. However, the corpus can only be used for questions based on comparisons within the Lost Generation in a broad sense, and it does not allow the user to perform comparisons along other lines of inquiry, for instance, carrying out comparisons between different genres of fiction. So far, the categorization is fairly intuitive and creates no overlapping tags, that is, a text is never both a short story and a novel. From a qualitative perspective and using a close-reading method, as is often the case in literature research, the use of overlapping genre categorizations can provide a way of connecting the content and style of the text being read to different movements or periods from which the author might have drawn their inspiration or within which they might have been an active participant. The genre features found in a piece of fiction can also be used to extend an argument regarding the intertextual properties of a text in such a way that the intellectual context of the writing can be connected to the written text. An example of this could be the use of genre tropes to create connection through allusions to other texts or even direct references to previous works that add to the narrative being conveyed, for instance a character coming across a different text within the story and having
129
130
Daniel Ocic Ihrmark
it influence the narrative. Genre plays a role in these examples as it draws on the expected previous reading of the audience, their context, and readers of one genre can more often be expected to be familiar with other texts of the same genre.1 This is a very different use-case than that of genre categorization as a sorting mechanism for achieving portioning of large datasets, which would often be the purpose of genre categorization of literary materials or fiction in the distantreading approach commonly used within corpus stylistics or digital humanities (Oberhelman 2015). The different approaches to the idea of a text categorization are also often connected to the perceived usefulness of these two different modes of research: externally defined for the close-reading relying on genre and based on internal features for the distant-reading approaches relying on text types, which is occasionally understood synonymously (Smitterberg & Kytö 2015: 119). This becomes an issue when performing research in-between the two fields, and inbetween the two modes of engaging with the texts. This is, of course, a simplified take on the issue, as overlapping categorization is also sometimes used for further analysis in computational methods, for instance analysing the distribution of topics (as done by Murakami et al. 2017). However, the use of genre categorization within literary studies differs from the uses of these more objective categorizations in distant-reading, and the two approaches are not mutually interchangeable, even though they might both be interested in overlapping features. Turning briefly to the world of fiction in film, combinations of genres have been shown to be viable from a commercial perspective. Star Wars (1977) stands out as a good example of this, as it has been described to combine aspects of multiple genres of film and fiction, such as westerns, samurai films, and science fiction, while still being extremely successful commercially (Taylor 2015). However, Berliner (2017) identifies a discrepancy in the quality evaluation of Star Wars amongst expert viewers and the general audience, which he traces back to the act of combining genres and the expectations set in the different groups based on a priori understanding. This led to the critics viewing the genre bending in Star Wars as problematic, while regular viewers did not. The Star Wars dilemma could be seen as the content of the work not matching the expert viewers’ expectations set by the genre categorization. The reception of Star Wars as an example of how genre has functioned in fiction prompts the question of how important 1. This is, of course, not the only criteria for successfully implementing these devices. Phenomena such as national literary canons and other expectations on the reader’s previous reading play a large role. This could also take the form of an intertextual connection made with the intent of being an “easter egg” for avid readers of what the author considers the context of their work.
Genre categories: Between linguistics and literature
the intuitive taxonomy inferred by genre labelling in fiction could be when applying genre as a sorting mechanism. Smitterberg and Kytö (2015) explored the use of the genre parameter in diachronic corpora of historical English and highlighted it as an important part of corpus compilation: “If the researcher is compiling a single-genre corpus, delimiting the genre sampled is crucial in order to reach reliable results. And if the corpus project includes several genres, considering genres in relation to one another is a key issue when the compiler decides what research questions studies based on the corpus can hope to answer” (Smitterberg & Kytö 2015: 118). Genres are described as “fuzzy sets”, meaning that some of the content in the genre category will be very close to the archetype while other texts will deviate more (Smitterberg & Kytö 2015: 118), similarly to what was described by Biber (1989). Several novels are highlighted as examples by Smitterberg and Kytö (2015), showcasing some of the issues one can run into when dealing with fiction. For instance, a novel could take the form of a diary or a collection of letters but still remain a novel (Smitterberg & Kytö 2015: 119). The basis of their showcase is the problematic sometimes synonymous use of the terms text type and genre, which gives rise to much internal “fuzziness” in the categorizations. The terms used by Smitterberg and Kytö to discuss the topic are genre as referring to text-external grounds for categorization and text type as referring to for categorizations based on the text itself. Genre is seen as a combination of intended function and audience expectation (Smitterberg & Kytö 2015: 119), connecting well to the dilemma regarding fiction genres showcased by Star Wars; that is, a mismatch between expectation and content created by the genre tag, although for a corpus linguist it could rather result in a mismatch between the actual features of a text and the features assumed by the user based on the tag.
3.
Text genre categorization in literature
Let us first start by discussing genre as a literary concept. Todorov (1990) places the genre designation of a text as a function of the choices made by the author but emphasizes the importance of contextual salience as a sorting mechanism for possible choices: “The literary genres, indeed, are nothing but such choices among discursive possibilities, choices that a given society has made conventional. For example, the sonnet is a type of discourse characterized by supplementary constraints governing its meter and its rhymes.” (Todorov 1990: 10). Todorov’s description emphasizes the influence of conventional choices on the structure of a piece, in this case a sonnet, but the concept of choices made conventional by society is an interesting one for the discussion carried in this chapter.
131
132
Daniel Ocic Ihrmark
Seeing the literary text as a social event adhering to (and understood through) the expectations of the receiving audience does connect well to intertextual ideas of literature, for instance, the work of Barthes proclaiming the death of the author as the birth of the reader (Barthes & Heath 1977). As stated by Chandler, genre categorization cannot be considered an objective procedure and is rather a “theoretical minefield” (Chandler 1997: 1). Genre being a theoretical minefield does not begin in the 1990s, however, as Levin points out in his 1984 review of Fowler’s Kinds of Literature: An Introduction to the Theory of Genres and Modes (1982). Literary genres such as drama, poetry and prose are at the core of the argument, but sub-categorization beyond that becomes complicated as the idea of texts moving between genres, or genres themselves undergoing change over time, is brought up. The relationship between literary scholars and genre could be seen in Levin’s statement that Fowler “[…] shies away from the merest hint of classification-perhaps a little too far; for, as soon as we admit that there can be different kinds of anything, we recognize the possibility of a taxonomic approach” (Levin 1984: 258). However, genre in literature has also been used for more applied reasons. Littlefair (1991) explored genre in a classroom setting with a focus on implementation in language teaching. The categories used by Littlefair were literary, expository, procedural, and reference genres. The literary genre is described as containing texts where authors provide an imaginative or personal experience through description and narrative (Stamboltzis & Pumfrey 2000: 59). Description and narrative align with the idea of plot and character being central to the genre taxonomy, but description also encompasses the language traits that go into the describing the contents of the text. While this provides an idea of what would place a text in the literary genre as a whole, it offers little help on the topic of how literary texts are then further divided into genres like fantasy, detective fiction, horror or science fiction. Kirk and Pearson (1996) suggest that genre should be seen as a combination of the structure of a text and the author’s intent. While the structure of the text can be understood from the textual artifact, in most cases, the author’s intent becomes more problematic as a text categorization tool. In terms of literary genres, the intent of the author can be said to manifest through the selection of devices used for the crafting of the narrative, or through the use of genre tropes, but whether these are employed as a form of pastiche or homage, or simply as a symptom of the text being written for the specific genre, opens up for a very subjective categorization approach.
Genre categories: Between linguistics and literature
4.
Text genre categorization in linguistics
It is important to start out with a distinction between text type and genre, as the terms occasionally fulfill a similar function (Paltridge 1996: 237). Text types are defined by Colwell and Tune (1969: 59) as being “the largest group of sources which can be generally identified”. While Colwell and Tune were mainly discussing the use of text type taxonomy in their studies of New Testament manuscripts, their definition is echoed by later descriptions of categorization procedures on other media types. An example providing a practice-oriented analogy can be seen in Allen (1989: 44), who describes the approach to the genre categorization of narrative media (specifically televised soap operas) as being “…much as the botanist divides the realm of flora into varieties of plants”. Biber (1989: 5) refers to this kind of implementation as “the folk-typology of ‘genres’”, where the distinct categories are easily identified. In this typology he includes texts such as novels, newspaper articles, academic articles and editorials, but also spoken language productions such as radio broadcasts, everyday speech and public speeches (Biber 1989: 6). The categories are defined by external format, such as location, as well as by purpose and situation. Biber (1989) also highlights the fuzzy nature of such categories as the texts within them might vary in their closeness to the archetype and describes the difference between text type and genre as being the former relying on text-internal features, and the latter having more to do with text-external or sociocultural ones. The use of genres as a sorting mechanism also has practical roots amongst linguists, as Lee (2001: 38) points out that “it is impossible to make many useful generalisations about ‘the English language’ or ‘general English’ since these are abstract constructions”. Genres thus provide the basis for categorizing language in a way that allows for more useful generalizations. Genres have also been defined from a more theoretical perspective within linguistics, some of which see genre as being dependent on the intended social context of meaning-making. Based on that conceptualization and the earlier work of Halliday (1978), genre was seen as a result of the language used in the meaningmaking process and the discourse community for which a text is intended (Martin 1999) or as a communicative event involving members of a discourse community and adhering to the communication purposes therein (Swales 1990). This is drastically different from the conceptualization seen in the previous descriptions, as the categorization here explicitly informs the authoring of the text. Kress and Knapp (2008) take a similar perspective on the concept of genre, and describe it as being the result of specific social events and their specific demands on form and function. Form and function become interesting distinctions as they could be considered to match with the separating features between text type and genre as presented by Paltridge (1996).
133
134
Daniel Ocic Ihrmark
Paltridge explores both text type and genre in the classroom setting and indicates a division according to generic structures and text structures, with the former being tied to genre and the latter to text type (1996: 239). The terms used are taken from Biber (1988), who saw the terms as complementary to each other and fruitful to use in tandem. Examples of genres brought up by Paltridge include recipes, advertisements, formal letters, personal letters, film reviews and student essays, while examples of text types are things such as procedures, anecdotes, expositions, and descriptions (Paltridge 1996: 239). These examples also serve to outline what could be considered a generic structure (genre) and a text structure (text type). Turning to corpora, text categorization often takes place first at a higher level, where genre can, for instance, be defined as “Fiction”, as is seen in the widely used COHA or the COCA corpora. The two corpora then provide further granularity by introducing sub-categorization. In the case of COHA, these sub-categories are explicitly tied to the Library of Congress taxonomy of non-fiction and academic materials (Figure 1), and fiction is not provided with further categorization. Fiction as a broad genre thus becomes the only filtering, with no further categorization available, although the Library of Congress taxonomy does provide a myriad of sub-categories to the genre heading.2
Figure 1. Text categories in the Corpus of Historical American English (COHA)
COCA provides a more granular sorting possibility for the material broadly defined as fiction, which includes categories relating more to what one would more intuitively consider as genres of fiction, such as “Scifi/fantasy” (Figure 2). However, the corpus also provides categorization in very broad terms, such as “General (book)” or “General (journal). We also see a category that could be directly tied to intended readership, “Juvenile”, although this could also be considered connected to youth fiction as a genre. The sub-categorization does not seem to follow a specific approach to the idea of genres, as some categories are defined by medium, some by content and some by target audience. 2. See list of genre terms for 2022 at 〈https://www.loc.gov/aba/publications/FreeLCGFT /GENRE.pdf〉
Genre categories: Between linguistics and literature
Figure 2. Fiction sub-categories of the Corpus of Contemporary American English (COCA)
Moving on to the BNC, fiction is categorized according to the tags drama, poetry, and prose (Figure 3). Drama refers to narratives intended to be acted out, poetry to composition using rhyme and/or meter, and prose to the “ordinary” written language without a meter. Prose is often further divided into fiction and non-fiction, although the higher-level category of fiction present in the BNC indicates that it is only the former that is included in the category here. Lee (2001) refers to the highest-level categories within the BNC’s written content as supergenres, which are used for larger groupings that are later subcategorized further. Based on these super-genres the underlying categories can be seen as assuming prose for non-fiction writing. An example of this is the Academic prose supergenre consisting of field-specific publications simply labelled as “ac”, or academic, combined with the field without prose being specified in the tag (Lee 2001: 57).
Figure 3. Genre and domain categories in the British National Corpus (BNC)
Lee (2001) provides the perspective used for the categories applied to the BNC, and suggests an approach based on pragmatic needs. The use of categories such as audience age is tied to possible users intending to create materials for use in classrooms with specific age groups of learners. This makes distinctions such as the one between adult prose fiction and children’s prose fiction important (Lee 2001: 55). In general, the approach described by Lee involves fairly broad main categorizations supplemented by category options connected to additional data. In summary: “When compiling a sub-corpus for the purpose of research, classroom concordancing, genre-based learning, and so forth, you need all the available information you can get” (Lee 2001: 55).
135
136
Daniel Ocic Ihrmark
5.
The genre category pitfall
To approach the potential genre category pitfall, this paper will first consider the position of corpus linguists and literature scholars based on the previous sections. In general, their positions on genre could be described to as defined by their relationship to the text as an object. The components highlighted as interesting from a linguistic perspective are features such as style, intended readership, intended function and the context in which the text is intended to fulfill a communicative purpose. The different purposes of the text are also of interest, especially when seen in relation to the other features. The position could be oversimplified as being one that is interested in the purpose and form of a text, and often focused on the intersection between the two. The occasional overlap between the idea of a genre and that of a text type serves to further muddy the waters in linguistic usage of the genre term. The perspective afforded by the field of literature, on the other hand, seems more interested in the intersection between style and content. Style is close to the idea of form as used in the definitions from the linguistics side of things, but it brings with it quite different connotations (see Aquilina 2014 for a thorough discussion). While style in linguistics can be discussed as having to do with formality, register and fitness-for-purpose, style in literature takes on an intertextual meaning connected to the idea of a genre. In addition, genres are often defined by their content to a higher degree, as the setting is used for genre descriptions in genres such as science fiction or westerns, whereas style and plot can play a larger defining role in genres such as horror or fantasy. Returning to the Star Wars dilemma, or the issue of expectations tied to the genre term amongst different audiences raised by Berliner (2017), these positions are exceptionally well set up for communication issues. This pitfall is different from the situation described there as both of the audience groups are expert readers with plentiful a priori knowledge of what is expected from the genre tag used. From the perspective of the Star Wars dilemma, not only would the use of a genre description be inviting issues with the individual groups as the complex theoretical backgrounds are likely to have set quite specific expectations regarding form, style, content and purpose, but using the genre categorization for a resource intended for both audiences is likely to result in a best-case scenario where only one group is dissatisfied or confused by the labelling, and a more likely worst-case scenario where neither group will find the categorization appropriate or useful. Genres as a mode of categorization in corpus resource creation vis-à-vis use as descriptive tags used for fiction should also be considered an issue in the same pit, as the expected rules governing the use are different in nature. Overlapping genre descriptions in literature or fiction are not considered an issue. In fact, the
Genre categories: Between linguistics and literature
combination of genre tropes and features within a piece can often be seen as a strength, as indicated by Berliner (2017). However, when genre is applied as a taxonomy for sorting materials for corpus linguistic methods, overlapping tags are less common. Turning from the issues to do with the positions of the audiences to the practice-oriented issues of the intersection between corpus linguistics and literature, there must first be an understanding of which tasks are performed therein. Biber (2011) has indicated that tasks such as keyword analysis, n-gram searches, collocation searches and lexical clusters are often at the center of projects within corpus stylistics. N-grams, collocations, and lexical clusters can all be said to have to do with the recurring language patterns of an author, and could be tied to form, style and content in that they indicate to the researcher what is commonly occurring in the materials. However, the projects in which these tasks are performed often deal with determining what is especially characteristic for a specific author or a specific text (Biber 2011), meaning that a reference sample must be used. Keyword searches would fall within a similar category, as keywords are often defined as being prominent in one text, or one collection of texts, when compared to another. This is where the genre pitfall opens up, and issues start to arise. To what, exactly, are we supposed to compare the language use of our author, or the language used in our texts? From a distant-reading linguistic perspective, the categorizations relying on text-internal features, such as the one presented by Biber (1989), might make more sense. These typologies rely on more objective metrics and can be agreed on based on the texts themselves.3 However, as shown by the reference corpora, categorizations used for corpus compilation do not adhere strictly to these taxonomies and instead include genres such as novels and newspaper articles, which both belong to the genre folk-typology. Moving on with using novels as the genre category, as I did in my work, results in a very broad category that can, for the most part, be commonly agreed upon. From a close-reading literature perspective, the resulting category “Novels” becomes too broad to be useful. It does not influence the close-reading method in a meaningful way and does not indicate information that could tie the text(s) to a broader intellectual context. Instead, a more useful taxonomy would be tied to fiction genres, allowing the researcher to connect the texts to their contexts and content more clearly.
3. As pointed out by one of the reviewers, this could also lead to circularity if the texts are later explored for linguistic features tied to the genre. For example, if the genre “Reports” is defined by text-internal features, those features would be very frequent within the resulting category.
137
138
Daniel Ocic Ihrmark
For those stuck between linguistics and literature, it becomes a decision regarding the intended audience of the corpus. In the case of a resource created for the linguistics community it would make sense to include both the text type and the broader genre, as both have a theoretical background through which results could be connected to previous and future research, as well as inform about the nature of the item in an objective way. It is also important to consider which methods are going to be applied, as the methods relying on comparisons exemplified by Biber (2011) would produce different results depending on the materials being compared to one another. In those cases, it might make more sense to consider a text type comparison for a detailoriented study or a genre comparison for things such as a discourse analysis of a specific medium. However, if the comparisons are intended to inform a closereading approach or are being conducted against a backdrop of previous research conducted using such an approach, the use of tags more closely connected to the use of genre in literary scholarship would be necessary, or at least desirable. Returning to my own pit for a moment, the use of the “Novel” and “Short story” categorization in the Lost Generation Corpus starts to appear more palatable to me than it was earlier. While very broad categories, they still provide some clear idea of their content in terms of it being fiction of a certain length, albeit with a large helping of internal variation as exemplified by Smitterberg and Kytö (2015). As previously stated, opting for more granular categorizations beyond that would necessitate decisions regarding the approach to the material and assumptions regarding the intended audience. However, depending on how the corpus is used decisions regarding the implementation of further categorization can be made on a study-by-study basis, with the demands of the methods applied being the main point of consideration. This would also most likely result in the need for further categorization of the materials to which the corpus is being compared, or at least explicit discussion of how the two corpora might differ in terms of the text categories being compared.
6.
Conclusion
The main conclusion drawn is that fiction is difficult to categorize according to genre for corpus compilation, and especially so when trying to communicate useful information to two distinct fields of research with complex theoretical backgrounds connected to the term. In practice, this means that work taking place at the intersection between literature and linguistics must be explicit about how the term “genre” is being used, and what exactly is meant by it. The previous research referred to in this chapter could be said to highlight the communication taking
Genre categories: Between linguistics and literature
place implicitly through the use of the term “genre” as the main culprit through its setting of different expectations depending on the audience. However, the connection between the approach to genre labelling and the methods being applied provides a concrete bridge across that rift. Considering genre categorization of fiction as a part of the methodology in a study could open up for both an explicit argumentation regarding why one has decided on the labelling being implemented, as well as a clear communication to the reader about what is actually intended by the labels. Moving the act of genre labelling for increased granularity to the methodologies of individual studies according to their specific needs, the broader categorizations make a lot of sense due to the corpora then having a wider applicability. In summary, Lee’s (2001) position and rationale would likely hold for corpora aimed at literary researchers as well, with broader genres at the higher levels of categorization and as much additional information included as possible. This does not mean forcing increased granularity, but rather making the additional information available for filtering according to user needs.
References Allen, Robert C. 1989. Bursting bubbles: “Soap opera” audiences and the limits of genre. In Remote Control: Television, Audiences and Cultural Power, Ellen Seiter, Hans Borchers, Gabriele Kreutzner & Eva-Maria Warth (eds), 44–55. London: Routledge. Aquilina, Mario. 2014. The Event of Style in Literature. Houndmills: Palgrave Macmillan. Barthes, Roland. & Heath, Stephen. 1977. Image, Music, Text: Essays. London: Fontana. Berliner, Todd. 2017. Hollywood Aesthetic: Pleasure in American Cinema. Oxford: OUP. Biber, Douglas. 1988. Variation across Speech and Writing, 1st edn. Cambridge: CUP. Biber, Douglas. 1989. A typology of English texts. Linguistics 27: 3–44. Biber, Douglas. 2011. Corpus linguistics and the study of literature: Back to the future? Scientific Study of Literature 1: 15–23. BNC Consortium. 2007. The British National Corpus, XML Edition. Oxford Text Archive.
(20 May 2024). Chandler, Daniel. 1997. An introduction to genre theory. (15 May 2023) Colwell, Ernest C. & Tune, Ernest. 1969. Studies in Methodology in Textual Criticism of the New Testament. Leiden: E. J. Brill. Davies, Mark. 2008. The Corpus of Contemporary American English (COCA). (20 May 2024). Davies, Mark. 2010. The Corpus of Historical American English (COHA). (20 May 2024).
139
140
Daniel Ocic Ihrmark
Donahue, Peter. 2003. The genre which is not one: Hemingway’s in our time, difference, and the short story cycle. In The Postmodern Short Story: Forms and Issues, Farhat Iftekharrudin, Joseph Boyden, Mary Rohrberger & Jaie Claudet (eds), 161–172. Westport CT: Praeger. Fowler, Alastair. 1982. Kinds of Literature: An Introduction to the Theory of Genres and Modes. Cambridge MA: Harvard University Press. Granger, Sylviane, Dupont, Maïté, Meunier, Fanny, Naets, Hubert & Paquot, Magali. 2020. The International Corpus of Learner English, Version 3. Louvain-la-Neuve: Presses universitaires de Louvain. (20 May 2024). Halliday, Michael A. K. 1978. Language as Social Semiotic: The Social Interpretation of Language and Meaning [Open University Set Book]. London: Arnold. Ihrmark, Daniel. 2018. ‘Cultivating one true sentence’: A corpus stylistic analysis of Hemingway’s language. Presented at the XVIII Hemingway Society Conference, 22–28 July. Paris, France. Ihrmark, Daniel. 2019. ‘O Fudge, the looks of the girls’: A corpus-driven analysis of the female role in F. Scott Fitzgerald’s fiction. Presented at the 15th F. Scott Fitzgerald Society Conference, 24–29 June. Toulouse, France. Ihrmark, Daniel & Nilsson, Johan. 2021. A corpus stylistic analysis of development in Hemingway’s literary production. The Hemingway Review 40: 71–93. Kirk, Lynne & Pearson, Henry. 1996. Genres and learning to read. Literacy 30: 37–41. Kress, Gunther & Knapp, Peter. 2008. Genre in a social theory of language. English in Education 26(2): 4–15. Lee, David Y. W. 2001. Genres, registers, text types, domains, and styles: Clarifying the concepts and navigating a path through the BNC jungle. Language Learning & Technology 5(3): 37–72. Leech, Geoffrey & Short, Michael. 2007. Style in Fiction: A Linguistic Introduction to English Fictional Prose [English Language Series], 2nd edn. New York NY: Pearson Longman. Levin, Harry. 1984. Review of Kinds of Literature: An Introduction to the Theory of Genres and Modes, by Alastair Fowler. Comparative Literature 36(3), 258–260. Littlefair, Alison B. 1991. Reading All Types of Writing: The Importance of Genre and Register for Reading Development [Rethinking Reading]. Milton Keynes: Open University Press. Mahlberg, Michaela. 2007. Corpus stylistics: Bridging the gap between linguistic and literary Studies. In Text, Discourse and Corpora: Theory and Analysis, Michael Hoey, Michaela Mahlberg, Michael Stubbs & Wolfgang Teubert (eds), 219–246. London: Continuum. Martin, James R. 1999. Mentoring semogenesis: “Genre-based” literacy pedagogy. In Pedagogy and the Shaping of Consciousness: Linguistic and Social Processes, Frances Christie (ed.), 123–155. London: Continuum. Murakami, Akira, Thompson, Paul, Hunston, Susan & Vajn, Dominik. 2017. “What is this corpus about?” Using topic modelling to explore a specialised corpus. Corpora 12: 243–277. Oberhelman, Daniel D. 2015. Distant reading, computational stylistics, and corpus linguistics: The critical theory of Digital Humanities for literature subject librarians. In Digital Humanities in the Library: Challenges and Opportunities for Subject Specialists, Arianne Hartsell-Gundy, Laura Braunstein & Liorah Golomb (eds), 53–66. Chicago IL: Association of College and Research Libraries.
Genre categories: Between linguistics and literature
Özgür, Arzucan, Özgür, Levent & Güngör, Tunga. 2005. Text categorization with class-based and corpus-based keyword selection. In Computer and Information Sciences – ISCIS 2005 [Lecture Notes in Computer Science 3733], Pinar Yolum, Tunga Güngör, Fikret Gürgen & Can Özturan (eds), 606–615. Berlin: Springer. Paltridge, Brian. 1996. Genre, text type, and the language learning classroom. ELT Journal 50: 237–243. Sahin, H. Bahadir, Tirkaz, Caglar, Yildiz, Eray, Eren, Mustafa Tolga & Sonmez, Ozan. 2017. Automatically annotated Turkish corpus for named entity recognition and text categorization using large-scale gazetteers. arXiv. (20 May 2024). Sebastiani, Fabrizio. 2002. Machine learning in automated text categorization. ACM Computing Surveys 34: 1–47. Smitterberg, Erik & Kytö, Merja. 2015. English genres in diachronic corpus linguistics. In From Clerks to Corpora: Essays on the English Language Yesterday and Today, Philip Shaw, Britt Erman, Gunnel Melchers & Peter Sundkvist (eds), 117–133. Stockholm: Stockholm University Press. Stamboltzis, Aglaia & Pumfrey, Peter. 2000. Reading across genres: A review of literature. Support for Learning 15: 58–61. Swales, John. 1990. Genre Analysis: English in Academic and Research Settings [Cambridge Applied Linguistics Series]. Cambridge: CUP. Taylor, Chris. 2015. How Star Wars Conquered the Universe: The Past, Present, and Future of a Multibillion Dollar Franchise, revised and expanded. New York NY: Basic Books. Todorov, Tzvetan. 1990. Genres in Discourse. Cambridge: CUP. Tognini-Bonelli, Elena. 2010. Theoretical overview of the evolution of corpus linguistics. In The Routledge Handbook of Corpus Linguistics, Anne O’Keeffe & Michael McCarthy (eds), 14–27. London: Routledge.
141
Modeling fine-grained sociolinguistic variation The promises and pitfalls of Twitter corpora and neural word embeddings Filip Miletić,1,2 Anne Przewozny-Desriaux1 & Ludovic Tanguy1 1
CLLE, CNRS & Université Toulouse - Jean Jaurès | 2 IMS, Universität Stuttgart
This chapter examines the use of recent data sources and computational methods to study fine-grained sociolinguistic phenomena. We deploy a custom-built corpus of tweets (Miletić et al. 2020) and neural word embeddings to investigate the use of contact-induced semantic shifts in Quebec English. Drawing on an analysis of 40 lexical items, we show that our approach is beneficial in facilitating manual inspection of vast amounts of data and establishing fine-grained patterns of language variation. While it is affected by a range of noise-related issues, which we describe in detail, coarse-grained annotation provides an efficient way of circumventing them. We use the results filtered in this way to conduct a quantitative analysis of sociolinguistic constraints on contact-induced semantic shifts, further confirming the relevance of our approach. Keywords: semantic shifts, language contact, Twitter corpora, word embeddings, large language models, Quebec English
1.
Introduction
In this chapter, we deploy a novel corpus-based approach to the study of a complex type of language variation. Our focus is on contact-induced semantic shifts in Quebec English, that is, preexisting English words used with a meaning typical of a similar French word. Consider the following example taken from a tweet posted by a speaker from Montreal: (1) I really want to go to an art museum or an art exposition. https://doi.org/10.1075/scl.118.09mil © 2024 John Benjamins Publishing Company
Modeling fine-grained sociolinguistic variation
Here, the word exposition refers to what is usually known as an art exhibition. This meaning is not conventionally used in English; it is instead associated with the homographic French word exposition. This phenomenon is explained by the local sociolinguistic context: Quebec is the only predominantly French-speaking Canadian province. As of 2021, 74.8% of its inhabitants – close to 6.3 million people – report that their mother tongue is French. Ten times fewer Quebecers – 7.6% of the population, or just under 640,000 individuals – are native speakers of English (Statistics Canada 2022). This constitutes an ongoing situation of language contact between the varieties of English and French spoken in Quebec. As we more extensively discuss in Section 2.1, its effects on the Quebec English lexicon are documented in the sociolinguistic and corpus linguistic literature, but these reports are often qualitative or anecdotal. We aim to complement them with a more comprehensive empirical description of contact-induced semantic shifts. From a methodological standpoint, we draw both on well-established principles of variationist sociolinguistics (Labov 1972; Tagliamonte 2006) and on new types of corpora and computational tools. Our analyses rely on a particular type of linguistic data – a large, custombuilt corpus of tweets – as well as a recent computational approach to modeling lexical semantic phenomena – neural word embeddings. Large-scale computational studies of language variation increasingly rely on both of these methodological choices (Nguyen 2021; Tahmasebi et al. 2021). Although they are routinely presented as promising alternatives to smaller-scale and often manual corpus analyses, they come with their own challenges, and their potential for linguistic description is not yet clear (Boleda 2020: 218). In applying them to a highly specific linguistic phenomenon, we aim to provide a more thorough understanding of their contributions, recurrent practical challenges, and potential solutions. More specifically, we analyze patterns of regional variation in our corpus of tweets; drawing on underlying demographic distinctions, we aim to isolate instances of contact-induced semantic shifts and further characterize their use. For a given lexical item, our method produces computational representations of its occurrences in the corpus and splits them into groups based on semantic similarity, prioritizing those that are expected to reflect the influence of French. Despite extensive data filtering and a carefully adapted implementation of a recent neural language model, our approach highlights not only contact-induced semantic shifts, but also a range of noise-related phenomena. While this means that it cannot provide reliable results in an unsupervised manner, we show that a coarse manual annotation – conducted on the automatically identified clusters of tweets rather than individual occurrences – provides an efficient way of eliminating false positives. Complementing an earlier analysis of the technical impact of
143
144
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
these issues (Miletić et al. 2021), we provide a more extensive, qualitative description of recurrent problems in the data. We further use the filtered quantitative information in a case study exploring the sociolinguistic constraints on contactinduced semantic shifts. The remainder of this chapter is organized as follows. We first introduce a summary of related work (Section 2) and a more detailed description of the data and methods we deployed (Section 3). We then present the key results of our analysis (Section 4) and conclude with a discussion and main takeaways (Section 5).
2.
Theoretical and methodological background
In this section, we contextualize our work with respect to research on semantic shifts in Quebec English, which represent our descriptive focus. We further discuss our key methodological choices, thus addressing the use of Twitter corpora and neural word embeddings.
2.1 Semantic shifts in Quebec English: The need for corpus studies The use of English in Quebec is influenced by ongoing contact with French, particularly on the lexical level. Evidence for this claim comes from sociolinguistic studies (e.g., Boberg 2005, 2010, 2012; Boberg & Hotton 2015; Chambers & Heisler 1999; McArthur 1989; Rouaud 2019), as well as corpus-based research conducted mainly on newspaper texts (e.g., Fee 1991, 2008; Grant-Russell 1999; Miletić 2019). The reported contact-related phenomena include loanwords (e.g., dépanneur ‘corner store’), loan translations (e.g., all-dressed ‘(pizza) with all the toppings’, cf. Fr. toute garnie), and semantic shifts (e.g., animator ‘group leader’, cf. Fr. animateur) (Boberg 2012: 501). The importance of these phenomena has been contested due to their overall rarity in recorded speech (Poplack et al. 2006), but they may carry high symbolic value (Boberg 2012: 495). Moreover, large-scale dialect surveys have investigated regional lexical preferences more systematically, reaffirming that “Montreal appears to be the most lexically distinct region in Canada” (Boberg 2005: 36). This is largely due to French lexical influence, with similar patterns reported elsewhere in Quebec (Boberg 2010: 170–188; Boberg & Hotton 2015). These observations, however, are mainly based on studies of French loanwords; descriptions of other types of contact-related lexical influence are considerably more limited. This is particularly the case for the previously mentioned issue of semantic shifts, on which we focus in this paper. We understand this
Modeling fine-grained sociolinguistic variation
phenomenon as the presence of a sense in a preexisting English word that is explained by the presence of the equivalent sense in a formally and/or semantically similar French word. We know from McArthur’s (1989) written survey that the acceptability of semantic shifts varies across lexical items; cases such as animator ‘group leader’ and collectivity ‘people as a whole, community’ enjoy broader support than library ‘bookstore’ and demand ‘to ask for something’ (McArthur 1989: 42).1 Corpus-based research has further described a range of linguistic mechanisms at play, including effects on the word’s connotation (e.g., functionary shifting from negative to neutral connotation) and degree of formality (e.g., ameliorate shifting from formal to neutral register; Fee 2008: 181). There is also some evidence that the use of semantic shifts is subject to regional variation (Fee 1991; Miletić 2019) and affected by the speakers’ degree of bilingualism (McArthur 1989). Most of these studies rely on traditional sociolinguistic methods, which involve recording the speech production of carefully sampled speakers using face-to-face interviews (Labov 1972; Tagliamonte 2006) or written questionnaires (Dollinger 2015). However, this provides insufficient data for a systematic analysis of lexical semantic phenomena in spontaneous communication. For instance, Rouaud (2019: 245) describes the use of campaign with the sense of the French lexical item campagne ‘countryside’, which occurs only once in her sociolinguistic interviews, precluding generalizations. Another line of work discussed above overcomes this issue of data sparsity by analyzing larger collections of written documents, but it still faces challenges inherent in manual semantic analyses of corpus data. This is a crucial but highly time-consuming process, aggravated by the fact that human annotators may struggle to determine the attested meaning of a lexical item (Miletić 2019). In summary, the existing studies provide dozens of valuable examples of semantic shifts from diverse data sources, but given the methodological challenges outlined above, they are generally qualitative and often anecdotal in nature. As a result, we do not have reliable quantitative estimates of the impact of specific factors on the use of semantic shifts, or of their diffusion in the local speech community. Other types of corpora and analytical approaches may therefore be better suited to this issue.
1. The examples are related to the French lexical items animateur, collectivité , librairie, and demander, respectively. Their meanings are affected by language contact to variable degrees; the most fine-grained distinction is related to demand being used in very general contexts, without the notion of authority with which it is usually associated in English.
145
146
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
2.2 Twitter-based corpora for language variation The development of social media has enabled large-scale analyses of language variation relying on publicly available posts from these websites. This is particularly true of Twitter, a social network created in 2006, where users can post 280-character messages known as tweets.2 In addition to their linguistic content, tweets carry metadata such as the automatically identified language of the message and the user’s geographic location at the time of posting (if the user chooses to include it). Each user has a profile page, where they can provide further information, such as a description and an additional, profile-level location. Users can interact with the content posted by others by liking it or replying to it, as well as reproduce it by retweeting it. They can form ties with other users by following them, and as a result regularly see their posts in their own timelines. The public availability of vast amounts of linguistic data, coupled with geographic information and the possibility to analyze patterns of interaction, is a key reason driving the use of Twitter-based corpora in studies of language variation. These kinds of studies frequently consist in analyzing regional patterns of language use reflected by geotagged posts, often in a bottom-up manner. As an example, Grieve et al. (2018) examine significant rises in frequency over time to semi-automatically identify lexical innovations in American Twitter (e.g., amirite ‘am I right?’). They then map the spread of these items across the United States, observing five distinct geographic patterns of diffusion. Other related studies have similarly focused on identifying regionally specific lexical items (Donoso & Sánchez 2017; Shoemark et al. 2017) or interpreting language variation in terms of demographic factors (Bamman et al. 2014; Jones 2015). The sheer amount of data available on Twitter is routinely presented as a key advantage compared to traditional sociolinguistic studies. However, this is counterbalanced by issues such as a lack of reliable demographic information – a mainstay of sociolinguistic research – as well as sources of bias inherent to the platform, affecting key information such as user location (Pavalanathan & Eisenstein 2015). Despite this, the patterns of regional variation derived from Twitter-based corpora broadly coincide with those obtained using traditional dialect surveys (Grieve et al. 2019). This suggests that Twitter corpora are a promising source of information in studies of language variation, especially where large amounts of data are required.
2. Twitter was rebranded as X in 2023. Since this occurred after the conclusion of our study, we refer to the website’s original name. Note moreover that the maximum length of tweets was 140 characters until November 2017.
Modeling fine-grained sociolinguistic variation
2.3 Vector space models for lexical semantic variation As previously suggested, the persistent challenges in systematic analyses of lexical semantic variation – in sociolinguistic studies in general (Durkin 2012) and in those on Quebec English in particular – can be attributed not only to the large amount of data required for meaningful quantitative estimates of lexical semantic phenomena, but also to practical difficulties in investigating those phenomena at scale. It may be possible to overcome these issues by automatizing key aspects of semantic analyses. One way to do so consists in using vector space models (VSMs), computational tools that represent a word’s meaning as a vector, essentially a list of numbers reflecting the word’s cooccurrence statistics in a corpus (Turney & Pantel 2010). These models are rooted in the principles of distributional semantics, and most importantly the assumption that words appearing in similar linguistic contexts have a similar meaning (Firth 1957; Harris 1954). Importantly, VSMs provide the ability to measure the distance between two vectors, which is taken to reflect the difference in meaning of the words represented by them; this in turn facilitates systematic corpus-based studies of lexical semantics (Boleda 2020). While different implementations exist, we focus on a recent approach based on deep neural networks pretrained on vast amounts of generic text data, the foremost among them being BERT (Bidirectional Encoder Representations from Transformers) (Devlin et al. 2019). Although originally designed for more complex NLP tasks, they also provide an efficient way of obtaining vector representations of word meanings, also known as word embeddings. They produce a contextually-informed vector for each occurrence of a given word, allowing for analyses of phenomena such as polysemy. Methods such as these constitute the cornerstone of recent computational approaches to semantic change. Starting from a diachronic corpus, different VSMs have been used to quantify the change in meaning of all words in a corpus (or any subset of them) over time (see Tahmasebi et al. 2021 for a detailed overview). This includes representations produced by BERT and related approaches, which generally estimate semantic change by analyzing differences in sense usage across time (Giulianelli et al. 2020; Martinc et al. 2020; Montariol et al. 2021). A variety of VSMs have been used to analyze meaning variation in different types of synchronic data (Del Tredici & Fernández 2017; Schlechtweg et al. 2019) and from a contact-linguistic perspective (Takamura et al. 2017; Uban et al. 2019). However, most of these studies focus on computational issues. Existing descriptive applications include assessing longstanding hypotheses on semantic change (Xu & Kemp 2015) and facilitating corpus analyses by domain experts (De Pascale 2019; Rodda et al. 2017). This type of work is indicative of the descriptive potential of these methods, but their usability in linguistic research is yet to be fully demonstrated (Boleda 2020: 218). And while recent quantitative evaluations have identified the best-performing computational setups (e.g., Schlechtweg et al. 2020),
147
148
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
practical issues such as the impact of noise in the data have not been extensively addressed. Echoing this situation, our previous work has shown that standard semantic change models can be applied to contact-induced semantic shifts in Quebec English, but that the resulting corpus-based analyses nevertheless entail important practical challenges (Miletić et al. 2021). In this chapter, we examine the issues we encountered in more detail and discuss potential ways to overcome them.
3.
Data and method
Our approach relies on contrasting synchronic data from different Canadian regions, under the assumption that linguistic behaviors that are specific to Quebec but absent from areas where the use of French is limited, are likely to reflect the influence of language contact. This is inspired by the comparative sociolinguistic approach (Tagliamonte 2002), where differences in speech across communities are taken to reflect the sociodemographic differences in their composition. This section presents the main data and methodological choices on which we relied: a corpus of tweets posted in Canada; a curated list of contact-induced semantic shifts; an implementation of a neural word embedding model; and a manual annotation procedure. A high-level overview of our method is presented in Figure 1.
Figure 1. Method overview. For each semantic shift, examples are extracted from a regionally stratified corpus, modeled in the form of vector representations, clustered based on semantic similarity, and annotated for contact influence
Modeling fine-grained sociolinguistic variation
3.1 A corpus of tweets We use a previously created corpus of Canadian English tweets published by users from Montreal, Toronto, and Vancouver (Miletić et al. 2020). Unlike other corpora of Canadian English, it is both sufficiently large for data-intensive methods such as vector space models (VSMs), and it contains information on the regional origin and authors of individual tweets. The three cities are roughly comparable in terms of broad demographic properties which may impact sociolinguistic dynamics, such as the size and diversity of the population. A key differentiator is the proportion of French speakers: they constitute a strong majority in Montreal, and a fraction of the population in the other two, predominantly Englishspeaking, cities. We therefore expect linguistic patterns specific to Montreal to be related to the influence of French. However, we cannot be certain of that, so we also rely on the structure of our corpus to distinguish contact phenomena from unrelated types of variation, such as regional trends that do not originate in Montreal (which can be identified using the two control regions) and strong idiolectal preferences (which we can isolate by tracing individual speakers). The data were collected from January to November 2019. We initially used Twitter’s Search API to look up tweets tagged as written in English and associated with the geographic area of one of the target cities. The users identified in this way were narrowed down to those whose free-text profile location strictly corresponded to one of the three target cities. We then crawled their profiles, collecting up to 3,200 most recent tweets per user; this allowed us to increase the amount of data and obtain basic sociolinguistic information. For instance, we stored the distribution of language tags in the users’ tweet production and subsequently used it as a rough estimate of their degree of bilingualism. Finally, we only retained the tweets tagged by Twitter as written in English, and we automatically removed near-duplicates posted by individual users. Exploratory analyses have shown that the retained data are both specific to the target cities and comparable across them (for more details, see Miletić et al. 2020).3 In addition to the preprocessing decisions applied to the original corpus, we introduced additional filtering for the experiments presented in this paper. First, we removed the content posted before 2016 in order to limit the likelihood of picking up diachronic effects; the tweets in the original corpus date back to 2006. In determining the cut-off point, our aim was to find a reasonable tradeoff between a reduction in time span and the remaining amount of data. We then 3. In accordance with Twitter’s developer terms, the corpus is released in the form of tweet IDs, which can be used with off-the-shelf software to collect the underlying data: 〈http://redac .univ-tlse2.fr/corpus/canen.html〉
149
150
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
excluded all users with fewer than 10 tweets in the corpus. A maximum of 1,000 tweets per user were retained, with random subsampling performed where this was exceeded. This ensured that the corpus was not dominated by very few highly active individuals: an average of 229 tweets were retained per user, with the top 1% of users accounting for 4% of tweets. The corpus was tokenized and POS-tagged using twokenize (Gimpel et al. 2011; Owoputi et al. 2013) and lemmatized using the NLTK WordNet lemmatizer (Bird et al. 2009). The structure of the final corpus is presented in Table 1. Table 1. Corpus structure Subcorpus
Users
Tweets
Tokens
Montreal
54,726
11,318,184
193,228,246
Toronto
51,245
12,465,659
222,508,471
Vancouver
47,697
11,381,080
213,200,523
Total
153,668
35,164,923
628,937,240
3.2 A set of semantic shifts in Quebec English In order to examine a manageable number of semantic shifts in greater detail, we started from a previously established list of lexical items whose meaning is subject to contact-related influence in Quebec English and which are attested with that meaning in our data. We specifically used an 80-item test set constructed for a technical evaluation of semantic change models on our data (Miletić et al. 2021).4 Since we investigated the models’ ability to distinguish words that do or do not undergo change, half of the items correspond to contact-induced semantic shifts, and the remaining half to stable words; given the scope of this paper, we focus on the former. In identifying the semantic shifts, we relied on descriptions provided in the literature on Quebec English (Boberg 2012; Fee 1991, 2008; Grant 2010a; Josselin 2001; McArthur 1989; Rouaud 2019) as well as initial manual exploration of the Twitter corpus. In order to ensure the availability of sufficient data for a meaningful quantitative interpretation, we only retained the items occurring at least 100 times in each subcorpus. A concordance-based analysis was then used to determine if the items presented at least one contact-related occurrence in the Montreal subcorpus; those that did not were excluded. In establishing the potential contact-related use, the existing descriptions and corpus-based observations were
4. The full test set is available at 〈http://redac.univ-tlse2.fr/misc/canenTestset.html〉
Modeling fine-grained sociolinguistic variation
complemented with lexicographic evidence.5 This process resulted in a final list of 40 semantic shifts retained after all filtering steps; their mean frequency in the entire corpus is 5,268 (min = 345, max = 97,188). The items are summarized in Table 2. Table 2. Set of semantic shifts with posited contact-related senses and sociolinguistic sources on which we principally relied to identify them. Items without provided sources were identified through manual corpus exploration Item
Shifted sense
Sources Item
Shifted sense
Sources
affirmation
claim, statement
deliberate btw. options
–
atmosphere, vibe
F1
hesitate
ambiance
laureate
winner
–
animator
activity/team leader
B F1 G1 local (n) M
room, site, premises
M
availability
(pl.) available times
–
manifestation protest, demonstration
boutique
shop, store
–
merit
chalet
summer cottage
activist, campaigner
circulation
traffic
B F2 G1 militant
nomination
appointment to a role
coordinate
(pl.) contact details
occasion
chance, opportunity
–
deceive
disappoint
G1
B
disappointment
F2 M R pass (v)
stop by
deception
–
permit
driver’s license
–
definitively
definitely, certainly
–
population
people, general public
deputy
member of parliament
R
portable
cell phone, laptop
F1 F2 G1 R
dossier
issue, portfolio family circle, friends
F2 G1 R proposition
suggestion, proposal
entourage
–
prudent
careful
–
exchange (v) talk with someone
–
remark (v)
notice
M
exploration
study of sth. little known
–
reparation
repairs (of a device)
–
exposition
art exhibition
summarize
MR
training, course
F1
resume
formation
BMR
souvenir
memory
MR
formidable
great, terrific
–
terrace
patio, eating area
grave (adj)
highly important
–
trio
sandwich-fries-drink
F2 R
–
F1
deserve, be worthy of
G1 –
G1
G1
G2 –
B
Sources: B = Boberg (2012); F1 = Fee (1991); F2 = Fee (2008); G1 = Grant (2010a); G2 = Grant (2010b); M = McArthur (1989); R = Rouaud (2019).
The meaning of these items, as attested in a range of occurrences from the corpus of tweets, was then computationally modeled. Our aim was to automat5. Canadian Oxford Dictionary (Barber 2004); Dictionary of Canadianisms on Historical Principles (Dollinger & Fee 2017); Trésor de la langue française informatisé (Dendien & Pierrel 2003); Usito (Cajolet-Laganière et al. 2014).
151
152
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
ically identify occurrences used in similar contexts and quantify sense distributions, focusing on both regional and user-level patterns.
3.3 Neural word embeddings For each of the 40 lexical items from Section 3.2, we first produced word embeddings for their individual occurrences. Each corresponds to a slightly different vector, which incorporates general distributional information captured during model pretraining and is further informed by the target item’s immediate linguistic context. We then used these representations to automatically group the occurrences into clusters, which were expected to reflect similar contexts (and thereby similar uses of the target lexical item). This allows for a more efficient subsequent analysis of the full range of uses exhibited by a lexical item: for instance, the fact that similar occurrences are grouped together means that it is not necessary to disambiguate them one at a time. Word embeddings were produced using the previously discussed BERT model, and specifically the Hugging Face implementation (Wolf et al. 2020) of bert-base-uncased, a 12-layer, 768-dimension version pretrained on generic English data.6 Models such as this are often adapted to specific applications using finetuning, that is, partial retraining on a specific task. Similarly to some existing work (e.g., Giulianelli et al. 2020), we did not perform fine-tuning given the assumption that word senses are reflected by differences in immediate linguistic context, which the pretrained model should be able to capture. For each analyzed lexical item, we extracted the tweets in which it appears in all three regional subcorpora. In order to limit processing and memory requirements, we retained no more than 1,000 total occurrences per word and used a random sample for more frequent items. We fed each tweet in its raw text form as a single sequence into BERT, which then produced context-informed vectors for each token in the tweet. The model outputs multiple vector representations per token, each corresponding to a different hidden layer in the neural network architecture. Similarly to other recent studies (e.g., Laicher et al. 2021), we averaged over the last four hidden layers to obtain a single token-level vector. BERT’s tokenizer splits some words into subparts (also known as wordpieces) with separate representations; when this occurred, we averaged over the subparts to produce a
6. A version of the model specifically adapted to tweet processing (BERTweet; Nguyen et al. 2020) was released after our experiments were implemented, but is relevant for future work using the same type of data.
Modeling fine-grained sociolinguistic variation
single vector. The embedding computation procedure is schematically presented in Figure 2.
Figure 2. Embedding computation for an occurrence of entourage. Each computational layer (grey bars) outputs contextualized embeddings for each token in the sequence7
3.4 Clustering and annotating the uses of a lexical item Similar uses of a lexical item were automatically identified by clustering its tokenlevel vectors using affinity propagation, an algorithm which performed well in other semantic change studies (e.g., Martinc et al. 2020). It does not require that the number of clusters be specified beforehand, and it produces clusters of variable size, making it well-suited for studying sense distributions. We used the scikitlearn implementation (Pedregosa et al. 2011) with default parameters, which is based on the negative squared Euclidean distance. Data from the three regional subcorpora were clustered at the same time, meaning that a single cluster may contain tweets from all three cities. This allows us to examine how each cluster’s occurrences are distributed across the regions. In analyzing the output of the analysis, we considered the clusters containing at least five tweets, and retained them if more than half of the tweets were from the Montreal subcorpus. This is because of the focus on the uses which are clearly 7. Figure inspired by Jay Alammar’s illustrations available at 〈http://jalammar.github.io /illustrated-bert/〉
153
154
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
more frequent in Montreal than elsewhere, but which may occasionally appear in other regions. Up to 10 such clusters were retained for each lexical item, starting with those with the highest proportion of Montreal tweets. The data for the 40 retained lexical items were then manually annotated at the level of clusters. We used binary labels and established if a cluster presented a contact-related sense based on the majority usage in it. More specifically, a target item’s use in a cluster was annotated as contactrelated if it was regionally specific to Montreal and potentially explained by the influence of a formally and/or semantically related French word. This determination relied on the same evidence used to select the target set of semantic shifts, that is, previous sociolinguistic studies and lexicographic sources (see Section 3.2). Recurrent phenomena that were not annotated as contact-related included a range of noise-related issues; these will be discussed in more detail below. A 15-word sample (91 clusters) was annotated by two annotators in order to test the reliability of the general procedure, obtaining a reasonably high interannotator agreement (Cohen’s kappa coefficient of 0.55). On average, 8 clusters per word (min = 3, max = 10) were retained for annotation. The mean number of tweets per cluster stands at 13 (averaged over the means for individual lexical items; min = 8, max = 20). As shown by the examples discussed below, the clusters are largely homogeneous; although some are occasionally difficult to interpret, this is overall rare. Our analysis also allows for a degree of uncertainty, as the annotation targets the predominant use in a cluster. The utility of this approach is confirmed by the fact that it led to the identification of at least one contact-related cluster for each of the 40 target items. From a practical standpoint, using cluster-level annotations was an order of magnitude faster than analyzing individual tweets. This is due to the lower number of required decisions and the comparative ease in determining the meaning of a larger number of similar examples appearing together.
4. Results This section discusses the results derived from the annotated data. It first presents a general overview of cluster types across lexical items; it then illustrates a range of true and false positives observed in the data; and it concludes with a case study examining the link of contact-induced semantic shifts with bilingualism.
Modeling fine-grained sociolinguistic variation
4.1 An overview of regionally specific clusters A global overview of the analysis (Figure 3) outlines the distribution of annotated tweets for the 40 target lexical items based on the annotated usage types. Regionally specific clusters may capture the effect of language contact, but this is not systematic: contact-related use was observed in at least one cluster for each item, but it was prevalent in a minority of cases. The proportion of contact-related tweets ranges from seemingly clear-cut cases such as definitively ‘definitely’ (100%) and terrace ‘outside eating area; patio’ (90%) on one end of the spectrum, to animator ‘group leader’ (6%) and reparation ‘repairs’ (4%) on the other; the median value is 26%. In other words, regional variation appears to be helpful in guiding the analysis of contact-induced phenomena, but it is not sufficient in isolating linguistic behaviors of interest.
Figure 3. Distribution of annotated tweets from Montreal-specific clusters across the binary labels. For example, terrace had the contact-related meaning ‘patio’ in 90% of tweets (blue bar), and other meanings in 10% of tweets (red bar)
Furthermore, a wide variety of linguistic phenomena is subsumed under the binary distinction in the overview plot. We now provide a more detailed, qualitative analysis of both contact-related and non-contact-related uses, which can respectively be viewed as true and false positives output by our computational system.
4.2 Types of variation captured by the analysis This section discusses examples of tweets extracted from our corpus using the clustering analysis described above. Sample clusters of tweets are presented in the keyword-in-context format, for ease of reading as well as to illustrate the effect of this approach on manual perusal of corpus data, as observed during the manual annotation. Each sample cluster contains three representative tweets published in Montreal and occurring in a single original cluster output by our analysis. Further information on the size and regional composition of the clusters is also provided.
155
156
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
In order to protect the privacy of tweet authors, we only reproduce textual content without any metadata. For the same reason, usernames, hashtags, URLs, and names of individuals are redacted from the tweets, except for widely known public figures or if necessary to interpret the meaning of the tweet. 4.2.1 True positives We begin by examining positive contributions of our computational system, focusing on the help it provided in distinguishing between conventional and contact-related uses based on documented patterns from the corpus. This was beneficial across different semantic mechanisms and degrees of granularity of contact-related influence.
A clear-cut distinction Perhaps the prototypical mechanism underlying contact-induced semantic shifts involves using an English lexical item to denote a referent conventionally designated by a formally similar French lexical item. One such example is manifestation, which is generally used to signify ‘a display of the existence of something’, but is also attested in Quebec English with the sense of ‘protest, demonstration’, typical of the homographic French lexical item manifestation. This sense is absent from the Canadian Oxford Dictionary (COD), but it is anecdotally reported by Grant (2010a: 187). It is also recorded in the Oxford English Dictionary (OED), but the most recent example (1978) comes from New Brunswick, an officially bilingual Canadian province bordering Quebec, and contains metalinguistic commentary supporting a link with language contact. Our corpus-based analysis, partly presented in Table 3, provides empirical evidence of the ongoing use of both senses; the French-related one is strongly regionally-specific to Montreal. The first cluster in the table corresponds to the contact-related sense of ‘demonstration’, which can be interpreted based on spatio-temporal references as well as nearsynonyms (e.g., walk); the second cluster corresponds to the conventional sense of ‘existence’, reflected by the medical topic. Table 3. Sample concordance lines for manifestation. Contact-related use (top): 9 out of 9 tweets (100%) from Montreal; conventional use (bottom): 8 out of 13 tweets (62%) from Montreal Montreal’s manifestation in progress at the immigration office Pro refugee manifestation at the Big O . Quebec’s history . This walk is the biggest manifestation for this week . And 52 more locating colonial trauma in genetic manifestation of health issues , all the funding goes into ask yourself how they might affect clinical manifestation of disease . is an emergency problem it’s local manifestation for a systemic issue
Modeling fine-grained sociolinguistic variation
A subtler distinction A more nuanced type of contact-related semantic influence is illustrated by the case of entourage. In English, this lexical item generally has the highly specific sense of ‘retinue, people surrounding an important person’, whereas in French it is also used with the more general sense of ‘family circle’ or ‘group of friends’. This sense is not recorded in either the COD or the OED, and we are unaware of any sociolinguistic descriptions of it. However, our corpus data point to the existence of both senses in Quebec English (Table 4). They additionally illustrate the fact that the contact-related sense involves a generalization of the preexisting English sense. This implies a partial overlap, hence a less striking difference between the two, which is often only apparent from contextual cues. The first cluster in the table captures the ‘family circle’ sense; this is suggested by the possessives associated with entourage, which point to the speaker or their interlocutor. The second cluster corresponds to the ‘retinue’ sense, as shown by third person references as well as background knowledge on the people whose entourage is being discussed (e.g., sports players and politicians). Fine-grained distinctions such as these underscore the complexity of the phenomenon under study and the need for linguistic expertise; the fact that they can be clearly established based on the output of our computational system confirms its value in guiding empirically grounded analyses. Table 4. Sample concordance lines for entourage. Contact-related use (top): 12 out of 13 tweets (92%) from Montreal; conventional use (bottom): 17 out of 21 tweets (81%) from Montreal learn it on your own and have people in your entourage help you practice . Personally , I tried taking have more than 30 people (coworkers) in my entourage who are new fan of Harry but they don’t know hateful shit . But when people from your entourage excuse someone’s homophobic actions ; it @MariaSharapova where are her entourage? Managers …? Don’t they check for her ? Damage control brought to you by POTUS entourage. He’s been golfing and hamming it up with I really feel like somebody from his entourage told him “oh yes , best we have ever seen ”
4.2.2 False positives We have so far focused on the informativeness of our semi-automated analysis in understanding often fine-grained patterns of contact-related language use, but this process was complicated by different types of false positives. This section provides a detailed analysis of the most frequent patterns that we encountered. We distinguish between the following types of locally-specific word usage which do not constitute contact-induced semantic shifts:
157
158
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
– – – –
cultural effects, where word usage is related to the local cultural context of Montreal; the use of common nouns as proper names denoting locally specific referents; French homographs of English words, attested in codeswitched tweets; structural patterns, such as the position of the target item across tweets, which accidentally affect model performance.
For each category that we discuss, sample clusters of tweets are provided in order to illustrate the contrast between the contact-related use – the target of our analysis – and the noise that we identified along the way.
Cultural effects The regionally specific character of some clusters output by our analysis is not related to the use of the target lexical items with a French-related sense, but rather to the local cultural context of Montreal. Take for example formation, whose English senses include ‘the action or process of forming’ and ‘arrangement or disposition’. In our data, it is also attested with the sense of ‘course, training program’, typical of the French homograph formation. This sense is absent from the COD and the OED, but its existence is noted in the sociolinguistic literature (McArthur 1989: 57; Boberg 2012: 497). As shown in Table 5, our computational system enabled us to identify corpus-based evidence of its use (top cluster). But it also accidentally highlighted its involvement in topical variation (bottom cluster), related to the widespread use of formation ‘disposition of players in a sports team’. The regional specificity of this cluster is not explained by language contact; the one apparent characteristic shared by these occurrences is discussion of local sports teams. This is evidenced by the names mentioned in the tweets: at the time of posting, Dominic Oduro and Donny Toia were players for the Montreal Impact soccer team; Mauro Biello was manager of the same team; and Carey Price and Al Montoya were goaltenders for the Montreal Canadiens hockey team. This is also a striking illustration of a more general practical challenge in analyzing Twitter data: detailed background information on a variety of topics – in this case, sports – is often required to fully grasp the local significance of the attested language use. A different but related effect was observed in the case of animator. This lexical item is generally used with the sense of ‘creator of animated films’, whereas the formally similar French equivalent animateur also includes the sense ‘group leader; organizer; facilitator’. The use of animator with the former senses is attested in the sociolinguistic literature (McArthur 1989: 53; Grant 2010a: 187; Boberg 2012: 497). It is also noted in dictionaries, including the OED, and is marked as being related to French (COD) and specific to Quebec and New Brunswick (Dictionary of Canadianisms on Historical Principles). Our system found some occurrences of
Modeling fine-grained sociolinguistic variation
Table 5. Sample concordance lines for formation. Contact-related use (top): 5 out of 7 tweets (71%) from Montreal; cultural effects (bottom): 9 out of 14 tweets (64%) from Montreal , the company I work at gave us a quick formation and that was it ! Little sister is taking on a formation to be a game tester and I am going to univ to realize that ??? Well i have internet on my formation but its a shitty internet .. so i probably not be Oduro isn’t strong enough defensively for this formation. Toia not strong enough offensively , but room for only 1 deep ball retriever . What formation should Biello use with a full lineup ? 2/2 The goalie formation is solid . even if Price gets injured , Montoya
the contact-related use, but these were limited to a single cluster; the remaining eight regionally-specific clusters reflected the conventional sense (Table 6). This is likely due to Montreal’s status as a global center for the animation industry, which would explain a larger number of conversations on this topic compared to the two control cities, making this use more prominent for our computational system. Table 6. Sample concordance lines for animator. Contact-related use (top): 4 out of 6 tweets (67%) from Montreal; cultural effects (bottom): 10 out of 14 tweets (71%) from Montreal Mr. , Spiritual and Community Animator, is organizing a Costa Rica trip for … Ms. , Spiritual Animator, has been busy with the annual #PoppyDrive created by , Spiritual Community Animator. Firefighter , Fire House Station 54 for an animation tech director , technical animator, and a VFX artist . I don’t know what any of SO many different jobs in animation , not just animator, and they’re all essential to every single prod . I’m a graphic and web designer turned 3d animator an motion designer . What I’d really like to do
Proper names The regional specificity of some clusters is explained by the target lexical item being used as a proper name, generally to denote a regionally-specific referent. Take for example deception, which in English refers to ‘the action of misleading someone’, but whose French homograph also means ‘disappointment’. This use is not recorded in the OED or the COD, nor is it described in the sociolinguistic literature we reviewed.8 Potential influence of French is, however, attested in our data (Table 7). But in addition to those reflecting the contact-related sense, some clusters specific to Montreal involve occurrences of the target lexical item in the 8. This pattern is however anecdotally reported for the verb deceive ‘mislead’ and the French equivalent décevoir ‘disappoint’ (McArthur 1989: 17), as well as for the corresponding deverbal adjective deceived ‘disappointed’ (Rouaud 2019: 167).
159
160
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
phrase Deception Bay, the name of a song by the Montreal band Milk & Bone. This is explained by the origin of the band in question and is entirely unrelated to language contact. Table 7. Sample concordance lines for deception. Contact-related use (top): 9 out of 10 tweets (90%) from Montreal; proper names (bottom): 7 out of 7 tweets (100%) from Montreal be my year , but so far it’s been nothing but deceptions and heartbreaks . That won’t stop me from Great expectations , few deceptions and stunning debuts make a unique I had some very bad deceptions lately … * coughs * The Technomancer …. Deception Bay , the title track from @milknbone’s The new song Deception Bay , from Milk & Bone’s second album , is Deception Bay on repeat !! Can’t wait for the whole
French homographs in codeswitched tweets The performance of our computational system is occasionally affected by crosslingual homographs of the target lexical items. They are generally used in a span of French text within a codeswitched tweet where most tokens are in English; this explains why the tweets were tagged as written in English and retained in the corpus. Codeswitching is overall rare in our corpus, but its relative frequency is considerably higher in Montreal, as can be expected given the prevalence of bilingual speakers in the city. The practical implications are illustrated by the case of souvenir. Our analysis focused on the conventional English sense ‘keepsake, memento’ and the potential presence of the more abstract sense ‘memory’, typical of the formally identical French equivalent. The use of the English lexical item with the French-associated sense is not recorded in the COD. It is however attested in the OED, though only as “chiefly literary”, as well as in the sociolinguistic literature (McArthur 1989: 69; Rouaud 2019: 167). It is also present in our data, providing a corpus-based contribution to the existing descriptions (Table 8). But in addition to this positive result, three out of the ten regionally-specific clusters are entirely composed of codeswitched tweets containing the French homograph of the target lexical item. These instances do not correspond to our definition of semantic shifts – we are interested in English, not French, lexical items – and as such represent noise detrimental to our analysis. Structural patterns affecting model performance A final recurrent issue is that of clusters where tweets appear to be grouped together based solely on structural regularities. This was observed in the case of trio, which conventionally means ‘a group of three’, but is also used with the
Modeling fine-grained sociolinguistic variation
Table 8. Sample concordance lines for souvenir. Contact-related use (top): 12 out of 12 tweets (100%) from Montreal; French homographs (bottom): 12 out of 16 tweets (75%) from Montreal which I had kept no memories . The only fond souvenir which still haunted me , her hand on my jaw . ago Winter 2003 with my son Raphaël . Great souvenir . Let’s…
🤗❤️ year ago . Wonderful memories ! Quel beau souvenir ! 💋💋
Montreal . We had VIP tickets . An amazing souvenir
Francois Fournier a partagé un souvenir . 1 h · 7 years ago , i played my third gig with
Old memories . Very old . Oh , les vieux souvenirs!
sense of its Quebec French homograph, denoting a ‘sandwich-fries-soda special, combo’. This specific use is not attested in the lexicographic sources we consulted, but it is described in the sociolinguistic literature (Boberg 2012: 500). Our analysis provided corpus-based evidence for this use; however, the second cluster shown in Table 9 illustrates the problem of uninformative results mentioned above. Here, the only characteristic common to the tweets is the fact that the target lexical item occurs at the end of the tweet, whose content is otherwise ambiguous. This is potentially explained by underlying issues with the way in which BERT calculates some vector representations. This may be exacerbated by short sequences of text, where positional information may carry excessive influence. Table 9. Sample concordance lines for trio. Contact-related use (top): 8 out of 8 tweets (100%) from Montreal; structural patterns (bottom): 7 out of 10 tweets (70%) from Montreal I could legit eat 3 Big Mac trios right now I’m so hungry
😭
Customer : can i have a number 3 trio ? Cashier who used to be a middle school I rarely order fries anymore , let alone a trio . If I did , however , I usually go fries first . (But oh you know it man ! This is the new trio got mr beefcake yay now I have the full ossan trio 2 Years ago , still the same goofy trio
The examples discussed so far illustrate the variety of patterns of contactrelated use attested in our data, for which we obtained valuable qualitative evidence thanks to the computational analysis we implemented. However, a range of regionally-specific uses were explained by different types of noise rather than, as we had anticipated, language contact. The system in its current implementation cannot reliably identify contact-related use in an unsupervised way. However, the coarse annotation we conducted facilitates manual corpus analysis, as shown throughout this section; it validates the existence of contact-related uses; and it
161
162
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
helps exclude the patterns that are related to noise. We now draw on the resulting data to present a broader, quantitative overview of contact-related language use in the corpus.
4.3 Deploying coarsely annotated data for linguistic description The structure of the clusters output by our analysis shows that lexical items differ in terms of the diffusion of contact-related usage (how many tweets are related to contact, out of all those retained in the regionally-specific clusters) as well as its regional specificity (how many tweets in contact-related clusters come from Montreal). These patterns may be indicative of different degrees and factors of diffusion of semantic shifts within the local speech community. To explore the descriptive relevance of this information, we calculated scores reflecting the two points raised above for each of the 40 manually annotated lexical items: a diffusion score, corresponding to the proportion of tweets tagged as contact-related, out of all manually annotated tweets; and a regionality score, corresponding to the proportion of tweets posted in Montreal, out of all tweets tagged as contact-related. In order to explore the potential impact of the degree of bilingualism on the use of semantic shifts, for each lexical item we also calculated a bilingualism score, corresponding to the mean proportion of tweets in English (out of tweets in English in French) posted by users who used the contact-related sense in the clusters tagged as such. It ranges from 0 for users tweeting only in French to 1 for users tweeting only in English, with intermediate values indicating a production of tweets in both languages. We first checked the relationship between the three scores by calculating Spearman’s rank correlation coefficient. The diffusion score is uncorrelated with both the regionality score (ρ = −0.13, p = 0.42) and the bilingualism score (ρ = 0.02, p = 0.90). However, the regionality and bilingualism scores exhibit a moderate negative correlation (ρ = −0.53, p < 0.001); this link is explored in more detail in Figure 4. The plotted results indicate that contact-related semantic shifts which are more regionally-specific (i.e., attested in Montreal to a higher extent) are also more directly related to the effects of bilingualism (i.e., a lower proportion of English, and hence a higher proportion of French, tweets). A typical example (bottom right) is the case of circulation, attested in the Quebec English data with the sense of ‘traffic’, which is associated with the corresponding French homograph. All of the tweets from clusters tagged as contact-related come from Montreal; moreover, the mean proportion of English tweets stands at 0.75 per user. This may appear to be a relatively high value, but it is in fact just above the 10th percentile for all users in the corpus (0.73); at least within this dataset, this is suggestive of a comparatively
Modeling fine-grained sociolinguistic variation
Figure 4. Regression plot showing the relationship between the regionality score (x-axis) and the bilingualism score (y-axis). Dotted lines show the 10th and 20th percentile for the bilingualism score, for all users in the corpus
important influence of bilingualism. Patterns at the other end of the spectrum (upper left) are illustrated by the verb remark; we focused on the sense ‘notice’, with which the French verb remarquer is widely used. It is less regionally-specific (62% of contact-related tweets posted in Montreal) and less strongly associated with bilingualism (higher mean proportion of English tweets per user, at 0.99). Unlike in the previous example, however, the contact-related sense is attested in dictionaries, but the OED marks it as rare in some syntactic contexts. While it is likely accessible to most English speakers, cross-linguistic influence might facilitate its wider use; this scenario is consistent with our data. It is also relevant to look at the outliers from the general trend. For instance, in the previously mentioned case of trio ‘sandwich-fries-soda special, combo’ (upper right in the plot above) all contact-related tweets similarly come from Montreal. However, the mean proportion of English tweets is higher, at 0.99 per user. This is indicative of a use which is regionally-specific, but is widespread in the local linguistic community, including among monolingual speakers. This is further sup-
163
164
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
ported by existing descriptions which have shown it to be typical of the speech of native English-speaking Quebecers (Boberg 2005: 36; Boberg & Hotton 2015: 307). These observations indicate that, barring some exceptions, the more region specific the contact-related use is, the more strongly it is associated with use of French. Once again, it is important to note that the manual annotation was conducted on the level of clusters, rather than individual tweets, meaning that some non-contact-related occurrences may have been included in the counts. Moreover, the information on the use of French has the benefit of being empirically grounded in the attested use of languages by individual Twitter users, but it is only a very rough approximation of their linguistic profiles; for instance, there is no reliable way to determine their native language. That said, our analysis identified clear trends regarding the use of semantic shifts based on a large amount of data, further confirming the potential that corpus-based analyses have in understanding the patterns behind complex linguistic behaviors. It also constituted the basis of a face-to-face sociolinguistic survey we conducted in January 2022, whose initial results confirm the overall relevance of our approach, but also highlight the distinct – and complementary – nature of corpus-based and in-person estimates of semantic change (Miletić et al. 2023).
5.
Discussion and conclusion
We have presented an analysis of a fine-grained sociolinguistic phenomenon – contact-induced semantic shifts in Quebec English – using a large, custom-built corpus of tweets and a recent pretrained language model relying on a deep neural network architecture. This approach has paved the way for a more detailed account of previously reported semantic shifts, contributing extensive empirical evidence where original descriptions often consisted in a single anecdotal mention of a lexical item of interest; our approach was also beneficial in more comprehensively characterizing previously undescribed semantic shifts, initially observed in isolated tweets. The computational tools we used facilitated manual inspection of vast amounts of data, directing our attention to the most relevant subsets of occurrences; they also enabled broad quantitative estimates of the use of semantic shifts, highlighting possible interpretations and informing the design of subsequent studies. The results have also broadly confirmed our highlevel assumption that regional variation in synchrony can be used as a proxy for detecting contact-induced phenomena. More generally, the overall setup – data extraction, clustering based on semantic similarity, and analysis of the distribution of occurrences over an explanatory factor – can be generalized to other descriptive issues.
Modeling fine-grained sociolinguistic variation
However, we cannot gloss over the fact that our computational system provided actionable results only once it was complemented with extensive manual analyses. The challenges that we encountered are related to several distinct issues: (i) a strong assumption on regional variation underpinning the methodological design – while some language use specific to Montreal is related to language contact, not all is; (ii) inherent limitations of the methods we used, with BERT occasionally capturing phenomena unrelated to lexical semantics; (iii) inherent limitations of the data we used, with a carefully filtered Twitter corpus representing an improvement on highly generic datasets, but still suffering from the 280-character limit and the limited ability to validate user descriptions, among other issues; (iv) the complexity of the phenomenon under study, which often involves very subtle – but nevertheless perceptible and socially meaningful – differences in language use. Some of the described false positives, such as French codeswitching and referents typical of Montreal, are specific to our corpus; however, they echo the observation that semantic change models capture different types of variation in word usage, also raised in other recent studies (Giulianelli et al. 2020; Hengchen et al. 2021; Tahmasebi et al. 2021). Despite these challenges, data-intensive computational approaches to lexical semantic phenomena, and to language variation in general, have an important role to play in descriptive linguistic research. They can provide meaningful quantitative accounts of lexical phenomena, including for the whole vocabulary, based on data obtained in an unobtrusive way; this is clearly complementary to traditional sociolinguistic methods. While methods such as those we implemented still require adaptations to the task at hand as well as some manual analysis, they simplify the tasks required of the linguist. One example of this approach is our analysis based on coarse cluster-level annotations; its relevance is confirmed by the fact that, together with the results presented in Miletić et al. (2021), this line of work represents the first systematic corpus-based analysis of contact-induced semantic shifts in Quebec English. The initial results of a face-to-face sociolinguistic survey that this work enabled highlight a complementary relationship between computational and in-person descriptions (Miletić et al. 2023), while more extensive forthcoming analyses will shed further light on the real-life sociolinguistic dynamics at play.
Funding The computational analyses presented in this paper were carried out using the OSIRIM computing platform, administered by the IRIT research laboratory and supported by the CNRS, the Ré gion Occitanie, the French Government and the European Regional Development Fund (see 〈https://osirim.irit.fr〉).
165
166
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
References Bamman, David, Eisenstein, Jacob & Schnoebelen, Tyler. 2014. Gender identity and lexical variation in social media. Journal of Sociolinguistics 18(2): 135–160. Barber, Katherine (ed.). 2004. Canadian Oxford dictionary. Oxford: OUP. Bird, Steven, Loper, Edward & Klein, Ewan. 2009. Natural Language Processing with Python. Sebastopol CA: O’Reilly Media. Boberg, Charles. 2005. The North American Regional Vocabulary Survey: New variables and methods in the study of North American English. American Speech 80(1): 22–60. Boberg, Charles. 2010. The English Language in Canada: Status, History and Comparative Analysis. Cambridge: CUP. Boberg, Charles. 2012. English as a minority language in Quebec. World Englishes 31(4): 493–502. Boberg, Charles & Hotton, Jenna. 2015. English in the Gaspé region of Quebec. English WorldWide 36(3): 277–314. Boleda, Gemma. 2020. Distributional semantics and linguistic theory. Annual Review of Linguistics 6: 213–234. Cajolet-Laganière, Hélène, Martel, Pierre, Masson, Chantal-Édith & Mercier, Louis. 2014. Usito. (20 May 2024). Chambers, J. K. & Heisler, Troy. 1999. Dialect topography of Québec City English. Canadian Journal of Linguistics/Revue Canadienne de Linguistique 44(1): 23–48. De Pascale, Stefano. 2019. Token-based Vector Space Models as Semantic Control in Lexical Lectometry. PhD dissertation, KU Leuven. Del Tredici, Marco & Fernández, Raquel. 2017. Semantic variation in online communities of practice. In IWCS 2017 – 12th International Conference on Computational Semantics – Long papers. (20 May 2024). Dendien, Jacques & Pierrel, Jean-Marie. 2003. Le trésor de la langue française informatisé. Un exemple d’informatisation d’un dictionnaire de langue de référence. Traitement Automatique des Langues 44(2): 11–37. Devlin, Jacob, Chang, Ming-Wei, Lee, Kenton & Toutanova, Kristina. 2019. BERT: Pretraining of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. Minneapolis MN: Association for Computational Linguistics. Dollinger, Stefan. 2015. The Written Questionnaire in Social Dialectology: History, Theory, Practice. Amsterdam: John Benjamins. Dollinger, Stefan & Fee, Margery. 2017. DCHP-2: The Dictionary of Canadianisms on Historical Principles, 2nd edn. (20 May 2024). Donoso, Gonzalo & Sánchez, David. 2017. Dialectometric analysis of language variation in Twitter. In Proceedings of the Fourth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial), 16–25. Valencia: Association for Computational Linguistics. Durkin, Philip. 2012. Variation in the lexicon: The ‘Cinderella’ of sociolinguistics? Why does variation in word forms and word meanings present such challenges for empirical research? English Today 28(4): 3–9.
Modeling fine-grained sociolinguistic variation
Fee, Margery. 1991. Frenglish in Quebec English newspapers. In Papers of the Fifteenth Annual Meeting of the Atlantic Provinces Linguistic Association, 12–23. New Brunswick: Atlantic Provinces Linguistic Association. Fee, Margery. 2008. French borrowing in Quebec English. Anglistik: International Journal of English Studies 19(2): 173–188. Firth, John R. 1957. A synopsis of linguistic theory, 1930–1955. In Studies in Linguistic Analysis, 1–32. Oxford: Blackwell. Gimpel, Kevin, Schneider, Nathan, O’Connor, Brendan, Das, Dipanjan, Mills, Daniel, Eisenstein, Jacob, Heilman, Michael, Yogatama, Dani, Flanigan, Jeffrey & Smith, Noah A. 2011. Part-of-speech tagging for Twitter: Annotation, features, and experiments. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 42–47. Portland OR: Association for Computational Linguistics. Giulianelli, Mario, Del Tredici, Marco & Fernández, Raquel. 2020. Analysing lexical semantic change with contextualised word representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 3960–3973. Stroudsburg PA: Association for Computational Linguistics. Grant, Pamela. 2010a. English usage in contemporary Quebec: Reflections of the local. In Canadian English: A Linguistic Reader [Strathy Occasional Papers on Canadian English 6], Elaine Gold & Janice McAlpine (eds), 177–197. Kingston ON: Queen’s University. Grant, Pamela. 2010b. Is Quebec English distinct? English usage in contemporary Quebec [lecture slides]. (20 May 2024). Grant-Russell, Pamela. 1999. The influence of French on Quebec English: Motivation for lexical borrowing and integration of loanwords. In LACUS Forum 26, Shin Ja J. Hwang & Arle R. Lommel (eds), 473–486. Fullerton CA: The Linguistic Association of Canada and the United States. Grieve, Jack, Montgomery, Chris, Nini, Andrea, Murakami, Akira & Guo, Diansheng. 2019. Mapping lexical dialect variation in British English using Twitter. Frontiers in Artificial Intelligence 2: 11. Harris, Zellig S. 1954. Distributional structure. Word 10(2–3): 146–162. Hengchen, Simon, Tahmasebi, Nina, Schlechtweg, Dominik & Dubossarsky, Haim. 2021. Challenges for computational lexical semantic change. In Computational Approaches to Semantic Change, Nina Tahmasebi, Lars Borin, Adam Jatowt, Yang Xu & Simon Hengchen (eds), 341–372. Berlin: Language Science Press. Jones, Taylor. 2015. Toward a description of African American Vernacular English dialect regions using “Black Twitter.” American Speech 90(4): 403–440. Josselin, Amélie. 2001. L’emprunt lexical en France et au Canada: Le cas particulier des anglicismes et des gallicismes et leur traitement lexicographique. DEA thesis, Université de Lyon II. Labov, William. 1972. Sociolinguistic Patterns. Philadelphia PA: University of Pennsylvania Press. Laicher, Severin, Kurtyigit, Sinan, Schlechtweg, Dominik, Kuhn, Jonas & Schulte im Walde, Sabine. 2021. Explaining and improving BERT performance on lexical semantic change detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, 192–202. Stroudsburg PA: Association for Computational Linguistics.
167
168
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
Martinc, Matej, Montariol, Syrielle, Zosa, Elaine & Pivovarova, Lidia. 2020. Capturing evolution in word usage: Just add more clusters? In Companion Proceedings of the Web Conference 2020 (WWW ’20), 343–349. New York NY: Association for Computing Machinery. McArthur, Tom. 1989. The English Language as Used in Quebec: A Survey [Strathy Occasional Papers on Canadian English 3]. Kingston ON: Queen’s University. Miletić, Filip. 2019. Contact-induced lexical variation in Quebec English: An accountable description. In RJC2019 – 22èmes rencontres des jeunes chercheurs en sciences du langage, Paris, France. Miletić, Filip, Przewozny-Desriaux, Anne & Tanguy, Ludovic. 2020. Collecting tweets to investigate regional variation in Canadian English. In Proceedings of the 12th Language Resources and Evaluation Conference, 6255–6264. Marseille: European Language Resources Association. Miletić, Filip, Przewozny-Desriaux, Anne & Tanguy, Ludovic. 2021. Detecting contact-induced semantic shifts: What can embedding-based methods do in practice? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 10852–10865. Punta Cana, Dominican Republic: Association for Computational Linguistics. Miletić, Filip, Przewozny-Desriaux, Anne & Tanguy, Ludovic. 2023. Understanding computational models of semantic change: New insights from the speech community. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9209–9220. Singapore: Association for Computational Linguistics. Montariol, Syrielle, Martinc, Matej & Pivovarova, Lidia. 2021. Scalable and interpretable semantic change detection. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4642–4652. Stroudsburg PA: Association for Computational Linguistics. Nguyen, Dong. 2021. Dialect variation on social media. In Similar Languages, Varieties, and Dialects. A Computational Perspective, Marcos Zampieri & Preslav Nakov (eds.), 204–218. Cambridge: CUP. Nguyen, Dat Quoc, Vu, Thanh & Tuan Nguyen, Anh. 2020. BERTweet: A pre-trained language model for English Tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 9–14. Stroudsburg PA: Association for Computational Linguistics. Owoputi, Olutobi, O’Connor, Brendan, Dyer, Chris, Gimpel, Kevin, Schneider, Nathan & Smith, Noah A. 2013. Improved part-of-speech tagging for online conversational text with word clusters. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 380–390. Atlanta GA: Association for Computational Linguistics. Pavalanathan, Umashanthi & Eisenstein, Jacob. 2015. Confounds and consequences in geotagged Twitter data. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 2138–2148. Lisbon: Association for Computational Linguistics. Pedregosa, Fabian, Varoquaux, Gaël, Gramfort, Alexandre, Michel, Vincent, Thirion, Bertrand, Grisel, Olivier & Blondel, Mathieu et al. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research 12: 2825–2830.
Modeling fine-grained sociolinguistic variation
Poplack, Shana, Walker, James A. & Malcolmson, Rebecca. 2006. An English ‘like no other’? Language contact and change in Quebec. Canadian Journal of Linguistics/Revue Canadienne de Linguistique 51(2–3): 185–213. Rodda, Martina A., Lenci, Alessandro & Senaldi, Marco S. G. 2017. Panta rei: Tracking semantic change with distributional semantics in Ancient Greek. Italian Journal of Computational Linguistics 3(1): 11–24. Rouaud, Julie. 2019. Lexical and Phonological Integration of French Loanwords into Varieties of Canadian English Since the Seventeenth Century. PhD dissertation, Université Toulouse – Jean Jaurès. Schlechtweg, Dominik, Hätty, Anna, Del Tredici, Marco & Schulte im Walde, Sabine. 2019. A wind of change: Detecting and evaluating lexical semantic change across times and domains. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 732–746. Florence: Association for Computational Linguistics. Schlechtweg, Dominik, McGillivray, Barbara, Hengchen, Simon, Dubossarsky, Haim & Tahmasebi, Nina. 2020. SemEval-2020 task 1: Unsupervised lexical semantic change detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, A. Herbelot, X. Zhu, A. Palmer, N. Schneider, J. May & E. Shutova (eds), 1–23. Barcelona: International Committee for Computational Linguistics. Shoemark, Philippa, Sur, Debnil, Shrimpton, Luke, Murray, Iain & Goldwater, Sharon. 2017. Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Vol. 1: Long Papers, 1239–1248. Valencia: Association for Computational Linguistics. Statistics Canada. 2022. Table 98-10-0218-01. Mother tongue by age: Canada, provinces and territories. (20 May 2024). Tagliamonte, Sali A. 2002. Comparative sociolinguistics. In The Handbook of Language Variation and Change, Jack K. Chambers, Peter Trudgill & Natalie Schilling-Estes (eds), 729–763. Malden MA: Blackwell. Tagliamonte, Sali A. 2006. Analysing Sociolinguistic Variation. Cambridge: CUP. Tahmasebi, Nina, Borin, Lars & Jatowt, Adam. 2021. Survey of computational approaches to lexical semantic change. In Computational Approaches to Semantic Change, Nina Tahmasebi, Lars Borin, Adam Jatowt, Yang Xu & Simon Hengchen (eds), 1–91. Berlin: Language Science Press. Takamura, Hiroya, Nagata, Ryo & Kawasaki, Yoshifumi. 2017. Analyzing semantic change in Japanese loanwords. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Vol. 1: Long Papers, 1195–1204. Valencia: Association for Computational Linguistics. Turney, Peter D. & Pantel, Patrick. 2010. From frequency to meaning: Vector space models of semantics. Journal of Artificial Intelligence Research 37: 141–188. Uban, Ana, Ciobanu, Alina Maria & Dinu, Liviu P. 2019. Studying laws of semantic divergence across languages using cognate sets. In Proceedings of the 1st International Workshop on Computational Approaches to Historical Language Change, 161–166. Florence: Association for Computational Linguistics.
169
170
Filip Miletić, Anne Przewozny-Desriaux & Ludovic Tanguy
Wolf, Thomas, Debut, Lysandre, Sanh, Victor, Chaumond, Julien, Delangue, Clement, Moi, Anthony & Cistac, Pierric et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45. Stroudsburg PA: Association for Computational Linguistics. Xu, Yang & Kemp, Charles. 2015. A computational evaluation of two laws of semantic change. In Proceedings of the 37th Annual Meeting of the Cognitive Science Society, 2703–2708. Austin TX: Cognitive Science Society.
Subject index A accessibility 5, 30, 73, 75, 77–79, 81, 89, 91–97, 100, 109–110, 112 affinity propagation 153 annotation 11–12, 16, 28–29, 40, 52, 55–58, 61–64, 75, 83, 90, 94–95, 142–143, 145, 148, 153–155, 161, 164–165 contact influence 148, 153–154, 161–165 compounds 17 learner errors 61–62 multilingual language use 57–58, 64 named entities 35–40, 42 part-of-speech (POS) 11–17, 28, 75, 83
computational linguistics 39, 51 Constituent Likelihood Automatic Word-tagging System (CLAWS) 12, 15, 40, 45 contact-induced semantic shifts 62, 142–165 copyright 92, 96–100 corpus exploration 151 Corpus of Contemporary American English (COCA) 14, 21, 29, 45, 94, 127, 134–135 Corpus of Global Web-based English (GloWbE) 42–44 Corpus of Historical American English (COHA) 13–14, 18–23, 29, 94, 127, 134
B balance 11, 22–23, 29, 74, 76, 78, 84, 93 big-data (see very large corpora) bilingualism 145, 149, 154, 156, 160–163 British Library Newspapers database 69–85 British National Corpus (BNC) 4, 21, 40–43, 45–48, 51, 98, 100, 127, 135 Brown Corpus 39, 72
D databases 4–5, 10–11, 23, 29–30, 69, 72, 74–75, 85 deep learning 27 diachronic corpora 9, 12, 18, 22, 29, 45, 129, 131, 147 Digital Humanities 24, 30, 68–73, 75–76, 81, 85, 130 discourse of deficit 55–56, 61 dispersion 2, 20–21, 29, 42, 52, 116
C calquing 56, 58, 61, 63 Clean Corpus of Historical American English 51 codeswitching 56, 58, 158, 160, 165 collocation 36–37, 48–51, 83, 126, 137 common nouns 37–38, 40–45, 158 compilation 2, 4–6, 9–10, 16, 22, 29–30, 56, 59, 63–65, 68–69, 76–79, 85, 93–94, 98–101, 109–110, 131, 135, 137–138
E Eighteenth Century Collections Online (ECCO) 11, 23–27, 30 EF Cambridge Open Language Database (EFCAMDAT) 59 F false positives 57, 91, 143, 154–155, 157–162, 165 formality 134, 136, 145, 154, 156, 158, 160 foreignizing 56, 58, 61, 63
G genre 4, 6, 10–11, 22–23, 29, 58, 78–79, 93, 108–110, 115, 121, 126–139 genre categorization/ classification 18, 126–139 genre evolution 15–16, 22, 132 God’s truth fallacy 1, 3, 24, 36, 52, 89–90, 93, 99 H hapax legomena 25–26 Helsinki Corpus 9, 22 historical corpora 9–13, 15–16, 22, 27–30, 69, 131 historical corpus linguistics 4, 9–30, 36, 69–85 historical lexis / historical spelling 12, 17, 26–27 historical text databases 10–11, 23–27, 29–30 homographs 143, 156, 158–162 I information extraction 39 interlanguage 56–58, 60–63 International Corpus of English (ICE) 59, 62 International Corpus of Learner English (ICLE) 59, 62, 127 K keyword analysis 74, 126, 137 L language contact 56, 143, 145, 148, 155–165 language variation 10, 142–143, 146, 165 learner corpus (research) 55–65, 127 lengthwise analysis 106, 111, 118–120 lengthwise scaling 118–120 lexical bias 60–61
172
Challenges in corpus linguistics
lexical diversity 108, 114, 122–123 lexical innovation 62–63 literary corpora 127 literary studies 130 loan translations 144 Lost Generation Corpus 129, 138 Louvain International Database of Spoken English Interlanguage (LINDSEI) 57 M metadata 11, 22–24, 26–27, 29–30, 64–65, 69, 75, 81, 93, 95, 98, 116–117, 146, 156 multilingualism 56–58, 61–64 multiple correspondence analysis 120–121 multi-word units 36–37, 39–41, 45 Mystery of the Vanishing Reliability 1, 89–90, 94, 99 N n-grams 59, 75, 100, 126, 137 named entities 36–51 named entity recognition 36, 39, 51, 75, 127 natural language processing (NLP) 6, 27, 36, 39, 51, 60, 147 neural networks 27, 147, 152, 164 neural word embeddings 142–144, 147–148, 152–153 News on the Web Corpus (NOW) 49, 51 normalization 26, 107–109, 118, 121 O Open American National Corpus 98, 100
optical character recognition (OCR) 11–12, 24–27, 29, 74, 81–85 open research 89–102 P Parsed Corpus of Early English Correspondence (PCEEC) 15–17, 28 Philologist’s dilemma 1, 9–11, 28, 89, 92–93, 99 POS (part-of-speech) category change (see word class change) precision 2, 25, 29, 35, 39, 82–83 proper nouns / proper names 12, 17, 36–48, 158–160 Q Quebec English 142–165 query building 15–16, 25–27, 29, 51 quotations 36–37, 58, 97–98, 101 R recall 25–26, 29, 82–83 reference corpora 93, 96, 128 regional variation 43, 49, 143, 145–146, 155, 164–165 register 69, 73, 78–81, 84–85, 112, 136 register analysis 79, 84, 111, 120 reliability 1, 11, 24, 36, 51, 90, 94, 115, 120, 146 remediation 69, 71 replicability 64, 73, 89–92, 99, 101–102 representativeness 2–3, 10, 22–23, 29, 36, 68, 76–81, 84–85, 90, 92–94, 109, 127 resampling 93, 120–121, 123
S sampling 2, 9–11, 18, 20–22, 24–27, 29, 59, 65, 69, 76–78, 80, 84, 92–93, 121–122, 145 semantic change (see also contact-induced semantic shifts) 144–145, 147–148, 150–151, 165 Spanish Learner Language Oral Corpora (SPLLOC) 58 special corpora 85, 127 stylistics 130, 137 (near-)synonyms 41, 45–50, 156 syntactic parsing 16 T task instruction / task effects 55–56, 59–61, 64 text categorization (see text type) text length 106–123 text type 46–47, 127–135 text sampling (see sampling) transparency 91–94 Twitter 110, 117, 146, 149 Twitter corpora 101, 121, 142–165 type-token ratio 122 V vector space models (VSMs) 147–149 very large corpora 9–11, 18, 27–28, 30, 93, 112, 120 Vocabulary-Based Discourse Unit 117 W word-class change 12–15, 28 writing prompt 55–56, 59, 64
This book contributes to the discussion of challenges faced in different areas of corpus linguistics, namely the compilation, annotation, and analysis of linguistic corpora. In a field of growing corpus sizes and expanding possibilities of gathering data, some old issues persist, while at the same time new problems have emerged. As the compilation and study of language corpora gets increasingly sophisticated and complex, continuous attention on ways of dealing with the data in question and challenges in text selection and interpretation is needed. The contributions to this volume address problems relating to a variety of areas in corpus linguistic study, including corpus annotation, data variability, learner language, social media texts, and database utilization. The authors provide critical overviews and research-based analyses, discuss the nature of some of the common pitfalls, and offer solutions to existing problems.
isbn 978 90 272 1588 8
JOHN BENJAMINS PUBLISHING COMPANY