This textbook examines empirical linguistics from a theoretical linguist’s perspective. It provides both a theoretical discussion of what quantitative corpus linguistics entails and detailed, hands-on, step-by-step instructions to implement the techniques in the field. The statistical methodology and R-based coding from this book teach readers the basic and then more advanced skills to work with large data sets in their linguistics research and studies. Massive data sets are now more than ever the basis for work that ranges from usage-based linguistics to the far reaches of applied linguistics. This book presents much of the methodology in a corpus-based approach. However, the corpus-based methods in this book are also essential components of recent developments in sociolinguistics, historical linguistics, computational linguistics, and psycholinguistics. Material from the book will also be appealing to researchers in digital humanities and the many non-linguistic fields that use textual data analysis and text-based sensorimetrics. Chapters cover topics including corpus processing, frequencing data, and clustering methods. Case studies illustrate each chapter with accompanying data sets, R code, and exercises for use by readers. This book may be used in advanced undergraduate courses, graduate courses, and self-study.
Presentation
Delving into corpus linguistics, its theoretical relevance, and its applications
More
ORCID ID
orcid.org/0000-0003-4895-0788recent publications
Categories
tags
Artificial Intelligence
BNC.query()
BNC2014
British National Corpus
choropleth maps
componential semantics
corpus linguistics
data frames
data manipulation
debate
discourse
distant reading
distributional hypothesis
Eleanor Rosch
frequency
frequency
frequency list
Likert scale
linguistic footprint
LLMs
networks
POS tagging
Prototype semantics
Prototype Theory
R
R
readme
script
semantics
semantic vector space
Shiny
sociolinguistics
split infinitive
spoken
structuralist semantics
tidyverse
Transformers
UDPipe
Universal Dependencies
variation
voyant tools
what makes soup soup?
wordcloud
word embeddings
word vectors

