Quantitative Text Analysis and Evaluating Lexical Style in R

Offered By
Coursera Project Network
In this Guided Project, you will:

tokenize text documents to examine top words by frequency

examine the change in type to token ratio or level of text complexity over time

Clock1 hour
BeginnerBeginner
CloudNo download needed
VideoSplit-screen video
Comment DotsEnglish
LaptopDesktop only

By the end of this project, you will learn about the concept of lexical style in textual analysis in R. You will know how to load and pre-process a data set of text documents by converting the data set into a corpus and document feature matrix. You will know how to calculate the type to token ration which evaluates the level of complexity of a text, and know how to isolate terms of particular lexical interest in a text and visualize the variation in frequency of such terms in texts over time.

Skills you will develop

  • Descriptive Analysis
  • Text Analysis
  • Data Wrangling
  • Data Visualization (DataViz)
  • Text Corpus

Learn step-by-step

In a video that plays in a split-screen with your work area, your instructor will walk you through these steps:

  1. Load textual data into R and turn it into a corpus object and understand the concept of lexical style in textual analysis

  2. Extract meta-data from text document filenames and calculate the type to token ratio (TTR)

  3. Examine the change in the type to token ratio or level of text complexity over time

  4. Tokenize text documents to examine top words by frequency of appearance and isolate words of particular lexical interest in the text

  5. Visualize the change in the variation in the frequency of features of particular lexical interest in your text

How Guided Projects work

Your workspace is a cloud desktop right in your browser, no download required

In a split-screen video, your instructor guides you step-by-step

Frequently asked questions

Frequently Asked Questions

More questions? Visit the Learner Help Center.