stringi: Fast and Portable Character String Processing in R

33Citations
Citations of this article
56Readers
Mendeley users who have this article in their library.

Abstract

Effective processing of character strings is required at various stages of data analysis pipelines: from data cleansing and preparation, through information extraction, to report generation. Pattern searching, string collation and sorting, normalization, transliteration, and formatting are ubiquitous in text mining, natural language processing, and bioinformatics. This paper discusses and demonstrates how and why stringi, a mature R package for fast and portable handling of string data based on ICU (International Components for Unicode), should be included in each statistician’s or data scientist’s repertoire to complement their numerical computing and data wrangling skills.

Cite

CITATION STYLE

APA

Gagolewski, M. (2022). stringi: Fast and Portable Character String Processing in R. Journal of Statistical Software, 103(2). https://doi.org/10.18637/jss.v103.i02

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free