1. Read the data
My first blog is about the start operations to open and get the data:
- open the data
- parse a vector
- other types of the data to open
- cleaning the data
1.1. Open
Reading the data from a file:
readr:
read_csv()
read_csv2() -- reads semicolon separated files.
read_tsv()
read_delim()
read_fwf() -- reads fixed width files. You can specify fields either by their widths with fwf_widths() or their position with fwf_positions().
read_table() -- reads a common variation of fixed width files where columns are separated by white space.
read_log() -- reads Apache style log files.
Change the IV columns to a factor type:
data <-
read_csv("data/knobology_synthetic.csv", col_types = "cccd") %>%
mutate(
device = as.factor(device),
vision = as.factor(vision))
1.2. Parsing a vector
Using parsers is mostly a matter of understanding what’s available and how they deal with different types of input.
str(parse_logical(c("TRUE", "FALSE", "NA")))
str(parse_integer(c("1", "2", "3")))
parse_integer(c("1", "231", ".", "456"), na = ".")
str(parse_date(c("2010-01-01", "1979-10-14")))
parse_double("1,23", locale = locale(decimal_mark = ","))
parse_number("$100") -- ignores non-numeric characters before and after the number.
parse_number("123'456'789", locale = locale(grouping_mark = "'"))
parse_factor()
parse_character()
charToRaw("Hadley") -- encoding ASCII.
parse_character(x1, locale = locale(encoding = "Latin1")) -- to specify the encoding.
guess_encoding(charToRaw(x1)) -- detect the used encoding.
parse_date("1 janvier 2015", "%d %B %Y", locale = locale("fr")) -- French
1.3. Other type of data
-
haven reads SPSS, Stata, and SAS files.
-
readxl reads excel files (both .xls and .xlsx).
-
DBI, along with a database specific backend (e.g. RMySQL, RSQLite, RPostgreSQL etc) allows you to run SQL queries against a database and return a data frame.
-
For hierarchical data: use jsonlite (by Jeroen Ooms) for json, and xml2 for XML.
2. Cleaning the data
library(tidyverse)
dplyr::select()
tidyr
gather()
spread()
table2 %>%
spread(key = type, value = count)
stocks %>%
complete(year, qtr) -- takes a set of columns, and finds all unique combinations.
fill() -- takes a set of columns where you want missing values to be replaced by the most recent non-missing value.
who %>%
gather(key, value, new_sp_m014:newrel_f65, na.rm = TRUE) %>%
mutate(key = stringr::str_replace(key, "newrel", "new_rel")) %>%
separate(key, c("new", "var", "sexage")) %>%
select(-new, -iso2, -iso3) %>%
separate(sexage, c("sex", "age"), sep = 1)