1/118
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
layout of r studio
r script = top left
environment = top right
r console = bottom left
help and plots = bottom right
r script
write and run code
r console
where r puts executed code and outputs (as well as errors)
what colour is executed code
blue
what colour is output of executed code
black
[1] thingy meaning
means the output to its right is the first output in this case
r script ___ code
saves
<
less than
<=
less than or equal to
!=
not equal to
==
must use both of these for conditionals!!!!!!
like “is this true or nah”
and symbol
&
or symbol
|
what’s the caveat for the or symbol
it won’t tell you which one is true
what does creating objects do
it’s how r stores info
like a box that contains something
object operator
←
sequence for creating an object
object_name ← object_content
how do i find the content of the object
object_name in R!!!!!
it’s basically saying what is inside this
where will the content of the object show up
environment!!!!!
rules for naming objects
cannot begin with a number/have spaces/special symbols
for textual content, use quotes
why do we use quotes for text
if you don’t it’ll assume the text is an object. and it' won’t work b/c the object hasn’t been named
when do we use quotes when writing code
not for names of objects, functions, arguments or special values
u use it for all other texts unless this choice is intentional
what’s this weird thing that r will do (hint: override)
it’ll override objects if you assign new content to existing object name
also r is case sensitive
what are vectors
objects with a series of values
“multiple elements of the same type”
what function do vectors use
c() COMBINE BABY
what are functions
actions you require r to perform on particular object or data
takes inputs → action → output
functions can also be applied to vectors
what if you have a combined function with “na”
no value so if u wanna calculate mean or smth u gotta deal w that
mean/median when you have na
mean(y, na.rm = TRUE)
median(y, na.rm = TRUE)
what is a name of a function followed by
function_name
PARANTHESIS
how many required arguments are needed (at least)
1, for most functions
optional arguments
are optional. not rlly necessary unless needed
what do you do if you have multiple arguments
function_name(argument 1, argument 2)
how do you specify arguments
put priority or include number of arguments in in specifics
if more than 1 req arg, put em in that order
okay what are the two formats we r using
function_name (required_argument, optional_argumentname = optional_argument)
function_name(required_argument)
packages
extend the functionality of r
what do we do with packages
install them + load
(installing is only done once tbh, load everytime)
sequence for installing packages
install.packages (“causaldata”)
loading packages
library(causaldata)
#
commenting code; r ignores everything after # until end of line
what are datasets/data
they capture characteristics of a particular set of individuals or entities
how are datasets organzied
as dataframes
rows =
observations
columns
variables
what are observations
information collected from a particular entity or individual in the study
unit of observation
defines the individual or entity that each observation in a dataframe represents; each observation is a row number, denoted as i
what are variables
values of changing characteristics for various individuals and entities in the study
how do we refer to variables
by its name
notation for defining a new variable
x = {10, 5, 8}
where x is the name of the variable
and the numbers are the content of the variables; multiple observations
each individual observation is
iiiiiiii
the observation number
what is data wrangling
allows us to make data in the form we want
tibble
how tidyverse stores dataframes
what are the three methods for datasets within r
some are built in, others are acquired via loading r packages, and others — we read em in ourselves
names()
gives us the names of columns
head()
first 6 rows and observations in a dataset
dim()
dimensions of a dataset
number of rows and columns
n(row)
number of rows
n(col)
number of columns
View() (w a capital)
contents of the entire dataset
what tf is dplyr
gives us a consistent set of verbs for data manipulation
how to get that dplyr
install.packages (“dplyr”)
library(dplyr)
pipe
used to write code sequentially %>%
base r vs the pipe
base r sqrt(sum(range(student_ages)))
pipe
student_ages %>%
range()%>%
sum()%>%
sqrt()
filter()
selects rows that satisfy the argument we provide it
how do we save as a new dataset
we must create a new object
subsets in base r
flights_AA2 ← flights [flights $ carrier == “AA”]
filtering on multiple conditions
flights%>%
filter(dest “LAS” | dest “LAX)
(w two equal signs)
notice how you put dest TWICE
new dataframe
new object
what does - do
removes columns
flights%>%
select(-carrier, -dep_delay)
you can combine dplyr functions
by pipe’in, filter, select
rename
(new_name = old_name)
what does mutate do
operates on the existing column
adds new row
(new_var = operation on existing)
in base r….
flightsdistancekm.←flightsdistance * 1.6
summarize()
calculates the summaries of variables
like mean, max, min
it can also count the number of rows in a dataframe
group_by()
applies a function to a group of observations; divides data into groups based on variables
what’s the note about groupby function
it doesn’t change the data. subsequent functions will make that mess
what happens after you add summarize() after groupby function
you get a stack of descriptive statistics which is awesome
example
flights %>%
group_by (month) %>%
summarize(n_obs = n()))
steps for working directory
set wd to 380data
load dataset using read.csv
understand the data by using View and head
identify the number of observations and variables
read.csv()
reads csv files
the required argument is the name of the csv file in quotes
types of variables
character and numeric
type of character variables
categorical (2+ text categories)
types of numeric variables
binary - only two
non binary - more than two
binary
only two values, 1s and 0s
presence or absence of a trait
non binary
more than 2 values
character variables
contain text
often categorical; take on a limited number of values
binary numeric variables
numeric encodings tha could be coded as character variables with two categories
(1, 0) or (voted, didn’t vote)
avg/mean of a variable
sum of all values across all observations divided by total number of observations
x bar
avg of x
that weird e to the power of n and subscript 1 = 1 x i
sum of all xi (observations of x) from i = 1 to i = n, 1st observation of x to the last
xi
observation x where i = position of observation and n = total number of observations in the variable
non binary mean of a variable
as an average; same units as the variable (1 or 0)
binary mean of a variable
as a proportion in %; after multiplying by 100
why are binary means a proportion
bc mean of a binary variable = proportion to observations that have that trait
for categorical variables with more than 2 categories
there’s no straightforward interpretation
numeric/double variables
decimals
integer variables
whole numbers
factor variables
categorical
ggplot2 package (what does it do)
basically just the grammar of graphics framework
ggplot2
a statistical graphic; mapping of data variables to aesthetic attributes of gramatical objects
data
datasets w/ variables
geom
geometric objects we wish to plot (points + bars)