Exploring One Variable Categorical Data and Statistical Fundamentals
Fundamental Concepts in Statistical Studies
A statistical study is defined as a study in which data are collected from a sample to answer an investigative question about a population. Statistical studies provide the systematic empirical foundation for collecting evidence and drawing valid conclusions.
At the core of every statistical study is an investigative question. An investigative question is a question that can be answered using data and is designed to explore variability in a population. Formulating a precise investigative question directs the design of data collection and the selection of analytical methods.
A population encompasses the entire group of individuals or objects that we want information about. Because evaluating an entire population is frequently unfeasible or impossible due to practical constraints, statistical studies collect data from a sample, which is a smaller group selected from the population.
Within a sample, individual characteristics are measured. A variable is defined as a characteristic that can take on different values for individuals in a study. The specific values collected for each variable from the individuals in a sample constitute the data.
Analyzing statistical data requires understanding variability. Variability is defined as the extent to which data values differ from one another. Quantifying and explaining variability is the primary purpose of statistical analysis.
Classification of Variables and Summary Statistics
To measure and analyze data systematically, researchers must identify the entity being measured and classify the variable type. An observational unit is the individual person, object, or thing on which a variable is measured.
Summary values in statistics are distinguished by whether they describe an entire population or a sample:
A parameter is a fixed, usually unknown value that describes a population.
A statistic is a value calculated from a sample, used to estimate a parameter.
Variables are broadly categorized into qualitative and quantitative types, which determine the appropriate mathematical operations and statistical summaries:
A categorical variable, also known as a qualitative variable, is a variable that places an individual into one of several groups or categories. Categorical variables are summarized using proportions.
A quantitative variable is a variable that takes numerical values for which arithmetic operations make sense. Quantitative variables are summarized using numerical measures.
Quantitative variables are further classified based on the nature of their numerical values:
Discrete data are quantitative data that result from counting. Discrete data can only take specific, separate values, which are usually whole numbers.
Continuous data are quantitative data that result from measuring. Continuous data can take any value within a range or continuum.
Tabular Representation and Summary Statistics for Categorical Variables
Summarizing a single categorical variable requires tracking the occurrence of each category within a data set. Frequency, also referred to as a count, is defined as the number of times a particular category occurs in a data set.
Proportion is defined as the ratio of a part to the whole. Relative frequency is the proportion or percentage of the total data set that falls into a particular category. It is computed by dividing the frequency of a category by the total sample size:
Similarly, a categorical proportion is calculated as:
Categorical summary statistics are structured visually in tabular forms:
A frequency table is a table that displays the count of observations in each category of a categorical variable.
A relative frequency table is a table that displays the percentage of observations in each category of a categorical variable.
Graphical Representations for One Categorical Variable
Graphical representations illustrate the distribution of categorical variables using visually distinct displays for single groups and comparative groups:
A frequency bar graph is a graph that displays the count of observations in each category of a categorical variable.
A relative frequency bar graph is a graph that displays the percentage of observations in each category of a categorical variable.
A circle graph, also known as a pie chart, is a graph that displays relative frequencies as proportional slices of a circle, where the entire circle represents of the data.
When evaluating a categorical variable across multiple groups, specialized comparative graphs are used:
A double bar graph is a graph that compares the distribution of a single categorical variable across two or more groups, using side-by-side bars.
A segmented bar graph is a graph that compares the distribution of a categorical variable across two or more groups, where each bar is divided into segments representing relative frequency.
A mosaic plot is a variation of a segmented bar graph where the width of each bar is proportional to the sample size of that group, allowing both distribution and group size to be displayed in a single graph.
All tabular and graphical methods serve to reveal the underlying distribution of a variable. Distribution is defined as the pattern of values a variable takes, including which categories are more or less common.