Classification and Applications of Multivariate Techniques Study Guide
Foundational Judgments for Classifying Multivariate Techniques
The classification of multivariate techniques is essential for becoming familiar with the various specific methods available to researchers. This classification system is fundamentally based on three critical judgments regarding the research objective and the nature of the data:
Division of Variables: Can the variables be divided into independent and dependent classifications based on established theory?
Number of Dependent Variables: If the variables can be divided, how many variables are treated as dependent within a single focused analysis?
Measurement Scales: How are the variables, including both dependent and independent sets, measured (e.g., metric vs. nonmetric)?
Distinction Between Dependence and Interdependence Techniques
Multivariate methods are broadly categorized into two types: dependence techniques and interdependence techniques.
Dependence Techniques: These are defined as methods where a variable or a set of variables is identified as the dependent variable to be predicted or explained by other variables, which are known as independent variables.
Interdependence Techniques: In these methods, no single variable or group of variables is defined as being independent or dependent. The procedure involves the simultaneous analysis of all variables within the set to discover underlying structures.
Rules of Thumb for Statistical Power Analysis
When designing studies using multivariate techniques, researchers should adhere to specific guidelines regarding statistical power:
Power Level Target: Researchers should design studies to achieve a power level of at the desired significance level.
Significance Levels and Sample Size: Utilizing more stringent significance levels (e.g., opting for instead of ) necessitates larger sample sizes to maintain the desired power level.
Alpha Level Adjustments: Conversely, power can be increased by choosing a less stringent alpha level (e.g., instead of ).
Effect Size Impact: Smaller effect sizes require significantly larger sample sizes to achieve the target power level.
Sample Size Priority: The most likely way to achieve an increase in power is by increasing the sample size.
Functional Characteristics of Dependence Techniques
Dependence techniques are further categorized based on two primary characteristics:
Categorization by Number of Dependent Variables:
Single dependent variable.
Several dependent variables.
Several dependent/independent relationships (found in complex models like Structural Equation Modeling).
Categorization by Measurement Scale:
Metric: Quantitative or numerical data.
Nonmetric: Qualitative or categorical data.
Mathematical Relationships Between Multivariate Dependence Methods
The following equations represent the family of dependence techniques as defined by the nature and number of variables involved:
Canonical Correlation: Note: Canonical correlation is considered the general model because it places the fewest restrictions on types and numbers of variables.
Multivariate Analysis of Variance (MANOVA):
Analysis of Variance (ANOVA):
Multiple Discriminant Analysis:
Multiple Regression Analysis:
Conjoint Analysis:
Structural Equation Modeling (SEM):
Overview of Specific Multivariate Techniques
1. Principal Component and Common Factor Analysis
Definition: A statistical approach used to analyze interrelationships among many variables to explain them in terms of common underlying dimensions (factors).
Objective: To condense information from original variables into a smaller set of variates with minimal information loss. This provides an objective basis for creating summated scales.
Example: A researcher analyzing fast-food restaurant ratings on six variables () might find they combine into two factors: (taste, temperature, freshness) and (waiting time, cleanliness, friendliness).
2. Multiple Regression
Definition: Used for problems involving a single metric dependent variable related to two or more metric independent variables.
Objective: To predict changes in the dependent variable using the statistical rule of least squares.
Examples:
Predicting monthly dining expenditures () based on family income, size, and age of the head of household ().
Predicting company sales based on advertising expenditures, number of salespeople, and number of retail stores.
3. Multiple Discriminant Analysis (MDA) and Logistic Regression
Multiple Discriminant Analysis: Appropriate when the single dependent variable is nonmetric/categorical ( or ) and independent variables are metric.
Objective: Understand group differences and predict group membership (e.g., distinguishing high-credit-risk individuals from low-risk ones).
Logistic Regression (Logit Analysis): A combination of regression and MDA. It predicts a single nonmetric dependent variable but accommodates all types of independent variables () without requiring multivariate normality.
Example: Identifying financial and managerial data that differentiate between successful and unsuccessful start-up firms over a -year period to select future investment candidates.
4. Canonical Correlation
Definition: A logical extension of multiple regression.
Objective: To simultaneously correlate several metric dependent variables with several metric independent variables by maximizing the linear correlation between the two sets.
Example: Comparing customer perceptions of a company versus "world-class companies" across metrically measured questions to determine overall correlation and specific question-level correlations.
5. Multivariate Analysis of Variance (MANOVA) and Covariance (MANCOVA)
MANOVA: Explores the relationship between several categorical independent variables (treatments) and two or more metric dependent variables.
MANCOVA: Used to remove the effect of uncontrolled metric independent variables (covariates) from the dependent variables.
Example: Determining statistical differences in customer perceptions () between those who saw a humorous advertisement and those who saw a non-humorous one.
6. Conjoint Analysis
Definition: Assesses the importance of product attributes and the levels within those attributes.
Objective: Consumers evaluate a subset of product profiles, allowing the researcher to determine the attractiveness of specific combinations.
Example: Designing a product with three attributes () each at three levels. Instead of evaluating all () combinations, consumers evaluate a subset of or more. This data aids product design simulators.
7. Cluster Analysis
Definition: An interdependence technique for developing meaningful subgroups of individuals or objects.
Objective: To classify entities into mutually exclusive groups based on similarity. Unlike MDA, groups are not predefined.
Example: A restaurant owner grouping customers into clusters based on whether they are motivated by low prices or other factors like food quality.
8. Perceptual Mapping (Multidimensional Scaling)
Definition: Transforms consumer judgments of similarity or preference into distances in multidimensional space.
Objective: To show the relative positioning of objects (e.g., brands).
Example: Mapping fast-food brands to show that Burger King is most similar to Wendy’s, identifying Wendy’s as the primary competitor.
9. Correspondence Analysis
Definition: An interdependence technique for perceptual mapping of objects on a set of nonmetric attributes.
Objective: Quantifies qualitative nominal data using a contingency table (cross-tabulation) and performs dimensional reduction.
Example: Mapping brand preferences against demographic variables to see the "correspondence" between certain brands and the characteristics of the people who prefer them.
10. Structural Equation Modeling (SEM) and Confirmatory Factor Analysis (CFA)
SEM: Allows for separate, interrelated relationships between a set of dependent variables estimated simultaneously. It consists of a structural model (path model) and a measurement model (incorporating indicators).
CFA: A component of SEM used to assess how well scale items measure a concept and to account for measurement error.
Example: A study on worker satisfaction involving predictors like supervisor support, work environment, and job performance. SEM allows the researcher to analyze direct effects on satisfaction and indirect effects through job performance simultaneously while incorporating multi-item scales.
Selecting a Multivariate Technique (Decision Tree Summary)
Selection is guided by the structure of relationships:
Dependence (Predicted Variables):
One dependent variable (Metric: Regression/Conjoint; Nonmetric: Discriminant/Logistic).
Several dependent variables (Metric: MANOVA/Canonical Correlation; Nonmetric: Canonical Correlation with dummy variables).
Multiple relationships (SEM).
Interdependence (Structure of Relationships):
Structure among Variables: Factor Analysis.
Structure among Cases/Respondents: Cluster Analysis/CFA.
Structure among Objects (Attributes): Metric Scaling (Multidimensional Scaling) or Nonmetric Scaling (Correspondence Analysis).