Psychology and behavioral sciences - Theme
- Login of registreer om te kunnen reageren
- 37621 keer gelezen
Multivariate statistics consists of a range of statistical techniques that deal with the analysis and interpretation of data sets containing multiple variables simultaneously. Multivariate statistics provides analysis when there are many independent variables (IVs) and/or many dependent variables (DVs), all correlated with one another to some degree. Unlike univariate statistics, which focus on analyzing one variable at a time, multivariate statistics examines the relationships and interactions among multiple variables in a dataset. It is widely used in various disciplines, including psychology, economics, biology, sociology, marketing, and more. The main goal of multivariate statistics is to uncover patterns, associations, and dependencies among the variables in a dataset to gain a deeper understanding of the underlying structure and complexity of the data.
First, it is important to distinguish between dependent and independent variables. Independent variables (IVs) are the differing conditions (treatment vs. placebo) to which you expose your subjects, or the characteristics (tall or short) that the subjects themselves bring into the research experiment. IVs are usually considered predictor variables because they predict the dependent variables, also called the response or outcome variables. Note that independent and dependent variables are defined within a research context; a DV in one research setting may be an IV in another.
Multivariate statistics, univariate statistics, and bivariate statistics are different approaches to analyzing data based on the number of variables involved in the analysis:
Univariate Statistics: Univariate statistics deal with the analysis of a single dependent variable at a time. There may be, however, more than one IV. The main goal of univariate analysis is to describe and summarize the distribution, central tendency, and dispersion of a single variable. Common univariate techniques include calculating measures such as mean, median, mode, standard deviation, variance, and conducting graphical representations like histograms and box plots. Univariate analysis provides basic insights into the characteristics of individual variables but does not consider the relationships between variables.
Bivariate Statistics: Bivariate statistics involve the analysis of two variables simultaneously, where neither is an experimental IV. The primary focus of bivariate analysis is to explore the relationship between two variables. Common bivariate techniques include correlation analysis, which measures the strength and direction of the linear relationship between two continuous variables, and cross-tabulation or contingency tables, used for categorical variables. Bivariate statistics are useful in understanding how two variables are related, but they don't take into account the influence of other variables in the dataset.
Multivariate Statistics: Multivariate statistics deal with the simultaneous analysis of three or more variables in a dataset. With multivariate statistics, you analyze multiple dependent and independent variables. The main objective of multivariate analysis is to explore complex relationships and interactions between multiple variables. Unlike univariate and bivariate statistics, multivariate statistics can consider the joint variability of multiple variables and examine how they collectively influence the outcome. Multivariate techniques include methods like Principal Component Analysis (PCA), Factor Analysis, Multivariate Analysis of Variance (MANOVA), Multiple Regression Analysis, Canonical Correlation Analysis (CCA), Cluster Analysis, and Structural Equation Modeling (SEM). Multivariate analysis is particularly useful when trying to understand the interdependencies among multiple variables and the underlying structure of complex datasets. It is important in both nonexperimental (correlational or survey) and experimental research.
In summary, the key differences between these three types of statistics lie in the number of variables they handle and their respective goals. Univariate statistics analyze one dependent variable to describe its distribution, bivariate statistics examine the relationship between two variables, and multivariate statistics examine the relationships and interactions among three or more variables to gain a deeper understanding of complex data patterns.
Experimental research involves the deliberate manipulation of variables with random assignment to groups to establish causal relationships, while nonexperimental research relies on observation without manipulation and is more focused on exploring correlations and associations between variables. Each type of research design has its strengths and limitations and is appropriate for different research questions and contexts.
Key characteristics of experimental research include:
Examples of experimental research include drug trials, educational interventions, and studies investigating the impact of a specific treatment or therapy on a particular outcome.
Nonexperimental research, also known as observational or correlational research, involves the observation and analysis of variables without manipulating them. In nonexperimental studies, researchers do not intervene or control the independent variables but instead observe how naturally occurring variables are related to each other.
Key characteristics of nonexperimental research include:
Examples of nonexperimental research include cross-sectional studies, longitudinal studies, case-control studies, and surveys that explore relationships between variables without any manipulation.
Multivariate statistics is gaining popularity because the techniques are now accessible by computer. Two packages are demonstrated in this book. Examples are based on programs in IBM SPSS and SAS. However, the real challenge in multivariate statistics lies not in the computation itself, which can be done easily with the help of computer programs. The true difficulty lies in the following key aspects:
This book provides support in helping researchers to be knowledgeable and diligent in every step of multivariate analysis to ensure that the conclusions drawn from the results are valid, reliable, and truly reflect the underlying patterns in the data.
Interval or quantitative (continuous): Interval or quantitative data is a type of numerical data that represents measurements on a continuous scale, where the values have meaningful and consistent intervals. This means that the numerical difference between two values is meaningful and can be measured. Examples include temperature in Celsius or Fahrenheit, height, weight, and time. Continuous data can take any value within a range and often involves measurements or real numbers.
Nominal, categorical, or qualitative (dichotomous): Nominal, categorical, or qualitative data is a type of data that consists of categories or groups into which items are classified. It does not have any inherent numerical meaning or order. Dichotomous data is a special case of categorical data that has only two categories or levels. Examples of nominal data include gender (male/female), eye color, and categorical data like hair color (black, brown, blonde, etc.). Dichotomous data examples include yes/no, true/false, or pass/fail.
Discrete: Discrete data is a type of quantitative data that consists of whole numbers or counts and represents distinct, separate values. There are no intermediate values between data points, and the data often results from counting or categorizing. Examples include the number of children in a family, the number of students in a class, or the number of cars in a parking lot.
In summary, interval or quantitative data is continuous and measured on a scale, while nominal, categorical, or qualitative data involves categories or groups without inherent numerical meaning. Discrete data is also quantitative but represents distinct, separate values with no intermediate values.
In statistics, the population refers to the entire group of individuals, items, or events that share a common characteristic and are of interest to the researcher. It is the complete set of elements that the researcher wants to study and draw conclusions about. The population is usually large and may be difficult or impractical to study in its entirety. For example, if a researcher is interested in studying the average height of all adults in a country, then the entire adult population of that country would be the population of interest.
A sample is a subset or a smaller representative group of the population. It is selected from the larger population with the intention of drawing inferences or making generalizations about the entire population. The process of selecting a sample from the population is called sampling. The goal of sampling is to obtain a sample that is representative of the population, so the findings from the sample can be generalized back to the larger group. In the previous example of studying the average height of all adults in a country, the researcher might select a sample of, say, 1000 adults from different regions to represent the entire population.
Sampling is a crucial aspect of statistical research because it allows researchers to draw conclusions and make predictions about a population without having to study every single individual or item within that population. Properly conducted sampling ensures that the findings from the sample can be applied to the entire population with a certain level of confidence, provided that the sample is representative and properly chosen.
Descriptive statistics involves the summarization, organization, and presentation of data to provide a clear and concise understanding of its main features. It aims to describe and depict the basic characteristics of a dataset without making any generalizations or inferences to a larger population. Common measures used in descriptive statistics include measures of central tendency (e.g., mean, median, mode) and measures of variability (e.g., range, standard deviation, variance). Graphs and charts, such as histograms, bar graphs, and pie charts, are also commonly used to visually represent the data. Descriptive statistics is helpful in simplifying complex data sets and providing a comprehensive overview of the data.
Inferential statistics involves drawing conclusions, making predictions, and testing hypotheses about a larger population based on data collected from a sample. It uses probability theory and sampling techniques to estimate parameters of interest and assess the uncertainty associated with those estimates. Researchers use inferential statistics to make inferences beyond the sample and generalize the findings to the entire population. Techniques such as hypothesis testing, confidence intervals, and regression analysis are commonly used in inferential statistics. This type of statistics is essential for making decisions, drawing insights, and understanding the relationships between variables at a broader level than what the sample data alone can provide.
Orthogonality is a perfect nonassociation between variables. If two variables are orthogonal, knowing the value of one variable gives no clue as to the value of the other; the correlation between them is zero. In the context of multivariate data, orthogonality between variables refers to the absence of correlation or linear dependence between the variables. Orthogonal variables are statistically independent, and changes in one variable do not affect the values of the other variables in the set.
Usually, however, the variables are correlated with each other (nonorthogonal). When variables are correlated, they have shared or overlapping variance. In standard analysis, the overlapping variance contributes to the size of summary statistics of the overall relationship but is not assigned to either variable. Overlapping variance is disregarded in assessing the contribution of each variable to the solution. Sequential analyses differ, in that the researcher assigns priority for entry of variables into equations, and the first one to enter is assigned both unique variance and any overlapping variance it has with other variables. Lower-priority variables are then assigned on entry their unique and any remaining overlapping variance.
Selecting a strategy to handle overlapping variance in data is not a simple task. When variables are correlated, the overall relationship remains consistent, but the perceived significance of individual variables in the solution can differ based on whether a standard or sequential strategy is employed. Multivariate procedures may be considered unreliable because their outcomes can change significantly with different variable entry strategies. However, each strategy poses unique questions to the data, and it is the responsibility of the researcher to precisely define the question they want to address.
The combination that is formed depends on the relationships among the variables and the goals of analysis, but in most cases, the combination is linear. A linear combination is one in which each variable is assigned a weight (for example W1). Next, the products of weights and the variable scores are summed to predict a score on a combined variable. In the equation below, Y' (the predicted DV) is predicted by a linear combination of X1 and X2 (the IVs). The basic formula looks like:
\[ Y' = W_{1}X_{1} + W_{2}X_{2} \]
The magnitude of the W values (or a related measure) provides valuable insights into the relationship between dependent variables (DVs) and independent variables (IVs). A W value of zero for an IV indicates that the IV is not necessary for the best DV-IV relationship. Conversely, a large W value suggests that the IV plays a significant role in the relationship. While interpreting the multivariate solution based solely on the W values is not always straightforward due to complexities, these values remain crucial in most multivariate procedures to understand the importance of variables in the relationships.
A general rule is to get the best solution with the fewest variables. As more and more variables are included, the solution usually improves, but only slightly. Sometimes the improvement does not compensate for the cost in degrees of freedom of including more variables, so the power of the analyses diminishes. Overfitting occurs when too many variables are included in an analysis relative to the sample size. With smaller samples, very few variables can be analyzed. Generally, a researcher should include only a limited number of uncorrelated variables in each analysis, fewer with smaller samples.
Statistical power refers to the probability of correctly detecting an effect or relationship in a statistical test when it truly exists. There are many software packages available to help you estimate the power available with various sample sizes for various statistical techniques, and to help you determine necessary sample size given a desired level of power and expected sizes of relationships. One of these programs that estimates power for several techniques is NCSS PASS.
An appropriate data set for multivariate statistical methods consists of values on a number of variables for each of several subjects or cases. These data can be organized in different ways, for example:
In summary, a data matrix organizes data in a tabular form, a correlation matrix shows the relationships between variables, a variance-covariance matrix displays variances and covariances, an SSCP matrix is used in statistical computations, and residuals help assess model fit and identify model errors.
In the context of statistical analysis, residuals are the differences between observed data points and the corresponding predicted values obtained from a statistical model, such as a regression model. Residuals are used to assess the goodness-of-fit of the model and identify patterns or systematic errors that the model does not capture.
This chapter explains how to organize the statistical techniques in this book by major research question. In summary, a decision tree leads you to an appropriate analysis for your data. On the basis of your major research question and a few characteristics of your data set, you determine which statistical technique(s) is appropriate.
To determine which analytic strategy to use, you should consider the following steps:
Hence, the analytic strategy is to be determined based on the major research question, the number (and kind) of dependent and independent variables and the presence (or absence) of covariates. This chapter briefly introduces the statistical techniques. However, the focus is on when to choose which technique rather than to discuss each technique into detail. To do so, each paragraph describes the techniques belonging to one of the following major research questions:
If the major purpose of analysis is to assess the associations among two or more variables, there are five statistical techniques you could potentially use. The choice among these five different statistical techniques depends on the number of independent and dependent variables, the nature of the variables (continuous or discrete), and whether any of the independent variables is best conceptualized as covariate.
| Major research question | Number (and kind) of DVs | Number (and kind) of IVS | Covariates | Analytic strategy | Goal of analysis |
| Degree of relationship among variables | One (continuous) | One (continuous) | - | Bivariate r (see Chapter 17) | Create a linear combination of IVs to optimally predict DV. |
| ,, | ,, | Multiple (continuous) | None | Multiple R (see Chapter 5) | Create a linear combination of IVs to optimally predict DV. |
| ,, | ,, | ,, | Some | Sequential Multiple R (see Chapter 5) | Create a linear combination of IVs to optimally predict DV. |
| ,, | Multiple (continuous) | Multiple (continuous) | - | Canonical R (see Chapter 12) | Maximally correlate a linear combination of DVs with a linear combination of IVs. |
| ,, | One (may be repeated) | Multiple (continuous and discrete; cases and IVs are nested) | - | Multilevel modeling (see Chapter 16) | Create linear combinations of IVs at one level to serve as DVs at another level. |
| ,, | None | Multiple (discrete) | - | Multiway frequency analysis (see Chapter 15) | Create a log-linear combination of IVs to optimally predict category frequencies. |
In studies where participants are randomly assigned to different groups (treatments), the main focus is usually on determining if there are statistically significant mean differences on dependent variables (DVs) associated with group membership. Once significant differences are identified, researchers often evaluate the strength of the relationship (effect size) between independent variables (IVs) and DVs. The choice of statistical techniques depends on factors such as the number of IVs and DVs, the consideration of certain variables as covariates, whether all DVs are measured on the same scale, and how within-subjects IVs are handled.
| Major research question | Number (and kind) of DVs | Number (and kind) of IVS | Covariates | Analytic strategy | Goal of analysis |
| Significance of group differences | One (continuous) | One (discrete) | None | One-way ANOVA or t test (see Chapter 3) | Determine significance of mean group differences. |
| ,, | ,, | ,, | Some | One-way ANCOVA (see Chapter 3 and 6) | ,, |
| ,, | ,, | Multiple (discrete) | None | Factorial ANOVA (see Chapter 17) | ,, |
| ,, | ,, | ,, | Some | Factorial ANCOVA (see Chapter 17) | ,, |
| ,, | Multiple (continuous) | One (discrete) | None | One-way MANOVA or Hotelling’s T 2 (see Chapter 7 and 9) | Create a linear combination of DVs to maximize mean group differences. |
| ,, | ,, | ,, | Some | One-way MANCOVA (see Chapter 7) | ,, |
| ,, | ,, | Multiple (discrete) | None | Factorial MANOVA (see Chapter 7) | ,, |
| ,, | ,, | ,, | Some | Factorial MANCOVA (see Chapter 7) | ,, |
| ,, | One (continuous) | Multiple (one discrete within S) | - | Profile Analysis of Repeated Measures (see Chapter 8) | Create linear combinations of DVs to maximize mean group differences and differences between levels of within subjects IVs |
| ,, | Multiple (continuous/ commensurate) | One (discrete) | - | Profile Analysis (see Chapter 8) | ,, |
| ,, | Multiple (continuous) | Multiple (one discrete within S) | - | Doubly multivariate profile analysis (see Chapter 8) | ,, |
In studies involving identified groups, the primary focus often revolves around predicting which group a participant belongs to based on a set of variables. To achieve this prediction, researchers employ techniques like discriminant analysis, logit analysis, and logistic regression. Discriminant analysis is suitable when all independent variables (IVs) are continuous and follow a normal distribution, logit analysis is used when all IVs are discrete, and logistic regression is preferred when the IVs consist of a combination of continuous and discrete variables, and/or when their distribution is not ideal.
| Major research question | Number (and kind) of DVs | Number (and kind) of IVS | Covariates | Analytic strategy | Goal of analysis |
| Prediction of group membership | One (discrete) | Multiple (continuous) | None | One-way discriminant function (see Chapter 9) | Create a linear combination of IVs to maximize group differences |
| ,, | ,, | ,, | Some | Sequential one-way discriminant function (see Chapter 9) | ,, |
| ,, | ,, | Multiple (discrete) | - | Multiway frequency analysis (logit) (see Chapter 16) | Create a log-linear combination of IVs to optimally predict DV. |
| ,, | ,, | Multiple (continuous and/or discrete) | None | Logistic regression (see Chapter 10) | Create a linear combination of the log of the odds of being in one group |
| ,, | ,, | ,, | Some | Sequential logistic regression (see Chapter 10) | ,, |
| ,, | Multiple (discrete) | Multiple (continuous) | None | Factorial discriminant function (see Chapter 9) | Create a linear combination of IVs to maximize group differences (DVs) |
| ,, | ,, | ,, | Some | Sequential factorial discriminant function (often questions can be rephrased to factorial MANCOVA) | ,, |
Another set of inquiries revolves around exploring the underlying structure of a group of variables. The choice of statistical techniques depends on whether the search for structure is data-driven (empirical) or theory-driven. Principal components analysis is an empirical approach, while factor analysis and structural equation modeling (SEM) are typically used in theory-driven analyses.
| Major research question | Number (and kind) of DVs | Number (and kind) of IVS | Covariates | Analytic strategy | Goal of analysis |
| Structure | Multiple (continuous observed) | Multiple (latent) | - | Factor analysis (theoretical) (see Chapter 13) | Create linear combinations of observed variables to represent latent variables. |
| ,, | Multiple (latent) | Multiple (continuous observed) | - | Principal components (empirical) (see Chapter 13) | ,, |
| ,, | Multiple (continuous observed and/or latent) | Multiple (continuous observed and/or latent) | - | Structural equation modeling (SEM) (see Chapter 14) | Create linear combinations of observed and latent IVs to predict linear combinations of observed and latent DVs. |
Two techniques focus on the time course of events. Survival/failure analysis asks how long it takes for something to happen. Time-series analysis looks at the change in a DV over the course of time.
| Major research question | Number (and kind) of DVs | Number (and kind) of IVS | Covariates | Analytic strategy | Goal of analysis |
| Time course of events | One (time) | None | None | Survival analysis (life tables) (see Chapter 11) | Determine how long it takes for something to happen. |
| ,, | ,, | One or more | None or some | Survival analysis (with predictors) (see Chapter 11) | Create a linear combination of IVs and CVs to predict time to an event. |
| ,, | One (continuous) | Time | None or some | Time-series analysis (forecasting) (see Chapter 18) | Predict future course of DV on basis of past course of DV. |
| ,, | ,, | One or more (including time) | None or some | Time-series analysis (intervention) (see Chapter 18) | Determine whether course of DV changes with intervention. |
Statistics are essential for making informed decisions when dealing with uncertain situations. Inferences or decisions are drawn about entire populations based on data collected from smaller samples, which may not fully represent the population. As a result, any conclusions about the population carry some degree of risk. To address this issue, statistical decision theory is commonly used. It involves formulating two hypothetical scenarios, each represented by a probability distribution, as alternative explanations for the data. Based on the observed sample results, formal statistical rules are applied to determine the most likely scenario, which serves as the best estimate for the underlying truth of events. However, it is essential to recognize that uncertainty still exists, and any conclusions drawn are subject to some level of uncertainty or risk.
The one-sample z-test is a statistical test used to compare the mean of a single sample to a known population mean or a hypothesized value. It assesses whether there is a significant difference between the sample mean and the population mean, helping researchers draw conclusions about the population based on the observed sample data. The test relies on the standard normal distribution and assumes that the sample is drawn from a normally distributed population or has a sufficiently large sample size for the central limit theorem to apply.
Note that, in hypothesis testing, we examine means rather than individual scores. Hypothetical scenarios represent distributions of means, not individual scores. These "sampling distributions of means" differ from distributions of individual scores in a systematic manner. The focus is on comparing and drawing conclusions about mean values rather than individual data points. Sampling distributions have smaller standard deviations than distributions of scores, and the decrease is related to N, the sample size.
H0 (null hypothesis) and Ha (alternative hypothesis) represent two possible scenarios, only one of which is true. When researchers have to decide whether to accept or reject H0, four possible outcomes can occur. If H0 is true, a correct decision is made when retaining H0, and an error is made when rejecting it. The probability of making the wrong decision (Type I error) is denoted as "a," while the probability of making the correct decision is "1 - α." On the other hand, if Ha is true, a correct decision is made when rejecting H0, and an error occurs when retaining it. The probability of making the wrong decision (Type II error) is denoted as "β," and the probability of making the correct decision is "1 - β." These probabilities are summarized in a "confusion matrix," illustrating the chances of each of the four possible outcomes.
| Reality= H0 | Reality = Ha | |
| Statistical decision = H0 | 1 - α | β (type II error) |
| Statistical decision = Ha | α (type I error) | 1 - β (power) |
Statistical power refers to the probability of correctly detecting a true effect or relationship in a statistical analysis. In other words, it is the likelihood that a statistical test will correctly reject the null hypothesis when the alternative hypothesis is true. A high statistical power indicates that the test is more likely to identify real effects, whereas a low statistical power suggests a higher chance of missing true effects or failing to detect significant relationships. A well-powered study is essential for obtaining reliable and meaningful results, as it increases the chances of detecting real effects and making accurate conclusions based on the data. There is occasionally the danger of too much power. The null hypothesis is probably never exactly true and any sample is likely to be slightly different from the population value. With a large enough sample, rejection of H0 is virtually certain.
The z-test for the difference between a sample mean and a population mean can be extended to compare the difference between two sample means. Under the null hypothesis that the two means are equal, a sampling distribution is generated, and a decision axis is set up. The power of an alternative hypothesis is then calculated based on this decision axis. When population variances are unknown, it is preferable to use Student's t-test instead of the z-test, even for large samples.
Analysis of Variance (ANOVA) is a statistical technique used to compare the means of two or more groups to determine if there are significant differences among them. It assesses the variation between group means and within groups, providing insights into whether the observed differences are due to chance or meaningful effects. ANOVA is widely used in various fields to analyze experimental or observational data and identify significant factors that influence the outcome variable. The technique can be extended to include multiple factors and interactions, allowing researchers to explore more complex relationships in the data. Overall, ANOVA is a powerful tool for hypothesis testing and investigating group differences in a wide range of research settings.
When an IV has more than one degree of freedom (more than two levels) or when there is an interaction between two or more IVs, the overall test of the effect is ambiguous. The overall test, with (k - 1) degrees of freedom, is pooled over (k - 1) single-degree-of-freedom subtests. If the overall test is significant, so usually are one or more of the subtests, but there is no way to tell which one(s). To find out which single- degree-of-freedom subtests are significant, comparisons are performed. Comparison of treatment means begins by assigning a weighting factor (w) to each of the cell or marginal means so the weights reflect your null hypotheses. Once the weighting coefficients are chosen, an equation is used to obtain F for the comparison. If obtained F exceeds critical F, the null hypothesis for the comparison is rejected.
One-way between-subjects ANOVA is a statistical technique used to compare the means of three or more independent groups to determine if there are significant differences among them. It is called "one-way" because it involves a single factor (independent variable) with multiple levels or groups. The term "between-subjects" indicates that each participant belongs to only one group and is not exposed to multiple conditions.
The one-way between-subjects ANOVA tests the null hypothesis that there is no difference in the population means of the groups. It assesses the variation between group means and within groups, and if the observed differences are significant, it indicates that at least one group mean is different from the others. Post-hoc tests, such as Tukey's Honestly Significant Difference (HSD) or Bonferroni tests, may be conducted to identify which specific groups differ from one another if the overall ANOVA result is significant.
One-way between-subjects ANOVA can be used for two groups as well. When there are only two groups (i.e., a binary comparison), one-way between-subjects ANOVA essentially becomes equivalent to an independent samples t-test. In the context of two groups, the null hypothesis in one-way between-subjects ANOVA is the same as that in the independent samples t-test: it tests whether there is no difference in the population means between the two groups. The advantage of using one-way ANOVA for two groups is that it allows for the extension to more than two groups if needed, making it a more flexible and general statistical approach. However, when there are only two groups, both the one-way ANOVA and the independent samples t-test will yield the same p-value and significance level for the comparison between the two means. So, researchers can choose either method based on their preference and the potential for expanding the analysis to more groups in the future.
Factorial between-subjects ANOVA is a statistical technique used to analyze the effects of two or more independent variables (factors) on a single dependent variable in an experiment with independent groups. It extends the concept of one-way between-subjects ANOVA, which compares the means of multiple independent groups based on a single factor, to situations where there are multiple factors and their combined effects need to be examined.
In a factorial between-subjects ANOVA, each participant is assigned to only one combination of the levels of the factors. The independent variables can be categorical (e.g., different treatment conditions, gender, or age groups) or quantitative (e.g., different levels of dosage or exposure). The dependent variable is typically a continuous outcome that is measured or observed for each participant.
The analysis involves testing the main effects of each factor (the individual effects of each independent variable) and the interaction effects (the combined effects of two or more independent variables). The main effects represent the average differences between the levels of each factor, while the interaction effects demonstrate whether the effects of one factor depend on the levels of another factor.
Within-subjects ANOVA, also known as repeated-measures ANOVA, is a statistical technique used to analyze the effects of one or more independent variables (factors) on a single dependent variable in an experiment with related or paired observations. Unlike between-subjects ANOVA, where each participant is assigned to only one condition or group, within-subjects ANOVA involves measuring the same participants under different conditions or at multiple time points.
In within-subjects ANOVA, each participant serves as their control, and the same individuals are exposed to all the levels of the independent variable(s). This design is particularly useful when comparing the effects of different treatments or conditions within the same individuals, which reduces the impact of individual differences and increases statistical power.
The analysis involves testing the main effects of each factor (the average differences between the levels of each independent variable) and the interaction effects (the combined effects of two or more independent variables). By considering the variations within each participant, within-subjects ANOVA can detect smaller effects and increase the sensitivity of the statistical test.
Mixed between-within-subjects ANOVA, also known as mixed design ANOVA, is a statistical technique that combines features of both between-subjects and within-subjects ANOVA. It is used to analyze experiments with two or more independent variables, where at least one independent variable is manipulated between participants (between-subjects factor), and another independent variable is manipulated within participants (within-subjects factor).
In a mixed design ANOVA, participants are divided into different groups based on the between-subjects factor, and each participant is measured or observed under multiple conditions corresponding to the within-subjects factor. For example, in a study examining the effects of a new drug (between-subjects factor) on cognitive performance measured at different time points (within-subjects factor), some participants may receive the drug, and others may receive a placebo, while all participants are assessed at multiple time points.
The analysis involves testing the main effects of each independent variable (between-subjects and within-subjects factors) and their interaction effect. The between-subjects factor compares the differences between the groups, while the within-subjects factor compares the differences within each participant. The interaction effect examines whether the combined effects of the two independent variables are significant.
Mixed design ANOVA allows researchers to explore the unique and combined effects of different factors on the dependent variable, considering both between-subjects and within-subjects sources of variability. It is a powerful and versatile statistical approach, often used in experimental designs that involve a combination of independent groups and repeated measures to gain deeper insights into the experimental effects and interactions.
In real research settings, one may encounter difficulties in the design and data structure. A few of the more common types of design complexity are:
Nesting refers to the relationship between subjects and levels of independent variables (IVs) in between-subjects designs. It means that each subject is restricted to one level of an IV or a combination of IVs, and in some cases, levels of one IV are limited to only one level of another IV, rather than crossing over all possible combinations.
To get an uncontaminated look at the effects of the IV, it is important to counterbalance the effects of increasing experience, time of day, and the like, so that they are independent of levels of the IV. This can be done by a Latin-square design.
In a simple one-way between-subjects ANOVA, having slightly unequal group sizes is not a major concern and can be handled with ease using computer-based analysis. However, as the group sizes become significantly different, it becomes crucial to consider the assumption of homogeneity of variance. If the group with the smaller sample size also has a larger variance, the F-test becomes overly lenient, leading to an increased risk of Type I errors (false positives or inflated alpha level).
In factorial designs with more than one between- subjects IV, unequal sample sizes in each cell create difficulty in computation and ambiguity of results. With unequal n, a factorial design is nonorthogonal. The simplest strategy is to randomly delete cases from cells with greater n until all cells are equal. If unequal group size is due to random loss of a few subjects in an experimental design originally set up for equal size, deletion is often a good choice. An alternative strategy with random loss of subjects in an experimental design is an unweighted-means analysis (see Chapter 6).
Unequal sample sizes in nonexperimental research reflect true differences in subject characteristics. Artificially equalizing them may distort findings, so researchers adjust tests for overlapping variance to account for the unequal n and maintain validity.
In fixed-effects ANOVA, researchers select specific levels of IVs for significance testing. In random-effects ANOVA, levels are randomly chosen from the population to generalize results beyond the selected sample. Computer programs can analyze random-effects, but their use is less common.
While significance testing and comparisons reveal group differences, they do not measure the strength of the relationship between IV(s) and DV. Effect size is vital to avoid promoting insignificant results as practically meaningful. Effect size quantifies the proportion of DV variance associated with IV levels, assessing the overlap of circles representing their total variances. It complements statistical significance testing, offering a comprehensive understanding of the relationship between variables.
A rough estimate of effect size is available for any ANOVA through η2 (eta squared). The formula is:
\[ \eta^{2} = \frac{SS_{effect}}{SS_{total}} \]
When there are two levels of the IV, η2 represents the squared point biserial correlation between the continuous variable (DV) and the dichotomous variable (two levels of the IV). It shows the proportion of DV variance attributed to the effect after finding significant main effects or interactions. However, η2 has limitations as it depends on the number and significance of other IVs in the design, leading to partial η2 being proposed as an alternative measure, containing only variance attributable to the effect of interest plus error.
\[ Partial \eta^{2} = \frac{SS_{effect}}{SS_{effect} + SS_{error}} \]
However, η2 describes the proportion of systematic variance in a sample with no attempt to estimate the proportion of systematic variance in the population. A statistic developed to estimate effect size between IV and DV in the population is ω2 (omega squared).
\[ \omega^{2} = \frac{SS_{effect} - (df_{effect})(MS_{error})}{SS_{total} + MS_{error}} \]
This is the additive form of ω2, where the denominator represents total variance, not just variance due to effect plus error, and is limited to between-subjects analysis of variance designs with equal sample sizes in all cells.
Pearson correlation is a statistical measure that quantifies the linear relationship between two continuous variables. It ranges from -1 to +1, where -1 indicates a perfect negative linear relationship, +1 indicates a perfect positive linear relationship, and 0 indicates no linear relationship. The most interpretable equation for Pearson r is:
\[ r = \frac{ \Sigma Z_{X}Z_{Y} }{N - 1} \]
where Pearson r is the average cross- product of standardized X and Y variable scores.
\[ Z_{X} = \frac{X - \bar{X}}{S} and Z_{Y} = \frac{Y - \bar{Y}}{S} \]
Whereas correlation is used to measure the size and direction of the linear relationship between two variables, regression is used to predict a score on one variable from a score on the other. To find the best- fitting straight line, the following formula is used:
\[ Y'= A + BX \]
where Y' is the predicted score, A is the value of Y when X is zero, B is the slope of the line (change in Y divided by change in X), and X is the value from which Y is to be predicted. The intercept (denoted by A) is the mean of the observed value of the predicted variable minus the product of the regression coefficient times the mean of the predictor variable.
\[ A = \bar{Y} -B\bar{X} \]
The difference between the predicted and the observed values of Y at each value of X repre[1]sents errors of prediction or residuals. The best- fitting straight line is the line that minimizes the squared errors of prediction.
\[ B = \frac{N\Sigma{XY} - (\Sigma{X})(\Sigma{Y})}{N\Sigma{X}^{2} - (\Sigma{X})^{2}} \]
The bivariate regression coefficient, B, is the ratio of the covariance of the variables and the variance of the one from which predictions are made.
Chi-square analysis, also known as chi-square test, is a statistical method used to determine if there is a significant association between two categorical variables. It is used to compare observed frequencies of categories in a contingency table with the expected frequencies under the assumption of independence. The test assesses whether any deviation from expected frequencies is due to chance or if there is a meaningful relationship between the variables. Chi-square analysis is commonly used in various fields, such as social sciences, biology, and market research, to examine the association between categorical variables and make inferences about their relationship.
\[ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} \]
where:
The sum is taken over all categories or cells in the contingency table. The chi-square test statistic is used to assess the difference between the observed and expected frequencies and determine if there is a significant association between the two categorical variables.
Usually, the expected frequencies for a cell are generated from the sum of its row and the sum of the column.
\[ Cell \hspace{2mm} F_{e} = \frac{(row \hspace{2mm} sum)(column \hspace{2mm} sum)}{N} \]
Using this procedure to generate expected frequencies tests the null hypothesis of independence between the row variable (e.g., region of the country) and the column variable (e.g., attitude toward current political leadership). A good fit with a small X2 suggests independence, while a poor fit with a large X2 leads to rejecting the null hypothesis and concluding a relationship between the variables.
This chapter focuses on data preparation and resolution of various issues before the analysis to ensure an honest and accurate data analysis, which may be time-consuming but is crucial for the integrity of the results. The topics covered in this chapter include:
To ensure data accuracy in small data files, proofreading the original data against the computerized data file is recommended, while for large data sets, screening involves examining univariate descriptive statistics and graphic representations, checking for plausible values, and accurate coding for missing values.
Most multivariate procedures analyze patterns of correlation (or covariance) among variables. It is important that the correlations, whether between two continuous variables or between a dichotomous and continuous variable, be as accurate as possible. Under some rather common research conditions, correlations are larger or smaller than they should be.
Inflated correlation refers to a correlation between variables that appears higher than it should be due to the presence of shared items or overlapping components in the measurement of those variables. In social and behavioral sciences, variables are often composites of several items. If composite variables are used and they contain, in part, the same items, correlations are inflated; do not overinterpret a high correlation between two measures composed, in part, of the same items. If there is enough overlap, consider using only one of the composite variables in the analysis.
Deflated correlation refers to a correlation between variables that appears lower than it should be due to the absence or exclusion of shared items or overlapping components in the measurement of those variables. A falsely small correlation between two continuous variables is obtained if the range of responses to one or both of the variables is restricted in the sample. When a correlation is too small because of restricted range in sampling, you can estimate its magnitude in a nonrestricted sample.
One of the most pervasive problems in data analysis is missing data, which occurs when rats die, equipment malfunctions, respondents become recalcitrant, or somebody goofs. The pattern of missing data is crucial, with randomly scattered missing values being less problematic than nonrandomly missing ones, as nonrandom missing data can significantly affect result generalizability. In scenarios like questionnaires with demographic and attitudinal questions, refusing to answer specific questions (e.g., income) may be related to other variables (attitudes), necessitating methods for estimating the missing data to ensure accurate analysis and unbiased results.
MCAR is the best of all possible worlds if data must be missing, because the distribution of missing data is unpredictable. When data are MAR, the pattern of missing data is predictable from other variables. In NMAR, the missingness is related to the variable itself and, therefore, cannot be ignored. IBM SPSS MVA (Missing Values Analysis) is specifically designed to highlight patterns of missing values as well as to replace them in the data set. The decision about how to handle missing data is important. At best, the decision is among several bad alternatives, several of which are discussed in the subsections that follow.
One approach to dealing with missing values is to remove cases with missing data. If the number of cases with missing values is small and appears to be a random subset of the entire sample, deletion is a suitable option, and it is the default choice in many statistical software packages like IBM SPSS and SAS. When missing values are concentrated in a few non-critical variables or highly correlated with complete variables, dropping those variables is a reasonable option. However, if missing values are scattered throughout cases and variables, deleting cases can lead to significant subject loss, which can be particularly problematic in experimental designs requiring adjustment for unequal sample sizes. Additionally, deleting cases with non-random missing values can distort the sample, making it undesirable to discard data after considerable effort in data collection.
The second option for handling missing values is to estimate (impute) them using various methods, including using prior knowledge, mean values, regression, expectation-maximization, or multiple imputation. Prior knowledge involves educated guesses or downgrading continuous variables to predict missing values. Longitudinal data can use the last observed value to fill in missing data. Alternatively, software offers various imputation methods for handling missing values.
Other options to deal with missing data are:
An outlier is an extreme case with unusual values on one or more variables that can distort statistics, affecting both univariate and multivariate situations and leading to Type I and Type II errors with uncertain effects on the analysis.Outliers can be present due to incorrect data entry, failure to specify missing-value codes, cases not belonging to the intended population, or extreme distributions in the population. Identifying and addressing errors in data entry and missing values are straightforward, but deciding whether to delete or retain and alter outliers requires careful consideration.
Univariate outliers are cases with an outlandish value on one variable. They are easier to spot. Multivariate outliers are cases with an unusual combination of scores on two or more variables. To spot univariate outliers, you can use graphical methods like boxplots or histograms to identify extreme values. For multivariate outliers, techniques like scatterplots, Mahalanobis distance, or cluster analysis can be used to detect cases that deviate from the general pattern of data points.
Once multivariate outliers are identified, you need to discover why the cases are extreme (You already know why univariate outliers are extreme). To handle multivariate outliers, you should identify the variables on which the extreme cases deviate to determine their relevance to the sample, potential modifications, and implications for generalization. For a small number of outliers, individual examination is suitable, while for several outliers, group analysis can be done using dummy grouping variables in statistical techniques like discriminant analysis, logistic regression, or regression to identify distinguishing variables.
To reduce the impact of univariate outliers, first, check for data accuracy and consider removing variables responsible for most outliers or those that are highly correlated or non-critical. If outliers are part of the intended population, options include variable transformation or altering scores for outlying cases. For truly multivariate outliers, deletion is often necessary. All changes made to address outliers should be reported in the Results section with a rationale.
Underlying some multivariate procedures and most statistical tests of their outcomes is the assumption of multivariate normality. Multivariate normality is the assumption that each variable and all linear combinations of the variables are normally distributed. When the assumption is met, the residuals of analysis are also normally distributed and independent. The assumption of multivariate normality is not readily tested because it is impractical to test an infinite number of linear combinations of variables for normality. Those tests that are available are overly sensitive.
If there is multivariate normality in ungrouped data, each variable is itself normally distributed and the relationships between pairs of variables, if present, are linear and homoscedastic: the variance of one variable is the same at all values of the other variable. The assumption of multivariate normality can be partially checked by examining the normality, linearity, and homoscedasticity of individual variables or through examination of residuals in analyses involving prediction. The assumption is certainly violated, at least to some extent, if the individual variables (or the residuals) are not normally distributed or do not have pairwise linearity and homoscedasticity.
Screening continuous variables for normality is an essential early step for multivariate analysis, especially when making inferences. Normality is assessed through statistical or graphical methods, considering skewness (symmetry) and kurtosis (peakedness) of the distribution. Nonnormal variables can degrade the analysis, and their presence should be carefully addressed. Conventional but conservative (.01 or .001) alpha levels are used to evaluate the significance of skewness and kurtosis with small to moderate samples, but if the sample is large, it is a good idea to look at the shape of the distribution instead of using formal inference tests. Because the standard errors for both skewness and kurtosis decrease with larger N, the null hypothesis is likely to be rejected with large samples when there are only minor deviations from normality. In a large sample, the significance level of skewness is not as important as its actual size (worse the farther from zero) and the visual appearance of the distribution.
The assumption of linearity in statistical analysis assumes a straight-line relationship between two variables. Nonlinearity can be identified through residuals plots or bivariate scatterplots, and transformation or recoding may be needed to address nonlinearity in the data. Assessing linearity through bivariate scatterplots can be time-consuming, especially with numerous variables. To simplify the process, you can focus on pairs of variables that are likely to exhibit nonlinearity based on skewness statistics and examine them through bivariate scatterplots using statistical software like IBM SPSS GRAPH or SAS PLOT.
Homoscedasticity, also known as homogeneity of variance, is a statistical assumption that states that the variance of the dependent variable is constant across all levels of the independent variable(s) in a regression or analysis of variance (ANOVA) model. In simpler terms, it means that the spread of data points around the regression line or group means is consistent, regardless of the values of the independent variable(s). When homoscedasticity is met, it ensures that the model's errors have a consistent level of variability, which is essential for accurate statistical inference.
Heteroscedasticity, the failure of homoscedasticity, is caused either by nonnormality of one of the variables or by the fact that one variable is related to some transformation of the other. Another source of heteroscedasticity is a greater error of measurement at some levels of an IV. For example, people in the age range 25 to 45 might be more concerned about their weight than people who are younger or older. Older and younger people would, as a result, give less reliable estimates of their weight, increasing the variance of weight scores at those ages.
Multicollinearity refers to a high degree of correlation among two or more predictor variables in a regression model. When predictor variables are highly correlated, it becomes challenging for the regression model to distinguish the individual effects of each variable on the dependent variable. This can lead to unstable estimates of regression coefficients and difficulty in interpreting the significance of individual predictors. Multicollinearity does not directly affect the predictive power of the model, but it can make it harder to identify the true relationships between predictors and the outcome variable.
Singularity occurs when one or more predictor variables are perfectly linearly related to each other, leading to a rank-deficient design matrix. In other words, one or more predictor variables can be expressed as a linear combination of other predictor variables in the model. This situation makes it impossible for the regression model to estimate unique regression coefficients for each predictor. As a result, the model may not be identifiable, and the estimates become unreliable. Singularity is a severe problem and needs to be resolved before valid regression results can be obtained. One common approach to address singularity is to remove one of the linearly dependent predictor variables from the model.
Finally, a checklist is provided for screening data. The order in which screening takes place is important because the decisions that you make at one step influence the outcomes of later steps.
Multiple regression is a statistical technique used to analyze the relationship between a dependent variable and two or more independent variables. It extends simple linear regression to account for multiple predictors simultaneously. The method estimates the unique contribution of each independent variable to the dependent variable while controlling for the effects of other predictors. Multiple regression is widely employed in various fields to model and predict complex relationships between variables and is useful for hypothesis testing, prediction, and understanding the influence of multiple factors on an outcome.
The standard form of the equation for multiple regression with K independent variables (predictors) can be represented as follows:
Y' = A + B1X1 + B2X2 + ... + BKXK
Where:
The best-fitting regression coefficients produce a prediction equation for which squared differences between Y and Y' are at a minimum. Because squared errors of prediction, (Y - Y) 2 , are minimized, this solution is called a least- squares solution.
The aim of multiple regression is to model and analyze the relationship between a dependent variable (response variable) and two or more independent variables (predictor variables). The main objective is to understand how the independent variables jointly influence the dependent variable and to estimate the regression coefficients (weights) that represent the strength and direction of these relationships. This enables us to make predictions, understand the relative importance of different predictors, and identify which variables have a significant impact on the outcome of interest. Multiple regression is a fundamental tool for analyzing complex relationships.
Multiple regression can be used to answer a wide range of research questions. Some examples of research questions that can be addressed using multiple regression include:
Multiple regression, like any statistical technique, has its limitations. Some of the key limitations include:
Researchers should be aware of these limitations and consider them when interpreting the results of multiple regression analyses. Using additional techniques, such as model diagnostics and sensitivity analyses, can help address some of these limitations and provide more robust results.
There are three major analytic strategies in multiple regression: standard multiple regression, sequential (hierarchical) regression, and statistical (stepwise) regression.
In standard multiple regression, all predictor variables are entered into the regression equation simultaneously. The goal is to examine the unique contribution of each predictor variable to the outcome variable while controlling for the effects of other predictors. This method is straightforward and provides a baseline for assessing the relationship between the outcome and each predictor independently.
In sequential regression, predictor variables are entered into the equation in a predetermined order, often based on theoretical or practical considerations. Researchers can test the effects of different sets of predictors in stages, allowing them to evaluate how the model improves with the addition of each set of variables. This approach is useful for studying the incremental contribution of specific variables or sets of variables to the prediction of the outcome variable.
Statistical regression (stepwise regression) methods automatically select variables to be included in the model based on statistical criteria, such as the F-statistic, p-values, or measures of model fit. Stepwise regression proceeds by adding or removing variables from the model based on their statistical significance, usually starting with a full model and iteratively selecting the best predictors until a stopping criterion is met. While stepwise regression can be convenient, it has some drawbacks, such as potential overfitting and unstable results.
Each of these analytic strategies has its advantages and limitations, and the choice of which method to use depends on the research question, the nature of the data, and the underlying assumptions of the model. It is crucial for researchers to carefully consider the appropriate strategy for their specific research context and interpret the results with caution.
Analysis of variance is a part of the general linear model, as is regression. Indeed, ANOVA can be viewed as a form of multiple regression in which the IVs are levels of discrete variables rather than the more usual continuous variables of regression. In other words, Analysis of Variance (ANOVA) is a special case of regression with dichotomous independent variables (IVs), while multiple regression can handle both dichotomous and continuous IVs, making it more flexible for analyzing a wider range of data and preserving the full range of continuous IVs. Converting multiple regression problems into ANOVA may lead to information loss and unequal cell sizes, making regression a more versatile and powerful tool for certain analyses.
A research question one might answer using multiple regression: Does the number of hours spent studying (X1) and the level of test anxiety (X2) predict students' exam scores (Y)?
Hypotheses:
Data Collection: A sample of psychology students is recruited, and each participant is asked to provide the following information:
Data Analysis: The data collected from the sample is subjected to multiple regression analysis using appropriate statistical software (e.g., SPSS).
Interpretation of results: The results of the multiple regression analysis will provide insights into the relationship between the predictor variables (hours spent studying and test anxiety level) and the outcome variable (exam scores). The key components to interpret include:
Conclusion: Based on the results of the multiple regression analysis, the researchers can draw conclusions regarding the predictive power of the number of hours spent studying and test anxiety level on students' exam scores. If the predictors are found to be significant, it implies that studying more and managing test anxiety can be important factors for improving exam performance.
However, it's crucial to remember that multiple regression analysis cannot establish causation, and any conclusion should be interpreted with consideration of the study's limitations and the broader context of the research question.
Analysis of Covariance (ANCOVA) is a statistical technique that combines elements of both analysis of variance (ANOVA) and regression analysis. It is used to compare the means of two or more groups while controlling for the influence of one or more continuous covariates (also known as independent variables or predictors) on the outcome variable. In ANCOVA, the covariates are continuous variables that could potentially influence the outcome variable. ANCOVA extends the traditional ANOVA by accounting for the covariates' influence, thus allowing researchers to isolate the unique effect of the grouping variable on the outcome variable while controlling for potential confounding variables. This helps in reducing variability and increasing the statistical power to detect group differences accurately. The primary goal of ANCOVA is to examine whether there are significant differences between the group means on the outcome variable after adjusting for the effects of the covariates.
ANCOVA is a versatile statistical technique that can be used to answer various research questions in different fields. Some common research questions that can be addressed using ANCOVA include:
Overall, ANCOVA is a valuable tool for researchers who wish to investigate group differences or the impact of interventions while taking into account the influence of other variables that might affect the outcome. It is particularly useful when researchers want to draw more robust conclusions by statistically adjusting for potential sources of variability or bias in their data.
ANCOVA, like any statistical method, has its limitations. Some of the main limitations include:
Despite these limitations, ANCOVA remains a valuable tool in many research settings. Researchers should be cautious and use additional diagnostic tools, such as residual plots and sensitivity analyses, to assess the validity of ANCOVA results and ensure their conclusions are sound.
Because of the stringent limitations to ANCOVA and potential ambiguity in interpreting results of ANCOVA, alternative analytical strategies are often sought. The availability of alternatives depends on such issues as the scale of measurement of the CV(s) and the DV, the time that elapses between measurement of the CV and assignment to treatment, and the difficulty of interpreting results.
When the continuous covariate (CV) and the dependent variable (DV) are measured on the same scale, two alternatives are available for analysis:
Difference scores: Compute the difference between the pretest score (previous CV) and the posttest score (previous DV) for each subject. Use these difference scores as the DV in an ANOVA. This approach is suitable for research questions focused on "change" between pretest and posttest measurements.
Conversion of the pretest and posttest scores into a within-subjects IV: Treat the pretest scores as a within-subjects independent variable (IV) and posttest scores as the dependent variable. Conduct ANCOVA to examine the effects of the main independent variable (e.g., treatment group) while adjusting for pretest differences in the DV.
However, both approaches have limitations. Change scores may be affected by ceiling or floor effects and may be less reliable than pre- or posttest scores. When CVs are measured on a continuous scale, other alternatives include the randomized-block design and blocking, where subjects are matched into blocks based on CV scores and treated as if they were the same person in a within-subjects analysis. These alternatives may have their own assumptions and limitations, such as the strong assumption of sphericity and loss of degrees of freedom for error without commensurate loss of sums of squares if the blocking variables are not highly related to the DV. Researchers should carefully consider the appropriateness of each approach based on their research question, data characteristics, and assumptions.
Blocking is an alternative method to handle continuous covariates in research studies. In this approach, subjects are first measured on potential covariates, and then they are grouped based on their scores into different levels of another independent variable. These groups are then combined with the levels of the main independent variable in a factorial design. The main advantage of blocking is that it does not require the same assumptions as ANCOVA or within-subjects ANOVA, making it more versatile and robust. Additionally, blocking can handle non-linear relationships between the potential covariates and the dependent variable, providing greater flexibility in analysis. It also allows for the removal of variation due to potential covariates from the error variance estimation, and any violation of homogeneity of regression assumptions in ANCOVA can be detected through interactions in blocking. Overall, blocking is a powerful and flexible method, especially suited for experimental research and situations where the relationships between covariates and the dependent variable are not linear.
Blocking can be extended to handle multiple continuous covariates by developing new independent variables for each covariate and crossing them in factorial design. However, as the number of covariates increases, the design becomes cumbersome.
In some cases, ANCOVA may be preferred over blocking, especially when the relationship between the dependent variable and covariate is linear, and ANCOVA is more powerful. Additionally, practical limitations may hinder the measurement of covariates before treatment, leading to unequal sample sizes in cells.
A combination of blocking and ANCOVA may be the best approach for some applications, using some covariates as new independent variables and others as covariates. Another alternative is multilevel modeling, which handles heterogeneity of regression by creating a second level of analysis with groups and allowing different slopes and intercepts for each group.
Research Question: Does a new teaching method improve students' math test scores compared to the traditional teaching method, after controlling for their previous math scores?
Study design:
Data: Let's say the data looks like this:
Group | Pretest score | Posttest score |
A | 70 | 85 |
A | 65 | 80 |
A | 75 | 90 |
B | 68 | 78 |
B | 62 | 72 |
B | 72 | 82 |
Analysis: We will conduct an ANCOVA to compare the posttest scores of the two groups while controlling for the effect of pretest scores. The model will be:
Posttest score = β0 + β1 x Group + β2 x Pretest score + ε
Where:
Interpretation: If the coefficient for the group variable (β1) is statistically significant and positive, it would indicate that the new teaching method led to higher posttest scores compared to the traditional method. The covariate effect of pretest score (β2) will help control for the influence of students' initial math abilities.
By conducting ANCOVA, we can assess whether the observed differences in posttest scores between the two groups are significant after accounting for their initial math abilities, thus providing more accurate and reliable conclusions about the effectiveness of the new teaching method.
MANOVA and MANCOVA are statistical techniques for analyzing multiple dependent variables in relation to one or more independent variables. MANOVA examines how the independent variables affect the dependent variables simultaneously, providing efficiency and multivariate insights. On the other hand, MANCOVA extends MANOVA by including continuous covariates, offering the advantage of controlling for confounding variables and increasing precision in estimating the effects of independent variables on dependent variables. However, both methods have assumptions to consider, and MANCOVA may require a larger sample size.
Multivariate Analysis of Variance (MANOVA) is a statistical technique used to analyze the relationships among several dependent variables simultaneously when they are influenced by one or more independent variables. It extends the concept of univariate Analysis of Variance (ANOVA) to situations where there are multiple dependent variables. MANOVA is particularly useful when you want to understand how different independent variables impact multiple related dependent variables.
The advantages of using MANOVA include its efficiency, as it can provide more powerful results compared to conducting separate ANOVA tests for each dependent variable. By analyzing multiple dependent variables together, MANOVA offers a more comprehensive understanding of their interrelationships and how they respond to changes in the independent variables. Additionally, the statistical power of MANOVA tends to be higher than conducting univariate analyses.
However, MANOVA has some disadvantages. One significant assumption is that the data should follow a multivariate normal distribution, and the variance-covariance matrices should be homogeneous across groups. Violating these assumptions can lead to biased results. Moreover, interpreting MANOVA results can be more complex than interpreting ANOVA results due to the multivariate nature of the analysis. Lastly, MANOVA typically requires a larger sample size, especially when there are multiple dependent variables.
Multivariate Analysis of Covariance (MANCOVA) is a technique that extends MANOVA by incorporating one or more continuous covariates (control variables) in addition to the independent variables. Including covariates allows researchers to control for the influence of these continuous variables while examining the effects of the independent variables on multiple dependent variables.
The primary advantage of MANCOVA is that it enables researchers to control for confounding variables. By including covariates in the analysis, it becomes possible to isolate the effects of the independent variables on the dependent variables more accurately. This increased precision can be beneficial in drawing more accurate conclusions from the data.
However, just like MANOVA, MANCOVA has its own set of assumptions and limitations. The data still need to satisfy the assumptions of multivariate normality, homogeneity of variance-covariance matrices, and independence of observations. Violations of these assumptions can compromise the validity of the results. Moreover, including covariates in the analysis can add complexity to result interpretation. Additionally, using MANCOVA often requires a larger sample size to produce reliable and meaningful outcomes.
Multivariate Analysis of Variance (MANOVA) and Multivariate Analysis of Covariance (MANCOVA) are statistical methods used to answer research questions that involve multiple dependent variables (DV) and one or more independent variables (IV) or covariates. These techniques are suitable for studying the relationships between several DVs simultaneously while considering the effects of one or more IVs or covariates.
Examples of research questions that can be addressed using MANOVA or MANCOVA include:
In summary, MANOVA and MANCOVA are valuable tools for investigating the relationships between multiple dependent variables and independent variables or covariates, allowing researchers to examine complex interactions and patterns in their data.
MANOVA and MANCOVA are powerful statistical techniques, but they also have some limitations that researchers should be aware of:
Despite these limitations, MANOVA and MANCOVA remain valuable tools for analyzing data with multiple dependent variables and studying the effects of multiple independent variables or covariates. Researchers should be cautious and use these methods appropriately while considering the assumptions and potential issues associated with their data.
Wilks' Lambda (also known as Wilks' Λ) is a statistical test used to assess the significance of the overall multivariate effect. It is a measure of the proportion of variance in the dependent variables that is not accounted for by the independent variables (or covariates) included in the model.
Wilks' Lambda is computed as the ratio of the determinant of the residual variance-covariance matrix to the determinant of the total variance-covariance matrix. The formula for Wilks' Lambda is as follows:
Λ = |W| / |T|
where:
Wilks' Lambda is used as a test statistic in MANOVA and MANCOVA to determine whether there are significant differences between groups on the combined set of dependent variables. It follows the multivariate F-distribution and is used to assess the overall significance of the model. A significant Wilks' Lambda indicates that there are significant differences among the groups on the combined dependent variables, suggesting that the groups are not equal across all dependent variables.
Researchers use Wilks' Lambda in conjunction with other multivariate test statistics, such as Pillai's Trace, Hotelling's Trace, or Roy's Largest Root, to comprehensively evaluate the multivariate effects and make informed decisions about the significance of group differences in MANOVA and MANCOVA analyses.
A measure of effect size is readily available from Wilks’ lambda. For MANOVA:
η2 = 1 – Λ
This equation represents the variance accounted for by the best linear combination of DVs. However, unlike η2 in the analogous ANOVA design, the sum of η2 for all effects in MANOVA may be greater than 1.0 because DVs are recombined for each effect. This lessens the appeal of an interpretation in terms of proportion of variance accounted for, although the size of η2 is still a measure of the relative importance of an effect. Another difficulty in using this form of η2 is that effects tend to be much larger in the multivariate than in the univariate case. Therefore, a recommended alternative when s > 1.
Partial η2 = 1 - Λ1/s
The estimated effect size is reduced to with the use of partial η2 for data with s > 1, a more reasonable assessment hence.
MANOVA is most effective when the dependent variables (DVs) have moderate negative correlations (around -0.6) or are moderately correlated in either direction. For instance, two DVs like "time to complete a task" and "number of errors" that are expected to have moderate negative correlation are suitable for MANOVA. However, MANOVA becomes less suitable when DVs have very high positive correlations or are nearly uncorrelated. When DVs have very high positive correlations, the multivariate test might show significant effects, but it becomes difficult to interpret individual effects on each DV after considering the highest priority DV. Similarly, if DVs are uncorrelated or represented by factor/component scores, MANOVA becomes less powerful compared to separate univariate analyses. In cases where there is a mix of correlated and uncorrelated DVs, conducting separate MANOVAs on each set of moderately correlated DVs with proper adjustments for familywise error rate is a more meaningful approach. Alternatively, one set of DVs might be used as covariates in a single MANCOVA analysis.
In MANOVA, multiple statistics can be used to test the significance of main effects and interactions, such as Wilks’ lambda, Hotelling’s trace criterion, Pillai’s criterion, and Roy’s greatest characteristic root (gcr) criterion. When there is only one degree of freedom for an effect, these statistics produce identical results. However, when there are more than two levels for an effect, the F values may slightly differ, leading to confusion on which result to believe. When there are multiple degrees of freedom, these statistics pool the results from different ways of combining dependent variables to test the effect. Roy’s gcr criterion uses only the first dimension, while the others pool results from all dimensions.
Wilks’ lambda, Hotelling’s trace, and Roy’s gcr criterion are more powerful when most of the separation of groups occurs in the first dimension, whereas Pillai’s criterion is considered more robust when the separation is distributed over multiple dimensions. Each statistic has its advantages depending on the research design and assumptions. Wilks’ lambda is the most commonly available and used criterion, unless there is a specific reason to use Pillai’s criterion. However, conflicting significance tests may arise, with a significant univariate F for one dependent variable, but a nonsignificant multivariate F when multiple dependent variables are measured. In such cases, researchers may report both results for future research guidance.
When a main effect or interaction is significant in MANOVA, the researcher has usually planned to pursue the finding to discover which DVs are affected. But the problems of assessing DVs in significant multivariate effects are similar to the problems of assigning importance to IVs in multiple regression (Chapter 5). First, there are multiple significance tests so some adjustment is necessary for inflated Type I error. Second, if DVs are uncorrelated, there is no ambiguity in assignment of variance to them, but if DVs are correlated, assignment of overlapping variance to DVs is problematical. Three techniques are discussed to assess the DVS:
In summary, univariate F-tests are acceptable when DVs are uncorrelated, but stepdown analysis is preferable when DVs are correlated, prioritizing significant DVs and reporting adjustments for higher-priority DVs. In cases of nonsignificant univariate F and significant stepdown F, interpretation becomes more complex, and considering loading matrices in discriminant analysis or conducting a discriminant analysis may be helpful. Additionally, tests for contrasts among DVs can be applied to examine differences in the effects of independent variables on individual DVs.
Let's consider an example where we want to compare the academic performance of students in three different subjects: Math, Science, and English. We have test scores for each student in these three subjects as our dependent variables. The independent variable is the type of teaching method used: traditional teaching, online classes, and interactive workshops.
Research Question: Does the type of teaching method have a significant effect on the academic performance of students in Math, Science, and English?
Hypothesis: There will be a significant difference in the academic performance of students across the three subjects based on the type of teaching method.
Method: We collect test scores for a sample of students who were taught using the traditional method, another group of students who attended online classes, and a third group who participated in interactive workshops. We then conduct a MANOVA to analyze the multivariate effect of the teaching method on the test scores in Math, Science, and English.
Results: The MANOVA output will provide us with the main effect of the teaching method and its significance level. If the p-value is below the chosen significance level (e.g., 0.05), we can conclude that there is a statistically significant difference in the academic performance of students across the three subjects based on the teaching method.
Interpretation: Suppose the MANOVA reveals a significant effect. We can then perform follow-up univariate ANOVAs to examine which specific subjects show significant differences in academic performance among the different teaching methods.
Overall, this example demonstrates how MANOVA can be used to analyze the impact of different teaching methods on multiple dependent variables (Math, Science, and English test scores) and provides valuable insights for educators to make informed decisions about teaching strategies
Profile analysis is a multivariate statistical technique used to analyze repeated measures data with multiple dependent variables. In a repeated measures design, the same participants are measured on multiple occasions or under different conditions. Profile analysis allows researchers to examine the patterns of change or differences across the dependent variables over time or across conditions. The main goal of profile analysis is to test the significance of the interaction between the repeated measures factor (time or condition) and the dependent variables. If the interaction is significant, it indicates that the effect of the independent variable (time or condition) varies across the dependent variables.
Profile analysis is particularly suited for research questions that involve the examination of changes or differences in multiple dependent variables across repeated measures or conditions. Some specific research questions that can be answered using profile analysis include:
In all these research questions, the focus is on examining the patterns of change or differences in multiple dependent variables across repeated measures or conditions. Profile analysis allows researchers to gain a comprehensive understanding of the interrelationships and interactions among the dependent variables, providing valuable insights into the underlying processes and relationships being studied.
Profile analysis, like any statistical method, has its limitations and considerations that researchers should be aware of:
Despite these limitations, profile analysis remains a valuable tool for understanding patterns of change or differences in multiple dependent variables across repeated measures or conditions. Researchers should carefully consider the assumptions and limitations while planning and interpreting the results of profile analysis.
The standard formula for profile analysis is:
\[ Y_{ij} = \mu_j + \alpha_i + \beta_{ij} + \epsilon_{ij} \]
where:
Note: This formula assumes a simple repeated measures design with one between-subjects factor and one within-subjects factor. The actual formula may vary depending on the specific design and assumptions of the profile analysis.
Next, tests of parallelism and flatness are used to examine the equality of slopes (parallelism) and the equality of levels (flatness) of multiple treatment groups across different time points or conditions.
These tests are essential in profile analysis to understand how different treatment groups' patterns change over time and whether there are significant differences in the starting levels among the groups. They help researchers evaluate the treatment effects and compare the groups' developmental trajectories or responses across different conditions.
When using profile analysis, there are several important issues to consider to ensure a robust and meaningful interpretation of the results:
By addressing these issues, researchers can conduct a rigorous and informative profile analysis, leading to a better understanding of how dependent variables change over time or conditions and the impact of different treatments or interventions.
Let's consider an example of profile analysis using the Wechsler Intelligence Scale for Children (WISC), a commonly used intelligence test for children. The WISC assesses different cognitive abilities, and we want to examine how these abilities change over time as children age. Imagine we have three age groups of children: Group A (ages 6-8), Group B (ages 9-11), and Group C (ages 12-14). We administer the WISC to each group and measure their performance on four different cognitive subtests: Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed.
The research question is: How do the cognitive abilities measured by the WISC change across these three age groups?
To perform the profile analysis, we first need to organize our data. We create a matrix with the rows representing participants and the columns representing the four cognitive subtests. We then enter the scores of each participant on the corresponding subtests. Next, we conduct a multivariate analysis of variance (MANOVA) to test for significant differences in the cognitive profiles of the three age groups. The MANOVA examines whether there are overall differences in the patterns of scores on the four subtests among the groups. If the MANOVA indicates a significant overall effect, we can conduct follow-up tests to explore specific group differences. We may use univariate ANOVAs or t-tests to compare the means of each cognitive subtest between the age groups. These follow-up tests can help us identify which cognitive abilities show significant differences across the age groups. For example, we might find that Group A (ages 6-8) performs significantly lower on the Working Memory subtest compared to Group C (ages 12-14), but there are no significant differences in the other subtests. This would suggest that Working Memory improves with age, but the other cognitive abilities remain relatively stable.
Profile analysis allows us to examine how the cognitive abilities of the children change as they grow older, providing valuable insights into their cognitive development. By comparing the patterns of scores on multiple dependent variables (cognitive subtests) across different age groups, we can gain a more comprehensive understanding of the changes in cognitive abilities over time.
Discriminant analysis is a statistical method used to classify objects or cases into predefined groups based on a set of predictor variables. The goal of discriminant analysis is to find a linear combination of the predictor variables that maximally separates the groups and minimizes the within-group variability. It is commonly employed in situations where there are two or more groups, and the objective is to determine which combination of variables can best discriminate between them. Discriminant analysis is widely used in various fields, such as psychology, medicine, and marketing, to predict group membership or to identify the key variables that contribute to group differences.
Discriminant analysis can be used to address a variety of research questions related to group classification and prediction. Some common research questions that can be answered using discriminant analysis include:
Overall, discriminant analysis is a powerful tool for understanding group differences, predicting group membership, and selecting important variables in research studies involving multiple groups or categories.
Discriminant analysis has several limitations that researchers should be aware of:
Researchers should consider these limitations and thoroughly evaluate their data and research questions before applying discriminant analysis to their research.
In discriminant analysis, the fundamental equations are used to estimate discriminant functions and classify new observations into predefined groups. The two main equations are:
Dj = aj1X1 + aj2X2 + ... + ajpXp + cj
where:
group j = arg maxj ∗ Dj∗
where:
The coefficients (weights) aj1, aj2 ... ajp and the constant terms cj are estimated from the sample data during the training phase of discriminant analysis using techniques such as Fisher's linear discriminant analysis or quadratic discriminant analysis. These coefficients and constants are then used in the discriminant function to classify new observations into the appropriate groups.
The three types of discriminant analyses— standard (direct), sequential, and statistical (stepwise) are analogous to the three types of multiple regressions. Criteria for choosing among the three strategies are the same as those discussed for multiple regression. These three types differ in their approaches to selecting predictor variables and constructing the discriminant function.
Overall, the choice between these types of discriminant analysis depends on the research question, the number of predictor variables, and the desired level of control over variable selection. Standard discriminant analysis provides a comprehensive view, sequential analysis allows for step-by-step exploration, and statistical analysis offers automated and efficient variable selection.
When using discriminant analysis, several important issues should be considered:
By carefully considering these issues and addressing them appropriately, researchers can conduct effective and meaningful discriminant analyses to classify and predict group membership based on predictor variables.
Interpreting discriminant functions involves understanding how they contribute to group separation and which predictor variables play a significant role in distinguishing between groups. Here's a step-by-step guide to interpreting discriminant functions:
Overall, interpreting discriminant functions involves a combination of statistical analysis, understanding the direction and magnitude of predictor variable coefficients, and visualizing group separation. It is essential to interpret discriminant functions in the context of the research question and the specific dataset under analysis.
Let's consider an example of discriminant analysis to distinguish between three different species of flowers based on their petal length and width. Imagine a botanist collected data from three types of flowers: Iris setosa, Iris versicolor, and Iris virginica. For each flower, they measured the length and width of its petals. The dataset looks like this:
| Flower species | Petal length (cm) | Petal width (cm) |
| Iris setosa | 1.4 | 0.2 |
| Iris setosa | 1.3 | 0.3 |
| ... | ... | ... |
| Iris versicolor | 4.5 | 1.5 |
| Iris versicolor | 4.9 | 1.5 |
| ... | ... | ... |
| Iris virginica | 6.3 | 2.5 |
| Iris virginica | 6.5 | 2.0 |
| ... | ... | ... |
The objective is to classify new flowers into the correct species based on their petal measurements. To achieve this, we'll use discriminant analysis.
By using discriminant analysis, the botanist can predict the species of new flowers based on their petal measurements with a high degree of accuracy. This example demonstrates how discriminant analysis is useful in classifying observations into distinct groups based on multiple predictor variables.
Logistic regression is a statistical method used to model and analyze the relationship between a binary (dichotomous) dependent variable and one or more independent variables (predictors). It is particularly suitable for situations where the dependent variable takes one of two possible outcomes, such as "yes" or "no," "success" or "failure," "1" or "0."
The goal of logistic regression is to estimate the probability that the dependent variable belongs to a particular category based on the values of the independent variables. It uses the logistic function (also known as the sigmoid function) to model the relationship between the predictors and the probability of the event occurring. The output of logistic regression is the predicted probability of the event, which is then used to make a binary decision or classification.
In contrast to linear regression, which predicts continuous outcomes, logistic regression deals with discrete outcomes and is commonly used in various fields, such as medicine, social sciences, marketing, and machine learning, for tasks like predicting whether a customer will purchase a product, diagnosing diseases, or analyzing the likelihood of success in a given situation.
Logistic regression is well-suited for research questions involving binary outcomes or categorical data with two categories. It can answer questions related to the probability or likelihood of an event happening based on predictor variables. Some examples of research questions that can be addressed with logistic regression are:
In general, logistic regression is used when the outcome of interest is binary or can be dichotomized, and the researcher wants to understand how different independent variables or predictors influence the likelihood of that outcome occurring. It is a valuable tool for understanding relationships between categorical variables and making predictions in situations where the response variable is discrete.
Logistic regression, like any statistical method, has its limitations. First, there are theoretical limitations. Logistic regression is a valuable tool for analyzing binary outcomes and understanding the relationship between categorical predictors and the probability of an event occurring. However, caution is needed when interpreting results, as the analysis does not imply causation. While logistic regression provides probabilities between 0 and 1, discriminant analysis might be more powerful and efficient when assumptions are met. Selecting predictors based on a well-justified theoretical model is crucial, and using a large number of predictors without strong justification can lead to misleading results, especially in critical applications like medical policy and practice. Next, there are practical limitations. Some of the key limitations of logistic regression include:
Despite these limitations, logistic regression remains a widely used and valuable tool for analyzing binary outcomes and understanding the relationship between categorical predictors and the probability of an event occurring. Researchers should be aware of these limitations and use appropriate techniques to address them when applying logistic regression to their data.
The difference between multiple regression and logistic regression is that the linear portion of the equation (A + B1X1 + B2X2 + B3X3), the logit, is not the end in itself, but is used to find the odds of being in one of the categories of the DV given a particular combination of scores on the Xs. Similar to multiway frequency analysis, models are evaluated by assessing the (natural log) likelihood for each model. Models are then compared by calculating the difference between their log- likelihoods.
The outcome variable, Ŷ, is the probability of having one outcome or another based on a nonlinear function of the best linear combination of predictors; with two outcomes:
\[ Ŷ_{i} = \frac{e^{u}}{1 + e^{u}} \]
where Ŷi is the estimated probability that the ith case (i = 1, . . . , n) is in one of the categories and u is the usual linear regression equation:
\[ u = A + B_{1}X_{1} + B_{2}X_{2} + ... + B_{K}X_{K} \]
with constant A, coefficients Bj, and predictors, Xj for k predictors (j = 1, . . . , k).
This linear regression equation creates the logit or log of the odds:
\[ ln(\frac{Ŷ}{1 - Ŷ}) = A + \sum B_{j}X_{ij} \]
That is, the linear regression equation is the natural (loge ) of the probability of being in one group divided by the probability of being in the other group. The procedure for estimating coefficients is maximum likelihood, and the goal is to find the best linear combination of predictors to maximize the likelihood of obtaining the observed outcome frequencies.
As in multiple regression and discriminant analysis, there are three major types of logistic regression: direct (standard), sequential, and statistical. Logistic regression programs tend to have more options for controlling equation- building than discriminant programs, but fewer options than multiple regression.
In this approach, all predictor variables are entered into the model simultaneously to estimate the coefficients. It provides a straightforward analysis of the relationship between the predictor variables and the probability of the binary outcome. The direct logistic regression is commonly used when the researcher has a clear theoretical understanding of the predictors' relationships with the outcome and wants to assess their direct effects.
In this approach, predictor variables are entered into the model one at a time in a pre-determined order. Each predictor's contribution to the model is assessed sequentially, allowing for an examination of the incremental effect of each variable on the prediction of the binary outcome. This method is useful when researchers want to understand how individual predictors contribute to the prediction of the outcome and when the order of entry is guided by theoretical or practical considerations.
This approach is similar to sequential logistic regression but differs in how predictors are selected and retained in the model. In statistical logistic regression, variables are entered into the model one at a time, and at each step, the variable that contributes the most to the prediction is added or removed based on predefined criteria (e.g., significance level). This method is useful when researchers have a large number of predictor variables and want to identify the most relevant variables for predicting the outcome automatically.
Each of these types of logistic regression has its strengths and limitations, and the choice of approach depends on the research objectives, the available data, and the researcher's understanding of the relationships between predictors and the binary outcome.
Probit analysis is a statistical method used to model binary outcomes (dichotomous data) and estimate the probability of an event occurring. It is commonly used when the response variable is binary, meaning it can take one of two possible values, such as "success" or "failure," "yes" or "no," or "alive" or "dead."
The relationship between probit analysis and logistic regression lies in their similarities as approaches to modeling binary outcomes. Both methods aim to estimate the probability of an event occurring based on predictor variables. However, they use different link functions to relate the linear predictor to the probability.
In logistic regression, the logit link function (logarithm of the odds) is used to model the relationship between the linear predictor and the probability of success. On the other hand, in probit analysis, the probit link function (the cumulative standard normal distribution) is used to model the relationship between the linear predictor and the probability of success.
Despite this difference in the link functions, the two methods often provide very similar results in practice, especially when the sample size is large. Both probit analysis and logistic regression are widely used in various fields, such as biology, economics, psychology, and epidemiology, to model binary outcomes and estimate the effects of predictor variables on the probability of success or the occurrence of an event.
When using logistic regression, there are several important issues to consider to ensure the validity and reliability of the results:
By addressing these issues, researchers can ensure the reliability and validity of the logistic regression analysis and draw meaningful conclusions from their results.
Let's consider a scenario where a researcher wants to investigate the factors that influence the likelihood of households experiencing financial strain due to high energy prices. The researcher collects data from a sample of households and records the following variables for each household:
Dependent Variable:
Independent Variables:
The researcher then applies logistic regression to model the relationship between the dependent variable (Financial Strain) and the independent variables (Income Level, Energy Consumption, and Energy Price Index).
The logistic regression model provides the researcher with the following results:
The coefficient estimates for the independent variables: The coefficient estimates indicate the direction and strength of the relationship between each independent variable and the log-odds of experiencing financial strain.
Odds Ratios: The odds ratios for each independent variable represent the change in the odds of experiencing financial strain for a one-unit change in the corresponding independent variable. For example, an odds ratio of 1.5 for Energy Price Index means that a one-unit increase in the Energy Price Index is associated with a 1.5 times higher odds of experiencing financial strain.
p-values: The p-values associated with the coefficient estimates determine the statistical significance of each independent variable's contribution to the model.
The researcher can use these results to make conclusions about the significant predictors of financial strain due to high energy prices. For instance, if the p-value for the Energy Price Index is less than the chosen significance level (e.g., 0.05), it indicates that the Energy Price Index is a significant predictor of financial strain, and its odds ratio provides insight into its effect on the odds of experiencing financial strain. The same interpretation applies to other significant predictors such as Income Level and Energy Consumption.
Overall, logistic regression helps identify the factors that are associated with an increased likelihood of experiencing financial strain due to high energy prices, providing valuable insights for policymakers and households to address energy affordability issues.
Survival/failure analysis is the family of statistical techniques dealing with the time it takes for something to happen. The term survival analysis is based on medical applications in which time to relapse or death is studied between groups who have received different medical treatments.
One interesting feature of the analysis is that survival time (the DV) often is unknown for a number of cases at the conclusion of the study. Some of the cases are still in the study but have not yet failed: some employees have not yet left, some components are still functioning, some patients are still apparently well, or some patients are still living. For other cases, the outcome is simply unknown because they have withdrawn from the study or are for some reason lost to follow-up. Cases whose DV values, meaning survival times, are unknown for whatever reason are referred to as censored.
Within the family of survival-analysis techniques, different procedures are used depending on the nature of the data and the kinds of questions that are of greatest interest. Life tables are used to describe the survival or failure times for cases. They are often accompanied by a graphical representation of the survival rate as a function of time, called a survivor function. Survivor functions are frequently plotted side-by-side for two or more groups and statistical tests are used to test the differences between groups in survival time.
Another set of procedures is used when the goal is to determine if survival time is influenced by some other variables. These are regression procedures in which survival time is predicted from a set of variables, where the set may include one or more treatment variables.
A potential source of confusion is that all of the predictors are called covariates. In most previous analyses, the word 'covariates' was used for variables that enter a sequential equation early and the analysis adjusts for their relationship with the DV before variables of greater interest enter. There is a sequential survival analysis in which the term covariates is used in the traditional way, but covariates is also used for the treatment IVs even when the analysis is not sequential.
The main aim of one type of survival analysis is to describe the proportion of cases surviving at various times, within a single group or separately for different groups. The analysis extends to statistical tests of group differences. The primary goal of the other type of survival analysis is to assess the relationship between survival time and a set of covariates (predictors), with treatment considered one of the covariates, to determine whether treatment differences are present after statistically controlling for the other covariates.
Survival/failure analysis can be used to calculate surviving rates at various point in time, or to study group differences in survival. When looking at data including survival time with covariates, it can also be used to assess treatment effect, the importance of covariates, parameter estimates, contingencies among covariates, and effect size and power.
There are some limitations to the use of survival analysis. We distinguish theoretical issues and practical issues.
One problem with survival analysis is the nature of the outcome variable, time itself. Events must occur before survival or failure time can be analyzed: Components must fail, employees must leave, patients must succumb to the illness. However, the purpose of treatment often is to delay this occurrence or prevent it altogether. The more successful the treatment, then, the less able the researcher is to collect data in a timely fashion.
Survival analysis is subject to the usual cautions about causal inference. For example, a difference in survival rates among groups cannot be attributed to the treatment unless assignment to levels of treatment and implementation of those levels, with control, are properly experimental.
In the descriptive use of survival analysis, assumptions regarding the distributions of covariates and survival times are not required. However, in the regression forms of survival analysis in which covariates are assessed, multivariate normality, linearity, and homoscedasticity among covariates, although not required, often enhance the power of the analysis to form a useful linear equation of predictors.
Life tables are built around time intervals. The survivor function, P, is the cumulative proportion of cases surviving to the beginning of the i + 1st interval, estimated as:
Pi=1 = piPi
where: pi = 1- qi
and: qi = di / ri
and where: di = number responding (dropping out) in the interval and ri = ni - 1/2*ci
where ni = number entering the interval and ci = number censored in the interval (lost to follow up for reasons other than dropping out).
In words, the proportion of cases surviving to the (i + 1)st interval is the proportion who survived to the start of ith interval times the probability of surviving to the end of the ith interval (by not dropping out or being censored during that interval).
Various statistics and standard errors for them are developed to facilitate inferential tests of survival functions.
The standard error of a cumulative proportion of cases surviving an interval is approximately:

The hazard, also sometimes called the failure rate, is the rate of not surviving to the midpoint of an interval, given survival to the start of the interval.

where hi = the width of the ith interval.
The probability density is the probability of not surviving to the midpoint of an interval, given survival to the start of the interval:

Note the distinction between the hazard function and the probability density function. The hazard function is the instantaneous rate of dropping out at a specific time among cases who survived to at least the beginning of that time interval. The probability density function is the probability of a given case dropping out at a specified time point.
Life table plots are simply the cumulative proportion surviving (Pi), plotted as a function of each time interval.
Group differences in survival are tested through x2 with degrees of freedom equal to the number of groups minus 1. Of the several tests that are available, the one demonstrated here is labeled Log-Rank in SAS LIFETEST and IBM SPSS KM. When there are only two groups, the overall test is equated as follows:

x2 equals the squared value of the observed minus expected frequencies of number of survivors summed over all intervals for one of the groups (v2j) under the null hypothesis of no group differences, divided by Vj, the variance of the group. The degree of freedom for this test is (number of groups -1). When there are only two groups, the value in the numerator is the same for both groups, but opposite in sign. The value in the denominator is also the same for both groups. Therefore, computations for either group produce the same x2.
The value, vi, of observed minus expected frequencies is calculated separately for each interval for one of the groups.

The difference between observed and expected frequencies for a group, v0, is the summed differences over the intervals between the number of survivors in each interval, d0, minus the ratio of the number of cases at risk in the interval (n0i) times the total number of survivors in that interval summed over all groups (dTi) divided by the total number of cases at risk in that interval summed over all groups (nTi).
The variance, V, for a group is calculated with the following formula:

The variance for the control group, V0, is the sum over all intervals of the difference between the total number of survivors in an interval (nTi) times the number of survivors in the control group in the interval (n0i) minus the squared number of survivors in the control group in the interval; this difference is multiplied by the product of the total number of survivors in the interval (dTi) times sTi(= nTi - dTi); all of this is divided by the squared total number of survivors in the interval (nTi) times the total number of survivors in the interval minus 1.
In jargon, the total number of cases that have survived to an interval (nTi) is called the risk set.
There are two major types of survival analyses:
Both the actuarial method as well as the product-limit method can be used for calculating life tables. The two methods produce identical results when there is no censoring, and intervals contain no more than one time unit.
The output is organized by time at which events occur rather than by time interval. For example, there are two lines of output for the two control cancer patients who dropped out during the second month and for the two cancer patients who dropped out in the tenth month. Both mean and median survival time are given, along with their standard errors and 95% confidence intervals. The survival-function chart differs slightly from that of the actuarial method by including information about cases that are censored. Group differences are tested through the Log Rank test, among others that can be requested, rather than the Wilcoxon test produced by IBM SPSS SURVIVAL.
Prediction of survival or failure time from covariates is similar to logistic regression but with provision for censored data. This method also differs in analyzing the time between events rather than predicting the occurrence of events. Cox proportional hazards (Cox regression) is the most popular method. Accelerated failure-time models are also available for the more sophisticated user.
As in other forms of regression, analysis of survival can be direct, sequential, or statistical. A treatment IV, if present, is analyzed the same as any other discrete covariate. When there are only two levels of treatment, the treated group is usually coded 1 and the control group 0. If there are more than two levels of treatment, dummy variable coding is used to represent group membership. Successful prediction on the basis of this IV indicates significant treatment effects.
The three major analytic strategies in survival analysis with covariates are direct (standard), sequential (hierarchical), and statistical (stepwise or setwise). Differences among the strategies involve what happens to overlapping variability due to correlated covariates, including treatment groups, and who determines the order of entry of covariates into the equation.
The Cox proportional-hazards method models event rates as a log-linear function of predictors, called covariates. Regression coefficients give the relative effect of each covariate on the survivor function. Cox modeling is available through IBM SPSS COXREG and SAS PHREG.
Accelerated failure-time models replace the general hazard function of the Cox model with a specific distribution. However, greater user sophistication is required to choose the distribution. Accelerated failure- time models are handled by SAS LIFEREG. IBM SPSS has no program for accelerated failure- time modeling. Choice of distributions in accelerated failure- time models has implications for hazard functions, so that modeling based on different distributions may lead to different interpretations.
Choice of models is made on the basis of logic, graphical fit, or, in the case of nested models, goodness-of-fit tests. Allison has provided guidance for producing graphs of appropriate transformations to Kaplan– Meier estimates. A resulting linear plot indicates that the distribution providing the transformation is the appropriate one. Allison also illustrates procedures for applying goodness-of-fit statistics to statistically evaluate competing models. Accelerated failure- time analyses produce x2 log- likelihood values, in which negative values closer to zero indicate better fits of data to models. Thus, twice the difference between nested models provides a likelihood- ratio χ2 statistic. Of the nested models, gamma is the most general, followed by Weibull, then exponential, and finally lognormal. That is, log- normal is nested within exponential, which in turn is nested within Weibull, et cetera.
The most straightforward way to analyze survival data with covariates is Cox regression. It is more robust than accelerated failure-time methods and requires no choice among the distributions on the part of the researcher. However, the Cox model does have the assumption of proportionality of hazards over time.
Issues in survival analysis include testing the assumption of proportionality of hazards, dealing with censored data, assessing effect size of models and individual covariates, choosing among the variety of statistical tests for differences among treatment groups and contributions of covariates, and interpreting hazard ratios.
When Cox regression is used to analyze differences between levels of a discrete covariate such as treatment, it is assumed that the shapes of the survival functions are the same for all groups over time. That is, the time until failures begin to appear may be longer for one group than another, but once failures start, they proceed at the same rate for all groups. When the assumption is met, the lines for the survival functions for different groups are roughly parallel. Although inspection of the plots is helpful, a formal test of the assumption is also required. The assumption is similar to homogeneity of regression in ANCOVA, which requires that the relationship between the DV and the covariate(s) is the same for all levels of treatment.
The proportionality of hazards assumption is that the relationship between survival rate and time is the same for all levels of treatment (or any other covariate). In survival analysis, violation of proportionality signals interaction between time and levels of treatment (or any other covariate). To test the assumption, a time variable is constructed, and its interaction with levels of treatment (and other covariates) is tested.
Censored cases are those for whom the time of the event being studied is unknown or only vaguely known.
A data set can contain a mix of cases with several forms of censoring.
Cox and Snell have provided a measure of effect size for logistic regression that is demonstrated for survival analysis by Allison. It is based on G2 , a likelihood-ratio chi-square statistic that can be calculated from SAS PHREG and LIFEREG and IBM SPSS COXREG. Models are fit both with and without covariates, and a difference G2 is found by:
G2 = [(-2 log-likelihood for smaller model) - (-2 log-likelihood for larger model)]
Then, R2 is found by:

When applied to experiments, the R2 of greatest interest is the association between survival and treatment, after adjustment for other covariates. Therefore, the smaller model is the one that includes covariates but not treatment, and the larger model is the one that includes covariates and treatment. Allison points out that this R2 is not the proportion of variance in survival that is explained by the covariates, but merely represents relative association between survival and the covariates tested.
Power in survival analysis is, as usual, enhanced by larger sample sizes and covariates with stronger effects. Amount of censoring and patterns of entry of cases into the study also affect power. as does the relative size of treatment groups. Unequal sample sizes reduce power while equal sample sizes increase it.
Numerous statistical tests are available for evaluating group differences due to treatment effects from an actuarial life table or product-limit analysis.
Statistics for predicting survival from covariates require calculating regression coefficients for each covariate where one or more of the covariates may represent treatment. The regression coefficients give the relative effect of each covariate on the survival function, but the size depends on the scale of the covariate. These coefficients may be used to develop a regression equation for risk as a DV.
IBM SPSS and SAS have two or more programs that do different types of analysis. SAS has one program for survival functions and two for regression- type problems: one for proportional-hazards models and the other for various nonproportional-hazards models. IBM SPSS has three programs as well: one for proportional-hazards models and two for survival functions (one for actuarial and one for product-limit methods). SYSTAT has a single program for survival analysis.
The goal of canonical correlation is to analyze the relationships between two sets of variables. It may be useful to think of one set of variables as IVs and the other set as DVs, or it may not. In any event, canonical correlation provides a statistical analysis for research in which each subject is measured on two sets of variables and the researcher wants to know if and how the two sets relate to each other.
The easiest way to understand canonical correlation is to think of multiple regression. In regression, there are several variables on one side of the equation and a single variable on the other side. The several variables are combined into a predicted value to produce, across all subjects, the highest correlation between the predicted value and the single variable. The combination of variables can be thought of as a dimension among the many variables that predicts the single variable. In canonical correlation, the same thing happens except that there are several variables on both sides of the equation. Sets of variables on each side are combined to produce, for each side, a predicted value that has the highest correlation with the predicted value on the other side. The combination of variables on each side can be thought of as a dimension that relates the variables on one side to the variables on the other. There is a complication, however. In multiple regression, there is only one combination of variables because there is only a single variable to predict on the other side of the equation. In canonical correlation, there are several variables on both sides and there may be several ways to recombine the variables on both sides to relate them to each other.
Canonical correlation analysis is useful when the underlying dimensions representing the combinations of variables are unknown. Structural equations modelling may provide a more effective analysis when theory or prior research suggests the underlying dimensions. Thus, canonical analysis may be seen as an exploratory technique, with SEM as the parallel confirmatory technique. A good deal of the difficulty with canonical correlation is due to jargon. First, there are variables, then there are canonical variates, and, finally, there are pairs of canonical variates. Variables refer to the variables measured in research. Canonical variates are linear combinations of variables, one combination on the IV side and another combination on the DV side. These two combinations form a pair of canonical variates. However, there may be more than one significant pair of canonical variates.
Although a large number of research questions are answered by canonical analysis in one of its specialized forms, relatively few intricate research questions are readily answered through direct application of computer programs currently available for canonical correlation. In part, this has to do with the programs themselves, and in part, it has to do with the kinds of questions researchers consider appropriate in a canonical correlation. In its present stage of development, canonical correlation is best considered a descriptive technique or a screening procedure rather than a hypothesis– testing procedure.
Canonical correlation has several important theoretical limitations that help explain its scarcity in the literature. Perhaps the most critical limitation is interpretability; procedures that maximize correlation do not necessarily maximize interpretation of pairs of canonical variates. Therefore, canonical solutions are often mathematically elegant but uninterpretable. And, although it is common practice in factor analysis and principal components analysis to rotate a solution to improve interpretation, rotation of canonical variates is not common practice or even available in some computer programs.
The algorithm used for canonical correlation maximizes the linear relationship between two sets of variables. If the relationship is nonlinear, the analysis misses some or most of it. If a nonlinear relationship between dimensions in a pair is suspected, use of canonical correlation may be inappropriate unless variables are transformed or combined to capture the nonlinear component. It is especially important in canonical analysis to emphasize that the use of terms IV and DV does not imply a causal relationship.
An important concern is the sensitivity of the solution in one set of variables to the variables included in the other set. In canonical analysis, the solution depends both on correlations among variables in each set and on correlations among variables between sets. Changing the variables in one set may markedly alter the composition of canonical variates in the other set. To some extent, this is expected given the goals of analysis; yet, the sensitivity of the procedure to apparently minor changes is a cause for concern.
We distinguish five important practical issues of canonical analysis:
A data set that is appropriately analyzed through canonical correlation has several subjects, each measured on four or more variables. The variables form two sets with at least two variables in the smaller set.
The first step in a canonical analysis is generation of a correlation matrix. In this case, however, the correlation matrix is subdivided into four parts: the correlations between the DVs (Ryy), the correlations between the IVs (Rxx), and the two matrices of correlations between DVs and IVs (Rxy and Ryx). The equations are all variants on the following equation:
R = R-1yyRyxR-1xxRxy
The canonical correlation matrix is a product of four correlation matrices, between DVs (inverted), between IVs (inverted), and between DVs and IVs.
Although computing eigenvalues and eigenvectors is best left to the computer, the relationship between a canonical correlation and an eigenvalue is simple, namely:
λi = r2ci
Each eigenvalue, λi, is equal to the squared canonical correlation, r2ci, for the pair of canonical variates. Once the eigenvalue is calculated for each pair of canonical variates, canonical correlation is found by taking the square root of the eigenvalue. Canonical correlation, rci, is interpreted as an ordinary Pearson product- moment correlation coefficient. When rci is squared, it represents, as usual, overlapping variance between two variables, or, in this case, variates. Because r2ci = λi, the eigenvalues themselves represent overlapping variance between pairs of canonical variates.
Significance tests are available to test whether one or a set of rcs differs from zero:

The significance of one or more canonical correlations is evaluated as a chi- square variable, where N is the number of cases, kx is the number of variables in the IV set, ky is the number in the DV set, and the natural logarithm of lambda, see the following formula: This chi square has (kx)(ky) df.

Lambda, Λ, is the product of differences between eigenvalues and unity, generated across m canonical correlations.
Two sets of canonical coefficients (analogous to regression coefficients) are required for each canonical correlation, one set to combine the DVs and the other to combine the IVs. The canonical coefficients for the DVs are found as follows:

Canonical coefficients for the DVs are a product of (the transpose of the inverse of the square root of) the matrix of correlations between DVs and the normalized matrix of eigenvectors, Bn y, for the DVs. Once the canonical coefficients are computed, coefficients for the IVs are found using the following equation:

Coefficients for the IVs are a product of (the inverse of the square root of) the matrix of correlations between the IVs, the matrix of correlations between the IVs and DVs, and the matrix formed by the coefficients for the DVs - each divided by their corresponding canonical correlations.
The two matrices of canonical coefficients are used to estimate scores on canonical variates: X = ZxBx and Y = ZyBy. Scores on canonical variates are estimated as the product of the standardized scores on the original variates, Zx and Zy, and the canonical coefficients used to weight them, Bx and By.
Matrices of correlations between the variables and the canonical coefficients, called loading matrices, are used to interpret the canonical variates.
Ax = RxxBx and Ay = RyyBy
Correlations between variables and canonical variates are found by multiplying the matrix of correlations between variables by the matrix of canonical coefficients.
How much variance does each of the canonical variates extract from the variables on its own side of the equation? The proportion of variance extracted from the IVs by the canonical variates of the IVs is as follows:

and

The proportion of variance extracted from a set of variables by a canonical variate of the set is the sum of the squared correlations divided by the number of variables in the set.
Often, however, one is interested in knowing how much variance the canonical variates from the IVs extract from the DVs, and vice versa. In canonical analysis, this variance is called redundancy.
rd = (pv)(r2c)
The redundancy in a canonical variate is the proportion of variance it extracts from its own set of variables times the squared canonical correlation for the pair.
As in most statistical procedures, establishing significance is usually the first step in evaluating a solution. Conventional statistical procedures apply to significance tests for the number of canonical variate pairs. The number of statistically significant pairs of canonical variates is often larger than the number of interpretable pairs if N is at all sizable.
The only potential source of confusion is the meaning of the chain of significance tests. The first test is for all pairs taken together and is a test of independence between the two sets of variables. The second test is for all pairs of variates with the first and most important pair of canonical variates removed; the third is done with the first two pairs removed, and so forth.
Once significance is established, amount of variance accounted for is of critical importance. Because there are two sets of variables, several assessments of variance are relevant. First, there is variance overlap between variates in a pair; second is a variance overlap between a variate and its own set of variables; and third is a variance overlap between a variate and the other set of variables.
Canonical correlation creates linear combinations of variables, canonical variates, that represent mathematically viable combinations of variables. However, although mathematically viable, they are not necessarily interpretable. A major task for the researcher is to discern, if possible, the meaning of pairs of canonical variates.
Interpretation of significant pairs of canonical variates is based on the loading matrices, Ax and Ay. Each pair of canonical variates is interpreted as a pair, with a variate from one set of variables interpreted vis-à-vis the variate from the other set. A variate is interpreted by considering the pattern of variables highly correlated (loaded) with it. Because the loading matrices contain correlations, and because squared correlations measure overlapping variance, variables with correlations of .30 and above are usually interpreted as part of the variate, and variables with loadings below .30 are not. Deciding on a cutoff for interpreting loadings is, however, somewhat a matter of taste.
One program is available in the SAS package for canonical analyses. IBM SPSS has two programs that may be used for canonical analysis. If available, the program of choice is SAS CANCORR and, with limitations, the IBM SPSS CANCORR macro.
SAS CANCORR is a flexible program with abundant features and ease of interpretation. Along with the basics, you can specify easily interpretable labels for canonical variates and the program accepts several types of input matrices. Multivariate output is quite detailed, with several test criteria and voluminous redundancy analyses. Univariate output is minimal, however, and if plots are desired, case statistics such as canonical scores are written to a file to be analyzed by the SAS PLOT procedure.
If requested, the program does separate multiple regressions with each variable predicted from the other set. You can also do separate canonical correlation analyses for different groups.
IBM SPSS has two programs for canonical analysis, both available only through syntax: IBM SPSS MANOVA and a CANCORR macro. A complete canonical analysis is available through SPSS MANOVA, which provides loadings, proportions of variance, redundancy, and much more. But problems arise with reading the results, because MANOVA is not designed specifically for canonical analysis and some of the labels are confusing.
Canonical analysis is requested through MANOVA by calling one set of variables DVs and the other set covariates; no IVs are listed. Although IBM SPSS MANOVA provides a rather complete canonical analysis, it does not calculate canonical variate scores, nor does it offer multivariate plots.
Currently, canonical analysis is most readily done through SETCOR SYSTAT System. The program provides all of the basics of canonical correlation and several others. There is a test of overall association between the two sets of variables, as well as tests of prediction of each DV from the set of IVs. The program also provides analyses in which one set is partialed from the other set––useful for statistical adjustment of irrelevant sources of variance (as per covariates in ANCOVA) as well as representation of curvilinear relationships and interactions.
Canonical analysis may also be done through the multivariate general linear model GLM program in SYSTAT. But to get all the output, the analysis must be done twice, once with the first set of variables defined as the DVs, and a second time with the other set of variables defined as DVs. The advantages over SETCOR are that canonical variate scores may be saved in a data file, and that standardized canonical coefficients are provided.
Principal components analysis (PCA) and factor analysis (FA) are statistical techniques applied to a single set of variables when the researcher is interested in discovering which variables in the set form coherent subsets that are relatively independent of one another. Variables that are correlated with one another but largely independent of other subsets of variables are combined into factors. Factors are thought to reflect underlying processes that have created the correlations among variables. A major use of PCA and FA in psychology is in development of objective tests for measurement of personality and intelligence and the like.
The specific goals of PCA or FA are to summarize patterns of correlations among observed variables, to reduce a large number of observed variables to a smaller number of factors, to provide an operational definition (a regression equation) for an underlying process by using observed variables, or to test a theory about the nature of underlying processes. Some or all of these goals may be the focus of a particular research project. PCA and FA have considerable utility in reducing numerous variables down to a few factors. Mathematically, PCA and FA produce several linear combinations of observed variables, where each linear combination is a factor. The factors summarize the patterns of correlations in the observed correlation matrix and can be used, with varying degrees of success, to reproduce the observed correlation matrix. But since the number of factors is usually far fewer than the number of observed variables, there is considerable parsimony in using the factor analysis. Further, when scores on factors are estimated for each subject, they are often more reliable than scores on individual observed variables.
Steps in PCA or FA include selecting and measuring a set of variables, preparing the correlation matrix (to perform either PCA or FA), extracting a set of factors from the correlation matrix, determining the number of factors, (probably) rotating the factors to increase interpretability, and, finally, interpreting the results. Although there are relevant statistical considerations to most of these steps, an important test of the analysis is its interpretability.
Interpretation and naming of factors depend on the meaning of the particular combination of observed variables that correlate highly with each factor. A factor is more easily interpreted when several observed variables correlate highly with it and those variables do not correlate with other factors. Once interpretability is adequate, the last, and very large, step is to verify the factor structure by establishing the construct validity of the factors. The researcher seeks to demonstrate that scores on the latent variables (factors) covary with scores on other variables, or that scores on latent variables change with experimental conditions as predicted by theory.
There are some problems when conducting PCA or FA.
There are two major types of FA: exploratory and confirmatory.
There are some important terms used in PCA and FA.
FA produces factors, while PCA produces components. However, the processes are similar except in preparation of the observed correlation matrix for extraction and in the underlying theory. Mathematically, the difference between PCA and FA is in the variance that is analyzed. In PCA, all the variances in the observed variables are analyzed. In FA, only shared variance is analyzed; attempts are made to estimate and eliminate variance due to error and variance that is unique to each variable. The term factor is used here to refer to both components and factors unless the distinction is critical, in which case the appropriate term is used.
Theoretically, the difference between FA and PCA lies in the reason that variables are associated with a factor or component. Factors are thought to 'cause' variables - the underlying construct (the factor) is what produces scores on the variables. Thus, exploratory FA is associated with theory development and confirmatory FA is associated with theory testing.
The goal of research using PCA or FA is to reduce a large number of variables to a smaller number of factors, to concisely describe (and perhaps understand) the relationships among observed variables, or to test theory about underlying processes.
There are some theoretical considerations when using PCA or FA:
Because FA and PCA are exquisitely sensitive to the sizes of correlations, it is critical that honest correlations be employed. Sensitivity to outlying cases, problems created by missing data, and degradation of correlations between poorly distributed variables all plague FA and PCA.
PCA and FA have several assumptions: normality, linearity, absence of outliers, and absence of multicollinearity and singularity.
A matrix that is factorable should include several sizable correlations. The expected size depends, to some extent, on N (larger sample sizes tend to produce smaller correlations), but if no correlation exceeds .30, use of FA is questionable because there is probably nothing to factor analyze. Inspect R for correlations in excess of .30, and, if none is found, reconsider use of FA.
High bivariate correlations, however, are not ironclad proof that the correlation matrix contains factors. It is possible that the correlations are between only two variables and do not reflect the underlying processes that are simultaneously affecting several variables. For this reason, it is helpful to examine matrices of partial correlations where pairwise correlations are adjusted for effects of all other variables. If there are factors present, then high bivariate correlations become very low partial correlations. IBM SPSS and SAS produce partial correlation matrices
After FA, in both exploratory and confirmatory FA, variables that are unrelated to others in the set are identified. These variables are usually not correlated with the first few factors although they often correlate with factors extracted later. These factors are usually unreliable, both because they account for very little variance and because factors that are defined by just one or two variables are not stable. Therefore, one never knows whether these factors are 'real'.
If the variance accounted for by a factor defined by only one or two variables is high enough, the factor is interpreted with great caution or is ignored, as pragmatic considerations dictate. In confirmatory FA done through FA rather than SEM programs, the factor represents either a promising lead for future work or (probably) error variance, but its interpretation awaits clarification by more research. A variable with a low squared multiple correlation with all other variables and low correlations with all important factors is an outlier among the variables. The variable is usually ignored in the current FA and either deleted or given friends in future research.
The main variables that are used in PCA and FA are matrices of correlations (between variables, between factors, and between variables and factors), matrices of standard scores (on variables and on factors), matrices of regression weights (for producing scores on factors from scores on variables), and the pattern matrix of unique relationships between factors and variables after oblique rotation.
| Label | Name | Rotation | Size | Description |
|---|---|---|---|---|
| R | Correlation matrix | Both orthogonal and oblique | p*p | Matrix of correlations between variables |
| Z | Variable matrix | Both orthogonal and oblique | N*p | Matrix of standardised observed variable scores |
| F | Factor-score matrix | Both orthogonal and oblique | N*m | Matrix of standardised scores on factors or components |
| A | Factor loading | Orthogonal | p*m | Matrix of regression-like weights used to estimate the unique contribution of each factor to the variance in a variable. If orthogonal, also correlations between variables and factors. |
| B | Factor-score coefficients matrix | Both orthogonal and oblique | p*m | Matrix of regression-like weights used to generate factor scores from variables. |
| C | Structure matrix | Oblique | p*m | Matrix of correlations between variables and (correlated) factors. |
| Φ | Factor correlation matrix | Oblique | m*m | Matrix of correlations among factors. |
| L | Eigenvalue matrix | Both orthogonal and oblique | m*m | Diagonal matrix of eigenvalues, one per factor. |
| V | Eigenvector matrix | Both orthogonal and oblique | p*m | Matrix of eigenvectors, one vector per eigenvalue. |
An important theorem from matrix algebra indicates that, under certain conditions, matrices can be diagonalized. Correlation and covariance matrices are among those that often can be diagonalized. When a matrix is diagonalized, it is transformed into a matrix with numbers in the positive diagonal and zeros everywhere else. In this application, the numbers in the positive diagonal represent variance from the correlation matrix that has been repackaged as follows:
L = V'RV
Diagonalization of R is accomplished by post- and pre- multiplying it by the matrix V and its transpose. The columns in V are called eigenvectors, and the values in the main diagonal of L are called eigenvalues. The first eigenvector corresponds to the first eigenvalue, and so forth.
Because there are four variables in the example, there are four eigenvalues with their corresponding eigenvectors. However, because the goal of FA is to summarize a pattern of correlations with as few factors as possible, and because each eigenvalue corresponds to a different potential factor, usually only factors with large eigenvalues are retained. In a good FA, these few factors almost duplicate the correlation matrix.
The matrix of eigenvectors pre- multiplied by its transpose produces the identity matrix with ones in the positive diagonal and zeros elsewhere. Therefore, pre- and post- multiplying the correlation matrix by eigenvectors does not change it so much as repackage it:
V'V = I
The important point is that because correlation matrices often meet requirements for diagonalizability, it is possible to use on them the matrix algebra of eigenvectors and eigenvalues with FA as the result. When a matrix is diagonalized, the information contained in it is repackaged. FA, the variance in the correlation matrix is condensed into eigenvalues. The factor with the largest eigenvalue has the most variance and so on, down to factors with small or negative eigenvalues that are usually omitted from solutions. Calculations for eigenvectors and eigenvalues are extremely laborious and not particularly enlightening.
R = VLV'
The correlation matrix can be considered a product of three matrices— the matrices of eigenvalues and corresponding eigenvectors.
After reorganization, the square root is taken of the matrix of eigenvalues:

or

If V√L is called A and √LV is A then R=AA'. The correlation matrix can also be considered a product of two matrices - each a combination of eigenvectors and the square root of eigenvalues. The factor loading matrix is a matrix of correlations between factors and variables.
Rotation is ordinarily used after extraction to maximize high correlations between factors and variables and minimize low ones. Numerous methods of rotation are available but the most commonly used, and the one illustrated here, is varimax. Varimax is a variance-maximizing procedure. The goal of varimax rotation is to maximize the variance of factor loadings by making high loadings higher and low ones lower for each factor. This goal is accomplished by means of a transformation matrix Λ.
AunrotatedΛ = Arotated
The unrotated factor loading matrix is multiplied by the transformation matrix to produce the rotated loading matrix. Compare the rotated and unrotated loading matrices. Notice that in the rotated matrix the low correlations are lower and the high ones are higher than in the unrotated loading matrix. Emphasizing differences in loadings facilitates interpretation of a factor by making unambiguous the variables that correlate with it. The numbers in the transformation matrix have a spatial interpretation.

The transformation matrix is a matrix of sines and cosines of an angle ψ.
Once the rotated loading matrix is available, other relationships are found. The communality for a variable is the variance accounted for by the factors. It is the squared multiple correlation of the variable as predicted from the factors. Communality is the sum of squared loadings (SSL) for a variable across factors. The proportion of variance in the set of variables accounted for by a factor is the SSL for the factor divided by the number of variables (if rotation is orthogonal). The proportion of variance in the solution accounted for by a factor - the proportion of covariance - is the SSL for the factor divided by the sum of communalities (or, equivalently, the sum of the SSLs).
Notice that the reproduced correlation matrix differs slightly from the original correlation matrix. The difference between the original and reproduced correlation matrices is the residual correlation matrix:

The residual correlation matrix is the difference between the observed correlation matrix and the reproduced correlation matrix.
Scores on factors can be predicted for each case once the loading matrix is available. Regressionlike coefficients are computed for weighting variable scores to produce factor scores.
B = R-1A
Factor score coefficients for estimating factor scores from variable scores are a product of the inverse of the correlation matrix and the factor loading matrix. To estimate a subject’s score for the first factor, all of the subject’s scores on variables are standardized. In matrix form, F = ZB. Factor scores are a product of standardized scores on variables and factor score coefficients.
Predicting scores on variables from scores on factors is also possible. The equation for doing so is:
Z = FA'
Predicted standardized scores on variables are a product of scores on factors weighted by factor loadings. A score on an observed variable is conceptualized as a properly weighted and summed combination of the scores on factors that underlie it. The researcher believes that each subject has the same latent factor structure, but different scores on the factors themselves. A particular subject’s score on an observed variable is produced as a weighted combination of that subject’s scores on the underlying factors.
All the relationships mentioned thus far are for orthogonal rotation. Most of the complexities of orthogonal rotation remain and several others are added when oblique (correlated) rotation is used. In oblique rotation, the loading matrix becomes the pattern matrix. Values in the pattern matrix, when squared, represent the unique contribution of each factor to the variance of each variable but do not include segments of variance that come from overlap between correlated factors. Once the factor scores are determined, correlations among factors can be obtained. Among the equations used for this purpose is:

One way to compute correlations among factors is from cross- products of standardized factor scores divided by the number of cases minus one. The factor correlation matrix is a standard part of computer output following oblique rotation.
If oblique rotation is used, the structure matrix, C, is the correlations between variables and factors. These correlations assess the unique relationship between the variable and the factor (in the pattern matrix) plus the relationship between the variable and the overlapping variance among the factors. The equation for the structure matrix is:
C = AΦ
The structure matrix is a product of the pattern matrix and the factor correlation matrix. Most researchers interpret and report the pattern matrix rather than the structure matrix. However, if the researcher reports either the structure or the pattern matrix and also , then the interested reader can generate the other using the following formula. In oblique rotation, R, is produced as follows:
R = CA'
The reproduced correlation matrix is a product of the structure matrix and the transpose of the pattern matrix.
Numerous procedures for factor extraction and rotation are available. However, only those procedures available in IBM SPSS and SAS packages are summarized here.
Some of the issues raised in this section can be resolved through several different methods. Usually, different methods lead to the same conclusion; occasionally they do not. When they do not, results are judged by the interpretability and scientific utility of the solutions.
IBM SPSS, SAS, and SYSTAT each have a single program to handle both FA and PCA. The first two programs have numerous options for extraction and rotation and give the user considerable latitude in directing the progress of the analysis.
IBM SPSS FACTOR does a PCA or FA on a correlation matrix or a factor loading matrix, helpful to the researcher who is interested in higher- order factoring (extracting factors from previous FAs). Several extraction methods and a variety of orthogonal rotation methods are available. Oblique rotation is done using direct oblimin, one of the best methods currently available.
SAS FACTOR is another highly flexible, full- featured program for FA and PCA. About the only weakness is in screening for outliers. SAS FACTOR accepts rotated loading matrices, as long as factor correlations are provided, and can analyze a partial correlation or covariance matrix (with specification of variables to partial out). There are several options for extraction, as well as orthogonal and oblique rotation. Maximum- likelihood estimation provides a x2 test for number of factors. Standard errors may be requested for factor loadings with maximum- likelihood estimation and promax rotation. A target pattern matrix can be specified as a criterion for oblique rotation in confirmatory FA. Additional options include specification of proportion of variance to be accounted for in determining the number of factors to retain and the option to allow communalities to be greater than 1.0. The correlation matrix can be weighted to allow the generalized least squares method of extraction.
The current SYSTAT FACTOR program is less limited than earlier versions. Wilkinson (1990) advocated the use of PCA rather than FA because of the indeterminacy problem (Section 13.6.6). However, the program now does PFA (called IPA) as well as PCA and maximum likelihood (MLA) extraction. Four common methods of orthogonal rotation are provided, as well as provision for oblique rotation. SYSTAT FACTOR can accept correlation or covariance matrices as well as raw data.
The SYSTAT FACTOR program provides scree plots and plots of factor loadings and will optionally sort the loading matrix by size of loading to aid interpretation. Additional information is available by requesting that standardized component scores, their coefficients, and loadings be sent to a data file. Factor scores (from PFA or MLA) cannot be saved. Residual scores (actual minus predicted z-scores) also can be saved, as well as the sum of the squared residuals and a probability value for it.
Structural equation modeling (SEM) is a collection of statistical techniques that allow a set of relationships between one or more IVs, either continuous or discrete, and one or more DVs, either continuous or discrete, to be examined. Both IVs and DVs can be either factors or measured variables. Structural equation modeling is also referred to as causal modeling, causal analysis, simultaneous equation modeling, analysis of covariance structures, path analysis, or confirmatory factor analysis (CFA). The latter two are actually special types of SEM. SEM allows questions to be answered that involve multiple regression analyses of factors.
Several conventions are used in developing SEM diagrams. Measured variables, also called observed variables, indicators, or manifest variables, are represented by squares or rectangles. Factors have two or more indicators and are also called latent variables, constructs, or unobserved variables. Factors are represented by circles or ovals in path diagrams. Relationships between variables are indicated by lines; lack of a line connecting variables implies that no direct relationship has been hypothesized. Lines have either one or two arrows.
The part of the model that relates the measured variables to the factors is sometimes called the measurement model.
The first step in a SEM analysis is specification of a model, so this is a confirmatory rather than an exploratory technique. The model is estimated, evaluated, and perhaps modified. The goal of the analysis might be to test a model, to test specific hypotheses about a model, to modify an existing model, or to test a set of related models.
There are a number of advantages to the use of SEM. When relationships among factors are examined, the relationships are free of measurement error because the error has been estimated and removed, leaving only common variance. Reliability of measurement can be accounted for explicitly within the analysis by estimating and removing the measurement error. Additionally, complex relationships can be examined. When the phenomena of interest are complex and multidimensional, SEM is the only analysis that allows complete and simultaneous tests of all the relationships.
Unfortunately, there is a small price to pay for the flexibility that SEM offers. With the ability to analyze complex relationships among combinations of discrete and continuous variables - both observed and latent - comes more complexity and more ambiguity. Indeed, there is quite a bit of jargon and many choices of analytic techniques.
The data set is an empirical covariance matrix and the model produces an estimated population covariance matrix. The major question asked by SEM is, 'Does the model produce an estimated population covariance matrix that is consistent with the sample (observed) covariance matrix?' After the adequacy of the model is assessed, various other questions about specific aspects of the model are addressed.
SEM is a confirmatory technique in contrast to exploratory factor analysis. It is used most often to test a theory - maybe just a personal theory - but a theory nonetheless. Indeed, one cannot do SEM without prior knowledge of, or hypotheses about, potential relationships among variables. This is perhaps the largest difference between SEM and other techniques in this book and one of its greatest strengths. Planning, driven by theory, is essential to any SEM analysis.Although SEM is a confirmatory technique, there are ways to test a variety of different models after a model has been estimated. However, if numerous modifications of a model are tested in hopes of finding the best- fitting model, the researcher has moved to exploratory data analysis and appropriate steps need to be taken to protect against inflated Type I error levels. Searching for the best model is appropriate, provided significance levels are viewed cautiously and cross- validation with another sample is performed whenever possible.
SEM has developed a bad reputation in some circles, in part because of the use of SEM for exploratory work without the necessary controls. It may also be due, in part, to the use of the term causal modeling to refer to SEM. There is nothing causal, in the sense of inferring causality, about the use of SEM. Attributing causality is a design issue, not a statistical issue.
Covariances, like correlations, are less stable when estimated from small samples. SEM is based on covariances. Parameter estimates and chi-square tests of fit are also very sensitive to sample size. SEM, then, like factor analysis, is a large sample technique. Models with strong expected parameter estimates and reliable variables may require fewer participants. Although SEM is a large sample technique, new test statistics have been developed that allow for estimation of models with as few as 60 participants
Problems are associated with either deleting or estimating missing data. An advantage of structural modeling is that the missing data mechanism can be included in the model. Some of the software packages now include procedures for estimating missing data, including the EM algorithm.
Structural equation modelling has several assumption:
The idea behind SEM is that the hypothesized model has a set of underlying parameters which correspond to the regression coefficients, and the variances and covariances of the independent variables in the model. These parameters are estimated from the sample data to be a 'best guess' about population values. The estimated parameters are then combined by means of covariance algebra to produce an estimated population covariance matrix. This estimated population covariance matrix is compared with the sample covariance matrix and, ideally, the difference is very small and not statistically significant.
Covariance algebra is a helpful tool in calculating variances and covariances in SEM models; however, matrix methods are generally employed because covariance algebra becomes extremely tedious as models become increasingly complex. Covariance algebra is useful to demonstrate how parameter estimates are combined to produce an estimated population covariance matrix for a small example. The three basic rules in covariance algebra appear below where c is a constant and Xi is a random variable:
By the first rule, the covariance between a variable and a constant is zero. By the second rule, the covariance between two variables where one is multiplied by a constant is the same as the constant multiplied by the covariance between the two variables. By the third rule, the covariance between the sum (or difference) of two variables and a third variable is the sum of the covariance of the first variable and the third and the covariance of the second variable and the third.
To specify the model, a separate equation is written for each DV. For motivation, Y1:
Y1 = γ11X1 + ε1
To calculate the covariance between X1 (treatment group) and Y1 (dependent variable), the first step is substituting in the equation for Y1:
COV (X1, Y1) = COV (X1, γ11X1 + ε1)
The second step is distributing the first term, in this case X1:
COV (X1, Y1) = COV (X1γ11X1) + COV (X1ε1)
The last term in this equation, COV (X1ε1), is equal to zero by assumption because it is assumed that there are no covariances between errors and other variables. Now:
COV (X1, Y1) = γ11COV (X1X1)
by rule 2, and because the covariance of a variable with itself is just a variance,
COV (X1Y1) = γ11σX1X1
The estimated population covariance between X1 and Y1 is equal to the path coefficient times the variance of X1. This is the population covariance between X1 and Y1 as estimated from the model. If the model is good, the product of γ11σX1X1 produces a covariance that is very close to the sample covariance.
Following the same procedures, the covariance between Y1 and Y2 is:
COV (Y1, Y2) = COV (γ11X1 + ε1, β21Y1 + γ21X1 + ε2)
= COV (γ11X1β21Y1) + COV (γ11X1γ21X1) + COV (γ11X1ε2) + COV (ε1β21Y1) + COV (ε1γ21X1) + COV (ε1ε2) = COV (γ11β21σX1Y1) + COV (γ11γ21σX1Y1)
because, as can be seen in the diagram, the error terms e1 and e2 do not correlate with any other variables.
One method of model specification is the Bentler-Weeks method. In this method, every variable in the model, latent or measured, is either an IV or a DV. The parameters to be estimated are the (1) regression coefficients, and (2) the variances and the covariances of the independent variables in the model. Residual variables (errors) of measured variables are labeled E and errors of latent variables (called disturbances) are labeled D. It may seem odd that a residual variable is considered an IV but remember the familiar regression equation:
Y = Xβ + ε
where Y is the DV and X and ε are both IVs. In fact, the Bentler-Weeks model is a regression model, expressed in matrix algebra:
η = Bη + γξ
where, if q is the number of DVs and r is the number of IVs, then η (eta) is a q*1 vector of DVs, β (beta) is a q*q matrix of regression coefficients between DVs, γ (gamma) is a q *r matrix of regression coefficients between DVs and IVs, and ξ (xi) is an r*1 vector of IVs.
In the Bentler-Weeks model, only independent variables have covariances and these covariances are in Φ (phi), an r*r matrix. Therefore, the parameter matrices of the model are B, γ, and Φ. Unknown parameters in these matrices need to be estimated. The vectors of dependent variables, η, and independent variables, ξ, are not estimated.
To calculate the estimated population covariance matrix implied by the parameter estimates, selection matrices are first used to pull the measured variables out of the full parameter matrices. The selection matrix is simply labeled G and has elements that are either 1s or 0s. The resulting vector is labeled Y:

where Y is our name for the those measured variables that are dependent.
The independent measured variables are selected in a similar manner,
X = GX * ξ = V5
where X is our name for the independent measured variables. Computation of the estimated population covariance matrix proceeds by rewriting the basic structural modeling as:
η = (I - B)-1γξ
where I is simply an identity matrix the same size as B. This equation expresses the DVs as a linear combination of the IVs. At this point, the estimated population covariance matrix for the DVs, Σyy, is estimated using:

The estimated population covariance matrix between IVs and DVs is obtained similarly by:

Finally, the estimated population covariance matrix between IVs is estimated:

A χ2 statistic is computed based upon the function minimum when the solution has converged. The χ2 is evaluated with degrees of freedom equal to the difference between the total number of degrees of freedom and the number of parameters estimated. The degrees of freedom in SEM are equal to the amount of unique information in the sample variance/covariance matrix (variances and covariances) minus the number of parameters in the model to be estimated (regression coefficients and variances and covariances of independent variables). In a model with a few variables, it is easy to count the number of variances and covariances; however, in larger models, the number of data points is calculated as,
number of data points = (p(p + 1)) / 2
where p equals the number of measured variables.
Next, researchers usually examine the statistically significant relationships within the model. If the unstandardized coefficients in the three parameter matrices are divided by their respective standard errors, a z-score is obtained for each parameter that is evaluated in the usual manner:
z = parameter estimate / std error for estimate
Freeing a parameter means estimating the parameter.
In SEM, a model is specified, parameters for the model are estimated using sample data, and the parameters are used to produce the estimated population covariance matrix. But only models that are identified can be estimated. A model is said to be identified if there is a unique numerical solution for each of the parameters in the model. For example, say both that the variance of Y = 10 and that the variance of Y = α + β. Any two values can be substituted for a and b as long as they sum to 10. There is no unique numerical solution for either α or β, that is, there are an infinite number of combinations of two numbers that would sum to 10. Therefore, this single equation model is not identified. However, if we fix a to zero, then there is a unique solution for b, 10, and the equation is identified. It is possible to use covariance algebra to calculate equations and assess identification in very simple models; however, in large models, this procedure quickly becomes unwieldy.
After a model is specified, population parameters are estimated with the goal of minimizing the difference between the observed and estimated population covariance matrices. To accomplish this goal, a function, Q, is minimized where
Q = (s - σ(Θ))'W(s - σ(Θ))
where s is the vector of data (the observed sample covariance matrix stacked into a vector); σ is the vector of the estimated population covariance matrix (again, stacked into a vector); and Θ indicates that σ is derived from the parameters (the regression coefficients, variances, and covariances) of the model. W is the matrix that weights the squared differences between the sample and estimated population covariance matrix.
The trick is to select W to minimize the squared differences between observed and estimated population covariance matrices.
After the model has been specified and then estimated, the major question is, “Is it a good model?” One component of a “good” model is the fit between the sample covariance matrix and the estimated population covariance matrix. Like multiway frequency analysis and logistic regression, a good fit is sometimes indicated by a nonsignificant χ2. Unfortunately, assessment of fit is not always as straightforward as assessment of χ2. With large samples, trivial differences between sample and estimated population covariance matrices are often significant because the minimum of the function is multiplied by N – 1. With small samples, the computed χ2 may not be distributed as χ2, leading to inaccurate probability levels. Finally, when assumptions underlying the χ2test statistic are violated, the probability levels are inaccurate.
One method of conceptualizing goodness of fit is by thinking of a series of models all nested within one another. Nested models are like the hierarchical models in log- linear modeling discussed in Chapter 16. Nested models are models that are subsets of one another. At one end of the continuum is the independence model: the model that corresponds to completely unrelated variables. This model would have degrees of freedom equal to the number of data points minus the variances that are estimated. At the other end of the continuum is the saturated (full or perfect) model with zero degrees of freedom. Fit indices that employ a comparative fit approach place the estimated model somewhere along this continuum.
The Bentler-Bonett normed fit index (NFI) evaluates the estimated model by comparing the χ2 value of the model to the χ2 value of the independence model.
NFI = χ2indep - χ2model / χ2indep
This yields a descriptive fit index that lies in the 0-1 range. High values (greater than .95) are indicative of a good-fitting model.
There are at least two reasons for modifying a SEM model: to improve fit (especially in exploratory work) and to test hypotheses (in theoretical work). The three basic methods of model modification are chi- square difference tests, Lagrange multiplier tests (LM), and Wald tests. All are asymptotically equivalent under the null hypothesis (act the same as the sample size approaches infinity) but approach model modification differently.
Reliability is defined in the classic sense as the proportion of true variance relative to total variance (true plus error variance). Both the reliability and the proportion of variance of a measured variable are assessed through squared multiple correlation (SMC) where the measured variable is the DV and the factor is the IV. Each SMC is interpreted as the reliability of the measured variable in the analysis and as the proportion of variance in the variable that is accounted for by the factor, conceptually the same as a communality estimate in factor analysis. To calculate a SMC:
SMCvar i = λ2i / (λ2i + Θii)
The factor loading for variable i is squared and divided by that value plus the residual variance associated with the variable i.
This equation is applicable only for when there are no complex factor loadings or correlated errors. The proportion of variance in the variables accounted for by the factor is assessed as:
Rj2 = 1 - Dj2
The disturbance (residual) for the DV factor j is squared and subtracted from 1.
SEM assumes that measured variables are continuous and measured on an interval scale. Often, however, a researcher desires to include discrete and/or ordinally measured, categorical variables in an analysis. Because the data points in SEM are variances and covariances, the trick is to produce reasonable values for these types of variables for analysis.
Although each of the models estimated in this chapter uses data from a single sample, it is also possible to estimate and compare models that come from two or more samples, called multiple group models. The general null hypothesis tested in multiple group models is that the data from each group are from the same population. For example, if data from a sample of men and a sample of women are drawn for the small- sample example, the general null hypothesis tested is that the two groups are drawn from the same population. The analysis begins by developing good- fitting models in separate runs for each group. The models are then tested in one run with none of the parameters across models constrained to be equal. This unconstrained multiple group model serves as the baseline against which to judge more restricted models.
Following baseline model estimation, progressively more stringent constraints are specified by constraining various parameters across all groups. When parameters are constrained, they are forced to be equal to one another. After each set of constraints is added, a chi-square difference test is performed for each group between the less restrictive and more restrictive model. The goal is to not degrade the models by constraining parameters across the groups; therefore, you want a nonsignificant χ2. If a significant difference in χ2 is found between the models at any stage, the LM test is examined to locate the specific parameters that are different in the groups and these parameters are estimated separately in each group; that is, the specific across-group parameter constraints are released.
The four SEM programs discussed - EQS, LISREL, SAS CALIS, and AMOS - are full-service, multioption programs.
Multilevel (hierarchical) linear modeling (MLM) is used for research designs, where the data for participants are organized at more than one level. Different variables may be available for each level of analysis. MLM provides an alternative analysis to several different designs discussed elsewhere in this volume. Although the lowest level of data in MLM is usually an individual, it may instead be repeated measurements of individuals. Another useful application of MLM is as an alternative to ANCOVA, where DV scores are adjusted for covariates (individual differences) prior to testing treatment differences.
An advantage of MLM over alternative analyses is that independence of errors is not required. In fact, independence is often violated at each level of analysis. In addition, there may be interactions across levels of the hierarchy. Analyzing data organized into hierarchies as if they are all on the same level leads to both interpretational and statistical errors. A less common but equally misleading approach is to interpret individual-level analyses at the group level, leading to the atomistic fallacy. A multilevel model, on the other hand, permits prediction of individual scores adjusted for group differences as well as prediction of group scores adjusted for individual differences within groups. Statistically, if individual scores are used without taking into account the hierarchical structure, the Type I error rate is inflated because analyses are based on too many degrees of freedom that are not truly independent.
Multilevel models often are called random coefficient regression models because the regression coefficients (the intercepts and predictor slopes) may vary across groups (higher-level units), which are considered to be randomly sampled from a population of groups. In garden-variety (OLS) regression, it is the individual participants who are considered to be a random sample from some population; in multilevel modelling, the groups also are considered to be a random sample.
One advantage of the multilevel modelling approach over other ways of handling hierarchical data is the opportunity to include predictors at every level of analysis.
The primary questions in MLM are like the questions in multiple regression: degree of relationship among the DV and various IVs; importance of IVs; adding and changing IVs; contingencies among IVs; parameter estimates; and predicting DV scores for members of a new sample. However, additional questions can be answered when the hierarchical structure of the data is taken into account and random intercepts and slopes are permitted.
Correlated predictors are even more problematic in MLM than in simple linear regression. In MLM, equations at multiple levels are solved and correlations among predictors at all levels are taken into account simultaneously. Because effects of correlated predictors are all adjusted for each other, it becomes increasingly likely that none of their regression coefficients will be statistically significant. The best advice, then, is to choose a very small number of relatively uncorrelated predictors. A strong theoretical framework helps limit the number of predictors and facilitates decisions about how to treat them.
If interactions are formed, predictors in the interactions are bound to be correlated with their main effects. The problem of multicollinearity between interactions and their main effects can be solved by centering. Modeling high-level predictors at their own level is not always the best way to deal with them. If there are only a few of them, they are often best entered at the next lower level as categorical predictors.
The price of large, complex models is instability and a requirement for a substantial sample size at each level. Even small models, with only a few predictors, grow rapidly as equations are added at higher levels of analysis. Therefore, large samples are necessary even if there are only a few predictors. As a maximum likelihood technique, a sample size of at least 60 is required if only 5 or fewer parameters are estimated. Parameters to be estimated include intercepts and slopes as well as effects of interest at each level. In practice, convergence often is difficult even with larger sample sizes.
As in most analyses, increasing sample sizes increases power, while smaller effect sizes and larger standard errors decrease power.
Multilevel linear modelling has several assumptions that need to be met.
Small data sets are difficult to analyze through MLM. The maximum likelihood procedure fails to converge with multiple equations unless samples are large enough to support those equations. Nor is it convenient to apply MLM to a matrix (or a series of matrices). MLM is often conducted in a sequence of steps. Therefore, the equations and computer analyses are divided into three models of increasing complexity.
The MLM is expressed as a set of regression equations. In the level-1 equation for the interceptsonly model (a model without predictors), the response (DV score) for an individual is predicted by an intercept that varies across groups. The intercepts-only model is of singular importance in MLM because it provides information about the intraclass correlation, a value helpful in determining whether a multilevel model is required or not.
In the level-1 equation for the interceptsonly model (a model without predictors), the response (DV score) for an individual is predicted by an intercept that varies across groups.
Yij = β0j + eij
An individual score on the DV, Yij, is the sum of an intercept (mean), β0j, that can vary over the j groups and individual error, eij (the deviation of an individual from her or his group mean). The terms Y and e that in ordinary regression have a single subscript i, indicating case, now have two subscripts, ij, indicating case and group. The intercept, β0 now also has a subscript, j, indicating that the coefficient varies over groups. That is, each group could have a separate equation.
β0j is not likely to have a single value because its value depends on the group. Instead, a parameter estimate (τ00) and standard error are developed for the variance of a random effect. The parameter estimate reflects the degree of variance for a random effect; large parameter estimates reflect effects that are highly variable. These parameter estimates and their significance tests are shown in computer runs to follow. The z test of the random component, τ00 divided by its standard error, evaluates whether the groups vary more than that would be expected by chance.
The second-level analysis (based on groups as research units) for the intercepts-only model uses the level-1 intercept (group mean) as the DV. To predict the intercept for group j:
βoj = γ00 + uoj
An intercept for a run, βoj, is predicted from the average intercept over groups when there are no predictors, γ00, and group error, uoj (deviation from average intercept for group j). A separate equation could be written for each group.
Yij = γ00 + uoj + eij
The average intercept (mean) is γ00 (a fixed component). The two random components are uoj (the deviation in intercept for cases in group j) and ej (the deviation for case i from its group j).
The full model adds a second-level predictor to the model with a first-level predictor. The level-1 equation for a model with predictors at both levels does not differ from that of the level-1 equations for a model with a predictor only at the first level. That is, the inclusion of the second- level IV does not affect the equations for the first level of analysis. However, it can affect the results of those equations, because all effects are adjusted for all other effects.
The second- level analysis (based on groups as research units) uses all three variables: level-1 intercept, level-1 slope, and the level-2 IV. For these analyses, level-1 slope and intercept again are considered DVs. To predict the random intercept:
βoj = γoo + γ01Wj + u0j
An intercept for a run βoj, is predicted by γ00 (the average intercept over groups when all predictors are zero), the slope γ01 (for the relationship between the intercepts of the level-1 analysis and the levels of the fixed level-2 IV) multiplied by Wj (the average value of the IV for the group), and error, u0j (the deviation from average intercept for group j). The coefficient, γ01, is the relationship between the original DV, Yij, and the IV, Wj. To predict the random slope for the level-1 predictor:
β1j = γ10 + u1j
The slope for a run, β1j, is predicted by a single intercept, γ10 (the average level-1 predictor-DV slope), and error, u1j (the deviation from average slope for group j). So:
Yij = γoo + γ01Wj + u0j (γ10 + u1j )Xij + eij
Thus the entire two- level equation, with terms rearranged, is: Yij = γoo + γ01Wj + γ10Xij + u0j + u1jXij + eij
The three fixed components are γoo (the average intercept- grand mean), γ01Wj (the overall slope for the relationship between a level-2 predictor and the DV times the second- level predictor score for group j), and γ10Xij (the unique effect of group j on slope when the value of the level-2 predictor W is zero times the first- level predictor score for case i in group j). The three random components are u0j (the deviation in intercept for cases in group j), u1jXij (the deviation in slope for cases in group j times the first- level predictor score for case i in group j), and eij (the deviation for case i from its group j).
One of the more common uses of MLM is to analyze a repeated- measures design, which violates some of the requirements of repeated- measures ANOVA. Longitudinal designs (called growth curve data) are handled in MLM by setting measurement occasions as the lowest level of analysis with cases the grouping variable. However, the repeated measures need not be limited to the first level of analysis. A big advantage of MLM over repeated-measures ANOVA is that there is no requirement for complete data over occasions (although it is assumed that data are missing at random), nor is there need for equal numbers of cases or equal intervals of measurements for each case. Another important advantage of MLM for repeated- measures data is the opportunity to test individual differences in growth curves (or other patterns of responses over the repeated measure). Are the regression coefficients the same for all cases? Because each case has its own regression equation when random slopes and intercepts are specified, it is possible to evaluate whether individuals do indeed differ in their mean response and/or in their pattern of responses over the repeated measure.
An additional advantage is that sphericity (uncorrelated errors over time) is not an issue because, as a linear regression technique, MLM tests trends for individuals over time (if individuals are the grouping variable). Finally, you may create explicit time- related level-1 predictors, other than the time period itself. Time-related predictors (a.k.a. time- varying covariates) come in a variety of forms.
The intraclass correlation (ρ)10 is the ratio of variance between groups at the second level of the hierarchy (ski runs in the small-sample example) to variance within those groups. High values imply that the assumption of independence of errors is violated and that errors are correlated - that is, the grouping level matters. An intraclass correlation is like η2 in a one-way ANOVA, although in MLM the groups are not deliberately subjected to different treatments. The need for a hierarchical analysis depends partially on the size of the intraclass correlation. If ρ is trivial, there is no meaningful average difference among groups on the DV, and data may be analyzed at the individual (first) level, unless there are predictors and groups that differ in their relationships between the predictors and the DV.
The intraclass correlation is calculated when there is a random intercept but no random slopes (because then there would be different correlations for cases with different values of a predictor). Therefore, ρ is calculated from the two- level intercept-only model.
Interactions of interest in MLM can be within a level or across levels. The small-sample example has only one predictor at each level. Inclusion of interactions is straightforward in MLM and follows the conventions of multiple regression: continuous predictors, from which interactions are formed, are centered and the interaction term is added.
Three issues arise regarding statistical inference in MLM.
The programs discussed in this chapter vary widely in the kinds of models they analyze and even in the results of analyses of the same models. These programs have far less in common than those used for other statistical techniques. SAS MIXED is part of the SAS package and is used for many analyses other than MLM. Indeed, at the time of writing this chapter, neither the SAS manuals nor special SAS publications directly address the issue of MLM. IBM SPSS MIXED MODEL is part of the IBM SPSS package starting with Version 11 and has been substantially revised since Version 11.5. MLwiN and HLM are stand- alone programs for MLM. SYSTAT MIXED REGRESSION is part of the SYSTAT package starting with Version 10.
Relationships among three or more discrete (categorical, qualitative) variables are studied through multiway frequency analysis or an extension of it called log- linear analysis. Relationships between two discrete variables are studied through the two- way χ2 test of association. If a third variable is added two- and three- way associations are sought through multiway frequency analysis.
To do a multiway frequency analysis, tables are formed that contain the one- way, two- way, three- way, and higher order associations. A linear model of (the logarithm of) expected cell frequencies is developed. The log- linear model starts with all of the one-, two-, three-, and higher- way associations and then eliminates as many of them as possible while still maintaining an adequate fit between expected and observed cell frequencies. The researcher may consider one of the variables a DV whereas the others are considered IVs.
The purpose of multiway frequency analysis is to discover associations among discrete variables. Once a preliminary search for associations is complete, a model is fit that includes only the associations necessary to reproduce the observed frequencies. Each cell has its own combination of parameter estimates for the associations retained in the model. The parameter estimates are used to predict cell frequency, and they also reflect the importance of each effect to the frequency in that cell. If one of the variables is a DV, the odds that a case falls into one of its categories can be predicted from the cell’s combination of parameter estimates.
As a nonparametric statistical technique with no assumptions about population distributions, multiway frequency analysis is remarkably free of limitations. The technique can be applied almost universally, even to continuous variables that fail to meet distributional assumptions of parametric statistics if the variables are cut into discrete categories.
With the enormous flexibility of current programs for log-linear analysis, many of the questions posed by highly complex data sets can be answered. However, the greatest danger in the use of this analysis is inclusion of so many variables that interpretation boggles the mind - a danger frequently noted in multifactorial analysis of variance, as well.
The only limitations to using multiway frequency analysis are the requirement for independence, adequate sample size, and the size of the expected frequency in each cell. During interpretation, however, certain cells may turn out to be poorly predicted by the solution.
Analysis of multiway frequency tables typically requires three steps: (1) screening, (2) choosing and testing appropriate models, and (3) evaluating and interpreting the selected model.
If only a single association is of interest, as is usually the case in the analysis of a two- way table, the familiar χ2 statistic is used:

where f0 represents observed frequencies in each cell of the table and Fe represents the expected frequencies in each cell under the null hypothesis of independence (no association) between the two variables. Summation is over all cells in the two- way table. If the goodness-of-fit tests for the two marginal effects are also computed, the usual χ2 tests for the 2 one- way and 1 two-way effects do not sum to total x2 . This situation is similar to that of unequal-n ANOVA, where F tests of main effects and interactions are not independent. Because overlapping variance cannot be unambiguously assigned to effects, and because overlapping variance is repeatedly analyzed, interpretation of results is not clear- cut. In multiway frequency tables, as in ANOVA, nonadditivity of χ2 becomes more serious as additional variables produce higher order (e.g., three-way and four-way) associations.
An alternative strategy is to use the likelihood ratio statistic, G2. The likelihood ratio statistic is distributed as χ2 so the χ2 tables can be used to evaluate significance. However, under certain conditions, G2 has the property of additivity of effects. For example, in a twoway analysis:
G2T = G2A + G2B + G2AB
The test of overall association within a two- way table, G2T is the sum of the first- order goodness-of-fit tests, G2A and G2B and the G2AB test of association, G2, like χ2, has a single equation for its various manifestations that differ among themselves only in how the expected frequencies are found.
G2 = 2 Σ(f0)ln(f0>Fe)
For each cell, the natural logarithm of the ratio of obtained to expected frequency is multiplied by the obtained frequency. These values are summed over cells, and the sum is doubled to produce the likelihood ratio statistics.
Screening is done if the researcher is data snooping and wishes simply to identify statistically significant effects. Screening is also done if the researcher hypothesizes a full model, a model with all possible effects included. Screening is not done if the researcher has hypothesized an incomplete model, a model with some effects included and others eliminated; in this case, the hypothesized model is tested and evaluated, followed by, perhaps, post hoc analysis.
The first step in screening is to determine if there are any effects to investigate. If there are, then screening progresses to a computation of Fe for each effect, a test of the reliability (significance) of each effect (finding G2 for the first-order effects, the second-order or two-way associations, the third- order or three- way associations, and so on), and an estimation of the size of the statistically significant effects. Because the latter equation is used for all tests of the observed frequencies (f0), the trick is to find the Fe necessary to test the various hypotheses.
If done by hand, the process starts by calculation of overall G2T which is used to test the hypothesis of no effects in the table (the hypothesis that all cells have equal frequencies). If this hypothesis cannot be rejected, there is no point to proceeding further. (Note that when all effects are tested simultaneously, as in computer programs, one can test either G2T or G2 for each of the effects, but not both, because degrees of freedom limit the number of hypotheses to be tested.) For the test of total effect:
Fe = N/rsp
Expected frequencies, Fe, for testing the hypothesis of no effects are the same for each cell in the table and are found by dividing the total frequency (N) by the number of cells in the table.
In some applications of multiway frequency analysis, results of screening provide sufficient information for the researcher. In the current example, for instance, the results are clear- cut. One first- order effect, preference for reading type, is statistically significant, as is the sex-by-profession association. Often, however, the results are not so evident and consistent, and/or the goal is to find the best model for predicting frequencies in each cell of the design. A log- linear model is developed where an additive regression- type equation is written for (the log of) expected frequency as a function of the effects in the design. The procedure is similar to multiple regression, where a predicted DV is obtained by combining the effects of several IVs. A full model includes all possible effects in a multiway frequency analysis. The full model for a three-way design of the example is:
ln Feijk = θ + λAi + λBj + λCk + λABij+ λACik + λBCjk + λABCijk
For each cell (the natural logarithm of), the expected frequency, ln Fe, is an additive sum of the effect parameters, ls, and a constant, θ. For each effect in the design, there are as many values of λ - called effect parameters - as there are levels in the effect, and these values sum to zero.
The purpose of modeling is to find the incomplete model with the fewest effects that still closely mimics the observed frequencies. Screening is done to avoid the necessity of exploring all possible incomplete models, an inhumane effort with large designs, even with computers. Effects that are found to be nonsignificant during the screening process are often omitted during modelling.
Models come in two flavors, hierarchical and nonhierarchical.
For hierarchical models, the optimal model is one that is not significantly worse than the next most complex one. Therefore, the choice among hierarchical models is made with reference to statistical criteria. There are no statistical criteria for choosing among nonhierarchical models and they are not recommended.
The optimal model, once chosen, is evaluated in terms of both the degree of fit to the overall data matrix (as discussed in the previous section) and the amount of deviation from fit in each cell.
Once a model is chosen, expected frequencies are computed for each cell and the deviation between the expected and the observed frequencies in each cell (the residual) is used to assess the adequacy of the model for fitting the observed frequency in that cell. In some cases, a model predicts the frequencies in some cells well, and in others very poorly, to give an indication of the combination of levels of variables for which the model is and is not adequate. Rather than trying to interpret raw differences, residuals usually are standardized by dividing the difference between observed and expected frequencies by the square root of the expected frequency to produce a z value.
There is a different linear combination of parameters for most cells, and the sizes of the parameters in a cell reflect the contribution of each of the effects in the model to the frequency found in that cell. Parameters are estimated for the model from the Fe in a manner that closely follows ANOVA. In ANOVA, the size of an effect for a cell is expressed as a deviation from the grand mean. Each cell has a different combination of deviations that correspond to the particular combination of levels of the statistically significant effects for that cell. In MFA, deviations are derived from natural logarithms of proportions: ln (Pijk). Expected frequencies for the model are converted to proportions by dividing Fe for each cell by N and then the proportions are changed to natural logarithms.
A model is hierarchical, or nested, if it includes all the lower effects contained in the highest-order association that is retained in the model. A hierarchical model for a four-way design, ABCD, with a significant three-way association, ABC, is A*B*C, A*B, A*C, B*C, and A, B, and C. The hierarchical model might or might not also include some of the other two-way associations and the D first-order effect.
A nonhierarchical model derived from the same four-way design includes only the significant two-way associations and first-order effects along with the significant threeway association; that is, a nonsignificant B*C association is included in a hierarchical model that retains the ABC effect but is not included automatically in a nonhierarchical model.
In log-linear analysis of multiway frequency tables, hierarchical models are the norm. Nonhierarchical models are suspect because higher order effects are confounded with lower order components. Therefore, it is best to explicitly include component lower order associations when specifying models in general log-linear programs. One major advantage of hierarchical models is the availability of a significance test for the difference between models, so that the most parsimonious adequately fitting model can be identified using inferential procedures. With nonhierarchical models, a statistical test for the difference between models is not available unless one of the candidate models happens to be nested in the other. IBM SPSS LOGLINEAR and GENLOG, and SAS CATMOD have the Newton–Raphson algorithm for assessing models and are considered general log- linear programs because they do not automatically impose hierarchical modeling. IBM SPSS HILOGLINEAR is restricted to hierarchical models.
A potential source of confusion is that tests of models look for statistical nonsignificance while tests of effects look for statistical significance. Both kinds of tests commonly use the same statistics - forms of χ2.
If you have one or more models hypothesized a priori, then there is no need for the strategies discussed in this section. The techniques in this section are used if you are building a model, or trying to find the most parsimonious incomplete model. As in all exploratory modeling, care should be taken in overgeneralizing results which may be subject to overfitting and inflated Type I error. Strategies for choosing a model differ depending on whether you are using IBM SPSS or SAS. Options and features differ among programs. You may find it handy to use one program to screen and another to evaluate models. Recall that hierarchical programs automatically include lower order components of higher order associations; general log- linear programs require that you explicitly include lower order components when specifying candidate hierarchical models.
Currently there are two programs for handling multiway frequency tables in the package: HILOGLINEAR, which deals with only hierarchical models, and GENLOG, which deals with hierarchical and nonhierarchical models.
No inferential tests of model components are provided in the program. Parameter estimates and their z tests are available for any specified model, along with their 95% confidence intervals. However, the parameter estimates are reported by single degrees of freedom, so that a factor with more than two categories has no omnibus significance test reported for either its main effect or its association with other effects. No quick screening for k-way tests is available. Screening information can be gleaned from a full model run, but identifying an appropriate model may be tedious with a large number of factors. Both IBM SPSS programs offer residuals plots. IBM SPSS GENLOG is the only program offering specification of Poisson models, which do not require that the analysis be conditional on total sample size.
SAS CATMOD is a general program for modeling discrete data, of which log- linear modeling is only one type. The program is primarily set up for logit analyses where one variable is the DV but provision is made for log- linear models where no such distinction is made. The program offers simple designation of logit models, contrasts, and single df tests of parameters as well as maximum likelihood tests of more complex components. The program lacks provision for continuous covariates and stepwise model building procedures. SAS CATMOD uses different algorithms from the other three programs both for parameter estimation and model testing.
SYSTAT LOGLIN is a general program for log- linear analysis of categorical data. The program uses its typical MODEL statement to set up the full, saturated, model (i.e., observed frequencies) on the left- hand side of the equation, and the desired model to be tested on the right- hand side. Structural zeros can be specified, and several options are available for controlling the iterative processing of model estimation. All of the usual descriptive and parameter estimate statistics are available, as well as multiple tests of effects in the model, both hierarchical and nonhierarchical. The program also prints outlying cells, designated 'outlandish'. Estimated frequencies and parameter estimates can be saved to a file.
To facilitate choice of the most useful technique to answer your research question, the emphasis has been on differences among statistical methods. We have repeatedly hinted, however, that most of these techniques are special applications of the general linear model (GLM). The goal of this chapter is to introduce the GLM and to fit the various techniques into the model. In addition to the aesthetic pleasure provided by insight into the GLM, an understanding of it provides a great deal of flexibility in data analysis by promoting use of more sophisticated statistical techniques and computer programs. Most data sets are fruitfully analyzed by one or more of several techniques.
Linearity and additivity are important to the GLM. Pairs of variables are assumed to have a linear relationship with each other; that is, it is assumed that relationships between pairs of variables are adequately represented by a straight line. Additivity is also relevant, because if one set of variables is to be predicted by a set of other variables, the effects of the variables within the set are additive in the prediction equation. The second variable in the set adds predictability to the first one, the third adds to the first two, and so on. In all multivariate solutions, the equation relating sets of variables is composed of a series of weighted terms added together. These assumptions, however, do not prevent inclusion of variables with curvilinear or multiplicative relationships. As discussed throughout this book, variables can be multiplied together, raised to powers, dichotomized, transformed, or recoded so that even complex relationships are evaluated within the GLM.
The GLM is based on prediction or, in jargon, regression. A regression equation represents the value of a DV, Y, as a combination of one or more IVs, Xs, plus error. The simplest case of the GLM, then, is the familiar bivariate regression:
A + BX + e = Y
where, B is the change in Y associated with a one-unit change in X; A is a constant representing the value of Y when X is 0; and e is a random variable representing error of prediction.
If X and Y are converted to standard z-scores, zx and zy, they are now measured on the same scale and cross at the point where both z-scores equal 0. The constant A automatically becomes 0 because zy is 0 when zx is 0. Further, after standardization of variances to 1, slope is measured in equal units (rather than the possibly unequal units of X and Y raw scores) and now represents the strength of the relationship between X and Y; in bivariate regression with standardized variables, it is equal to the Pearson product- moment correlation coefficient. The closer β is to 1.00 or -1.00, the better the prediction of Y from X (or X from Y). The equation then simplifies to:
βzx + e = zy
One distinction that is sometimes important in statistics is whether data are continuous or discrete. There are, then, three forms of bivariate regression for situations where X and Y are (1) both continuous - analyzed by Pearson product-moment correlation, (2) mixed, with X dichotomous and Y continuous - analyzed by point biserial correlation, and (3) both dichotomous - analyzed by phi coefficient. In fact, these three forms of correlation are identical. If the dichotomous variable is coded 0– 1, all the correlations can be calculated using the equation for Pearson product- moment correlation.
The first generalization of the simple bivariate form of the GLM is to increase the number of IVs, Xs, used to predict Y. It is here that the additivity of the model first becomes apparent. In standardized form:

That is, Y is predicted by a weighted sum of Xs. The weights, bi, no longer reflect the correlation between Y and each X because they are also affected by correlations among the Xs. Here again there are special statistical techniques associated with whether all Xs are continuous; here also, with appropriate coding, the most general form of the equation can be used to solve all the special cases. If Y and all Xs are continuous, the special statistical technique is multiple regression.
The latter equation is used to describe the multiple regression problem. But if Y is continuous and all Xs are discrete, we have the special case of regression known as analysis of variance. The values of X represent 'groups' and the emphasis is on finding mean differences in Y among groups rather than on predicting Y, but the basic equation is the same. A significant difference among groups implies that knowledge of X can be used to predict performance on Y.
Analysis of variance problems can be solved through multiple regression computer programs. There are as many Xs as there are degrees of freedom for the effects. For example, in a one way design, three groups are recoded into two dichotomous Xs - one representing the first group versus the other two and the second representing the second group versus the other two. The third group is those who are not in either of the other two groups. Inclusion of a third X would produce singularity because it is perfectly predictable from the combination of the other two.
The GLM takes a major leap when the Y side of the equation is expanded because more than one equation may be required to relate the Xs to the Ys:

where m equals k or p, whichever is smaller, and γ are regression weights for the standardized Y variables.
In general, there are as many equations as the number of X or Y variables, whichever is smaller. When there is only one Y, Xs are combined to produce one straight- line relationship with Y. Once there is more than one Y, however, combined Ys and combined Xs may fit together in several different ways. Each combination of Ys and Xs is a root. Roots are called by other names in the special statistical technique in which they are developed: discriminant functions, principal components, canonical variates, and so forth. Full multivariate techniques need multidimensional space to describe relationships among variables. With 2 df, two dimensions might be needed. With 3 df, up to three dimensions might be needed, and so on.
The number of roots necessary to describe the relationship between two sets of variables may be smaller than the number of roots maximally available. For this reason, the error term for Equation 17.4 is not necessarily associated with the mth root. It is associated with the last necessary root, with “necessary” statistically or psychometrically defined.
As with simpler forms of the GLM, specialized statistical techniques are associated with whether variables are continuous.
For most data sets, there is more than one appropriate analytical strategy, and choice among them depends on considerations such as how the variables are interrelated, your preference for interpreting statistics associated with certain techniques, and the audience you intend to address.
A data set for which alternative strategies are appropriate has groups of people who receive one of three types of treatments: behavior modification, short- term psychotherapy, or a waiting- list control group. Suppose a great many variables are measured— self- reports of symptoms and moods, reports of family members, therapist reports, and a host of personality and attitudinal tests - the major goal of the analysis is probably to find out if, and on which variable(s), the groups differ after treatment. The obvious strategy is MANOVA, but a likely problem is that the number of variables exceeds the number of clients in some group, leading to singularity. Further, with so many variables, some are likely to be highly related to combinations of others. You could choose among them or combine them on some rational basis, or you might choose first to look at empirical relationships among them.
Time-series analysis is used when observations are made repeatedly over 50 or more time periods. Sometimes, the observations are from a single case, but more often they are aggregate scores from many cases. One goal of the analysis is to identify patterns in the sequence of numbers over time, which are correlated with themselves, but offset in time. Another goal in many research applications is to test the impact of one or more interventions (IVs). Time-series analysis is also used to forecast future patterns of events or to compare series of different kinds of events.
As in other regression analyses, a score is decomposed into several potential elements. One of the elements is a random process, called a shock. Shocks are similar to the error terms in other analyses. Overlaying this random element are numerous potential patterns.
The model described in this chapter is the auto-regressive, integrated, moving average, called an ARIMA (p, d, q) model.
The first three steps in the analysis are identification, estimation, and diagnosis. These steps are devoted to modelling the patterns in the data.
However, often the goal is to assess the impact of an intervention. The intervention is occasionally experimental, featuring random assignment of cases to the levels of treatment, control of other variables, and manipulation of the levels of the IV by the researcher.
An intervention is tested for significance in the traditional manner, once the endogenous patterns in the data have been reduced to error through modeling. The effects of interventions considered in this chapter vary in both onset and duration; the onset of the effect may be abrupt or gradual and the duration of the effect may be either permanent or temporary. Effects of interventions are superimposed on the other patterns in the data.
True experiments typically assess the DV far fewer than 50 times and therefore cannot use time-series analysis. More likely with experiments, there is a single measure before the intervention (a pretest) and, perhaps, a few measures after the intervention to assess immediate and follow-up effects. In this situation, repeated measures ANOVA is more appropriate as long as heterogeneity of covariance (sphericity) is taken into account.
The major research questions involve the patterns in the series, the predicted value of the scores in the near future, and the effect of an intervention (an IV). Less common questions address the relationships among time series. It should be understood that this chapter barely scratches the surface of the complex world of time-series analysis. Only those questions that are relatively easily addressed in IBM SPSS and SAS are discussed.
The usual cautions apply with regard to causal inference in time-series analysis. Only with random assignment to treatment conditions, control of extraneous variables, and manipulation of the intervention(s) can cause reasonably be inferred when differences associated with an intervention are observed. Quasi-experiments are designed to rule out as many alternative sources of influence on the DV as possible, but causal inference is much weaker in any design that falls short of the requirements of a true experiment.
The random shocks that perturb the system are considered to be independent and normally distributed with mean zero and constant variance over time. Contingencies among scores over time are part of the model that is developed during identification and estimation. If the model is good, all sequential contingencies are removed so that you are left with the randomly distributed shocks. The residuals, then, are a reflection of the random shocks: independent and normally distributed, with mean zero and homogeneity of variance. It is expressly assumed that there are correlations in the sequence of observations over time that have been adequately modeled. This assumption is tested during the diagnostic phase when remaining, as yet unaccounted for patterns are sought among the residuals. Outliers among scores are sought before modelling and among the residuals once the model is developed.
Time-series analysis has four assumptions that need to be met:
The ARIMA (auto-regressive, integrated, moving average) model of a time series is defined by three terms (p, d, q). Identification of a time series is the process of finding integer, usually very small (e.g., 0, 1, or 2), values of p, d, and q that model the patterns in the data. When the value is 0, the element is not needed in the model. The middle element, d, is investigated before p and q. The goal is to determine if the process is stationary and, if not, to make it stationary before determining the values of p and q. Recall that a stationary process has a constant mean and variance over the time period of the study.
The first step in the analysis is to plot the sequence of scores over weeks by selecting Graphs, and then Sequence. The two relevant features of the plot are central tendency and dispersion. Is the mean apparently shifting over the time period? Is the dispersion increasing or decreasing over the time period?
There are possible shifts in both the mean and the dispersion over time for this series. The mean may be edging upwards, and the variability may be increasing. If the mean is changing, the trend is removed by differencing once or twice. If the variability is changing, the process may be made stationary by logarithmic transformation. Differencing the scores is the easiest way to make a nonstationary mean stationary (flat). The number of times you have to difference the scores to make the process stationary determines the value of d. If d = 0, the model is already stationary and has no trend. When the series is differenced once, d = 1 and linear trend is removed. When the difference is then differenced, d = 2 and both linear and quadratic trend are removed. For nonstationary series, d values of 1 or 2 are usually adequate to make the mean stationary. Differencing simply means subtracting the value of an earlier observation from the value of a later observation.
In the simplest time series, an observation at a time period simply reflects the random shock at that time period, at, that is:
Yt = ai
The random shocks are independent, with constant mean and variance, and so are the observations. If there is trend in the data, however, the score also reflects that trend as represented by the slope of the process. In this slightly more complex model, the observation at the current time, Yt, depends on the value of the previous observation, Yt-1 the slope, and the random shock at the current time period:
Yt = θ0(Yt-1)+at
To see if the process is stationary after linear trend is removed, the first difference scores at lag 1 are plotted against weeks. If the process is now stationary, the line will be basically horizontal with constant variance.The series now appears stationary with respect to central tendency, so second differencing does not appear necessary. However, the variability seems to be increasing over time. Transformation is considered for series in which variance changes over time and differencing does not stabilize the variance.
The auto-regressive components represent the memory of the process for preceding observations. The value of p is the number of auto-regressive components in an ARIMA (p, d, q ) model. The value of p is 0 if there is no relationship between adjacent observations. When the value of p is 1, there is a relationship between observations at lag 1 and the correlation coefficient φ1 is the magnitude of the relationship. When the value of p is 2, there is a relationship between observations at lag 2 and the correlation coefficient φ2 is the magnitude of the relationship. Thus, p is the number of correlations you need to model the relationship. For example, a model with p = 2, ARIMA (2, 0, 0), is:
Yt = ϕ1Yt-1 + ϕ2Yt-2 + at
The moving average components represent the memory of the process for preceding random shocks. The value q indicates the number of moving average components in an ARIMA (p, d, q ) model. When q is zero, there are no moving average components. When q is 1, there is a relationship between the current score and the random shock at lag 1, and the correlation coefficient u1 represents the magnitude of the relationship. When q is 2, there is a relationship between the current score and the random shock at lag 2, and the correlation coefficient u2 represents the magnitude of the relationship. Thus, an ARIMA (0, 0, 2) model is:
Yt = at - θ1at-1 - θ2at-2
Somewhat rarely, a series has both auto-regressive and moving average components so both types of correlations are required to model the patterns. If both elements are present only at lag 1, the equation is:
Yt = ϕ1Yt-1 - θ1at-1 + at
Models are identified through patterns in their ACFs (autocorrelation functions) and PACFs (partial autocorrelation functions). Both autocorrelations and partial autocorrelations are computed for sequential lags in the series. The first lag has an autocorrelation between Yt-1 and Yt, the second lag has both an autocorrelation and partial autocorrelation between Yt-2 and Yt, and so on. ACFs and PACFs are the functions across all the lags. The equation for autocorrelation is similar to bivariate r except that the overall mean Y is subtracted from each Yt and from each Yt-k, and the denominator is the variance of the whole series.

where N is the number of observations in the whole series and k is the lag. Y is the mean of the whole series and the denominator is the variance of the whole series.
The standard error of an autocorrelation is based on the squared autocorrelations from all previous lags. At lag 1, there are no previous autocorrelations, so r 2 0 is set to 0.

The equations for computing partial autocorrelations are much more complex, and involve a recursive technique. However, the standard error for a partial autocorrelation is simple and the same at all lags:

The series used to compute the autocorrelation is the series that is to be analyzed. Because differencing is used with the small-sample example to remove linear trend, it is the differenced scores that are analyzed.
If an autocorrelation at some lag is significantly different from zero, the correlation is included in the ARIMA model. Similarly, if a partial autocorrelation at some lag is significantly different from zero, it is included in the ARIMA model as well. The significance of full and partial autocorrelations is assessed using their standard errors. Although you can look at the autocorrelations and partial autocorrelations numerically, it is standard practice to plot them. The center vertical (or horizontal) line for these plots represents full or partial autocorrelations of zero; then symbols such as * or _ or bars are used to represent the size and direction of the autocorrelation and partial autocorrelation at each lag. You compare these obtained plots with standard, and somewhat idealized, patterns that are shown by various ARIMA models.
The boundary lines around the functions are the 95% confidence bounds. The pattern here is a large, negative autocorrelation at lag 1 and a decaying PACF, suggestive of an ARIMA (0, 0, 1) model.
Estimating the values of parameters in models consists of estimating the f parameter(s) from an auto-regressive model or the parameter(s) from a moving average model. As indicated by McDowall and colleagues, the following rules apply:
ϕ1 + ϕ2 < 1 and ϕ2 - ϕ1 < 1
These are called the bounds of stationarity for the auto-regressive parameter(s).
θ1 + θ2 < 1 and θ2 - θ1 < 1
These are called the bounds of invertibility for the moving average parameter(s).
Complex and iterative maximum likelihood procedures are used to estimate these parameters. The reason for this is seen in an equation for θ1 :

How well does the model fit the data? Are the values of observations predicted from the model close to actual ones? If the model is good, the residuals (differences between actual and predicted values) of the model are a series of random errors. These residuals form a set of observations that are examined the same way as any time series. ACFs and PACFs for the residuals of the model are examined.
There are two major varieties of time-series analysis: time domain (including Box–Jenkins ARIMA analysis) and spectral domain (including Fourier spectral analysis).
Either time or spectral domain analyses can be used for identification, estimation, and diagnosis of a time series. However, current statistical software offers no assistance for intervention analysis using spectral methods. As a result, this chapter is limited to time domain analyses. Numerous complexities are available with these analyses, however—seasonal autocorrelation and one or more interventions (and with different effects), to name a few.
Seasonal autocorrelation is distinguished from local autocorrelation in that it is predictably spaced in time. Observations gathered monthly, for example, are often expected to have a spike at lag 12 because many behaviors vary consistently from month to month over the year. Similarly, observations made daily could easily have a spike at lag 7, and observations gathered hourly often have a spike at lag 24. These seasonal cycles can often be postulated a priori, while local cycles are inferred from the data.
Like local cycles, seasonal cycles show up in plots of ACFs and PACFs as spikes. However, they show up at the appropriate lag for the cycle. Like local cycles, these autocorrelations can also be auto- regressive or moving average (or both). And, like local autocorrelation, a seasonal autoregressive component tends to produce a decaying ACF function and spikes on the PACF while a moving average component tends to produce the reverse pattern. When there are both local and seasonal trends, d, multiple differencing is used. Seasonal models may be either additive or multiplicative.
Intervention, or interrupted, time-series analyses compare observations before and after some identifiable event. In quasi- experiments, the intervention is an attempt at experimental manipulation; however, the techniques are applicable to analysis of any event that occurs during the time series. The goal is to evaluate the impact of the intervention. Interventions differ in both the onset (abrupt or gradual) and duration (permanent or temporary) of the effects they produce.
The connection between an intervention and its effects is called a transfer function. In time-series jargon, an impulse intervention is also called an impulse or a pulse indicator and the effect is called an impulse or a pulse function. An effect with abrupt onset and permanent or long duration is called a step function. Because there are two levels of duration (permanent and temporary) and two levels of onset (abrupt and gradual), there are four possible combinations of effects, but the gradual-onset, short-term effect occurs rarely and requires curve fitting. Intervention analysis requires a column for the indicator variable that flags the occurrence of the event. With an impulse indicator, a code of 1 is applied in the indicator column at the single specific time period of the intervention and a code of 0 to the remaining time periods. When the intervention is of long duration or the effect is expected to persist, the column for the indicator contains 0s during the baseline time period and 1s at the time of and after the intervention.
Intervention analysis begins with identification, estimation, and diagnosis of the observations before intervention. The model, including the indicator variable, is then re-estimated for the entire series and the results diagnosed. The effect of the intervention is assessed by interpreting the coefficients for the indicator variable.
Transfer functions have two possible parameters. The first parameter, ω, is the magnitude of the asymptotic change in level after intervention. The second parameter, δ, reflects the onset of the change, the rate at which the post-intervention series approaches its asymptotic level. The ultimate change in level is:

Both ω and δ are tested for significance. If the null hypothesis that ω is 0 is retained, there is no impact of intervention. If ω is significant, the size of the change is ω. The d coefficient varies between 0 and 1. When δ is 0 (nonsignificant), the onset of the impact is abrupt. When d is significant but small, the series reaches its asymptotic level quickly. The closer δ is to 1, the more gradual the onset of change. SAS permits specification and evaluation of both ω and δ but IBM SPSS permits evaluation only of ω.
Abrupt, permanent effects, called step functions, are those that are expected to show an immediate impact and continue over the long term. All post-intervention observations are coded 1 and ω is specified, but not δ (or it is specified as 0).
Now suppose that the intervention has a strong, abrupt initial impact that dies out quickly. This is called a pulse effect.
Another possible effect of intervention is one that gradually reaches its peak effectiveness and then continues to be effective. The data set is the same as for an abrupt, permanent effect. However, now we are proposing that there is a linear growth in the effect of the intervention from the time it is implemented.
Multiple interventions are modeled by specifying more than one IV. Several interventions can be included in models when using either IBM SPSS ARIMA or SAS ARIMA.
Continuous variables are added to time-series analyses to serve several purposes.
These analyses have their own assumptions and limitations. As usual, use of a covariate assumes that it is not affected by the IV; for example, it is assumed that temperature is not affected by introduction of the profit-sharing plan in the company under investigation.
ARIMA models are identified by matching obtained patterns of ACF and PACF plots with idealized patterns. The best match often indicates which of the (p, d, q) parameters need to be included in the model, and at what size (0, 1, or 2). Sometimes, however, more than one pattern is suggested by the plots, or, like the small-sample example, the best pattern match does not reduce the residuals to random error. A first best guess is made on the basis of the ACF, PACF pattern and then, if the model fits the data poorly, another is tried out until the diagnostic phase is satisfactory.
If either an autocorrelation or a partial autocorrelation between observations k lags apart is statistically significant, the autocorrelation is included in the ARIMA model. The significance of an autocorrelation is evaluated from the 95% confidence intervals printed alongside it, or from the t distribution where the autocorrelation is divided by its standard error, or from the Box–Ljung statistic printed as standard IBM SPSS output. Alternatively, SAS ARIMA provides tests of sets of lags.
Effect size in time- series analysis is based on evaluation of residuals after finding an acceptable model. If there is an intervention, the residuals after the intervention is modeled are compared with the residuals before the intervention is modeled.
Two measures of effect size are available.


Forecasting refers to the process of predicting future observations from a known time series and is often the major goal in nonexperimental use of the analysis. However, prediction beyond the data is to be approached with caution. The farther the prediction beyond the actual data, the less reliable the prediction. And only a small percentage of the number of actual data points can be predicted before the forecast turns into a straight line.
Often the identification process suggests two or more candidate models. Sometimes, however, both models fail to meet the criteria, or both meet all of the criteria. Statistical tests are available to determine whether two models are reliably different. The difference in AIC for the two models is evaluated as x2 with df equal to the difference in the number of parameters for the two models.
x2 (df) = AIC(smaller model) - AIC(larger model)
SAS and IBM SPSS each have a single program for ARIMA time-series modeling. IBM SPSS has additional programs for producing time-series graphs and for forecasting seasonally adjusted time series. SAS has a variety of programs for time-series analysis, including a special one for seasonally adjusted time series. SYSTAT also has a time-series program, but it does not support intervention analysis.
Join with a free account for more service, or become a member for full access to exclusives and extra support of WorldSupporter >>
JoHo WorldSupporter mission and vision:
JoHo concept:
Volunteering: WorldSupporter moderators and Summary Supporters
Volunteering: Share your summaries or study notes
Student jobs: Part-time work as study assistant in Leiden


Search only via club, country, goal, study, topic or sector
Select any filter and click on Search to see results