Various methods for performing parallel analysis. This function uses
future_lapply() for which a parallel processing plan can
be selected. To do so, call library(future) and, for example,
plan(multisession); see examples.
Usage
efa_parallel(
x = NULL,
N = NA,
n_vars = NA,
n_datasets = 1000,
percent = 95,
eigen_type = c("PCA", "SMC", "EFA"),
use = c("pairwise.complete.obs", "all.obs", "complete.obs", "everything",
"na.or.complete"),
cor_method = c("pearson", "spearman", "kendall", "poly", "tetra"),
decision_rule = c("means", "percentile", "crawford"),
n_factors = 1,
estimate_control = NULL,
...
)Source
Braeken, J., & van Assen, M. A. (2017). An empirical Kaiser criterion. Psychological Methods, 22, 450 – 466. http://dx.doi.org/10.1037/ met0000074
Crawford, A. V., Green, S. B., Levy, R., Lo, W. J., Scott, L., Svetina, D., & Thompson, M. S. (2010). Evaluation of parallel analysis methods for determining the number of factors. Educational and Psychological Measurement, 70(6), 885-901.
Glorfeld, L. W. (1995). An improvement on Horn's parallel analysis methodology for selecting the correct number of factors to retain. Educational and Psychological Measurement, 55(3), 377-393.
Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30(2), 179–185. doi: 10.1007/BF02289447
Arguments
- x
matrix or data.frame. The real data to compare the simulated eigenvalues against. Must not contain variables of classes other than numeric. Can be a correlation matrix or raw data.
- N
numeric. The number of cases / observations to simulate. Only has to be specified if
xis either a correlation matrix orNULL. If x contains raw data,Nis found from the dimensions ofx.- n_vars
numeric. The number of variables / indicators to simulate. Only has to be specified if
xis left asNULLas otherwise the dimensions are taken fromx.- n_datasets
numeric. The number of datasets to simulate. Default is 1000.
- percent
numeric. The percentile to take from the simulated eigenvalues. Default is 95.
- eigen_type
character. On what the eigenvalues should be found. Can be either "SMC", "PCA", or "EFA". If using "SMC", the diagonal of the correlation matrix is replaced by the squared multiple correlations (SMCs) of the indicators. If using "PCA", the diagonal values of the correlation matrices are left to be 1. If using "EFA", eigenvalues are found on the correlation matrices with the final communalities of an EFA solution as diagonal.
- use
character. Passed to
stats::cor()if raw data is given as input. Default is "pairwise.complete.obs".- cor_method
character. One of
"pearson","spearman", or"kendall", passed tostats::cor()."poly"and"tetra"are not supported becausePARALLELcompares the data against simulated continuous reference data. Default is "pearson".- decision_rule
character. Which rule to use to determine the number of factors to retain. Default is
"means", which will use the average simulated eigenvalues."percentile", uses the percentiles specified in percent."crawford"uses the 95th percentile for the first factor and the mean afterwards (based on Crawford et al, 2010). The"means"rule retains a factor whenever its real eigenvalue exceeds the average simulated one and thus tends to retain more factors than the more conservative"percentile"rule (Glorfeld, 1995).- n_factors
numeric. Number of factors to extract if "EFA" is included in
eigen_type. Default is 1.- estimate_control
an
estimate_control()object with the estimation settings for theefa_fit()fits (of both the real and the simulated data) when"EFA"is included ineigen_type.NULL(default) uses theefa_fit()defaults. The fits are unrotated, so no rotation settings apply.- ...
Additional arguments passed to
efa_fit(). For example,estimator, to change the estimator (default is "PAF"). PAF is more robust, but it will take longer compared to the other estimators available ("ML" and "ULS"). The estimation tuning knobs are not passed here; they live inestimate_control.
Value
An object of class efa_retention (see print.efa_retention() and
plot.efa_retention() for the print and plot methods). Its main fields are:
- n_factors
A named numeric vector with the suggested number of factors for each requested eigenvalue type (
"PCA","SMC", and/or"EFA"). These areNAwhen no real data are supplied (i.e. onlyNandn_varsare given). When every real eigenvalue exceeds its reference (no crossing is found), alln_varscomponents are retained and a warning is issued.- results
A list with one record per eigenvalue type, each holding the real eigenvalues (when real data were supplied) and the simulated reference eigenvalues (means and percentiles) used for printing and plotting.
- settings
A list of the settings used.
Details
Parallel analysis (Horn, 1965) compares the eigenvalues obtained from
the sample
correlation matrix against those of null model correlation matrices (i.e.,
with uncorrelated variables) of the same sample size. This way, it accounts
for the variation in eigenvalues introduced by sampling error and thus
eliminates the main problem inherent in the Kaiser-Guttman criterion
(efa_kgc()).
Parallel analysis is often argued to be one of the most accurate factor retention criteria. However, for highly correlated factor structures it has been shown to underestimate the correct number of factors. The reason for this is that a null model (uncorrelated variables) is used as reference. However, when factors are highly correlated, the first eigenvalue will be much larger compared to the following ones, as later eigenvalues are conditional on the earlier ones in the sequence and thus the shared variance is already accounted in the first eigenvalue (e.g., Braeken & van Assen, 2017).
The reference eigenvalues are obtained from simulated data, so the suggested number
of factors varies slightly from run to run. Call set.seed() beforehand to make a
run reproducible; the result is then also independent of the parallel plan set via
future::plan(), so it can be reproduced on a machine with a different number of
cores.
The efa_parallel function can also be called together with other factor
retention criteria in the efa_retain() function.
See also
efa_retain() as a wrapper function for this and the other factor
retention criteria.
Other factor retention criteria:
efa_cd(),
efa_ekc(),
efa_hull(),
efa_kgc(),
efa_map(),
efa_nest(),
efa_retain(),
efa_scree(),
efa_smt()
Examples
# \donttest{
# example without real data
pa_unreal <- efa_parallel(N = 500, n_vars = 10)
# example with correlation matrix with all eigen_types and PAF estimation
pa_paf <- efa_parallel(test_models$case_11b$cormat, N = 500)
# example with correlation matrix with all eigen_types and ML estimation
# this will be faster than the above with PAF)
pa_ml <- efa_parallel(test_models$case_11b$cormat, N = 500, estimator = "ML")
# }
if (FALSE) { # \dontrun{
# for parallel computation
future::plan(future::multisession)
pa_faster <- efa_parallel(test_models$case_11b$cormat, N = 500)
} # }