Skip to contents

Among the most important decisions for an exploratory factor analysis (EFA) is the choice of the number of factors to retain. Several factor retention criteria have been developed for this. With this function, various factor retention criteria can be performed simultaneously. Additionally, the data can be checked for their suitability for factor analysis.

Usage

efa_retain(
  x,
  criteria = c("CD", "EKC", "HULL", "MAP", "NEST", "PARALLEL"),
  suitability = TRUE,
  N = NA,
  use = c("pairwise.complete.obs", "all.obs", "complete.obs", "everything",
    "na.or.complete"),
  cor_method = c("pearson", "spearman", "kendall", "poly", "tetra"),
  n_factors_max = NA,
  N_pop = 10000,
  N_samples = 500,
  alpha = 0.3,
  ...,
  max_iter_CD = 50,
  n_fac_theor = NA,
  estimator = c("ML", "PAF", "ULS"),
  gof = c("CAF", "CFI", "RMSEA"),
  eigen_type_HULL = c("SMC", "PCA", "EFA"),
  eigen_type_other = c("SMC"),
  n_factors = 1,
  n_datasets = 1000,
  percent = 95,
  decision_rule = c("means", "percentile", "crawford"),
  ekc_type = c("BvA2017"),
  n_datasets_nest = 1000,
  alpha_nest = 0.05,
  show_progress = FALSE,
  estimate_control = NULL
)

Arguments

x

data.frame or matrix. Dataframe or matrix of raw data or matrix with correlations. If "CD" is included as a criterion, x must be raw data.

criteria

character. A vector with the factor retention methods to perform. Possible inputs are: "CD", "EKC", "HULL", "KGC", "MAP", "NEST","PARALLEL", "SCREE", and "SMT" (see details). The values are matched case-insensitively. By default, a subset of often used, well-performing methods are performed.

suitability

logical. Whether the data should be checked for suitability for factor analysis using the Bartlett's test of sphericity and the Kaiser-Meyer-Olkin criterion (see details). Default is TRUE.

N

numeric. The number of observations. Only needed if x is a correlation matrix.

use

character. Passed to stats::cor() if raw data is given as input. Default is "pairwise.complete.obs".

cor_method

character. Correlation computed from raw data: "pearson", "spearman", or "kendall" (passed to stats::cor()), or "poly" / "tetra" for polychoric / tetrachoric correlations (a two-step estimator with no empty-cell continuity correction). CD, PARALLEL, NEST, and HULL compare against simulated continuous data, and SMT relies on a normal-theory chi-square test; none of these support "poly" / "tetra", so they are skipped in that case. Default is "pearson".

n_factors_max

numeric. Passed to efa_cd(). The maximum number of factors to test against. Larger numbers will increase the duration the procedure takes, but test more possible solutions. If left NA (default), the maximum number of factors for which the model is still over-identified (df > 0) is used.

N_pop

numeric. Passed to efa_cd(). Size of finite populations of comparison data. Default is 10000.

N_samples

numeric. Passed to efa_cd(). Number of samples drawn from each population. Default is 500.

alpha

numeric. Passed to efa_cd(). The alpha level used to test the significance of the improvement added by an additional factor. Default is .30.

...

Further arguments passed to efa_fit() in efa_parallel() (also within efa_hull()), efa_kgc(), efa_scree(), and efa_nest(). The estimation tuning knobs are not passed here; they live in estimate_control. Note that the arguments listed after ... must be given by their full name (R matches an abbreviated name only against the arguments before ...), so that a tuning knob such as max_iter cannot be mistaken for max_iter_CD.

max_iter_CD

numeric. Passed to efa_cd(). The maximum number of iterations to perform after which the iterative PAF procedure is halted. Default is 50.

n_fac_theor

numeric. Passed to efa_hull(). Theoretical number of factors to retain. The maximum of this number and the number of factors suggested by efa_parallel() plus one will be used in the Hull method.

estimator

character. Passed to efa_fit() in efa_hull(), efa_kgc(), efa_scree(), efa_parallel(), and efa_nest(). The estimator to use. One of "PAF", "ULS", or "ML", for principal axis factoring, unweighted least squares, and maximum likelihood, respectively. The value is matched case-insensitively. In efa_kgc(), efa_scree(), and efa_parallel() it only takes effect when the respective eigen_type includes "EFA".

gof

character. Passed to efa_hull(). The goodness of fit index to use. Either "CAF", "CFI", or "RMSEA", or any combination of them. With the "PAF" estimator, only the CAF can be used as goodness of fit index. For details on the CAF, see Lorenzo-Seva, Timmerman, and Kiers (2011).

eigen_type_HULL

character. Passed to efa_parallel() in efa_hull(). On what the eigenvalues should be found in the parallel analysis. Can be one of "SMC", "PCA", or "EFA". If using "SMC" (default), the diagonal of the correlation matrices is replaced by the squared multiple correlations (SMCs) of the indicators. If using "PCA", the diagonal values of the correlation matrices are left to be 1. If using "EFA", eigenvalues are found on the correlation matrices with the final communalities of an EFA solution as diagonal.

eigen_type_other

character. Passed to efa_kgc(), efa_scree(), and efa_parallel(). The same as eigen_type_HULL, but multiple inputs are possible here (any combination of "PCA", "SMC", and "EFA"). Default is "SMC".

n_factors

numeric. Passed to efa_parallel() (also within efa_hull()), efa_kgc(), and efa_scree(). Number of factors to extract if "EFA" is included in eigen_type_HULL or eigen_type_other. Default is 1.

n_datasets

numeric. Passed to efa_parallel() (also within efa_hull()). The number of datasets to simulate. Default is 1000.

percent

numeric. Passed to efa_parallel() (also within efa_hull()). The percentile to take from the simulated eigenvalues. Default is 95.

decision_rule

character. Passed to efa_parallel() (also within efa_hull()). Which rule to use to determine the number of factors to retain. Default is "means", which will use the average simulated eigenvalues. "percentile", uses the percentiles specified in percent. "crawford" uses the 95th percentile for the first factor and the mean afterwards (based on Crawford et al, 2010).

ekc_type

character. Passed to the type argument of efa_ekc(). Either "BvA2017" for the original implementation by Braeken and van Assen (2017), or "AM2019" for the adapted implementation by Auerswald and Moshagen (2019).

n_datasets_nest

numeric. The number of datasets to simulate in efa_nest(). Default is 1000.

alpha_nest

numeric. The alpha level to use in efa_nest() (i.e., 1-alpha percentile of eigenvalues is used for reference values).

show_progress

logical. Whether a progress bar should be shown in the console. Default is FALSE.

estimate_control

an estimate_control() object with the estimation settings for the efa_fit() fits run by the criteria that fit a model (efa_hull(), efa_kgc(), efa_scree(), efa_parallel(), efa_nest(), and efa_smt()). NULL (default) uses the efa_fit() defaults. efa_cd(), efa_ekc(), and efa_map() fit no model, so it does not apply to them, and efa_smt() fits with maximum likelihood by definition, so only start_method takes effect there. All fits are unrotated, so no rotation settings apply.

Value

A list of class c("efa_retain", "N_FACTORS"), the trailing class keeping inherits(x, "N_FACTORS") working for code written against the superseded name. It contains

suitability

A list with the results from efa_bartlett() and efa_kmo() (bartlett and kmo), or NULL if suitability = FALSE.

outputs

A named list with one efa_retention object per factor retention criterion that was run (see, e.g., efa_ekc()).

n_factors

A named numeric vector with the suggested number of factors per criterion and, where a criterion has several variants, per variant (e.g. EKC_BvA2017 or PARALLEL_SMC). Criteria without a numeric suggestion (the scree plot) are not included.

not_run

A named character vector with the criteria that were skipped or failed and the reason, or NULL if all requested criteria ran.

settings

A list of the settings used.

Details

By default, the entered data are checked for suitability for factor analysis using the following methods (see respective documentations for details):

The available factor retention criteria are the following (see respective documentations for details):

The comparison data, parallel analysis, and NEST criteria compare the data against simulated reference data, so their suggested numbers of factors vary slightly from run to run; the Hull method does too, because it calls efa_parallel() to set its upper bound. Call set.seed() before efa_retain() to make them reproducible.

See also

efa_screen() for data screening before retention, and efa_fit() to extract the chosen number of factors.

Other factor retention criteria: efa_cd(), efa_ekc(), efa_hull(), efa_kgc(), efa_map(), efa_nest(), efa_parallel(), efa_scree(), efa_smt()

Examples

# \donttest{
# Default criteria, with correlation matrix and estimator "ML" (where needed)
# This will throw a warning for CD, as no raw data were specified
# The simulation-based criteria are seeded to make the run reproducible
set.seed(42)
nfac_all <- efa_retain(test_models$baseline$cormat, N = 500, estimator = "ML")
#> Warning: `x` is a correlation matrix, but "CD" needs raw data.
#>  Skipping "CD".

# The same as above, but without "CD"
nfac_wo_CD <- efa_retain(test_models$baseline$cormat, criteria = c("EKC",
                         "HULL", "PARALLEL", "NEST"), N = 500,
                         estimator = "ML")

# Use PAF instead of ML (this will take longer). For this, gof has
# to be set to "CAF" for the Hull method.
nfac_PAF <- efa_retain(test_models$baseline$cormat, criteria = c("EKC",
                       "HULL", "PARALLEL", "NEST"), N = 500,
                       estimator = "PAF", gof = "CAF")

# Do KGC and PARALLEL with only "PCA" type of eigenvalues
nfac_PCA <- efa_retain(test_models$baseline$cormat, criteria = c("EKC",
                       "HULL", "PARALLEL", "NEST"), N = 500,
                       estimator = "ML", eigen_type_other = "PCA")

# Use raw data, such that CD can also be performed
nfac_raw <- efa_retain(GRiPS_raw, estimator = "ML")
#> Warning: The suggested maximum number of factors was 2, but the Hull method needs at
#> least 3.
#>  Setting it to 3.
# }