This function checks for nestedness in a data frame. It identifies columns that are nested within other columns, meaning that the values in one column are subsets of the values in another column. The function returns a list of identified nested columns and can print the results to the console.
Usage
cols_nested(data, na.rm = FALSE, ignore = c("constant", "unique", "bijective"))Arguments
- data
The data frame to be checked for nestedness.
- na.rm
Remove NA values when checking for nestedness.
- ignore
A character vector of trivial cases of nesting to ignore, i.e. not count as nested. Any of
"constant"(xoryhas a single value),"unique"(xhas all unique values) and"bijective"(xandyhave a one-to-one correspondence). UseNULLto count all cases.
Value
A list of identified nested columns. The first element of each list is the child variable, and the second element is the parent variable.
Details
By default, trivial cases of nesting (constant columns, columns with all
unique values and pairs of bijective columns) are not reported as nested.
Use ignore to change this (see is_nested()).
See also
Other quality checks:
cols_all_unique(),
cols_bijective(),
cols_constant(),
cols_identify_all(),
cols_missing()