SALib.analyze.delta module#
- SALib.analyze.delta.analyze(problem: Dict, X: ndarray, Y: ndarray, num_resamples: int = 100, conf_level: float = 0.95, print_to_console: bool = False, seed: int | None = None, y_resamples: int | None = None, method: str = 'all', bootstrap_savedf: str | None = None, bins_specs: Dict = {}) Dict[source]#
Perform Delta Moment-Independent Analysis on model outputs.
Returns a dictionary with keys: - ‘delta_balanced’ - ‘delta_balanced_conf’ - ‘delta_raw’ - ‘delta_raw_conf’ - ‘delta_step’ - ‘delta_step_conf’ - ‘S1’ - ‘S1_conf’ - ‘names’
where each entry is a list of size D (the number of parameters) containing the indices in the same order as the input file and ‘names’.
The different delta configurations/methods ‘balanced’, ‘step’, ‘raw’, their corresponding confidence scores (*_conf), and other keys in returned dictionary:
- ‘names’
input column/feature names in X
- ‘notes’
Notes and insights on analysis, warnings, errors, etc.
- ‘balanced’
Divides the input into bins during bootstrapping (default 10) and divides bootstrap subset equally among those bins
- ‘step’
Generates all entries in the input into 1 (X>0) or 0 (X<=0) and bootstrap subset is approximately 50%
- ‘raw’
Bootstraps the input column as is, with random bootstrap sampling with replacement
- ‘S1’
Sobol’ first indices, on raw dataset with no bootstrap manipulations (random, with replacement)
Notes
- Compatible with:
all samplers
- Interpretaion:
- ‘balanced’ tells us how important variation in that input feature is on the output feature.
Answers the question: “How important is input feature X on output Y?”
- ‘step’ tells us how important that input feature as active (versus inactive) is on the output feature.
- Answers the question: “What is the impact of input feature X when it is greater than zero/active,
versus inactive, regardless of its actual value?”
- ‘raw’ tells us how much the input feature affects variation in the output feature, specifically in this
dataset and its composition. This score may be skewed or biased toward frequently occurring values in the dataset. (eg. if most values of the input feature lie between 15-20 in your dataset, then the delta score will likely primarily be biased toward influence within that range, rather than its total influence over its entire input range) - Answers the question: “How important is input feature X, and how prominent is its influence specifically
in my dataset?”
- Is a valid interpretation for real-world scenario, if the input data is representative of the
real-world sample and variability without inconsistent or incomplete sampling (rare for real-world data capture)
- ‘notes’ contains flags, warnings, or error messages.
- Will raise a flag if ‘balanced’ and ‘raw’ delta scores differ by >0.1. This suggests that the input dataset is
likely skewed or biased with frequently occuring values in a specific range.
Raises errors if binning cannot be executed, if there aren’t enough/any zeroes/>0s present to do step analysis, etc.
Guides user with results interpretation
Examples
>>> X = latin.sample(problem, 1000) >>> Y = Ishigami.evaluate(X) >>> Si = delta.analyze(problem, X, Y, print_to_console=True)
- Parameters:
problem (dict) – The problem definition
X (numpy.matrix) – A NumPy matrix containing the model inputs
Y (numpy.array) – A NumPy array containing the model outputs
num_resamples (int) – The number of resamples when computing confidence intervals (default 100)
conf_level (float) – The confidence interval level (default 0.95)
print_to_console (bool) – Print results directly to console (default False)
y_resamples (int, optional) – Number of samples to use when resampling (bootstrap) (default None)
method ({"all", "delta", "sobol"}, optional) – Whether to compute “delta”, “sobol” or both (“all”) indices (default “all”)
bootstrap_savedf (str, optional) – User inputs path or filename if they want to save a bootstrap sample to inspect subset composition
bins_specs (dict, optional) –
- Dict with parameter title as key and bin_input as entry.
If bin_input is an int, number of bins. If bin_input is a list, it specifies bin boundaries ( number of bins = len(list)+1 ) Default is 10 bins equally distributed across input feature range
References
- Borgonovo, E. (2007). “A new uncertainty importance measure.”
Reliability Engineering & System Safety, 92(6):771-784, doi:10.1016/j.ress.2006.04.015.
- Plischke, E., E. Borgonovo, and C. L. Smith (2013). “Global
sensitivity measures from given data.” European Journal of Operational Research, 226(3):536-550, doi:10.1016/j.ejor.2012.11.047.
- SALib.analyze.delta.bias_reduced_delta(Y, Ygrid, X, m, num_resamples, conf_level, y_resamples, mode, bin_edges, paramname, min_class_size)[source]#
Plischke et al. 2013 bias reduction technique (eqn 30)
- SALib.analyze.delta.calc_delta(Y, Ygrid, X, m)[source]#
Plischke et al. (2013) delta index estimator (eqn 26) for d_hat.
- SALib.analyze.delta.check_specified_bininfo(bininfo, Xmin, Xmax, paramname)[source]#
Validate user-specified bin edges or number of bins; fallback to default if invalid.
- SALib.analyze.delta.custom_warning_formatter(message, category, filename, lineno, line=None)[source]#
- SALib.analyze.delta.obtain_bins_indices(bin_edges, X, min_class_size)[source]#
Returns indices for all specified bins, ensuring that all bins are at least minimum_class_size