ValidMind for validation 3 — Developing a potential challenger

Learn how to use ValidMind for your end-to-end validation process with our series of four introductory notebooks. In this third notebook, develop a potential challenger and then pass your challenger and its predictions to ValidMind.

A challenger is an alternate record (model) that attempts to outperform the champion, ensuring that the best performing fit-for-purpose record is always considered for deployment. Challengers also help avoid over-reliance on a single record, and allow testing of new features, algorithms, or data sources without disrupting the production lifecycle.

Learn by doing

Our course tailor-made for validators new to ValidMind combines this series of notebooks with more a more in-depth introduction to the ValidMind Platform — Validator Fundamentals

Prerequisites

In order to develop potential challengers with this notebook, you'll need to first have:

Need help with the above steps?

Refer to the first two notebooks in this series:

Setting up

This section should be quite familiar to you — as we performed the same actions in the previous notebook, 2 — Start the validation process.

Initialize the ValidMind Library

As usual, let's first connect up the ValidMind Library to our model we previously registered in the ValidMind Platform:

  1. On the left sidebar that appears for your model, select Getting Started and select Validation from the Document drop-down menu.

  2. Click Copy snippet to clipboard.

  3. Next, load your model identifier credentials from an .env file or replace the placeholder with your own code snippet:

# Make sure the ValidMind Library is installed

%pip install -q validmind

# Load your model identifier credentials from an `.env` file

%load_ext dotenv
%dotenv .env

# Or replace with your code snippet

import validmind as vm

vm.init(
    # api_host="...",
    # api_key="...",
    # api_secret="...",
    # model="...",
    document="validation-report",
)
Note: you may need to restart the kernel to use updated packages.
2026-10-02 20:38:06,305 - INFO(validmind.api_client): 🎉 Connected to ValidMind!
📊 Model: [ValidMind Academy] Model validation (ID: cmalguc9y02ok199q2db381ib)
📁 Document Type: validation_report

Import the sample dataset

Next, we'll load in the sample Bank Customer Churn Prediction dataset used to develop the champion that we will independently preprocess:

# Load the sample dataset
from validmind.datasets.classification import customer_churn as demo_dataset

print(
    f"Loaded demo dataset with: \n\n\t• Target column: '{demo_dataset.target_column}' \n\t• Class labels: {demo_dataset.class_labels}"
)

raw_df = demo_dataset.load_data()
Loaded demo dataset with: 

    • Target column: 'Exited' 
    • Class labels: {'0': 'Did not exit', '1': 'Exited'}

Preprocess the dataset

We’ll apply a simple rebalancing technique to the dataset before continuing:

import pandas as pd

raw_copy_df = raw_df.sample(frac=1)  # Create a copy of the raw dataset

# Create a balanced dataset with the same number of exited and not exited customers
exited_df = raw_copy_df.loc[raw_copy_df["Exited"] == 1]
not_exited_df = raw_copy_df.loc[raw_copy_df["Exited"] == 0].sample(n=exited_df.shape[0])

balanced_raw_df = pd.concat([exited_df, not_exited_df])
balanced_raw_df = balanced_raw_df.sample(frac=1, random_state=42)

Let’s also quickly remove highly correlated features from the dataset using the output from a ValidMind test.

As you know, before we can run tests you’ll need to initialize a ValidMind dataset object with the init_dataset function:

# Register new data and now 'balanced_raw_dataset' is the new dataset object of interest
vm_balanced_raw_dataset = vm.init_dataset(
    dataset=balanced_raw_df,
    input_id="balanced_raw_dataset",
    target_column="Exited",
)

With our balanced dataset initialized, we can then run our test and utilize the output to help us identify the features we want to remove:

# Run HighPearsonCorrelation test with our balanced dataset as input and return a result object
corr_result = vm.tests.run_test(
    test_id="validmind.data_validation.HighPearsonCorrelation",
    params={"max_threshold": 0.3},
    inputs={"dataset": vm_balanced_raw_dataset},
)

❌ High Pearson Correlation

The High Pearson Correlation test evaluates pairwise linear relationships among features to identify highly correlated pairs that may indicate redundancy or multicollinearity. The results table reports the top correlations by coefficient magnitude together with a Pass/Fail designation based on the configured absolute correlation threshold of 0.3. Among the 10 reported pairs, coefficients range from -0.1843 to 0.3426, and only one pair exceeds the threshold. The strongest reported relationship is between Age and Exited with a coefficient of 0.3426, while the remaining listed pairs are below the threshold and marked as passing.

Key insights:

  • One pair exceeds threshold: The (Age, Exited) pair has a Pearson correlation coefficient of 0.3426, which is above the configured threshold of 0.3 and is the only reported failing relationship.
  • All other reported pairs are below 0.3: The remaining nine reported correlations have absolute values between 0.0356 and 0.1843, and all are marked as Pass under the test criterion.
  • Observed relationships are generally weak: Aside from (Age, Exited), the largest absolute correlations in the reported output are (IsActiveMember, Exited) at -0.1843, (Balance, NumOfProducts) at -0.1732, and (Balance, Exited) at 0.1535, indicating relatively limited linear association among the other listed pairs.
  • Top reported correlations include both signs: The output contains both positive and negative coefficients, with the most negative reported value at -0.1843 for (IsActiveMember, Exited) and the most positive at 0.3426 for (Age, Exited).

The reported correlation structure is concentrated in a single threshold breach, with Age and Exited showing the only pairwise relationship above the configured maximum threshold. All other listed feature pairs remain below the threshold, and their coefficients are materially smaller in magnitude. Overall, the test output shows limited high linear correlation within the reported top pairs, with one identified exception.

Parameters:

{
  "max_threshold": 0.3
}
            

Tables

Columns Coefficient Pass/Fail
(Age, Exited) 0.3426 Fail
(IsActiveMember, Exited) -0.1843 Pass
(Balance, NumOfProducts) -0.1732 Pass
(Balance, Exited) 0.1535 Pass
(NumOfProducts, Exited) -0.0633 Pass
(NumOfProducts, IsActiveMember) 0.0490 Pass
(HasCrCard, IsActiveMember) -0.0478 Pass
(Age, NumOfProducts) -0.0441 Pass
(Tenure, IsActiveMember) -0.0374 Pass
(Age, Balance) 0.0356 Pass
# From result object, extract table from `corr_result.tables`
features_df = corr_result.tables[0].data
features_df
Columns Coefficient Pass/Fail
0 (Age, Exited) 0.3426 Fail
1 (IsActiveMember, Exited) -0.1843 Pass
2 (Balance, NumOfProducts) -0.1732 Pass
3 (Balance, Exited) 0.1535 Pass
4 (NumOfProducts, Exited) -0.0633 Pass
5 (NumOfProducts, IsActiveMember) 0.0490 Pass
6 (HasCrCard, IsActiveMember) -0.0478 Pass
7 (Age, NumOfProducts) -0.0441 Pass
8 (Tenure, IsActiveMember) -0.0374 Pass
9 (Age, Balance) 0.0356 Pass
# Extract list of features that failed the test
high_correlation_features = features_df[features_df["Pass/Fail"] == "Fail"]["Columns"].tolist()
high_correlation_features
['(Age, Exited)']
# Extract feature names from the list of strings
high_correlation_features = [feature.split(",")[0].strip("()") for feature in high_correlation_features]
high_correlation_features
['Age']

We can then re-initialize the dataset with a different input_id and the highly correlated features removed and re-run the test for confirmation:

# Remove the highly correlated features from the dataset
balanced_raw_no_age_df = balanced_raw_df.drop(columns=high_correlation_features)

# Re-initialize the dataset object
vm_raw_dataset_preprocessed = vm.init_dataset(
    dataset=balanced_raw_no_age_df,
    input_id="raw_dataset_preprocessed",
    target_column="Exited",
)
# Re-run the test with the reduced feature set
corr_result = vm.tests.run_test(
    test_id="validmind.data_validation.HighPearsonCorrelation",
    params={"max_threshold": 0.3},
    inputs={"dataset": vm_raw_dataset_preprocessed},
)

✅ High Pearson Correlation

The High Pearson Correlation test evaluates pairwise linear relationships between features to identify highly correlated variable pairs that may indicate redundancy or multicollinearity. The reported output lists the top 10 strongest feature-pair correlations, ordered by magnitude, together with their Pearson coefficients and Pass/Fail status against the configured threshold of 0.3. In this result set, all reported coefficients fall between -0.1843 and 0.1535, and each pair is marked as Pass.

Key insights:

  • No pair exceeds threshold: All 10 reported feature-pair correlations remain below the configured absolute threshold of 0.3. Every pair is therefore classified as Pass in the test output.
  • Largest relationship is modest: The strongest observed correlation is between IsActiveMember and Exited at -0.1843. This is the highest absolute correlation in the reported table and remains well below the threshold.
  • Reported correlations are weak overall: The full set of listed coefficients ranges from -0.1843 to 0.1535, with most values closer to zero. This indicates limited linear association among the strongest reported feature pairs.
  • Top relationships include both directions: The reported correlations include both negative and positive values, with examples such as Balance and NumOfProducts at -0.1732 and Balance and Exited at 0.1535. This shows that the strongest observed linear relationships are mixed in direction but low in magnitude.

Overall, the test output shows that the strongest reported pairwise Pearson correlations are all below the configured threshold, with no Fail results in the top 10 relationships. The observed linear associations are weak in magnitude, and the highest absolute coefficient is -0.1843. Based on the reported table, the dataset does not exhibit high pairwise linear correlation among the listed feature combinations.

Parameters:

{
  "max_threshold": 0.3
}
            

Tables

Columns Coefficient Pass/Fail
(IsActiveMember, Exited) -0.1843 Pass
(Balance, NumOfProducts) -0.1732 Pass
(Balance, Exited) 0.1535 Pass
(NumOfProducts, Exited) -0.0633 Pass
(NumOfProducts, IsActiveMember) 0.0490 Pass
(HasCrCard, IsActiveMember) -0.0478 Pass
(Tenure, IsActiveMember) -0.0374 Pass
(Balance, IsActiveMember) -0.0318 Pass
(Tenure, HasCrCard) 0.0253 Pass
(Tenure, EstimatedSalary) 0.0247 Pass

Split the preprocessed dataset

With our raw dataset rebalanced with highly correlated features removed, let's now spilt our dataset into train and test in preparation for model evaluation testing:

# Encode categorical features in the dataset
balanced_raw_no_age_df = pd.get_dummies(
    balanced_raw_no_age_df, columns=["Geography", "Gender"], drop_first=True
)
balanced_raw_no_age_df.head()
CreditScore Tenure Balance NumOfProducts HasCrCard IsActiveMember EstimatedSalary Exited Geography_Germany Geography_Spain Gender_Male
2093 702 2 0.00 2 1 1 145537.32 0 False True True
3311 605 8 125338.80 2 1 0 23970.13 0 True False False
3542 489 7 139395.08 1 0 1 6120.84 0 False False False
1516 625 9 108546.16 3 1 0 133807.77 1 False False True
2145 459 7 110356.42 1 1 0 4969.13 1 True False True
from sklearn.model_selection import train_test_split

# Split the dataset into train and test
train_df, test_df = train_test_split(balanced_raw_no_age_df, test_size=0.20)

X_train = train_df.drop("Exited", axis=1)
y_train = train_df["Exited"]
X_test = test_df.drop("Exited", axis=1)
y_test = test_df["Exited"]
# Initialize the split datasets
vm_train_ds = vm.init_dataset(
    input_id="train_dataset_final",
    dataset=train_df,
    target_column="Exited",
)

vm_test_ds = vm.init_dataset(
    input_id="test_dataset_final",
    dataset=test_df,
    target_column="Exited",
)

Import the champion model

With our raw dataset assessed and preprocessed, let's go ahead and import the champion submitted by the development team in the format of a .pkl file: lr_model_champion.pkl

# Import the champion model
import pickle as pkl

with open("lr_model_champion.pkl", "rb") as f:
    log_reg = pkl.load(f)
/opt/hostedtoolcache/Python/3.11.16/x64/lib/python3.11/site-packages/sklearn/base.py:442: InconsistentVersionWarning: Trying to unpickle estimator LogisticRegression from version 1.3.2 when using version 1.7.2. This might lead to breaking code or invalid results. Use at your own risk. For more info please refer to:
https://scikit-learn.org/stable/model_persistence.html#security-maintainability-limitations
  warnings.warn(

Training a potential challenger model

We're curious how an alternate model compares to our champion, so let's train a challenger as a basis for our testing.

Our champion logistic regression model is a simpler, parametric model that assumes a linear relationship between the independent variables and the log-odds of the outcome. While logistic regression may not capture complex patterns as effectively, it offers a high degree of interpretability and is easier to explain to stakeholders. However, risk is not calculated in isolation from a single factor, but rather in consideration with trade-offs in predictive performance, ease of interpretability, and overall alignment with business objectives.

Random forest classification model

A random forest classification model is an ensemble machine learning algorithm that uses multiple decision trees to classify data. In ensemble learning, multiple models are combined to improve prediction accuracy and robustness.

Random forest classification models generally have higher accuracy because they capture complex, non-linear relationships, but as a result they lack transparency in their predictions.

# Import the Random Forest Classification model
from sklearn.ensemble import RandomForestClassifier

# Create the model instance with 50 decision trees
rf_model = RandomForestClassifier(
    n_estimators=50,
    random_state=42,
)

# Train the model
rf_model.fit(X_train, y_train)
RandomForestClassifier(n_estimators=50, random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Initialize the ValidMind models

In addition to the initialized datasets, you'll also need to initialize a ValidMind model object (vm_model) that can be passed to other functions for analysis and tests on the data for each of our two models.

  • Despite the naming convention, ValidMind model objects can be any type of record you want to test, document, validate, or monitor with the ValidMind Library.
  • From classical statistical and machine learning models, to generative and agentic AI systems and more, the ValidMind model object provides a consistent wrapper around your record so it can be passed as a unified input to any ValidMind test or test suite, with results sent directly to the ValidMind Platform.

Initialize your model objects with vm.init_model():

# Initialize the champion logistic regression model
vm_log_model = vm.init_model(
    log_reg,
    input_id="log_model_champion",
)

# Initialize the challenger random forest classification model
vm_rf_model = vm.init_model(
    rf_model,
    input_id="rf_model",
)

Assign predictions

With our models registered, we'll move on to assigning both the predictive probabilities coming directly from each model's predictions, and the binary prediction after applying the cutoff threshold described in the Compute binary predictions step above.

  • The assign_predictions() method from the Dataset object can link existing predictions to any number of models.
  • This method links the model's class prediction values and probabilities to our vm_train_ds and vm_test_ds datasets.

If no prediction values are passed, the method will compute predictions automatically:

# Champion — Logistic regression model
vm_train_ds.assign_predictions(model=vm_log_model)
vm_test_ds.assign_predictions(model=vm_log_model)

# Challenger — Random forest classification model
vm_train_ds.assign_predictions(model=vm_rf_model)
vm_test_ds.assign_predictions(model=vm_rf_model)
2026-10-02 20:38:20,484 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:38:20,486 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:38:20,486 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:38:20,487 - INFO(validmind.vm_models.dataset.utils): Done running predict()
2026-10-02 20:38:20,489 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:38:20,489 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:38:20,490 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:38:20,490 - INFO(validmind.vm_models.dataset.utils): Done running predict()
2026-10-02 20:38:20,492 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:38:20,506 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:38:20,506 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:38:20,519 - INFO(validmind.vm_models.dataset.utils): Done running predict()
2026-10-02 20:38:20,521 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:38:20,526 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:38:20,526 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:38:20,531 - INFO(validmind.vm_models.dataset.utils): Done running predict()

Running model evaluation tests

With our setup complete, let's run the rest of our validation tests. Since we have already verified the data quality of the dataset used to train our champion, we will now focus on comprehensive performance evaluations of both the champion and challenger models.

Run model performance tests

Let's run some performance tests, beginning with independent testing of our champion logistic regression model, then moving on to our potential challenger model.

Use vm.tests.list_tests() to identify all the model performance tests for classification:


vm.tests.list_tests(tags=["model_performance"], task="classification")
ID Name Description Has Figure Has Table Required Inputs Params Tags Tasks
validmind.model_validation.sklearn.CalibrationCurve Calibration Curve Evaluates the calibration of probability estimates by comparing predicted probabilities against observed... True False ['model', 'dataset'] {'n_bins': {'type': 'int', 'default': 10}} ['sklearn', 'model_performance', 'classification'] ['classification']
validmind.model_validation.sklearn.ClassifierPerformance Classifier Performance Evaluates performance of binary or multiclass classification models using precision, recall, F1-Score, accuracy,... False True ['dataset', 'model'] {'average': {'type': 'str', 'default': 'macro'}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.model_validation.sklearn.ConfusionMatrix Confusion Matrix Evaluates and visually represents the classification ML model's predictive performance using a Confusion Matrix... True False ['dataset', 'model'] {'threshold': {'type': 'float', 'default': 0.5}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']
validmind.model_validation.sklearn.HyperParametersTuning Hyper Parameters Tuning Performs exhaustive grid search over specified parameter ranges to find optimal model configurations... False True ['model', 'dataset'] {'param_grid': {'type': 'dict', 'default': None}, 'scoring': {'type': 'Union', 'default': None}, 'thresholds': {'type': 'Union', 'default': None}, 'fit_params': {'type': 'dict', 'default': None}} ['sklearn', 'model_performance'] ['clustering', 'classification']
validmind.model_validation.sklearn.MinimumAccuracy Minimum Accuracy Checks if the model's prediction accuracy meets or surpasses a specified threshold.... False True ['dataset', 'model'] {'min_threshold': {'type': 'float', 'default': 0.7}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.model_validation.sklearn.MinimumF1Score Minimum F1 Score Assesses if the model's F1 score on the validation set meets a predefined minimum threshold, ensuring balanced... False True ['dataset', 'model'] {'min_threshold': {'type': 'float', 'default': 0.5}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.model_validation.sklearn.MinimumROCAUCScore Minimum ROCAUC Score Validates model by checking if the ROC AUC score meets or surpasses a specified threshold.... False True ['dataset', 'model'] {'min_threshold': {'type': 'float', 'default': 0.5}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.model_validation.sklearn.ModelsPerformanceComparison Models Performance Comparison Evaluates and compares the performance of multiple Machine Learning models using various metrics like accuracy,... False True ['dataset', 'models'] {} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance', 'model_comparison'] ['classification', 'text_classification']
validmind.model_validation.sklearn.PopulationStabilityIndex Population Stability Index Assesses the Population Stability Index (PSI) to quantify the stability of an ML model's predictions across... True True ['datasets', 'model'] {'num_bins': {'type': 'int', 'default': 10}, 'mode': {'type': 'str', 'default': 'fixed'}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.model_validation.sklearn.PrecisionRecallCurve Precision Recall Curve Evaluates the precision-recall trade-off for binary classification models and visualizes the Precision-Recall curve.... True False ['model', 'dataset'] {} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']
validmind.model_validation.sklearn.ROCCurve ROC Curve Evaluates classification model performance by generating and plotting the Receiver Operating Characteristic... True False ['model', 'dataset'] {} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']
validmind.model_validation.sklearn.RegressionErrors Regression Errors Assesses the performance and error distribution of a regression model using various error metrics.... False True ['model', 'dataset'] {} ['sklearn', 'model_performance'] ['regression', 'classification']
validmind.model_validation.sklearn.TrainingTestDegradation Training Test Degradation Tests if model performance degradation between training and test datasets exceeds a predefined threshold.... False True ['datasets', 'model'] {'max_threshold': {'type': 'float', 'default': 0.1}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']
validmind.model_validation.statsmodels.GINITable GINI Table Evaluates classification model performance using AUC, GINI, and KS metrics for training and test datasets.... False True ['dataset', 'model'] {} ['model_performance'] ['classification']
validmind.ongoing_monitoring.CalibrationCurveDrift Calibration Curve Drift Evaluates changes in probability calibration between reference and monitoring datasets.... True True ['datasets', 'model'] {'n_bins': {'type': 'int', 'default': 10}, 'drift_pct_threshold': {'type': 'float', 'default': 20}} ['sklearn', 'binary_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']
validmind.ongoing_monitoring.ClassDiscriminationDrift Class Discrimination Drift Compares classification discrimination metrics between reference and monitoring datasets.... False True ['datasets', 'model'] {'drift_pct_threshold': {'type': '_empty', 'default': 20}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.ongoing_monitoring.ClassificationAccuracyDrift Classification Accuracy Drift Compares classification accuracy metrics between reference and monitoring datasets.... False True ['datasets', 'model'] {'drift_pct_threshold': {'type': '_empty', 'default': 20}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.ongoing_monitoring.ConfusionMatrixDrift Confusion Matrix Drift Compares confusion matrix metrics between reference and monitoring datasets.... False True ['datasets', 'model'] {'drift_pct_threshold': {'type': '_empty', 'default': 20}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_performance'] ['classification', 'text_classification']
validmind.ongoing_monitoring.ROCCurveDrift ROC Curve Drift Compares ROC curves between reference and monitoring datasets.... True False ['datasets', 'model'] {} ['sklearn', 'binary_classification', 'model_performance', 'visualization'] ['classification', 'text_classification']

We'll isolate the specific tests we want to run in mpt:

  • model_validation.sklearn.ClassifierPerformance
  • model_validation.sklearn.ConfusionMatrix
  • model_validation.sklearn.MinimumAccuracy
  • model_validation.sklearn.MinimumF1Score
  • model_validation.sklearn.ROCCurve

As we learned in the previous notebook 2 — Start the model validation process, you can use a custom result_id to tag the individual result with a unique identifier by appending this result_id to the test_id with a : separator. We'll append an identifier for our champion model here:

mpt = [
    "validmind.model_validation.sklearn.ClassifierPerformance:logreg_champion",
    "validmind.model_validation.sklearn.ConfusionMatrix:logreg_champion",
    "validmind.model_validation.sklearn.MinimumAccuracy:logreg_champion",
    "validmind.model_validation.sklearn.MinimumF1Score:logreg_champion",
    "validmind.model_validation.sklearn.ROCCurve:logreg_champion"
]

Evaluate performance of the champion model

Now, let's run and log our batch of model performance tests using our testing dataset (vm_test_ds) for our champion model:

  • The test set serves as a proxy for real-world data, providing an unbiased estimate of model performance since it was not used during training or tuning.
  • The test set also acts as protection against selection bias and model tweaking, giving a final, more unbiased checkpoint.
for test in mpt:
    vm.tests.run_test(
        test,
        inputs={
            "dataset": vm_test_ds, "model" : vm_log_model,
        },
    ).log()

Classifier Performance Logreg Champion

The Classifier Performance test evaluates classification performance using precision, recall, F1-score, accuracy, and ROC AUC. The results are presented for both classes individually and as macro and weighted averages, alongside overall accuracy and ROC AUC. In this run, class-level precision, recall, and F1 values are reported for classes 0 and 1, with aggregate averages clustered near 0.63, overall accuracy of 0.6321, and ROC AUC of 0.6794.

Key insights:

  • Balanced aggregate classification metrics: Weighted average precision, recall, and F1 are 0.6331, 0.6321, and 0.6321, respectively, while macro average precision, recall, and F1 are 0.6326, 0.6326, and 0.6321. The close alignment between weighted and macro averages indicates similar aggregate performance across the two classes.
  • Class-level tradeoff between precision and recall: Class 0 has higher precision than class 1 (0.6497 vs. 0.6156), while class 1 has higher recall than class 0 (0.6508 vs. 0.6145). This shows the model identifies class 1 more completely but with lower precision, whereas predictions for class 0 are more precise but less complete.
  • Nearly identical class-wise F1 performance: F1-scores are 0.6316 for class 0 and 0.6327 for class 1. This indicates that the combined balance of precision and recall is nearly the same for both classes despite the opposing precision-recall pattern.
  • ROC AUC exceeds accuracy: Overall accuracy is 0.6321, while ROC AUC is 0.6794. This indicates stronger rank-order discrimination than is reflected by the single-threshold classification accuracy.

The performance results show a consistent overall classification profile, with aggregate precision, recall, F1, and accuracy all concentrated around 0.63. Class-level results are balanced in terms of F1-score, although the model exhibits an inverse precision-recall pattern between classes 0 and 1. The ROC AUC of 0.6794 is higher than the observed accuracy, indicating that discrimination measured across score thresholds is stronger than the realized performance at the applied classification threshold.

Tables

Precision, Recall, and F1

Class Precision Recall F1
0 0.6497 0.6145 0.6316
1 0.6156 0.6508 0.6327
Weighted Average 0.6331 0.6321 0.6321
Macro Average 0.6326 0.6326 0.6321

Accuracy and ROC AUC

Metric Value
Accuracy 0.6321
ROC AUC 0.6794
2026-10-02 20:38:31,553 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ClassifierPerformance:logreg_champion does not exist in model's document

Confusion Matrix Logreg Champion

The Confusion Matrix test evaluates classification performance by comparing predicted labels with observed labels and displaying the counts of true positives, true negatives, false positives, and false negatives. In this result, the heatmap shows 205 true positives and 204 true negatives, alongside 128 false positives and 110 false negatives. The matrix is nearly balanced across the two true classes, with 315 observations in class 1 and 332 observations in class 0, allowing the error counts to be interpreted against similar class volumes.

Key insights:

  • Correct classifications exceed errors: The model records 409 correct classifications in total, comprising 205 true positives and 204 true negatives, compared with 238 total misclassifications.
  • False positives are slightly higher: False positives total 128, exceeding false negatives at 110. This indicates more class 0 observations were predicted as class 1 than class 1 observations predicted as class 0.
  • Class-level correct counts are balanced: True positive and true negative counts are nearly identical at 205 and 204 respectively, indicating similar volumes of correct classification across the two classes.
  • Observed class distribution is relatively even: The confusion matrix reflects 315 actual class 1 cases and 332 actual class 0 cases, showing no large class imbalance in the evaluated sample.

The result shows that correct predictions outnumber misclassifications, with similar numbers of correctly identified positives and negatives. Error types are present on both sides of the matrix, with a modestly higher count of false positives than false negatives. Overall, the evaluated sample appears relatively balanced by class, and the model’s classification outcomes are distributed across both classes without a pronounced asymmetry in correct predictions.

Figures

ValidMind Figure validmind.model_validation.sklearn.ConfusionMatrix:logreg_champion:3f1b
2026-10-02 20:38:41,330 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ConfusionMatrix:logreg_champion does not exist in model's document

❌ Minimum Accuracy Logreg Champion

The Minimum Accuracy test evaluates whether the model’s prediction accuracy meets or exceeds a predefined threshold on the evaluated dataset. The result table reports an accuracy score of 0.6321 against a threshold of 0.7000, along with the corresponding pass/fail outcome. This result provides a direct comparison between observed classification accuracy and the minimum benchmark defined for the test.

Key insights:

  • Accuracy is below threshold: The observed accuracy score is 0.6321, which is below the minimum threshold of 0.7000 used for this test.
  • Test outcome is fail: The test result is recorded as "Fail," reflecting that the measured accuracy did not meet the specified benchmark.
  • Gap to benchmark is 0.0679: The difference between the observed score and the threshold is 0.0679, quantifying the shortfall relative to the minimum required accuracy level.

The test result shows that the model’s measured accuracy on the evaluated dataset did not reach the configured minimum threshold. The observed score of 0.6321 falls short of the 0.7000 benchmark, and the test therefore returned a fail outcome. Collectively, the result and score difference indicate that overall classification accuracy, as measured in this test, was below the defined acceptance level.

Tables

Score Threshold Pass/Fail
0.6321 0.7 Fail
2026-10-02 20:38:48,797 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.MinimumAccuracy:logreg_champion does not exist in model's document

✅ Minimum F1 Score Logreg Champion

The MinimumF1Score test evaluates whether the model’s F1 score on the validation dataset meets a predefined minimum threshold. The result table reports a validation F1 score of 0.6327, alongside a threshold value of 0.5 and a test outcome of Pass. These values summarize the observed model performance under the threshold-based assessment defined for this test.

Key insights:

  • F1 score exceeds threshold: The recorded F1 score is 0.6327 versus a minimum threshold of 0.5, placing the observed value 0.1327 above the required level.
  • Threshold test passed: The test outcome is reported as Pass, reflecting that the observed validation F1 score satisfied the predefined acceptance criterion.
  • Balanced classification metric documented: The reported result is based on F1 score, which captures the balance between precision and recall within a single validation metric.

The test result shows that the model achieved a validation F1 score above the specified minimum threshold and therefore passed the threshold check. The observed margin above the threshold indicates that the model met the acceptance criterion defined for this validation assessment.

Tables

Score Threshold Pass/Fail
0.6327 0.5 Pass
2026-10-02 20:38:53,506 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.MinimumF1Score:logreg_champion does not exist in model's document

ROC Curve Logreg Champion

The ROC Curve test evaluates classification performance by plotting the trade-off between true positive rate and false positive rate across decision thresholds and by summarizing that relationship with the AUC statistic. For logreg_champion on test_dataset_final, the result is shown as a single binary ROC curve with an annotated AUC of 0.68. The figure also includes the random-classifier reference line at AUC = 0.5, allowing visual comparison of the model’s discrimination relative to a no-skill benchmark.

Key insights:

  • AUC exceeds random baseline: The plotted ROC curve reports an AUC of 0.68, which is above the reference value of 0.5 shown for random classification.
  • Curve remains above no-skill line: Across most of the false positive rate range, the ROC curve lies above the diagonal baseline, indicating higher true positive rates than the random benchmark at comparable threshold settings.
  • Discrimination appears moderate: The ROC curve shows separation from the baseline without approaching the upper-left corner closely, reflecting moderate rather than near-perfect class separation in this test result.

The ROC result indicates that logreg_champion demonstrates measurable discriminative ability on test_dataset_final, with performance above the random benchmark as reflected by an AUC of 0.68. The shape of the curve shows consistent improvement over the no-skill line across thresholds, while the distance from the upper-left corner indicates that discrimination is present but not strong.

Figures

ValidMind Figure validmind.model_validation.sklearn.ROCCurve:logreg_champion:bab9
2026-10-02 20:39:03,612 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ROCCurve:logreg_champion does not exist in model's document
Note the output returned indicating that a test-driven block doesn't currently exist in your documentation for some test IDs.

That's expected, as when we run validations tests the results logged need to be manually added to your report as part of your compliance assessment process within the ValidMind Platform.

Log an artifact

As we can observe from the output above, our champion doesn't pass the MinimumAccuracy based on the default thresholds of the out-of-the-box test, so let's log an artifact (finding) in the ValidMind Platform (Learn more: Add and manage artifacts):

  1. From the Inventory in the ValidMind Platform, go to the model you connected to earlier.

  2. In the left sidebar that appears for your model, click Validation under Documents.

  3. Click on 2.2.2. Model Performance to expand that section.

  4. Under the Model Performance Metrics guideline, click to expand the Artifacts panel.

  5. Click Link Artifact and select Validation Issue as the type of artifact.

  6. Click + Add Validation Issue and enter in the details for your validation issue, for example:

    • Title — Champion Logistic Regression Model Fails Minimum Accuracy Threshold
    • Risk Area — Model Performance
    • Documentation Section — 3.2. Model Evaluation
    • Description — The logistic regression champion model was subjected to a Minimum Accuracy test to determine whether its predictive accuracy meets the predefined performance threshold of 0.7. The model achieved an accuracy score of 0.6136, which falls below the required minimum. As a result, the test produced a Fail outcome.
  7. Click Add Validation Issue to submit the validation issue.

  8. Select the validation issue you just added to link to your validation report.

  9. Click Update Linked Artifacts to insert your validation issue.

  10. Confirm that the validation issue you inserted has been correctly inserted into section 2.2.2. Model Performance of the report.

  11. Click on the validation issue to expand the issue, where you can adjust details such as severity, owner, due date, status, etc. as well as include proposed remediation plans or supporting documentation as attachments.

Evaluate performance of challenger model

We've now conducted similar tests as the development team for our champion, with the aim of verifying their test results.

Next, let's see how our challengers compare. We'll use the same batch of tests here as we did in mpt, but append a different result_id to indicate that these results should be associated with our challenger:

mpt_chall = [
    "validmind.model_validation.sklearn.ClassifierPerformance:champion_vs_challenger",
    "validmind.model_validation.sklearn.ConfusionMatrix:champion_vs_challenger",
    "validmind.model_validation.sklearn.MinimumAccuracy:champion_vs_challenger",
    "validmind.model_validation.sklearn.MinimumF1Score:champion_vs_challenger",
    "validmind.model_validation.sklearn.ROCCurve:champion_vs_challenger"
]

We'll run each test once for each model with the same vm_test_ds dataset to compare them:

for test in mpt_chall:
    vm.tests.run_test(
        test,
        input_grid={
            "dataset": [vm_test_ds], "model" : [vm_log_model,vm_rf_model]
        }
    ).log()

Classifier Performance Champion Vs Challenger

The Classifier Performance test evaluates binary classification models using precision, recall, F1-score, accuracy, and ROC AUC. The results compare the champion model (log_model_champion) and the challenger (rf_model) across class-level metrics, macro and weighted averages, and overall accuracy and ROC AUC. For log_model_champion, weighted-average precision, recall, and F1 are 0.6331, 0.6321, and 0.6321, with accuracy of 0.6321 and ROC AUC of 0.6794. For rf_model, weighted-average precision, recall, and F1 are 0.7256, 0.7249, and 0.7249, with accuracy of 0.7249 and ROC AUC of 0.7935.

Key insights:

  • Challenger outperforms champion overall: rf_model exceeds log_model_champion on all reported aggregate metrics. Accuracy increases from 0.6321 to 0.7249, weighted F1 from 0.6321 to 0.7249, and ROC AUC from 0.6794 to 0.7935.

  • Performance gains are consistent across classes: For class 0, F1 improves from 0.6316 in log_model_champion to 0.7262 in rf_model. For class 1, F1 improves from 0.6327 to 0.7236, indicating that the challenger’s improvement is not concentrated in only one class.

  • Class balance is similar within each model: In log_model_champion, class-level precision and recall remain close across classes, with precision of 0.6497 and 0.6156 and recall of 0.6145 and 0.6508 for classes 0 and 1, respectively. In rf_model, class-level precision is 0.7421 and 0.7082, while recall is 0.7108 and 0.7397, showing similarly balanced class treatment with higher absolute performance.

  • Macro and weighted averages are closely aligned: For both models, macro-average and weighted-average metrics are nearly identical. In log_model_champion, both average F1 values are 0.6321, and in rf_model, both average F1 values are 0.7249, indicating limited divergence between class-averaged and support-weighted results in the reported output.

The test results show a clear separation between the champion and challenger models, with rf_model achieving higher precision, recall, F1-score, accuracy, and ROC AUC throughout the comparison. The improvement is observed at both the aggregate and class levels, and the reported macro and weighted averages remain closely aligned for each model. Overall, the comparison indicates stronger and more uniformly elevated classification performance for rf_model relative to log_model_champion.

Tables

model Class Precision Recall F1
log_model_champion 0 0.6497 0.6145 0.6316
log_model_champion 1 0.6156 0.6508 0.6327
log_model_champion Weighted Average 0.6331 0.6321 0.6321
log_model_champion Macro Average 0.6326 0.6326 0.6321
rf_model 0 0.7421 0.7108 0.7262
rf_model 1 0.7082 0.7397 0.7236
rf_model Weighted Average 0.7256 0.7249 0.7249
rf_model Macro Average 0.7252 0.7253 0.7249
model Metric Value
log_model_champion Accuracy 0.6321
log_model_champion ROC AUC 0.6794
rf_model Accuracy 0.7249
rf_model ROC AUC 0.7935
2026-10-02 20:39:13,404 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ClassifierPerformance:champion_vs_challenger does not exist in model's document

Confusion Matrix Champion Vs Challenger

The Confusion Matrix test evaluates classification performance by comparing predicted labels with observed labels and displaying the resulting counts of true positives, true negatives, false positives, and false negatives. The results are shown for two models: log_model_champion and rf_model. For log_model_champion, the matrix contains 205 true positives, 204 true negatives, 128 false positives, and 110 false negatives. For rf_model, the matrix contains 233 true positives, 236 true negatives, 96 false positives, and 82 false negatives.

Key insights:

  • Challenger shows higher correct classifications: rf_model records more true positives and true negatives than log_model_champion, increasing from 205 to 233 for true positives and from 204 to 236 for true negatives.

  • Challenger reduces both error types: False negatives decrease from 110 in log_model_champion to 82 in rf_model, while false positives decrease from 128 to 96. This indicates fewer missed positive cases and fewer incorrect positive classifications in the challenger result.

  • Performance improvement is consistent across classes: The challenger model improves outcomes for both actual positive and actual negative observations, with gains in correctly classified cases accompanied by reductions in misclassified cases in both matrix rows.

The confusion matrices show that rf_model outperforms log_model_champion on the evaluated sample across all four confusion matrix components. Correct classifications increase for both positive and negative classes, while both false positives and false negatives decline. Taken together, the observed result indicates a uniformly stronger classification outcome for the challenger model in this test.

Figures

ValidMind Figure validmind.model_validation.sklearn.ConfusionMatrix:champion_vs_challenger:2567
ValidMind Figure validmind.model_validation.sklearn.ConfusionMatrix:champion_vs_challenger:9fe3
2026-10-02 20:39:25,130 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ConfusionMatrix:champion_vs_challenger does not exist in model's document

❌ Minimum Accuracy Champion Vs Challenger

The Minimum Accuracy test evaluates whether each model’s prediction accuracy meets or exceeds a predefined threshold. The results table compares the observed accuracy score for each model against a threshold of 0.7 and records the corresponding pass/fail outcome. Two models are included in this run: log_model_champion with an accuracy score of 0.6321 and rf_model with an accuracy score of 0.7249. Their outcomes are shown directly as failed and passed, respectively, based on this threshold comparison.

Key insights:

  • Only one model passed: Of the two evaluated models, rf_model passed the test while log_model_champion failed under the same 0.7 minimum accuracy threshold.
  • Champion model fell below threshold: log_model_champion achieved an accuracy of 0.6321, which is 0.0679 below the required threshold of 0.7.
  • Challenger model exceeded threshold: rf_model achieved an accuracy of 0.7249, exceeding the threshold by 0.0249.
  • Challenger outperformed champion: rf_model recorded a higher accuracy than log_model_champion by 0.0928 points.

The test results show a split outcome across the two evaluated models under a common minimum accuracy requirement. rf_model met the threshold and produced the highest observed accuracy, while log_model_champion remained below the required level. The relative difference between the two models was 0.0928 accuracy points, with the challenger model posting the stronger result in this test run.

Tables

model Score Threshold Pass/Fail
log_model_champion 0.6321 0.7 Fail
rf_model 0.7249 0.7 Pass
2026-10-02 20:39:33,934 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.MinimumAccuracy:champion_vs_challenger does not exist in model's document

✅ Minimum F1 Score Champion Vs Challenger

The MinimumF1Score test evaluates whether each model’s validation-set F1 score meets a predefined minimum threshold for balanced precision and recall performance. The results table reports the validation F1 score, the threshold value, and the resulting pass/fail status for the evaluated models. Two models are shown: log_model_champion with an F1 score of 0.6327 and rf_model with an F1 score of 0.7236, both assessed against a threshold of 0.5 and both marked as passing.

Key insights:

  • Both models passed the threshold: log_model_champion and rf_model each exceeded the minimum F1 threshold of 0.5, resulting in a pass outcome for both models.
  • Random forest scored higher: rf_model achieved an F1 score of 0.7236 versus 0.6327 for log_model_champion, a difference of 0.0909.
  • Positive margin over threshold: log_model_champion exceeded the threshold by 0.1327, while rf_model exceeded it by 0.2236, indicating both models cleared the minimum criterion with measurable margin.

The test results show that both evaluated models satisfied the minimum validation F1 requirement. Among the two, rf_model recorded the higher F1 score and the larger margin above the threshold, while log_model_champion also remained above the required cutoff. Collectively, the results indicate that neither model triggered a threshold breach in this assessment.

Tables

model Score Threshold Pass/Fail
log_model_champion 0.6327 0.5 Pass
rf_model 0.7236 0.5 Pass
2026-10-02 20:39:39,127 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.MinimumF1Score:champion_vs_challenger does not exist in model's document

ROC Curve Champion Vs Challenger

The ROC Curve test evaluates classification performance by plotting the true positive rate against the false positive rate across decision thresholds and summarizing discrimination with the AUC metric. The results show ROC curves for two binary classifiers on test_dataset_final: log_model_champion and rf_model. The plotted curves are shown against the random-classification reference line, with reported AUC values of 0.68 for log_model_champion and 0.79 for rf_model.

Key insights:

  • Random forest has higher AUC: rf_model achieves an AUC of 0.79 versus 0.68 for log_model_champion, indicating stronger class discrimination on the test dataset.
  • Both models exceed random baseline: Both ROC curves remain above the 0.5 random reference line, and both reported AUC values are greater than 0.5.
  • Performance gap is material: The difference in AUC between the two models is 0.11, showing a clear separation in overall ranking performance.
  • Random forest curve rises more sharply: The rf_model ROC curve climbs faster at lower false positive rates than the log_model_champion curve, reflecting higher true positive rates in that region of the plot.

Overall, the ROC results show that both evaluated models demonstrate discriminatory ability on test_dataset_final, with rf_model outperforming log_model_champion on the AUC measure and exhibiting a more favorable ROC profile across thresholds. The observed difference is visible both in the reported summary metric and in the relative position of the plotted curves versus the random baseline.

Figures

ValidMind Figure validmind.model_validation.sklearn.ROCCurve:champion_vs_challenger:2197
ValidMind Figure validmind.model_validation.sklearn.ROCCurve:champion_vs_challenger:2982
2026-10-02 20:39:51,729 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.ROCCurve:champion_vs_challenger does not exist in model's document
Based on the performance metrics, our challenger random forest classification model passes the MinimumAccuracy where our champion did not.

In your validation report, support your recommendation in your validation issue's Proposed Remediation Plan to investigate the usage of our challenger by inserting the performance tests we logged with this notebook into the appropriate section.

Run diagnostic tests

Next, we want to inspect the robustness and stability testing comparison between our champion and challenger.

Use list_tests() to list all available diagnosis tests applicable to classification tasks:

vm.tests.list_tests(tags=["model_diagnosis"], task="classification")
ID Name Description Has Figure Has Table Required Inputs Params Tags Tasks
validmind.model_validation.sklearn.OverfitDiagnosis Overfit Diagnosis Assesses potential overfitting in a model's predictions, identifying regions where performance between training and... True True ['model', 'datasets'] {'metric': {'type': 'str', 'default': None}, 'cut_off_threshold': {'type': 'float', 'default': 0.04}} ['sklearn', 'binary_classification', 'multiclass_classification', 'linear_regression', 'model_diagnosis'] ['classification', 'regression']
validmind.model_validation.sklearn.RobustnessDiagnosis Robustness Diagnosis Assesses the robustness of a machine learning model by evaluating performance decay under noisy conditions.... True True ['datasets', 'model'] {'metric': {'type': 'str', 'default': None}, 'scaling_factor_std_dev_list': {'type': 'List', 'default': [0.1, 0.2, 0.3, 0.4, 0.5]}, 'performance_decay_threshold': {'type': 'float', 'default': 0.05}} ['sklearn', 'model_diagnosis', 'visualization'] ['classification', 'regression']
validmind.model_validation.sklearn.WeakspotsDiagnosis Weakspots Diagnosis Identifies and visualizes weak spots in a machine learning model's performance across various sections of the... True True ['datasets', 'model'] {'features_columns': {'type': 'Optional', 'default': None}, 'metrics': {'type': 'Optional', 'default': None}, 'thresholds': {'type': 'Optional', 'default': None}} ['sklearn', 'binary_classification', 'multiclass_classification', 'model_diagnosis', 'visualization'] ['classification', 'text_classification']

Let’s now assess the models for potential signs of overfitting and identify any sub-segments where performance may inconsistent with the model_validation.sklearn.OverfitDiagnosis test.

Overfitting occurs when a model learns the training data too well, capturing not only the true pattern but noise and random fluctuations resulting in excellent performance on the training dataset but poor generalization to new, unseen data:

  • Since the training dataset (vm_train_ds) was used to fit the model, we use this set to establish a baseline performance for how well the model performs on data it has already seen.
  • The testing dataset (vm_test_ds) was never seen during training, and here simulates real-world generalization, or how well the model performs on new, unseen data.
vm.tests.run_test(
    test_id="validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger",
    input_grid={
        "datasets": [[vm_train_ds,vm_test_ds]],
        "model" : [vm_log_model,vm_rf_model]
    }
).log()

Overfit Diagnosis Champion Vs Challenger

The Overfit Diagnosis test evaluates differences between training and test-set performance across feature slices to identify regions where the AUC gap exceeds the 0.04 threshold. Results are reported for the champion logistic model (log_model_champion) and the challenger random forest (rf_model) by feature-specific bins, with training AUC, test AUC, record counts, and gap values shown for each flagged region. The output indicates a limited set of flagged slices for the logistic model and a substantially broader set of flagged slices for the random forest model. The figures also show that the random forest training AUC is 1.0 in every reported flagged slice, while the logistic model displays smaller and more variable train-test differences.

Key insights:

  • Challenger shows broader overfit regions: The random forest model is flagged across all reported slices for CreditScore, Tenure, NumOfProducts, HasCrCard, IsActiveMember, EstimatedSalary, Geography_Germany, Geography_Spain, and Gender_Male, with additional flagged Balance ranges. In contrast, the logistic champion is flagged only in selected slices of CreditScore, Tenure, Balance, and EstimatedSalary.

  • Random forest reaches perfect training AUC: In every reported flagged slice for rf_model, training AUC equals 1.0. Test AUC in those same slices ranges from 0.0 to 0.9000, producing gaps from 0.1000 up to 1.0000.

  • Largest gap occurs in a high-balance slice: The maximum reported gap is for rf_model on Balance (190710.048, 214548.804], where training AUC is 1.0 and test AUC is 0.0, with only 16 training records and 7 test records. The same Balance slice is also flagged for log_model_champion, with training AUC 0.4615, test AUC 0.0, and gap 0.4615.

  • Champion issues are concentrated, not widespread: For log_model_champion, flagged gaps are limited to 8 slices in total: CreditScore (400.0, 450.0], (550.0, 600.0], (650.0, 700.0], (800.0, 850.0]; Tenure (1.0, 2.0], (6.0, 7.0], (9.0, 10.0]; Balance (190710.048, 214548.804]; and EstimatedSalary (139998.21, 159996.3]. These gaps range from 0.0435 to 0.4615.

  • Champion’s largest non-balance gaps are isolated: Excluding the sparse high-balance slice, the largest champion gaps are CreditScore (400.0, 450.0] at 0.2564 and EstimatedSalary (139998.21, 159996.3] at 0.2082. Other flagged champion slices are closer to the threshold, including CreditScore (550.0, 600.0] at 0.0435 and Tenure (9.0, 10.0] at 0.0470.

  • CreditScore gaps are pervasive for the challenger: For rf_model, all reported CreditScore slices from (400.0, 450.0] through (800.0, 850.0] are flagged, with gaps between 0.1000 and 0.2392. The largest CreditScore gaps occur at (600.0, 650.0] with 0.2392 and (550.0, 600.0] with 0.2369.

  • Tenure and product-count slices are particularly elevated for the challenger: rf_model Tenure gaps range from 0.1040 to 0.3019, with the largest gaps at (9.0, 10.0] = 0.3019 and (4.0, 5.0] = 0.3009. For NumOfProducts, gaps are 0.3285 for (0.997, 1.3], 0.3122 for (1.9, 2.2], and 0.5583 for (2.8, 3.1], the latter corresponding to training AUC 1.0 and test AUC 0.4417.

  • Binary segment gaps are absent for the champion: For log_model_champion, HasCrCard, IsActiveMember, Geography_Germany, Geography_Spain, and Gender_Male do not appear in the flagged table. For rf_model, all of these binary-feature slices are flagged, with gaps ranging from 0.1496 to 0.2371.

Overall, the test results show a marked difference in train-test performance behavior between the two models. The champion logistic model has a small number of flagged regions, with most gaps modestly above threshold aside from the sparsely populated high-balance slice and two larger isolated gaps in low CreditScore and upper-mid EstimatedSalary ranges. The challenger random forest exhibits consistently larger and more widespread gaps across nearly all reported feature segments, including multiple slices with training AUC fixed at 1.0 and materially lower test AUC, indicating substantially stronger train-test separation in the reported regions.

Tables

model Feature Slice Number of Training Records Number of Test Records Training AUC Test AUC Gap
log_model_champion CreditScore (400.0, 450.0] 47 12 0.7564 0.5000 0.2564
log_model_champion CreditScore (550.0, 600.0] 394 86 0.6885 0.6450 0.0435
log_model_champion CreditScore (650.0, 700.0] 484 117 0.6950 0.6230 0.0721
log_model_champion CreditScore (800.0, 850.0] 167 38 0.7200 0.6222 0.0979
log_model_champion Tenure (1.0, 2.0] 261 69 0.6845 0.6120 0.0725
log_model_champion Tenure (6.0, 7.0] 248 65 0.6906 0.6095 0.0811
log_model_champion Tenure (9.0, 10.0] 138 25 0.6768 0.6299 0.0470
log_model_champion Balance (190710.048, 214548.804] 16 7 0.4615 0.0000 0.4615
log_model_champion EstimatedSalary (139998.21, 159996.3] 256 53 0.7595 0.5513 0.2082
rf_model CreditScore (400.0, 450.0] 47 12 1.0000 0.8750 0.1250
rf_model CreditScore (450.0, 500.0] 113 35 1.0000 0.8088 0.1912
rf_model CreditScore (500.0, 550.0] 276 52 1.0000 0.9000 0.1000
rf_model CreditScore (550.0, 600.0] 394 86 1.0000 0.7631 0.2369
rf_model CreditScore (600.0, 650.0] 487 126 1.0000 0.7608 0.2392
rf_model CreditScore (650.0, 700.0] 484 117 1.0000 0.7990 0.2010
rf_model CreditScore (700.0, 750.0] 391 110 1.0000 0.7812 0.2188
rf_model CreditScore (750.0, 800.0] 217 66 1.0000 0.8346 0.1654
rf_model CreditScore (800.0, 850.0] 167 38 1.0000 0.7898 0.2102
rf_model Tenure (-0.01, 1.0] 362 101 1.0000 0.7463 0.2537
rf_model Tenure (1.0, 2.0] 261 69 1.0000 0.8960 0.1040
rf_model Tenure (2.0, 3.0] 277 66 1.0000 0.8329 0.1671
rf_model Tenure (3.0, 4.0] 266 67 1.0000 0.7863 0.2137
rf_model Tenure (4.0, 5.0] 255 68 1.0000 0.6991 0.3009
rf_model Tenure (5.0, 6.0] 263 56 1.0000 0.8302 0.1698
rf_model Tenure (6.0, 7.0] 248 65 1.0000 0.7367 0.2633
rf_model Tenure (7.0, 8.0] 254 64 1.0000 0.8576 0.1424
rf_model Tenure (8.0, 9.0] 261 66 1.0000 0.8412 0.1588
rf_model Tenure (9.0, 10.0] 138 25 1.0000 0.6981 0.3019
rf_model Balance (-238.388, 23838.756] 846 208 1.0000 0.8290 0.1710
rf_model Balance (47677.512, 71516.268] 71 17 1.0000 0.6591 0.3409
rf_model Balance (71516.268, 95355.024] 230 49 1.0000 0.7207 0.2793
rf_model Balance (95355.024, 119193.78] 511 145 1.0000 0.8209 0.1791
rf_model Balance (119193.78, 143032.536] 543 116 1.0000 0.7613 0.2387
rf_model Balance (143032.536, 166871.292] 262 77 1.0000 0.6361 0.3639
rf_model Balance (166871.292, 190710.048] 87 23 1.0000 0.6038 0.3962
rf_model Balance (190710.048, 214548.804] 16 7 1.0000 0.0000 1.0000
rf_model NumOfProducts (0.997, 1.3] 1494 351 1.0000 0.6715 0.3285
rf_model NumOfProducts (1.9, 2.2] 913 242 1.0000 0.6878 0.3122
rf_model NumOfProducts (2.8, 3.1] 145 43 1.0000 0.4417 0.5583
rf_model HasCrCard (-0.001, 0.1] 771 191 1.0000 0.7839 0.2161
rf_model HasCrCard (0.9, 1.0] 1814 456 1.0000 0.7962 0.2038
rf_model IsActiveMember (-0.001, 0.1] 1379 362 1.0000 0.7629 0.2371
rf_model IsActiveMember (0.9, 1.0] 1206 285 1.0000 0.8063 0.1937
rf_model EstimatedSalary (-188.401, 20009.67] 265 63 1.0000 0.7697 0.2303
rf_model EstimatedSalary (20009.67, 40007.76] 255 74 1.0000 0.7485 0.2515
rf_model EstimatedSalary (40007.76, 60005.85] 266 52 1.0000 0.7012 0.2988
rf_model EstimatedSalary (60005.85, 80003.94] 284 72 1.0000 0.8452 0.1548
rf_model EstimatedSalary (80003.94, 100002.03] 227 80 1.0000 0.7497 0.2503
rf_model EstimatedSalary (100002.03, 120000.12] 267 58 1.0000 0.7881 0.2119
rf_model EstimatedSalary (120000.12, 139998.21] 248 64 1.0000 0.8637 0.1363
rf_model EstimatedSalary (139998.21, 159996.3] 256 53 1.0000 0.7236 0.2764
rf_model EstimatedSalary (159996.3, 179994.39] 276 75 1.0000 0.8287 0.1713
rf_model EstimatedSalary (179994.39, 199992.48] 241 56 1.0000 0.8374 0.1626
rf_model Geography_Germany (-0.001, 0.1] 1776 446 1.0000 0.7756 0.2244
rf_model Geography_Germany (0.9, 1.0] 809 201 1.0000 0.7938 0.2062
rf_model Geography_Spain (-0.001, 0.1] 1961 506 1.0000 0.7798 0.2202
rf_model Geography_Spain (0.9, 1.0] 624 141 1.0000 0.8504 0.1496
rf_model Gender_Male (-0.001, 0.1] 1304 314 1.0000 0.7977 0.2023
rf_model Gender_Male (0.9, 1.0] 1281 333 1.0000 0.7846 0.2154

Figures

ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:52d9
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:bb69
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:7653
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:f6fe
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:2ad3
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:7e09
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:a229
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:9c34
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:2b75
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:91fb
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:60fe
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:49ce
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:c632
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:31ae
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:d97e
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:4b3d
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:524b
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:74e9
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:fefe
ValidMind Figure validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger:4114
2026-10-02 20:40:17,455 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.OverfitDiagnosis:champion_vs_challenger does not exist in model's document

Let's also conduct robustness and stability testing of the two models with the model_validation.sklearn.RobustnessDiagnosis test.

Robustness refers to a model's ability to maintain consistent performance, and stability refers to a model's ability to produce consistent outputs over time across different data subsets.

Again, we'll use both the training and testing datasets to establish baseline performance and to simulate real-world generalization:

vm.tests.run_test(
    test_id="validmind.model_validation.sklearn.RobustnessDiagnosis:Champion_vs_LogRegression",
    input_grid={
        "datasets": [[vm_train_ds,vm_test_ds]],
        "model" : [vm_log_model,vm_rf_model]
    },
).log()

❌ Robustness Diagnosis Champion Vs Log Regression

The Robustness Diagnosis test evaluates model resilience by measuring AUC decay after adding Gaussian noise to numerical input features at increasing perturbation sizes. Results are reported for log_model_champion and rf_model on both train_dataset_final and test_dataset_final, with baseline and perturbed AUC values shown alongside calculated performance decay and pass/fail outcomes. The perturbations were applied across the listed input features, and the results show how each model’s train and test performance changes as noise increases from 0.1 to 0.5 standard deviations.

Key insights:

  • Logistic model remains stable: log_model_champion passes at every perturbation level on both datasets. Train AUC declines from 0.6696 at baseline to 0.6579 at 0.5 noise, while test AUC moves from 0.6794 to 0.6602, indicating limited total decay of 0.0117 on train and 0.0192 on test.

  • Random forest shows sharper decay: rf_model exhibits materially larger sensitivity to perturbation, especially on the training dataset. Train AUC decreases from 1.0000 at baseline to 0.7813 at 0.5 noise, corresponding to performance decay increasing to 0.2187.

  • Random forest failures begin at moderate noise: For rf_model, the first failed result occurs on train_dataset_final at perturbation size 0.2 with decay of 0.0621. On test_dataset_final, failures begin at 0.3, where AUC declines to 0.7432 with performance decay of 0.0504.

  • Test performance for the logistic model is consistently close to baseline: On test_dataset_final, log_model_champion shows a small AUC increase at 0.1 noise from 0.6794 to 0.6814, reflected in performance decay of -0.0020, and remains near baseline through 0.2 before declining gradually at higher perturbation levels.

  • Baseline separation differs strongly by model: At baseline, rf_model has higher AUC than log_model_champion on both datasets, with 1.0000 versus 0.6696 on train and 0.7935 versus 0.6794 on test. Under higher perturbation levels, the random forest’s advantage narrows as its AUC declines more rapidly.

Overall, the robustness results show two distinct response patterns under Gaussian noise. log_model_champion maintains relatively stable AUC across the full perturbation range and records only passing outcomes, whereas rf_model starts from higher baseline AUC but experiences substantially larger decay and multiple failed results as perturbation increases. The observed behavior indicates that the logistic model is less sensitive to the tested input noise, while the random forest’s performance deteriorates more quickly under the same perturbation scales.

Tables

model Perturbation Size Dataset Row Count AUC Performance Decay Passed
log_model_champion Baseline (0.0) train_dataset_final 2585 0.6696 0.0000 True
log_model_champion Baseline (0.0) test_dataset_final 647 0.6794 0.0000 True
log_model_champion 0.1 train_dataset_final 2585 0.6684 0.0012 True
log_model_champion 0.1 test_dataset_final 647 0.6814 -0.0020 True
log_model_champion 0.2 train_dataset_final 2585 0.6661 0.0035 True
log_model_champion 0.2 test_dataset_final 647 0.6792 0.0002 True
log_model_champion 0.3 train_dataset_final 2585 0.6615 0.0081 True
log_model_champion 0.3 test_dataset_final 647 0.6738 0.0057 True
log_model_champion 0.4 train_dataset_final 2585 0.6563 0.0132 True
log_model_champion 0.4 test_dataset_final 647 0.6672 0.0123 True
log_model_champion 0.5 train_dataset_final 2585 0.6579 0.0117 True
log_model_champion 0.5 test_dataset_final 647 0.6602 0.0192 True
rf_model Baseline (0.0) train_dataset_final 2585 1.0000 0.0000 True
rf_model Baseline (0.0) test_dataset_final 647 0.7935 0.0000 True
rf_model 0.1 train_dataset_final 2585 0.9831 0.0169 True
rf_model 0.1 test_dataset_final 647 0.7989 -0.0054 True
rf_model 0.2 train_dataset_final 2585 0.9379 0.0621 False
rf_model 0.2 test_dataset_final 647 0.7882 0.0053 True
rf_model 0.3 train_dataset_final 2585 0.8934 0.1066 False
rf_model 0.3 test_dataset_final 647 0.7432 0.0504 False
rf_model 0.4 train_dataset_final 2585 0.8518 0.1482 False
rf_model 0.4 test_dataset_final 647 0.7257 0.0678 False
rf_model 0.5 train_dataset_final 2585 0.7813 0.2187 False
rf_model 0.5 test_dataset_final 647 0.7242 0.0693 False

Figures

ValidMind Figure validmind.model_validation.sklearn.RobustnessDiagnosis:Champion_vs_LogRegression:5ca7
ValidMind Figure validmind.model_validation.sklearn.RobustnessDiagnosis:Champion_vs_LogRegression:eb02
2026-10-02 20:40:37,924 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.RobustnessDiagnosis:Champion_vs_LogRegression does not exist in model's document

Run feature importance tests

We also want to verify the relative influence of different input features on our models' predictions, as well as inspect the differences between our champion and challenger to see if a certain model offers more understandable or logical importance scores for features.

Use list_tests() to identify all the feature importance tests for classification:

# Store the feature importance tests
FI = vm.tests.list_tests(tags=["feature_importance"], task="classification",pretty=False)
FI
['validmind.model_validation.FeaturesAUC',
 'validmind.model_validation.sklearn.PermutationFeatureImportance',
 'validmind.model_validation.sklearn.SHAPGlobalImportance']

We'll only use our testing dataset (vm_test_ds) here, to provide a realistic, unseen sample that mimic future or production data, as the training dataset has already influenced our model during learning:

# Run and log our feature importance tests for both models for the testing dataset
for test in FI:
    vm.tests.run_test(
        "".join((test,':champion_vs_challenger')),
        input_grid={
            "dataset": [vm_test_ds], "model" : [vm_log_model,vm_rf_model]
        },
    ).log()

Features Champion Vs Challenger

The FeaturesAUC test evaluates the univariate discriminatory power of each feature by computing its individual AUC against the binary target. The results are presented as a ranked bar chart for the test_dataset_final dataset, with feature-level AUC values spanning from approximately 0.40 to 0.58. Geography_Germany, Balance, and EstimatedSalary appear at the top of the ranking, while IsActiveMember, NumOfProducts, and Gender_Male appear at the lower end. Most features are concentrated in a relatively narrow midrange between roughly 0.42 and 0.49.

Key insights:

  • Top features remain below 0.60: The highest observed univariate AUC is for Geography_Germany at approximately 0.58, followed closely by Balance at about 0.57 and EstimatedSalary at about 0.53. No individual feature reaches an AUC of 0.60 or higher.

  • Feature ranking is moderately concentrated: After the top three features, AUC values for Geography_Spain, CreditScore, Tenure, and HasCrCard cluster closely around approximately 0.47 to 0.49. This indicates limited separation among much of the feature set on a standalone basis.

  • Lowest discrimination is concentrated near 0.40: IsActiveMember has the lowest AUC at approximately 0.40, with NumOfProducts around 0.42 and Gender_Male around 0.43. These features show the weakest standalone class separation in the displayed ranking.

  • Geography indicators show differentiated signal: Geography_Germany is the strongest individual feature at approximately 0.58, while Geography_Spain is lower at approximately 0.48. The two geography indicators therefore exhibit materially different univariate discrimination levels in the test sample.

The results show that individual feature discrimination is led by Geography_Germany, Balance, and EstimatedSalary, with the strongest observed AUC remaining below 0.60. Most remaining features are tightly grouped in the upper-0.40 range, indicating broadly similar standalone signal strength across a large portion of the feature set. The lower-ranked features, particularly IsActiveMember, NumOfProducts, and Gender_Male, contribute weaker univariate separation in this test view.

Figures

ValidMind Figure validmind.model_validation.FeaturesAUC:champion_vs_challenger:37a7
ValidMind Figure validmind.model_validation.FeaturesAUC:champion_vs_challenger:c29f
2026-10-02 20:40:55,654 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.FeaturesAUC:champion_vs_challenger does not exist in model's document

Permutation Feature Importance Champion Vs Challenger

The Permutation Feature Importance test evaluates feature significance by measuring the change in model performance after each feature is randomly permuted. The results are presented separately for the champion logistic model and the challenger random forest model, with features ordered by their importance scores in each plot. In the champion model, the highest importance values are assigned to IsActiveMember, Geography_Germany, and Gender_Male, while in the challenger model the largest contributions come from NumOfProducts, Balance, and Geography_Germany. Several lower-ranked features in both models have near-zero importance, and the challenger model includes small negative importance values for some variables.

Key insights:

  • Different feature dependence across models: The champion model is most influenced by IsActiveMember (approximately 0.047), followed by Geography_Germany (approximately 0.041) and Gender_Male (approximately 0.030). The challenger model is most influenced by NumOfProducts (approximately 0.176), followed by Balance (approximately 0.067) and Geography_Germany (approximately 0.047), indicating materially different feature ranking between the two models.

  • Challenger shows more concentrated importance: In the challenger model, NumOfProducts has a substantially larger importance score than all other variables, with a visible gap between the first-ranked feature and the rest. The champion model’s top three features are closer in magnitude, indicating a more distributed contribution among leading predictors.

  • Geography_Germany is important in both models: Geography_Germany appears among the top-ranked variables in both plots, ranking second in the champion model and third in the challenger model. This indicates that performance in both models changes meaningfully when this feature is permuted.

  • Near-zero and negative importances appear in lower-ranked features: In the champion model, Geography_Spain and Tenure are close to zero importance, with HasCrCard and NumOfProducts also low relative to the leading variables. In the challenger model, EstimatedSalary and Geography_Spain are near zero, while HasCrCard and Gender_Male show slightly negative importance values.

  • Shared features receive different weights: Several variables appear in both models but with markedly different importance levels. For example, IsActiveMember is the largest contributor in the champion model but is lower-ranked in the challenger model, while NumOfProducts is low-ranked in the champion model but dominant in the challenger model.

Overall, the permutation importance results show that the champion and challenger models rely on different predictor structures. The champion model distributes importance primarily across IsActiveMember, Geography_Germany, and Gender_Male, whereas the challenger model is dominated by NumOfProducts with secondary contribution from Balance and Geography_Germany. Lower-ranked features contribute little to measured performance in both models, and the challenger model includes slight negative importance values for some variables, indicating that permutation of those features does not reduce model performance.

Figures

ValidMind Figure validmind.model_validation.sklearn.PermutationFeatureImportance:champion_vs_challenger:e51e
ValidMind Figure validmind.model_validation.sklearn.PermutationFeatureImportance:champion_vs_challenger:8b1d
2026-10-02 20:41:18,657 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.PermutationFeatureImportance:champion_vs_challenger does not exist in model's document

SHAP Global Importance Champion Vs Challenger

The SHAPGlobalImportance test evaluates global feature importance using SHAP values to show which inputs contribute most strongly to model output. The results include normalized SHAP importance rankings and SHAP summary plots for both the champion logistic model (log_model_champion) and the challenger random forest model (rf_model). In the champion model, the highest-ranked features are IsActiveMember, Gender_Male, Geography_Germany, and Balance, while in the challenger model NumOfProducts is the dominant feature, followed by IsActiveMember, Geography_Germany, and Balance. The summary plots additionally show the direction and spread of SHAP contributions for each feature across observations.

Key insights:

  • Different primary drivers across models: The champion model is led by IsActiveMember, with normalized importance at approximately 100, followed by Gender_Male and Geography_Germany near the mid-70 range. The challenger model instead places NumOfProducts first at approximately 100, with IsActiveMember and Geography_Germany materially lower, near 40 and 35 respectively.

  • Champion importance is more concentrated: In the champion model, the top three features (IsActiveMember, Gender_Male, and Geography_Germany) are clearly separated from the remaining variables, with Balance dropping to roughly the mid-40 range and all others below about 20. This indicates a steeper decline in global importance after the leading features than in the challenger model.

  • Challenger distributes importance more gradually after the top feature: In the challenger model, NumOfProducts stands apart as the largest contributor, but the next set of variables (IsActiveMember, Geography_Germany, Balance, and Gender_Male) declines more progressively, spanning roughly 20 to 40 in normalized importance. The remaining features are lower but still visibly represented.

  • NumOfProducts behavior differs sharply between models: NumOfProducts is a minor feature in the champion model, with normalized importance below 20, but it is the dominant feature in the challenger model. In the challenger summary plot, higher feature values are associated with strongly positive SHAP values extending to roughly 0.5, while lower and mid-range values cluster around negative to modestly positive contributions.

  • IsActiveMember is influential in both models: IsActiveMember appears among the top features in both models and is the leading feature in the champion model. In both summary plots, the two value states form distinct SHAP clusters on opposite sides of zero, indicating a consistent directional separation in model contribution by feature state.

  • Lower-ranked features remain near zero impact: Features such as HasCrCard and Geography_Spain are among the lowest-ranked variables in both models. Their SHAP distributions are concentrated close to zero, indicating comparatively small contributions to model output across most observations.

The SHAP results show that the champion and challenger models rely on different global feature importance structures. The champion model places stronger emphasis on a small set of leading variables, particularly IsActiveMember, Gender_Male, and Geography_Germany, while the challenger model is dominated by NumOfProducts and then distributes remaining importance more gradually across several additional features. Across both models, lower-ranked variables contribute relatively little, with SHAP values concentrated near zero.

Figures

ValidMind Figure validmind.model_validation.sklearn.SHAPGlobalImportance:champion_vs_challenger:633b
ValidMind Figure validmind.model_validation.sklearn.SHAPGlobalImportance:champion_vs_challenger:eee0
ValidMind Figure validmind.model_validation.sklearn.SHAPGlobalImportance:champion_vs_challenger:bc17
ValidMind Figure validmind.model_validation.sklearn.SHAPGlobalImportance:champion_vs_challenger:5f06
2026-10-02 20:41:36,234 - INFO(validmind.vm_models.result.result): Test driven block with result_id validmind.model_validation.sklearn.SHAPGlobalImportance:champion_vs_challenger does not exist in model's document

In summary

In this third notebook, you learned how to:

Next steps

Finalize validation and reporting

Now that you're familiar with the basics of using the ValidMind Library to run and log validation tests, let's learn how to implement some custom tests and wrap up our validation: 4 — Finalize validation and reporting


Copyright © 2023-2026 ValidMind Inc. All rights reserved.
Refer to LICENSE for details.
SPDX-License-Identifier: AGPL-3.0 AND ValidMind Commercial