ADM NB Explained

This notebook shows exactly how all the values in an ADM model report are calculated. It also shows how the propensity is calculated for a particular customer.

We use one of the shipped datamart exports for the example. This is a model very similar to one used in some of the ADM PowerPoint/Excel deep dive examples. You can change this notebook to apply to your own data.

📦 Optional dependencies

This article uses features from the pdstools adm, healthcheck extras. Install with your favorite package manager, e.g. uv pip install "pdstools[adm,healthcheck]".

[2]:
import polars as pl
import numpy as np

import plotly.express as px
from math import log
from great_tables import GT
from pdstools import datasets
from pdstools.utils import cdh_utils
[3]:
model_name = "AutoNew84Months"
predictor_name = "IH.Email.Outbound.Accepted.pyHistoricalOutcomeCount"
channel = "Web"

For the example we pick one particular model over a channel. To explain the ADM model report, we use one of the active predictors as an example. Swap for any other predictor when using different data.

[4]:
dm = datasets.cdh_sample()

model = dm.combined_data.filter(
    (pl.col("Name") == model_name) & (pl.col("Channel") == channel)
)

modelpredictors = (
    dm.combined_data.join(
        model.select(pl.col("ModelID").unique()), on="ModelID", how="inner"
    )
    .filter(pl.col("EntryType") != "Inactive")
    .with_columns(
        Action=pl.concat_str(["Issue", "Group"], separator="/"),
        PredictorName=pl.col("PredictorName").cast(pl.Utf8),
    )
    .collect()
)

predictorbinning = modelpredictors.filter(
    pl.col("PredictorName") == predictor_name
).sort("BinIndex")

Model Overview

The selected model is shown below. Only the currently active predictors are used for the propensity calculation, so only showing those.

[6]:
Overview
Action Sales/AutoLoans
Channel Web
Name AutoNew84Months
Active Predictors Classifier, Customer.Age, Customer.AnnualIncome, Customer.BusinessSegment, Customer.CLV, Customer.CLV_VALUE, Customer.CreditScore, Customer.Date_of_Birth, Customer.Gender, Customer.MaritalStatus, Customer.NetWealth, Customer.NoOfDependents, Customer.Prefix, Customer.RelationshipStartDate, Customer.RiskCode, Customer.WinScore, Customer.pyCountry, IH.Email.Outbound.Accepted.pxLastGroupID, IH.Email.Outbound.Accepted.pxLastOutcomeTime.DaysSince, IH.Email.Outbound.Accepted.pyHistoricalOutcomeCount, IH.Email.Outbound.Churned.pyHistoricalOutcomeCount, IH.Email.Outbound.Loyal.pxLastOutcomeTime.DaysSince, IH.Email.Outbound.Rejected.pyHistoricalOutcomeCount, IH.SMS.Outbound.Accepted.pxLastGroupID, IH.SMS.Outbound.Accepted.pyHistoricalOutcomeCount, IH.SMS.Outbound.Churned.pxLastOutcomeTime.DaysSince, IH.SMS.Outbound.Loyal.pxLastOutcomeTime.DaysSince, IH.SMS.Outbound.Loyal.pyHistoricalOutcomeCount, IH.SMS.Outbound.Rejected.pxLastGroupID, IH.SMS.Outbound.Rejected.pyHistoricalOutcomeCount, IH.Web.Inbound.Accepted.pxLastGroupID, IH.Web.Inbound.Accepted.pyHistoricalOutcomeCount, IH.Web.Inbound.Loyal.pxLastGroupID, IH.Web.Inbound.Loyal.pyHistoricalOutcomeCount, IH.Web.Inbound.Rejected.pxLastGroupID, IH.Web.Inbound.Rejected.pyHistoricalOutcomeCount, Param.ExtGroupCreditcards
Model Performance (AUC) 77.4901

Binning of the selected Predictor

The Model Report in Prediction Studio for this model will have a predictor binning plot like below.

All numbers can be derived from just the number of positives and negatives in each bin that are stored in the ADM Data Mart. The next sections will show exactly how that is done.

Predictor information
Predictor Name IH.Email.Outbound.Accepted.pyHistoricalOutcomeCount
# Responses 1636
# Bins 5
Selected predictor AUC (safe Pega scale) 60.40
[8]:
Binning statistics
Range/Symbol Responses (%) Positives Positives (%) Negatives Negatives (%) Propensity (%) ZRatio Lift
MISSING 13.81% 17 8.25% 209 14.62% 7.52% −2.98 0.60
<2.0 23.29% 30 14.56% 351 24.55% 7.87% −3.69 0.63
[2.0, 2.02> 21.27% 44 21.36% 304 21.26% 12.64% 0.03 1.00
[2.02, 5.02> 33.68% 89 43.20% 462 32.31% 16.15% 2.97 1.28
>=5.02 7.95% 26 12.62% 104 7.27% 20.00% 2.22 1.59
Total 100.00% 206 100.00% 1430 100.00% 12.59% 0.00 1.00

Bin Statistics

Positive and Negative ratios

Internally, ADM only keeps track of the total counts of positive and negative responses in each bin. Everything else is derived from those numbers. The percentages and totals are trivially derived, and the propensity is just the number of positives divided by the total. The numbers calculated here match the numbers from the datamart table exactly.

[9]:
binning_derived = predictorbinning.select(
    pl.col("BinSymbol").alias("Range/Symbol"),
    BinPositives.alias("Positives"),
    BinNegatives.alias("Negatives"),
    ((BinPositives + BinNegatives) / (sumPositives + sumNegatives)).alias(
        "Responses %"
    ),
    (BinPositives / sumPositives).alias("Positives %"),
    (BinNegatives / sumNegatives).alias("Negatives %"),
    (BinPositives / (BinPositives + BinNegatives)).round(4).alias("Propensity"),
)

pcts = ["Responses %", "Positives %", "Negatives %", "Propensity"]
GT(binning_derived).tab_header("Derived binning statistics").tab_style(
    style=style.text(weight="bold"), locations=loc.body(columns="Range/Symbol")
).tab_style(
    style=style.text(color="blue"),
    locations=loc.body(columns=pcts),
).fmt_percent(pcts).tab_options(table_margin_left=0)

[9]:
Derived binning statistics
Range/Symbol Positives Negatives Responses % Positives % Negatives % Propensity
MISSING 17.0 209.0 13.81% 8.25% 14.62% 7.52%
<2.0 30.0 351.0 23.29% 14.56% 24.55% 7.87%
[2.0, 2.02> 44.0 304.0 21.27% 21.36% 21.26% 12.64%
[2.02, 5.02> 89.0 462.0 33.68% 43.20% 32.31% 16.15%
>=5.02 26.0 104.0 7.95% 12.62% 7.27% 20.00%

Lift

Lift is the ratio of the propensity in a particular bin over the average propensity. So a value of 1 is the average, larger than 1 means higher propensity, smaller means lower propensity:

[10]:
positives = pl.col("Positives")
negatives = pl.col("Negatives")
sumPositives = pl.sum("Positives")
sumNegatives = pl.sum("Negatives")
GT(
    binning_derived.select(
        "Range/Symbol",
        "Positives",
        "Negatives",
        (
            (positives / (positives + negatives))
            / (sumPositives / (positives + negatives).sum())
        )
        .round(4)
        .alias("Lift"),
    )
).tab_style(
    style=style.text(weight="bold"), locations=loc.body(columns="Range/Symbol")
).tab_style(
    style=style.text(color="blue"), locations=loc.body(columns=["Lift"])
).tab_options(table_margin_left=0)
[10]:
Range/Symbol Positives Negatives Lift
MISSING 17.0 209.0 0.5974
<2.0 30.0 351.0 0.6253
[2.0, 2.02> 44.0 304.0 1.0041
[2.02, 5.02> 89.0 462.0 1.2828
>=5.02 26.0 104.0 1.5883

Z-Ratio

The Z-Ratio is also a measure of the how the propensity in a bin differs from the average, but takes into account the size of the bin and thus is statistically more relevant. It represents the number of standard deviations from the average, so centers around 0. The wider the spread, the better the predictor is.

\[\frac{\text{posFraction} - \text{negFraction}}{\sqrt{\frac{\text{posFraction}(1 - \text{posFraction})}{\sum \text{positives}} + \frac{\text{negFraction}(1 - \text{negFraction})}{\sum \text{negatives}}}}\]

See the calculation here, which is also included in cdh_utils’ z_ratio().

[11]:
def z_ratio(
    pos_col: pl.Expr = pl.col("BinPositives"), neg_col: pl.Expr = pl.col("BinNegatives")
) -> pl.Expr:
    def get_fracs(pos_col=pl.col("BinPositives"), neg_col=pl.col("BinNegatives")):
        return pos_col / pos_col.sum(), neg_col / neg_col.sum()

    def z_ratio_impl(
        pos_fraction_col=pl.col("posFraction"),
        neg_fraction_col=pl.col("negFraction"),
        positives_col=pl.sum("BinPositives"),
        negatives_col=pl.sum("BinNegatives"),
    ):
        return (
            (pos_fraction_col - neg_fraction_col)
            / (
                (pos_fraction_col * (1 - pos_fraction_col) / positives_col)
                + (neg_fraction_col * (1 - neg_fraction_col) / negatives_col)
            ).sqrt()
        ).alias("ZRatio")

    return z_ratio_impl(*get_fracs(pos_col, neg_col), pos_col.sum(), neg_col.sum())


GT(
    binning_derived.select(
        "Range/Symbol", "Positives", "Negatives", "Positives %", "Negatives %"
    ).with_columns(z_ratio(positives, negatives).round(4))
).tab_style(
    style=style.text(weight="bold"), locations=loc.body(columns="Range/Symbol")
).tab_style(
    style=style.text(color="blue"), locations=loc.body(columns=["ZRatio"])
).fmt_percent(pl.selectors.ends_with("%")).tab_options(table_margin_left=0)
[11]:
Range/Symbol Positives Negatives Positives % Negatives % ZRatio
MISSING 17.0 209.0 8.25% 14.62% -2.9836
<2.0 30.0 351.0 14.56% 24.55% -3.6858
[2.0, 2.02> 44.0 304.0 21.36% 21.26% 0.0329
[2.02, 5.02> 89.0 462.0 43.20% 32.31% 2.9721
>=5.02 26.0 104.0 12.62% 7.27% 2.2161

Predictor AUC

The predictor AUC is the univariate performance of the selected predictor against the outcome. It answers: if we rank cases using only this predictor’s bins, how well does that single predictor separate positives from negatives?

This is different from the model AUC below. The model AUC is calculated from the Classifier bins after all active predictor contributions have been combined into the final Naive Bayes score.

This predictor-level value can be derived from the positives and negatives in the predictor bins. Because this is the first AUC calculation in the notebook, we also calculate its confidence interval from those same binned counts. The helper prints the raw 0-to-1 AUC first, then the safe Pega 50-to-100 value and its safe confidence interval.

The AUC point estimate and confidence interval are implemented in cdh_utils: cdh_utils.auc_from_bincounts() and cdh_utils.auc_ci_from_bincounts().

[12]:
predictor_auc_ci = cdh_utils.auc_ci_from_bincounts(
    pos=binning_derived.get_column("Positives"),
    neg=binning_derived.get_column("Negatives"),
    probs=binning_derived.get_column("Propensity"),
    confidence_level=0.95,
)
predictor_performance = predictorbinning.get_column("PredictorPerformance")[0]

if predictor_auc_ci["ci_available"]:
    safe_predictor_ci = f"[{100 * predictor_auc_ci['safe_ci_lower']:.2f}, {100 * predictor_auc_ci['safe_ci_upper']:.2f}]"
else:
    safe_predictor_ci = f"unavailable ({predictor_auc_ci['ci_reason']})"

print(
    f"Selected predictor AUC from predictor bins: {predictor_auc_ci['auc']:.6f} | "
    f"safe Pega scale: {100 * predictor_auc_ci['auc']:.2f} | "
    f"95% safe CI: {safe_predictor_ci}"
)
print(
    f"Predictor Performance column: {predictor_performance:.6f} | "
    f"safe Pega scale: {100 * predictor_performance:.2f}"
)

Selected predictor AUC from predictor bins: 0.603987 | safe Pega scale: 60.40 | 95% safe CI: [56.51, 64.29]
Predictor Performance column: 0.603987 | safe Pega scale: 60.40

Model AUC

The model’s overall AUC — shown as Model Performance (AUC) in the model overview above — is derived from the Classifier row, not from the selected predictor and not as an average of the predictor AUCs.

The Classifier is the PAVA-calibrated score distribution: after all predictor log-odds contributions are summed into the final score, the Classifier records the observed positive and negative response counts per score interval. Because those bins carry the same \(pos\) / \(neg\) structure as any predictor bin, the model AUC and confidence interval can be calculated from the Classifier bins with the same auc_ci_from_bincounts() helper used above.

\[ \begin{align}\begin{aligned}AUC_{\text{model}} = \text{auc\_from\_bincounts}(\,pos_{\text{classifier}},\; neg_{\text{classifier}}\,)\\**Portfolio AUC:** Each Naïve Bayes model covers one action / channel / direction combination. The natural portfolio-level summary is a response-count-weighted average of the individual model AUCs — which is what ``cdh_utils.weighted_performance_polars()`` computes across a set of models. The same principle applies to Adaptive Gradient Boosting models, where each action has its own independently validated AUC.\end{aligned}\end{align} \]
[13]:
classifier_bins = (
    modelpredictors
    .filter(pl.col("EntryType") == "Classifier")
    .sort("BinIndex")
)

classifier_pos = classifier_bins.get_column("BinPositives").to_numpy()
classifier_neg = classifier_bins.get_column("BinNegatives").to_numpy()
classifier_scores = classifier_bins.get_column("BinIndex").to_numpy()
classifier_order = np.argsort(classifier_scores)[::-1]

FPR = np.cumsum(classifier_neg[classifier_order]) / np.sum(classifier_neg[classifier_order])
TPR = np.cumsum(classifier_pos[classifier_order]) / np.sum(classifier_pos[classifier_order])
TPR = np.insert(TPR, 0, 0, axis=0)
FPR = np.insert(FPR, 0, 0, axis=0)

model_auc_ci = cdh_utils.auc_ci_from_bincounts(
    pos=classifier_bins.get_column("BinPositives"),
    neg=classifier_bins.get_column("BinNegatives"),
    probs=classifier_bins.get_column("BinIndex"),
    confidence_level=0.95,
)
model_performance = modelpredictors.select(pl.col("Performance").unique()).item()

fig = px.line(
    x=FPR,
    y=TPR,
    labels=dict(x="False positive rate", y="Sensitivity"),
    title=f"Overall model AUC = {model_auc_ci['auc']:.3f}",
    width=525,
    height=525,
    range_x=[0, 1],
    range_y=[0, 1],
    template="none",
)
fig.update_traces(
    fill="tozeroy",
    fillcolor="rgba(0, 31, 95, 0.15)",
    line=dict(color="#001F5F", width=3),
    marker=dict(size=7),
)
fig.add_shape(
    type="line",
    line=dict(color="#63666F", dash="dash"),
    x0=0,
    x1=1,
    y0=0,
    y1=1,
    layer="above",
)
fig.show()

if model_auc_ci["ci_available"]:
    safe_model_ci = f"[{100 * model_auc_ci['safe_ci_lower']:.2f}, {100 * model_auc_ci['safe_ci_upper']:.2f}]"
else:
    safe_model_ci = f"unavailable ({model_auc_ci['ci_reason']})"

print(
    f"Overall model AUC from Classifier bins: {model_auc_ci['auc']:.6f} | "
    f"safe Pega scale: {100 * model_auc_ci['auc']:.2f} | "
    f"95% safe CI: {safe_model_ci}"
)
print(
    f"Model Performance column: {model_performance:.6f} | "
    f"safe Pega scale: {100 * model_performance:.2f}"
)

Overall model AUC from Classifier bins: 0.778734 | safe Pega scale: 77.87 | 95% safe CI: [74.22, 81.52]
Model Performance column: 0.774901 | safe Pega scale: 77.49

Predictor Contributions

Naive Bayes and Log Odds

The basis for the Naive Bayes algorithm is Bayes’ Theorem:

\[\Pr(C_k \mid x) = \frac{\Pr(x \mid C_k)\Pr(C_k)}{\Pr(x)}\]

with \(C_k\) the outcome and \(x\) the customer. Bayes’ theorem turns the question “what’s the probability to accept this action given a customer” around to “what’s the probability of this customer given an action”. With the independence assumption, and after applying a log odds transformation we get a log odds score that can be calculated efficiently and in a numerically stable manner:

\[\text{log odds score} = \sum_{p \in \text{active predictors}} \log(\Pr(x_p \mid \text{Positive})) + \log(p_{\text{positive}}) - \sum_p \log(\Pr(x_p \mid \text{Negative})) - \log(p_{\text{negative}})\]

note that the prior can be written as:

\[\log(p_{\text{positive}}) - \log(p_{\text{negative}}) = \log\left(\frac{\text{TotalPositives}}{\text{Total}}\right) - \log\left(\frac{\text{TotalNegatives}}{\text{Total}}\right) = \log(\text{TotalPositives}) - \log(\text{TotalNegatives})\]

Predictor Contribution

The contribution (conditional log odds) of an active predictor \(p\) for bin \(i\) with the number of positive and negative responses in \(\text{Positives}_i\) and \(\text{Negatives}_i\) is calculated as (note the “Laplace smoothing” to avoid log 0 issues):

\[\text{contribution}_p = \log\left(\text{Positives}_i + \frac{1}{\text{nBins}}\right) - \log\left(\text{Negatives}_i + \frac{1}{\text{nBins}}\right) - \log\left(1 + \sum_{i = 1}^{\text{nBins}} \text{Positives}_i\right) + \log\left(1 + \sum_i \text{Negatives}_i\right)\]
[14]:
N = binning_derived.shape[0]
GT(
    binning_derived.with_columns(
        LogOdds=(pl.col("Positives %") / pl.col("Negatives %")).log().round(5),
        ModifiedLogOdds=(
            ((positives + 1 / N).log() - (positives.sum() + 1).log())
            - ((negatives + 1 / N).log() - (negatives.sum() + 1).log())
        ).round(5),
    ).drop("Responses %", "Propensity")
).tab_style(
    style=style.text(weight="bold"), locations=loc.body(columns="Range/Symbol")
).tab_style(
    style=style.text(color="blue"),
    locations=loc.body(columns=["LogOdds", "ModifiedLogOdds"]),
).fmt_percent(pl.selectors.ends_with("%")).tab_options(table_margin_left=0)
[14]:
Range/Symbol Positives Negatives Positives % Negatives % LogOdds ModifiedLogOdds
MISSING 17.0 209.0 8.25% 14.62% -0.57157 -0.56497
<2.0 30.0 351.0 14.56% 24.55% -0.52204 -0.5201
[2.0, 2.02> 44.0 304.0 21.36% 21.26% 0.00472 0.00445
[2.02, 5.02> 89.0 462.0 43.20% 32.31% 0.29063 0.28829
>=5.02 26.0 104.0 12.62% 7.27% 0.55126 0.55286

Propensity mapping

Log odds contribution for all the predictors

\[\text{score} = \frac{\log(1 + \text{TotalPositives}) - \log(1 + \text{TotalNegatives}) + \sum_p \text{contribution}_p}{1 + \text{nActivePredictors}}\]

Here, \(\text{TotalPositives}\) and \(\text{TotalNegatives}\) are the total number of positive and negative responses in the model, and \(\text{contribution}_p\) is the conditional log odds contribution of each active predictor.

[15]:
PredictorName Value Bin Positives Negatives LogOdds
Customer.Age 34.56 4 9.0 198.0 -1.1459
Customer.AnnualIncome -24043.049 1 74.0 1166.0 -0.8197
Customer.BusinessSegment middleSegmentPlus 1 96.0 970.0 -0.3764
Customer.CLV NON-MISSING 1 111.0 570.0 0.3009
Customer.CLV_VALUE 1345.52 4 31.0 297.0 -0.3227
Customer.CreditScore 518.92 3 33.0 205.0 0.1105
Customer.Date_of_Birth 18773.504 5 28.0 152.0 0.2446
Customer.Gender U 1 52.0 481.0 -0.2855
Customer.MaritalStatus No Resp+ 1 67.0 745.0 -0.4708
Customer.NetWealth 17992.398 4 51.0 179.0 0.6796
Customer.NoOfDependents 0.0 1 111.0 850.0 -0.0997
Customer.Prefix Mrs. 1 64.0 552.0 -0.2167
Customer.RelationshipStartDate 1426.4596 4 16.0 117.0 -0.0502
Customer.RiskCode R4 1 36.0 329.0 -0.2709
Customer.WinScore 66.600006 4 39.0 102.0 0.9738
Customer.pyCountry USA 1 99.0 776.0 -0.1227
IH.Email.Outbound.Accepted.pxLastGroupID HomeLoans 3 25.0 218.0 -0.2272
IH.Email.Outbound.Accepted.pxLastOutcomeTime.DaysSince -55.88436 2 145.0 881.0 0.1305
IH.Email.Outbound.Accepted.pyHistoricalOutcomeCount 1.5 2 30.0 351.0 -0.5201
IH.Email.Outbound.Churned.pyHistoricalOutcomeCount None 1 143.0 898.0 0.1015
IH.Email.Outbound.Loyal.pxLastOutcomeTime.DaysSince None 1 129.0 1071.0 -0.1788
IH.Email.Outbound.Rejected.pyHistoricalOutcomeCount 83.16 3 24.0 218.0 -0.2678
IH.SMS.Outbound.Accepted.pxLastGroupID Account 4 45.0 316.0 -0.0133
IH.SMS.Outbound.Accepted.pyHistoricalOutcomeCount 9.02 4 6.0 96.0 -0.822
IH.SMS.Outbound.Churned.pxLastOutcomeTime.DaysSince -20.5537 2 9.0 27.0 0.8492
IH.SMS.Outbound.Loyal.pxLastOutcomeTime.DaysSince None 1 165.0 1240.0 -0.0797
IH.SMS.Outbound.Loyal.pyHistoricalOutcomeCount None 1 165.0 1240.0 -0.0797
IH.SMS.Outbound.Rejected.pxLastGroupID Account 2 47.0 357.0 -0.0905
IH.SMS.Outbound.Rejected.pyHistoricalOutcomeCount 102.72 4 12.0 117.0 -0.3356
IH.Web.Inbound.Accepted.pxLastGroupID DepositAccounts 3 53.0 397.0 -0.0779
IH.Web.Inbound.Accepted.pyHistoricalOutcomeCount 11.04 5 25.0 164.0 0.0558
IH.Web.Inbound.Loyal.pxLastGroupID MISSING 1 100.0 857.0 -0.2119
IH.Web.Inbound.Loyal.pyHistoricalOutcomeCount 4.52 3 30.0 212.0 -0.0172
IH.Web.Inbound.Rejected.pxLastGroupID Account 2 81.0 546.0 0.0279
IH.Web.Inbound.Rejected.pyHistoricalOutcomeCount 111.08 4 35.0 306.0 -0.2317
Param.ExtGroupCreditcards NON-MISSING 1 136.0 721.0 0.2684
Final Score None None None None -0.1493

From Log Odds Score to Final Propensity

The log odds score is mapped to a propensity using a score distribution, which is built using the PAV(A) (Pool Adjacent Violators Algorithm), a form of isotonic regression that ensures the mapping is monotonically increasing.

For each bin (score interval) the final propensity is calculated from the observed positive and negative counts:

\[Final\ Propensity = \frac{0.5 + BinPositives}{1 + BinPositives + BinNegatives}\]

The addition of 0.5 is for Laplace smoothing. An empty model (no responses yet) returns a neutral propensity of 0.5, rather than 0 or undefined. This quickly converges to the observed success rate as more responses are received.

The transformation is a calibrated, isotonic mapping from log odds to propensity, grounded in the actual observed response rates of the model.

[16]:
Index Bin Positives Negatives Cum. Total (%) Propensity (%) Final Propensity (%) Cum Positives (%) ZRatio Lift(%)
1 <-0.21 17.0 443.0 28.12% 3.70% 3.80% 8.25% -9.9945 1.96%
2 [-0.21, -0.185> 8.0 133.0 36.74% 5.67% 5.99% 12.14% -3.4954 3.00%
3 [-0.185, -0.175> 3.0 48.0 39.85% 5.88% 6.73% 13.59% -1.9775 3.11%
4 [-0.175, -0.105> 28.0 370.0 64.18% 7.04% 7.14% 27.18% -4.6281 3.72%
5 [-0.105, -0.095> 4.0 51.0 67.54% 7.27% 8.04% 29.13% -1.5054 3.85%
6 [-0.095, -0.09> 2.0 19.0 68.83% 9.52% 11.36% 30.10% -0.4788 5.04%
7 [-0.09, -0.065> 9.0 77.0 74.08% 10.47% 10.92% 34.47% -0.6578 5.54%
8 [-0.065, -0.02> 30.0 154.0 85.33% 16.30% 16.49% 49.03% 1.4644 8.63%
9 [-0.02, 0.03> 37.0 65.0 91.56% 36.27% 36.41% 66.99% 4.913 19.21%
10 [0.03, 0.06> 20.0 29.0 94.56% 40.82% 41.00% 76.70% 3.664 21.61%
11 [0.06, 0.12> 30.0 33.0 98.41% 47.62% 47.66% 91.26% 4.9229 25.21%
12 [0.12, 0.125> 2.0 2.0 98.66% 50.00% 50.00% 92.23% 1.2039 26.47%
13 [0.125, 0.13> 4.0 2.0 99.02% 66.67% 64.29% 94.17% 1.8644 35.30%
14 [0.13, 0.995> 8.0 3.0 99.69% 72.73% 70.83% 98.06% 2.7182 38.51%
15 >=0.995 4.0 1.0 100.00% 80.00% 75.00% 100.00% 1.9418 42.36%

The chart below visualizes the score distribution from the table above. The score bin that contains the calculated final score for the example customer is highlighted.

Feature Importance

Feature importance in Naive Bayes models represents how strongly a predictor differentiates between positive and negative outcomes. It measures the average magnitude of log odds contributions across the predictor’s bins.

A predictor with high importance has bins with strong log odds values (far from zero), indicating strong predictive power. A predictor with low importance has bins with weak log odds values (close to zero), indicating weak predictive power.

Formula

For a predictor with bins \(i = 1, \ldots, n\):

Step 1: Log odds per bin (with Laplace smoothing \(\frac{1}{n}\)):

\[\text{LogOdds}_i = \log\left(\text{Pos}_i + \frac{1}{n}\right) - \log\left(\sum_i \text{Pos}_i + 1\right) - \left[\log\left(\text{Neg}_i + \frac{1}{n}\right) - \log\left(\sum_i \text{Neg}_i + 1\right)\right]\]

Step 2: Absolute log odds per bin:

\[|\text{LogOdds}_i|\]

Step 3: Feature importance (weighted average of absolute log odds):

\[\text{Importance} = \frac{\sum_i |\text{LogOdds}_i| \times \text{Responses}_i}{\sum_i \text{Responses}_i}\]

Step 4: Optional scaling (default):

\[\text{Scaled Importance} = \frac{\text{Importance}}{\max(\text{Importance})} \times 100\]

Interpretation

  • Higher values: Predictor has strong log odds values (strong predictive power)

  • Lower values: Predictor has weak log odds values (weak predictive power)

  • Zero: All bins have zero log odds (predictor adds no information)

The absolute log odds captures the magnitude of each bin’s contribution to the model score, weighted by how many responses are in that bin. This matches the Pega platform implementation.

[18]:
# Calculate feature importance for all active predictors
predictor_importance = (
    modelpredictors.filter(pl.col("PredictorName") != "Classifier")
    .with_columns(
        # Calculate both unscaled and scaled importance
        cdh_utils.feature_importance(scaled=False).alias("Importance"),
        cdh_utils.feature_importance(scaled=True).alias("ScaledImportance"),
    )
    .group_by("PredictorName", "ModelID", "PredictorCategory")
    .agg(pl.first("Importance"), pl.first("ScaledImportance"))
    .sort("ScaledImportance", descending=True)
)

GT(predictor_importance.drop("ModelID")).tab_header("Feature Importance by Predictor").fmt_number(
    "Importance", decimals=4
).fmt_number("ScaledImportance", decimals=2)

[18]:
Feature Importance by Predictor
PredictorName PredictorCategory Importance ScaledImportance
Customer.AnnualIncome Customer 0.9043 100.00
Customer.NetWealth Customer 0.8399 92.88
Customer.WinScore Customer 0.6859 75.85
Customer.CreditScore Customer 0.6494 71.82
Customer.RelationshipStartDate Customer 0.5316 58.79
Customer.Age Customer 0.5054 55.89
Customer.CLV_VALUE Customer 0.4982 55.09
IH.Web.Inbound.Accepted.pyHistoricalOutcomeCount IH 0.4781 52.88
Customer.MaritalStatus Customer 0.4281 47.35
IH.SMS.Outbound.Accepted.pyHistoricalOutcomeCount IH 0.4277 47.29
Customer.BusinessSegment Customer 0.4219 46.66
IH.Email.Outbound.Accepted.pyHistoricalOutcomeCount IH 0.3411 37.73
Param.ExtGroupCreditcards Param 0.3194 35.32
IH.Web.Inbound.Rejected.pyHistoricalOutcomeCount IH 0.3132 34.64
IH.Email.Outbound.Rejected.pyHistoricalOutcomeCount IH 0.3114 34.44
Customer.CLV Customer 0.2799 30.96
IH.Email.Outbound.Accepted.pxLastGroupID IH 0.2794 30.90
IH.SMS.Outbound.Rejected.pyHistoricalOutcomeCount IH 0.2398 26.52
Customer.Date_of_Birth Customer 0.2379 26.31
IH.Email.Outbound.Loyal.pxLastOutcomeTime.DaysSince IH 0.2332 25.79
Customer.Gender Customer 0.2291 25.33
IH.Web.Inbound.Loyal.pxLastGroupID IH 0.2273 25.14
IH.Web.Inbound.Loyal.pyHistoricalOutcomeCount IH 0.2249 24.87
IH.SMS.Outbound.Accepted.pxLastGroupID IH 0.2133 23.59
Customer.RiskCode Customer 0.2130 23.56
IH.Email.Outbound.Accepted.pxLastOutcomeTime.DaysSince IH 0.2045 22.61
IH.Web.Inbound.Accepted.pxLastGroupID IH 0.1761 19.47
Customer.Prefix Customer 0.1520 16.81
IH.SMS.Outbound.Loyal.pxLastOutcomeTime.DaysSince IH 0.1466 16.22
IH.Email.Outbound.Churned.pyHistoricalOutcomeCount IH 0.1393 15.41
IH.SMS.Outbound.Loyal.pyHistoricalOutcomeCount IH 0.1256 13.89
Customer.pyCountry Customer 0.1250 13.82
IH.Web.Inbound.Rejected.pxLastGroupID IH 0.1154 12.77
Customer.NoOfDependents Customer 0.1134 12.54
IH.SMS.Outbound.Churned.pxLastOutcomeTime.DaysSince IH 0.1052 11.63
IH.SMS.Outbound.Rejected.pxLastGroupID IH 0.1007 11.14