Follow the complete path from ambient air and fuel to gas-turbine power,
recovered exhaust heat, steam-turbine power, and the final net electrical
output used in the Combined Cycle Power Plant dataset.
Gas turbineSteam turbineHRSGCondenserNet output
Purpose of Tutorial 02.
This article explains the physical system behind the dataset. It does not yet
analyze feature ranges, correlations, missing values, or Python code. Those
topics belong to Tutorials 03 and 04.
1. What Does “Combined Cycle” Mean?
A combined cycle power plant joins two thermodynamic power-conversion systems.
The first is a gas-turbine cycle, commonly described by the
Brayton-cycle framework. The second is a steam-turbine cycle,
commonly described by the Rankine-cycle framework. The two systems are connected
by a heat-recovery steam generator, usually abbreviated
HRSG.
Fuel is burned only in the gas-turbine combustor in a basic unfired configuration.
The gas turbine produces electricity, but the exhaust leaving it remains very
hot. Instead of rejecting that thermal energy directly to the atmosphere, the
plant transfers much of it to water in the HRSG. The resulting steam drives a
second turbine and generator.
This reuse of exhaust heat is the central reason combined-cycle systems can
convert fuel to electricity more efficiently than a comparable simple-cycle gas
turbine. The U.S. Department of Energy describes the HRSG as a boiler that
captures gas-turbine exhaust heat to produce high-pressure steam for additional
power generation.[1]
Combined cycle is not the same as combined heat and power.
Combined cycle combines gas- and steam-turbine power cycles. Combined heat and
power, or cogeneration, intentionally supplies useful heat to an external
process or building in addition to generating electricity. A plant can use
both concepts, but the terms are not interchangeable.
2. Complete Energy Flow Through the Plant
Figure 1. Main energy and working-fluid paths. The upper path
represents the gas-turbine system; the lower path recovers exhaust heat and
circulates water and steam through the steam-turbine system.
The full process can be understood as ten linked steps.
1
Air enters the compressor. Ambient air is drawn into the
gas turbine through the intake and filtration system.
2
The compressor raises air pressure. Compressing the air
requires a substantial share of the turbine's mechanical work.
3
Fuel is injected and burned. Fuel mixes with compressed
air in the combustion system, creating a high-temperature, high-pressure gas
stream.
4
Hot gas expands through the gas turbine. The expanding gas
rotates turbine blades and the main shaft.
5
The first generator produces electricity. Shaft power
drives the gas-turbine generator.
6
Hot exhaust enters the HRSG. The exhaust still carries
useful thermal energy after leaving the gas turbine.
7
The HRSG produces steam. Exhaust heat preheats feedwater,
evaporates it, and can superheat the resulting steam.
8
Steam expands through the steam turbine. The steam's
enthalpy is converted into shaft work.
9
The second generator produces electricity. The steam
turbine contributes additional power without requiring a second primary fuel
stream in the basic arrangement.
10
Steam is condensed and recirculated. The condenser turns
exhaust steam back into water, and a feedwater pump returns it to the HRSG.
3. The Gas-Turbine System
The gas turbine is the first power-producing system and the source of heat for
the steam cycle. The U.S. Department of Energy separates a gas turbine into three
principal sections: the compressor, the combustion system, and the turbine
section.[1]
Figure 2. Gas-turbine sequence. Expansion creates shaft
power, while the exhaust remains hot enough to supply the HRSG.
3.1 Compressor
The compressor draws in ambient air and raises its pressure before combustion.
This is not a minor auxiliary process: the compressor consumes a large fraction
of the turbine's gross mechanical work. The useful gas-turbine shaft output is
therefore the difference between turbine work and compressor work.
ẆGT,net = Ẇturbine − Ẇcompressor
Air density and compressor inlet conditions are consequently important. When
the inlet air is less dense, the machine may process less air mass for a similar
volumetric flow, which can reduce gas-turbine power. Compressor efficiency,
pressure ratio, inlet pressure loss, fouling, and control settings also influence
performance.
3.2 Combustion system
Fuel is injected into compressed air and burned. The combustion system must
provide a stable high-temperature gas stream while controlling emissions,
pressure loss, flame stability, and component temperatures. Turbine inlet
temperature is a major performance parameter, but it is constrained by materials,
cooling technology, emissions requirements, and equipment life.
3.3 Turbine and generator
The hot gas expands across stationary and rotating blade rows. The rotating
blades drive the shaft, which simultaneously powers the compressor and the
electrical generator. The gas leaving the turbine has a lower pressure and
temperature than at the inlet, yet it still contains enough heat to support the
second cycle.
4. The Heat-Recovery Steam Generator
The HRSG is the thermal bridge between the gas turbine and steam turbine. It does
not convert heat directly into electricity. Instead, it transfers heat from
gas-turbine exhaust to the water–steam circuit.
Figure 3. Simplified HRSG functional zones. The ordering and
complexity vary by plant, but economizing, evaporation, and superheating are
the essential heat-recovery functions.
HRSG section
Main purpose
Water–steam condition
Economizer
Uses lower-temperature exhaust heat to warm feedwater before boiling
Pressurized liquid water
Evaporator
Supplies latent heat required to convert water into saturated steam
Water–steam mixture and saturated steam
Superheater
Raises steam temperature above saturation before turbine admission
Superheated steam
Reheater, when used
Reheats partially expanded steam between turbine sections
Intermediate-pressure steam
Large HRSGs can contain multiple pressure levels. A three-pressure HRSG may
generate high-, intermediate-, and low-pressure steam to recover heat more
effectively across the exhaust-temperature range. Some units use supplementary
firing, where additional fuel is burned in the exhaust path to increase steam
production. Those design choices improve flexibility or output but alter fuel
use, emissions, and efficiency.
Heat transfer is limited by practical temperature differences, pressure losses,
heat-exchanger surface area, material constraints, and the risk of thermal stress.
The HRSG must therefore balance maximum heat recovery against cost, size,
durability, and operating flexibility.
5. The Steam-Turbine System
Steam generated in the HRSG expands through the steam turbine. As pressure and
temperature fall, the steam transfers energy to turbine blades and creates
mechanical shaft power. The connected generator converts that shaft power into
electricity.
The U.S. Department of Energy explains that a steam turbine is driven by
high-pressure steam produced by a boiler or HRSG; the steam turbine itself does
not directly consume fuel.[3]
Figure 4. Closed steam–water loop. Water is repeatedly
pressurized, heated, expanded as steam, condensed, and returned to the HRSG.
A large steam turbine may have high-pressure, intermediate-pressure, and
low-pressure sections. Reheat can return partially expanded steam to the HRSG
before it enters a later turbine section. These arrangements improve heat use and
help control steam quality near the turbine exit.
The steam-turbine contribution depends on steam flow, steam pressure and
temperature, turbine efficiency, condenser pressure, generator efficiency, and
the amount of recoverable heat entering the HRSG.
6. The Condenser and Feedwater Pump
6.1 Why the condenser is necessary
Steam leaving the final turbine section enters the condenser, where heat is
rejected to a cooling medium and the steam becomes liquid water. Condensation
serves two purposes. It recovers the working fluid for reuse and maintains a low
steam-turbine exhaust pressure.
A lower exhaust pressure allows steam to expand through a larger pressure range,
which can increase turbine work. This is why steam-side vacuum is relevant to
plant output and why the CCPP dataset includes an exhaust-vacuum variable.
Detailed interpretation of that variable and its units is reserved for
Tutorial 03.
6.2 Cooling system
The condenser must transfer rejected heat to the environment. Cooling can be
provided by once-through water, recirculating wet cooling towers, air-cooled
condensers, or hybrid systems. Cooling-system design affects water use, auxiliary
power, condenser pressure, sensitivity to weather, and overall plant performance.
6.3 Feedwater pump
After condensation, the feedwater pump raises the liquid pressure so the water
can return to the HRSG. Pumping a liquid requires much less work than compressing
the same substance as a vapor, which is one of the reasons the closed Rankine
cycle is practical.
7. Why Combined Cycle Is More Efficient
A simple-cycle gas turbine generates electricity once and then rejects its hot
exhaust. A combined-cycle plant generates electricity in the gas turbine and
then uses the exhaust heat to produce additional steam-turbine power. More of the
fuel's available energy is therefore converted into useful electrical output.
The U.S. Energy Information Administration reported average 2020 operating heat
rates of approximately 10,000 Btu/kWh for simple-cycle systems
and 7,146 Btu/kWh for combined-cycle systems.[2] Heat rate is fuel energy input divided by electrical
energy output, so a lower value indicates better conversion efficiency.
Heat rate = fuel-energy input ÷ net electrical-energy output
Figure 5. Lower heat rate means less fuel energy is required
per kilowatt-hour. The values are EIA's reported 2020 averages, not universal
ratings for every plant.
Using 3,412 Btu as the electrical-energy equivalent of one kilowatt-hour, those
heat rates correspond to approximate higher-heating-value efficiencies of 34.1%
and 47.8%, respectively. These calculated figures are illustrative fleet
averages. Modern design-point performance can differ substantially according to
turbine class, cooling, pressure levels, supplementary firing, fuel, ambient
conditions, age, maintenance, and operating load.
Do not treat a single efficiency value as a plant constant.
Thermal efficiency changes with load, weather, equipment condition, control
mode, and auxiliary consumption. Vendor design values, annual fleet averages,
and measured hourly efficiency answer different questions.
8. Common Combined-Cycle Configurations
A power block can contain one or more gas turbines connected thermally to one or
more steam turbines. The notation 1×1, 2×1, or
3×1 describes the number of combustion turbines and steam turbines
within the block.
Configuration
Meaning
General characteristic
1×1
One gas turbine and one steam turbine
Smaller block and comparatively simple integration
2×1
Two gas turbines supplying one steam turbine
Common utility-scale arrangement
3×1
Three gas turbines supplying one steam turbine
Larger block with shared steam-cycle equipment
Single shaft
Gas turbine, steam turbine, and generator share a shaft train
Compact arrangement with tightly coupled operation
Multi-shaft
Gas and steam turbines drive separate generators
Greater separation of turbine-generator trains
EIA reported that the predominant U.S. combined-cycle power-block configuration
was two combustion turbines with one steam turbine.[2] However, actual layouts vary, and the public CCPP
dataset should not be used to infer an undocumented physical configuration of the
specific plant.
9. How Environmental Conditions Affect Plant Operation
The CCPP dataset is built around three ambient measurements and one steam-side
vacuum measurement. These variables matter because the plant exchanges mass and
energy with its surroundings. Their effects do not occur through a single
isolated equation; they propagate through compressor intake, gas-turbine flow,
heat recovery, cooling, steam expansion, and auxiliary loads.
Figure 6. Principal physical pathways linking recorded
conditions to net output. The arrows represent mechanisms, not an assumption
that each effect is independent or perfectly linear.
9.1 Ambient temperature
Ambient temperature influences inlet-air density, compressor behavior, cooling
performance, and the mass flow processed by the gas turbine. Warmer air is
generally less dense at similar pressure and humidity, so a fixed-volume intake
can admit less air mass. Higher cooling-water or air-cooler temperatures can also
raise condenser pressure. For many gas-turbine and combined-cycle plants, these
mechanisms make high ambient temperature unfavorable for net power output.
9.2 Ambient pressure
Atmospheric pressure influences air density and therefore the mass of air entering
the compressor. Site altitude and weather both contribute to pressure variation.
Lower inlet pressure can reduce gas-turbine mass flow and power, although the
magnitude depends on the machine and controls.
9.3 Relative humidity
Humidity changes moist-air properties and can interact with combustion, inlet
cooling, compressor flow, and cooling-system behavior. Its influence is usually
more design- and condition-dependent than a simple statement such as “more
humidity always increases output” or “always decreases output.” The dataset can
reveal an empirical relationship, but that relationship must be interpreted in
the context of the other variables.
9.4 Exhaust vacuum
Steam-turbine exhaust vacuum is linked to condenser pressure. Better vacuum means
lower absolute pressure at the turbine exhaust, giving steam a larger expansion
range. The measured variable can therefore carry information about steam-cycle
conditions and cooling performance. Care is required with sign conventions and
the practical meaning of “higher vacuum”; Tutorial 03 handles the dataset column
in detail.
10. Gross Output, Auxiliary Loads, and Net Output
The two generators produce gross electrical power, but the plant also consumes
electricity. Pumps circulate feedwater and cooling water. Fans move air through
cooling systems. Fuel systems, control equipment, lubrication systems,
transformers, emissions controls, and other services require power.
Figure 7. Net output is the power available after internal
plant consumption has been subtracted from gross generation.
The UCI dataset describes PE as net hourly electrical energy output
and lists the unit as megawatts. Because megawatts measure power, the most precise
interpretation is an hourly averaged net electrical power output.
The word “hourly” refers to the averaging interval rather than changing MW into
an energy unit.
11. Connecting Plant Physics to the CCPP Dataset
The physical explanation now gives the five dataset variables a clear place in
the plant.
Dataset variable
Plant-side connection
Detailed treatment
AT
Gas-turbine inlet and plant cooling environment
Tutorial 03
AP
Atmospheric condition affecting inlet-air density and mass flow
Tutorial 03
RH
Moist-air and cooling-related environmental condition
Tutorial 03
V
Steam-turbine exhaust and condenser-side condition
Tutorial 03
PE
Hourly averaged net electrical power output
Tutorial 03
The original dataset paper explains that ambient temperature, pressure, and
relative humidity are major factors for gas-turbine performance, while exhaust
vacuum is measured from the steam-turbine side.[5] This structure is what makes the dataset particularly
interesting: four compact measurements summarize conditions that influence two
coupled power cycles.
A machine-learning model does not explicitly simulate every compressor stage,
heat-exchanger surface, steam property, or control loop. It learns the empirical
relationship between the four recorded inputs and measured net output. Plant
physics remains essential because it helps us judge whether the learned
relationship is plausible, where it may fail, and what information is missing.
12. Key Takeaways
Two power cycles
A gas turbine and steam turbine generate electricity in one integrated plant.
Exhaust heat is reused
The HRSG converts gas-turbine exhaust heat into steam-cycle input.
Cooling matters
Condenser pressure and cooling conditions influence steam-turbine output.
Net is not gross
Internal auxiliary consumption is subtracted from generator output.
Weather reaches both cycles
Ambient conditions affect intake, compression, heat recovery, and cooling.
U.S. Energy Information Administration. (2022).
Most Combined-Cycle Power Plants Employ Two Combustion Turbines with
One Steam Turbine.
EIA Today in Energy
.
U.S. Department of Energy. (2016).
Combined Heat and Power Technology Fact Sheet Series: Steam Turbines.
DOE Steam Turbine Fact Sheet
.
Tüfekci, P., & Kaya, H. (2014).
Combined Cycle Power Plant [Dataset].
UCI Machine Learning Repository.
https://doi.org/10.24432/C5002N
.
Kaya, H., Tüfekci, P., & Gürgen, F. S. (2012).
Local and global learning methods for predicting power of a combined gas
and steam turbine. International Conference on Emerging Trends in
Computer and Electronics Engineering, 13–18.
Tüfekci, P. (2014).
Prediction of full load electrical power output of a base load operated
combined cycle power plant using machine learning methods.
International Journal of Electrical Power & Energy Systems,
60, 126–140.
https://doi.org/10.1016/j.ijepes.2014.02.027
.
Next tutorial
03 — Understanding the CCPP Dataset Features
A detailed explanation of ambient temperature, exhaust vacuum, atmospheric
pressure, relative humidity, net electrical output, units, ranges, expected
relationships, and interpretation cautions.
Meet the real-world energy dataset that will be used throughout this tutorial
series. This introductory article explains what the dataset contains, what
quantity will eventually be predicted, and why the problem is valuable for
learning regression with Python.
Scope of Tutorial 01.
This post is an overview only. The operation of a combined cycle power plant,
detailed interpretation of individual features, Python-based data inspection,
and formal experimental design are reserved for Tutorials 02–05.
1. Dataset Overview
The Combined Cycle Power Plant dataset, usually shortened to
the CCPP dataset, is a real-world regression dataset available
through the UCI Machine Learning Repository. It contains measurements collected
from a combined cycle power plant while the plant was operating at full load.
According to the UCI repository, the dataset contains 9,568 data
points collected over six years, from 2006 to 2011.
Each observation combines four measured input variables with one measured value
of net hourly electrical energy output.
9,568
recorded observations
4
continuous predictors
1
continuous target
2006–2011
collection period
Dataset property
Overview
Application area
Electrical-power generation
Machine-learning task
Supervised regression
Number of observations
9,568
Number of input variables
4
Number of target variables
1
Operating condition
Full-load plant operation
Collection period
2006–2011
Repository
UCI Machine Learning Repository
Dataset DOI
10.24432/C5002N
The dataset is especially attractive for teaching because the table is compact,
the variables are numeric, and the prediction target has an immediate engineering
meaning. At the same time, the data are sufficiently rich to support comparisons
between simple statistical models and advanced machine-learning algorithms.
Figure 1. High-level structure of the CCPP learning problem.
Four measured variables describe each observation, and net electrical power
output is the value to be estimated.
2. What Is the Prediction Problem?
The general goal is to estimate the plant's net hourly electrical energy
output from four available measurements. The output is represented by
the variable PE and is measured in megawatts.
Because electrical output is a continuous numerical value, the problem belongs
to regression. A future machine-learning model will receive the
four input values for an observation and produce an estimated value of
PE.
At this introductory stage, the problem can be summarized as:
Use ambient and plant-related measurements to estimate the net
electrical output produced during full-load operation.
This simple statement will later be converted into a complete machine-learning
workflow. Subsequent tutorials will define the predictors and target formally,
prepare the data, establish validation rules, train regression algorithms, and
compare their performance.
3. Input and Target Variables at a Glance
The dataset contains four predictors and one target. Only a short orientation is
provided here; each variable will be examined properly in Tutorial 03.
Variable
Role
High-level meaning
Unit
AT
Input
Ambient temperature
°C
V
Input
Exhaust vacuum
cm Hg
AP
Input
Ambient pressure
mbar
RH
Input
Relative humidity
%
PE
Target
Net hourly electrical energy output
MW
The four inputs describe environmental or operational conditions associated
with a recorded hour. The target records the corresponding electrical output.
The central machine-learning question is whether the relationship between these
inputs and the target can be learned accurately enough to make useful predictions
for observations that were not used during model training.
The feature names are easy to list, but their engineering interpretation should
not be reduced to one sentence. Tutorial 03 will examine what each variable
represents, how it is measured, and why it may be related to output.
4. Why Predicting Electrical Output Is Important
Accurate output prediction can support a clearer understanding of plant
performance under changing recorded conditions. In practical energy analysis,
predicted output may contribute to planning, operational assessment, performance
monitoring, scenario comparison, and the early identification of unexpected
behavior.
The value of the dataset is not limited to power-plant engineering. It also
provides an excellent educational bridge between machine-learning theory and a
physical system. A regression score is easier to interpret when the target is
measured in megawatts and the error represents a tangible difference between
predicted and measured electrical production.
The dataset has consequently been used in research on machine-learning methods
for full-load power-output prediction. The 2014 study by Tüfekci examined several
regression approaches for predicting hourly full-load electrical power output,
helping establish the dataset as a recognized benchmark for this problem.
5. Why This Dataset Is Useful for Machine-Learning Tutorials
Figure 2. The CCPP dataset combines a real application with
a manageable structure, making it appropriate for a progressive tutorial
series.
It represents a real engineering problem
The observations originate from an operating power plant rather than an
artificial formula. This allows every modeling decision to be connected to a
recognizable energy application.
It has a clear target
The objective is not vague: estimate net hourly electrical output. This makes it
straightforward to explain regression predictions and later evaluate prediction
errors.
It has a compact feature space
Only four predictors are required. Readers can therefore focus on the modeling
process without first managing hundreds of variables, images, text fields, or
complex categorical encodings.
It supports many algorithms
The same dataset can be used with linear regression, polynomial regression,
regularized models, support-vector regression, decision trees, random forests,
gradient boosting, neural networks, symbolic regression, and other approaches.
This makes comparisons easier because the underlying prediction task remains
unchanged.
It is suitable for progressive learning
A beginner can start with dataset structure and a linear baseline. More advanced
readers can later investigate nonlinear relationships, hyperparameter tuning,
feature importance, uncertainty, interpretability, symbolic equations, and
ensemble learning.
6. Position of This Post in the CCPP Tutorial Series
Figure 3. Tutorial 01 provides orientation only. The
engineering, data-analysis, and methodological details are intentionally
separated into later posts.
Tutorial
Title
Primary purpose
01
Introduction to the Combined Cycle Power Plant Dataset
Introduce the dataset, prediction target, and practical importance
02
How a Combined Cycle Power Plant Works
Explain the energy-conversion system and major plant components
03
Understanding the CCPP Dataset Features
Interpret all predictors and the electrical-output target in detail
04
Loading and Inspecting the CCPP Dataset with Python
Load, validate, and inspect the table with pandas
05
Defining the CCPP Machine Learning Problem
Formalize the regression task, assumptions, questions, and evaluation plan
7. What This Introduction Does—and Does Not—Cover
Covered here
Identity and origin of the dataset
Number of observations and variables
High-level prediction objective
Names and roles of the variables
Practical and educational importance
Position within the tutorial series
Reserved for later tutorials
Detailed power-plant operation
Gas and steam turbine mechanics
Detailed feature interpretation
Data ranges and distributions
Python loading and inspection
Missing values and duplicates
Train-test splitting and validation
Regression metrics and model training
Keeping these subjects separate gives every tutorial one clear learning goal.
Readers can follow the series in order, while experienced practitioners can open
the specific article that addresses the topic they need.
8. References
Tüfekci, P., & Kaya, H. (2014).
Combined Cycle Power Plant [Dataset].
UCI Machine Learning Repository.
https://doi.org/10.24432/C5002N
Tüfekci, P. (2014).
Prediction of full load electrical power output of a base load operated
combined cycle power plant using machine learning methods.
International Journal of Electrical Power & Energy Systems,
60, 126–140.
https://doi.org/10.1016/j.ijepes.2014.02.027
Next tutorial
02 — How a Combined Cycle Power Plant Works
Gas turbines, steam turbines, heat-recovery steam generators,
environmental conditions, and net power output.
Feature Importance in Random Forests: How to Interpret Your ModelPythonholics Machine Learning Tutorial
A detailed practical guide to impurity-based feature importance, permutation
importance, correlated predictors, importance stability, feature selection,
and responsible interpretation using Python and scikit-learn.
Two core methodsMean decrease in impurity and permutation importance are explained and compared.
Complete Python workflowThe script generates all figures, result tables, metrics, and CSV files.
Random Forests are among the most useful machine-learning algorithms for structured
and tabular data. They can model nonlinear relationships, capture complex interactions,
handle classification and regression problems, and often provide strong performance
without requiring extensive data transformation. Nevertheless, a high-performing
model is not automatically an understandable model.
After a Random Forest has been trained, an important practical question remains:
which input features influence the model most strongly? Feature
importance attempts to answer this question by assigning a numerical score to every
predictor. These scores can help us understand the fitted model, identify potentially
useful or redundant variables, communicate results, and design additional experiments.
Feature importance must still be interpreted cautiously. A high importance value does
not prove that a feature causes the predicted outcome. It only indicates that the
fitted model used the information associated with that feature in a particular way.
Rankings can also be affected by correlated predictors, data leakage, high-cardinality
variables, the selected scoring metric, sampling variation, and model configuration.
Main lesson: Random Forest feature importance is a model-inspection
tool, not a causal analysis. The most reliable interpretation combines several
importance methods with predictive evaluation, correlation analysis, stability
checks, and domain knowledge.
1. What Does Feature Importance Mean?
A machine-learning dataset contains input variables, commonly called
features, and an output that the model attempts to predict. For example,
a medical classification dataset may contain measurements of a tumor, while the
target indicates whether the tumor is benign or malignant. A Random Forest may use
some measurements repeatedly and depend only weakly on others.
Feature importance provides a model-based estimate of this dependence. In broad
terms, an important feature is one that helps the model reduce uncertainty, separate
classes, lower prediction error, or maintain predictive performance. However,
different feature-importance methods operationalize this idea differently.
Table 1. Practical questions feature importance can help answer
Question
How feature importance helps
Important limitation
Which variables does the model use most?
Ranks predictors according to a defined importance measure.
The ranking explains the fitted model, not necessarily the real-world mechanism.
Can the feature set be reduced?
Identifies candidates for feature-selection experiments.
Low individual importance does not prove that a feature is useless in interactions.
Is the model relying on suspicious information?
Can expose identifiers, proxies, leakage variables, or unexpected predictors.
A leakage audit still requires understanding how every variable was created.
How can the model be explained to stakeholders?
Provides a compact global summary of model reliance.
Global importance does not explain every individual prediction.
Which measurements deserve further study?
Suggests variables or feature groups for additional analysis.
Predictive association must not be presented as causation.
Global and local interpretation
Standard Random Forest feature importance is a global
interpretation method. It summarizes model behavior across many observations.
It does not directly explain why one specific observation received one specific
prediction. Local explanations require methods such as local permutation,
SHAP values, local surrogate models, or careful examination of decision paths.
Importance is method-dependent
There is no single universal definition of feature importance. A feature can be
important because it frequently creates strong decision-tree splits, because
shuffling it damages test performance, because removing it reduces cross-validation
performance, or because it contributes strongly to individual predictions.
Consequently, two valid methods can produce different rankings.
2. Why Can a Random Forest Rank Features?
A Random Forest combines many decision trees. For classification, each tree predicts
a class and the forest aggregates those predictions, usually through majority voting
or averaged class probabilities. For regression, the outputs of individual trees are
averaged.
Randomness is introduced in two important places. First, each tree is trained on a
bootstrap sample drawn from the training data. Second, only a random subset of
features is considered at each candidate split. This creates diversity among trees
and reduces the tendency of a single tree to overfit the training observations.
Every internal node of a classification tree selects a feature and threshold that
reduce impurity. Since the forest records which features were used and how much
impurity they reduced, the reductions can be aggregated into an impurity-based
importance score. Alternatively, the fitted forest can be treated as a black box and
evaluated by measuring how much its predictions deteriorate when individual features
are randomly permuted.
Figure 1. A practical workflow for interpreting a Random Forest.
Predictive performance is evaluated before impurity-based importance, permutation
importance, stability, and correlation are examined.
Figure 1 emphasizes that feature importance should not be calculated in isolation.
The model must first be trained using a valid experimental design and evaluated on
observations that were not used to fit it. Only then should importance rankings be
interpreted and validated.
3. Impurity-Based Feature Importance
The built-in importance returned by
RandomForestClassifier.feature_importances_ is commonly described as
impurity-based feature importance, mean decrease in impurity, or MDI. It measures
the accumulated reduction in node impurity attributed to each feature across all
trees.
3.1 Impurity reduction at one node
Consider a node containing a subset of training observations. A candidate split
divides those observations into a left child and a right child. For a classification
tree using Gini impurity, the impurity of node t can be written as:
Gini(t) = 1 − Σk=1K p(k|t)2
Here, K is the number of classes and p(k|t) is the proportion of observations
belonging to class k in node t. A pure node has a Gini value of zero.
The weighted impurity decrease produced by a split can be represented conceptually as:
ΔI(t) = w(t)I(t) − w(tL)I(tL) − w(tR)I(tR)
The symbols t, tL, and
tR denote the parent, left-child, and right-child nodes.
The function I denotes impurity, and w represents the proportion
or weighted proportion of observations reaching a node.
3.2 Aggregation across the forest
Whenever feature j is used for a split, its weighted impurity reduction is
added to its total. The totals are averaged across the trees and normalized so that
the final importance values sum to one:
MDI(j) = normalized sum of weighted impurity decreases produced by feature j
A feature obtains a high MDI value when it is used frequently, produces substantial
impurity reductions, affects many training observations, or combines these properties.
3.3 Advantages of MDI
It is available immediately after the forest has been fitted.
It is inexpensive because no additional model evaluation is required.
It provides a simple ranking whose values sum to one.
It can be calculated separately for every tree to study stability.
It is useful as a fast initial overview of the forest structure.
3.4 Limitations and bias
MDI is not a neutral measure. It is computed from the same training process that built
the forest, so it describes internal split behavior rather than an independently
measured loss of predictive performance. It can favor continuous variables and
variables with many possible split points. Correlated predictors can share importance
unevenly, and a variable can look important even when its contribution is redundant.
Do not interpret the largest MDI value as proof of causality.
The score only indicates that the trained forest assigned substantial split-based
importance to the feature.
The clearest first visualization is a horizontal bar chart in which features are
sorted by MDI. Error bars can display the standard deviation of importance across
individual trees. This adds information that the single forest-level vector does
not provide.
Figure 2. Top ten features according to mean decrease in
impurity. Longer bars indicate a greater aggregated contribution to weighted
impurity reduction. Error bars summarize variation across individual trees.
Figure 2 should be interpreted as a description of the forest's internal structure.
A long bar indicates that the feature contributed strongly to splits across the
ensemble. The error bar indicates whether that contribution was consistent among
trees or concentrated in only part of the forest.
4. Permutation Feature Importance
Permutation importance measures the extent to which predictive performance depends
on a feature. It can be applied to Random Forests and many other fitted estimators,
which makes it a model-agnostic inspection method.
4.1 Core idea
First, the fitted model is evaluated on a validation or test dataset to obtain a
baseline score. Next, the values of one feature are randomly shuffled across rows.
This destroys the association between that feature and the target while leaving the
feature's marginal distribution unchanged. The model is evaluated again, and the
decrease in performance becomes the importance estimate.
PI(j) = baseline score − score after permuting feature j
The procedure is repeated several times because one random permutation can produce
an unusually large or small change. The mean score decrease and standard deviation
are then reported.
4.2 Step-by-step algorithm
Fit the Random Forest using the training data.
Evaluate the fitted forest on an independent validation or test set.
Select one feature and randomly permute its values across observations.
Evaluate the model using the modified dataset.
Record the decrease in the selected performance score.
Repeat the permutation several times and calculate the mean and standard deviation.
Restore the feature and repeat the process for every remaining predictor.
4.3 Choosing the scoring metric
Permutation importance depends on the chosen score. Accuracy may be suitable for a
balanced classification problem, but balanced accuracy, F1-score, ROC-AUC, average
precision, or a cost-sensitive score may be more meaningful in other settings.
Regression problems may use R², negative mean squared error, or negative mean
absolute error.
This tutorial uses balanced accuracy. Consequently, the permutation
importance of a feature is the mean decrease in balanced accuracy after that feature
is shuffled.
4.4 Interpreting positive, zero, and negative values
Shuffling the feature substantially damages model performance.
Treat the feature as important, then verify stability and correlation.
Small positive value
The feature has a limited measurable contribution under the selected score.
Compare uncertainty and test feature removal experimentally.
Value near zero
The model can largely maintain performance when the feature is shuffled.
Check whether correlated features preserve the same information.
Negative value
The model performed slightly better after shuffling the feature.
Consider sampling variation, noise, instability, or harmful reliance.
The mean permutation importance should be shown with an uncertainty estimate.
A feature whose mean is positive but whose error bar crosses zero may not have a
reliably positive contribution under the current test sample and scoring metric.
Figure 3. Permutation importance calculated on the held-out
test subset. Each bar represents the mean decrease in balanced accuracy, while
each error bar represents the standard deviation across repeated permutations.
Figure 3 is more directly connected to predictive performance than Figure 2.
Nevertheless, it is not automatically free from interpretation problems. In
particular, correlated predictors can cause permutation importance to underestimate
the importance of information that is duplicated across multiple features.
5. MDI vs Permutation Importance
MDI and permutation importance answer related but different questions. MDI asks how
strongly the feature contributed to impurity reduction inside the trained forest.
Permutation importance asks how much test performance deteriorates when the feature's
information is disrupted.
Table 3. Comparison of impurity-based and permutation importance
Property
Impurity-based importance
Permutation importance
Main quantity
Aggregated weighted impurity reduction
Decrease in a selected predictive score
Data normally used
Training process and fitted tree structure
Preferably validation or test data
Computational cost
Low
Higher because predictions are repeated
Model dependence
Specific to tree-based models
Applicable to many fitted models
High-cardinality bias
Can be substantial
Generally less direct, but not immune to dataset problems
Correlated predictors
Importance may be unevenly shared or concentrated
Importance may be masked by substitute predictors
Best use
Fast structural overview
Performance-based validation of model reliance
5.1 Why the numerical scales cannot be compared directly
MDI values are normalized to sum to one. Permutation values represent changes in a
performance metric and do not generally sum to one. Therefore, a value of 0.10 in
MDI is not numerically equivalent to a 0.10 decrease in balanced accuracy.
The Python script creates a normalized version of positive permutation importance
only for visual comparison. This normalization makes broad patterns easier to see,
but it does not transform the two measures into the same mathematical quantity.
Figure 4. Visual comparison of MDI and positive permutation
importance after separate normalization. The normalization is used only to
compare broad ranking patterns; it does not make the methods equivalent.
5.2 Understanding disagreement
A feature can have high MDI and low permutation importance when it is used frequently
by the trees but is redundant with other predictors. When shuffled, the remaining
variables preserve enough information for the model to maintain performance.
A feature can also have moderate MDI and comparatively high permutation importance.
This suggests that it may not dominate the split structure, but disrupting it still
damages predictive performance. Such disagreement is not necessarily an error. It is
a signal that the data structure should be examined more carefully.
6. Practical Experiment with the Breast Cancer Dataset
The complete Python example uses the Breast Cancer Wisconsin diagnostic dataset
distributed with scikit-learn. The problem is binary classification. Each observation
contains measurements derived from a digitized image of a fine-needle aspirate of a
breast mass. The target classes are malignant and benign.
Table 4. Dataset summary
Property
Value
Observations
569
Input features
30 numerical variables
Target classes
2
Malignant observations
212
Benign observations
357
Missing values
0
Train/test split
70% / 30%, stratified
6.1 Random Forest configuration
Table 5. Random Forest configuration used in the example
Hyperparameter
Value
Purpose
n_estimators
500
Builds a stable ensemble of 500 trees.
criterion
gini
Uses Gini impurity for classification splits.
max_depth
None
Allows trees to expand until other stopping conditions apply.
min_samples_split
2
Minimum observations needed to split an internal node.
min_samples_leaf
1
Minimum observations required in a leaf.
max_features
sqrt
Considers a random square-root-sized feature subset at each split.
bootstrap
True
Trains each tree using a bootstrap sample.
oob_score
True
Calculates an out-of-bag performance estimate.
random_state
42
Makes the experiment reproducible.
n_jobs
-1
Uses available processor cores where supported.
6.2 Evaluate the model before interpreting it
Feature-importance rankings are meaningful only when the underlying model has learned
a useful predictive relationship. A poor model can still produce rankings, but those
rankings explain a model whose predictions are unreliable.
Table 6. Representative performance from the reproducible example
Metric
Value
Accuracy
0.9474
Balanced accuracy
0.9391
Precision
0.9455
Recall
0.9720
F1-score
0.9585
Matthews correlation coefficient
0.8872
ROC-AUC
0.9917
Out-of-bag score
0.9673
The confusion matrix provides more detail than accuracy alone. It distinguishes
correctly identified malignant cases, malignant cases incorrectly classified as
benign, benign cases incorrectly classified as malignant, and correctly identified
benign cases.
Figure 5. Confusion matrix for the held-out test set. In the
representative run, 58 malignant and 104 benign observations were classified
correctly, while six malignant and three benign observations were misclassified.
In a medical context, the two error types may have different consequences. For this
reason, model evaluation should not rely on one aggregate score. The example reports
several complementary metrics and retains the confusion matrix as part of the
interpretation workflow.
7. How Correlated Features Affect Importance
Correlation is one of the main reasons why feature-importance rankings can be
misleading. Suppose that two predictors contain nearly the same information. A tree
may use the first feature in one branch and the second in another branch. Across the
forest, the importance may be split between them or concentrated unevenly.
7.1 Effect on MDI
MDI may assign most of the importance to one feature from a correlated group because
that feature happened to produce slightly better splits. Another nearly equivalent
feature can then appear unimportant even though it contains similar predictive
information.
7.2 Effect on permutation importance
Permuting one correlated feature may produce only a small performance decrease
because another feature still provides similar information. As a result, the
importance of both features can be underestimated when each is permuted separately.
A correlation matrix should therefore be examined alongside the importance rankings.
The Breast Cancer dataset contains several radius, perimeter, area, concavity, and
concave-point measurements that are strongly related.
Figure 6. Correlation among the ten highest-ranked MDI features.
Strongly correlated predictors can share, redistribute, or mask importance and
should often be interpreted as related feature groups.
7.3 Better approaches for correlated predictors
Grouped permutation: shuffle a related group of predictors together.
Feature clustering: cluster correlated variables and retain representative features.
Conditional permutation: permute a feature conditionally on related variables.
Domain grouping: interpret related measurements as one conceptual factor.
Ablation experiments: remove groups of predictors and retrain the entire pipeline.
8. Stability and Uncertainty of Feature Importance
A single importance vector can create a false impression of certainty. Random Forests
are stochastic models, and individual trees use different bootstrap samples and
feature subsets. Consequently, importance can vary from one tree to another and from
one fitted forest to another.
8.1 Tree-to-tree variability
The importance values from all 500 individual trees can be collected and displayed
as boxplots. A narrow distribution suggests that many trees assign a similar level
of importance to the feature. A wide distribution indicates that the feature is
highly important in some trees but less important in others.
Figure 7. Distribution of feature importance across individual
trees. Wide distributions indicate that a feature's role varies substantially
across bootstrap samples and random feature subsets.
8.2 Stability across repeated model fits
Tree-to-tree variability is informative, but it does not replace repeated model
fitting. A stronger analysis repeats the complete training procedure using different
cross-validation folds or random seeds. The resulting feature ranks can then be
summarized using means, standard deviations, confidence intervals, or rank
correlations.
Recommended publication practice: report importance variability
across repeated cross-validation runs rather than presenting one importance ranking
from one fitted model as definitive.
9. Cumulative Importance and Feature Selection
Feature importance is often used to create a reduced feature subset. A common
approach sorts features from highest to lowest MDI and calculates cumulative
importance. This shows how quickly the total reaches thresholds such as 80% or 90%.
Figure 8. Cumulative impurity-based feature importance. The curve
indicates how many top-ranked features account for a selected proportion of total
MDI, but it should not be used as an automatic deletion rule.
9.1 Why threshold selection is not enough
Keeping enough features to reach 90% cumulative MDI does not guarantee that the
reduced model will retain 90% of predictive performance. Importance values are not
percentages of accuracy. A low-ranked feature may provide unique information,
contribute only through interactions, or become important after another feature is
removed.
9.2 Correct feature-selection procedure
Calculate importance using only the training process or training folds.
Choose one or more candidate subsets, such as the top 5, 10, 15, and 20 features.
Retrain the complete preprocessing and modeling pipeline for each subset.
Evaluate every candidate using cross-validation or a nested validation design.
Select the smallest subset whose performance is acceptably close to the full model.
Evaluate the selected pipeline once on the untouched test set.
Avoid leakage: do not calculate importance on the complete dataset
and then report cross-validation performance using the selected variables. Feature
selection must be performed inside the training folds.
10. Complete Python Code
The following script generates every figure and CSV table referenced in this article.
Save it as random_forest_feature_importance_tutorial.py and run it from
Spyder, a terminal, or another Python environment.
10.1 Required packages
pip install numpy pandas matplotlib scikit-learn
10.2 Full reproducible script
"""
Random Forest Feature Importance Tutorial
Generates all tables and figures used in the Pythonholics article:
"Feature Importance in Random Forests: How to Interpret Your Model"
Tested workflow:
- Breast Cancer Wisconsin diagnostic dataset
- RandomForestClassifier
- Impurity-based feature importance
- Permutation feature importance
- Tree-to-tree importance stability
- Correlation analysis
- Cumulative feature importance
"""
from pathlib import Path
import warnings
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
from sklearn.metrics import (
accuracy_score,
balanced_accuracy_score,
classification_report,
confusion_matrix,
f1_score,
matthews_corrcoef,
precision_score,
recall_score,
roc_auc_score,
)
from sklearn.model_selection import train_test_split
# =============================================================================
# 1. Configuration
# =============================================================================
RANDOM_STATE = 42
TEST_SIZE = 0.30
N_ESTIMATORS = 500
N_PERMUTATION_REPEATS = 30
TOP_N = 10
OUTPUT_DIR = Path("rf_feature_importance_outputs")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
# Times New Roman is used when it is installed on the computer.
# Matplotlib will use a fallback font when it is unavailable.
plt.rcParams["font.family"] = "Times New Roman"
plt.rcParams["font.size"] = 12
plt.rcParams["axes.titlesize"] = 15
plt.rcParams["axes.labelsize"] = 13
plt.rcParams["xtick.labelsize"] = 11
plt.rcParams["ytick.labelsize"] = 11
plt.rcParams["legend.fontsize"] = 11
plt.rcParams["figure.titlesize"] = 15
warnings.filterwarnings("ignore", category=UserWarning)
def save_figure(filename: str) -> None:
"""Save the active Matplotlib figure and close it."""
plt.tight_layout()
plt.savefig(
OUTPUT_DIR / filename,
dpi=300,
bbox_inches="tight",
)
plt.close()
# =============================================================================
# 2. Load and inspect the dataset
# =============================================================================
dataset = load_breast_cancer()
X = pd.DataFrame(
dataset.data,
columns=dataset.feature_names,
)
y = pd.Series(
dataset.target,
name="target",
)
dataset_summary = pd.DataFrame(
{
"property": [
"Number of observations",
"Number of input features",
"Number of target classes",
"Class 0 observations",
"Class 1 observations",
"Missing values",
],
"value": [
X.shape[0],
X.shape[1],
y.nunique(),
int((y == 0).sum()),
int((y == 1).sum()),
int(X.isna().sum().sum()),
],
}
)
dataset_summary.to_csv(
OUTPUT_DIR / "table_01_dataset_summary.csv",
index=False,
)
# =============================================================================
# 3. Create a stratified training/test split
# =============================================================================
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=TEST_SIZE,
random_state=RANDOM_STATE,
stratify=y,
)
# =============================================================================
# 4. Train the Random Forest classifier
# =============================================================================
model = RandomForestClassifier(
n_estimators=N_ESTIMATORS,
criterion="gini",
max_depth=None,
min_samples_split=2,
min_samples_leaf=1,
max_features="sqrt",
bootstrap=True,
oob_score=True,
class_weight=None,
random_state=RANDOM_STATE,
n_jobs=-1,
)
model.fit(X_train, y_train)
# =============================================================================
# 5. Evaluate predictive performance
# =============================================================================
y_pred = model.predict(X_test)
y_probability = model.predict_proba(X_test)[:, 1]
performance = pd.DataFrame(
{
"metric": [
"Accuracy",
"Balanced accuracy",
"Precision",
"Recall",
"F1-score",
"Matthews correlation coefficient",
"ROC-AUC",
"Out-of-bag score",
],
"value": [
accuracy_score(y_test, y_pred),
balanced_accuracy_score(y_test, y_pred),
precision_score(y_test, y_pred),
recall_score(y_test, y_pred),
f1_score(y_test, y_pred),
matthews_corrcoef(y_test, y_pred),
roc_auc_score(y_test, y_probability),
model.oob_score_,
],
}
)
performance.to_csv(
OUTPUT_DIR / "table_02_model_performance.csv",
index=False,
)
classification_report_table = pd.DataFrame(
classification_report(
y_test,
y_pred,
target_names=dataset.target_names,
output_dict=True,
zero_division=0,
)
).transpose()
classification_report_table.to_csv(
OUTPUT_DIR / "table_03_classification_report.csv"
)
cm = confusion_matrix(y_test, y_pred)
confusion_matrix_table = pd.DataFrame(
cm,
index=["Actual malignant", "Actual benign"],
columns=["Predicted malignant", "Predicted benign"],
)
confusion_matrix_table.to_csv(
OUTPUT_DIR / "table_04_confusion_matrix.csv"
)
# =============================================================================
# 6. Calculate impurity-based feature importance
# =============================================================================
# Importance produced by the complete forest.
mdi_table = pd.DataFrame(
{
"feature": X.columns,
"mdi_importance": model.feature_importances_,
}
)
# Importance values from individual trees allow us to estimate stability.
tree_importances = np.array(
[tree.feature_importances_ for tree in model.estimators_]
)
mdi_table["tree_importance_std"] = tree_importances.std(axis=0)
mdi_table = (
mdi_table
.sort_values("mdi_importance", ascending=False)
.reset_index(drop=True)
)
mdi_table["mdi_rank"] = np.arange(1, len(mdi_table) + 1)
mdi_table["cumulative_mdi"] = mdi_table["mdi_importance"].cumsum()
mdi_table.to_csv(
OUTPUT_DIR / "table_05_mdi_feature_importance.csv",
index=False,
)
# =============================================================================
# 7. Calculate permutation feature importance
# =============================================================================
permutation_result = permutation_importance(
model,
X_test,
y_test,
scoring="balanced_accuracy",
n_repeats=N_PERMUTATION_REPEATS,
random_state=RANDOM_STATE,
n_jobs=-1,
)
permutation_table = pd.DataFrame(
{
"feature": X.columns,
"permutation_mean": permutation_result.importances_mean,
"permutation_std": permutation_result.importances_std,
}
)
permutation_table = (
permutation_table
.sort_values("permutation_mean", ascending=False)
.reset_index(drop=True)
)
permutation_table["permutation_rank"] = np.arange(
1,
len(permutation_table) + 1,
)
permutation_table.to_csv(
OUTPUT_DIR / "table_06_permutation_feature_importance.csv",
index=False,
)
# =============================================================================
# 8. Create a combined comparison table
# =============================================================================
comparison_table = pd.merge(
mdi_table,
permutation_table,
on="feature",
how="inner",
)
comparison_table["absolute_rank_difference"] = (
comparison_table["mdi_rank"]
- comparison_table["permutation_rank"]
).abs()
# Normalization below is used only to make the two methods easier to compare
# visually. It does not make the methods mathematically equivalent.
positive_permutation = comparison_table["permutation_mean"].clip(lower=0)
if positive_permutation.sum() > 0:
comparison_table["normalized_positive_permutation"] = (
positive_permutation / positive_permutation.sum()
)
else:
comparison_table["normalized_positive_permutation"] = 0.0
comparison_table.to_csv(
OUTPUT_DIR / "table_07_importance_comparison.csv",
index=False,
)
# =============================================================================
# 9. Figure 1: Conceptual Random Forest interpretation workflow
# =============================================================================
fig, ax = plt.subplots(figsize=(12, 4.8))
ax.axis("off")
workflow_nodes = [
(0.10, "Training data"),
(0.32, "Random Forest\nclassifier"),
(0.55, "Model\nperformance"),
(0.76, "Feature-importance\nmethods"),
(0.94, "Interpretation and\nvalidation"),
]
for x_position, label in workflow_nodes:
ax.text(
x_position,
0.52,
label,
ha="center",
va="center",
transform=ax.transAxes,
bbox={
"boxstyle": "round,pad=0.7",
"facecolor": "white",
"edgecolor": "black",
},
)
for index in range(len(workflow_nodes) - 1):
x_start = workflow_nodes[index][0] + 0.07
x_end = workflow_nodes[index + 1][0] - 0.08
ax.annotate(
"",
xy=(x_end, 0.52),
xytext=(x_start, 0.52),
xycoords=ax.transAxes,
textcoords=ax.transAxes,
arrowprops={"arrowstyle": "->", "linewidth": 1.5},
)
ax.text(
0.76,
0.23,
"MDI | permutation | stability | correlation",
ha="center",
va="center",
transform=ax.transAxes,
)
ax.set_title("A Practical Workflow for Interpreting Random Forest Feature Importance")
save_figure("figure_01_random_forest_interpretation_workflow.png")
# =============================================================================
# 10. Figure 2: Top impurity-based feature importances
# =============================================================================
mdi_top = mdi_table.head(TOP_N).sort_values(
"mdi_importance",
ascending=True,
)
plt.figure(figsize=(11, 7))
plt.barh(
mdi_top["feature"],
mdi_top["mdi_importance"],
xerr=mdi_top["tree_importance_std"],
capsize=3,
)
plt.xlabel("Mean decrease in impurity")
plt.ylabel("Feature")
plt.title("Top 10 Features by Impurity-Based Importance")
save_figure("figure_02_mdi_feature_importance.png")
# =============================================================================
# 11. Figure 3: Top permutation importances
# =============================================================================
permutation_top = permutation_table.head(TOP_N).sort_values(
"permutation_mean",
ascending=True,
)
plt.figure(figsize=(11, 7))
plt.barh(
permutation_top["feature"],
permutation_top["permutation_mean"],
xerr=permutation_top["permutation_std"],
capsize=3,
)
plt.axvline(0, linewidth=1)
plt.xlabel("Mean decrease in balanced accuracy")
plt.ylabel("Feature")
plt.title("Top 10 Features by Permutation Importance")
save_figure("figure_03_permutation_feature_importance.png")
# =============================================================================
# 12. Figure 4: Normalized MDI/permutation comparison
# =============================================================================
comparison_top_features = (
comparison_table
.sort_values("mdi_importance", ascending=False)
.head(TOP_N)
.copy()
)
plot_positions = np.arange(len(comparison_top_features))
bar_width = 0.38
plt.figure(figsize=(13, 7))
plt.bar(
plot_positions - bar_width / 2,
comparison_top_features["mdi_importance"],
width=bar_width,
label="Normalized MDI",
)
plt.bar(
plot_positions + bar_width / 2,
comparison_top_features["normalized_positive_permutation"],
width=bar_width,
label="Normalized positive permutation importance",
)
plt.xticks(
plot_positions,
comparison_top_features["feature"],
rotation=75,
ha="right",
)
plt.ylabel("Normalized importance used for visual comparison")
plt.xlabel("Feature")
plt.title("Comparison of MDI and Permutation Feature Importance")
plt.legend()
save_figure("figure_04_mdi_vs_permutation_comparison.png")
# =============================================================================
# 13. Figure 5: Confusion matrix
# =============================================================================
plt.figure(figsize=(7, 6))
plt.imshow(cm, interpolation="nearest", aspect="auto")
plt.colorbar()
class_labels = ["Malignant", "Benign"]
plt.xticks([0, 1], class_labels)
plt.yticks([0, 1], class_labels)
plt.xlabel("Predicted class")
plt.ylabel("True class")
plt.title("Random Forest Confusion Matrix")
threshold = cm.max() / 2
for row in range(cm.shape[0]):
for column in range(cm.shape[1]):
plt.text(
column,
row,
str(cm[row, column]),
ha="center",
va="center",
)
save_figure("figure_05_confusion_matrix.png")
# =============================================================================
# 14. Figure 6: Correlation heatmap for the top MDI features
# =============================================================================
top_feature_names = mdi_table.head(TOP_N)["feature"].tolist()
correlation_matrix = X[top_feature_names].corr(method="pearson")
plt.figure(figsize=(11, 9))
image = plt.imshow(
correlation_matrix,
interpolation="nearest",
aspect="auto",
vmin=-1,
vmax=1,
)
plt.colorbar(image, label="Pearson correlation coefficient")
plt.xticks(
np.arange(len(top_feature_names)),
top_feature_names,
rotation=75,
ha="right",
)
plt.yticks(
np.arange(len(top_feature_names)),
top_feature_names,
)
plt.title("Correlation Among the Top 10 MDI Features")
save_figure("figure_06_top_feature_correlation_heatmap.png")
# =============================================================================
# 15. Figure 7: Tree-to-tree importance stability
# =============================================================================
top_indices = [
X.columns.get_loc(feature)
for feature in mdi_table.head(TOP_N)["feature"]
]
tree_importance_top = tree_importances[:, top_indices]
tree_importance_labels = mdi_table.head(TOP_N)["feature"].tolist()
plt.figure(figsize=(13, 7))
plt.boxplot(
tree_importance_top,
tick_labels=tree_importance_labels,
showfliers=False,
)
plt.xticks(rotation=75, ha="right")
plt.xlabel("Feature")
plt.ylabel("Importance across individual trees")
plt.title("Tree-to-Tree Variability of the Top Feature Importances")
save_figure("figure_07_tree_importance_stability.png")
# =============================================================================
# 16. Figure 8: Cumulative MDI importance
# =============================================================================
feature_count = np.arange(1, len(mdi_table) + 1)
plt.figure(figsize=(10, 6))
plt.plot(
feature_count,
mdi_table["cumulative_mdi"],
marker="o",
)
plt.axhline(0.80, linestyle="--", label="80% importance")
plt.axhline(0.90, linestyle=":", label="90% importance")
plt.xlabel("Number of features included")
plt.ylabel("Cumulative MDI importance")
plt.title("Cumulative Impurity-Based Feature Importance")
plt.grid(True, alpha=0.3)
plt.legend()
save_figure("figure_08_cumulative_feature_importance.png")
# =============================================================================
# 17. Determine the numbers of features required for thresholds
# =============================================================================
features_for_80 = int(
np.argmax(mdi_table["cumulative_mdi"].to_numpy() >= 0.80) + 1
)
features_for_90 = int(
np.argmax(mdi_table["cumulative_mdi"].to_numpy() >= 0.90) + 1
)
threshold_summary = pd.DataFrame(
{
"cumulative_importance_threshold": [0.80, 0.90],
"number_of_features_required": [features_for_80, features_for_90],
}
)
threshold_summary.to_csv(
OUTPUT_DIR / "table_08_cumulative_importance_thresholds.csv",
index=False,
)
# =============================================================================
# 18. Print the main results
# =============================================================================
print("\nDATASET SUMMARY")
print(dataset_summary.to_string(index=False))
print("\nMODEL PERFORMANCE")
print(performance.to_string(index=False))
print("\nCONFUSION MATRIX")
print(confusion_matrix_table)
print("\nTOP 10 MDI FEATURES")
print(mdi_table.head(10).to_string(index=False))
print("\nTOP 10 PERMUTATION FEATURES")
print(permutation_table.head(10).to_string(index=False))
print("\nCUMULATIVE IMPORTANCE THRESHOLDS")
print(threshold_summary.to_string(index=False))
print(f"\nAll tables and figures were saved to:\n{OUTPUT_DIR.resolve()}")
10.3 Generated output files
Table 7. Files generated by the Python script
File
Content
table_01_dataset_summary.csv
Dataset dimensions, class counts, and missing-value count.
In the representative run, the highest MDI values were associated with measurements
such as worst concave points, worst perimeter, worst area, worst radius, and mean
concave points. These variables describe characteristics of the cell nuclei and
provide strong splitting opportunities for the fitted forest.
Table 8. Top ten MDI features from the representative run
Rank
Feature
MDI importance
1
worst concave points
0.136424
2
worst perimeter
0.132857
3
worst area
0.131652
4
worst radius
0.088794
5
mean concave points
0.080803
6
mean radius
0.063168
7
mean perimeter
0.053367
8
mean area
0.043312
9
mean concavity
0.042598
10
worst concavity
0.036355
11.2 Top permutation features
The permutation ranking differs considerably. Worst texture and mean texture appear
at the top, while several highly ranked MDI variables have smaller permutation values.
This is an excellent example of why one method should not be interpreted alone.
Several geometric measurements are strongly correlated, so shuffling one of them
may not damage performance substantially because related features remain available.
Table 9. Top ten permutation features from the representative run
Rank
Feature
Mean decrease in balanced accuracy
Standard deviation
1
worst texture
0.009813
0.003271
2
mean texture
0.005763
0.003119
3
mean concavity
0.004829
0.003301
4
worst radius
0.004461
0.005277
5
mean concave points
0.004206
0.003486
6
worst concavity
0.003271
0.004849
7
area error
0.003115
0.003267
8
radius error
0.002336
0.002893
9
worst smoothness
0.001558
0.002203
10
worst perimeter
0.001341
0.005475
Interpretation: the disagreement does not mean that one method is
wrong. MDI emphasizes the features used to build strong splits, while permutation
importance measures additional performance loss after disrupting one feature while
all correlated substitutes remain available.
11.3 Why permutation values may be small
Small permutation values do not necessarily mean that the model contains no useful
features. The forest can distribute predictive information across many related
variables. If several predictors are substitutes, shuffling one variable at a time
may cause only a modest score reduction. This is precisely why correlation analysis,
grouped permutation, and feature-group ablation can be valuable.
12. Common Interpretation Mistakes
12.1 Treating importance as causality
Predictive importance is not a causal effect. A feature may be a proxy for another
variable, reflect a selection process, or become important because of leakage.
12.2 Interpreting only the highest bar
The difference between the first and second feature may be unstable. Features should
be interpreted in groups, with uncertainty estimates and repeated fitting.
12.3 Ignoring correlated predictors
Correlation can redistribute MDI and mask permutation importance. Always inspect
relationships among the highest-ranked predictors.
12.4 Calculating permutation importance on training data
Training-set permutation importance can reflect overfitting. Validation or test data
provide a more meaningful estimate of model reliance on unseen observations.
12.5 Selecting features before splitting the data
Using the complete dataset to select features leaks information into model
evaluation. Importance-based selection belongs inside cross-validation or the
training portion of the experiment.
12.6 Assuming importance reveals direction
A positive importance score does not tell us whether higher feature values increase
or decrease the predicted probability. Partial dependence, accumulated local effects,
SHAP dependence plots, or targeted conditional analysis are needed to study direction.
12.7 Ignoring class imbalance and the scoring metric
Permutation rankings can change when accuracy is replaced by balanced accuracy,
recall, F1-score, ROC-AUC, or a domain-specific cost function. The score should match
the actual objective of the model.
12.8 Removing all low-ranked features automatically
A feature can have low marginal importance while contributing through interactions.
Removal decisions must be validated by retraining and evaluating the complete model.
13. Advanced Methods for Deeper Interpretation
13.1 Partial dependence plots
Partial dependence plots estimate the average predicted response as one or two
features vary. They help show whether the model response is increasing, decreasing,
nonlinear, or threshold-like. They can be misleading when the plotted feature is
strongly correlated with other predictors.
13.2 Individual conditional expectation
Individual conditional expectation curves show how predictions change for individual
observations rather than only displaying an average. They can reveal heterogeneous
effects that a partial dependence curve hides.
13.3 Accumulated local effects
Accumulated local effects plots are designed to reduce some problems caused by
correlated features. They estimate local changes in predictions within regions
supported by the observed data.
13.4 SHAP values
SHAP-based methods assign contributions to features for individual predictions and
can be aggregated into global summaries. They provide richer local explanations but
require additional assumptions, computation, and careful treatment of dependent
predictors.
13.5 Drop-column importance
Drop-column importance removes one feature, retrains the model, and measures the
resulting performance change. It can be more faithful to the retrained modeling
process than simple permutation, but it is computationally expensive because the
model must be refitted for every feature or feature group.
13.6 Grouped feature importance
When several variables describe the same concept, grouped permutation or grouped
removal may be more meaningful than ranking each column independently. For example,
radius, perimeter, and area measurements might be analyzed as a geometric feature
group.
Table 10. Choosing an interpretation method
Goal
Suitable method
Main caution
Fast global overview
MDI
Training-based bias and correlated predictors
Performance-based global reliance
Permutation importance
Depends on scoring metric and correlation
Importance after retraining
Drop-column importance
High computational cost
Average response shape
Partial dependence
Can extrapolate into unrealistic combinations
Observation-specific contribution
SHAP or local explanation
Requires careful background and dependence assumptions
Related predictors
Grouped permutation or grouped ablation
Feature groups require domain justification
14. Random Forest Regression
The same workflow applies to RandomForestRegressor. The primary
difference is the impurity criterion and the evaluation metric. Instead of Gini
impurity, regression trees typically reduce squared error, absolute error, or another
regression criterion.
Permutation importance should then be calculated using a regression score such as
R², negative mean squared error, or negative mean absolute error. The interpretation
remains the same: a feature is important when disrupting its information causes the
selected predictive score to deteriorate.
15. Best-Practice Checklist
Evaluate predictive performance before interpreting feature importance.
Use a held-out validation or test subset for permutation importance.
Select a permutation scoring metric that matches the real modeling objective.
Compare MDI with at least one performance-based method.
Inspect correlations among highly ranked predictors.
Report standard deviations or repeated-fit variability.
Interpret groups of related predictors rather than overemphasizing tiny rank differences.
Perform feature selection inside cross-validation to prevent leakage.
Retrain and reevaluate the model after removing features.
Use domain knowledge to judge whether the ranking is plausible.
Do not claim that predictive importance proves causality.
Use local explanation methods when individual predictions must be explained.
16. Conclusion
Feature importance makes Random Forests easier to inspect, but it does not provide
one final and unquestionable explanation. Impurity-based importance is fast and
reveals how features contribute to the structure of the forest. Permutation
importance measures how strongly predictive performance depends on preserving the
information in each feature.
The two methods can disagree because they measure different properties. Correlated
predictors, redundant information, model instability, scoring choices, and sampling
variation all affect the resulting rankings. A responsible analysis therefore
combines importance measures with performance evaluation, correlation analysis,
uncertainty estimates, repeated model fitting, and domain knowledge.
Used carefully, Random Forest feature importance can help identify meaningful
predictors, detect suspicious dependencies, guide feature-selection experiments,
and communicate model behavior. Used carelessly, it can produce confident but
misleading conclusions. The objective is not simply to generate a ranked bar chart;
it is to understand what the ranking measures, why it may change, and how it should
influence the next modeling decision.
17. References
Breiman, L. (2001). Random Forests. Machine Learning, 45,
5–32. https://doi.org/10.1023/A:1010933404324
Louppe, G., Wehenkel, L., Sutera, A., and Geurts, P. (2013).
Understanding Variable Importances in Forests of Randomized Trees.
Advances in Neural Information Processing Systems, 26.
Strobl, C., Boulesteix, A.-L., Zeileis, A., and Hothorn, T. (2007).
Bias in Random Forest Variable Importance Measures: Illustrations,
Sources and a Solution. BMC Bioinformatics, 8, 25.
https://doi.org/10.1186/1471-2105-8-25
Altmann, A., Toloşi, L., Sander, O., and Lengauer, T. (2010).
Permutation Importance: A Corrected Feature Importance Measure.
Bioinformatics, 26(10), 1340–1347.
https://doi.org/10.1093/bioinformatics/btq134
Fisher, A., Rudin, C., and Dominici, F. (2019). All Models Are Wrong,
but Many Are Useful: Learning a Variable's Importance by Studying an
Entire Class of Prediction Models Simultaneously.
Journal of Machine Learning Research, 20(177), 1–81.
Hooker, G., Mentch, L., and Zhou, S. (2021). Unrestricted Permutation
Forces Extrapolation: Variable Importance Requires at Least One More
Model, or There Is No Free Variable Importance.
Statistics and Computing, 31, 82.
https://doi.org/10.1007/s11222-021-10057-z
Pedregosa, F., Varoquaux, G., Gramfort, A., et al. (2011).
Scikit-learn: Machine Learning in Python.
Journal of Machine Learning Research, 12, 2825–2830.
Lundberg, S. M., and Lee, S.-I. (2017). A Unified Approach to
Interpreting Model Predictions. Advances in Neural Information
Processing Systems, 30.
Molnar, C. (2022). Interpretable Machine Learning,
second edition.