Tuesday, August 4, 2026

How a Combined Cycle Power Plant Works

CCPP Tutorial 02

Follow the complete path from ambient air and fuel to gas-turbine power, recovered exhaust heat, steam-turbine power, and the final net electrical output used in the Combined Cycle Power Plant dataset.

Gas turbine Steam turbine HRSG Condenser Net output
Purpose of Tutorial 02. This article explains the physical system behind the dataset. It does not yet analyze feature ranges, correlations, missing values, or Python code. Those topics belong to Tutorials 03 and 04.

1. What Does “Combined Cycle” Mean?

A combined cycle power plant joins two thermodynamic power-conversion systems. The first is a gas-turbine cycle, commonly described by the Brayton-cycle framework. The second is a steam-turbine cycle, commonly described by the Rankine-cycle framework. The two systems are connected by a heat-recovery steam generator, usually abbreviated HRSG.

Fuel is burned only in the gas-turbine combustor in a basic unfired configuration. The gas turbine produces electricity, but the exhaust leaving it remains very hot. Instead of rejecting that thermal energy directly to the atmosphere, the plant transfers much of it to water in the HRSG. The resulting steam drives a second turbine and generator.

Combined-cycle output = gas-turbine output + steam-turbine output

This reuse of exhaust heat is the central reason combined-cycle systems can convert fuel to electricity more efficiently than a comparable simple-cycle gas turbine. The U.S. Department of Energy describes the HRSG as a boiler that captures gas-turbine exhaust heat to produce high-pressure steam for additional power generation.[1]

Combined cycle is not the same as combined heat and power. Combined cycle combines gas- and steam-turbine power cycles. Combined heat and power, or cogeneration, intentionally supplies useful heat to an external process or building in addition to generating electricity. A plant can use both concepts, but the terms are not interchangeable.

2. Complete Energy Flow Through the Plant

Complete combined cycle power plant diagram with compressor, combustor, gas turbine, HRSG, steam turbine, condenser, pump, two generators, and grid.
Figure 1. Main energy and working-fluid paths. The upper path represents the gas-turbine system; the lower path recovers exhaust heat and circulates water and steam through the steam-turbine system.

The full process can be understood as ten linked steps.

1
Air enters the compressor. Ambient air is drawn into the gas turbine through the intake and filtration system.
2
The compressor raises air pressure. Compressing the air requires a substantial share of the turbine's mechanical work.
3
Fuel is injected and burned. Fuel mixes with compressed air in the combustion system, creating a high-temperature, high-pressure gas stream.
4
Hot gas expands through the gas turbine. The expanding gas rotates turbine blades and the main shaft.
5
The first generator produces electricity. Shaft power drives the gas-turbine generator.
6
Hot exhaust enters the HRSG. The exhaust still carries useful thermal energy after leaving the gas turbine.
7
The HRSG produces steam. Exhaust heat preheats feedwater, evaporates it, and can superheat the resulting steam.
8
Steam expands through the steam turbine. The steam's enthalpy is converted into shaft work.
9
The second generator produces electricity. The steam turbine contributes additional power without requiring a second primary fuel stream in the basic arrangement.
10
Steam is condensed and recirculated. The condenser turns exhaust steam back into water, and a feedwater pump returns it to the HRSG.

3. The Gas-Turbine System

The gas turbine is the first power-producing system and the source of heat for the steam cycle. The U.S. Department of Energy separates a gas turbine into three principal sections: the compressor, the combustion system, and the turbine section.[1]

Five-stage gas turbine sequence from air intake through compression, combustion, expansion, and mechanical output.
Figure 2. Gas-turbine sequence. Expansion creates shaft power, while the exhaust remains hot enough to supply the HRSG.

3.1 Compressor

The compressor draws in ambient air and raises its pressure before combustion. This is not a minor auxiliary process: the compressor consumes a large fraction of the turbine's gross mechanical work. The useful gas-turbine shaft output is therefore the difference between turbine work and compressor work.

GT,net = Ẇturbine − Ẇcompressor

Air density and compressor inlet conditions are consequently important. When the inlet air is less dense, the machine may process less air mass for a similar volumetric flow, which can reduce gas-turbine power. Compressor efficiency, pressure ratio, inlet pressure loss, fouling, and control settings also influence performance.

3.2 Combustion system

Fuel is injected into compressed air and burned. The combustion system must provide a stable high-temperature gas stream while controlling emissions, pressure loss, flame stability, and component temperatures. Turbine inlet temperature is a major performance parameter, but it is constrained by materials, cooling technology, emissions requirements, and equipment life.

3.3 Turbine and generator

The hot gas expands across stationary and rotating blade rows. The rotating blades drive the shaft, which simultaneously powers the compressor and the electrical generator. The gas leaving the turbine has a lower pressure and temperature than at the inlet, yet it still contains enough heat to support the second cycle.

4. The Heat-Recovery Steam Generator

The HRSG is the thermal bridge between the gas turbine and steam turbine. It does not convert heat directly into electricity. Instead, it transfers heat from gas-turbine exhaust to the water–steam circuit.

HRSG diagram showing superheater, evaporator, and economizer zones between hot exhaust and cooler stack gas.
Figure 3. Simplified HRSG functional zones. The ordering and complexity vary by plant, but economizing, evaporation, and superheating are the essential heat-recovery functions.
HRSG section Main purpose Water–steam condition
Economizer Uses lower-temperature exhaust heat to warm feedwater before boiling Pressurized liquid water
Evaporator Supplies latent heat required to convert water into saturated steam Water–steam mixture and saturated steam
Superheater Raises steam temperature above saturation before turbine admission Superheated steam
Reheater, when used Reheats partially expanded steam between turbine sections Intermediate-pressure steam

Large HRSGs can contain multiple pressure levels. A three-pressure HRSG may generate high-, intermediate-, and low-pressure steam to recover heat more effectively across the exhaust-temperature range. Some units use supplementary firing, where additional fuel is burned in the exhaust path to increase steam production. Those design choices improve flexibility or output but alter fuel use, emissions, and efficiency.

Heat transfer is limited by practical temperature differences, pressure losses, heat-exchanger surface area, material constraints, and the risk of thermal stress. The HRSG must therefore balance maximum heat recovery against cost, size, durability, and operating flexibility.

5. The Steam-Turbine System

Steam generated in the HRSG expands through the steam turbine. As pressure and temperature fall, the steam transfers energy to turbine blades and creates mechanical shaft power. The connected generator converts that shaft power into electricity.

The U.S. Department of Energy explains that a steam turbine is driven by high-pressure steam produced by a boiler or HRSG; the steam turbine itself does not directly consume fuel.[3]

Closed steam-water loop connecting HRSG, steam turbine, condenser, and feedwater pump.
Figure 4. Closed steam–water loop. Water is repeatedly pressurized, heated, expanded as steam, condensed, and returned to the HRSG.

A large steam turbine may have high-pressure, intermediate-pressure, and low-pressure sections. Reheat can return partially expanded steam to the HRSG before it enters a later turbine section. These arrangements improve heat use and help control steam quality near the turbine exit.

The steam-turbine contribution depends on steam flow, steam pressure and temperature, turbine efficiency, condenser pressure, generator efficiency, and the amount of recoverable heat entering the HRSG.

6. The Condenser and Feedwater Pump

6.1 Why the condenser is necessary

Steam leaving the final turbine section enters the condenser, where heat is rejected to a cooling medium and the steam becomes liquid water. Condensation serves two purposes. It recovers the working fluid for reuse and maintains a low steam-turbine exhaust pressure.

A lower exhaust pressure allows steam to expand through a larger pressure range, which can increase turbine work. This is why steam-side vacuum is relevant to plant output and why the CCPP dataset includes an exhaust-vacuum variable. Detailed interpretation of that variable and its units is reserved for Tutorial 03.

6.2 Cooling system

The condenser must transfer rejected heat to the environment. Cooling can be provided by once-through water, recirculating wet cooling towers, air-cooled condensers, or hybrid systems. Cooling-system design affects water use, auxiliary power, condenser pressure, sensitivity to weather, and overall plant performance.

6.3 Feedwater pump

After condensation, the feedwater pump raises the liquid pressure so the water can return to the HRSG. Pumping a liquid requires much less work than compressing the same substance as a vapor, which is one of the reasons the closed Rankine cycle is practical.

7. Why Combined Cycle Is More Efficient

A simple-cycle gas turbine generates electricity once and then rejects its hot exhaust. A combined-cycle plant generates electricity in the gas turbine and then uses the exhaust heat to produce additional steam-turbine power. More of the fuel's available energy is therefore converted into useful electrical output.

The U.S. Energy Information Administration reported average 2020 operating heat rates of approximately 10,000 Btu/kWh for simple-cycle systems and 7,146 Btu/kWh for combined-cycle systems.[2] Heat rate is fuel energy input divided by electrical energy output, so a lower value indicates better conversion efficiency.

Heat rate = fuel-energy input ÷ net electrical-energy output
Bar chart comparing illustrative 2020 U.S. average operating heat rates for simple-cycle and combined-cycle systems.
Figure 5. Lower heat rate means less fuel energy is required per kilowatt-hour. The values are EIA's reported 2020 averages, not universal ratings for every plant.

Using 3,412 Btu as the electrical-energy equivalent of one kilowatt-hour, those heat rates correspond to approximate higher-heating-value efficiencies of 34.1% and 47.8%, respectively. These calculated figures are illustrative fleet averages. Modern design-point performance can differ substantially according to turbine class, cooling, pressure levels, supplementary firing, fuel, ambient conditions, age, maintenance, and operating load.

Do not treat a single efficiency value as a plant constant. Thermal efficiency changes with load, weather, equipment condition, control mode, and auxiliary consumption. Vendor design values, annual fleet averages, and measured hourly efficiency answer different questions.

8. Common Combined-Cycle Configurations

A power block can contain one or more gas turbines connected thermally to one or more steam turbines. The notation 1×1, 2×1, or 3×1 describes the number of combustion turbines and steam turbines within the block.

Configuration Meaning General characteristic
1×1 One gas turbine and one steam turbine Smaller block and comparatively simple integration
2×1 Two gas turbines supplying one steam turbine Common utility-scale arrangement
3×1 Three gas turbines supplying one steam turbine Larger block with shared steam-cycle equipment
Single shaft Gas turbine, steam turbine, and generator share a shaft train Compact arrangement with tightly coupled operation
Multi-shaft Gas and steam turbines drive separate generators Greater separation of turbine-generator trains

EIA reported that the predominant U.S. combined-cycle power-block configuration was two combustion turbines with one steam turbine.[2] However, actual layouts vary, and the public CCPP dataset should not be used to infer an undocumented physical configuration of the specific plant.

9. How Environmental Conditions Affect Plant Operation

The CCPP dataset is built around three ambient measurements and one steam-side vacuum measurement. These variables matter because the plant exchanges mass and energy with its surroundings. Their effects do not occur through a single isolated equation; they propagate through compressor intake, gas-turbine flow, heat recovery, cooling, steam expansion, and auxiliary loads.

Diagram linking ambient temperature, pressure, humidity, and steam-side vacuum to gas and steam turbine power and net plant output.
Figure 6. Principal physical pathways linking recorded conditions to net output. The arrows represent mechanisms, not an assumption that each effect is independent or perfectly linear.

9.1 Ambient temperature

Ambient temperature influences inlet-air density, compressor behavior, cooling performance, and the mass flow processed by the gas turbine. Warmer air is generally less dense at similar pressure and humidity, so a fixed-volume intake can admit less air mass. Higher cooling-water or air-cooler temperatures can also raise condenser pressure. For many gas-turbine and combined-cycle plants, these mechanisms make high ambient temperature unfavorable for net power output.

9.2 Ambient pressure

Atmospheric pressure influences air density and therefore the mass of air entering the compressor. Site altitude and weather both contribute to pressure variation. Lower inlet pressure can reduce gas-turbine mass flow and power, although the magnitude depends on the machine and controls.

9.3 Relative humidity

Humidity changes moist-air properties and can interact with combustion, inlet cooling, compressor flow, and cooling-system behavior. Its influence is usually more design- and condition-dependent than a simple statement such as “more humidity always increases output” or “always decreases output.” The dataset can reveal an empirical relationship, but that relationship must be interpreted in the context of the other variables.

9.4 Exhaust vacuum

Steam-turbine exhaust vacuum is linked to condenser pressure. Better vacuum means lower absolute pressure at the turbine exhaust, giving steam a larger expansion range. The measured variable can therefore carry information about steam-cycle conditions and cooling performance. Care is required with sign conventions and the practical meaning of “higher vacuum”; Tutorial 03 handles the dataset column in detail.

10. Gross Output, Auxiliary Loads, and Net Output

The two generators produce gross electrical power, but the plant also consumes electricity. Pumps circulate feedwater and cooling water. Fans move air through cooling systems. Fuel systems, control equipment, lubrication systems, transformers, emissions controls, and other services require power.

Pnet = Pgas-turbine generators + Psteam-turbine generators − Pauxiliaries
Diagram showing gas-turbine and steam-turbine generator output combined as gross output, with auxiliary loads subtracted to obtain net output.
Figure 7. Net output is the power available after internal plant consumption has been subtracted from gross generation.

The UCI dataset describes PE as net hourly electrical energy output and lists the unit as megawatts. Because megawatts measure power, the most precise interpretation is an hourly averaged net electrical power output. The word “hourly” refers to the averaging interval rather than changing MW into an energy unit.

11. Connecting Plant Physics to the CCPP Dataset

The physical explanation now gives the five dataset variables a clear place in the plant.

Dataset variable Plant-side connection Detailed treatment
AT Gas-turbine inlet and plant cooling environment Tutorial 03
AP Atmospheric condition affecting inlet-air density and mass flow Tutorial 03
RH Moist-air and cooling-related environmental condition Tutorial 03
V Steam-turbine exhaust and condenser-side condition Tutorial 03
PE Hourly averaged net electrical power output Tutorial 03

The original dataset paper explains that ambient temperature, pressure, and relative humidity are major factors for gas-turbine performance, while exhaust vacuum is measured from the steam-turbine side.[5] This structure is what makes the dataset particularly interesting: four compact measurements summarize conditions that influence two coupled power cycles.

A machine-learning model does not explicitly simulate every compressor stage, heat-exchanger surface, steam property, or control loop. It learns the empirical relationship between the four recorded inputs and measured net output. Plant physics remains essential because it helps us judge whether the learned relationship is plausible, where it may fail, and what information is missing.

12. Key Takeaways

Two power cycles A gas turbine and steam turbine generate electricity in one integrated plant.
Exhaust heat is reused The HRSG converts gas-turbine exhaust heat into steam-cycle input.
Cooling matters Condenser pressure and cooling conditions influence steam-turbine output.
Net is not gross Internal auxiliary consumption is subtracted from generator output.
Weather reaches both cycles Ambient conditions affect intake, compression, heat recovery, and cooling.
Physics supports machine learning Engineering knowledge helps interpret patterns and detect implausible models.

13. References

  1. U.S. Department of Energy. How Gas Turbine Power Plants Work. Department of Energy .
  2. U.S. Energy Information Administration. (2022). Most Combined-Cycle Power Plants Employ Two Combustion Turbines with One Steam Turbine. EIA Today in Energy .
  3. U.S. Department of Energy. (2016). Combined Heat and Power Technology Fact Sheet Series: Steam Turbines. DOE Steam Turbine Fact Sheet .
  4. Tüfekci, P., & Kaya, H. (2014). Combined Cycle Power Plant [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5002N .
  5. Kaya, H., Tüfekci, P., & Gürgen, F. S. (2012). Local and global learning methods for predicting power of a combined gas and steam turbine. International Conference on Emerging Trends in Computer and Electronics Engineering, 13–18.
  6. Tüfekci, P. (2014). Prediction of full load electrical power output of a base load operated combined cycle power plant using machine learning methods. International Journal of Electrical Power & Energy Systems, 60, 126–140. https://doi.org/10.1016/j.ijepes.2014.02.027 .

Next tutorial

03 — Understanding the CCPP Dataset Features
A detailed explanation of ambient temperature, exhaust vacuum, atmospheric pressure, relative humidity, net electrical output, units, ranges, expected relationships, and interpretation cautions.

Introduction to the Combined Cycle Power Plant Dataset

CCPP Tutorial 01

Meet the real-world energy dataset that will be used throughout this tutorial series. This introductory article explains what the dataset contains, what quantity will eventually be predicted, and why the problem is valuable for learning regression with Python.

9,568 observations 4 input variables 1 target variable Regression dataset Energy application
Scope of Tutorial 01. This post is an overview only. The operation of a combined cycle power plant, detailed interpretation of individual features, Python-based data inspection, and formal experimental design are reserved for Tutorials 02–05.

1. Dataset Overview

The Combined Cycle Power Plant dataset, usually shortened to the CCPP dataset, is a real-world regression dataset available through the UCI Machine Learning Repository. It contains measurements collected from a combined cycle power plant while the plant was operating at full load.

According to the UCI repository, the dataset contains 9,568 data points collected over six years, from 2006 to 2011. Each observation combines four measured input variables with one measured value of net hourly electrical energy output.

9,568 recorded observations
4 continuous predictors
1 continuous target
2006–2011 collection period
Dataset property Overview
Application area Electrical-power generation
Machine-learning task Supervised regression
Number of observations 9,568
Number of input variables 4
Number of target variables 1
Operating condition Full-load plant operation
Collection period 2006–2011
Repository UCI Machine Learning Repository
Dataset DOI 10.24432/C5002N

The dataset is especially attractive for teaching because the table is compact, the variables are numeric, and the prediction target has an immediate engineering meaning. At the same time, the data are sufficiently rich to support comparisons between simple statistical models and advanced machine-learning algorithms.

Overview of the CCPP dataset showing four input variables leading to a regression problem and net electrical output.
Figure 1. High-level structure of the CCPP learning problem. Four measured variables describe each observation, and net electrical power output is the value to be estimated.

2. What Is the Prediction Problem?

The general goal is to estimate the plant's net hourly electrical energy output from four available measurements. The output is represented by the variable PE and is measured in megawatts.

Because electrical output is a continuous numerical value, the problem belongs to regression. A future machine-learning model will receive the four input values for an observation and produce an estimated value of PE.

At this introductory stage, the problem can be summarized as:

Use ambient and plant-related measurements to estimate the net electrical output produced during full-load operation.

This simple statement will later be converted into a complete machine-learning workflow. Subsequent tutorials will define the predictors and target formally, prepare the data, establish validation rules, train regression algorithms, and compare their performance.

3. Input and Target Variables at a Glance

The dataset contains four predictors and one target. Only a short orientation is provided here; each variable will be examined properly in Tutorial 03.

Variable Role High-level meaning Unit
AT Input Ambient temperature °C
V Input Exhaust vacuum cm Hg
AP Input Ambient pressure mbar
RH Input Relative humidity %
PE Target Net hourly electrical energy output MW

The four inputs describe environmental or operational conditions associated with a recorded hour. The target records the corresponding electrical output. The central machine-learning question is whether the relationship between these inputs and the target can be learned accurately enough to make useful predictions for observations that were not used during model training.

The feature names are easy to list, but their engineering interpretation should not be reduced to one sentence. Tutorial 03 will examine what each variable represents, how it is measured, and why it may be related to output.

4. Why Predicting Electrical Output Is Important

Accurate output prediction can support a clearer understanding of plant performance under changing recorded conditions. In practical energy analysis, predicted output may contribute to planning, operational assessment, performance monitoring, scenario comparison, and the early identification of unexpected behavior.

The value of the dataset is not limited to power-plant engineering. It also provides an excellent educational bridge between machine-learning theory and a physical system. A regression score is easier to interpret when the target is measured in megawatts and the error represents a tangible difference between predicted and measured electrical production.

The dataset has consequently been used in research on machine-learning methods for full-load power-output prediction. The 2014 study by Tüfekci examined several regression approaches for predicting hourly full-load electrical power output, helping establish the dataset as a recognized benchmark for this problem.

5. Why This Dataset Is Useful for Machine-Learning Tutorials

Diagram explaining why the CCPP dataset is useful: real engineering data, a clear target, a compact feature space, and suitability for many algorithms.
Figure 2. The CCPP dataset combines a real application with a manageable structure, making it appropriate for a progressive tutorial series.

It represents a real engineering problem

The observations originate from an operating power plant rather than an artificial formula. This allows every modeling decision to be connected to a recognizable energy application.

It has a clear target

The objective is not vague: estimate net hourly electrical output. This makes it straightforward to explain regression predictions and later evaluate prediction errors.

It has a compact feature space

Only four predictors are required. Readers can therefore focus on the modeling process without first managing hundreds of variables, images, text fields, or complex categorical encodings.

It supports many algorithms

The same dataset can be used with linear regression, polynomial regression, regularized models, support-vector regression, decision trees, random forests, gradient boosting, neural networks, symbolic regression, and other approaches. This makes comparisons easier because the underlying prediction task remains unchanged.

It is suitable for progressive learning

A beginner can start with dataset structure and a linear baseline. More advanced readers can later investigate nonlinear relationships, hyperparameter tuning, feature importance, uncertainty, interpretability, symbolic equations, and ensemble learning.

6. Position of This Post in the CCPP Tutorial Series

Roadmap showing Tutorial 01 followed by tutorials on plant operation, feature understanding, Python loading, and machine-learning problem definition.
Figure 3. Tutorial 01 provides orientation only. The engineering, data-analysis, and methodological details are intentionally separated into later posts.
Tutorial Title Primary purpose
01 Introduction to the Combined Cycle Power Plant Dataset Introduce the dataset, prediction target, and practical importance
02 How a Combined Cycle Power Plant Works Explain the energy-conversion system and major plant components
03 Understanding the CCPP Dataset Features Interpret all predictors and the electrical-output target in detail
04 Loading and Inspecting the CCPP Dataset with Python Load, validate, and inspect the table with pandas
05 Defining the CCPP Machine Learning Problem Formalize the regression task, assumptions, questions, and evaluation plan

7. What This Introduction Does—and Does Not—Cover

Covered here

  • Identity and origin of the dataset
  • Number of observations and variables
  • High-level prediction objective
  • Names and roles of the variables
  • Practical and educational importance
  • Position within the tutorial series

Reserved for later tutorials

  • Detailed power-plant operation
  • Gas and steam turbine mechanics
  • Detailed feature interpretation
  • Data ranges and distributions
  • Python loading and inspection
  • Missing values and duplicates
  • Train-test splitting and validation
  • Regression metrics and model training

Keeping these subjects separate gives every tutorial one clear learning goal. Readers can follow the series in order, while experienced practitioners can open the specific article that addresses the topic they need.

8. References

  1. Tüfekci, P., & Kaya, H. (2014). Combined Cycle Power Plant [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5002N
  2. Tüfekci, P. (2014). Prediction of full load electrical power output of a base load operated combined cycle power plant using machine learning methods. International Journal of Electrical Power & Energy Systems, 60, 126–140. https://doi.org/10.1016/j.ijepes.2014.02.027

Next tutorial

02 — How a Combined Cycle Power Plant Works
Gas turbines, steam turbines, heat-recovery steam generators, environmental conditions, and net power output.

Wednesday, July 29, 2026

Feature Importance in Random Forests: How to Interpret Your Model

Feature Importance in Random Forests: How to Interpret Your Model
Pythonholics Machine Learning Tutorial

A detailed practical guide to impurity-based feature importance, permutation importance, correlated predictors, importance stability, feature selection, and responsible interpretation using Python and scikit-learn.

Two core methods Mean decrease in impurity and permutation importance are explained and compared.
Complete Python workflow The script generates all figures, result tables, metrics, and CSV files.

Random Forests are among the most useful machine-learning algorithms for structured and tabular data. They can model nonlinear relationships, capture complex interactions, handle classification and regression problems, and often provide strong performance without requiring extensive data transformation. Nevertheless, a high-performing model is not automatically an understandable model.

After a Random Forest has been trained, an important practical question remains: which input features influence the model most strongly? Feature importance attempts to answer this question by assigning a numerical score to every predictor. These scores can help us understand the fitted model, identify potentially useful or redundant variables, communicate results, and design additional experiments.

Feature importance must still be interpreted cautiously. A high importance value does not prove that a feature causes the predicted outcome. It only indicates that the fitted model used the information associated with that feature in a particular way. Rankings can also be affected by correlated predictors, data leakage, high-cardinality variables, the selected scoring metric, sampling variation, and model configuration.

Main lesson: Random Forest feature importance is a model-inspection tool, not a causal analysis. The most reliable interpretation combines several importance methods with predictive evaluation, correlation analysis, stability checks, and domain knowledge.

1. What Does Feature Importance Mean?

A machine-learning dataset contains input variables, commonly called features, and an output that the model attempts to predict. For example, a medical classification dataset may contain measurements of a tumor, while the target indicates whether the tumor is benign or malignant. A Random Forest may use some measurements repeatedly and depend only weakly on others.

Feature importance provides a model-based estimate of this dependence. In broad terms, an important feature is one that helps the model reduce uncertainty, separate classes, lower prediction error, or maintain predictive performance. However, different feature-importance methods operationalize this idea differently.

Table 1. Practical questions feature importance can help answer
Question How feature importance helps Important limitation
Which variables does the model use most? Ranks predictors according to a defined importance measure. The ranking explains the fitted model, not necessarily the real-world mechanism.
Can the feature set be reduced? Identifies candidates for feature-selection experiments. Low individual importance does not prove that a feature is useless in interactions.
Is the model relying on suspicious information? Can expose identifiers, proxies, leakage variables, or unexpected predictors. A leakage audit still requires understanding how every variable was created.
How can the model be explained to stakeholders? Provides a compact global summary of model reliance. Global importance does not explain every individual prediction.
Which measurements deserve further study? Suggests variables or feature groups for additional analysis. Predictive association must not be presented as causation.

Global and local interpretation

Standard Random Forest feature importance is a global interpretation method. It summarizes model behavior across many observations. It does not directly explain why one specific observation received one specific prediction. Local explanations require methods such as local permutation, SHAP values, local surrogate models, or careful examination of decision paths.

Importance is method-dependent

There is no single universal definition of feature importance. A feature can be important because it frequently creates strong decision-tree splits, because shuffling it damages test performance, because removing it reduces cross-validation performance, or because it contributes strongly to individual predictions. Consequently, two valid methods can produce different rankings.

2. Why Can a Random Forest Rank Features?

A Random Forest combines many decision trees. For classification, each tree predicts a class and the forest aggregates those predictions, usually through majority voting or averaged class probabilities. For regression, the outputs of individual trees are averaged.

Randomness is introduced in two important places. First, each tree is trained on a bootstrap sample drawn from the training data. Second, only a random subset of features is considered at each candidate split. This creates diversity among trees and reduces the tendency of a single tree to overfit the training observations.

Every internal node of a classification tree selects a feature and threshold that reduce impurity. Since the forest records which features were used and how much impurity they reduced, the reductions can be aggregated into an impurity-based importance score. Alternatively, the fitted forest can be treated as a black box and evaluated by measuring how much its predictions deteriorate when individual features are randomly permuted.

Figure 1. A practical workflow for interpreting a Random Forest. Predictive performance is evaluated before impurity-based importance, permutation importance, stability, and correlation are examined.

Figure 1 emphasizes that feature importance should not be calculated in isolation. The model must first be trained using a valid experimental design and evaluated on observations that were not used to fit it. Only then should importance rankings be interpreted and validated.

3. Impurity-Based Feature Importance

The built-in importance returned by RandomForestClassifier.feature_importances_ is commonly described as impurity-based feature importance, mean decrease in impurity, or MDI. It measures the accumulated reduction in node impurity attributed to each feature across all trees.

3.1 Impurity reduction at one node

Consider a node containing a subset of training observations. A candidate split divides those observations into a left child and a right child. For a classification tree using Gini impurity, the impurity of node t can be written as:

Gini(t) = 1 − Σk=1K p(k|t)2

Here, K is the number of classes and p(k|t) is the proportion of observations belonging to class k in node t. A pure node has a Gini value of zero.

The weighted impurity decrease produced by a split can be represented conceptually as:

ΔI(t) = w(t)I(t) − w(tL)I(tL) − w(tR)I(tR)

The symbols t, tL, and tR denote the parent, left-child, and right-child nodes. The function I denotes impurity, and w represents the proportion or weighted proportion of observations reaching a node.

3.2 Aggregation across the forest

Whenever feature j is used for a split, its weighted impurity reduction is added to its total. The totals are averaged across the trees and normalized so that the final importance values sum to one:

MDI(j) = normalized sum of weighted impurity decreases produced by feature j

A feature obtains a high MDI value when it is used frequently, produces substantial impurity reductions, affects many training observations, or combines these properties.

3.3 Advantages of MDI

  • It is available immediately after the forest has been fitted.
  • It is inexpensive because no additional model evaluation is required.
  • It provides a simple ranking whose values sum to one.
  • It can be calculated separately for every tree to study stability.
  • It is useful as a fast initial overview of the forest structure.

3.4 Limitations and bias

MDI is not a neutral measure. It is computed from the same training process that built the forest, so it describes internal split behavior rather than an independently measured loss of predictive performance. It can favor continuous variables and variables with many possible split points. Correlated predictors can share importance unevenly, and a variable can look important even when its contribution is redundant.

Do not interpret the largest MDI value as proof of causality. The score only indicates that the trained forest assigned substantial split-based importance to the feature.

The clearest first visualization is a horizontal bar chart in which features are sorted by MDI. Error bars can display the standard deviation of importance across individual trees. This adds information that the single forest-level vector does not provide.

Figure 2. Top ten features according to mean decrease in impurity. Longer bars indicate a greater aggregated contribution to weighted impurity reduction. Error bars summarize variation across individual trees.

Figure 2 should be interpreted as a description of the forest's internal structure. A long bar indicates that the feature contributed strongly to splits across the ensemble. The error bar indicates whether that contribution was consistent among trees or concentrated in only part of the forest.

4. Permutation Feature Importance

Permutation importance measures the extent to which predictive performance depends on a feature. It can be applied to Random Forests and many other fitted estimators, which makes it a model-agnostic inspection method.

4.1 Core idea

First, the fitted model is evaluated on a validation or test dataset to obtain a baseline score. Next, the values of one feature are randomly shuffled across rows. This destroys the association between that feature and the target while leaving the feature's marginal distribution unchanged. The model is evaluated again, and the decrease in performance becomes the importance estimate.

PI(j) = baseline score − score after permuting feature j

The procedure is repeated several times because one random permutation can produce an unusually large or small change. The mean score decrease and standard deviation are then reported.

4.2 Step-by-step algorithm

  1. Fit the Random Forest using the training data.
  2. Evaluate the fitted forest on an independent validation or test set.
  3. Select one feature and randomly permute its values across observations.
  4. Evaluate the model using the modified dataset.
  5. Record the decrease in the selected performance score.
  6. Repeat the permutation several times and calculate the mean and standard deviation.
  7. Restore the feature and repeat the process for every remaining predictor.

4.3 Choosing the scoring metric

Permutation importance depends on the chosen score. Accuracy may be suitable for a balanced classification problem, but balanced accuracy, F1-score, ROC-AUC, average precision, or a cost-sensitive score may be more meaningful in other settings. Regression problems may use R², negative mean squared error, or negative mean absolute error.

This tutorial uses balanced accuracy. Consequently, the permutation importance of a feature is the mean decrease in balanced accuracy after that feature is shuffled.

4.4 Interpreting positive, zero, and negative values

Table 2. Interpreting permutation-importance values
Observed value General interpretation Recommended response
Large positive value Shuffling the feature substantially damages model performance. Treat the feature as important, then verify stability and correlation.
Small positive value The feature has a limited measurable contribution under the selected score. Compare uncertainty and test feature removal experimentally.
Value near zero The model can largely maintain performance when the feature is shuffled. Check whether correlated features preserve the same information.
Negative value The model performed slightly better after shuffling the feature. Consider sampling variation, noise, instability, or harmful reliance.

The mean permutation importance should be shown with an uncertainty estimate. A feature whose mean is positive but whose error bar crosses zero may not have a reliably positive contribution under the current test sample and scoring metric.

Figure 3. Permutation importance calculated on the held-out test subset. Each bar represents the mean decrease in balanced accuracy, while each error bar represents the standard deviation across repeated permutations.

Figure 3 is more directly connected to predictive performance than Figure 2. Nevertheless, it is not automatically free from interpretation problems. In particular, correlated predictors can cause permutation importance to underestimate the importance of information that is duplicated across multiple features.

5. MDI vs Permutation Importance

MDI and permutation importance answer related but different questions. MDI asks how strongly the feature contributed to impurity reduction inside the trained forest. Permutation importance asks how much test performance deteriorates when the feature's information is disrupted.

Table 3. Comparison of impurity-based and permutation importance
Property Impurity-based importance Permutation importance
Main quantity Aggregated weighted impurity reduction Decrease in a selected predictive score
Data normally used Training process and fitted tree structure Preferably validation or test data
Computational cost Low Higher because predictions are repeated
Model dependence Specific to tree-based models Applicable to many fitted models
High-cardinality bias Can be substantial Generally less direct, but not immune to dataset problems
Correlated predictors Importance may be unevenly shared or concentrated Importance may be masked by substitute predictors
Best use Fast structural overview Performance-based validation of model reliance

5.1 Why the numerical scales cannot be compared directly

MDI values are normalized to sum to one. Permutation values represent changes in a performance metric and do not generally sum to one. Therefore, a value of 0.10 in MDI is not numerically equivalent to a 0.10 decrease in balanced accuracy.

The Python script creates a normalized version of positive permutation importance only for visual comparison. This normalization makes broad patterns easier to see, but it does not transform the two measures into the same mathematical quantity.

Figure 4. Visual comparison of MDI and positive permutation importance after separate normalization. The normalization is used only to compare broad ranking patterns; it does not make the methods equivalent.

5.2 Understanding disagreement

A feature can have high MDI and low permutation importance when it is used frequently by the trees but is redundant with other predictors. When shuffled, the remaining variables preserve enough information for the model to maintain performance.

A feature can also have moderate MDI and comparatively high permutation importance. This suggests that it may not dominate the split structure, but disrupting it still damages predictive performance. Such disagreement is not necessarily an error. It is a signal that the data structure should be examined more carefully.

6. Practical Experiment with the Breast Cancer Dataset

The complete Python example uses the Breast Cancer Wisconsin diagnostic dataset distributed with scikit-learn. The problem is binary classification. Each observation contains measurements derived from a digitized image of a fine-needle aspirate of a breast mass. The target classes are malignant and benign.

Table 4. Dataset summary
Property Value
Observations569
Input features30 numerical variables
Target classes2
Malignant observations212
Benign observations357
Missing values0
Train/test split70% / 30%, stratified

6.1 Random Forest configuration

Table 5. Random Forest configuration used in the example
Hyperparameter Value Purpose
n_estimators500Builds a stable ensemble of 500 trees.
criterionginiUses Gini impurity for classification splits.
max_depthNoneAllows trees to expand until other stopping conditions apply.
min_samples_split2Minimum observations needed to split an internal node.
min_samples_leaf1Minimum observations required in a leaf.
max_featuressqrtConsiders a random square-root-sized feature subset at each split.
bootstrapTrueTrains each tree using a bootstrap sample.
oob_scoreTrueCalculates an out-of-bag performance estimate.
random_state42Makes the experiment reproducible.
n_jobs-1Uses available processor cores where supported.

6.2 Evaluate the model before interpreting it

Feature-importance rankings are meaningful only when the underlying model has learned a useful predictive relationship. A poor model can still produce rankings, but those rankings explain a model whose predictions are unreliable.

Table 6. Representative performance from the reproducible example
Metric Value
Accuracy0.9474
Balanced accuracy0.9391
Precision0.9455
Recall0.9720
F1-score0.9585
Matthews correlation coefficient0.8872
ROC-AUC0.9917
Out-of-bag score0.9673

The confusion matrix provides more detail than accuracy alone. It distinguishes correctly identified malignant cases, malignant cases incorrectly classified as benign, benign cases incorrectly classified as malignant, and correctly identified benign cases.

Figure 5. Confusion matrix for the held-out test set. In the representative run, 58 malignant and 104 benign observations were classified correctly, while six malignant and three benign observations were misclassified.

In a medical context, the two error types may have different consequences. For this reason, model evaluation should not rely on one aggregate score. The example reports several complementary metrics and retains the confusion matrix as part of the interpretation workflow.

7. How Correlated Features Affect Importance

Correlation is one of the main reasons why feature-importance rankings can be misleading. Suppose that two predictors contain nearly the same information. A tree may use the first feature in one branch and the second in another branch. Across the forest, the importance may be split between them or concentrated unevenly.

7.1 Effect on MDI

MDI may assign most of the importance to one feature from a correlated group because that feature happened to produce slightly better splits. Another nearly equivalent feature can then appear unimportant even though it contains similar predictive information.

7.2 Effect on permutation importance

Permuting one correlated feature may produce only a small performance decrease because another feature still provides similar information. As a result, the importance of both features can be underestimated when each is permuted separately.

A correlation matrix should therefore be examined alongside the importance rankings. The Breast Cancer dataset contains several radius, perimeter, area, concavity, and concave-point measurements that are strongly related.

Figure 6. Correlation among the ten highest-ranked MDI features. Strongly correlated predictors can share, redistribute, or mask importance and should often be interpreted as related feature groups.

7.3 Better approaches for correlated predictors

  • Grouped permutation: shuffle a related group of predictors together.
  • Feature clustering: cluster correlated variables and retain representative features.
  • Conditional permutation: permute a feature conditionally on related variables.
  • Domain grouping: interpret related measurements as one conceptual factor.
  • Ablation experiments: remove groups of predictors and retrain the entire pipeline.

8. Stability and Uncertainty of Feature Importance

A single importance vector can create a false impression of certainty. Random Forests are stochastic models, and individual trees use different bootstrap samples and feature subsets. Consequently, importance can vary from one tree to another and from one fitted forest to another.

8.1 Tree-to-tree variability

The importance values from all 500 individual trees can be collected and displayed as boxplots. A narrow distribution suggests that many trees assign a similar level of importance to the feature. A wide distribution indicates that the feature is highly important in some trees but less important in others.

Figure 7. Distribution of feature importance across individual trees. Wide distributions indicate that a feature's role varies substantially across bootstrap samples and random feature subsets.

8.2 Stability across repeated model fits

Tree-to-tree variability is informative, but it does not replace repeated model fitting. A stronger analysis repeats the complete training procedure using different cross-validation folds or random seeds. The resulting feature ranks can then be summarized using means, standard deviations, confidence intervals, or rank correlations.

Recommended publication practice: report importance variability across repeated cross-validation runs rather than presenting one importance ranking from one fitted model as definitive.

9. Cumulative Importance and Feature Selection

Feature importance is often used to create a reduced feature subset. A common approach sorts features from highest to lowest MDI and calculates cumulative importance. This shows how quickly the total reaches thresholds such as 80% or 90%.

Figure 8. Cumulative impurity-based feature importance. The curve indicates how many top-ranked features account for a selected proportion of total MDI, but it should not be used as an automatic deletion rule.

9.1 Why threshold selection is not enough

Keeping enough features to reach 90% cumulative MDI does not guarantee that the reduced model will retain 90% of predictive performance. Importance values are not percentages of accuracy. A low-ranked feature may provide unique information, contribute only through interactions, or become important after another feature is removed.

9.2 Correct feature-selection procedure

  1. Calculate importance using only the training process or training folds.
  2. Choose one or more candidate subsets, such as the top 5, 10, 15, and 20 features.
  3. Retrain the complete preprocessing and modeling pipeline for each subset.
  4. Evaluate every candidate using cross-validation or a nested validation design.
  5. Select the smallest subset whose performance is acceptably close to the full model.
  6. Evaluate the selected pipeline once on the untouched test set.
Avoid leakage: do not calculate importance on the complete dataset and then report cross-validation performance using the selected variables. Feature selection must be performed inside the training folds.

10. Complete Python Code

The following script generates every figure and CSV table referenced in this article. Save it as random_forest_feature_importance_tutorial.py and run it from Spyder, a terminal, or another Python environment.

10.1 Required packages

pip install numpy pandas matplotlib scikit-learn

10.2 Full reproducible script

"""
Random Forest Feature Importance Tutorial
Generates all tables and figures used in the Pythonholics article:
"Feature Importance in Random Forests: How to Interpret Your Model"

Tested workflow:
- Breast Cancer Wisconsin diagnostic dataset
- RandomForestClassifier
- Impurity-based feature importance
- Permutation feature importance
- Tree-to-tree importance stability
- Correlation analysis
- Cumulative feature importance
"""

from pathlib import Path
import warnings

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd

from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
from sklearn.metrics import (
    accuracy_score,
    balanced_accuracy_score,
    classification_report,
    confusion_matrix,
    f1_score,
    matthews_corrcoef,
    precision_score,
    recall_score,
    roc_auc_score,
)
from sklearn.model_selection import train_test_split


# =============================================================================
# 1. Configuration
# =============================================================================

RANDOM_STATE = 42
TEST_SIZE = 0.30
N_ESTIMATORS = 500
N_PERMUTATION_REPEATS = 30
TOP_N = 10

OUTPUT_DIR = Path("rf_feature_importance_outputs")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

# Times New Roman is used when it is installed on the computer.
# Matplotlib will use a fallback font when it is unavailable.
plt.rcParams["font.family"] = "Times New Roman"
plt.rcParams["font.size"] = 12
plt.rcParams["axes.titlesize"] = 15
plt.rcParams["axes.labelsize"] = 13
plt.rcParams["xtick.labelsize"] = 11
plt.rcParams["ytick.labelsize"] = 11
plt.rcParams["legend.fontsize"] = 11
plt.rcParams["figure.titlesize"] = 15

warnings.filterwarnings("ignore", category=UserWarning)


def save_figure(filename: str) -> None:
    """Save the active Matplotlib figure and close it."""
    plt.tight_layout()
    plt.savefig(
        OUTPUT_DIR / filename,
        dpi=300,
        bbox_inches="tight",
    )
    plt.close()


# =============================================================================
# 2. Load and inspect the dataset
# =============================================================================

dataset = load_breast_cancer()

X = pd.DataFrame(
    dataset.data,
    columns=dataset.feature_names,
)
y = pd.Series(
    dataset.target,
    name="target",
)

dataset_summary = pd.DataFrame(
    {
        "property": [
            "Number of observations",
            "Number of input features",
            "Number of target classes",
            "Class 0 observations",
            "Class 1 observations",
            "Missing values",
        ],
        "value": [
            X.shape[0],
            X.shape[1],
            y.nunique(),
            int((y == 0).sum()),
            int((y == 1).sum()),
            int(X.isna().sum().sum()),
        ],
    }
)

dataset_summary.to_csv(
    OUTPUT_DIR / "table_01_dataset_summary.csv",
    index=False,
)


# =============================================================================
# 3. Create a stratified training/test split
# =============================================================================

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=TEST_SIZE,
    random_state=RANDOM_STATE,
    stratify=y,
)


# =============================================================================
# 4. Train the Random Forest classifier
# =============================================================================

model = RandomForestClassifier(
    n_estimators=N_ESTIMATORS,
    criterion="gini",
    max_depth=None,
    min_samples_split=2,
    min_samples_leaf=1,
    max_features="sqrt",
    bootstrap=True,
    oob_score=True,
    class_weight=None,
    random_state=RANDOM_STATE,
    n_jobs=-1,
)

model.fit(X_train, y_train)


# =============================================================================
# 5. Evaluate predictive performance
# =============================================================================

y_pred = model.predict(X_test)
y_probability = model.predict_proba(X_test)[:, 1]

performance = pd.DataFrame(
    {
        "metric": [
            "Accuracy",
            "Balanced accuracy",
            "Precision",
            "Recall",
            "F1-score",
            "Matthews correlation coefficient",
            "ROC-AUC",
            "Out-of-bag score",
        ],
        "value": [
            accuracy_score(y_test, y_pred),
            balanced_accuracy_score(y_test, y_pred),
            precision_score(y_test, y_pred),
            recall_score(y_test, y_pred),
            f1_score(y_test, y_pred),
            matthews_corrcoef(y_test, y_pred),
            roc_auc_score(y_test, y_probability),
            model.oob_score_,
        ],
    }
)

performance.to_csv(
    OUTPUT_DIR / "table_02_model_performance.csv",
    index=False,
)

classification_report_table = pd.DataFrame(
    classification_report(
        y_test,
        y_pred,
        target_names=dataset.target_names,
        output_dict=True,
        zero_division=0,
    )
).transpose()

classification_report_table.to_csv(
    OUTPUT_DIR / "table_03_classification_report.csv"
)

cm = confusion_matrix(y_test, y_pred)

confusion_matrix_table = pd.DataFrame(
    cm,
    index=["Actual malignant", "Actual benign"],
    columns=["Predicted malignant", "Predicted benign"],
)

confusion_matrix_table.to_csv(
    OUTPUT_DIR / "table_04_confusion_matrix.csv"
)


# =============================================================================
# 6. Calculate impurity-based feature importance
# =============================================================================

# Importance produced by the complete forest.
mdi_table = pd.DataFrame(
    {
        "feature": X.columns,
        "mdi_importance": model.feature_importances_,
    }
)

# Importance values from individual trees allow us to estimate stability.
tree_importances = np.array(
    [tree.feature_importances_ for tree in model.estimators_]
)

mdi_table["tree_importance_std"] = tree_importances.std(axis=0)

mdi_table = (
    mdi_table
    .sort_values("mdi_importance", ascending=False)
    .reset_index(drop=True)
)

mdi_table["mdi_rank"] = np.arange(1, len(mdi_table) + 1)
mdi_table["cumulative_mdi"] = mdi_table["mdi_importance"].cumsum()

mdi_table.to_csv(
    OUTPUT_DIR / "table_05_mdi_feature_importance.csv",
    index=False,
)


# =============================================================================
# 7. Calculate permutation feature importance
# =============================================================================

permutation_result = permutation_importance(
    model,
    X_test,
    y_test,
    scoring="balanced_accuracy",
    n_repeats=N_PERMUTATION_REPEATS,
    random_state=RANDOM_STATE,
    n_jobs=-1,
)

permutation_table = pd.DataFrame(
    {
        "feature": X.columns,
        "permutation_mean": permutation_result.importances_mean,
        "permutation_std": permutation_result.importances_std,
    }
)

permutation_table = (
    permutation_table
    .sort_values("permutation_mean", ascending=False)
    .reset_index(drop=True)
)

permutation_table["permutation_rank"] = np.arange(
    1,
    len(permutation_table) + 1,
)

permutation_table.to_csv(
    OUTPUT_DIR / "table_06_permutation_feature_importance.csv",
    index=False,
)


# =============================================================================
# 8. Create a combined comparison table
# =============================================================================

comparison_table = pd.merge(
    mdi_table,
    permutation_table,
    on="feature",
    how="inner",
)

comparison_table["absolute_rank_difference"] = (
    comparison_table["mdi_rank"]
    - comparison_table["permutation_rank"]
).abs()

# Normalization below is used only to make the two methods easier to compare
# visually. It does not make the methods mathematically equivalent.
positive_permutation = comparison_table["permutation_mean"].clip(lower=0)

if positive_permutation.sum() > 0:
    comparison_table["normalized_positive_permutation"] = (
        positive_permutation / positive_permutation.sum()
    )
else:
    comparison_table["normalized_positive_permutation"] = 0.0

comparison_table.to_csv(
    OUTPUT_DIR / "table_07_importance_comparison.csv",
    index=False,
)


# =============================================================================
# 9. Figure 1: Conceptual Random Forest interpretation workflow
# =============================================================================

fig, ax = plt.subplots(figsize=(12, 4.8))
ax.axis("off")

workflow_nodes = [
    (0.10, "Training data"),
    (0.32, "Random Forest\nclassifier"),
    (0.55, "Model\nperformance"),
    (0.76, "Feature-importance\nmethods"),
    (0.94, "Interpretation and\nvalidation"),
]

for x_position, label in workflow_nodes:
    ax.text(
        x_position,
        0.52,
        label,
        ha="center",
        va="center",
        transform=ax.transAxes,
        bbox={
            "boxstyle": "round,pad=0.7",
            "facecolor": "white",
            "edgecolor": "black",
        },
    )

for index in range(len(workflow_nodes) - 1):
    x_start = workflow_nodes[index][0] + 0.07
    x_end = workflow_nodes[index + 1][0] - 0.08

    ax.annotate(
        "",
        xy=(x_end, 0.52),
        xytext=(x_start, 0.52),
        xycoords=ax.transAxes,
        textcoords=ax.transAxes,
        arrowprops={"arrowstyle": "->", "linewidth": 1.5},
    )

ax.text(
    0.76,
    0.23,
    "MDI  |  permutation  |  stability  |  correlation",
    ha="center",
    va="center",
    transform=ax.transAxes,
)

ax.set_title("A Practical Workflow for Interpreting Random Forest Feature Importance")
save_figure("figure_01_random_forest_interpretation_workflow.png")


# =============================================================================
# 10. Figure 2: Top impurity-based feature importances
# =============================================================================

mdi_top = mdi_table.head(TOP_N).sort_values(
    "mdi_importance",
    ascending=True,
)

plt.figure(figsize=(11, 7))
plt.barh(
    mdi_top["feature"],
    mdi_top["mdi_importance"],
    xerr=mdi_top["tree_importance_std"],
    capsize=3,
)
plt.xlabel("Mean decrease in impurity")
plt.ylabel("Feature")
plt.title("Top 10 Features by Impurity-Based Importance")
save_figure("figure_02_mdi_feature_importance.png")


# =============================================================================
# 11. Figure 3: Top permutation importances
# =============================================================================

permutation_top = permutation_table.head(TOP_N).sort_values(
    "permutation_mean",
    ascending=True,
)

plt.figure(figsize=(11, 7))
plt.barh(
    permutation_top["feature"],
    permutation_top["permutation_mean"],
    xerr=permutation_top["permutation_std"],
    capsize=3,
)
plt.axvline(0, linewidth=1)
plt.xlabel("Mean decrease in balanced accuracy")
plt.ylabel("Feature")
plt.title("Top 10 Features by Permutation Importance")
save_figure("figure_03_permutation_feature_importance.png")


# =============================================================================
# 12. Figure 4: Normalized MDI/permutation comparison
# =============================================================================

comparison_top_features = (
    comparison_table
    .sort_values("mdi_importance", ascending=False)
    .head(TOP_N)
    .copy()
)

plot_positions = np.arange(len(comparison_top_features))
bar_width = 0.38

plt.figure(figsize=(13, 7))
plt.bar(
    plot_positions - bar_width / 2,
    comparison_top_features["mdi_importance"],
    width=bar_width,
    label="Normalized MDI",
)
plt.bar(
    plot_positions + bar_width / 2,
    comparison_top_features["normalized_positive_permutation"],
    width=bar_width,
    label="Normalized positive permutation importance",
)
plt.xticks(
    plot_positions,
    comparison_top_features["feature"],
    rotation=75,
    ha="right",
)
plt.ylabel("Normalized importance used for visual comparison")
plt.xlabel("Feature")
plt.title("Comparison of MDI and Permutation Feature Importance")
plt.legend()
save_figure("figure_04_mdi_vs_permutation_comparison.png")


# =============================================================================
# 13. Figure 5: Confusion matrix
# =============================================================================

plt.figure(figsize=(7, 6))
plt.imshow(cm, interpolation="nearest", aspect="auto")
plt.colorbar()

class_labels = ["Malignant", "Benign"]
plt.xticks([0, 1], class_labels)
plt.yticks([0, 1], class_labels)
plt.xlabel("Predicted class")
plt.ylabel("True class")
plt.title("Random Forest Confusion Matrix")

threshold = cm.max() / 2

for row in range(cm.shape[0]):
    for column in range(cm.shape[1]):
        plt.text(
            column,
            row,
            str(cm[row, column]),
            ha="center",
            va="center",
        )

save_figure("figure_05_confusion_matrix.png")


# =============================================================================
# 14. Figure 6: Correlation heatmap for the top MDI features
# =============================================================================

top_feature_names = mdi_table.head(TOP_N)["feature"].tolist()
correlation_matrix = X[top_feature_names].corr(method="pearson")

plt.figure(figsize=(11, 9))
image = plt.imshow(
    correlation_matrix,
    interpolation="nearest",
    aspect="auto",
    vmin=-1,
    vmax=1,
)
plt.colorbar(image, label="Pearson correlation coefficient")
plt.xticks(
    np.arange(len(top_feature_names)),
    top_feature_names,
    rotation=75,
    ha="right",
)
plt.yticks(
    np.arange(len(top_feature_names)),
    top_feature_names,
)
plt.title("Correlation Among the Top 10 MDI Features")
save_figure("figure_06_top_feature_correlation_heatmap.png")


# =============================================================================
# 15. Figure 7: Tree-to-tree importance stability
# =============================================================================

top_indices = [
    X.columns.get_loc(feature)
    for feature in mdi_table.head(TOP_N)["feature"]
]

tree_importance_top = tree_importances[:, top_indices]
tree_importance_labels = mdi_table.head(TOP_N)["feature"].tolist()

plt.figure(figsize=(13, 7))
plt.boxplot(
    tree_importance_top,
    tick_labels=tree_importance_labels,
    showfliers=False,
)
plt.xticks(rotation=75, ha="right")
plt.xlabel("Feature")
plt.ylabel("Importance across individual trees")
plt.title("Tree-to-Tree Variability of the Top Feature Importances")
save_figure("figure_07_tree_importance_stability.png")


# =============================================================================
# 16. Figure 8: Cumulative MDI importance
# =============================================================================

feature_count = np.arange(1, len(mdi_table) + 1)

plt.figure(figsize=(10, 6))
plt.plot(
    feature_count,
    mdi_table["cumulative_mdi"],
    marker="o",
)
plt.axhline(0.80, linestyle="--", label="80% importance")
plt.axhline(0.90, linestyle=":", label="90% importance")
plt.xlabel("Number of features included")
plt.ylabel("Cumulative MDI importance")
plt.title("Cumulative Impurity-Based Feature Importance")
plt.grid(True, alpha=0.3)
plt.legend()
save_figure("figure_08_cumulative_feature_importance.png")


# =============================================================================
# 17. Determine the numbers of features required for thresholds
# =============================================================================

features_for_80 = int(
    np.argmax(mdi_table["cumulative_mdi"].to_numpy() >= 0.80) + 1
)

features_for_90 = int(
    np.argmax(mdi_table["cumulative_mdi"].to_numpy() >= 0.90) + 1
)

threshold_summary = pd.DataFrame(
    {
        "cumulative_importance_threshold": [0.80, 0.90],
        "number_of_features_required": [features_for_80, features_for_90],
    }
)

threshold_summary.to_csv(
    OUTPUT_DIR / "table_08_cumulative_importance_thresholds.csv",
    index=False,
)


# =============================================================================
# 18. Print the main results
# =============================================================================

print("\nDATASET SUMMARY")
print(dataset_summary.to_string(index=False))

print("\nMODEL PERFORMANCE")
print(performance.to_string(index=False))

print("\nCONFUSION MATRIX")
print(confusion_matrix_table)

print("\nTOP 10 MDI FEATURES")
print(mdi_table.head(10).to_string(index=False))

print("\nTOP 10 PERMUTATION FEATURES")
print(permutation_table.head(10).to_string(index=False))

print("\nCUMULATIVE IMPORTANCE THRESHOLDS")
print(threshold_summary.to_string(index=False))

print(f"\nAll tables and figures were saved to:\n{OUTPUT_DIR.resolve()}")

10.3 Generated output files

Table 7. Files generated by the Python script
File Content
table_01_dataset_summary.csvDataset dimensions, class counts, and missing-value count.
table_02_model_performance.csvAccuracy, balanced accuracy, precision, recall, F1, MCC, ROC-AUC, and OOB score.
table_03_classification_report.csvClass-specific precision, recall, F1-score, and support.
table_04_confusion_matrix.csvNumerical confusion matrix.
table_05_mdi_feature_importance.csvMDI values, tree-level standard deviations, ranks, and cumulative MDI.
table_06_permutation_feature_importance.csvMean and standard deviation of repeated permutation importance.
table_07_importance_comparison.csvMerged rankings and normalized values used for comparison.
table_08_cumulative_importance_thresholds.csvNumbers of features required to reach 80% and 90% cumulative MDI.
figure_01_random_forest_interpretation_workflow.pngInterpretation workflow diagram.
figure_02_mdi_feature_importance.pngTop MDI importance chart.
figure_03_permutation_feature_importance.pngTop permutation importance chart.
figure_04_mdi_vs_permutation_comparison.pngNormalized comparison chart.
figure_05_confusion_matrix.pngConfusion matrix visualization.
figure_06_top_feature_correlation_heatmap.pngTop-feature correlation heatmap.
figure_07_tree_importance_stability.pngTree-to-tree importance distributions.
figure_08_cumulative_feature_importance.pngCumulative MDI curve.

11. How to Interpret the Example Results

11.1 Top impurity-based features

In the representative run, the highest MDI values were associated with measurements such as worst concave points, worst perimeter, worst area, worst radius, and mean concave points. These variables describe characteristics of the cell nuclei and provide strong splitting opportunities for the fitted forest.

Table 8. Top ten MDI features from the representative run
Rank Feature MDI importance
1worst concave points0.136424
2worst perimeter0.132857
3worst area0.131652
4worst radius0.088794
5mean concave points0.080803
6mean radius0.063168
7mean perimeter0.053367
8mean area0.043312
9mean concavity0.042598
10worst concavity0.036355

11.2 Top permutation features

The permutation ranking differs considerably. Worst texture and mean texture appear at the top, while several highly ranked MDI variables have smaller permutation values. This is an excellent example of why one method should not be interpreted alone. Several geometric measurements are strongly correlated, so shuffling one of them may not damage performance substantially because related features remain available.

Table 9. Top ten permutation features from the representative run
Rank Feature Mean decrease in balanced accuracy Standard deviation
1worst texture0.0098130.003271
2mean texture0.0057630.003119
3mean concavity0.0048290.003301
4worst radius0.0044610.005277
5mean concave points0.0042060.003486
6worst concavity0.0032710.004849
7area error0.0031150.003267
8radius error0.0023360.002893
9worst smoothness0.0015580.002203
10worst perimeter0.0013410.005475
Interpretation: the disagreement does not mean that one method is wrong. MDI emphasizes the features used to build strong splits, while permutation importance measures additional performance loss after disrupting one feature while all correlated substitutes remain available.

11.3 Why permutation values may be small

Small permutation values do not necessarily mean that the model contains no useful features. The forest can distribute predictive information across many related variables. If several predictors are substitutes, shuffling one variable at a time may cause only a modest score reduction. This is precisely why correlation analysis, grouped permutation, and feature-group ablation can be valuable.

12. Common Interpretation Mistakes

12.1 Treating importance as causality

Predictive importance is not a causal effect. A feature may be a proxy for another variable, reflect a selection process, or become important because of leakage.

12.2 Interpreting only the highest bar

The difference between the first and second feature may be unstable. Features should be interpreted in groups, with uncertainty estimates and repeated fitting.

12.3 Ignoring correlated predictors

Correlation can redistribute MDI and mask permutation importance. Always inspect relationships among the highest-ranked predictors.

12.4 Calculating permutation importance on training data

Training-set permutation importance can reflect overfitting. Validation or test data provide a more meaningful estimate of model reliance on unseen observations.

12.5 Selecting features before splitting the data

Using the complete dataset to select features leaks information into model evaluation. Importance-based selection belongs inside cross-validation or the training portion of the experiment.

12.6 Assuming importance reveals direction

A positive importance score does not tell us whether higher feature values increase or decrease the predicted probability. Partial dependence, accumulated local effects, SHAP dependence plots, or targeted conditional analysis are needed to study direction.

12.7 Ignoring class imbalance and the scoring metric

Permutation rankings can change when accuracy is replaced by balanced accuracy, recall, F1-score, ROC-AUC, or a domain-specific cost function. The score should match the actual objective of the model.

12.8 Removing all low-ranked features automatically

A feature can have low marginal importance while contributing through interactions. Removal decisions must be validated by retraining and evaluating the complete model.

13. Advanced Methods for Deeper Interpretation

13.1 Partial dependence plots

Partial dependence plots estimate the average predicted response as one or two features vary. They help show whether the model response is increasing, decreasing, nonlinear, or threshold-like. They can be misleading when the plotted feature is strongly correlated with other predictors.

13.2 Individual conditional expectation

Individual conditional expectation curves show how predictions change for individual observations rather than only displaying an average. They can reveal heterogeneous effects that a partial dependence curve hides.

13.3 Accumulated local effects

Accumulated local effects plots are designed to reduce some problems caused by correlated features. They estimate local changes in predictions within regions supported by the observed data.

13.4 SHAP values

SHAP-based methods assign contributions to features for individual predictions and can be aggregated into global summaries. They provide richer local explanations but require additional assumptions, computation, and careful treatment of dependent predictors.

13.5 Drop-column importance

Drop-column importance removes one feature, retrains the model, and measures the resulting performance change. It can be more faithful to the retrained modeling process than simple permutation, but it is computationally expensive because the model must be refitted for every feature or feature group.

13.6 Grouped feature importance

When several variables describe the same concept, grouped permutation or grouped removal may be more meaningful than ranking each column independently. For example, radius, perimeter, and area measurements might be analyzed as a geometric feature group.

Table 10. Choosing an interpretation method
Goal Suitable method Main caution
Fast global overview MDI Training-based bias and correlated predictors
Performance-based global reliance Permutation importance Depends on scoring metric and correlation
Importance after retraining Drop-column importance High computational cost
Average response shape Partial dependence Can extrapolate into unrealistic combinations
Observation-specific contribution SHAP or local explanation Requires careful background and dependence assumptions
Related predictors Grouped permutation or grouped ablation Feature groups require domain justification

14. Random Forest Regression

The same workflow applies to RandomForestRegressor. The primary difference is the impurity criterion and the evaluation metric. Instead of Gini impurity, regression trees typically reduce squared error, absolute error, or another regression criterion.

Permutation importance should then be calculated using a regression score such as R², negative mean squared error, or negative mean absolute error. The interpretation remains the same: a feature is important when disrupting its information causes the selected predictive score to deteriorate.

15. Best-Practice Checklist

  • Evaluate predictive performance before interpreting feature importance.
  • Use a held-out validation or test subset for permutation importance.
  • Select a permutation scoring metric that matches the real modeling objective.
  • Compare MDI with at least one performance-based method.
  • Inspect correlations among highly ranked predictors.
  • Report standard deviations or repeated-fit variability.
  • Interpret groups of related predictors rather than overemphasizing tiny rank differences.
  • Perform feature selection inside cross-validation to prevent leakage.
  • Retrain and reevaluate the model after removing features.
  • Use domain knowledge to judge whether the ranking is plausible.
  • Do not claim that predictive importance proves causality.
  • Use local explanation methods when individual predictions must be explained.

16. Conclusion

Feature importance makes Random Forests easier to inspect, but it does not provide one final and unquestionable explanation. Impurity-based importance is fast and reveals how features contribute to the structure of the forest. Permutation importance measures how strongly predictive performance depends on preserving the information in each feature.

The two methods can disagree because they measure different properties. Correlated predictors, redundant information, model instability, scoring choices, and sampling variation all affect the resulting rankings. A responsible analysis therefore combines importance measures with performance evaluation, correlation analysis, uncertainty estimates, repeated model fitting, and domain knowledge.

Used carefully, Random Forest feature importance can help identify meaningful predictors, detect suspicious dependencies, guide feature-selection experiments, and communicate model behavior. Used carelessly, it can produce confident but misleading conclusions. The objective is not simply to generate a ranked bar chart; it is to understand what the ranking measures, why it may change, and how it should influence the next modeling decision.

17. References

  1. Breiman, L. (2001). Random Forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324
  2. Louppe, G., Wehenkel, L., Sutera, A., and Geurts, P. (2013). Understanding Variable Importances in Forests of Randomized Trees. Advances in Neural Information Processing Systems, 26.
  3. Strobl, C., Boulesteix, A.-L., Zeileis, A., and Hothorn, T. (2007). Bias in Random Forest Variable Importance Measures: Illustrations, Sources and a Solution. BMC Bioinformatics, 8, 25. https://doi.org/10.1186/1471-2105-8-25
  4. Altmann, A., Toloşi, L., Sander, O., and Lengauer, T. (2010). Permutation Importance: A Corrected Feature Importance Measure. Bioinformatics, 26(10), 1340–1347. https://doi.org/10.1093/bioinformatics/btq134
  5. Fisher, A., Rudin, C., and Dominici, F. (2019). All Models Are Wrong, but Many Are Useful: Learning a Variable's Importance by Studying an Entire Class of Prediction Models Simultaneously. Journal of Machine Learning Research, 20(177), 1–81.
  6. Hooker, G., Mentch, L., and Zhou, S. (2021). Unrestricted Permutation Forces Extrapolation: Variable Importance Requires at Least One More Model, or There Is No Free Variable Importance. Statistics and Computing, 31, 82. https://doi.org/10.1007/s11222-021-10057-z
  7. Pedregosa, F., Varoquaux, G., Gramfort, A., et al. (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
  8. Lundberg, S. M., and Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems, 30.
  9. Molnar, C. (2022). Interpretable Machine Learning, second edition.

Continue Learning

This article belongs to the Pythonholics Tree-Based Models roadmap. The recommended next tutorial is Gradient Boosting Machines Explained: A Practical Introduction.