| Type: | Package |
| Title: | Partial Dependence Plots |
| Version: | 0.10.0 |
| Description: | A general framework for constructing partial dependence (i.e., marginal effect) plots from various types of machine learning models in R. |
| License: | GPL-2 | GPL-3 [expanded from: GPL (≥ 2)] |
| URL: | https://github.com/bgreenwell/pdp, https://bgreenwell.github.io/pdp/, https://bgreenwell.r-universe.dev/pdp |
| BugReports: | https://github.com/bgreenwell/pdp/issues |
| Depends: | R (≥ 3.6.0) |
| Suggests: | C50, caret, covr, doParallel, e1071, earth, foreach, gbm, gridExtra, ICEbox, kernlab, knitr, magrittr, MASS, Matrix, mlbench, nnet, party, randomForest, ranger, rmarkdown, rpart, tinytest, xgboost (≥ 3.1.3.1) |
| Imports: | graphics, grDevices, lattice, methods, stats, tinyplot, utils |
| LazyData: | TRUE |
| RoxygenNote: | 7.3.3 |
| Encoding: | UTF-8 |
| VignetteBuilder: | knitr |
| NeedsCompilation: | yes |
| Packaged: | 2026-10-08 16:18:09 UTC; greenwbm |
| Author: | Brandon M. Greenwell
|
| Maintainer: | Brandon M. Greenwell <greenwell.brandon@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-08 17:10:02 UTC |
pdp: A general framework for constructing partial dependence (i.e., marginal effect) plots from various types of machine learning models in R.
Description
Partial dependence plots (PDPs) help visualize the relationship between a subset of the features (typically 1-3) and the response while accounting for the average effect of the other predictors in the model. They are particularly effective with black box models like random forests and support vector machines.
Details
The development version can be found on GitHub: https://github.com/bgreenwell/pdp. As of right now, pdp exports the following functions:
-
partial()- construct partial dependence functions (i.e., objects of class"partial") from various fitted model objects; -
plot()- plot partial dependence functions (i.e., objects of class"partial") using lightweight base R graphics viatinyplot::tinyplot()(or lattice graphics wheneverlattice = TRUE); -
exemplar()- construct a single "exemplar" record from a data frame.
Author(s)
Maintainer: Brandon M. Greenwell greenwell.brandon@gmail.com (ORCID)
See Also
Useful links:
Report bugs at https://github.com/bgreenwell/pdp/issues
Boston Housing Data
Description
Data on median housing values from 506 census tracts in the suburbs of Boston
from the 1970 census. This data frame is a corrected version of the original
data by Harrison and Rubinfeld (1978) with additional spatial information.
The data were taken directly from mlbench::BostonHousing2() and
unneeded columns (i.e., name of town, census tract, and the uncorrected
median home value) were removed.
Usage
data(boston)
Format
A data frame with 506 rows and 16 variables.
-
lonLongitude of census tract. -
latLatitude of census tract. -
cmedvCorrected median value of owner-occupied homes in USD 1000's -
crimPer capita crime rate by town. -
znProportion of residential land zoned for lots over 25,000 sq.ft. -
indusProportion of non-retail business acres per town. -
chasCharles River dummy variable (= 1 if tract bounds river; 0 otherwise). -
noxNitric oxides concentration (parts per 10 million). -
rmAverage number of rooms per dwelling. -
ageProportion of owner-occupied units built prior to 1940. -
disWeighted distances to five Boston employment centers. -
radIndex of accessibility to radial highways. -
taxFull-value property-tax rate per USD 10,000. -
ptratioPupil-teacher ratio by town. -
b$1000(B - 0.63)^2$ where B is the proportion of blacks by town. -
lstatPercentage of lower status of the population.
References
Harrison, D. and Rubinfeld, D.L. (1978). Hedonic prices and the demand for clean air. Journal of Environmental Economics and Management, 5, 81-102.
Gilley, O.W., and R. Kelley Pace (1996). On the Harrison and Rubinfeld Data. Journal of Environmental Economics and Management, 31, 403-405.
Newman, D.J. & Hettich, S. & Blake, C.L. & Merz, C.J. (1998). UCI Repository of machine learning databases https://archive.ics.uci.edu/. Irvine, CA: University of California, Department of Information and Computer Science.
Pace, R. Kelley, and O.W. Gilley (1997). Using the Spatial Configuration of the Data to Improve Estimation. Journal of the Real Estate Finance and Economics, 14, 333-340.
Friedrich Leisch & Evgenia Dimitriadou (2010). mlbench: Machine Learning Benchmark Problems. R package version 2.1-1.
Examples
head(boston)
Exemplar observation
Description
Construct a single "exemplar" record from a data frame. For now, all numeric columns (including Date objects) are replaced with their corresponding median value and non-numeric columns are replaced with their most frequent value.
Usage
exemplar(object, ...)
## S3 method for class 'data.frame'
exemplar(object, ...)
## S3 method for class 'matrix'
exemplar(object, cats = NULL, ...)
## S3 method for class 'dgCMatrix'
exemplar(object, cats = NULL, ...)
Arguments
object |
A data frame, matrix, or
|
... |
Additional optional arguments (currently ignored). |
cats |
Character string indicating which columns of |
Value
A data frame with the same number of columns as object and a
single row.
Examples
set.seed(1554) # for reproducibility
train <- data.frame(
x = rnorm(100),
y = sample(letters[1L:3L], size = 100, replace = TRUE,
prob = c(0.1, 0.1, 0.8))
)
exemplar(train)
Partial Dependence Functions
Description
Compute partial dependence functions (i.e., marginal effects) for various model fitting objects.
Usage
partial(object, ...)
## Default S3 method:
partial(
object,
pred.var,
pred.grid,
pred.fun = NULL,
grid.resolution = NULL,
ice = FALSE,
center = FALSE,
approx = FALSE,
quantiles = FALSE,
probs = 1:9/10,
trim.outliers = FALSE,
type = c("auto", "regression", "classification"),
inv.link = NULL,
which.class = 1L,
prob = FALSE,
recursive = TRUE,
plot = FALSE,
plot.engine = c("tinyplot", "lattice"),
smooth = FALSE,
rug = FALSE,
chull = FALSE,
levelplot = TRUE,
contour = FALSE,
contour.color = "white",
alpha = 1,
train,
cats = NULL,
check.class = TRUE,
batch.size = NULL,
progress = FALSE,
parallel = FALSE,
paropts = NULL,
frac = 1,
...
)
## S3 method for class 'model_fit'
partial(object, ...)
Arguments
object |
A fitted model object of appropriate class (e.g., |
... |
Additional optional arguments to be passed onto
|
pred.var |
Character string giving the names of the predictor variables
of interest. For reasons of computation/interpretation, this should include
no more than three variables. Can be omitted whenever |
pred.grid |
Data frame containing the joint values of interest for the
variables listed in |
pred.fun |
Optional prediction function that requires two arguments:
|
grid.resolution |
Integer giving the number of equally spaced points to
use for the continuous variables listed in |
ice |
Logical indicating whether or not to compute individual
conditional expectation (ICE) curves. Default is |
center |
Logical indicating whether or not to produce centered ICE
curves (c-ICE curves). Only used when |
approx |
Logical indicating whether or not to compute a faster, but
approximate, marginal effect plot (similar in spirit to the
plotmo package). If |
quantiles |
Logical indicating whether or not to use the sample
quantiles of the continuous predictors listed in |
probs |
Numeric vector of probabilities with values in |
trim.outliers |
Logical indicating whether or not to trim off outliers
from the continuous predictors listed in |
type |
Character string specifying the type of supervised learning.
Current options are |
inv.link |
Function specifying the transformation to be applied to the
predictions before the partial dependence function is computed
(experimental). Default is |
which.class |
Integer specifying which column of the matrix of predicted
probabilities to use as the "focus" class. Default is to use the first class.
Only used for classification problems (i.e., when
|
prob |
Logical indicating whether or not partial dependence for
classification problems should be returned on the probability scale, rather
than the centered logit. If |
recursive |
Logical indicating whether or not to use the weighted tree
traversal method described in Friedman (2001). This only applies to objects
that inherit from class |
plot |
Logical indicating whether to return a data frame containing the
partial dependence values ( |
plot.engine |
Character string specifying which plotting engine to use
whenever |
smooth |
Logical indicating whether or not to overlay a LOESS smooth.
Default is |
rug |
Logical indicating whether or not to include a rug display on the
predictor axes. The tick marks indicate the min/max and deciles of the
predictor distributions. This helps reduce the risk of interpreting the
partial dependence plot outside the region of the data (i.e., extrapolating).
Only used when |
chull |
Logical indicating whether or not to restrict the values of the
first two variables in |
levelplot |
Logical indicating whether or not to use a false color level
plot ( |
contour |
Logical indicating whether or not to add contour lines to the
level plot. Only used when |
contour.color |
Character string specifying the color to use for the
contour lines when |
alpha |
Numeric value in |
train |
An optional data frame, matrix, or sparse matrix containing the
original training data. This may be required depending on the class of
|
cats |
Character string indicating which columns of |
check.class |
Logical indicating whether or not to make sure each column
in |
batch.size |
Optional positive integer specifying the (approximate)
maximum number of rows to score per call to |
progress |
Logical indicating whether or not to display a text-based
progress bar. Default is |
parallel |
Logical indicating whether or not to run |
paropts |
List containing additional options to be passed onto
|
frac |
Numeric value in (0, 1] specifying the fraction of the training
data to randomly sample (without replacement) before computing the partial
dependence function. Default is |
Value
By default, partial returns an object of class
c("data.frame", "partial"). If ice = TRUE and
center = FALSE then an object of class c("data.frame", "ice")
is returned. If ice = TRUE and center = TRUE then an object of
class c("data.frame", "cice") is returned. These three classes
determine the behavior of the plotting functions that are automatically
called whenever plot = TRUE. Specifically, when plot = TRUE
and plot.engine = "tinyplot" (the default), the plot is drawn
directly (as a side effect) and the data frame of partial dependence values
is returned invisibly. When plot = TRUE and
plot.engine = "lattice", a "trellis" object is returned (see
lattice for details); the "trellis" object
will also include an additional attribute, "partial.data", containing
the data displayed in the plot.
Note
In some cases it is difficult for partial to extract the original
training data from object. In these cases an error message is
displayed requesting the user to supply the training data via the
train argument in the call to partial. In most cases where
partial can extract the required training data from object,
it is taken from the same environment in which partial is called.
Therefore, it is important to not change the training data used to construct
object before calling partial. This problem is completely
avoided when the training data are passed to the train argument in the
call to partial.
It is recommended to call partial with plot = FALSE and store
the results. This allows for more flexible plotting, and the user will not
have to waste time calling partial again if the default plot is not
sufficient.
It is possible to retrieve the last printed "trellis" object, such as
those produced by plotPartial, using trellis.last.object().
If ice = TRUE or the prediction function given to pred.fun
returns a prediction for each observation in newdata, then the result
will be a curve for each observation. These are called individual conditional
expectation (ICE) curves; see Goldstein et al. (2015) and
ICEbox::ice() for details.
References
J. H. Friedman. Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29: 1189-1232, 2001.
Goldstein, A., Kapelner, A., Bleich, J., and Pitkin, E., Peeking Inside the Black Box: Visualizing Statistical Learning With Plots of Individual Conditional Expectation. (2014) Journal of Computational and Graphical Statistics, 24(1): 44-65, 2015.
Examples
## Not run:
#
# Regression example (requires randomForest package to run)
#
# Fit a random forest to the boston housing data
library(randomForest)
data (boston) # load the boston housing data
set.seed(101) # for reproducibility
boston.rf <- randomForest(cmedv ~ ., data = boston)
# Using randomForest's partialPlot function
partialPlot(boston.rf, pred.data = boston, x.var = "lstat")
# Using pdp's partial function
head(partial(boston.rf, pred.var = "lstat")) # returns a data frame
partial(boston.rf, pred.var = "lstat", plot = TRUE, rug = TRUE)
# The partial function allows for multiple predictors
partial(boston.rf, pred.var = c("lstat", "rm"), grid.resolution = 40,
plot = TRUE, chull = TRUE, progress = TRUE)
# The plot method produces lightweight base R graphics via the tinyplot
# package by default; set `lattice = TRUE` for lattice graphics (e.g., for
# 3-D surfaces or paneled three-predictor displays)
pd <- partial(boston.rf, pred.var = c("lstat", "rm"), grid.resolution = 40)
plot(pd, contour = TRUE)
plot(pd, lattice = TRUE, levelplot = FALSE, zlab = "cmedv", drape = TRUE,
colorkey = FALSE, screen = list(z = -20, x = -60))
#
# Individual conditional expectation (ICE) curves
#
# Use partial to obtain ICE/c-ICE curves
rm.ice <- partial(boston.rf, pred.var = "rm", ice = TRUE)
plot(rm.ice, rug = TRUE, train = boston, alpha = 0.2)
plot(rm.ice, center = TRUE, alpha = 0.2, rug = TRUE, train = boston)
#
# Classification example (requires randomForest package to run)
#
# Fit a random forest to the (synthetic) diabetes data
data (pima) # load the synthetic diabetes data
set.seed(102) # for reproducibility
pima.rf <- randomForest(diabetes ~ ., data = pima, na.action = na.omit)
# Partial dependence of positive test result on glucose (default logit scale)
partial(pima.rf, pred.var = "glucose", plot = TRUE, chull = TRUE,
progress = TRUE)
# Partial dependence of positive test result on glucose (probability scale)
partial(pima.rf, pred.var = "glucose", prob = TRUE, plot = TRUE,
chull = TRUE, progress = TRUE)
## End(Not run)
Synthetic Diabetes Data
Description
A fully synthetic diabetes data set created by Matthias Templ to mimic the
Pima Indians diabetes data analyzed by Smith et al. (1988). Every value is
synthetic and no row corresponds to a real person. The data were taken
directly from mlbench::SynthDiabetes2, which mimics the missing-value
pattern of the original data (physically impossible zeros are coded as
NA).
Usage
data(pima)
Format
A data frame with 768 observations on 9 variables.
-
pregnantNumber of times pregnant. -
glucosePlasma glucose concentration (glucose tolerance test). -
pressureDiastolic blood pressure (mm Hg). -
tricepsTriceps skin fold thickness (mm). -
insulin2-Hour serum insulin (mu U/ml). -
massBody mass index (weight in kg/(height in m)^2). -
pedigreeDiabetes pedigree function. -
ageAge (years). -
diabetesFactor indicating the diabetes test result (neg/pos).
Details
Earlier versions of pdp shipped a copy of the original data (taken from
mlbench::PimaIndiansDiabetes2) under this name. The original data had
most likely been shared without the consent of the participants, and both
the UCI repository and mlbench (as of version 2.1-11) have stopped
distributing it. The pima name is kept so existing code continues to run,
but results will differ from those based on the original data.
References
Smith, J.W., Everhart, J.E., Dickson, W.C., Knowler, W.C., and Johannes, R.S. (1988). Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus. In Proceedings of the Symposium on Computer Applications and Medical Care, 261-265.
Brian D. Ripley (1996), Pattern Recognition and Neural Networks, Cambridge University Press, Cambridge.
Grace Whaba, Chong Gu, Yuedong Wang, and Richard Chappell (1995), Soft Classification a.k.a. Risk Estimation via Penalized Log Likelihood and Smoothing Spline Analysis of Variance, in D. H. Wolpert (1995), The Mathematics of Generalization, 331-359, Addison-Wesley, Reading, MA.
Friedrich Leisch & Evgenia Dimitriadou (2026). mlbench: Machine Learning Benchmark Problems. R package version 2.1-11.
Examples
head(pima)
Plotting Partial Dependence Functions
Description
Plot partial dependence functions (i.e., marginal effects) and individual
conditional expectation (ICE) curves using lightweight base R graphics via
the tinyplot package, or
lattice graphics whenever lattice = TRUE.
Usage
## S3 method for class 'partial'
plot(
x,
center = FALSE,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
smooth = FALSE,
rug = FALSE,
contour = FALSE,
contour.color = "white",
train = NULL,
alpha = 1,
color.by = NULL,
bars = FALSE,
legend.title = "yhat",
lattice = FALSE,
...
)
## S3 method for class 'ice'
plot(
x,
center = FALSE,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
rug = FALSE,
train = NULL,
alpha = 1,
color.by = NULL,
lattice = FALSE,
...
)
## S3 method for class 'cice'
plot(
x,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
rug = FALSE,
train = NULL,
alpha = 1,
color.by = NULL,
lattice = FALSE,
...
)
Arguments
x |
An object that inherits from class |
center |
Logical indicating whether or not to produce centered ICE
curves (c-ICE curves). Only useful when |
plot.pdp |
Logical indicating whether or not to plot the partial
dependence function on top of the ICE curves. Default is |
pdp.col |
Character string specifying the color to use for the partial
dependence function when |
pdp.lwd |
Integer specifying the line width to use for the partial
dependence function when |
pdp.lty |
Integer or character string specifying the line type to use
for the partial dependence function when |
smooth |
Logical indicating whether or not to overlay a LOESS smooth.
Default is |
rug |
Logical indicating whether or not to include rug marks (i.e.,
the min/max and deciles of the predictor distribution) on the predictor
axes. Not currently supported for faceted displays (i.e., partial dependence
of two predictors where at least one is a factor). Default is |
contour |
Logical indicating whether or not to add contour lines to the
false color level plot used for two continuous predictors. Default is
|
contour.color |
Character string specifying the color to use for the
contour lines when |
train |
Data frame containing the original training data. Only required
if |
alpha |
Numeric value in |
color.by |
Optional character string specifying the name of a column in
|
bars |
Logical indicating whether or not to use a bar plot (rather than
points) whenever the predictor of interest is a factor. Default is
|
legend.title |
Character string specifying the text for the legend
title of the false color level plot used for two continuous predictors.
Default is |
lattice |
Logical indicating whether or not to draw the display using
lattice graphics instead of tinyplot/base graphics.
The lattice engine additionally supports three-predictor (paneled) displays
and 3-D surfaces; see Details. Default is |
... |
Additional optional arguments to be passed on to
|
Details
When lattice = TRUE, the display is constructed with
lattice graphics (this subsumes the now-deprecated
plotPartial() interface). In that case, additional lattice-specific
options can be supplied via ...:
-
levelplot- use a false color level plot (TRUE; default) or a 3-Dlattice::wireframe()surface (FALSE) for two continuous predictors; -
chull- overlay the convex hull of the first two predictors (requirestrain); -
col.regions- color palette for level/wireframe plots; -
number/overlap- number of conditioning intervals (and their fraction of overlap) used to panel a third (continuous) predictor; any other argument accepted by
lattice::xyplot(),lattice::levelplot(),lattice::wireframe(), orlattice::dotplot()(e.g.,screenordrape). The tinyplot-specific argumentscolor.by,bars, andlegend.titleare ignored whenlattice = TRUE.
Value
Draws a plot as a side effect. The tinyplot engine (invisibly)
returns x; the lattice engine (lattice = TRUE) (invisibly)
returns the "trellis" object, which can be captured for further
manipulation (e.g., arranging multiple displays with
gridExtra::grid.arrange()).
Examples
## Not run:
#
# Regression example (requires randomForest package to run)
#
# Fit a random forest to the Boston housing data
library(randomForest)
data (boston) # load the boston housing data
set.seed(101) # for reproducibility
boston.rf <- randomForest(cmedv ~ ., data = boston)
# Partial dependence of cmedv on lstat
pd <- partial(boston.rf, pred.var = "lstat")
plot(pd, rug = TRUE, train = boston)
# Partial dependence of cmedv on lstat and rm
pd2 <- partial(boston.rf, pred.var = c("lstat", "rm"), chull = TRUE)
plot(pd2, contour = TRUE)
# ICE and c-ICE curves
rm.ice <- partial(boston.rf, pred.var = "rm", ice = TRUE)
plot(rm.ice, rug = TRUE, train = boston, alpha = 0.2)
plot(rm.ice, center = TRUE, alpha = 0.2)
## End(Not run)
Plotting Partial Dependence Functions (deprecated)
Description
Plots partial dependence functions (i.e., marginal effects) using lattice graphics.
Usage
plotPartial(object, ...)
## S3 method for class 'ice'
plotPartial(
object,
center = FALSE,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
rug = FALSE,
train = NULL,
...
)
## S3 method for class 'cice'
plotPartial(
object,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
rug = FALSE,
train = NULL,
...
)
## S3 method for class 'partial'
plotPartial(
object,
center = FALSE,
plot.pdp = TRUE,
pdp.col = "red2",
pdp.lwd = 2,
pdp.lty = 1,
smooth = FALSE,
rug = FALSE,
chull = FALSE,
levelplot = TRUE,
contour = FALSE,
contour.color = "white",
col.regions = NULL,
number = 4,
overlap = 0.1,
train = NULL,
...
)
Arguments
object |
An object that inherits from the |
... |
Additional optional arguments to be passed onto |
center |
Logical indicating whether or not to produce centered ICE
curves (c-ICE curves). Only useful when |
plot.pdp |
Logical indicating whether or not to plot the partial
dependence function on top of the ICE curves. Default is |
pdp.col |
Character string specifying the color to use for the partial
dependence function when |
pdp.lwd |
Integer specifying the line width to use for the partial
dependence function when |
pdp.lty |
Integer or character string specifying the line type to use
for the partial dependence function when |
rug |
Logical indicating whether or not to include rug marks on the
predictor axes. Default is |
train |
Data frame containing the original training data. Only required
if |
smooth |
Logical indicating whether or not to overlay a LOESS smooth.
Default is |
chull |
Logical indicating whether or not to restrict the first two
variables in |
levelplot |
Logical indicating whether or not to use a false color level
plot ( |
contour |
Logical indicating whether or not to add contour lines to the
level plot. Only used when |
contour.color |
Character string specifying the color to use for the
contour lines when |
col.regions |
Vector of colors to be passed on to
|
number |
Integer specifying the number of conditional intervals to use
for the continuous panel variables. See |
overlap |
The fraction of overlap of the conditioning variables. See
|
Details
Deprecated: plotPartial() is deprecated and will be removed
in a future release; please use plot(..., lattice = TRUE) instead,
which produces the same displays through a single interface (see
plot.partial() for details).
Examples
## Not run:
#
# Regression example (requires randomForest package to run)
#
# Load required packages
library(gridExtra) # for `grid.arrange()`
library(magrittr) # for forward pipe operator `%>%`
library(randomForest)
# Fit a random forest to the Boston housing data
data (boston) # load the boston housing data
set.seed(101) # for reproducibility
boston.rf <- randomForest(cmedv ~ ., data = boston)
# Partial dependence of cmedv on lstat
boston.rf %>%
partial(pred.var = "lstat") %>%
plotPartial(rug = TRUE, train = boston)
# Partial dependence of cmedv on lstat and rm
boston.rf %>%
partial(pred.var = c("lstat", "rm"), chull = TRUE, progress = TRUE) %>%
plotPartial(contour = TRUE, legend.title = "rm")
# ICE curves and c-ICE curves
age.ice <- partial(boston.rf, pred.var = "lstat", ice = TRUE)
p1 <- plotPartial(age.ice, alpha = 0.1)
p2 <- plotPartial(age.ice, center = TRUE, alpha = 0.1)
grid.arrange(p1, p2, ncol = 2)
## End(Not run)
Retrieve the last trellis object
Description
See lattice::trellis.last.object() for more details.
Usage
trellis.last.object(..., prefix)