fireSense_spreadFit Module

made-with-Markdown

Authors:

Eliot McIntire [aut, cre], Tati Micheletti [aut], Ian Eddy [aut], Jean Marchal [aut], Alex M. Chubaty [ctb]

Module Overview

Module summary

Fit statistical models that can be used to parameterize the fire spread component of simulation models (e.g., fireSense (Marchal, Steve G. Cumming, et al. 2017b; Marchal, Steve G. Cumming, et al. 2017a; Marchal et al. 2019)). This module implement a Pattern Oriented Modelling (POM) approach to derive spread probabilities from final fire sizes. Spread probabilities can vary between pixels, and thus reflect local heterogeneity in environmental conditions.

The fit is a differential evolution search (DEoptim, run by fireSenseUtils::runDEoptim() on a cluster built by the clusters package). Each candidate parameter set is scored by simulating the historical fires and comparing simulated with observed fire sizes (fireSenseUtils::.objfunSpreadFit()). The 5 best parameter sets, and the covariate ranges used to rescale the covariates, are written as one row per polygon to a shared “fit ledger” on Google Drive (spreadFitFilename in spreadFitGoogleDriveFolder), keyed by .ELFind. If the ledger already holds a row for the polygon, the module does nothing unless refitExisting = TRUE. By default (stopIfNoPreRunFit = TRUE) the module stops rather than start a fit; set it to FALSE to fit.

refitExisting

refitExisting = TRUE forces a fit for this polygon even when the ledger already holds a row for it. Use it when the fit’s INPUTS have changed – new land cover, new vegetation parameters, a new objective function – so the stored row is stale and the polygon must be fitted again.

refitExisting overrides stopIfNoPreRunFit: with refitExisting = TRUE, init schedules the fit rather than stopping, whatever stopIfNoPreRunFit is set to. Setting stopIfNoPreRunFit = FALSE is only needed when the polygon has no ledger row.

refitExisting is intended for developers who have access to at least 40 cores: it triggers a full DEoptim run (see cores and nCoresNeeded), which is not practical on a small machine.

Module inputs and parameters

The covariate tables, fire buffers, fire points and formula are made by fireSense_dataPrepFit. fireSense_spreadFormula must be supplied; .ELFind defaults to .runName.

Table 1 shows the full list of module inputs.

Table 1: Table 2: List of fireSense_spreadFit input objects and their description.
objectName objectClass desc sourceURL
.runName character Some descriptive, short name for this fitting, e.g., ELF14.1 NA
.ELFind character Identifier of the polygon being fit, e.g. ‘6.1.1’. This becomes the polygonID of the row this module writes to the shared cloud fit ledger (spreadFitFilename in spreadFitGoogleDriveFolder), which fireSense_dataPrepFit matches against the polygon ids carried by rasterToMatchELF. It must therefore be the polygon’s identity, not a run label: .runName encodes the whole scenario (climate period, GCM, SSP, rep) in some projects, and keying the ledger on it writes rows no other run can find and trips dataPrepFit’s id match. Defaults to .runName for backwards compatibility. NA
fireBufferedListDT list list of data.tables with fire id, pixelID, and buffer status NA
fireSense_annualSpreadFitCovariates data.table table of climate and/or veg covariates, burn status, polyID, and pixelID NA
fireSense_nonAnnualSpreadFitCovariates data.table table of veg covariates, burn status, polyID, and pixelID NA
spreadFitAdditionalColNames character Names of the list-columns of the ledger row. Reset to fireSenseUtils::spreadFitAdditionalColNamesTxt if different. NA
fireSense_spreadFormula character a formula that contains the annual and non-annual covariates e.g. ~ 0 + MDC + class2 + class3 + youngAge. NA
parsKnown numeric Optional vector of known parameters, e.g., from a previous DEoptim run. If this is supplied, then ‘mode’ will be automatically converted to ‘debug’ NA
rasterToMatch SpatRaster template raster for study area NA
spreadFirePoints sf list of sf points, one element per year, of fire ignition locations NA
studyArea sf Polygon being fit; its geometry and crs go in the ledger row. Defaults to NWT. https://drive.google.com/open?id=1LUxoY2-pgkCmmNH5goagBp3IMpj6YrdU

Summary of user-visible parameters (Table 3)

Table 3: Table 4: List of fireSense_spreadFit parameters and their description.
paramName paramClass default min max paramDesc
.plots characte…. NA NA Plot types passed to Plots(), e.g. ‘png’ or ‘screen’; NULL or NA for none.
.plotInterval numeric 25 NA NA DEoptim generations between DEoptim progress figures; the final figures are always drawn. Passed to fireSenseUtils::runDEoptim() as plotEvery.
.plotSize list 1600, 2000 NA NA List specifying height and width of plotting device (in pixels) used to plot DEoptim histograms when visualizeDEoptim is TRUE.
.runInitialTime numeric 0 NA NA when to start this module? By default, the start time of the simulation.
.studyAreaName character NA NA NA Human-readable name for the study area used.
.useCache logical,…. init NA NA Should this entire module be run with caching activated? This is generally intended for data-type modules, where stochasticity and time are not relevant.
cores integer 1 NA NA Passed to cores in fireSenseUtils::runDEoptim(): a number of local cores, or a character vector of machine names, one element per core wanted on that machine.
DEoptimTests character adTest, …. NA NA Currently either 'SNLL_FS' or 'adTest' or a length 2 character vector of both. Passed to tests in fireSenseUtils::.objfunSpreadFit().
doObjFunAssertions logical TRUE NA NA Passed to fireSenseUtils::.objfunSpreadFit(); TRUE runs diagnostics but is slower; FALSE for operational runs
initialpop numeric NA NA A numeric matrix of dimensions NCOL = length(lower) and NROW = NP. This will be passed into DEoptim through control$initialpop = P(sim)$initialpop if it is not NULL
iterDEoptim integer 5000 NA NA integer defining the maximum number of iterations allowed (DEoptim optimizer). A ceiling: clusters (>= 0.0.46) stops the fit earlier, once the population’s median value has stopped improving.
iterThresh integer 96 NA NA Number of random parameter sets tried when calibrating SNLL_FS_thresh.
libPathDEoptim character /home/ru…. NA NA Absolute path specifying R package directory location to use when running DEotpim. NOTE: this path must be read/write accessible on ALL machines used for fitting (identified in cores). Therefore, it’s best use a directory in your user’s ~ directory. If the directory does not exist at this path, will attempt to create it.
lower numeric NA NA NA see ?DEoptim. Lower limits for the logistic function parameters (lower bound, upper bound, slope, asymmetry) and the statistical model parameters (named in the order they appear in the formula). Do not include hillSlope1: it is fixed at 1, not fitted (see estimateSpreadParams()); supplying it is an error.
maxFireSpread numeric 0.28 NA NA optional. Maximum fire spread average to be passed to the .objFun. This puts an upper limit on spreadProb during optimization.
link character logistic3p NA NA The spread link. ‘logistic3p’, or ‘logistic3pUpper’: the same curve with Stukel’s upper tail, one more parameter upperTail1 that changes only how the curve approaches its ceiling (fireSenseUtils::logistic3pUpper()). Its default bounds are upperTailBounds.
mode character fit NA NA Options: debug, fit, visualize, validate. Can use multiples. ‘debug’ runs the objective function with visuals instead of DEoptim; ‘fit’ runs DEoptim; ‘visualize’ adds the debug and plot events after the fit; ‘validate’ adds crossValidate, two more fits, each on half the years, predicting the other half (sim$spreadFitHeldOut). Validation never writes the ledger, but does write sim$spreadFitHeldOut to outputPath(sim), since a batch run typically stops after crossValidate and the simList is discarded. See heldOutFold to run a single fold as its own job instead of both folds together.
heldOutFold integer NA NA NA NA (default): unchanged behaviour, governed by mode. 1 or 2: run ONLY that cross-validation fold, as its own job. init then schedules spreadFitPrepare, estimateThreshold and crossValidate – never run, so the full fit and the ledger write never happen, and the ledger (stopIfNoPreRunFit/refitExisting) is not consulted. crossValidate fits on the OTHER fold’s years and scores this fold’s held-out years (cvFolds() in R/fitSpread.R), and writes spreadFitHeldOut_<.runName>_fold<heldOutFold>.rds instead of spreadFitHeldOut_<.runName>.rds. A run script stops after crossValidate: events = list(.stopAfter = list(fireSense_spreadFit = "crossValidate")). Any other value is an error.
profileReps integer 10 NA NA After the fit, each covariate coefficient in turn is set to 0 and to 5 values across the final population, the others held at the best member, and each point is evaluated this many times (fireSenseUtils::profileCoefficients()). About 6 nCoefficients profileReps evaluations, on the fit’s workers. It decides which coefficients are identified in isolation (sim$spreadFitIdentifiability). 0 skips it.
mutuallyExclusiveCols list c(“class…. NA NA a named list of mutually exclusive covariates - see fireSenseUtils::makeMutuallyExclusive
nCoresNeeded integer NA NA How many workers to request for the DEoptim cluster. This IS the population size: clusters::clusterSetup() sets NP to the workers it builds. NULL leaves fireSenseUtils::runDEoptim()’s default of 10 per estimated parameter. A generation costs the slowest of NP evaluations and that barely falls as NP falls, so a smaller NP buys throughput by allowing more fits at once rather than by shortening generations (measured 2026-09-16).
simulateMembers integer 10 NA NA After the fit, this many best members simulate the observed fires without the size cap, for sim$spreadFitSizes and sim$spreadFitLinkSaturation; also the members each crossValidate fold predicts with. 0 skips it after the fit.
sizeLik character t NA NA Likelihood of fire size in the objective, ‘kde’ or ‘t’, passed to fireSenseUtils::runDEoptim(). ‘t’ with weighted = FALSE predicted held-out years best in the 2026-09-21 cross-validation.
escapeSizeHa numeric 50 NA NA Size (ha) a fire must reach to count as escaped. The spread model is fitted to escaped fires only, and each simulated fire first grows to this size with its own spread probabilities (its burning cells stay active until it gets there), then spreads normally. NULL or NA gives the old fit (any fire over 1 pixel). Passed to fireSenseUtils::runDEoptim().
sizeLikDf numeric 5 NA NA Degrees of freedom of the ‘t’ size likelihood.
weighted logical|…. FALSE NA NA Weight of each fire in the size likelihood: FALSE (none), TRUE (log size) or ‘sqrt’. Passed to fireSenseUtils::runDEoptim().
adWeight characte…. auto NA NA Weight of the Anderson-Darling term against the size likelihood; ‘auto’ is fireSenseUtils::adWeightAuto().
yearAreaWeight numeric|…. auto NA NA Weight of the annual-area term in the objective: each fit year’s observed area burned is scored against that year’s simulated totals (one per replicate) with the size likelihood. 0 leaves it out; ‘auto’ (default) is (number of fitted fires) / (number of fit years), so the year view and the per-fire view weigh the same. Passed to fireSenseUtils::runDEoptim().
areaDistWeight numeric|…. auto NA NA Weight of the area-weighted size-distribution term: simulated and observed fires compared by the share of area burned that fires up to each size make up (fireSenseUtils::areaWeightedCvM()). 0 leaves it out; ‘auto’ (default) uses the Anderson-Darling term’s weight (fireSenseUtils::adWeightAuto()).
jumpTries numeric 20 NA NA With escapeSizeHa: how many attempts a simulated fire that is still below the escape size, with no burnable neighbour left, may make to jump to burnable land nearby (SpaDES.tools::spreadCpp()). Default 20; 0 is off.
jumpMeanDist numeric 3 NA NA Mean jump distance (pixels) for jumpTries; distances are exponential, truncated to 1.5-20 pixels. No effect while jumpTries is 0.
objFunCoresInternal integer 1 NA NA Integer defining the number of cores to pass to mcmapply(mc.cores = ...) This will fork this many to do the years loop internally. This would be in addition to cores and is effecively a multiplier. The computer needs to have cores objFunCoresInternal threads or it will stall.
objfunFireReps integer 50 NA NA integer defining the number of replicates the objective function will attempt each fire.
rep integer 1 NA NA An optional integer indicating which replicate run this represents. This is used to identify unique runs of runDEoptim, from a Cache perspective. For example, if this module is run twice with all the same data, Cache will think that the second run should recover the cache result, unless this rep is modified
.c numeric 0 NA NA the c argument passed to DEoptim.control. iterStep is hard-coded to 1, so DEoptim’s adaptation restarts every generation, and c has no effect either way.
DEoptimControl list 0.1 NA NA Further DEoptim.control() settings, e.g. list(CR = 0.7, F = 0.6), passed through fireSenseUtils::runDEoptim() to DEoptim. Names must be DEoptim.control() arguments. strategy, trace, initialpop and .c have their own parameters; NP is the number of workers the cluster gets. The default p = 0.1 is for strategy = 6.
rescaleAll logical TRUE NA NA rescale covariates for DEOptim
spreadFitGoogleDriveFolder character https://…. NA NA Google Drive folder url holding the shared fit ledger (spreadFitFilename).
spreadFitFilename character latest NA NA File name of the shared fit ledger: an sf object with one row of fitted parameters per polygon. "latest" (the default) writes to the file named for this fit’s fire years and model, fireSenseUtils::spreadFitFilenameFor(), e.g. fireSenseParams_1985-2024_linearFuel.rds; readers then find it with fireSenseUtils::latestSpreadFits().
strategy integer 6 NA NA Passed to DEoptim.control. 6 (DE/current-to-p-best/1) with p = 0.1 did best in a settings study (see NEWS).
SNLL_FS_thresh integer NA NA Threshold multiplier used in objective function SNLL fire size test.
refitExisting logical FALSE NA NA FOR DEVELOPERS ONLY: a re-fit is a full DEoptim run and is only practical with access to at least 40 cores. Fit this polygon even when the ledger already holds parameters for it. A ledger row normally means the fit is done, and the run event skips it. Set this when the fit’s INPUTS have changed – new land cover, new vegetation parameters, a new objective – so the stored row is stale and the polygon must be fitted again. When TRUE it OVERRIDES stopIfNoPreRunFit: init schedules the fit instead of stopping.
stopIfNoPreRunFit logical TRUE NA NA If TRUE, init stops with an error when this polygon would have to be fitted, instead of fitting it. Ignored when refitExisting is TRUE.
trace numeric 1 NA NA non-negative integer. If > 0, tracing information on the progress of the optimization are printed every trace iteration. Default is 1, i.e. every iteration. Setting to 0 turns off tracing.
upper numeric NA NA NA see ?DEoptim. Upper limits for the logistic function parameters (lower bound, upper bound, slope, asymmetry) and the statistical model parameters (named in the order they appear in the formula). Do not include hillSlope1: it is fixed at 1, not fitted (see estimateSpreadParams()); supplying it is an error.
useCache_DE logical TRUE NA NA should DEoptim use Cache? to do multiple independent runs, use FALSE
verbose numeric 1 NA NA optional. With increasing number, more verbosity. Level 1 is normal reproducible (e.g., Cache), level 2 includes objective function e.g., print median of spreadProb during calculations
visualizeDEoptim Path /tmp/Rtm…. NA NA Directory where runDEoptim saves parameter plots every .plotInterval generations. Reset to figurePath(sim) unless its last folder is the module name.
covFixedRange list c(0, 100…. NA NA Named list of c(min, max): covariates rescaled with this FIXED range and not with the range of this polygon’s data. CMDsm = c(0, 100) makes the covariate CMDsm / 100 in every polygon. With the data’s range, 1 meant a CMDsm of 104 in one polygon and 297 in another, so the coefficient could not be compared across polygons, and a polygon that never gets dry stretched its small range over [0, 1]. Names not among the covariates are ignored. fireSense_spreadPredict rescales with the stored covMinMax_spread, so it follows. CMD, CMDsp and cumMDC (also mm) are the other candidates of fireSense_dataPrepFit’s spread = 'auto', so an ELF that picks one of them gets the same fixed scale.
yearSpreadSDBounds numeric 0, 1 NA NA Bounds of yearSpreadSD, the sd of a per-year random effect on logit spread probability (fireSenseUtils::.objfunSpreadFit()), when lower/upper are not supplied. A seasonal departure: each year draws one eps, so all of a year’s fires burn hotter or cooler together, which widens the simulated fire-size distribution. NA turns it off.
upperTailBounds numeric -1, 1 NA NA Bounds of upperTail1 when link is ‘logistic3pUpper’ and lower/upper are not supplied.
upperAndLowerVal numeric 50 NA NA Bound given to each covariate coefficient (upper = this, lower = minus this) when upper or lower is not supplied. Bounds should be wide enough that they do not influence the fitted value; only the sign of drought-index and youngAge terms is constrained (see estimateSpreadParams()). A held-out experiment (7 ELFs x 2 folds) found estimates up to 25.7 and youngAge medians down to -23.0 with the previous default of 9.
upperAndLowerValFuel numeric 100 NA NA As upperAndLowerVal, for the fuel biomass covariates. They are biomass / 1e4, so their coefficients are larger than those of covariates rescaled to [0, 1]. The same held-out experiment found fuel estimates up to 54.5 (29 of 82 above 25) with the previous default of 60.

Events

  • init: schedules spreadFitPrepare and, if the polygon needs a fit, estimateThreshold then run (or debug when mode includes "debug"; debug and plot after run when it includes "visualize");
  • spreadFitPrepare: sets default lower/upper, computes covMinMax_spread and lociList, converts covariates to integers (x 1000);
  • estimateThreshold: uses SNLL_FS_thresh, or calibrates it from iterThresh random parameter sets (cached, with a seed derived from .ELFind);
  • run: runs DEoptim (cached per generation) and writes the polygon’s row to the ledger;
  • debug: evaluates the objective function without DEoptim;
  • plot: histograms of the final population;

Plotting

With .plots set, histograms of the annual and non-annual covariates. During the fit, runDEoptim() saves parameter histograms and trace plots to visualizeDEoptim after each generation.

Saving

The fit’s row in the cloud ledger, and the DEoptim cache entries. Nothing else is saved.

Module outputs

Description of the module outputs (Table 5).

Table 5: Table 6: List of fireSense_spreadFit outputs and their description.
objectName objectClass desc
covMinMax_spread data.table data.table of covariates min and max
DE data.table list of DEoptim objects, one per generation, ordered by best objective value
studyAreaWithSpreadParams sf Rows of the shared fit ledger that intersect studyArea, including the row this fit writes: studyArea geometry, polygonID, and list-columns named by spreadFitAdditionalColNames (the 5 best parameter sets are in params).
fsSpreadFit_hists ggplot histograms of each parameter used in DEoptim fitting.
lociList list per-year data.tables of fire start cells and sizes, from fireSenseUtils::makeLociList()
spreadFitConvergence data.table The objective across the fit’s generations (fireSenseUtils::fitConvergence()).
spreadFitRescore data.table The final population, one row per member, with the mean and sd of its replicated re-scores (reMean, reSD). The ledger’s parameter sets are the best of these.
spreadFitIdentifiability data.table One row per covariate coefficient: how tightly the population pins it (fireSenseUtils::coefIdentifiability()) and, with profileReps > 0, whether dropping it worsens the fit; identified = identified in isolation (fireSenseUtils::identifiedInIsolation()).
spreadFitProfile data.table The one-at-a-time profile around the best member (fireSenseUtils::profileCoefficients()).
spreadFitSizes data.table Observed against simulated fire sizes of the fitted years, without the size cap (fireSenseUtils::scoreFireSizes()): bias, error, quantiles.
spreadFitLinkSaturation data.table Per member, the share of pixel-years at the spread-probability ceiling and the quantiles of spread probability (fireSenseUtils::linkSaturation()).
spreadFitHeldOut list mode ‘validate’, or heldOutFold in 1:2: sims, the held-out years simulated from the fit to the other years (column fold), and score, from fireSenseUtils::scoreFireSizes(). With heldOutFold, sims holds only that fold. Also written to file.path(outputPath(sim), currentModule(sim), "spreadFitHeldOut_<.runName>.rds") (mode ‘validate’) or "...spreadFitHeldOut_<.runName>_fold<heldOutFold>.rds" (heldOutFold).

Inputs come from fireSense_dataPrepFit. fireSense_spreadPredict uses studyAreaWithSpreadParams (the ledger rows) to predict spread probabilities.

Usage

## in a project that also runs fireSense_dataPrepFit
params <- list(fireSense_spreadFit = list(
  stopIfNoPreRunFit = FALSE,          # allow a fit to start
  cores = rep(c("hostA", "hostB"), each = 20), # host name repeated once per worker; or a number for localhost
  nCoresNeeded = 40,                  # = DEoptim population size (NP)
  iterDEoptim = 500,
  mode = "fit"
))
objects <- list(.ELFind = "6.1.1")    # polygon id used as the ledger key

References

Marchal, Jean, Steve G. Cumming, and Eliot J. B. McIntire. 2017a. “Exploiting Poisson Additivity to Predict Fire Frequency from Maps of Fire Weather and Land Cover in Boreal Forests of Québec, Canada.” Ecography 40 (1): 200–209. https://doi.org/10.1111/ecog.01849.
Marchal, Jean, Steve G Cumming, and Eliot J B McIntire. 2017b. “Land Cover, More Than Monthly Fire Weather, Drives Fire-Size Distribution in Southern Québec Forests: Implications for Fire Risk Management.” PLoS ONE 12 (6): 1–17. https://doi.org/10.1371/journal.pone.0179294.
Marchal, Jean, Steven G. Cumming, and Eliot J. B. McIntire. 2019. “Turning Down the Heat: Vegetation Feedbacks Limit Fire Regime Responses to Global Warming.” Ecosystems, ahead of print, May. https://doi.org/10.1007/s10021-019-00398-2.