Changelog
Source:NEWS.md
randomPlantedForest 0.4.0
It is now possible to serialize and de-serialize an rpf object via marshalling, meaning a fitted model can be saved to disk or dispatched to a worker process for parallelization or encapsulation, as is done in mlr3. This also re-opens the door for the mlr3extralearners wrapper, which was removed due to the lack of serialization support.
- New
rpf_marshal()/rpf_unmarshal()serialize a fitted forest to a plain R list and back, makingsaveRDS()-based storage of rpf models possible (#52). Purified forests restore their purified state directly; training data is only embedded withinclude_data = TRUE(required topurify()after restoring). Blobs record the blob format version and the package version they were created with; restoring under an older package than the one that saved the model warns. Malformed or corrupt blobs error instead of reading out of bounds. - New
rpf_is_valid()checks whether an rpf object’s internal model is usable,predict(),purify()andpredict_components()now give an actionable error forrpfobjects restored viareadRDS()without marshaling. - New
bundle::bundle()method for rpf models wrapping the marshaling API.
Behavior changes
- The default for
deltachanged from 0 to 0.001. Withdelta = 0, splits producing single-class nodes have infinite logit loss and are always rejected, which also degraded binary logit fits.
Fixes and improvements
- Fixed multiclass classification with
loss = "logit", which effectively never split and predicted near-uniform class probabilities (#40). The C++ logit loss is a reference-class multinomial formulation expectingK-1indicator columns, but the R wrapper passed a fullK-column one-hot matrix, pinning the implicit reference-class probability to zero. Multiclass logit outcomes are now encoded with the first factor level as reference class:-
predict(type = "prob")reconstructs allKclass probabilities from theK-1logits; rows sum to 1 exactly. -
predict(type = "numeric"/"link")now returnsK-1columns named after the non-reference levels (previouslyKcolumns). -
predict_components()on multiclass logit fits returns per-class components for theK-1non-reference levels;target_levelsreflects this.
-
randomPlantedForest 0.3.0
Major changes (#61)
- New
rpf()arguments controlling split-candidate sampling:-
split_structure = "leaves": Defines what a split candidate is and how candidates are drawn. One of"leaves"(default),"hist","cur_trees_1","cur_trees_2", or"res_trees"; see?rpffor details. -
max_candidates = 50: Maximum number of split candidates sampled per iteration. -
split_decay_rate = 0.1: Exponential aging of repeatedly drawn but unchosen split candidates.split_decay_rate = 0corresponds to no aging and uniform sampling. -
delete_leaves = TRUE: Whether a parent leaf is deleted when splitting along an existing dimension.
-
- Fitting results change: The new candidate-sampling defaults and a reworked internal RNG mean that fits are not reproducible against previous versions, even with the same seed. Install an older commit if exact reproduction of previous results is required.
- Seeded fits are now reproducible regardless of
nthreads: per-tree seeds are drawn from R’s RNG, soset.seed()gives identical forests for serial and multithreaded fits. - Substantial speedups in fitting (cached per-leaf orderings, prefix sums) and reduced memory use (training-only buffers are released after each tree family is built).
-
purify()gains arguments:-
mode = 2: Purification algorithm;2is a new fast exact method,1is the legacy grid-based path. -
nthreads = NULL: Purification is now multithreaded, defaulting to the fit’snthreadssetting. -
maxp_interaction = NULL: Optionally only compute purified components up to this interaction order.
-
- New
rpf()argumentexport_forest = FALSE: The flattened forest is no longer stored in the fitted object by default, sorpf_object$forestisNULLunlessexport_forest = TRUE. This reduces object size;predict(),purify(), andpredict_components()are unaffected. -
preprocess_predictors_predict()is now exported. - Fixed a memory bug in the legacy purification path where the grid was sized one element too large, causing out-of-bounds reads (crashes on Windows, silently wrong purification results elsewhere).
- Fixed a crash on Windows when fitting with
nthreads > 1, caused by athread_localbuffer with a non-trivial destructor being destroyed at thread exit.
Other changes
- Internals in
src/have been refactored into modular sub-files (#53) -
rpf()now errors if a regression target is combined with alossother than"L2". - Allow features of type
logical, which are now converted viaas.integer. - The
parallel = TRUE|FALSEargument inrpf()has been substituted by annthreads = 1Largument, allowing for more flexible parallelization. The previous behavior only allowed for either no parallelization or using n-1 of n available cores. The new implementation should be reasonably robust and the default behavior remains serial execution. - Remove
SystemRequirementsfield fromDESCRIPTION: Now the default C++ version is C++17 and with a minor change to internal use of random numbers,randomPlantedForestis now compatible with C++11 through C++23. - Add
remainderterm topredict_componentsoutput for case wheremax_interactionsupplied is smaller thanmax_interactioninrpffit. In that case, themvalues don’t sum up to the global predictions, so we add a remainder to allow reconstruction of that property.
randomPlantedForest 0.2.1
- Add
glexclass to output ofpredict_components(), for extended functionality available withglex. - Add
target_levelsvector to output ofpredict_components()to aid multiclass handling. Keeping track of levels is somewhat awkward since column names of$mneed to be identifiable regarding the target level.