0

Why Random Forest Needs to Be This Random

https://towardsdatascience.com/why-random-forest-needs-to-be-this-random/(towardsdatascience.com)
The Random Forest algorithm's second layer of randomness, feature subsampling, is necessary because simple bagging has a performance ceiling that more trees cannot break. Mathematically, the variance of an ensemble of correlated models does not approach zero but instead converges to a floor determined by the average pairwise correlation (ρ) between the individual trees. As the number of trees increases, the ensemble variance approaches ρσ², meaning that correlated errors cannot be fully averaged away. The additional randomness at each split deliberately decorrelates the trees, reducing ρ and thereby lowering the variance floor to improve the model's predictive power beyond what bagging alone can achieve.
0 pointsby hdt1 hour ago

Comments (0)

No comments yet. Be the first to comment!

Want to join the discussion?