scieee AI-readable full text Open interactive document viewer

Boosted Neural Networks for Tabular Regression

Musaliyarakath, Rizeen; A.S Abbas, Jessica

Full text

Student Research Project Research Topic: Boosted Neural Networks for Tabular Regression Authors: 1747556 Rizeen Musaliyrakath, 1748971 Jessica A.S Abbas, 1749059 Gayara Gunasekera, Supervisor Kiran Madhusudhanan 15th March 2025 Contents I abstract .............................. 2 1 Introduction 3 I Motivation ............................ 3 I.1 Problem Setting . . . . . . . . . . . . . . . . . . . . . 4 I.2 ResearchIdea....................... 5 I.3 Objective ......................... 6 2 Related Works 7 I The Boosting Framework . . . . . . . . . . . . . . . . . . . . 7 II Gradient-Boosted Decision Trees (GBDTs) . . . . . . . . . . 8 III Boosted Neural Network Architectures . . . . . . . . . . . . . 9 3 Methodology 22 I Integrated NODE + PLE Architecture . . . . . . . . . . . . . 22 II Boosted Fully Connected Networks (BFCN) . . . . . . . . . . 24 III Boosted Residual Networks . . . . . . . . . . . . . . . . . . . 26 4 Experiments and Results 29 I Datasets.............................. 29 II Analysis of Results . . . . . . . . . . . . . . . . . . . . . . . . 30 5 Conclusion 33 I Contributions........................... 35 1 I abstract Tabular data represent one of the most prevalent forms of data in machine learning context. Despite recent advancements in using neural nets (NNs) to handle tabular data, there remains an active and ongoing debate about whether NNs outperform gradient-boosted decision trees (GBDTs) when examined with respect to tabular data, with some recent work suggesting either that GBDTs are consistently better than NNs or that NNs are consistently better than GBDTs. In this work, we attempt to bridge this gap by exploring various boosted neural network architectures on tabular data. We propose and evaluate three frameworks: Integrated Neural Oblivious Decision Ensembles (NODE) with Piecewise Linear Encoding (PLE), Boosted Fully Connected Networks (BFCN), and Boosted Residual Networks. Our experimental results reveal a mixed performance landscape for the proposed boosted neural network architectures when compared to established baselines. Against gradient-boosted decision trees, NODE+PLE demonstrates competitive performance primarily on regression tasks. However, significant underperformance is evident across classification tasks, with particularly poor results on HELOC and moderate performance on Adult Datasets. BFCN consistently underperforms GBDTs across all metrics, failing to achieve competitive results on any dataset. When evaluated against deep learning baselines, NODE+PLE shows more promising results, achieving state-of-the-art performance on California Housing regression and competitive accuracy on Adult classification. The results underscore the persistent challenge of achieving GBDT-level performance with neural architectures on tabular data, while simultaneously demonstrating that boosted neural networks can advance the state-of-the-art within the deep learning paradigm for specific dataset characteristics. Keywords Deep Learning (DL), Neural Oblivious Decision Ensembles (NODE), Deep Neural Networks (DNNs),Piecewise Linear Encoding (PLE), Gradient-boosted decision trees (GBDTs), Boosted Dynamic Neural Networks (BoostNet) Codebase: https://www.uni-hildesheim.de/gitlab/srp-group 2 Chapter 1 Introduction IMotivation Deep neural networks (DNNs) have achieved exceptional performance across a wide array of domains, including computer vision, natural language processing, and speech recognition, particularly when dealing with homogeneous data such as images, audio, and text [1, 8, 11]. However, their effectiveness on heterogeneous tabular data remains a significant and persistent challenge [14, 46, 37]. Tabular data, unlike image or language data, is inherently heterogeneous, often comprising a mix of dense numerical and sparse categorical features, with correlations among these features typically weaker and more irregular [7]. This challenge is critical because tabular data represents the most commonly used form of data and is indispensable for numerous vital and computationally demanding applications [3, 7, 46]. Despite the proven success of DNNs in other domains, gradient-boosted decision trees (GBDTs), such as XGBoost [9], LightGBM [27], and CatBoost [38], still largely outperform deep learning models on supervised learning tasks involving tabular data [7]. This indicates a potential stagnation in research progress for competitive deep learning models in this domain [7]. Empirical comparisons often show that GBDTs o!er superior accuracy, training e”ciency, inference speed, and hyperparameter optimization time [7]. There is an active debate on whether neural networks or GBDTs generally outperform each other on tabular data, with various works arguing for [4, 26, 30, 36, 41] or against [7, 20, 21, 46] NNs. However, McElfresh et al. [32] reported that this ”NN vs. GBDT” debate may be overemphasized, as for a significant number of datasets, either the performance di!erence is 3 negligible, or light hyperparameter tuning on a GBDT is more impactful than the choice between NN and GBDT. Nevertheless, deep neural networks possess several inherent advantages that make their adaptation to tabular data a compelling research direction. DNNs are highly flexible [42],allow for e”cient and iterative training, and are particularly valuable in AutoML contexts [23, 45]. Furthermore, neural networks can be deployed for multimodal learning problems where tabular data serves as one input modality [45], for tabular data distillation [33, 31], and in federated learning scenarios [40]. Given the persistent performance gap and the unique characteristics of tabular data, this project aims to bridge the divide by exploring and enhancing boosted neural network frameworks for tabular data prediction. The core idea is to investigate how an ensemble approach to modeling neural networks can adapt the strengths of traditional tree-boosted models to achieve improved predictive performance. I.1 Problem Setting Given a dataset X→RM→Nand corresponding labels y→R1→N,whereN represents the number of training instances and Mrepresents the number of input features. Let ω:RN↑RN↓Rdenote a loss function, and fω: RM↓Rrepresent a neural network parameterized by learnable parameters ω. Our objective is to find the optimal parameter vector ω↑that minimizes the empirical risk: ω↑= arg min ωL(ω) = arg min ωω(y,f ω(X)) where fω(X)→RNrepresents the network’s predictions over all training instances, and ωencompasses all trainable parameters including weights and biases across all layers of the neural network. The loss function is defined for di!erent task types by the following: Regression: Mean Squared Error (MSE) is applied LMSE =1 n n ! j=1 (yj↔ˆyj)2. 4 Binary Classification: Binary Cross-Entropy Loss LBCE =↔1 n n ! j=1 [yjlog ˆyj+(1↔yj) log(1 ↔ˆyj)] . Multiclass classification: Cross-Entropy Loss LCE =↔ n ! j=1 yjlog ˆyj I.2 Research Idea Gradient boosting techniques, particularly models like XGBoost, LightGBM, and CatBoost, have become the established standard for tabular data prediction, consistently achieving state-of-the-art performance [7, 32]. This stands in contrast to the exceptional performance of deep neural networks (DNNs) in other homogeneous data domains such as computer vision and natural language processing [48, 18]. The e!ectiveness of DNNs on heterogeneous tabular data, which often consists of a mix of dense numerical and sparse categorical features with weaker and more irregular correlations, remains a significant challenge [43, 52, 48]. Indeed, tabular datasets have been called the ”last ’unconquered castle’” for deep neural network models [7]. While the ”NN vs. GBDT” debate can be overemphasized for many datasets where performance di!erences are negligible or hyperparameter tuning is more impactful, GBDTs are generally better at handling skewed or heavytailed feature distributions and other data irregularities, and tend to perform better on larger datasets [32]. This project is grounded in the hypothesis that employing neural networks within a boost-like structure can enhance their predictive accuracy on tabular datasets . Specifically, we propose to investigate the possibility of combining structured feature engineering, such as Piecewise Linear Encoding (PLE) [19], with hierarchical, boost-like learning principles through neural networks. We aspire to explore how a thoughtful combination of neural networks in a consecutive boosting format can adapt the strengths of traditional tree-boosted models, aiming for improved predictive performance over conventional methods for tabular data prediction. The ultimate goal is to demonstrate that architectures employing these boosted neural network frameworks could potentially surpass the performance of both traditional GBDTs and standalone deep neural networks, o!ering versatile and potentially interpretable solutions for complex tabular data challenges. 5 I.3 Objective The primary objective of this study is to explore and advance the use of boosted neural network techniques for tabular data prediction. To achieve this, we define the following goals: •Conduct a thorough review of recent advancements in boosted neural network architectures and ensemble strategies. •Reproduce and validate the reported results of these approaches to establish a reliable baseline. •Adapt the identified neural network architectures specifically for tabular data applications where applicable. •Benchmark and evaluate the adapted models against state-of-the-art gradient-boosted tree algorithms as well as existing neural network baselines. •Enhance model performance through structured feature engineering and optimization of ensemble configurations. •Analyze the resulting performance to develop a deeper understanding of the suitability, strengths, and limitations of boosted neural networks for tabular data. 6 Chapter 2 Related Works The concept of boosting, an ensemble algorithm that combines multiple weak learners into a single strong learner has significantly influenced predictive performance on tabular data.This section reviews the foundational gradient boosted decision tree models and the main neural architectures that inform our approach, by considering the models that incorporate boosting-inspired methodologies[32] I The Boosting Framework The algorithmic foundation for boosting was established with the AdaBoost (Adaptive Boosting) algorithm by Freund & Schapire [16]. AdaBoost works by adaptively changing the weights of training instances, making subsequent weak learners (e.g., decision stumps) to focus on previously misclassified examples. This concept was expanded by Friedman[17] with the introduction to gradient boosting by modeling it as a numerical optimization problem in function space. In this, the model is built sequentially in a greedy, stage-wise manner where each new learner is trained to minimize the loss by correcting the errors of its predecessors. The core idea is to iteratively add weak learners to an ensemble, with each new model trained to predict the negative gradients (the ”pseudo-residuals”) of the current ensemble’s loss. Formally, given a di!erentiable loss function L(y,F(x)), the gradient boosting algorithm proceeds as follows[17] 1. Initialize the model with a constant value: F0(x) = arg min ω n ! i=1 L(yi,ε) (2.1) 7 2. For m=1to M(number of boosting rounds), perform the following steps: (a) Compute the pseudo-residuals for each instance i: r(m) i=↔"ϑL(yi,F(xi)) ϑF(xi)#F(x)=Fm→1(x) (2.2) This equation defines the “error” that the new weak learner must fit. (b) Fit a weak learner hm(x) (e.g., a decision tree) to the pseudoresiduals {r(m) i}. (c) Find the optimal step size εmvia line search: εm= arg min ω n ! i=1 L(yi,F m↓1(xi)+εh m(xi)) (2.3) (d) Update the model: Fm(x)=Fm↓1(x)+ϖ·εmhm(x) (2.4) Here, ϖ→(0,1] is the shrinkage or learning rate, a hyperparameter that controls the contribution of each weak learner to prevent overfitting.[17] II Gradient-Boosted Decision Trees (GBDTs) GBDT models form the baselines for any research on tabular data regression. These models build an ensemble of decision trees iteratively, and correct the errors made by the existing ensemble of trees. Each new tree is trained to model the gradient of the loss function.[32] Gradient-boosted decision tree(GBDT) models such as XGBoost [9], LightGBM[27], and CatBoost[38] are well known for their robust performance, e”ciency, and ability to handle heterogeneous features which made them dominate the area of tabular data prediction. XGBoost (Extreme Gradient Boosting) is an implementation of the gradient boosting framework that enhances the core GBDT algorithm through a second-order approximation of the loss function for more e”cient boosting, regularization (L1/L2) to prevent overfitting, and a sparsity-aware algorithm that e”ciently handles missing values. It learns the best direction to send a data point with a missing value at each split, allowing it to process sparse data without expensive preprocessing [9]. 8 at the last layer, F(x)=w↘hT(x), where wis the linear classifier attached to the output of the final block. BoostResNet modifies this by decomposing the network into weak module classifiers. Each module consists of a residual block ftpaired with its own linear classifier wt. Formally, the module classifier at step tis defined as ot(x)=w↘ tht(x), where ht(x) is the output of the residual mapping at that stage. The final prediction of the network is then represented as a telescoping sum of these module classifiers: F(x)= T ! t=0 ↼tot(x), with coe”cients ↼tchosen such that the ensemble exactly reconstructs the standard ResNet output [24]. To support the theory, Huang et al. (2018) conduct extensive experiments on CIFAR-10, CIFAR-100, and SVHN. The results show that BoostResNet “achieves test performance comparable to that of end-to-end ResNet training” while requiring substantially less GPU memory (p. 2065). Because the method trains blocks sequentially, only one shallow block needs to be loaded into memory at a time, making it e”cient for very deep architectures. The experiments also confirm the theoretical predictions: as the number of modules increases, training error decreases steadily, and test accuracy improves correspondingly. Overall, BoostResNet o!ers a theoretically grounded reinterpretation of residual networks by formalizing them as a boosting ensemble, with a telescoping-sum representation that preserves the ResNet output. Its sequential training procedure provides both computational savings and provable convergence guarantees, positioning it as a unique contribution in the literature on deep residual learning. Neural Oblivious Decision Ensembles(NODE) Neural Oblivious Decision Ensembles, is a deep learning architecture that combines the strengths of tree-based models and deep neural networks. Its main innovation is the creation of a di!erentiable ensemble of oblivious decision trees (ODTs), that makes end-to-end training via gradient descent possible.[36] An Oblivious Decision Tree (ODT) is a type of decision tree that has all nodes at the same depth and must use the same feature and the 15 same threshold for splitting. For a tree of depth dthis homogeneity transforms the tree from a standard branching structure into a decision table with 2dentries, where each entry represents a unique combination of binary decisions. While this constraint reduces the capacity of a single tree, it makes ensembles of ODTs highly e”cient for inference and remarkably resistant to overfitting.[36] The CatBoost algorithm [38], uses ODTs as weak learners in gradient boosting and this has contributed significantly to its success [38] This idea is expanded upon in the NODE architecture. Several di!erentiable ODTs make up a NODE layer. The key to di!erentiability lies in replacing the hard, non-di!erentiable operations of a standard decision tree (feature selection and binary routing) with soft, learnable alternatives: The NODE architecture generalizes ensembles of Oblivious Decision Trees(ODTs) into a fully di!erentiable framework that can be trained end-to-end via gradient descent. An ODT is a decision tree where all nodes at a given depth duse the same feature and threshold for splitting, e!ectively forming a decision table.[36] For the di!erentiable ODT,in a single NODE layer, there are m di!erentiable ODTs and the forward pass for one tree is built to approximate the function of a classical ODT while maintaining di!erentiability. In a classical ODT, the non-di!erentiable output is given by: h(x)=R(1(f1(x)↔b1),1(f2(x)↔b2),...,1(fd(x)↔bd))(2.10) where 1(·) is the Heaviside step function, fiis the selected feature at the i-th split, biis the corresponding threshold, and Ris a d-dimensional response tensor that holds the leaf values. To allow di!erentiability, NODE replaces these hard operations with soft, learnable alternatives.The hard selection of a single feature fiis replaced by a sparse, weighted combination of all features using the ↼-entmax transformation[35] applied to a learnable feature selection matrix F→Rd→n: ˆ fi(x)= n ! j=1 xj·entmaxε(Fij) (2.11) Here, ↼=1.5 is used to induce sparsity, ensuring the output closely mimics a hard feature selection. The classical Heaviside step function is replaced by a scaled, two-class variant of entmax, defined as: ↽ε(x) = entmaxε([x, 0]) (2.12) The soft routing probability for the i-th split is then: ci(x)=↽ε*fi(x)↔bi φi+,(2.13) 16 where biand φiare learnable scaling parameters. The value ci(x)→[0,1] represents the soft probability of taking the right branch at the i-th split. The final tree output is computed as a weighted sum over all leaves. The weights are given by the outer product of the soft choice vectors for all d splits, forming a ”choice tensor” C(x): C(x)="c1(x) 1↔c1(x)#⇐"c2(x) 1↔c2(x)#⇐···⇐"cd(x) 1↔cd(x)#(2.14) The final prediction of a single tree is: ˆ h(x)= ! i1,...,id≃{0,1}d Ri1,...,id·Ci1,...,id(x) (2.15) The output of a NODE layer is the concatenation of the outputs of all m trees: (ˆ h1(x),ˆ h2(x),...,ˆ hm(x)) Figure 2.4: Architecture of a single di!erentiable Oblivious Decision Tree (ODT) within a NODE layer.The single ODT inside the NODE layer. The splitting features and the splitting thresholds are shared across all the internal nodes of the same depth. The output is a sum of leaf responses scaled by the choice weights[36] Multi-Layer Hierarchical Architecture Several NODE layers stacked in a denseNet-like fashion forms the full NODE model[51] where each layer uses a concatenation of all previous layers.so,the input to the k-th layer is a concatenation of the original input features and the outputs from all previous layers 0 to k↔1. Due to this design,the model is able to learn hierarchical feature interactions, where trees in deeper layers can learn complex rules 17 based on high-level representations extracted by earlier layers. The final prediction is computed as the average of the outputs from all trees across all NODE layers: ˆy=1 K K ! k=1 ˆ hk(x) (2.16) Figure 2.5: The multi-layer NODE architecture with DenseNet-style feature reuse across layers.The NODE architecture, consisting of densely connected NODE layers. Each layer contains several trees whose outputs are concatenated and serve as input for the subsequent layer. Thefinal prediction is obtained by averaging the outputs of all trees from all the layers[36]. On Embeddings for Numerical Features in Tabular Deep Learning This researched introduced an underexplored domain of deep learning (DL) for tabular data being the embedding of numerical features. The authors introduce two approaches for constructing embeddings for numerical features: Piecewise Linear Encoding(PLE) and Periodic Activation(P) functions.[19] The PLE method is inspired by classical feature binning techniques, where the value range of a numerical feature is divided into intervals (bins), and the feature values are encoded in a piecewise linear manner. The results of the research demonstrate that the technique helps to improve the performance of deep learning models on tabular data. This approach allows simple MLP models to compete with more complex Transformer-based models.They also show that the integration of this approach into the deep learning pipeline produces state-of-the-art results on tabular DeepLearning closing the performance gap with GBDTs.[19] 18 Piecewise Linear Encoding(PLE): The design of PLE is motivated by the limitations of deep learning for tabular data. While Multilayer Perceptrons are known to be a universal approximator[24][19],their learning capabilities in practice are often hampered by optimization di”culties[39][19] Recent work by Tancik et al.[47] demonstrated that transforming the input space can significantly solve these optimization issues. This finding directly inspires the core premise of PLE: that altering the representation of original scalar numerical feature values can enhance the learning capabilities of tabular deep learning models[16]. The authors use the one-hot encoding algorithm, a method that is wildly successful for representing discrete entities (e.g., categorical features, NLP tokens)[19]. The one-hot representation sits at the opposite end of the spectrum from a scalar representation in the trade-o!between parameter e”- ciency and expressivity[19]. To test if a one-hot-like approach could benefit deep learning models on numerical data, PLE is designed as a continuous alternative to one-hot encoding, making it applicable to numerical features.[16][19] For a given numerical feature x, PLE defines Tbins(intervals) with boundaries b0,b 1,...,b T. The encoding transforms the scalar value xinto a Tdimensional vector[19] where the t-th element(bin) is calculated as: et=         0,if x<b t↓1and t>1, 1,if x≃btand t<T, x↔bt↓1 bt↔bt↓1 ,otherwise. The full PLE encoding vector is: PLE(x)=[e1,e 2,...,e T], 19 Figure 2.6: The Piecewise Linear Encoding (PLE) in action for T= 4[19]. The scalar input value xis mapped to a 4-dimensional vector [e1,e 2,e 3,e 4] based on its position within the bins, creating a structured and interpretable representation. The authors rely on the classic binning algorithms [12] and one of the two algorithms is unsupervised, while another one utilizes labels for constructing bins. Obtaining bins from quantiles (Unsupervised Binning): A natural baseline way to construct the bins is by splitting value ranges according to uniformly chosen empirical quantiles of the corresponding individual feature distributions[PLEPaper]: bt=qt T{xj i}for j→training set. Trivial bins of zero size are removed. Supervised Target-aware Binning(Building target-aware bins): This supervised approach employs training labels for constructing bins, identical in spirit to the C4.5 Discretization [29] algorithm[19]. For each feature, we recursively split its value range in a greedy manner using the target as guidance. This is equivalent to building a decision tree (which uses for growing only this one feature and the target) and treating the regions corresponding to its leaves as the bins for PLE[19]. we define bi 0=min j≃Jtrain xj iand bi T= max j≃Jtrain xj i[?]. Periodic Activation Function(P) This design projects scalar values into a periodic space using learnable frequencies.[47] For this method, a feature x, the embedding is constructed as: fi(x)=Periodic(x) = concat(sin(v),cos(v)),(2.17) 20 where v=(2⇀c1x, 2⇀c2x, . . . , 2⇀ckx),(2.18) and ciare trainable parameters initialized from a normal distribution ci⇒N(0,↽).(2.19) The hyperparameters ↽(initial frequency scale) and k(number of frequencies) are crucial and are tuned on the validation set. 21 Chapter 3 Methodology This project aims to bridge the performance gap between deep learning models and Gradient Boosted Decision Trees (GBDTs) on heterogeneous tabular data. To achieve this, we explore advanced concepts of: •Integrating a deep learning architecture that mimics tree ensembles and introducing a novel method for representing numerical features into this architecture. The core of our methodology is the integration of Piecewise Linear Encoding (PLE) into the Neural Oblivious Decision Ensembles (NODE) architecture, creating a powerful and di!erentiable model for tabular regression and classification. •Adapting deep neural network models to be more comparative with gradient boosting decision tree models(GBDTs) on heterogenous tabular data. I Integrated NODE + PLE Architecture The integration of the piecewise linear encoding(PLE)for numerical embeddings with the Neural Oblivious Decision Ensembles(NODE) architecture is an innovation of this work with the aim of addressing the limitations of deep learning on tabular data. While the components are powerful individually, we believe their combination will result in a more robust model. The integrated architecture leverages the strengths of both NODE and PLE to process data. The PLE model processes each numerical feature xiseparately by its own PLE module. Based on the chosen strategy (unsupervised PLE or supervised ), the scalar value is transformed into a high-dimensional, piecewise-linear representation PLE(xi)→RT. This vector is then passed 22 through a feature-specific linear layer to obtain a final dense embedding[37]: enum i=Wi·PLE(xi)+bi. Each categorical feature is processed through a standard embedding layer, mapping each category to a dense vector ecat j. All resulting numerical embeddings enum iand categorical embeddings ecat jare concatenated to form a joint, rich input representation vector z. The vector zis fed into the multi-layer NODE architecture[37]. The di!erentiable oblivious decision trees within each NODE layer now operate on this pre-enriched embedding space rather than on raw, normalized scalars. The DenseNet-like structure allows subsequent layers to use the transformed, high-level features learned by earlier layers, enabling the learning of complex hierarchical interactions between the PLE-encoded features. The final output is a simple average of the outputs from all trees across all NODE layers. Figure 3.1: Architecture of Integrated NODE + PLE. from the top left,input features are encoded via Piecewise Linear Encoding (PLE).All embeddings are concatenated and processed by a multi-layer NODE architecture, which learns hierarchical interactions through its ensembles of di!erentiable oblivious decision trees. The final prediction is an average of all tree outputs[37] Why PLE for NODE Integration The choice of Piecewise Linear Encoding (PLE) over other embeddings for integration with NODE is due to a shared inductive bias. Both methods are fundamentally based on the concept of splitting data on feature thresholds. This is the core operational principle that makes tree-based models like 23 Gradient Boosted Decision Trees (GBDTs) powerful on tabular data.NODE explicitly mimics this principle by constructing a di!erentiable ensemble of oblivious decision trees that learn di!erentiable splits. PLE directly encodes this principle into the input representation. It transforms a scalar value into a vector based on its position within learned bins (or intervals), defined by thresholds (b0,b 1,...,b T). Therefore, PLE provides an input signal that is already structured in the form that NODE is designed to understand. The tree-like output of PLE is a fit for the tree-based learning of NODE, creating a more powerful and aligned model than with other, less compatible embeddings. Why This Combination is Innovative Previous attempts to make deep learning competitive with GBDTs focused primarily on designing novel backbone architectures (e.g., Transformers[48, 20], other attention mechanisms) or on creating completely di!erentiable tree structures (e.g., NODE[37]). This work is unique by focusing on the critical but underexplored input representation layer for numerical features. Innovation lies in recognizing that: •The input representation (simple scalars) is a key bottleneck. •A technique (PLE) exists to create a much richer representation. •A specific architecture (NODE) exists that is perfectly suited to exploit this richer representation due to its tree-like nature. By integrating PLE with NODE, we aim to create a unified architecture that directly addresses the core weaknesses of Deep Neural Networks on tabular data, pushing their performance closer to and beyond that of state-of-the-art GBDTs II Boosted Fully Connected Networks (BFCN) A concerted attempt was made to extend the advantages of dynamic inference to heterogeneous tabular data by integrating the Boosted Dynamic Neural Network (BoostNet) model into a fully-connected network framework, which will be referred to as Boosted Fully Connected Networks (BFCN). This modification was essential for resolving the well-known issues that deep neural networks encounter when working with tabular datasets, 24 Table 4.3: Performance metrics of NODE + PLE and GBDTs and Machine Learning Models Method HELOC Adult HIGGS Covertype Cal. Housing Acc⇑AUC⇑Acc⇑AUC⇑Acc⇑AUC⇑Acc⇑AUC⇑MSE⇓ Linear Model 73.0±0.0 80.1±0.1 82.5±0.2 85.4±0.2 64.1±0.0 68.4±0.0 72.4±0.0 92.8±0.0 0.528±0.008 KNN 72.2±0.0 79.0±0.1 83.2±0.2 87.5±0.2 62.3±0.1 67.1±0.0 70.2±0.1 90.1±0.2 0.421±0.009 Decision Tree 80.3±0.0 89.3±0.1 85.3±0.2 89.8±0.1 71.3±0.0 78.7±0.0 79.1±0.0 95.0±0.0 0.404±0.007 Random Forest 82.1±0.3 90.0±0.2 86.1±0.2 91.7±0.2 71.9±0.0 79.7±0.0 78.1±0.1 96.1±0.0 0.272±0.006 XGBoost 83.5±0.2 92.2±0.0 87.3±0.2 92.8±0.1 77.6±0.0 85.9±0.0 97.3±0.0 99.9±0.0 0.206±0.005 LightGBM 83.5±0.1 92.3±0.0 87.4±0.2 92.9±0.1 77.1±0.0 85.5±0.0 93.5±0.0 99.7±0.0 0.195±0.005 CatBoost 83.6±0.3 92.4±0.1 87.2±0.2 92.8±0.1 77.5±0.0 85.8±0.0 96.4±0.0 99.8±0.0 0.196±0.004 Model Trees 82.6±0.2 91.5±0.0 85.0±0.2 90.4±0.1 69.8±0.0 76.7±0.0 --0.385±0.019 NODE + PLE 78.6±0.0 72.2±0.0 86.1±0.5 91.4±0.4 73.3±0.0 81.4±0.0 74.2±0.0 95.1±0.0 0.215±0.009 BFCN 71.0±1.1 77.6±1.3 85.1±0.5 90.9±0.4 70.6±1.1 77.7±0.9 73.4±0.2 93.2±0.6 0.277±0.017 NODE + PLE demonstrates that its performance suggests that it is highly dataset-dependent. NODE + PLE proves highly e!ective on the California Housing regression task, where it achieves an MSE (0.215) that is highly competitive with top-performing GBDTs like LightGBM (0.195), this shows its potential on numerical data. It also performs comparably well on the Adult dataset but is still outperformed by GBDTs. But its performance collapses on the HELOC dataset, where its AUC (72.2) is a huge outlier, falling below even simple linear models and suggesting probably a misconfiguration specific to that data. NODE + PLE on HIGGS and Covertype even though good but not the best, which is likely due to the subsampling of datasets and reduction of embedding dimensions in other to solve the issue of memory running out during training. Table 4.4: Performance metrics of NODE +PLE and Deep Learning Models. Method HELOC Adult HIGGS Covertype Cal. Housing Acc⇑AUC⇑Acc⇑AUC⇑Acc⇑AUC⇑Acc⇑AUC⇑MSE⇓ MLP 73.2±0.3 80.3±0.1 84.8±0.1 90.3±0.2 77.1±0.0 85.6±0.0 91.0±0.4 76.1±3.0 0.263±0.008 VIME 72.7±0.0 79.2±0.0 84.8±0.2 90.5±0.2 76.9±0.2 85.5±0.1 90.9±0.1 82.9±0.7 0.275±0.007 DeepFM 73.6±0.2 80.4±0.1 86.1±0.2 91.7±0.1 76.9±0.0 83.4±0.0 --0.260±0.006 DeepGBM 78.0±0.4 84.1±0.1 84.6±0.3 90.8±0.1 74.5±0.0 83.0±0.0 --0.856±0.065 NAM 73.3±0.1 80.7±0.3 83.4±0.1 86.6±0.1 53.9±0.6 55.0±1.2 --0.725±0.022 Net-DNF 82.6±0.4 91.5±0.2 85.7±0.2 91.3±0.1 76.6±0.1 85.1±0.1 94.2±0.1 99.1±0.0 - TabNet 81.0±0.1 90.0±0.1 85.4±0.2 91.1±0.1 76.5±1.3 84.9±1.4 93.1±0.2 99.4±0.0 0.346±0.007 TabTransformer 73.3±0.1 80.1±0.2 85.2±0.2 90.6±0.2 73.8±0.0 81.9±0.0 76.5±0.3 72.9±2.3 0.451±0.014 SAINT 82.1±0.3 90.7±0.2 86.1±0.3 91.6±0.2 79.8±0.0 88.3±0.0 96.3±0.1 99.8±0.0 0.226±0.004 RLN 73.2±0.4 80.1±0.4 81.0±1.6 75.9±8.2 71.8±0.2 79.4±0.2 77.2±1.5 92.0±0.9 0.348±0.013 STG 73.1±0.1 80.0±0.1 85.4±0.1 90.9±0.1 73.9±0.1 81.9±0.1 81.8±0.3 96.2±0.0 0.285±0.006 NODE 79.8±0.2 87.5±0.2 85.6±0.3 91.1±0.2 76.9±0.1 85.4±0.1 89.9±0.1 98.7±0.0 0.276±0.005 NODE + PLE 78.6±0.0 72.2±0.0 86.1±0.5 91.4±0.4 73.3±0.0 81.4±0.0 74.2±0.0 95.1±0.0 0.215±0.009 BFCN 71.0±1.1 77.6±1.3 85.1±0.5 90.9±0.4 70.6±1.1 77.7±0.9 73.4±0.2 93.2±0.6 0.277±0.017 BoostResNet 69.6.0±0.2 - 85.7±0.1 - ---- - Compared to the deep learning models,NODE + PLE achieves state of the art results on the California Housing regression task (MSE: 0.215) and competitive accuracy on the Adult dataset (Acc: 86.1). However it fails 31 greatly on the Heloc dataset and its results on Higgs and Covertype is good but not the best likely due to the subsampling of datasets and reduction of embedding dimensions in other to solve the issue of memory running out during training. 32 Chapter 5 Conclusion We explored three di!erent model architectures, its integration and tuning for tabular data. our model; NODE+PLE shows a significant potential of achieving competitive performance with GBDTs especially on California Housing regression task and also on the Adult dataset which validates the e!ectiveness of PLE. However its performance on the other datasets were not the best. Limitations •The results suggests that NODE + PLE’s performance is subjected to the kind of dataset as its performance was inconsistent across all the datasets. •The complexity of NODE+PLE, resulted in longer training times and higher memory requirements compared to Gradient Boosted Decision Trees (GBDTs) and simpler MLPs. An instance is with the Covertype and Higgs datasets which consistently run out of memory during training. •The PLE layer relies on fixed binning strategies. In an environment where feature distributions keeps on changing, this binning may not adapt quickly resulting to a decline in model performance. Future Works •It is still unclear when to use NODE + PLE. large datset or not? and what features? 33 •The highly invalid results of HELOC data set need to be further investigated. It suggests that PLE can memorize noise and need to be well regulated. 34 Bibliography [1] Rishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang, Ben Lengerich, Rich Caruana, and Geo!rey E. Hinton. Neural additive models: Interpretable machine learning with neural nets. In Advances in Neural Information Processing Systems, 2021. [2] Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1–10, Jul. 2019. [3] Edesio Alcoba¸ca, Felipe Siqueira, Adriano Rivolli, Lu´ıs P. F. Garcia, Je!erson T. Oliva, and Andr´e C. P. L. F. de Carvalho. Mfe: Towards reproducible meta-feature extraction. Journal of Machine Learning Research, 21(111):1–5, 2020. [4] Sercan ¨ O Arik and Tomas Pfister. Tabnet: Attentive interpretable tabular learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 6679–6687, 2021. [5] Sabuhi Badirli, Xiaowen Liu, Zhao Xing, Arko Bhowmik, Khanh Doan, and Sathiya S. Keerthi. Gradient boosting neural networks: Grownet. arXiv preprint, arXiv:2002.07971, 2020. [6] Pierre Baldi, Peter Sadowski, and Daniel Whiteson. Searching for exotic particles in high-energy physics with deep learning. Nature Communications, 5(1):1–9, Sep. 2014. [7] R.V. Borisov, T. Leemann, K. Seßler, J. Haug, M. Pawelczyk, and G. Kasneci. Deep neural networks and tabular data: A survey. 2022. [8] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish 37 Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020. [9] Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016. [10] Corinna Cortes, Xavier Gonzalvo, Vitaly Kuznetsov, Mehryar Mohri, and Scott Yang. Adanet: Adaptive structural learning of artificial neural networks. In International Conference on Machine Learning, pages 874–883. PMLR, 2017. [11] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. [12] James Dougherty, Ron Kohavi, and Mehran Sahami. Supervised and unsupervised discretization of continuous features. In Proceedings of the 12th International Conference on Machine Learning (ICML), pages 194–202, 1995. [13] Dheeru Dua and Casey Gra!. Uci machine learning repository. Online, 2017. [14] S. Elsayed, D. Thyssens, A. Rashed, H. S. Jomaa, and L. SchmidtThieme. Do we really need deep learning models for time series forecasting? arXiv preprint, arXiv:2101.02118, 2021. [15] FICO. Home equity line of credit (heloc) dataset, 2019. Accessed: Jun. 15, 2022. [Online]. Available: https://community.fico.com/s/ explainable-machine-learning-challenge. [16] Yoav Freund and Robert E. Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119–139, 1997. [17] Jerome H. Friedman. Greedy function approximation: a gradient boosting machine. Annals of Statistics, 29(5):1189–1232, 2001. [18] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. 38 [19] Yury Gorishniy, Ivan Rubachev, and Artem Babenko. On embeddings for numerical features in tabular deep learning. In Advances in Neural Information Processing Systems, volume 35, pages 24991–25004, 2022. [20] Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Revisiting deep learning models for tabular data. In Advances in Neural Information Processing Systems, volume 34, pages 18932–18943, 2021. [21] Lucien Grinsztajn, Edouard Oyallon, and Ga¨el Varoquaux. Why do tree-based models still outperform deep learning on typical tabular data? Advances in Neural Information Processing Systems, 35:507– 520, 2022. [22] Xiaojuan Qi Ruigang Yang Gao Huang Hao Li, Hong Zhang. Improved techniques for training adaptive deep networks. In Computer Vision and Pattern Recognition, 2019. [23] Xiang He, Ke Zhao, and Xiaowen Chu. Automl: A survey of the stateof-the-art. Knowledge-Based Systems, 212:106622, 2021. [24] Furong Huang, Jordan Ash, and Robert Schapire. Learning deep resnet blocks sequentially using boosting theory. In International Conference on Machine Learning, 2018. [25] Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. Multi-scale dense convolutional networks for e”cient prediction. arXiv preprint, arXiv:1703.09844, 2017. [26] Arlind Kadra, Marius Lindauer, Frank Hutter, and Josif Grabocka. Well-tuned simple nets excel on tabular datasets. In Advances in Neural Information Processing Systems, volume 34, 2021. [27] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. Lightgbm: A highly e”cient gradient boosting decision tree. 2017. [28] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [29] Ron Kohavi and Mehran Sahami. Error-based and entropy-based discretization of continuous features. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD), pages 114–119. AAAI Press, 1996. 39 [30] Roman Levin, Valeriia Cherepanova, Avi Schwarzschild, Arpit Bansal, C Bayan Bruss, Tom Goldstein, Andrew Gordon Wilson, and Micah Goldblum. Transfer learning with deep tabular models. In International Conference on Learning Representations, 2023. [31] J. Li, Y. Li, X. Xiang, S.-T. Xia, S. Dong, and Y. Cai. Tnt: An interpretable tree-network-tree learning framework using knowledge distillation. Entropy, 22(11):1203, 2020. [32] Duncan McElfresh, Saurabh Khandagale, Jose Valverde, Chaitanya V. Prasad, Goutham Ramakrishnan, Micah Goldblum, and Colin White. When do neural nets outperform boosted trees on tabular data? In Advances in Neural Information Processing Systems, volume 36, pages 76336–76369, 2023. [33] D. Medvedev and A. D’yakonov. New properties of the data distillation method when working with tabular data. arXiv preprint, arXiv:2010.09839, 2020. [34] Christopher Z. Mooney. Monte Carlo Simulation. SAGE, Newbury Park, CA, USA, 1997. [35] Ben Peters, Vlad Niculae, and Andr´e FT Martins. Sparse sequence-tosequence models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1504–1519. Association for Computational Linguistics, 2019. [36] Sergei Popov, Stanislav Morozov, and Artem Babenko. Neural oblivious decision ensembles for deep learning on tabular data. In International Conference on Learning Representations, 2020. [37] Sergey Popov, Stanislav Morozov, and Andrey Babenko. Neural oblivious decision ensembles for deep learning on tabular data. arXiv preprint, arXiv:1909.06312, 2019. [38] Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna V. Dorogush, and Andrey Gulin. Catboost: unbiased boosting with categorical features. 2018. [39] Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), pages 5301–5310. PMLR, 2019. 40 [40] D. Roschewitz, M.-A. Hartley, L. Corinzia, and M. Jaggi. Ifedavg: Interpretable data-interoperability for federated learning. arXiv preprint, arXiv:2107.06580, 2021. [41] Ivan Rubachev, Artem Alekberov, Yury Gorishniy, and Artem Babenko. Revisiting pretraining objectives for tabular deep learning. arXiv preprint, arXiv:2207.03208, 2022. [42] Debo Sahoo, Quang Pham, Jing Lu, and Steven C. Hoi. Online deep learning: Learning deep neural networks on the fly. arXiv preprint, arXiv:1711.03705, 2017. [43] J¨urgen Schmidhuber. Deep learning in neural networks: An overview. Neural Networks, 61:85–117, 2015. [44] Shai Shalev-Shwartz. Selfieboost: A boosting algorithm for deep learning. In International Conference on Machine Learning, 2014. [45] Xiang Shi, Johannes Mueller, Nicholas Erickson, Ming Li, and Alexander Smola. Multimodal automl on structured tables with text fields. In 8th ICML Workshop on Automated Machine Learning (AutoML), 2021. [46] Ravid Shwartz-Ziv and Amit Armon. Tabular data: Deep learning is not all you need. 2021. [47] Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In Advances in Neural Information Processing Systems (NeurIPS), 2020. [48] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, #Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017. [49] Le Yang, Xiaoyang Huang, Hao Zhang, Yu Wang, Zhiwei Liu, and Gao Huang. Resolution adaptive networks for e”cient inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 41 [50] Hanting Yu, Hao Li, Gang Hua, Gao Huang, and Honghui Shi. Boosted dynamic neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, 2023. [51] Hanting Yu, Hao Li, Gang Hua, Gao Huang, and Honghui Shi. Boosted dynamic neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 10989–10997, 2023. [52] Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. Deep learning based recommender system: A survey and new perspectives. ACM Computing Surveys, 52(1):1–38, 2019. 42