Full text
ChronoTab: Forecasting Multivariate Time Series with Tabular LLMs Alexandros Zeakis National Kapodistrian University of Athens, Athena Research Center Athens, Greece [email protected], [email protected] Giorgos Chatzigeorgakidis Athena Research Center Athens, Greece [email protected] Konstantinos Lentzos Athena Research Center Athens, Greece [email protected] Dimitrios Skoutas Athena Research Center Athens, Greece [email protected] Abstract—Forecasting future values in multivariate time series is a critical challenge in many application domains, such as agriculture, transportation, energy, etc. Recently, Large Language Models (LLMs) have been used for time series analysis tasks. However, they are typically limited to handling univariate time series, since their input has the form of a single sequence. The problem of understanding and processing multiple dimensions also arises when using LLMs to analyze tabular data, where each table has multiple columns. To improve the capability of LLMs to handle tabular data, some recent works have relied on fine-tuning an LLM to perform various table-related tasks, such as missing values imputation, on large-scale table corpora. In this paper, our goal is to investigate whether such an LLM that has been fine-tuned on tabular data can exhibit better performance when used for multivariete time series forecasting. In particular, we present ChronoTab, an approach that utilizes the tabular format of multivariate time series to create prompts with a specific context and then via model inference enables zero-shot multivariate forecasting. Our experiments show that ChronoTab improves forecasting accuracy, outperforming in most cases both pre-trained LLMs and state-of-the-art methods. Index Terms—large language models, multivariate time series, tabular data, forecasting I. INTRODUCTION Time series data are ubiquitous in modern society, driven by the vast increase of large-scale data recording across various domains. A time series is a sequence of data points recorded at successive intervals of time, either equally, or non-uniformly spaced. Forecasting is arguably the most critical time series analysis task, with applications spanning finance, healthcare, energy and climate science. Accurate time series forecasting can improve decision making by predicting future values of a given time series based on historical data patterns. Before the rise of Machine Learning, time series forecasting relied on traditional statistical methods that captured temporal dependencies and patterns [1]–[4]. The ARMA family, including AR [5], MA [6], ARMA [7], ARIMA [8], and SARIMA [9], were among the most widely used models. Over the past two decades, Machine and Deep Learning methods © A. Zeakis, G. Chatzigeorgakidis, K. Lentzos, D. Skoutas, 2025. This is the author’s version of the work. It is posted here for your personal use. Not for redistribution. The definitive version was published in 2025 IEEE 41st International Conference on Data Engineering Workshops (ICDEW), https://ieeexplore.ieee.org/document/11108161 have emerged, offering greater flexibility in modeling complex, nonlinear patterns [10]–[13]. The introduction of transformers [14] revolutionized NLP, enabling Large Language Models (LLMs) to leverage vast amounts of text data for pretraining and fine-tuning [15]. Notably, LLMs exhibit emergent abilities [16], enhancing problem-solving across diverse tasks without explicit training. Recently, the scientific community has focused on leveraging the potential of LLMs in domains other than NLP. One such family of methods are the tabular LLMs [17]– [20], which include pre-trained LLMs fine-tuned to handle tabular data, demonstrating promising capabilities on tasks such as the imputation of missing values. Furthermore, there are various approaches that attempt to employ LLMs in time series forecasting tasks by taking advantage of their potential to capture temporal relationships and time series dynamics [21] or use external knowledge about the data gained from LLMs to also enhance their performance [22]. However, most such approaches focus on univariate time series forecasting, applied on data containing a single observation per timestamp. A recent work [23] has attempted to extend the applicability of LLMs to forecasting multivariate time series data, by employing a multiplexing scheme to effectively reduce the dimensions of the time series. However, the model did not demonstrate significant improvements over traditional baselines when applied on time series of more than two dimensions. This strengthens the point that LLMs might not be the appropriate tool to solve time series forecasting [24], rather to be used in other tasks, such as time series reasoning or social understanding. In this paper, we examine the application of tabular LLMs on zero-shot multivariate time series forecasting, by utilizing their tabular nature, which can be leveraged by such models. An example can be found in Figure 1, where the table represents a multivariate time series that has three variables, each one being a column. Using the previous five rows, we would like to estimate the next row. To that end, we introduce ChronoTab, an approach that leverages the tabular format of multivariate time series to construct context-aware prompts, enabling zero-shot multivariate forecasting through model inference. With ChronoTab, we aim to answer the following key questions:
Date Variable 1 Variable 2 Variable 3 Context 2018-01-16 11.043 3.170 4.914 2018-01-19 11.839 3.118 5.176 2018-01-22 11.895 3.140 2.611 2018-01-25 13.863 2.729 -1.456 2018-01-28 12.254 2.376 0.957 Actual 2018-01-31 11.699 3.424 2.916 Predicted 2018-01-31 11.943 2.715 0.570 Fig. 1: An example of forecasting the next row of a multivariate time series that has 3 variables. •Q1: How does table size impact model accuracy? •Q2: How does incorporating time data affect accuracy? •Q3: Does additional domain-specific information improve accuracy? •Q4: Can the model effectively predict multiple rows at once? •Q5: Does the model produce consistent results when given the same prompt repeatedly? •Q6: Does fine-tuning on tabular data improve performance over a pre-trained model? •Q7: How does ChronoTab compare to baseline and stateof-the-art methods in multivariate time series forecasting? The remainder of this paper is organized as follows: Section II reviews related work. In Section III-B we describe the framework under examination and Section IV presents the experimental evaluation. Finally, Section V concludes the paper. II. RELATED WORK Time Series Analysis with LLMs. The recent introduction of LLMs has triggered a surge of interest in the scientific community, with many attempts to leverage their potential in other domains and contexts than NLP, e.g., in healthcare [25], [26], [27], financial modeling [28], [29], [30], and education and research [31], [32]. Time series analysis is no exception; the LLMs applicability has been studied in various time series analysis tasks [33] and application domains [34]. Univariate Time Series Forecasting. TIME-LLM [35] repurposes LLMs for time series forecasting without retraining by reprogramming input data into text-based representations. It introduces the Prompt-as-Prefix (PaP) mechanism, enriching inputs with natural language instructions to improve reasoning. The model remains frozen, updating only input and output parameters, and supports shortand long-term forecasting, including fewand one-shot learning. LLMTIME [21] applies zero-shot forecasting by encoding time series as sequences of numerical digits, treating forecasting as a next-token prediction task. The model processes tokenized, rescaled data, restricting outputs to digits and commas. Predictions are generated by sampling multiple outputs and selecting the median. This approach demonstrates that LLMs can achieve forecasting performance comparable to specialized models without fine-tuning. Another approach is presented in [22], where the authors confirm the observation that time series with clear patterns and trends can be accurately predicted using LLM in a zero-shot setting. However, they attempt to enhance the performance of such models by incorporating external knowledge and transforming numerical data into natural language. These enhancements lead to an improvement in model performance even when applied on datasets lacking periodicity. Chronos [36] introduces a pretrained probabilistic time series framework that tokenizes values via scaling and quantization, training transformer-based models on tokenized sequences. Pretrained on diverse datasets, including synthetic data, Chronos demonstrates strong performance on seen datasets and competitive zero-shot forecasting on new ones, highlighting the potential of pretrained models for time series forecasting. Recently, however, the authors in [24] critically evaluate the use of LLMs for time series forecasting tasks. They conducted extensive ablation studies on such approaches, concluding that replacing LLMs with simpler models (e.g., basic attention layers) yields comparable or even better performance. Multivariate Time Series Forecasting. The above-mentioned approaches for applying LLMs to time series forecasting tasks focus solely on univariate time series. MultiCast [23] extends LLMTIME for multivariate time series forecasting using three token multiplexing techniques to reduce input dimensionality while preserving patterns. It also integrates SAX-based [37] quantization to enhance computational efficiency. MOIRAI [38] introduces architectural enhancements to time series Transformers to enable universal forecasting across diverse tasks. Trained on the Large-scale Open Time Series Archive (LOTSA), it achieves competitive zero-shot performance against full-shot models. Tabular LLMs. The wide use of LLMs led to the utilization of LLMs in tasks related to tabular data. Pre-trained Tabular LLMs: Leveraging pre-trained LLMs, CHORUS [39] is a framework that focuses on three tasks: (i) Table-class selection assigns an ontology class to the whole table; (ii) Column-Type Annotation assigns an ontology class to a single column; and (iii) Join-Column prediction: given a second table, it can find the pair of columns on which these two tables could be linked. All corresponding prompts were created with sampling and in-context learning and were used on a pre-trained GPT-3.5. Fine-tuned Tabular LLMs: For tasks involving fine-tuned models, UNIFIEDSKG [40] unifies heterogeneous structured knowledge grounding (SKG) tasks by proposing a framework that transforms 21 SKG tasks into a text-to-text format. By fine-tuning T5 models with multi-task learning and prefixtuning, it achieves state-of-the-art results on most tasks and highlights the challenges of zero-shot and few-shot SKG for models like GPT-3 and Codex. To handle the long-context challenge, TableLlama [17] fine-tunes Llama 2 (7B) using
instruction tuning with LongLoRA on data from diverse tablebased tasks. TableLlama achieves competitive performance on in-domain tasks and significant gains on out-of-domain datasets, demonstrating its adaptability and generalizability. TableGPT [18], apart from table-related tasks, it also enables tasks like question answering, data manipulation and report generation. To do so, it utilizes global tabular representations, that allows LLMs to understand the entire table beyond its meta-information. Extending TableGPT, TableGPT2 [19], based on the Qwen-2.5 architecture, improves upon the original TableGPT by incorporating continual pretraining (CPT) with domain-specific data, enhancing its ability to handle business intelligence (BI) tasks. This pretraining, followed by finetuning, strengthens its capacity for complex data analysis and code generation. Finally, regarding proprietary models, TableGPT [20] fine-tunes GPT-3.5 and ChatGPT and showcases better performance compared to their pre-trained counterparts and stronger generalizability. III. CHRONOTAB A. Problem Formulation Multivariate time series forecasting aims to predict future values of multiple interdependent variables based on past observations. Given a time series dataset with Nvariables observed over Ttime steps, the objective is to learn a function fthat maps past observations to future values. ˆ Xt+1:t+τ=f(Xt−w+1:t) where Xt−w+1:t∈Rw×Nrepresents the past wtime steps of Nvariables, and ˆ Xt+1:t+τ∈Rτ×Ndenotes the predicted future values for the next τsteps. B. Framework In this subsection we introduce ChronoTab as a framework and its corresponding steps. ChronoTab includes three main steps: •Context Definition: In this step we define the context of the table that will be used in the prompt. There are two main features here: –Context Window c: This parameter involves how many rows preceding the testing row will be used. It aims to improve the performance of the model, by balancing the knowledge about the data and the prompt size, since limiting the value of context window offers the model very little knowledge about the data and might result in poor performance, while, on the other hand, increasing the window might result in a large prompt which will disorientate the model. –Time-Awareness t: Time-Awareness aims to improve the performance of the model by omitting time references. In tabular data, time is typically stored in a separate column, which can be excluded during serialization. Since models are trained to generate the next possible row(s), treating time as a column may #Task Description: Based on the following table that contains multivariate time-series, predict the values only for the next row for each column. Do not explain your answer and do not provide any code, return only the result. #Input: **Table:** {} Return the final result as JSON in the format {}. #Output:" Fig. 2: A domain-agnostic prompt template for ChronoTab with h= 1. The corresponding table is serialized and added in the prompt, as well as a JSON containing the keys / columns of the specific table. impact their performance. Notably, this affects only the inclusion or exclusion of the time column, not the order of the data, which preserves the original time sequence. These two features can also be considered as a horizontal and vertical filtering of the table accordingly. •Prompt Construction: In this section we construct the prompt that will be the input to the model. We observe two features here: –Domain-Awareness d: Domain-Awareness (d) attempts to improve the performance of the model, by using extra information specific about the nature of the data. This information involve the nature of the dataset and each column separately. This is something attempted before in [22], but they used generated information about the data, whereas we used human-curated information per dataset. –Horizon Length h: Horizon Length defines how many rows will be predicted per prompt. While predicting a single row each time might seem more focused on effectiveness, it requires more prompts, whereas asking for more rows per prompt can help improve efficiency. After these two features, the table is serialized and inserted into the prompt. Templates of the prompt can be showcased in Figures 2 and 3. •Model Inference This final step involves the generation of the response. We define two features: –Prompt Repetitions r: Prompt repetitions (r) help assess model reliability. Since models are not deterministic, even small variations in predicted values can lead to different errors. Models that generate similar answers with the same prompt are more preferable, since the results are then more reliable. –Model: For ChronoTab we select TableGPT2 [19], which is a model fine-tuned to comprehend tabular data and assist in many table-related tasks, such as missing values imputation. The framework of ChronoTab can be found in Figure 4.
#Task Description: Based on the following table that contains multivariate time-series, predict the next {} values as an array for each column. You can use the Table Information as you see fit. Do not explain your answer and do not provide any code, return only the result. #Input: **Table:** {} **Table Information:** {} Return the final result as JSON in the format {}. #Output:" Fig. 3: A domain-aware prompt template for ChronoTab with h > 1. The corresponding table is serialized and added in the prompt, as well as a JSON containing the keys / columns of the specific table. IV. EVALUATION This section presents the results of our experiments. We first explain how we set up our tests and assess the suggested methods. A. Experimental Setup 1) System: Our code and datasets used are publicly available.1All experiments were executed on a server with Ubuntu 20.04, AMD Ryzen Threadripper 3960X 24Core processor, 256 GB RAM and an RTX 4090 GPU. 2) Datasets: We employ four real-world multivariate time series datasets. Their characteristics are summarized in Table I. TABLE I: Characteristics of Datasets. Dataset Domain Dimensions Length Test Gas Rate Energy 2 296 60 ETT Temperature 3 242 49 Weather Weather 4 109 22 ILI Healthcare 8 217 55 Gas Rate: This is a 2-dimensional dataset containing carbon dioxide (CO2) emissions. The first dimension contains the input CO2 measurements (ft3/min) in a gas furnace. The second dimension contains the output CO2 percentage. The dataset is obtained from the darts library.2Of course, the two dimensions are correlated, which makes this dataset ideal for multivariate forecasting. ETT: This multivariate time series is part of the Electricity Transformer Dataset (ETDataset).3It contains hourly measurements of various metrics, which were resampled on a 3-day basis, for a total of 242 timestamps. From this dataset, we extracted 3 dimensions of electricity measurements, specifically the High UseFul Load (HUFL), 1https://github.com/alexZeakis/ChronoTab 2https://unit8co.github.io/darts 3https://github.com/zhouhaoyi/ETDataset High UseLess Load(HULL), and Oil Temperature (OT). Again, the dimensions are correlated. Weather: The weather dataset4was generated by the Max Planck Institute and contains 21 weather-related metrics obtained from a weather station located in Germany. From the 21 variables, we extracted the air temperatures (Tlog) measured in Celsius degrees, the water vapor concentration (H2OC) measured in mmol/mol, the saturation water vapor pressure (VPmax), measured in mbar, and the potential temperature (Tpot) measured in Kelvin degrees. Again, being weather-related, all dimensions are correlated. ILI: ILI5includes the weekly recorded influenza-like illness (ILI) patients data from Centers for Disease Control and Prevention of the United States, which describes the ratio of patients seen with ILI and the total number of the patients. We used data between 2020 and 2025 and 8 columns: 5 age-related incidents numbers, total patients examined and 2 columns with the weighted and unweighted ILI score. 3) Parameters: The parameters utilized in our experimental assessment are listed in Table II. For each parameter, we performed tuning tests to establish their ranges and default values, which are highlighted in bold within the table. TABLE II: Characteristics of Parameters. Parameter Range Context Window c[5, 10, 20] Prompt Repetitions r[1, 5, 10] Domain-awareness d[Agnostic, Aware] Time-awareness t[Agnostic, Aware] Horizon Length h[1, 3, 5] 4) Metrics: We used two established evaluation metrics in time series forecasting: •Root Mean Squared Error (RMSE): RMSE = q1 nPn i=1(yi−ˆyi)2 •Mean Absolute Percentage Error (MAPE): MAPE = 1 nPn i=1 yi−ˆyi yi ×100 where yiis the actual value, ˆyiis the predicted value at timestamp iand nis the number of timestamps on which forecasting was applied. To be noted, while RMSE operates on unnormalized values, MAPE normalizes errors relative to the actual values, making the combination of both metrics necessary for a comprehensive evaluation of the results. As a consequence, for certain cases—such as ILI variables—where the original values are large, the error can be substantial and is denoted as “>100”. Similarly, when the original values are very small, MAPE can be large and is also denoted as “>100”. 5) Answer Engineering: TableGPT2 is fine-tuned to return results in JSON format with different keys per task. Pre4https://www.bgc-jena.mpg.de/wetter/ 5https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html
Test Context Definition Time-Awareness Context-Window Test Prompt Construction Domain-Awareness Horizon Length Domain #Task Description: #Input: *Table*: #Output: Model Inference Prompt Repetition Model Forecast 1 Forecast 2 Forecast n Fig. 4: The steps of ChronoTab. trained models, that were used in Section IV-C, generate responses in natural language and thus sometimes require parsing. For example, sometimes they fail to return the proper column names in the response or they add irrelevant characters in the response, thus making the cleaning necessary. In some rare occurrences, the response cannot be parsed effectively; in these cases, it is disregarded. 6) Interface: ChronoTab also offers a Graphical User Interface, which was implemented with Streamlit6. The user can load their dataset as a dataframe, inspect the data and visualize each time series. Then, they specify all the parameters that were used in the experiments, i.e. Context Window c, Time-Awareness t, DomainAwareness d, Horizon Length hand Prompt Repetition r, as also the range in which they want to forecast. Finally, the user can also specify which model they want to run, as well as settings as to where a model is hosted, in case they want to run external models, such as OpenAI7models. After the forecasting is completed, the user can save their results as a CSV locally and also visualize the time series - original and forecasted - per variable in the table. B. Impact of parameters In this set of experiments we evaluate each parameter of our framework, by maintaining the default value for all parameters and tweaking only one at a time. The default settings of all parameters can be found in Table II. Context Window Context Window (c) is responsible for the number of rows of the table that is included in the prompt. TableGPT2 is sensitive to sample size, so increasing the table size might affect the effectiveness of the model. We tried three values for Context Window: 5, 10 and 20. The results can be found at Table III. Regarding the Gas Rate dataset, Context Window has limited effect on the column GasRate, since the RMSE is very small and does not vary much, while the CO2 appears to perform better when c= 10. This can be explained by the fact 6https://streamlit.io/ 7https://openai.com/ that c= 5 is too limited to supply sufficient information, while c= 20 increases the context within the prompt and leads to a poorer performance. Similarly, within ETT dataset, a balance occurs when c= 10 for all three columns. On the other hand, on the Weather dataset, in two of the columns, VPmax and Tpot, we notice that the best behavior occurs when c= 20, which may be due to the model requiring more rows and data to capture relevant patterns. Nonetheless, for the other two rows, Tlog and H2OC, we notice the same optimal behaviour for c= 10. Regarding MAPE, we do not observe significant differences across different values of c. However, an interesting observation is that the ILI dataset exhibits the lowest MAPE values compared to the other three datasets. TABLE III: Performance on Context-Window. Context Window 5 10 20 Dataset Variable RMSE MAPE RMSE MAPE RMSE MAPE Gas GasRate 0.29 >100 0.3 >100 0.3 >100 CO2 0.93 1 0.85 1 0.89 1 Mean 0.61 >100 0.57 >100 0.6 >100 ETT HUFL 2.72 39 2.43 36 2.59 40 HULL 1.1 42 1.04 41 0.97 36 OT 2.46 29 2.41 29 2.59 32 Mean 2.09 37 1.96 35 2.05 36 Weat. Tlog 2.86 19 2.53 17 2.54 15 H2OC 1.75 24 1.68 24 1.69 23 VPmax 2.5 29 2.37 29 2.2 24 Tpot 4.92 1 5.24 1 4.61 1 Mean 3.01 18 2.96 17 2.76 16 ILI W. ILI 0.51 6 0.52 6 0.52 6 U. ILI 0.48 5 0.5 6 0.49 6 AGE-1 >100 7 >100 8 >100 8 AGE-2 >100 11 >100 10 >100 11 AGE-3 >100 7 >100 7 >100 8 AGE-4 >100 9 >100 10 >100 12 AGE-5 >100 8 >100 11 >100 12 TOTAL >100 5 >100 5 >100 4 Mean >100 7 >100 8 >100 8 Time-awareness Time-Awareness (t) attempts to improve the performance of the model, by omitting the time references. For each dataset we have two variations: Time-Aware and TimeAgnostic. The results can be found at Table IV. While in Gas Rate, ETT and ILI we observe very small differences between the two variations (favoring either strategy), in Weather we have a distinct improvement when using Time-Aware, since using a Time-Agnostic strategy leads to
0 50 100 150 200 250 46 48 50 52 54 56 58 60 c=5 c=10 c=20 Original Time CO2% (a) Gas Rate dataset – CO2%. Accurate predictions with small RMSE. Jul 2016 Jan 2017 Jul 2017 Jan 2018 0 5 10 15 c=5 c=10 c=20 Original Time HUFL (b) Eletricity dataset – HUFL. Inaccurate predictions with medium RMSE. Jul 2023 Aug 2023 Sep 2023 Oct 2023 Nov 2023 Dec 2023 Jan 2024 Feb 2024 265 270 275 280 285 290 295 300 c=5 c=10 c=20 Original Time Tpot (K) (c) Weather dataset – Tpot. Inaccurate predictions with large RMSE. Fig. 5: Forecasted time series for different values of Context Window c. an average 14% increase in RMSE and 30% in MAPE. This indicates that weather data are more time-dependent, likely due to seasonality or trends. TABLE IV: Performance on Time-Awareness. Time-Awareness agnostic aware Dataset Variable RMSE MAPE RMSE MAPE Gas GasRate 0.31 >100 0.3 >100 CO2 0.96 1 0.85 1 Mean 0.64 >100 0.57 >100 ETT HUFL 2.4 37 2.43 36 HULL 1.14 52 1.04 41 OT 2.42 28 2.41 29 Mean 1.99 39 1.96 35 Weat. Tlog 2.95 20 2.53 17 H2OC 2.02 30 1.68 24 VPmax 2.69 36 2.37 29 Tpot 5.97 1 5.24 1 Mean 3.41 22 2.96 17 ILI W. ILI 0.55 6 0.52 6 U. ILI 0.53 6 0.5 6 AGE-1 >100 8 >100 8 AGE-2 >100 11 >100 10 AGE-3 >100 8 >100 7 AGE-4 >100 9 >100 10 AGE-5 >100 9 >100 11 TOTAL >100 3 >100 5 Mean >100 7 >100 8 Domain-awareness Domain-Awareness (d) attempts to improve the performance of the model, by using specific additional information about the nature of the data. For each dataset, we added the information provided earlier in Section IV-A. Thus, we have two variations: Domain-Aware and Domain-Agnostic. The results can be found at Table V. With the exception of two columns, HUFL and OT in ETT, where the RMSE decreases by 2% and 1% accordingly, all other columns do not show any improvement when using a domain-aware prompt, with the major difference spotted in Tpot with an increase in 20%. In ILI, MAPE indicates a similar behaviour, with an averange increase of 13%. While the additional information did not prove beneficial, exploring other types of domain knowledge may yield improvements in future work. Nonetheless, for that reason, we adopt a domainagnostic strategy for the rest of the experiments. Horizon Length Horizon Length (h) has a goal to increase the efficiency of ChronoTab, by batching multiple timestamps in a single prompt. For each dataset we used three values: predicting 1, 3 and 5 ahead. Within each group of hpredictions, we calculated the RMSE. The results can be found at Table VI. It is apparent that in all datasets, predicting more than one timestamps yields a larger error. For example, in the Gas Rate dataset, we observe a 98% average increase in RMSE from h= 1 to h= 3 and a 175% average increase in RMSE from h= 1 to h= 5. A similar trend follows ETT and Weather, while on ILI, regarding MAPE, we notice an average increase of 125% and 225% accordingly. This suggests that while TableGPT2 can generate a single row given a specific context and error, generating multiple rows based on predicted values accumulates errors, leading to poor performance. Prompt Repetitions Prompt repetitions (r) helps assess the reliability of a model. Since models are not deterministic,
TABLE V: Performance on Domain-Awareness. Domain-Awareness aware agnostic Dataset Variable RMSE MAPE RMSE MAPE Gas GasRate 0.39 >100 0.3 >100 CO2 0.9 1 0.85 1 Mean 0.64 >100 0.57 >100 ETT HUFL 2.38 36 2.43 36 HULL 1.11 44 1.04 41 OT 2.37 30 2.41 29 Mean 1.95 37 1.96 35 Weat. Tlog 2.75 17 2.53 17 H2OC 1.85 26 1.68 24 VPmax 2.52 27 2.37 29 Tpot 6.25 1 5.24 1 Mean 3.34 18 2.96 17 ILI W. ILI 0.57 7 0.52 6 U. ILI 0.55 7 0.50 6 AGE-1 >100 8 >100 8 AGE-2 >100 12 >100 10 AGE-3 >100 8 >100 7 AGE-4 >100 11 >100 10 AGE-5 >100 12 >100 11 TOTAL >100 4 >100 5 Mean >100 9 >100 8 TABLE VI: Performance on Horizon Length. Horizon Length 1 3 5 Dataset Variable RMSE MAPE RMSE MAPE RMSE MAPE Gas GasRate 0.3 >100 0.66 >100 0.77 >100 CO2 0.85 1 1.6 2 2.36 3 Mean 0.57 >100 1.13 >100 1.57 >100 ETT HUFL 2.43 36 2.51 39 2.54 40 HULL 1.04 41 1.4 61 1.64 74 OT 2.41 29 2.61 31 2.96 37 Mean 1.96 35 2.17 44 2.38 50 Weat. Tlog 2.53 17 3.36 23 6.55 30 H2OC 1.68 24 2.31 36 2.56 42 VPmax 2.37 29 3.2 42 3.67 51 Tpot 5.24 1 7.39 2 8.81 2 Mean 2.96 17 4.06 25 5.4 31 ILI W. ILI 0.52 6 0.87 12 1.09 15 U. ILI 0.50 6 0.83 11 1.04 14 AGE-1 >100 8 >100 14 >100 17 AGE-2 >100 10 >100 17 >100 21 AGE-3 >100 7 >100 16 >100 22 AGE-4 >100 10 >100 32 >100 52 AGE-5 >100 11 >100 34 >100 58 TOTAL >100 5 >100 6 >100 7 Mean >100 8 >100 18 >100 26 we average responses across repetitions for a more reliable evaluation. We tried three values for Prompt Repetitions: 1, 5 and 10. The results can be found at Table VII. In almost all datasets we observe that both RMSE and MAPE improve when using r > 1. For example, in ETT we notice an average drop of 9% and 10% from r= 1 to r= 5 and r= 10 accordingly. In ILI, a similar average behaviour occurs for MAPE, with an average decrease of 11% and 22%. While r= 10 seems to have the best behaviour in both RMSE and MAPE in almost all datasets, with the exception of Weather data, the extra cost of running twice as many prompts and thus increasing the cost of Model Inference should be avoided, thus we adopt r= 5 for the rest of the experiments. TABLE VII: Performance on Prompt Repetitions. Prompt Repetitions 1 5 10 Dataset Variable RMSE MAPE RMSE MAPE RMSE MAPE Gas GasRate 0.31 >100 0.3 >100 0.29 >100 CO2 0.87 1 0.85 1 0.86 1 Mean 0.59 >100 0.57 >100 0.58 >100 ETT HUFL 2.71 39 2.43 36 2.41 38 HULL 1.03 38 1.04 41 1.06 42 OT 2.71 33 2.41 29 2.32 28 Mean 2.15 37 1.96 35 1.93 36 Weat. Tlog 2.97 19 2.53 17 2.51 16 H2OC 2.15 33 1.68 24 1.65 23 VPmax 2.56 31 2.37 29 2.18 25 Tpot 5.97 1 5.24 1 6.84 1 Mean 3.41 21 2.96 17 3.3 16 ILI W. ILI 0.61 7 0.52 6 0.54 6 U. ILI 0.57 6 0.50 6 0.52 6 AGE-1 >100 9 >100 8 >100 8 AGE-2 >100 11 >100 10 >100 11 AGE-3 >100 7 >100 7 >100 7 AGE-4 >100 11 >100 10 >100 8 AGE-5 >100 12 >100 11 >100 8 TOTAL >100 6 >100 5 >100 5 Mean >100 9 >100 8 >100 7 C. Comparison with Baselines In this set of experiments we compare ChronoTab with other approaches or models as baselines to establish ChronoTab’s effectiveness. Pre-trained Models In this section we study the performance of TableGPT2 as our core model, i.e. a model that is fine-tuned on comprehending tabular data within a prompt and solving various table-related tasks, such as missing values imputation. To evaluate its performance we compare it against two pretrained models, Llama-3.1:8b8and Qwen-2.5:32b.9We use the former as a popular smaller LLM, used in many comparisons, while the latter has a more promising performance due to its higher parameters. The results can be found in Table IX. Comparing Qwen-2.5 and Llama-3.1, Qwen2.5 seems to have a better performance in almost all datasets. This can be explained by the fact that Qwen is a larger model, thus it understands better the assignment than Llama-3.1 due to its emergent abilities. TableGPT2 remains consistently competitive across all datasets, suggesting that fine-tuning can reduce the dependence on pre-trained model size, leading to more stable performance. State-of-the-art In this section we compare ChronoTab with baselines and state-of-the-art methods: •ARIMA [8]: Autoregressive Integrated Moving Average (ARIMA) is one of the most widely used univariate time series forecasting methods. •MultiCast [23]: A method that also utilizes LLMs for multivariate time series, where we used Llama-3.1:8b as the inference model. •Chronos [36]: A pretrained approach for univariate time series forecasting. We used the ”chronos-t5-small” model. 8https://ollama.com/library/llama3.1:8b 9https://ollama.com/library/qwen2.5:32b
TABLE VIII: Comparison with State-of-the-art. Method ChronoTab MultiCast Chronos MOIRAI ARIMA Dataset Variable RMSE MAPE RMSE MAPE RMSE MAPE RMSE MAPE RMSE MAPE Gas GasRate 0.3 >100 0.62 >100 0.26 >100 0.31 >100 0.9 >100 CO2 0.85 1 2.79 4 19.38 6 2.74 3 4.74 6 Mean 0.57 >100 1.71 >100 9.82 >100 1.52 >100 2.82 >100 ETT HUFL 2.43 36 5.27 >100 2.57 39 3.25 45 9.49 >100 HULL 1.04 41 1.52 76 0.74 26 0.88 31 2.7 83 OT 2.41 29 7.87 95 2.53 31 2.88 34 11.21 >100 Mean 1.96 35 4.89 92 1.94 32 2.34 37 7.8 >100 Weat. Tlog 2.53 17 3.11 22 2.16 13 2.63 17 3.78 29 H2OC 1.68 24 1.86 26 1.61 22 2.81 30 3.73 63 VPmax 2.37 29 2.74 33 2.11 24 2.63 25 4.64 48 Tpot 5.24 1 6.43 1 6.08 1 49.04 7 4.88 1 Mean 2.96 17 3.54 20 2.99 15 14.28 20 4.26 35 ILI W. ILI 0.52 6 2.42 >100 0.47 6 0.72 12 1.8 68 U. ILI 0.50 6 2.29 >100 0.44 6 0.83 14 2.38 >100 AGE-1 >100 8 >100 87 >100 8 >100 12 >100 63 AGE-2 >100 10 >100 >100 >100 9 >100 15 >100 >100 AGE-3 >100 7 >100 >100 >100 6 >100 12 >100 64 AGE-4 >100 10 >100 97 >100 11 >100 13 >100 94 AGE-5 >100 11 >100 90 >100 11 >100 10 >100 62 TOTAL >100 5 >100 6 >100 3 >100 5 >100 4 Mean >100 8 >100 92 >100 8 >100 12 >100 75 TABLE IX: Comparison with pre-trained models. Model Llama-3.1 Qwen-2.5 TableGPT2 Dataset Variable RMSE MAPE RMSE MAPE RMSE MAPE Gas GasRate 0.78 >100 0.25 >100 0.3 >100 CO2 1.67 2 0.7 0 0.85 1 Mean 1.22 >100 0.47 >100 0.57 >100 ETT HUFL 2.84 46 2.36 39 2.43 36 HULL 1.03 41 1.04 41 1.04 41 OT 2.47 32 2.29 28 2.41 29 Mean 2.11 40 1.9 36 1.96 35 Weat. Tlog 2.8 18 2.77 18 2.53 17 H2OC 1.73 22 1.91 28 1.68 24 VPmax 2.4 27 2.46 32 2.37 29 Tpot 4.95 1 5.74 1 5.24 1 Mean 2.97 17 3.22 20 2.96 17 ILI W. ILI 1.02 9 0.46 5 0.52 6 U. ILI 0.95 9 0.45 5 0.50 6 AGE-1 >100 11 >100 6 >100 8 AGE-2 >100 14 >100 9 >100 10 AGE-3 >100 9 >100 6 >100 7 AGE-4 >100 8 >100 6 >100 10 AGE-5 >100 9 >100 5 >100 11 TOTAL >100 5 >100 4 >100 5 Mean >100 9 >100 6 >100 8 •MOIRAI [38]: A transformer-based model for multivariate time series. We used the ”moirai-moe-1.0-R-small” model. The results can be found in Table VIII. First, regarding multivariate methods, MOIRAI is almost always better to MultiCast in both RMSE and MAPE, with the exception of Weather, where MOIRAI exhibits a 4x larger RMSE on average, but balances it with an equal average MAPE. Nonetheless, compared to ChronoTab, they both perform worse in every dataset, since it reaches a better RMSE and MAPE on average. Regarding univariate methods, Chronos has a similar behaviour, but keeping in mind that we use inference as many times as the original columns, ChronoTab is more efficient. Finally, we observe that ARIMA has the worst performance in almost all datasets with the highest average RMSE and MAPE. V. CONCLUSIONS In this paper we have introduced ChronoTab, a framework for forecasting values for multivariate time series with the use of Tabular LLMs. ChronoTab consists of three components: Context Definition, Prompt Construction and Model Inference. For each component we have addressed corresponding parameters, which were evaluated experimentally: In Context Definition, we showed that Time-Awareness is essential for datasets that contain multiple variables and Context Window needs a balanced choice, since too small or too large might affect the performance of the model. In Prompt Construction, we noticed that Domain-Awareness did not offer any improvement of the model and, regarding Horizon Length, a model performs best when predicting only one row per prompt. Finally, in Model Inference, we observe that Prompt Repetition strengthens the performance of the model. Regarding comparison with pretrained models and state-of-the-art, ChronoTab shows better performance in most occasions. As a next step, we aim to explore few-shot techniques for multivariate time series forecasting by leveraging RAG strategies to retrieve relevant examples. ACKNOWLEDGEMENT This work was partially funded by the EU Horizon Europe projects STELAR (101070122) and DT4GS (101056799). REFERENCES [1] Z. Liu, Z. Zhu, J. Gao, and C. Xu, “Forecast methods for time series data: A survey,” IEEE Access, vol. 9, pp. 91 896–91 912, 2021. [2] J. G. De Gooijer and R. J. Hyndman, “25 years of time series forecasting,” International Journal of Forecasting, vol. 22, no. 3, pp. 443–473, 2006. [3] J. Kuvulmaz, S. Usanmaz, and S. N. Engin, “Time-series forecasting by means of linear and nonlinear models,” in MICAI 2005: Advances in Artificial Intelligence, A. Gelbukh, ´ A. de Albornoz, and H. TerashimaMar´ ın, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2005, pp. 504–513.
[4] C. Cheng, A. Sa-Ngasoongsong, O. Beyca, T. Le, H. Yang, Z. Kong, and S. Bukkapatnam, “Time series forecasting for nonlinear and nonstationary processes: a review and comparative study,” IIE Transactions, vol. 47, no. 10, pp. 1053–1071, 2015. [5] F. Canova, “Vector autoregressive models: specification, estimation, inference, and forecasting,” Handbook of applied econometrics volume 1: Macroeconomics, pp. 53–110, 1999. [6] J. Durbin, “Efficient estimation of parameters in moving-average models,” Biometrika, vol. 46, no. 3/4, pp. 306–316, 1959. [7] S. H. Holan, R. Lund, and G. Davis, “The arma alphabet soup: A tour of arma model variants,” 2010. [8] G. E. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time series analysis: forecasting and control. John Wiley & Sons, 2015. [9] A. K. Dubey, A. Kumar, V. Garc´ ıa-D´ ıaz, A. K. Sharma, and K. Kanhaiya, “Study and analysis of sarima and lstm in forecasting time series data,” Sustainable Energy Technologies and Assessments, vol. 47, p. 101474, 2021. [10] K. e. a. Benidis, “Deep learning for time series forecasting: Tutorial and literature survey,” ACM Comput. Surv., vol. 55, no. 6, dec 2022. [Online]. Available: https://doi.org/10.1145/3533382 [11] J. F. Torres, D. Hadjout, A. Sebaa, F. Mart´ ınez- ´ Alvarez, and A. Troncoso, “Deep learning for time series forecasting: A survey,” Big Data, vol. 9, no. 1, pp. 3–21, 2021. [12] A. Mahmoud and A. Mohammed, A Survey on Deep Learning for TimeSeries Forecasting. Springer International Publishing, 2021, pp. 365– 392. [13] S. Du, T. Li, Y. Yang, and S.-J. Horng, “Multivariate time series forecasting via attention-based encoder–decoder framework,” Neurocomputing, vol. 388, pp. 269–279, 2020. [14] A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, 2017. [15] W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023. [16] J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler et al., “Emergent abilities of large language models,” arXiv preprint arXiv:2206.07682, 2022. [17] T. Zhang, X. Yue, Y. Li, and H. Sun, “Tablellama: Towards open large generalist models for tables,” in NAACL-HLT. Association for Computational Linguistics, 2024, pp. 6024–6044. [18] L. Zha, J. Zhou, L. Li, R. Wang, Q. Huang, S. Yang, J. Yuan, C. Su, X. Li, A. Su, T. Zhang, C. Zhou, K. Shou, M. Wang, W. Zhu, G. Lu, C. Ye, Y. Ye, W. Ye, Y. Zhang, X. Deng, J. Xu, H. Wang, G. Chen, and J. Zhao, “Tablegpt: Towards unifying tables, nature language and commands into one GPT,” CoRR, vol. abs/2307.08674, 2023. [19] A. Su, A. Wang, C. Ye, C. Zhou, G. Zhang, G. Chen, G. Zhu, H. Wang, H. Xu, H. Chen, H. Li, H. Lan, J. Tian, J. Yuan, J. Zhao, J. Zhou, K. Shou, L. Zha, L. Long, L. Li, P. Wu, Q. Zhang, Q. Huang, S. Yang, T. Zhang, W. Ye, W. Zhu, X. Hu, X. Gu, X. Sun, X. Li, Y. Yang, and Z. Xiao, “Tablegpt2: A large multimodal model with tabular data integration,” CoRR, vol. abs/2411.02059, 2024. [20] P. Li, Y. He, D. Yashar, W. Cui, S. Ge, H. Zhang, D. R. Fainman, D. Zhang, and S. Chaudhuri, “Table-gpt: Table fine-tuned GPT for diverse table tasks,” Proc. ACM Manag. Data, vol. 2, no. 3, p. 176, 2024. [21] N. Gruver, M. Finzi, S. Qiu, and A. G. Wilson, “Large language models are zero-shot time series forecasters,” in NeurIPS, 2023. [22] H. Tang, C. Zhang, M. Jin, Q. Yu, Z. Wang, X. Jin, Y. Zhang, and M. Du, “Time series forecasting with llms: Understanding and enhancing model capabilities,” SIGKDD Explor., vol. 26, no. 2, pp. 109–118, 2024. [23] G. Chatzigeorgakidis, K. Lentzos, and D. Skoutas, “Multicast: Zeroshot multivariate time series forecasting using llms,” in 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW). IEEE, 2024, pp. 119–127. [24] M. Tan, M. A. Merrill, V. Gupta, T. Althoff, and T. Hartvigsen, “Are language models actually useful for time series forecasting?” in NeurIPS, 2024. [25] P. Schmiedmayer, A. Rao, P. Zagar, V. Ravi, A. Zahedivash, A. Fereydooni, and O. Aalami, “Llm on fhir–demystifying health records,” arXiv preprint arXiv:2402.01711, 2024. [26] A. J. Nashwan, A. A. AbuJaber, and A. AbuJaber, “Harnessing the power of large language models (llms) for electronic health records (ehrs) optimization,” Cureus, vol. 15, no. 7, 2023. [27] M. Wornow, A. Lozano, D. Dash, J. Jindal, K. W. Mahaffey, and N. H. Shah, “Zero-shot clinical trial patient matching with llms,” arXiv preprint arXiv:2402.05125, 2024. [28] Y. Li, S. Wang, H. Ding, and H. Chen, “Large language models in finance: A survey,” in Proceedings of the Fourth ACM International Conference on AI in Finance, 2023, pp. 374–382. [29] H. Yang, X.-Y. Liu, and C. D. Wang, “Fingpt: Open-source financial large language models,” arXiv preprint arXiv:2306.06031, 2023. [30] H. Zhao, Z. Liu, Z. Wu, Y. Li, T. Yang, P. Shu, S. Xu, H. Dai, L. Zhao, G. Mai et al., “Revolutionizing finance with llms: An overview of applications and insights,” arXiv preprint arXiv:2401.11641, 2024. [31] M. Hosseini, C. A. Gao, D. M. Liebovitz, A. M. Carvalho, F. S. Ahmad, Y. Luo, N. MacDonald, K. L. Holmes, and A. Kho, “An exploratory survey about using chatgpt in education, healthcare, and research,” medRxiv, pp. 2023–03, 2023. [32] S. Moore, R. Tong, A. Singh, Z. Liu, X. Hu, Y. Lu, J. Liang, C. Cao, H. Khosravi, P. Denny et al., “Empowering education with llms-the next-gen interface and content generation,” in International Conference on Artificial Intelligence in Education. Springer, 2023, pp. 32–37. [33] Y. Jiang, Z. Pan, X. Zhang, S. Garg, A. Schneider, Y. Nevmyvaka, and D. Song, “Empowering time series analysis with large language models: A survey,” 2024. [34] X. Zhang, R. R. Chowdhury, R. K. Gupta, and J. Shang, “Large language models for time series: A survey,” 2024. [35] M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P.-Y. Chen, Y. Liang, Y.-F. Li, S. Pan, and Q. Wen, “Time-LLM: Time series forecasting by reprogramming large language models,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=Unb5CVPtae [36] A. F. Ansari, L. Stella, C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. Rangapuram, S. P. Arango, S. Kapoor et al., “Chronos: Learning the language of time series,” 2024. [37] J. Shieh and E. J. Keogh, “iSAX: indexing and mining terabyte sized time series,” in SIGKDD, 2008, pp. 623–631. [38] G. Woo, C. Liu, A. Kumar, C. Xiong, S. Savarese, and D. Sahoo, “Unified training of universal time series forecasting transformers,” in Forty-first International Conference on Machine Learning, 2024. [Online]. Available: https://openreview.net/forum?id=Yd8eHMY1wz [39] M. Kayali, A. Lykov, I. Fountalis, N. Vasiloglou, D. Olteanu, and D. Suciu, “CHORUS: foundation models for unified data discovery and exploration,” Proc. VLDB Endow., vol. 17, no. 8, pp. 2104–2114, 2024. [40] T. Xie, C. H. Wu, P. Shi, R. Zhong, T. Scholak, M. Yasunaga, C. Wu, M. Zhong, P. Yin, S. I. Wang, V. Zhong, B. Wang, C. Li, C. Boyle, A. Ni, Z. Yao, D. Radev, C. Xiong, L. Kong, R. Zhang, N. A. Smith, L. Zettlemoyer, and T. Yu, “Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models,” in EMNLP. Association for Computational Linguistics, 2022, pp. 602– 631.