A Benchmark for Early Time-Series Classification (Extended Abstract)
Abstract
A framework that allows the comparison of ETSC algorithms, and a new method that is based on the selective truncation of time-series principle are developed, which includes a bundle of datasets originating from real-world time-critical applications.
Full text
A Benchmark for Early Time-Series Classification1 Petro-Foti Kamberi !2 Institute of Informatics & Telecommunications, NCSR ‘Demokritos’, Greece3 Evgenios Kladis !4 Institute of Informatics & Telecommunications, NCSR ‘Demokritos’, Greece5 Charilaos Akasiadis !6 Institute of Informatics & Telecommunications, NCSR ‘Demokritos’, Greece7 Abstract 8 The objective of Early Time-Series Classification (ETSC) is to predict the class of incoming time9 series by observing the fewest time-points possible. Although many approaches have been proposed 10 in the past, not all techniques are suitable for every problem type. In particular, the characteristics of 11 the input data may impact performance. To aid researchers and developers with deciding which kind 12 of method suits their needs best, we developed a framework that allows the comparison of five existing 13 ETSC algorithms, and also introduce a new method that is based on the selective truncation of 14 time-series principle. To promote results reproducibility and the alignment of algorithm comparisons, 15 we also include a bundle of datasets originating from real-world time-critical applications, and for 16 which the application of ETSC algorithms can be considered quite valuable.17 2012 ACM Subject Classification Computing methodologies →Machine learning algorithms18 Keywords and phrases Time-series analysis, Classification, Benchmark19 Digital Object Identifier 10.4230/LIPIcs.TIME.2023.1820 Category Extended Abstract21 Related Version Previous Version:https://arxiv.org/abs/2203.0162822 Supplementary Material Software:https://github.com/xarakas/ETSC23 Funding This work has received funding from the EU project CREXDATA (under grant agreement 24 No. 101092749).25 1 Extended Abstract26 Latest technological advancements drive the generation of large volumes of time-series 27 data. Sea vessels, for example, utilize integrated sensory and telecommunication devices to 28 continuously report trajectory information in the form of time-series data. This abundance 29 of data is leveraged by machine learning techniques to address various problems. In the life 30 sciences domain, simulators are incorporated to test the effectiveness of new experimental 31 drugs “in-silico”. Such simulations often require long time and large amounts of computational 32 resources, which, in the case of unsuccessful drug treatment cases being simulated, are 33 consumed in vain [ 1 ]. It would be desirable to be able to predict such outcomes early on by 34 observing the simulations course as time-series, and to terminate not interesting trials so to 35 speed up the whole drug discovery process. To this end, the domain of Early Time-Series 36 Classification (ETSC) has an objective to classify time-series at the earliest point possible, 37 before the entire series is observed [5].38 Meanwhile, despite the numerous proposed methods for ETSC, there is a notable absence 39 of a dedicated experimental evaluation and comparison framework in this field. Furthermore, 40 ETSC methods are predominantly evaluated and compared against only a limited set of 41 alternative algorithms. In order to fill this gap, we have developed a framework that allows 42 ©Petro-Foti Kamberi, Evgenios Kladis, Charilaos Akasiadis; licensed under Creative Commons License CC-BY 4.0 30th International Symposium on Temporal Representation and Reasoning (TIME 2023). Editors: Alexander Artikis, Florian Bruse, and Luke Hunsberger; Article No. 18; pp. 18:1–18:3 Leibniz International Proceedings in Informatics Schloss Dagstuhl – Leibniz-Zentrum für Informatik, Dagstuhl Publishing, Germany
18:2 A Benchmark for Early Time-Series Classification to empirically compare five existing approaches as well as a newly introduced one, using 43 a curated bundle of datasets from real-world applications. This framework is utilized to 44 highlight the ETSC algorithms merits and shortcomings when applied to cases with different 45 features, e.g. dataset size, observation variability, class imbalance, etc. The framework can 46 be easily extended to include more datasets and algorithms, and is openly available online [ 2 ]. 47 Evaluation metrics : In the ETSC domain, apart from the predictive performance 48 (accuracy and F 1 -score), there is also the objective to optimize the earliness of the generated 49 predictions. These two metrics can also be combined into a single metric as a harmonic 50 mean [ 9 ]. Training and testing times are also of interest to consider in real-world applications. 51 Algorithms : Our framework includes five existing ETSC algorithms, i.e. ECEC [ 7 ], 52 ECONOMY-K [ 3 ], ECTS [ 10 ], EDSC [ 11 ], and TEASER [ 9 ]. ECEC calculates confidence 53 thresholds above which a class label prediction is considered to be reliable. ECONOMY-K 54 performs clustering on the training data, and estimates the cost of having to observe more 55 time points to generate a prediction. ECTS utilizes nearest neighbors and reverse nearest 56 neighbors sets for its decisions. EDSC extracts shapelets of the training data that are then57 matched with the input test data. TEASER trains classifiers on overlapping prefixes of the 58 training data, and then applies an one-class SVM that validates the class label prediction.59 In addition, we propose a new method that can be configured to utilize different state-of60 the-art full time-series classification algorithms for ETSC. It relies on iteratively truncating 61 time-series into prefixes of gradually increasing length and then applying Minirocket [ 4 ], 62 MLSTM [ 6 ], and WEASEL [ 8 ] to the truncated examples for predicting the corresponding 63 class labels. Thus, for each dataset, a fixed earliness is determined throughout the training 64 phase. Although this might be suboptimal in the sense that for particular instances the 65 prediction could have been generated earlier than for others, it constitutes a comprehensive 66 baseline for diagnosing if applying ETSC to particular datasets and domains would be 67 successful, i.e. if accurate predictions can be generated earlier, before the whole time-series 68 is observed. Our approach can incorporate any full time-series classification algorithm.69 Datasets : We have collected 10 publicly available datasets from the UEA & UCR 70 Time-series Classification Repository, and also introduced two new cases, one originating 71 from the drug discovery domain, and the other from the field of maritime intelligence. In 72 turn, these datasets are categorized according to particular characteristics: size, coefficient of 73 variability, levels of class imbalance, number of distinct class labels, and number of variables. 74 Comparison results : Judging by our experimental results, we can see that ECEC 75 achieves the most accurate and early predictions for datasets with lengthy time-series, but 76 it requires higher training times. When the number of examples in the dataset increases, 77 the MLSTM variant of our method competes with ECEC and TEASER in terms of the 78 harmonic mean between accuracy and earliness, but it has higher training times compared 79 to both. In applications with high variance in measurements and high class imbalance, 80 ECEC and MLSTM achieve the highest harmonic mean scores. In multi-class classification 81 cases, MLSTM is the best choice with the lowest earliness scores, followed by Minirocket, 82 which has high accuracy and reduced training times. For the rest of the datasets we tested, 83 Minirocket is the most suitable algorithm for ETSC in terms of harmonic mean. It has very 84 low earliness scores and training times, although its predictive accuracy is worse than ECEC, 85 ECONOMY-K, and ECTS, which however achieve higher earliness scores.86 As concerns future work, we are already in the process of further enriching our framework 87 with additional ETSC algorithms and more datasets. We believe that establishing consolidated 88 benchmarking procedures will be of great benefit for the ETSC community.89
P.-F. Kamberi et al. 18:3 References 90 1 C. Akasiadis, M. Ponce-de Leon, A. Montagud, E. Michelioudakis, A. Atsidakou, E. Alevizos, 91 A. Artikis, A. Valencia, and G. Paliouras. Parallel model exploration for tumor treatment 92 simulations. Computational Intelligence, 38(4):1379–1401, 2022. doi:https://doi.org/10.93 1111/coin.12515.94 2 Charilaos Akasiadis, Evgenios Kladis, Evangelos Michelioudakis, Elias Alevizos, and Alexander 95 Artikis. Early time-series classification algorithms: An empirical comparison. arXiv preprint96 arXiv:2203.01628, 2022.97 3 A. Dachraoui, A. Bondu, and A. Cornuéjols. Early classification of time series as a non 98 myopic sequential decision making problem. In Joint European Conf. on Machine Learning 99 and Knowledge Discovery in Databases, pages 433–447. Springer, 2015.100 4 A. Dempster, D. F. Schmidt, and G. I. Webb. Minirocket: A very fast (almost) deterministic 101 transform for time series classification. In Proc. of the 27th ACM SIGKDD Conf. on Knowledge 102 Discovery & Data Mining, pages 248–257, 2021.103 5 A. Gupta, H. P. Gupta, B. Biswas, and T. Dutta. Approaches and applications of early 104 classification of time series: A review. IEEE Trans. Artif. Intell., 1(1):47–61, 2020.105 6 F. Karim, S. Majumdar, H. Darabi, and S. Harford. Multivariate LSTM-FCNs for time series 106 classification. Neural Networks, 116:237–245, 2019.107 7 J. Lv, X. Hu, L. Li, and P.-P. Li. An effective confidence-based early classification of time 108 series. IEEE Access, 7:96113–96124, 2019.109 8 P. Schäfer and U. Leser. Fast and accurate time series classification with WEASEL. In Proc. 110 of the 2017 ACM Conf. on Information and Knowledge Management, pages 637–646, 2017.111 9 P. Schäfer and U. Leser. Teaser: early and accurate time series classification. Data mining 112 and knowledge discovery, 34(5):1336–1362, 2020.113 10 Z. Xing, J. Pei, and P. S. Yu. Early classification on time series. Knowledge and information 114 systems, 31(1):105–127, 2012.115 11 Z. Xing, J. Pei, P. S. Yu, and K. Wang. Extracting interpretable features for early classification 116 on time series. In Proceedings of the 2011 SIAM international conference on data mining, 117 pages 247–258. SIAM, 2011.118 TIME 2023