scieee AI-readable full text Open interactive document viewer

Multitask deep learning for native language identification

Habic, Vuk,Semenov, Alexander,Pasiliao, Eduardo L.

Full text

This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY-NC-ND 4.0 https://creativecommons.org/licenses/by-nc-nd/4.0/ Multitask deep learning for native language identification © 2020 Published by Elsevier B.V. Accepted version (Final draft) Habic, Vuk; Semenov, Alexander; Pasiliao, Eduardo L. Habic, V., Semenov, A., & Pasiliao, E. L. (2020). Multitask deep learning for native language identification. Knowledge-Based Systems, 209, Article 106440. https://doi.org/10.1016/j.knosys.2020.106440 2020 Journal Pre-proof Multitask deep learning for native language identification Vuk Habic, Alexander Semenov, Eduardo L. Pasiliao PII: S0950-7051(20)30569-4 DOI: https://doi.org/10.1016/j.knosys.2020.106440 Reference: KNOSYS 106440 To appear in: Knowledge-Based Systems Received date : 26 March 2020 Revised date : 7 September 2020 Accepted date : 8 September 2020 Please cite this article as: V. Habic, A. Semenov and E.L. Pasiliao, Multitask deep learning for native language identification, Knowledge-Based Systems (2020), doi: https://doi.org/10.1016/j.knosys.2020.106440. This is a PDF file of an article that has undergone enhancements after acceptance, such as the addition of a cover page and metadata, and formatting for readability, but it is not yet the definitive version of record. This version will undergo additional copyediting, typesetting and review before it is published in its final form, but we are providing this version to give early visibility of the article. Please note that, during the production process, errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain. ©2020 Published by Elsevier B.V. Multitask deep learning for native language identification Vuk Habic Department of Electrical Engineering, University of Florida. Email: [email protected] Alexander Semenov Faculty of Information Technology, University of Jyv¨askyl¨a . Email: [email protected] Eduardo L. Pasiliao Munitions Directorate, Air Force Research Laboratory. Email: [email protected] Abstract Identifying the native language of a person by their text written in English (L1 identification) plays an important role in such tasks as authorship profiling and identification. With the current proliferation of misinformation in social media, these methods are especially topical. Most studies in this field have focused on the development of supervised classification algorithms, that are trained on a single L1 dataset. Although multiple labeled datasets are available for L1 identification, they contain texts authored by speakers of different languages and do not completely overlap. Current approaches achieve high accuracy on available datasets, but this is attained by training an individual classifier for each dataset. Studies show that joint training of multiple classifiers on different datasets can result in sharing information between the classifiers, leading to an increase in the accuracy of both tasks. In this study, we develop a novel deep neural network (DNN) architecture for L1 classification; it is based on an adversarial multitask learning method that integrates shared knowledge from multiple L1 datasets. We propose several variants of the architecture and rigorously evaluate their ∗Corresponding author Preprint submitted to Elsevier September 7, 2020 Revised Manuscript (Clean Version) Journal Pre-proof Journal Pre-proof performance on multiple datasets. Our results indicate the proposed multitask architecture is more efficient in terms of classification accuracy than previously proposed methods. Keywords: multitask learning; text classification; natural language processing; deep learning 1. Introduction With the rapid development of the World Wide Web and social media during the early 21st century, the amount of textual data has grown exponentially. With the growing availability of data and computational resources, the past 20 years have seen increasingly rapid advances in the field of computational5 linguistics and natural language processing; typical tasks in these fields include text classification, document summarization, text understanding, named entity recognition, and others. The majority of textual data on the Internet consist of text written in English; it is estimated that English is used by 54% of websites1, and it is the most common language used for communication on the Internet2,10 including such social media platforms as Twitter3. In contrast, only about 5% of people use English as their first language. Indeed, there are approximately 6,500 languages in the world, each of them being unique in their own way, but only ten languages of these are adopted as a mother tongues by approximately 3.5 billion of people (about 50% of the world population), as estimated by Ethnologue4.15 The top three languages by number of speakers are Mandarin Chinese with 1.3 billion speakers, Spanish with 460 million speakers, and English with 379 million speakers. Many people speak more than one language; the most common second language is English, with about 753 million speakers. Further, we refer to the first language as L1 and the second language as L2.20 Recently, there has been increased interest in the automatic determination 1https://w3techs.com/technologies/overview/content_language/all 2https://www.statista.com/statistics/262946/share-of-the-most-common-languages-on-the-internet/ 3https://www.technologyreview.com/s/522376/the-many-tongues-of-twitter/ 4https://www.ethnologue.com/statistics/size 2 Journal Pre-proof Journal Pre-proof of authors’ L1 based on their text written in an L2. It is especially important to determine the L1 based on text written in English. In part, this is motivated by the proliferation of misinformation on social media, information that is often alleged to be authored by individuals whose L1 is not English. Knowing the25 L1 can also facilitate such tasks as authorship profiling, authorship attribution, and authorship identification. Indeed, writing in English is sometimes difficult for non-native speakers. Even after they study the grammar and mechanics, it does not have the same natural flow as their mother tongue. There are errors or certain ways that non-native speakers write in English that are often consistent30 among speakers of the same native tongue. Below, we provide examples of English texts written by native Arabic speakers, extracted from the Test of English as a Foreign Language 11 (TOEFL11) dataset. Arabic: “I agree that most advertisements make products seems to be much more better than they really are concerning the material used and satisfaction35 to consumers. Useing a below average materials in some products and advertise it the best” Chinese: “Now advertisements have come to pervade every aspect of our lifes and, as a result, we can see advertisements every where, where in the school or in the store, where on the TV or on the newspaper. Some people think that most40 advertisements make products seem much better that they really are”. A considerable amount of literature has been published on L1 identification. These studies usually approach the problem using supervised classification algorithms; for example, a report on a native language identification shared task [1] described the results of a competition between 19 teams. The goal of the45 competition was to develop a machine learning method that would be capable of classifying L1 with the highest accuracy for the given dataset. The teams used the TOEFL11 dataset containing essays authored by speakers of 11 languages who were taking the TOEFL exam. A predefined set of 11,000 entities were used for training, and 1,100 entities were used for evaluation; the test entities50 were chosen by the competition organizers. The top team achieved about 88% accuracy, with a baseline accuracy of 71%. The baseline was achieved using 3 Journal Pre-proof Journal Pre-proof support vector machines (SVMs) trained on word unigrams. Data from several other studies also suggest that L1 classification can be performed with rather high accuracy.55 Aside from L1 classification, there has been a growing recent trend toward the application of deep machine learning for various natural language processing (NLP) tasks; it is estimated [2] that about 70% of papers in top-level NLP conferences employ deep learning. Deep learning has been successfully applied for text classification, machine translation, text understanding, part-of-speech60 (PoS) tagging, and other tasks. Indeed, many teams from the competition reported in [1] used deep learning for L1 prediction. It has also been shown that very deep neural networks (DNNs) often outperform other architectures in text classification tasks [3], such as sentiment analysis or text topic prediction. Current approaches used for L1 identification [1] demonstrate high accuracy;65 however, all existing studies develop these methods for a specific dataset at a time, either TOEFL11, International Corpus of Learner English (ICLE), or others; treating the problem as single-task learning. The existing body of research on multitask learning [4] suggests that learning efficiency can increase if multiple learners of different but related tasks would70 learn simultaneously, provided that they share the learned information between themselves. This way, the learners would exploit the commonalities between the tasks, allowing them to mutually increase their performance. In practice, different learning tasks are represented as different but related datasets, and the goal of learning is to improve performance on all the datasets at the same time.75 Multitask learning architectures may be implemented as a special class of neural networks with shared parameters, at the same time having separate inputs and outputs for different datasets. Multitask learning has been previously applied to NLP; for example, adversarial multitask learning has been successfully applied for the segmentation of Chinese texts based on multiple heterogeneous datasets80 [5]. However, to the best of our knowledge, multitask learning has not yet been applied to the L1 identification problem. The purpose of this paper is to develop novel L1 multitask learning ar4 Journal Pre-proof Journal Pre-proof chitectures, based on DNNs and to perform a rigorous evaluation of them on multiple datasets. Our results show the effectiveness of multitask learning, with85 increased the classification accuracy compared to previously proposed models and the baselines. To summarize, this paper makes the following contributions: •We develop a multitask learning method that improves L1 classification accuracy on both TOEFL11 and ICLE datasets.90 •We conduct extensive experiments and demonstrate the performance of our method. •We propose multiple architectures with shared layers and experiment with adversarial strategies to force shared layers to be dataset-invariant. The remainder of the paper is structured as follows. Section 2 discusses re-95 lated literature. Section 3 discusses the objectives and introduces the problem. Section 4 describes our suggested architectures, and Section 5 describes and analyzes the experimental results. Section 6 concludes the paper. 2. Literature overview A number of studies have proposed different architectures of DNNs for text100 classification tasks [3]. Representative examples include convolutional neural networks (CNNs) for sentence classification [6], and character-level CNNs [7]. These architectures are conceptually similar to CNNs used in image classification, but they use one-dimensional convolutions. In addition deep CNNs have been proposed for the categorization of texts in other papers, such as [8], [9]; a105 recent article [10] proposes a graph CNN for text classification, and [7] describes architecture combining a CNN and a recurrent neural network. A related area of text classification is language recognition. In [11], the authors propose sparse representation of letters and words; their design shows accuracy of 95.4%. Language recognition aims at identifying the language of a text written using that110 language. It is fundamentally different from L1 classification, the goal of which 5 Journal Pre-proof Journal Pre-proof is to identify the native language of the person using text written in English as an input. In recent years, there has been an increasing amount of literature on L1 classification. All of it deals with classification using a single dataset, and as stated, none of it deals with multitask learning. A related area, the application115 of transfer learning for NLP, is described in a survey [12]. There is a general consensus on what elements to use as features for text processing, and they are character N-grams, word N-grams, and PoS N-grams. In one paper [13], the authors evaluated several architectures while considering several different versions of the N-gram features. They achieved accuracy of 0.83120 using an ensemble residual neural network with word uni-grams and character 5-grams and 6-grams as input features. This was the most accurate model in the competition. Authors of [13] also outlined the different options for features and architectures that can be utilized. Bjerva et al. [13] discussed PoS tagged sentences, continuous bag of words features, and spelling features. One important125 note to take from [13] is that the performance of any model is lower when using external sources, such as those features mentioned above. Another takeaway is that DNNs are able to perform this classification, although traditional methods such as SVMs still appear to be better. L1 classification has also been done in [14] using SVM on N-grams, achieving 83% accuracy. In fact, most works130 in this field sport relatively simple model architectures and basic character and word N-grams. In another paper [15], the authors illustrate the different features used for native language identification, including character, word, and PoS N-grams. They achieved a maximum accuracy of 81.17% .135 In [16], the authors reference several techniques to achieve natural language identification. They look at unique character string identification, frequent word recognition, and bigraph/trigraph-based recognition. They found that the trigraph approach was the best. They define trigraphs the same way we define character 3-grams.140 [2] discusses natural language identification using spectrogramand cochleagrambased features for very short speech utterances. That paper differs from our 6 Journal Pre-proof Journal Pre-proof approach in that they use speech as opposed to text. Nevertheless, it was advantageous for us to look at their experiments. For example, they use SDC features that include the number of cepstral coefficients in each frame, time ad-145 vance and delay for delta computation, and time shift between between consecutive frames. They use a bidirectional long short-term memory neural network however, they only achieved about 75% accuracy. In [17], the authors create a model in order to classify the TOEFL11 dataset. They achieve about 83% classification accuracy. The features they used include150 words, characters, and external features such as PoS tags. The word category includes lexemes, lemmas, and PoS tags. Character features included character N-grams from 1–9 grams. Their trial method contains several models, including SVM and logistic regression. Information from the sources cited above suggests that character N-grams155 and words are the most important features and that shallow neural network architectures and methods such as SVM perform well for L1 identification. None of the sources go further than the standard text prepossessing, and external sources seem to be ineffective. 3. Objectives160 In this study, we address the following research questions: •What is the architecture of a DNN, leveraging multitask learning for a joint training machine learning model on multiple L1 datasets? •Is multitask learning beneficial for L1 identification in terms of classification accuracy?165 •How does the resulting classifier compare to the state-of-the art methods? 4. Approach Consider a dataset Dt={(xt i, yt i)}nt i=1, where xt i∈ X is the ith observed variable, yt i∈ Y is the corresponding ith label, and tdenotes a learning task, 7 Journal Pre-proof Journal Pre-proof a ReLU activation function as hyperparameters. Further, each softmax output had a dropout layer with p = 0.5 before it. We trained the models for 40 epochs285 and used adam as the optimization algorithm. The code was executed on a machine with four Tesla V100-SXM2-16GB GPUs and 32 CPU cores. We used k-fold validation with k = 5, that is, a traditional 80/20 split. Different from our setup, competition results listed in [1] were evaluated using about 9% of the TOEFL11 data; the rest was used for training.290 5.3. Experimental results and discussion # Name inputs No adversarial loss 1 Architecture 1 with Dense words, 4-grams 2 Architecture 1 with Dense words, 3-grams, 4-grams 3 Architecture 2 with Dense words, 4-grams 4 Architecture 2 with Dense words, 3-grams, 4-grams 5 Architecture 3 with Dense words, 4-grams 6 Architecture 3 with Dense words, 3-grams, 4-grams 7 Architecture 1 w/o Dense words, 4-grams 8 Architecture 1 w/o Dense words, 3-grams, 4-grams 9 Architecture 2 w/o Dense words, 4-grams 10 Architecture 2 w/o Dense words, 3-grams, 4-grams 11 Architecture 3 w/o Dense words, 4-grams 12 Architecture 3 w/o Dense words, 3-grams, 4-grams With adversarial loss 13 Architecture 2 with Dense words, 4-grams 14 Architecture 3 with Dense words, 4-grams 15 Architecture 2 with Dense words, 3-grams, 4-grams 16 Architecture 3 with Dense words, 3-grams, 4-grams Baselines 17 Shallow ANN (w/o Dense) words, 4-grams 18 Shallow ANN (with Dense) words, 4-grams 19 Shallow ANN(w/o Dense) words, 3-grams, 4-grams 20 Shallow ANN (with Dense) words, 3-grams,4-grams 21 Shallow ANN (with Convolutional layer) words, 3-grams,4-grams 22 Random Forest 3-grams, 4-grams 23 Random Forest words 24 SVM 3-grams, 4-grams Table 2: Description of the implemented neural network architectures, and baselines; and input data types 14 Journal Pre-proof Journal Pre-proof #Accuracy,TOEFL F1, TOEFL Accuracy, ICLE F1, ICLE No adversarial loss 1 0.753 ±0.00015 0.752 ±0.00017 0.751 ±0.00015 0.702 ±0.00037 2 0.753 ±0.00015 0.753 ±0.00014 0.746 ±0.00011 0.700 ±0.00019 3 0.806 ±0.00016 0.805 ±0.00019 0.846 ±0.00038 0.815 ±0.00047 4 0.812 ±0.00018 0.810 ±0.00022 0.834 ±0.00015 0.803 ±0.00022 5 0.850 ±0.00006 0.849 ±0.00006 0.960 ±0.00005 0.952 ±0.00008 6 0.803 ±0.00017 0.803 ±0.00016 0.819 ±0.00093 0.783 ±0.00107 7 0.851 ±0.00003 0.851 ±0.00004 0.898 ±0.00011 0.875 ±0.00019 80.851 ±0.00002 0.851 ±0.00002 0.905 ±0.00007 0.885 ±0.00015 9 0.839 ±0.00007 0.839 ±0.00008 0.973 ±0.00001 0.968 ±0.00002 10 0.847 ±0.00003 0.847 ±0.00003 0.958 ±0.00001 0.949 ±0.00002 11 0.848 ±0.00003 0.848 ±0.00004 0.956 ±0.00005 0.946 ±0.00008 12 0.847 ±0.00006 0.846 ±0.00006 0.961 ±0.00004 0.954 ±0.00008 With adversarial loss 13 0.830 ±0.00003 0.830 ±0.00004 0.864 ±0.00035 0.837 ±0.00056 14 0.827 ±0.00010 0.827 ±0.00009 0.851 ±0.00013 0.820 ±0.00015 15 0.830 ±0.00006 0.830 ±0.00007 0.871 ±0.00013 0.841 ±0.00027 16 0.837 ±0.00006 0.836 ±0.00009 0.878 ±0.00032 0.853 ±0.00048 Baselines 17 0.790 ±0.00004 0.790 ±0.00006 0.950 ±0.00001 0.938 ±0.00002 18 0.602 ±0.00166 0.605 ±0.00165 0.933 ±0.00163 0.915 ±0.00340 19 0.789 ±0.00001 0.789 ±0.00002 0.962 ±0.00003 0.953 ±0.00003 20 0.631 ±0.00052 0.634 ±0.00043 0.954 ±0.00003 0.943 ±0.00005 21 0.657 ±0.00008 0.656 ±0.00010 0.701 ±0.00086 0.654 ±0.00071 22 0.500 ±0.00020 0.493 ±0.00021 0.668 ±0.00006 0.599 ±0.00001 23 0.570 ±0.00038 0.566 ±0.00034 0.716 ±0.00014 0.659 ±0.00024 24 0.339 ±0.00011 0.336 ±0.00010 0.300 ±0.00025 0.147 ±0.00006 Table 3: Experimental results: accuracy and F1 score ±standard deviation for the TOEFL11 and ICLE datasets (k-fold validation with k= 5) for proposed models and baselines, listed in table 2. Maximal values of corresponding metrics are highlighted in bold. The experimental results are presented in Table 3. The table lists the accuracy for TOEFL11 averaged over 5 folds, the F1 score for TOEFL11, after macro averaging over 5 folds, and corresponding values for the ICLE dataset with standard deviations. We can make the following observations:295 1. The best results in terms of accuracy vis-`a-vis the TOEFL11 dataset are achieved by Architecture 1 without a dense layer with words, 3-grams, and 4-grams as inputs (model number 8); its accuracy is 0.851. The best performing model for ICLE is Architecture 2, also without the final dense 15 Journal Pre-proof Journal Pre-proof layer, with words and 3-grams as inputs (model number 9). Its accuracy300 is 0.973. The accuracy of model 8 vis-`a-vis TOEFL11 dataset is 7.7% higher than the best-performing baseline (model number 17, 0.79 accuracy). The accuracy of model 9 for the ICLE dataset is higher than the best-performing ICLE baseline (model number 19) by 1.11%. Both multitask architectures significantly outperform the baselines, with the p−value305 of the t−test for comparing TOEFL11 model 8 accuracy with baseline 17 being to 1.58 ×10−7and the corresponding p−value of ICLE accuracy for model 9 and that of baseline 19 being 0.0047. Our experiment shows that neural network-based multitask architectures outperform single-task architectures.310 2. For the TOEFL11 dataset all, but two of the proposed multitask architectures outperform all baselines. Table 1 shows that model numbers 1 and 2 are outperformed by baselines 17 and 19. Models 1 and 2 implement architecture 1 with a dense layer at the end of the neural network. We can see from Table 3, architectures 1–3 without a dense layer demonstrate higher315 accuracy than the baselines for TOEFL11; however, the baseline demonstrates higher accuracy than the proposed models for the ICLE dataset using 3and 4gram inputs. Otherwise proposed models perform better. We can also see that the shallow neural network with convolutional layers (baseline 21) performs worse than the baselines without a dense layer320 (baselines 17 and 19). 3. Most of the models having a dense layer at the end perform worse than the same models without this layer. However, adversarial loss greatly increases the accuracy of these models, but unexpectedly, architectures with no dense layer are generally better.325 4. For the TOEFL11 dataset, Architecture 1 without a final dense layer is the best, followed by Architecture 3 without a final dense layer, followed by Architecture 2, without a final dense layer. For the ICLE dataset, the second best accuracy is shown by baseline 19, and the third best performing model is architecture 3 without a final dense layer. However,330 16 Journal Pre-proof Journal Pre-proof the difference between accuracy values for the latter two models is not statistically significant, based on a t-test. 5. In general, the number of parameters of one multitask architecture equals to combined number of parameters of both corresponding single-task TOEFL11 and ICLE baselines. This is also reflected in the time of training, which335 is longer for multitask models. We have observed that time of training of multitask models is in average twice longer. 6. Improvement in F1 score goes at par with the accuracy. To study the best-performing model further, we trained the model implementing Architecture 1 without the dense layer, accepting words and character340 4-grams as input for 80 epochs on the 90.9% of training data, and tested it on the remaining 9.1%. This roughly corresponds to the setup of the task of L1 classification competition described in [1]. The maximum accuracy demonstrated by the developed model for the TOEFL11 dataset is 88.75 %, while the accuracy of the top-performing model from [1] was 88.18 %, the model is345 described in [19]. Figure 4 depicts the dependence of accuracy on the training epoch for training and validation datasets. In summary, these results show that multitask learning is beneficial for L1 identification tasks and that it is capable of outperforming state-of-the-art methods. In order to analyze the results, we constructed confusion matrices from the350 validation dataset containing 9.1% of the data. Figure 5 shows the confusion matrix for the TOEFL11 dataset, Figure 6 shows confusion matrix for the ICLE dataset. Matrices are visualized as heatmaps. We can see that for TOEFL11 the most often confusing for the model is Japanese versus Korean, and Japanese vs Chinese. Italian speakers are sometimes confused with Spanish speakers.355 Although Korean and Japanese are not related [20], they belong to the Japonic language family. Therefore, similar mistakes in English could be typical for speakers of these languages; this may explain the observation that texts written by speakers of these languages are confused by the model. Italian and Spanish are both Romance languages with similarities in their vocabulary; this may360 17 Journal Pre-proof Journal Pre-proof Figure 4: Accuracy of the model implementing Architecture 2 (without a dense layer) with words and 4-grams as input, trained for 80 epochs. The top figure shows training accuracy, and the figure on the bottom shows validation accuracy. The blue line represents accuracy for the TOEFL11 dataset, and the orange one represents accuracy for the ICLE dataset. We can see that training accuracy tends to be 100% at around the 50th training epoch. explain the confusion of the model. As opposed to TOEFL11, there is no clear indication of specific confusion of the model for the ICLE dataset (see Figure 6). This may be explained by the fact that accuracy vis-`a-vis ICLE is much higher than vis-`a-vis TOEFL11, and even simple baselines achieve high accuracy (see Table 3) on this dataset. Interestingly, there is no confusion between related365 languages, such as Bulgarian, Russian, Polish, and Czech; also different from TOEFL11, there is no confusion between Chinese and Japanese. 18 Journal Pre-proof Journal Pre-proof ARA DEU FRA HIN ITA JPN KOR SPA TEL TUR ZHO ARADEUFRAHINITAJPNKORSPATELTURZHO 102 1351144010 076 020001010 0 3 87 02003022 31083 0000310 024093 002000 11100100 12 2 0 0 0 20000682 1000 213171080 020 1000011088 0 0 02011116087 3 010006302198 0 20 40 60 80 100 Figure 5: Confusion matrix of the model on TOEFL part of validation dataset (9.1% of TOEFL11). 6. Conclusions This paper presents novel multitask learning methods for L1 language classification. We proposed three neural network architectures that are capable of370 sharing the knowledge between multiple heterogeneous input datasets. We evaluate our architectures using TOEFL11 and ICLE datasets; based on the results, we conclude that our approach shows the benefit of multitask learning is more accurate than state-of-the-art methods. In future work, we would like to investigate further tuning of the models’375 hyper parameters to increase the classification accuracy. We would also like to experiment with additional input features, such as N-grams with different values of N. 19 Journal Pre-proof Journal Pre-proof BG CN CZ DB DN FI FR GE IT JP NO PO RU SP SW TR TS BGCNCZDBDNFIFRGEITJPNOPORUSPSWTRTS 470000000000110000 0165 000001000000010 005300100003000000 000310000001000000 000122000000000000 101005000000010100 000300570000000000 000000084 000000000 000000006900001000 000000000660000000 000002000058001001 000002020205910000 000001010000410000 000000000000049010 000003001000025800 000000000000000550 000000000010000099 0 30 60 90 120 150 Figure 6: Confusion matrix of the model on ICLE part of validation dataset (9.1% of ICLE) 7. Acknowledgements This work was funded in part by the US Air Force Research Laboratory380 (AFRL) European Office of Aerospace Research and Development (grant no. FA9550-17-1-0030), and the AFRL Mathematical Modeling and Optimization Institute. References [1] S. Malmasi, K. Evanini, A. Cahill, J. Tetreault, R. Pugh, C. Hamill,385 D. Napolitano, Y. Qian, A report on the 2017 native language identification shared task, in: Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, Association 20 Journal Pre-proof Journal Pre-proof for Computational Linguistics, Copenhagen, Denmark, 2017, pp. 62–75. doi:10.18653/v1/W17-5007.390 URL https://www.aclweb.org/anthology/W17-5007 [2] T. Young, D. Hazarika, S. Poria, E. Cambria, Recent trends in deep learning based natural language processing [review article], IEEE Computational Intelligence Magazine 13 (3) (2018) 55–75. doi:10.1109/MCI.2018.2840738. [3] A. Conneau, H. Schwenk, L. Barrault, Y. Lecun, Very deep convolutional395 networks for text classification, in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, Association for Computational Linguistics, Valencia, Spain, 2017, pp. 1107–1116. URL https://www.aclweb.org/anthology/E17-1104400 [4] R. Caruana, Multitask learning, Machine Learning 28 (1) (1997) 41–75. doi:10.1023/A:1007379606734. URL https://doi.org/10.1023/A:1007379606734 [5] X. Chen, Z. Shi, X. Qiu, X. Huang, Adversarial multi-criteria learning for Chinese word segmentation, in: Proceedings of the 55th Annual Meeting of405 the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Vancouver, Canada, 2017, pp. 1193–1203. doi:10.18653/v1/P17-1110. URL https://www.aclweb.org/anthology/P17-1110 [6] Y. Kim, Convolutional neural networks for sentence classification, in: Pro-410 ceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar, 2014, pp. 1746–1751. doi:10.3115/v1/D14-1181. URL https://www.aclweb.org/anthology/D14-1181 [7] X. Zhang, J. Zhao, Y. LeCun, Character-level convolutional networks for415 text classification, in: Proceedings of the 28th International Conference on 21 Journal Pre-proof Journal Pre-proof Neural Information Processing Systems - Volume 1, NIPS’15, MIT Press, Cambridge, MA, USA, 2015, p. 649–657. [8] C. dos Santos, M. Gatti, Deep convolutional neural networks for sentiment analysis of short texts, in: Proceedings of COLING 2014, the 25th Interna-420 tional Conference on Computational Linguistics: Technical Papers, Dublin City University and Association for Computational Linguistics, Dublin, Ireland, 2014, pp. 69–78. URL https://www.aclweb.org/anthology/C14-1008 [9] P. Wang, J. Xu, B. Xu, C. Liu, H. Zhang, F. Wang, H. Hao, Seman-425 tic clustering and convolutional neural network for short text categorization, in: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Association for Computational Linguistics, Beijing, China, 2015, pp. 352–357.430 doi:10.3115/v1/P15-2058. URL https://www.aclweb.org/anthology/P15-2058 [10] L. Yao, C. Mao, Y. Luo, Graph convolutional networks for text classification, CoRR abs/1809.05679. arXiv:1809.05679. URL http://arxiv.org/abs/1809.05679435 [11] M. Imani, J. Hwang, T. Rosing, A. Rahimi, J. M. Rabaey, Low-power sparse hyperdimensional encoder for language recognition, IEEE Design Test 34 (6) (2017) 94–101. [12] L. Qi, Literature survey: domain adaptation algorithms for natural language processing, Department of Computer Science The Graduate Center,440 The City University of New York, 2012. [13] J. Bjerva, G. Grigonyte, R. ¨ Ostling, B. Plank, Neural networks and spelling features for native language identification, in: EMNLP, Association for Computational Linguistics, 2017, pp. 235–239. 22 Journal Pre-proof Journal Pre-proof [14] A. Kulmizev, B. Blankers, J. Bjerva, M. Nissim, G. van Noord, B. Plank,445 M. Wieling, The power of character n-grams in native language identification, in: Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Copenhagen, Denmark, 2017, pp. 382–389. doi:10.18653/v1/ W17-5043.450 URL https://www.aclweb.org/anthology/W17-5043 [15] T. Mizumoto, Y. Hayashibe, K. Sakaguchi, M. Komachi, Y. Matsumoto, Naist at the nli 2013 shared task, in: Proceedings of the Eighth Workshop on Innovative Use of NLP for Building Educational Applications, 2013, pp. 134–139.455 [16] C. Souter et al., Natural language identification using corpus-based models, HERMES - Journal of Language and Communication in Business 7 (13) (2017) 183–203. doi:10.7146/hjlcb.v7i13.25083. URL https://tidsskrift.dk/her/article/view/25083 [17] S. Jarvis, Y. Bestgen, S. Pepper, Maximizing classification accuracy in460 native language identification, in: Proceedings of the Eighth Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Atlanta, Georgia, 2013, pp. 111–118. URL https://www.aclweb.org/anthology/W13-1714 [18] Y. Sari, M. Rifqi Fatchurrahman, M. Dwiastuti, A shallow neural net-465 work for native language identification with character n-grams, in: Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Copenhagen, Denmark, 2017, pp. 249–254. doi:10.18653/v1/W17-5027. URL https://www.aclweb.org/anthology/W17-5027470 [19] A. Cimino, F. Dell’Orletta, Stacked sentence-document classifier approach for improving native language identification, in: BEA@EMNLP, 2017. 23 Journal Pre-proof Journal Pre-proof