scieee AI-readable full text Open interactive document viewer

General compound hawkes processes in limit order books

Sviščuk, Anatolij,Huffman, Aiden

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Sviščuk, Anatolij; Huffman, Aiden Article General compound hawkes processes in limit order books Risks Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Sviščuk, Anatolij; Huffman, Aiden (2020) : General compound hawkes processes in limit order books, Risks, ISSN 2227-9091, MDPI, Basel, Vol. 8, Iss. 1, pp. 1-25, https://doi.org/10.3390/risks8010028 This Version is available at: https://hdl.handle.net/10419/257983 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ risks Article General Compound Hawkes Processes in Limit Order Books Anatoliy Swishchuk *,† and Aiden Huffman † Department of Mathematics and Statistics, Faculty of Science, Calgary, AL T2N1N4, Canada; [email protected] *Correspondence: [email protected]; Tel.: +1-(403)-220-3274 † These authors contributed equally to this work. Received: 29 January 2020; Accepted: 11 March 2020; Published: 14 March 2020   Abstract: In this paper, we study various new Hawkes processes. Specifically, we construct general compound Hawkes processes and investigate their properties in limit order books. With regard to these general compound Hawkes processes, we prove a Law of Large Numbers (LLN) and a Functional Central Limit Theorems (FCLT) for several specific variations. We apply several of these FCLTs to limit order books to study the link between price volatility and order flow, where the volatility in mid-price changes is expressed in terms of parameters describing the arrival rates and mid-price process. Keywords: Hawkes processes; general compound Hawkes processes; limit order books; functional central limit theorems; LOBSTER data 1. Introduction The Hawkes process (HP) is named after its creator, Hawkes (1971); Hawkes and Oakes (1974). The HP is a simple point process equipped with a self-exciting property, clustering effect and long run memory. Through its dependence on the history of the process, the HP captures the temporal and cross sectional dependence of the event arrival process as well as the ’self-exciting’ property observed in our empirical data on limit order books. Self-exciting point processes have recently been applied to high frequency data for price changes Bacry et al. (2011) or order arrival times Embrechts et al. (2011). HPs have seen their application in many areas, like genetics Carstensen (2010), occurrence of crime Mohler et al. (2011), bank defaults Azizpour et al. (2018) and earthquakes Ogata (1988). Point processes gained a significant amount of attention in statistics during the 1950s and 1960s. Cox (1955) introduced the notion of a doubly stochastic Poisson process (called the Cox process now) and Bartlett (1963) investigated statistical methods for point processes based on their power spectral densities. Lewis (1964) formulated a point process model (for computer power failure patterns) which was a step in the direction of the HP. A nice introduction to the theory of point processes can be found in Daley and Vere-Jones (2003). The first type of point process in the context of market microstructure is the autoregressive conditional duration (ACD) model introduced by Engle and Russell (1998). A recent application of HP is in financial analysis, in particular limit order books. In this paper, we study various new Hawkes processes, namely general compound Hawkes processes to model the price process in limit order books. We prove a Law of Larges Numbers (LLN) and a Functional Central Limit Theorem (FCLT) for specific cases of these processes. Several of these FCLTs are applied to limit order books where we use asymptotic methods to study the link between price volatility and order flow in our models. The volatility of the price changes is expressed in terms of parameters describing the arrival rates and price changes. We also present some numerical Risks 2020,8, 28; doi:10.3390/risks8010028 www.mdpi.com/journal/risks Risks 2020,8, 28 2 of 25 examples. The general compound Hawkes process was first introduced in Swishchuk (2017) to model the risk process in insurance and studied in detail there. In the paper Swishchuk et al. (2019), we obtained functional CLTs and LLNs for general compound Hawkes processes with dependent orders and regime-switching compound Hawkes processes. Bowsher (2007) was the first one who applied the HP to financial data modelling. Cartea et al. (2014) applied HP to model market order arrivals. Filimonov et al. (2014) and Filimonov and Sornette (2012) applied the HPs to estimate the percentage of price changes caused by endogenous self-generated activity rather than by the exogenous impact of news or novel information. Bauwens and Hautsch (2009) used a five-dimensional HP to estimate multivariate volatility between five stocks, based on price intensities. Hewlett (2006) used the instantaneous jump in the intensity caused by the occurrence of an event to qualify the market impact of that event, taking into account the cascading effect of secondary events causing further events. Hewlett (2006) also used the Hawkes model to derive optimal pricing strategies for market makers and optimal trading strategies for investors given that the rational market makers have the historic trading data. Large (2007) applied a Hawkes model for the purpose of investigating market impact, with a specific interest in order book resiliency. Specifically, he considered limit orders, market orders and cancellations on both the buy and sell side, and further categorizes these events based on their level of aggression, resulting in a ten-dimensional Hawkes process. Other econometric models based on marked point processes with stochastic intensity include autoregressive conditional intensity (ACI) models with the intensity depending on its history. Hasbrouck (1999) introduced a multivariate point process to model the different events of an order book but did not parametrize the intensity. We note that Brémaud and Massoulié (1996) generalized the HP to its nonlinear form. In addition, a functional central limit theorem for nonlinear Hawkes processes was obtained in Zheng et al. (2013). The ’Hawkes diffusion model’ introduced in Aït-Sahalia et al. (2015) attempted to extend previous models of stock prices to include financial contagion. Chavez-Demoulin and McGill (2012) used Hawkes processes to model high-frequency financial data. An application of affine point processes to portfolio credit risk may be found in Errais et al. (2010). Some applications of Hawkes processes to financial data are also given in Embrechts et al. (2011). Cohen and Elliott (2013) derived an explicit filter for Markov modulated Hawkes processes. Vinkovskaya (2014) considered a regime-switching Hawkes process to model its dependency on the bid-ask spread in limit order book. Regime-switching models for pricing of European and American options were considered in Buffington and Elliott (2002b) and Buffington and Elliott (2002a), respectively. Semi-Markov processes were applied to limit order books in Swishchuk and Vadori (2017) to model the mid-price. We also note that level-1 limit order books with time dependent arrival rates λ(t) were studied in Chávez-Casillas et al. (2019), including the asymptotic distribution of the price process. The paper by Bacry et al. (2015) proposes an overview of the recent academic literature devoted to the applications of Hawkes processes in finance. It is a nice survey of applications of Hawkes processes in finance. In general, the main models in high-frequency finance can be divided into univariate models, price models, impact models, order-book models and some systemic risk models, models accounting for news, high-dimensional models and clustering with graph models. The book by Cartea et al. (2015) developed models for algorithmic trading such as methods for executing large orders, market making, trading pairs of collections of assets, and executing in the dark pool. This book also contains a link from which several datasets can be downloaded, along with MATLAB code to assist in experimentation with the data. A detailed description of the mathematical theory of Hawkes processes is given in Liniger (2009). The paper by Laub et al. (2015) provides background, introduces the field and historical development, and touches upon all major aspects of Hawkes processes. The results of the current paper were first announced in Swishchuk and Huffman (2018). The main contribution and novelty of He and Swishchuk (2019) paper consists of considering different types of general compound Hawkes processes and their diffusive limits to model the mid-prices of six different stocks. Namely, EBAY, FB, Risks 2020,8, 28 3 of 25 MU, PCAR, SMH, and CSCO were used for quantitative and comparative analysis, and for defining the error rates to estimate the models fitting accuracy. The paper Swishchuk et al. (2017) deals with compound and regime-switching Hawkes processes to model the mid-price processes in limit order books. Diffusive limits were used for both models to study the link between the price volatility and the order flow. Numerical examples were presented using CISCO data dated from 3 November to 7 November 2014. Thus, the present paper contains not only new results with proofs for different general compound Hawkes processes, but also another data set where those results were applied compared with previous papers. The paper is organized as follows. A definition of a Hawkes process and a description of its properties are given in Section 2. Law of Large Numbers (LLN) and Functional Central Limit Theorems (FCLT) for various general compound Hawkes processes, including nonlinear, in limit order books are proved in Section 3. Descriptions of the data set and empirical results are presented in Section 4. Section 5gives a quantitative analysis of our results and Section 6concludes the paper. 2. Definition and Some Properties of Hawkes Processes (HP) 2.1. Definition of Hawkes Processes (HPs) In this section, we give various definitions and some properties of Hawkes processes which can be found in the existing literature (see, e.g., Hawkes (1971), Hawkes and Oakes (1974), Embrechts et al. (2011) and Zheng et al. (2013), to name a few). They include in particular one-dimensional and nonlinear Hawkes processes. Definition 1 (Counting Process) . A counting process is a stochastic process N(t) with t≥ 0, where N(t) takes positive integer values and satisfies N( 0 ) = 0. It is almost surely finite and a right-continuous step function with increments of size +1. Denote by FN(t) , t≥ 0, the history of the arrival up to time t; that is, FN(t) , t≥ 0, is a filtration (an increasing sequence of σ-algebras). A counting process N(t) can be interpreted as a cumulative count of the number of arrivals into a system up to the current time t . The counting process can also be characterized by the sequence of random arrival times (T1 , T2 , . . . ) at which the counting process N(t) has jumped. The process defined by these arrival times is called a point process (see Daley and Vere-Jones 2003). Definition 2 (Point Process) . If a sequence of random variables (T1 , T2 , . . . ) , taking values in [ 0, ∞) has P( 0 ≤T1≤T2≤. . . ) = 1, and the number of points in a bounded region is almost surely finite, then (T1,T2, . . . )is called a point process. Definition 3 (Conditional Intensity Function).Consider a counting process N(t)with associated histories FN(t), t ≥0. If a non-negative function λ(t)exists such that λ(t) = lim h→0 E[N(t+h)−N(t)| FN(t)] h(1) Then, it is called the conditional intensity function of N(t) (see Laub et al. 2015). We note that originally this function was called the hazard function (see Cox 1955). Definition 4 (One-dimensional Hawkes Process) . The one-dimensional Hawkes process (see Laub et al. 2015; Hawkes and Oakes 1974) is a point process N(t) which is characterized by its intensity λ(t) with respect to its natural filtration: λ(t) = λ+Zt 0µ(t−s)dN(s)(2) where λ>0, and the response function µ(t)is a positive function that satisfies R∞ 0µ(s)ds <1. Risks 2020,8, 28 4 of 25 The constant λ is called the background intensity and the function µ(t) is sometimes called the excitation function. To avoid the trivial case of a homogeneous Poisson process, we assume µ(t)6= 0. Thus, the Hawkes process is a non-Markovian extension of the Poisson process. With respect to the Definitions of λ(t)in 3and N(t)in 4, it follows that P(N(t+h)−N(t) = m| FN(t)) =        λ(t)h+o(h)m=1 o(h)m>1 1−λ(t)h+o(h)m=0 The interpretation of Equation (2) is that the events occur according to an intensity with a background intensity λ which increases by µ( 0 ) at each new event, eventually decaying back to the background intensity value according to the evolution of the function µ(t) . Choosing µ( 0 )> 0 leads to a jolt in the intensity at each new event, and this feature is often called the self-exciting feature. In other words, if an arrival causes the conditional intensity function λ(t) in Equations (1) and (2) to increase, then the process is called self-exciting. We would like to mention that the conditional intensity function λ(t) in Equations (1) and (2) can be associated with the compensator Λ(t)of the counting process N(t), that is, Λ(t) = Zt 0λ(s)ds (3) We note that Λ(t) is the unique non-decreasing, FN(t) , t≥ 0, predictable function, with Λ( 0 ) = 0 such that N(t) = M(t) + Λ(t)a.s., where M(t) is an FN(t) , t≥ 0, local martingale (existence of which is guaranteed by the Doob–Meyer decomposition). A common choice for the function µ(t) in Equation (2) is the one of exponential decay (see Hawkes 1971) µ(t) = αe−βt(4) with parameters α , β> 0. In this case, the Hawkes process is called the Hawkes process with exponentially decaying intensity. In the case of Equation (4), Equation (2) becomes λ(t) = λ+Zt 0αe−β(t−s)dN(s)(5) We note that, in the case of Equation (4) , the process (N(t) , λ(t)) is a continuous-time Markov process, which is not the case for a general choice of excitation function in Equation (1). With some inititial condition λ( 0 ) = λ0 , the conditional intensity λ(t) in Equation (5) with exponential decay in Equation (4) satisfies the SDE dλ(t) = β(λ−λ(t))dt +αdN(t),t≥0 (6) which can be solved using stochastic calculus as λ(t) = e−βt(λ0−λ) + λ+Zt 0αe−β(t−s)dN(s), (7) which is an extension of Equation (5). Risks 2020,8, 28 5 of 25 Another choice for µ(t)is a power law function λ(t) = λ+Zt 0 k (c+ (t−s))pdN(s)(8) with positive parameters (c , k , p) . This power law form for λ(t) in Equation (8) was applied in the geological model called Omori’s law, and used to predict the rate of aftershocks caused by an earthquake. Definition 5 (D-dimensional Hawkes Process) . The D-dimensional Hawkes process (see Embrechts et al. 2011) is a point process ~ N(t) = (Ni(t))D i=1which is characterized by its intensity vector~ λ(t) = (λi(t))D i=1such that: λi(t) = λi+Zt 0µij(t−s)dNj(s)(9) where λi>0, and M(t) = (µij(t)) is a matrix-valued kernel such that: 1. it is component-wise non-negative: (µij(t)) ≥0for each 1≤i,j≤D 2. it is component-wise L1-integrable In matrix-convolution form, Equation (9)can be written as ~ λ(t) = ~ λ+M∗d~ N(t)(10) where~ λ(t) = (λi)D i=1. Definition 6 (Nonlinear Hawkes Process) . The nonlinear Hawkes process (see, e.g., Zheng et al. 2013) is defined by the intensity function in the following form: λ(t) = hλ+Zt 0µ(t−s)dN(s)(11) where h(·) is a nonlinear function with support in R+ . Typical examples for h(·) are h(x) = 1x∈R+ and h(x) = ex. Remark 1. Many other generalizations of Hawkes processes have been proposed. They include mixed diffusion–Hawkes models Errais et al. (2010), Hawkes models with shot noise exogenous events Dassios and Zhao (2011), and Hawkes processes with generation dependent kernels Mehrdad and Zhu (2014), to name a few. 2.2. Compound Hawkes Processes In this section, we define nonlinear compound Hawkes process with N -state dependent orders. The dependent orders means the dependency of both, the type of a book event and its corresponding inter-arrival times, on the type of the previous book event. We also consider special cases of this general compound Hawkes process. Definition 7 (Nonlinear Compound Hawkes Process with n -state Dependent Orders (NLCHPnSDO) in Limit Order Books).Consider the price process St St=S0+ N(t) ∑ k=1 a(Xk)(12) Risks 2020,8, 28 6 of 25 where Xk is a continuous time n -state Markov chain, a(x) is a continuous and bounded function on the state space X:={ 1, 2, ..., n} , N(t) is the nonlinear Hawkes process (see, e.g., Zheng et al. (2013) defined by the intensity function in the following form (see Equation (11)): λ(t) = hλ+Zt 0µ(t−s)dN(s) where h(·) is a nonlinear increasing function with support in R+ . We note that in Brémaud and Massoulié (1996) it was shown that, if h(·) is α -Lipschitz (see Brémaud and Massoulié 1996) such that α||h||L1< 1, then there exists a unique stationary and ergodic Hawkes process satisfying the dynamics of Equation (11) . We shall refer to the process in Equation (12)as a Nonlinear Compound Hawkes Process with n-State Dependent Orders (NLCHPnSDO). This nonlinear compound Hawkes process will be the foundation for our studies throughout this paper. In the following subsection, we will introduce four specific examples, which will be used for our empirical investigations of the mid-price processes. 2.3. General Compound Hawkes Processes Definition 8 (General Compound Hawkes Process With N-state Dependent Orders (GCHPnSDO)) . Suppose that Xk is an ergodic continuous-time Markov chain, independent of N(t) , with state space X={1, 2, ..., n}, N(t) is a one-dimensional Hawkes process defined in Definition 4and a(x) is any bounded and continuous function on X . We define the General Compound Hawkes process with N-state Dependent Orders (GCHPnSDO) by the following process: St=S0+ N(t) ∑ k=1 a(Xk)(13) Note that this process can be recovered from Equation (12)by letting h(x) = x. Definition 9 (General Compound Hawkes Process with Two-State Dependent Orders (GCHP2SDO)) . Suppose that Xk is an ergodic continuous time Markov chain, independent of N(t) , with two states { 1, 2 } . Then, Equation (13)becomes St=S0+ N(t) ∑ k=1 a(Xk)(14) where a(Xk) takes only the values a( 1 ) and a( 2 ) . Of course, we can view this as a special case of the n-state case, where n= 2. This model was used in Swishchuk et al. (2017) for the mid-price process in limit order books with non-fixed tick δand two-valued price changes. Definition 10 (General Compound Hawkes Process with Dependent Orders (GCHPDO)) . Suppose that Xk∈ {−δ,δ}and that a(x) = x, then Stin Equation (13)becomes St=S0+ N(t) ∑ k=1 Xk(15) This type of process can be a model for the mid-price = (bidprice +askprice)/ 2in limit order books, where δ is a fixed tick size and N(t) is the number of order arrivals up to time t . We shall call this process a General Compound Hawkes Process with Dependent Orders (GCHPDO). This is a generalization of the previous process, obtained by letting a(1) = −δand a(2) = δ. Having defined several modifications of Hawkes processes, we now prove diffusion limit theorems and LLNs for each price process in the following section. These diffusion processes will be used for our exploration of the applicability of this model to real world limit order book data. Risks 2020,8, 28 7 of 25 Since order arrivals and cancellations are very frequent, with many of the applications and practical needs occurring at the millisecond time scale (e.g., order liquidations, order executions, etc.), we are interested in the dynamics that occur on a large time scale, such as tens of second or minutes. Therefore, in the following section, we consider the scale nt instead of t , where t is in milliseconds and n can be 1000, 10,000, making nt large with respect to t. 3. Diffusion Limits and LLNs for GCHP 3.1. Diffusion Limit and LLN for NLCHPnSDO We consider the mid-price process Stdefined in Definition 7, namely St=S0+ N(t) ∑ k=1 a(Xk) where Xk is a continuous time N -state Markov chain and a(x) is a continuous bounded function on the state space X={ 1, 2, ..., n} . Nt is the number of price changes up to time t , described by the nonlinear Hawkes process given in Equation (11). Theorem 1 (Diffusion Limit for NLCHPnSDO) . Let Xk be an ergodic Markov chain with N -states {1, 2, ..., n}and with ergodic probabilities (π∗ 1,π∗ 2, ..., π∗ n). Let also Stbe as defined in Definition 7, then Snt −N(nt)·ˆ a∗ √n n→∞ −−−→ ˆ σ∗qE[N[0,1]]Wt(16) where Wt is a standard Wiener process and E[N[ 0,1 ]] is the mean of the number of arrivals on a unit interval under the stationary and ergodic measure. Furthermore, 0<ˆ µ:=Z∞ 0µ(s)ds <1and Z∞ 0sµ(s)ds <∞(17) (ˆ σ∗)2:=∑ i∈X π∗ iv(i) v(i) = b(i)2+∑ j∈X (g(j)−g(i))2P(i,j)−2b(i)∑ j∈X (g(j)−g(i))P(i,j) b= (b(1),b(2),..., b(n))0 b(i):=a(Xi)−a∗:=a(i)−a∗ g:= (P+Π∗−I)−1b ˆ a∗:=∑ i∈X π∗ ia(Xi) (18) P is the transition probability matrix for Xk , i.e., P(i , j) = P(Xk+1=j|Xk=i) . Π∗ denotes the matrix of stationary distributions of P and g(j)is the jth entry of g. Proof. From Equation (12), we have Snt =S0+ N(nt) ∑ k=1 a(Xk)(19) and Snt =S0+ N(nt) ∑ i=1 (a(Xk)−ˆ a∗) + N(nt)ˆ a∗. (20) Risks 2020,8, 28 8 of 25 Therefore, Snt −N(nt)ˆ a∗ √n=S0+∑N(nt) k=1(a(Xk)−ˆ a∗) √n. (21) As long as S0 √n n→∞ −−−→ 0, we need only find the limit for a(Xk)−ˆ a∗ √n when n→+∞. Consider the following sums ˆ R∗ n:= n ∑ k=1 (a(Xk)−ˆ a∗)(22) and ˆ U∗ n(t):=n−1/2[(1−(nt −bntc))] ˆ R∗ bntc+ (nt −bntc)ˆ R∗ bntc+1](23) where b·c is the floor function. Following the martingale method from Vadori and Swishchuk (2015), we have the following weak convergence in the Skorokhod topology (see Getoor 1967): ˆ U∗ n(t)n→∞ −−−→ ˆ σ∗W(t)(24) We note that the results from Brémaud and Massoulié (1996) imply by the ergodic theorem that Nt t t→∞ −−→ E[N[0,1]] (25) or Nnt nt n→∞ −−−→ tE[N[0,1]]. (26) Using the change of time t→Nnt/n, we find that ˆ U∗ n(Nnt/n)n→∞ −−−→ ˆ σ∗W(tE[N[0,1]]) (27) or ˆ U∗ n(Nnt/n)n→∞ −−−→ ˆ σ∗qE[N[0,1]]W(t)(28) The result now follows from Equations (19)–(21). Lemma 1 (LLN for NLCHPnSDO) . The process Snt in Equation (19) satisfies the following weak convergence in the Skorokhod topology (see Getoor 1967): Snt n n→∞ −−−→ ˆ a∗E[N[0,1]]t(29) where ˆ a∗is defined in Equation (18), respectively. Proof. From Equation (12), we have Snt n=S0 n+ N(nt) ∑ k=1 a(Xk) n(30) Risks 2020,8, 28 15 of 25 In this case, −δ will be state one, and δ with be state two. This results in the transition matrix P given below: P="pdd 1−pdd 1−puu puu # After determining our parameters and transition probabilities, we calculate a∗ and σ in Table 4 together with puu and pdd. Table 4. Provided above are the values for s∗ , σ as well as the probabilities of an upward/downward movement given an upward/downward movement for each of the five stocks in question. pdd puu σa∗ AAPL 0.4956 0.4933 0.0049 −1.1463 ×10−5 AMZN 0.4635 0.4576 0.0046 −2.7373 ×10−5 GOOG 0.4769 0.4461 0.0046 −1.4301 ×10−4 MSFT 0.6269 0.5827 0.0062 −2.7956 ×10−4 INTC 0.6106 0.5588 0.0059 −3.1185 ×10−4 Provided these values, the asymptotic method from Section 3and its Corollaries can be used to study the link between order flow in our model and price volatility. The estimated volatility of price changes is expressed in terms of parameters describing arrival rate of limit orders. Allowing us to test our claim that our model accurately describes the mid-price process. We recall that the mid-price = (bidprice +askprice)/2, and the following diffusion limit Snt −N(nt)a∗ √n n→∞ −−−→ σsλ 1−α/βW(t). (48) If the data satisfy our proposed model, then, after considering large windows of time (5 min, 10 min, 20 min), we would expect to see the empirical and theoretical standard deviations to follow each other closely. To test this, we compare the equivalent process, constructed by multiplying the LHS and RHS by √n . Then, cutting our data into disjoint windows of size n , specifically [in , (i+ 1 )n] with t= 1 and by setting the left bound as our starting time, we can calculate Snt −N(nt)a∗ for each individual window and give a generalized formula for this below: S∗ i=S(i+1)nt −Sint −(N((i+1)nt)−N(int))a∗(49) This gives a collection of values {S∗ i} over which we compute the standard deviation. If our model is accurate, we would would expect that std{S∗ i} ≈ √ntσsλ 1−α/β, where t=1. (50) We plot the empirical standard deviation against the theoretical one for various window sizes starting at 10 s and increasing in steps of 10 s until we reach 20 min; this is illustrated in Figure 3. Several important remarks should be made at this point. It is clear that, while the model accurately predicts the overall trend for MSFT and INTC, we severely underestimate the variability in the mid-price process for APPL, AMZN and GOOG. Furthermore, as the window size increases, the overall spread in the data increases. We attribute this to the decreasing sample size imposed on us as we increase the window size. For example, when we consider a 20 min window, we can only construct 27 disjoint windows in the 9 h trading day forcing us to deal with the problem of predicting a ’population’ standard deviation from a increasingly small sample. We remedy this by using a variance stabilizing transformation in later sections. Specifically, a popular method for a Poisson process is to Risks 2020,8, 28 16 of 25 take the square root of our empirical and theoretical standard deviations. This makes it possible to qualitatively view the overall trend in the data, gaining a clearer idea of goodness of fit from there. Figure 3. Each figure compares the empirical standard deviation for a fixed window size to the theoretical standard deviation. We have plotted an empirical standard deviation for all n from 10 s to 20 min in step sizes of 10 s. Each empirical standard deviation corresponds to a single point in the scatter plot, and the plotted curve corresponds to the predicted theoretical value. 4.3. General Compound Hawkes Process with Two Dependent Orders As of now, we have considered a fixed δ related to the trading tick size. However, if we consider the mid-price changes for APPL, AMZN and GOOG the assumption of a fixed tick size is violated. In fact, we observe that approximately 61%, 53% and 71% of all mid-price changes are larger than half Risks 2020,8, 28 17 of 25 a tick size, which is opposed to what we observe for MSFT and INTC where all mid-price changes occur at the half tick size, we illustrate this for AAPL, AMZN and GOOG in Figure 4. Figure 4. We can see clearly that the change in the mid-price is often larger than a half tick. These mid-price changes make up a significant portion of the actual data, contradicting the assumption needed for the CHPDO model that the mid-price changes occur on average at a half tick size. It is clear in Figure 4that additional considerations need to be made. A simple way to include the variability in mid-price movements in our model is to introduce a(Xi) as described in Definition 9. It is of course necessary to determine the values of a(·) for each state of our Markov chain. A naive method is to take the mean of the downward and upward mid-price movements and assign them to a(1)and a(2), respectively. We provide these values in Table 5. Table 5. a(i) is the average of the upward or downward mid-price movements. Following our previous convention, the first state will be associated with the mean of all downward mid-price movements and the second state will be associated with the mean of all upward mid-price movements. a(1)a(2) AAPL −0.0172 0.0170 AMZN −0.0134 0.0133 GOOG −0.0302 0.0308 In this step, we have only endeavoured to better realize the actual price movements in our data. Therefore, when we observe a downward mid-price movement, we continue to assign it to state one and similarly for an upward price movement we continue to assign it to state two. It follows that our transition matrix will remain the same. Then, using these new state values, we recalculate a∗ and σ , Risks 2020,8, 28 18 of 25 providing them in Table 6. The effect of these changes is investigated in Figure 5. Note that in Figure 5 we have used the variance stabilizing transformation discussed earlier in order to better visualize the overall trend in our data. Table 6. Above, we have the values for a∗ , σ , as well as the probabilities of an upwards/downwards movement, given an upwards/downwards movement for the three stocks of interest. pdd puu σa∗ AAPL 0.4956 0.4933 0.0169 −1.5624 ×10−4 AMZN 0.4635 0.4576 0.0123 −1.0475 ×10−4 GOOG 0.4769 0.4461 0.0282 −5.5095 ×10−4 Notice that there is a significant qualitative improvement in the fits for AAPL and GOOG in Figure 5, but the variability in mid-price movements for AMZN is still clearly underestimated by our model. The unexplained variance may be captured by investigating an N -state Markov chain since the additional transition probabilities could explain the variability missing in the 2-state case. AAPL AMZN GOOG Figure 5. A comparison of the empirical standard deviation for a fixed window size n to the theoretical standard deviation for AAPL, AMZN and GOOG using the 2-state dependent order model. We have plotted the empirical standard deviation for all n from 10 s to 20 min in step sizes of 10 s. Each empirical standard deviation corresponds to a single point in the scatter plot and the plotted curve corresponds to the predicted theoretical value. Visually, there is a significant improvement for all stocks, although the theoretical standard deviation for AMZN is still underestimating the empirical variability. Risks 2020,8, 28 19 of 25 4.4. General Compound Hawkes Process with NDependent Orders We recall the N -state model described in Definition 8. The immediate question becomes how best to choose the state values. We modify the quantile based approach from Swishchuk et al. (2017). After calculating the mid-price changes, we separate the data into upward and downward price movements. Then, we calculate evenly distributed quantiles for both data sets. Depending on the data, several quantiles may be identical, we reject any duplicates. We thus obtain a list of bounds which we complete by adding the minimum observed value if necessary. To determine the state values a(Xi) , we take the average of all mid-price changes located between two neighbouring boundary values. Furthermore, we assign a mid-price change to state i if it is greater than or equal to the (i− 1 ) th boundary and strictly less than the i th boundary. An exception is made for the largest upper bound where equality is permitted at both ends. As we could not capture the full variability of mid-price changes for AMZN in the previous method, we investigate it for this case. Furthermore, for tractability, we only consider 14 boundary values from which we obtain a 12-state Markov chain. Instead of providing the transition matrix, we provide the ergodic probabilities for the transition matrix and the associated states in Table 7. Table 7. Above, we have provided the state, associated ergodic probabilities and state values a(i) for AMZN, given a 12-state Markov chain that was obtained from choosing a 16 quantile method. AMZN iπ∗ ia(Xi) 1 0.0275 −0.0524 2 0.0281 −0.0318 3 0.0264 −0.0250 4 0.0382 −0.0200 5 0.0576 −0.0150 6 0.3249 −0.0064 7 0.2321 0.0050 8 0.0923 0.0100 9 0.0578 0.0150 10 0.0353 0.0200 11 0.0412 0.0271 12 0.0387 0.0476 In order to compare the two state and N -state approaches, we first take a qualitative approach and plot the two theoretical and empirical standard deviations against each other in Figure 6. When we compare the mean squared residuals, the 2-state model discussed before has mean squared error 0.0208 while our 12-state Markov chain brings that down to 0.0125. Considering an even larger Markov chain with 24-states, we are only able to obtain a meager improvement to 0.0123, suggesting that there is some underlying variance in the mid-price process not captured by our model. We investigate these more quantitative measures in the following subsection. First, comparing the models against a numerical best fit, and then investigating the mean squared error, we gain a better quantitative understanding of the overall improvement obtained from each model increasing the number of quantiles. While we do not propose a method for selecting an N -state model in general, a possible piece of criteria is to halt when there is no appreciable gain in the MSE for more information (see Swishchuk et al. 2017, p. 17). In light of this, in the following section, we use k-fold cross-validation to make a determination. Risks 2020,8, 28 20 of 25 AMZN Figure 6. We consider the N-state model for AMZN discussed previously in the paper. While there is a slight improvement against the original fit, the model still struggles to perfectly predict the variability in the mid-price changes of our data. 5. Quantitative Analysis While our model does visually appear to fit the expected variability in four of the five cases, it still fails to capture the complete dynamics of mid-price changes seen in our AMZN data. We investigate the mean square error of our models with a varying number of quantiles in Table 8. This gives a good indication as to whether the N -state model is a better fit for our data. If we look closely at AAPL and GOOG, we see that the N -state case can still improve our results from the 2-state case. For AAPL, we constructed a 17-state Markov chain by taking 16 quantiles on the downward movements and 16 quantiles upward movements. This resulted in a mean squared error of 0.0036 that is approximately a 28% improvement to the 2-state case where the mean squared error was 0.0050. Even more extreme, using a 25-state Markov chain for GOOG which was constructed similarly, we observed a mean squared error of 0.0046, which is a 60% improvement from the 2-state case with a mean squared error of 0.0115. We conclude with Table 8, which provides the mean residuals for AAPL, AMZN and GOOG with several Markov chains constructed from various numbers of quantiles. We also include mean residuals for INTC and MSFT for comparison. Another quantitative measure of our fit would be to compare them to the one which best minimizes the residuals. Notice that each of our models assumes that the standard deviation proportionally to the square root of the time step. Therefore, we can estimate the best possible coefficient by minimizing L2 -norm or equivalently the mean squared error. We provide plots of these hypothetical best fits against the empirical data and theoretical fits in Figure 7. We also provide the coefficients for the theoretical fits and regression in Table 9in order to have a more quantitative comparison. They are calculated using least-square regression. Risks 2020,8, 28 21 of 25 Table 8. We list the mean residuals for several Markov chains with varying numbers of states. These were generated using our modified quantile approach choosing to start with 2, 8, 16 or 32 quantiles. We see that, in general, the mean residual decreases to some lower limit where we can no longer perform any better. Recall that the only observed mid-price changes for INTC and MSFT were of a half tick size, and any increase in the number of quantiles will result in the same performance. CHPDO 2 8 16 32 AAPL 0.2679 0.0050 0.0036 0.0036 0.0036 AMZN 0.1122 0.0208 0.0131 0.0124 0.0123 GOOG 0.4036 0.0115 0.0048 0.0045 0.0047 INTC 1.7917 ×10−51.7917 ×10−51.7917 ×10−51.7917 ×10−51.7917 ×10−5 MSFT 1.0586 ×10−41.0586 ×10−41.0586 ×10−41.0586 ×10−41.0586 ×10−4 We notice that in each case the errors are close to, or under five percent, with AMZN being the biggest offender. This is consistent with the discussion provided throughout our analysis and highlights the general applicability of our model. Table 9. The coefficients calculated for AAPL, AMZN and GOOG are generated using a Markov chain created by 16 quantiles on the upward and downward movements, while the coefficients for INTC and MSFT are obtained from the CHPDO case. Theoretical Coefficient Regression Coefficient Percent Error AAPL 0.02868 0.02828 1.42% AMZN 0.01450 0.01831 20.8% GOOG 0.02883 0.03023 4.63% INTC 0.00186 0.00193 3.4% MSFT 0.00231 0.00246 6.4% Cross-Validation We conclude this subsection using a method similar to k-fold cross-validation for model selection. Generally, the data are shuffled randomly and then partitioned into k sets of approximately equal size. From these k sets, one is selected for testing the model while the remaining are kept for training it. For our data, we are unable to perform this shuffling as there is an important ordering imposed by the arrival times. Fortunately, since we are considering high-frequency models for the mid-price changes, we expect a minimal amount of correlation in the volatility of the price process over large windows of time. Therefore, dividing the data into twelve 30 min windows selecting them one at a time should still be meaningful. The training data will need to be glued back together after the thirty-minute window is removed. This is done by assuming the events never occurred. We provide an example below for how one would remove the second thirty-minute window in Tables 10 and 11. This introduces a small error into the arrival times, but only at a single point in a data set of several thousand arrival times. Had we randomly removed points, or shuffled the data, we could introduce many errors which could significantly change the underlying arrival process. At this point, we perform the same fitting procedure on the training data and after the fitting is completed, we can compute the empirical standard deviations on the testing data to obtain a mean squared error. This determines how well the model was able to predict the training data, calculating this for each window of time and averaging provides a score for how well the model was able to make predictions. We recall that, for MSFT and INTC, the only observed prices changes were a single tick, and it is not possible to outperform the original CHPDO model. Instead, we consider AAPL, AMZN and GOOG calculating scores for CHPDO, GCHP2SDO and various GCHPnSDO models. Specifically the set of GCHPnSDO models are constructed using the quantile method, allowing for 2–14 quantiles. Risks 2020,8, 28 22 of 25 This was chosen only to make the recording more tractable; the resulting scores are provided in Table 12. AAPL AMZN GOOG INTC MSFT Figure 7. A qualitative comparison of the regression to the theoretical model. For APPL, AMZN and GOOG, we have used a Markov chain generated from 16 quantiles taken on the upward movements and downward movements. For INTC and MSFT, we have taken the CHPDO coefficient since a different coefficient is not possible with the other models. Risks 2020,8, 28 23 of 25 Table 10. An imagined sequence of price change events before the second thirty minute window has been removed. Time (s) ··· 1789 1795 1803 ··· 3193 3601 3608 ··· Event ··· n−1 n n + 1 ··· n+k n+k+1 n+k+2 ··· Table 11. An imagined sequence of price change events after the second thirty minute window has been removed. Time (s) ··· 1789 1795 1801 1808 ··· Event ··· n−1 n n+k+1 n+k+2 ··· Table 12. A table of testing scores for various models. Computations were performed for up to the 40-quantile case for each stock, and this table represents a sample of that data. We note that the models stop obtaining appreciable performance gains at around 4–7 quantiles, fluctuating around the same scores. AAPL AMZN GOOG CHPDO 0.0348 0.0142 0.0436 GCHP2SDO 0.0049 0.0048 0.0044 2-Quantiles 0.0042 0.0041 0.0036 3-Quantiles 0.0041 0.0042 0.0035 4-Quantiles 0.0040 0.0040 0.0035 5-Quantiles 4.0341 ×10−33.9878 ×10−33.4593 ×10−3 6-Quantiles 4.0362 ×10−33.8971 ×10−33.4168 ×10−3 7-Quantiles 4.0184 ×10−33.8971 ×10−33.3977 ×10−3 8-Quantiles 4.0097 ×10−33.8656 ×10−33.3881 ×10−3 9-Quantiles 4.0026 ×10−33.8124 ×10−33.3976 ×10−3 10-Quantiles 3.9984 ×10−33.8124 ×10−33.425 ×10−3 11-Quantiles 3.9971 ×10−33.81594 ×10−33.3867 ×10−3 12-Quantiles 3.9826 ×10−33.81463 ×10−33.3906 ×10−3 13-Quantiles 3.9804 ×10−33.81517 ×10−33.3562 ×10−3 14-Quantiles 3.9725 ×10−33.81427 ×10−33.3549 ×10−3 Overall, there is a meaningful gain in performance for the N -state models, especially when we take into account the previous discussion. The cross-validation we’ve performed provides a reasonable criteria for model rejection, similar to when we stop observing appreciable gains in the mean squared error. From our analysis, the best model would be one constructed from 4–7 quantiles, depending on the stock of interest. 6. Conclusions and Future Work Overall, the N -state model outperforms the others when the number of states is kept small. It generates fits for four out of the five datasets that are reasonable and providing reductions in the mean squared error by upwards of 25%. While not able to capture the full dynamics observed in AMZN, it appears to be a strong candidate for a simple model of price dynamics observed in our data. Further investigation would be necessary to determine what causes the additional volatility observed in AMZN, and potentially implement a more robust model which captures this. The potential users of our models are practitioners working in the financial industry who wish to implement our results in high-frequency and algorithmic trading combined with their industrial knowledge. Author Contributions: Both authors have contributed equally to the paper. All authors have read and agreed to the published version of the manuscript. Funding: This research was funded by NSERC Grant No. RT732266. Risks 2020,8, 28 24 of 25 Conflicts of Interest: The authors declare no conflict of interest. References Azizpour, Shahriar, Kay Giesecke, and Gustavo Schwenkler. 2018. Exploring the sources of default clustering. Journal of Financial Economics 129: 154–83. doi10.1016/j.jfineco.2018.04.008. [CrossRef] Aït-Sahalia, Yacine, Julio Cacho-Diaz, and Roger J.A. Laeven. 2015. Modeling financial contagion using mutually exciting jump processes. Journal of Financial Economics 117: 585–606. [CrossRef] Bacry, Emmanuel, Sylvain Delattre, Marc Hoffmann, and Jean-François Muzy. 2011. Modeling microstructure noise with mutually exciting point processes. arXiv. arXiv:1101.3422. Bacry, Emmanuel, Iacopo Mastromatteo, and Jean-François Muzy. 2015. Hawkes processes in finance. arXiv. arXiv:1502.04592. Bartlett, Maurice Stevenson. 1963.The spectral analysis of point processes. Journal of the Royal Statistical Society. Series B (Methodological) 25: 264–296. [CrossRef] Bauwens, Luc, and Nikolaus Hautsch. 2009. Modelling Financial High Frequency Data Using Point Processes. Berlin/Heidelberg: Springer, pp. 953–79._41. [CrossRef] Bowsher, Clive G. 2007. Modelling security market events in continuous time: Intensity based, multivariate point process models. Journal of Econometrics 141: 876–912. [CrossRef] Brémaud, Pierre, and Laurent Massoulié. 1996. Stability of nonlinear hawkes processes. The Annals of Probability 24: 1563–88. [CrossRef] Buffington, John, and Robert J. Elliott. 2002a. American options with regime switching. International Journal of Theoretical and Applied Finance 5: 497–514. [CrossRef] Buffington, John, and Robert J. Elliott. 2002b. Regime switching and european options. In Stochastic Theory and Control. Edited by B. Pasik-Duncan. Berlin/Heidelberg: Springer, pp. 73–82. Carstensen, Lisbeth. 2010. Hawkes Processes and Combinatiorial Transcriptional Regulation. København: Museum Tusculanum. Cartea, Álvaro, Sebastian Jaimungal, and José Penalva. 2015. Algorithmic and High-Frequency Trading. Cambridge: Cambridge University Press. Cartea, Álvaro, Sebastian Jaimungal, and Jason Ricci. 2014. Buy low, sell high: A high frequency trading perspective. SIAM Journal on Financial Mathematics 5: 415–44. [CrossRef] Chávez-Casillas, Jonathan A., Robert J. Elliott, Bruno Rémillard, and Anatoliy V. Swishchuk. 2019. A level-1 limit order book with time dependent arrival rates. Methodology and Computing in Applied Probability 21: 699–719. [CrossRef] Chavez-Demoulin, Valérie, and J. A. McGill. 2012. High-frequency financial data modeling using hawkes processes. Journal of Banking & Finance 36: 3415–26. Cohen, Samuel N., and Robert J. Elliott. 2013. Filters and smoothers for self-exciting Markov modulated counting processes. arXiv. arXiv:1311.6257. Cox, David R. 1955. Some statistical methods connected with series of events. Journal of the Royal Statistical Society Series B (Methodological) 17: 129–64. [CrossRef] Daley, Daryl J., and David Vere-Jones. 2003. An Introduction to The Theory of Point Processes: Volume I: Elementary Theory and Methods. Berlin/Heidelberg: Springer Science & Business Media. Dassios, Angelos, and Hongbiao Zhao. 2011. A dynamic contagion process. Advances in Applied Probability 43: 814–46. [CrossRef] Embrechts, Paul, Thomas Liniger, and Lu Lin. 2011. Multivariate hawkes processes: An application to financial data. Journal of Applied Probability 48A: 367–78. [CrossRef] Engle, Robert F., and Jeffrey R. Russell. 1998. Autoregressive conditional duration: A new model for irregularly spaced transaction data. Econometrica 1998: 1127–62. [CrossRef] Errais, Eymen, Kay Giesecke, and Lisa R. Goldberg. 2010. Affine point processes and portfolio credit risk. SIAM Journal on Financial Mathematics 1: 642–65. [CrossRef] Filimonov, Vladimir, David Bicchetti, Nicolas Maystre, and Didier Sornette. 2014. Quantification of the high level of endogeneity and of structural regime shifts in commodity markets. Journal of International Money and Finance 42: 174–92. doi:10.1016/j.jimonfin.2013.08.010. [CrossRef]