Full text
Title: Volatility Estimation With Microstructure Noise Author: Ferran Vilar Pagès Advisor: Luis Ortiz Supervisor: Simona Sanfelici Department: Statistics and Operations Research University: Universitat Politècnica de Catalunya Academic year: 2021-2022 Interuniversity Master in Statistics and Operations Research UPC-UB
Universitat Polit`ecnica de Catalunya Facultat de Matem`atiques i Estad´ıstica Tesi de m`aster Volatility Estimation With Microstructure Noise Ferran Vilar Pag`es Advisor: Luis Ortiz Supervisor: Simona Sanfelici Estad´ıstica i Inverstigaci´o Operativa
I would like to express my sincere gratitude to my supervisor Simona Sanfelici. Without her daily guidance and motivation this project would not have been possible. Also thank my friends for their support through these academic years. Special and sincere mention to my family for investing and believing in me no matter what.
Abstract Keywords: High Frequency Data, Integrated Volatility, Quadratic Variation, Realized Volatility, Geometric Brownian Motion, Microstructure Effects, Sub-sampling, Two Scale, Kernel, Pre-Averaging, Mean Square Error MSC2000: 62G05 Aquest treball revisa i estudia la literatura econom`etrica sobre l'estimaci´o de la variaci´o quadr`atica fent ´us de la vari`ancia realitzada (RV). Quan el proc´es subeditat es comporta com una semimartingale l'estad´ıstic RV estima de manera consistent la variaci´o quadr`atica. Es presenten les bases i assumpcions necess`aries per construir aquest estimador. El raonament per`o, canvia dr`asticament quan es treballa amb preus observats. La pres`encia de soroll en les dades provoca que l'estimador ja no sigui consistent. Tot i aix`o, hi ha maneres de filtrar aquest soroll i eliminar el biaix que la pres`encia de soroll provoca en les estimacions. En particular es treballa amb tres estimadors de la vari`ancia quadr`atica. Cada estimador dep`en d’uns par`ametres que s'estimaran mitjan¸cant minimitzaci´o de la vari`an¸ca asimpt`otica i mitjan¸cant criteris de minimitzac´o del Error Quadr`atic Mig (EQM). Provarem els estimators en una base de dades amb preus Bitcoin minut a minut entre els per´ıodes 7 de febrer 2022 i 27 d'abril 2022. Discutirem si els estimadors son eficients al estimar la variaci´o quadr`atica, i en concret el m`etode Pre-Averaging ser`a el que millors resultats oferir`a.
Abstract Keywords: High Frequency Data, Integrated Volatility, Quadratic Variation, Realized Volatility, Geometric Brownian Motion, Microstructure Effects, Sub-sampling, Two Scale, Kernel, Pre-Averaging, Mean Square Error MSC2000: 62G05 This thesis revises and looks the econometric literature on estimating the quadratic variation of financial prices. When the underlying process is a semimartingale, the very well known statistic Realized Volatility (RV) consistently estimates the quadratic variation. We present the fundamentals and additional assumptions to build the estimator. Moreover the position changes dramatically with observed asset prices in high frequency data, since the estimator is no longer consistent due to the presence of noise. Nevertheless, the quadratic variation can be still estimated correcting the bias due to the presence of noise. In particular, three estimators based on RV are presented to filter out the noise and estimate the quadratic variation. Optimal parameters through minimizing the asymptotic variance of the estimation error and by minimizing the Mean Square Error (MSE) are also discussed. We test the estimators on a minute-by-minute sample returns from Bitcoin prices over the period between the 7th of February, 2022 and the 27th of April, 2022. We argue that the methods proposed produce efficient estimates of the quadratic variation, with the Pre-Averaging giving the best estimations.
i Acronyms GBM Geometric Brownian Motion RV Realized Volatility IV Integrated Volatility TTS Two Time Scales KB Kernel Based PA Pre-Averaging CLT Central Limit Theorem MSE Mean Square Error
Contents List of Tables 1 List of Figures 3 Chapter 1. Introduction 5 1. Overview 5 2. Bitcoin 6 Chapter 2. Mathematical Modeling 9 1. The Geometric Brownian Motion and Features 9 2. High Frequency Data Framework 12 3. General Pricing Model 13 Chapter 3. Microstructure Noise and RV Estimation 17 1. Observable Prices 17 2. Effects of Noise on IV Estimation 18 3. Sparse RV 20 4. Average Realized Volatility 22 Chapter 4. Estimation Methods 25 1. Two Time Scales 25 2. Kernel-Based Estimator 27 3. Pre-Averaging Estimator 29 Chapter 5. Bitcoin Realized Volatility Analysis 33 1. Data Description 33 2. Data Analysis 35 3. Estimators Implementation 37 4. Comparison 42 Chapter 6. Conclusions 47 References 49 iii
2. BITCOIN 7 49% of the market, followed by Ethereum, which holds a 19% share. In December 2016, Bitcoin accounted for 89% of the total market capitalization. Not only the private sector have taken advantage of the new features, furthermore, people were the main target for Bitcoins platform. Bitcoin aim was, and still is, for every person to have access to financial services without relying on third party institutions as traditionally. Then Nakamoto tries to get rid of the central organisations as we know them, creating a network that works as a trusted third party in transactions. Cryptocurrencies have also raised a continuous interest for the financial econometric world (see for instance Catania et al. 2019, Zheng et al. (2017, 2018) or Christidis 2016). A large amount of literature can be found about volatility measures, the long term memory of the financial series, forecasting on returns, and many other fields. Every Bitcoin network user should be interested in the behavior of Bitcoin due to the possible systemic risk they would face by entering in this new market.
Chapter 2 Mathematical Modeling This section aims to approach the main tools and understanding of the mathematics behind the estimation methods (See A.Mykland and Lan Zhang, 2012). Section 2.1 motivates the RV estimation method for the returns variance. Section 2.2 covers estimation in the high frequency framework. Section 2.3 describes how market prices are generally modelled as Itˆo process. 1. The Geometric Brownian Motion and Features In the financial world it is very common to generate prices following the very well known Geometric Brownian Motion (GBM) process. The GBM is a theoretical price modelling which allows to generate prices under an additive model on the log-scale, although it has several limitations and drawbacks. It will be shown in later sections that market prices don’t follow a GBM process, but it is widely used to price assets, options, futures and many other financial instruments in a very simple way. Consider a stochastic process (St)t≥0describing the price of a financial asset and let Xt=logSt be its logarithmic price. The GBM model is (1) Xt=X0+µt +σWt with µcalled drift and σknown as volatility. They measure the expected growth and standard deviation of the annual log returns, respectively. The GBM model assumes that both µand σare constants. On the other hand, Wtis the so-called Brownian Motion (BM) and is defined as 9
10 2. MATHEMATICAL MODELING (1) W0= 0. (2) t→Wtis a continuous function of t. (3) Whas independent increments: if t>s>u>v, then Wt−Wsis independent of Wu−Wv (4) for t>s,Wt−Wsis normal with mean zero and variance t−s, N(0,t−s). 1.1. Estimation with GBM prices. Under some regularity conditions, prices can be generated under GBM process in an arbitrary and friction-less environment. It is instructive first to consider a discrete framework for estimation in this model where t= 0 is the begining of the trading day, and t=Tthe end of the trading day. Later on we will assess the continuous model to give the general sense. Some assumptions have to be made when dealing with GBM prior to estimating the parameters. Let’s assume our data have nobservations at time tn,i, i = 0,1, . . . , n and, for the sake of simplicity, that they all are equally spaced. This consideration is not mandatory, one can actually not consider that prices are observed equally spaced. Although on a theoretic framework this consideration makes sense for simplicity because it allows us to build ∆tn=T n, i.e the time interval between prices. Log-prices at time tn,i are defined as Xtn,i and transactions take place every ∆tn. If we take the log-returns, i.e the differences on the log-prices ∆Xtn,i =Xtn,i −Xtn,i−1, it can be proven that ∆Xtn,i are independently and identically distributed following aN(µ∆tn, σ2∆tn) from the former GBM model above (1). One can then find unbiased and minimum variance estimators for the drift, µand variance parameter σ2of the process. ˆµn,MLE,UMV U =1 n∆tn n X i=1 ∆Xtn,i ˆσ2 n,MLE =1 n∆tn n X i=1 (∆Xtn,i −∆Xtn)2 (2) ˆσ2 n,UMV U =1 (n−1)∆tn n X i=1 (∆Xtn,i −∆Xtn)2 Here, MLE stands for the maximum likelihood estimator, and UMVU for the uniformly minimum variance unbiased estimator. Also note
1. THE GEOMETRIC BROWNIAN MOTION AND FEATURES 11 ∆Xtn=1 n n X i=1 ∆Xtn,i = ˆµn∆tn . The estimators above not only estimate the parameters, also clarify some basics. It follows immediately that ˆµn= (XT−X0)/T, i.e it depends only on the log-prices taken into consideration at the beginning and at the end of the time period T. Therefore, µcannot be consistently estimated for a fixed Tas n→ ∞. On the contrary, it is perhaps surprising that σ2can be estimated consistently for a fixed time period T as n→ ∞. To prove that, set the following statistic: Un,i = ∆Xtn,i /(σ∆t1/2 n). This will allow us to standardize the log-price difference variance. Then Un,i is iid with distribution N((µ/σ)∆t1/2 n,1). We can now adapt the variance estimator in (2) with data from the new statistic Un,i as follows: ˆσ2 n=σ2∆tn 1 (n−1)∆tn | {z } A n X i=1 (Un,i −Un,·)2 | {z } B where Un,·=1 n n X i=1 Un,i Part A is to adapt the estimator with the statistic we have built and B is the sum of the difference between standard normal random variables to the square. Further considerations on B have to be done prior to estimation regarding its distribution. B is distributed as χ2with n−1 degrees of freedom which yields to the following identity in law: ˆσ2 n L =σ2χ2 n−1 n−1, i.e. the two random variables have the same distribution. Moreover, E(ˆσ2 n) = σ2
12 2. MATHEMATICAL MODELING V ar(ˆσ2 n) = 2σ4 n−1 given that the the Expectation of a χ2with mdegrees of freedom is Eχ2 m=mand its Variance is V ar(χ2 m)=2m. Hence, it follows that the UMVU estimator for the variance parameter is consistent, namely ˆ σ2 n→σ2in probability as n→ ∞ for fixed T. 2. High Frequency Data Framework Above we have discussed and presented the classical estimators for the drift and variance parameter on an iid basis. When dealing with high frequency financial data, it is commonly used the non-centered estimator for the variance rather than the classical ones in (2). The reason is due to the fact that the mean of intraday returns are considered negligible in a context where n→ ∞ in a finite interval [0,T]. The non-centered estimator is the following: (3) ˆσ2 n,non−center =1 n∆tn n X i=1 (∆Xtn,i )2 and if we recall the ˆσ2 n,MLE estimator for the variance and the non-centered estimator in (3), we can show that they both yield to the same results in the limit. In fact ˆσ2 n,MLE =1 n∆tn n X i=1 (∆Xtn,i −∆Xtn)2 =1 n∆tnn X i=1 (∆Xtn,i )2−n(∆Xtn)2 = ˆσ2 n,non−center −∆tnˆµ2 n = ˆσ2 n,non−center −T nˆµ2 n (4) The last term tends to zero as n→ ∞, since ˆµndoes not depend on n. Hence, ˆσ2 n,non−center consistently estimates σ2and has the same asymptotic properties as σ2 n,MLE estimator. This holds for σ2 n,UMV U as well, since it is a rescalation of the maximum likelihood estimator for the variance.
3. GENERAL PRICING MODEL 13 The non-centered estimator is also known as the Realized Variance estimator for the Spot or Instantaneous Variance σ2. We remark that in the econometric literature this estimator is commonly denoted as Realized Volatility estimator, i.e the terms volatility and variance are considered as synonyms. From now on we will adopt this convention. 3. General Pricing Model We are about to address the main part of this chapter. In the financial econometric world, the drift and volatility of the log returns cannot be defined as a constant as stated until now. Pricing theory requires to make them time varying. Prior to present the general pricing model we define mathematical concepts to help the interpretation and the understanding of the new general process (Mykland and Zhang (2012)). First consider a stochastic process (Xt)t≥0, where the time variable t∈[0, T ]. Information sets. σ−fields Information is usually described by σ−fields. Consider a probability space (Ω,F,P), where Ω is the set of all possible outcomes w,Fis a collection of subsets, and Pis a probability measure for wto occur with values in [0,1]. Fis required to be a σ−field. Definition 2.1 A collection of events Fof Ω is a σ−field if (1) ∅, Ω ∈ F; (2) if F∈ F, then Fc= Ω −F∈ F; (3) if Fn, n = 1,2, ... are all in F, so is the event U∞ n=1Fn. Fcan therefore be considered as a collection of decidable events. Random Variables A random variable Xis a function Ω → ℜ that is measurable with respect to F, i.e such that ∀x∈ ℜ :{w∈Ω : X(w)≤x} ∈ F. Filtrations
14 2. MATHEMATICAL MODELING The evolution of our knowledge is defined by the filtration (Ft)t≥0up to time t. Since increasing time makes more sets decidable, Ftsatisfies that if s≤tthen Fs⊆ Ft. Adapted Processes One can say a process (Xt)t≥0is adapted to Ftif for all t∈[0, T], Xtis Ftmeasurable. Wiener Process We come back again to the Brownian Motion, but adapted to Ft. From now on, when referring to Wtit will be relative to Ftand will be called Wiener process. Specifically: Definition 2.2 The process (Wt)0≤t≤Tis an Ft-Wiener process if it is adapted to the information in Ftand (1) W0= 0. (2) t→Wtis a continuous function of t. (3) W has independent increments relative to Ft: if t>s, then Wt−Wsis independent of Fs. (4) for t>s,Wt−Wsis normal with mean zero and variance t−s, N(0,t−s). Quadratic Variation Consider a uniform decomposition of the interval [0,t] 0 = t0< t1< . . . < tn=t. For any Xtprocess, we define the quadratic variation of Xas the stochastic process (5) [X, X]t:= p−lim n→∞ n X i=1 (Xti−Xti−1)2 where the convergence is in probability. Stochastic Integrals
3. GENERAL PRICING MODEL 15 The central concept of a stochastic integral relies on a generalization of the Riemann integral where both integrand and integrator are stochastic processes. The Riemann integral is defined as follows: Zt 0 f(s)ds := lim n→∞ n−1 X i=0 f(si)∆s where 0 = s0< s1< . . . < sn=tand ∆s=t n and the stochastic integral generalisation would be to make f(si) and ∆sstochastic processes. Therefore, given a Wiener process Wt, the Itˆo stochastic integral is defined as: (6) Zt 0 f(s)dWs:= p−lim n→∞ n−1 X i=0 f(si)(Wsi+1 −Wsi) 3.1. The Itˆo Process. With all the statements above we are able to present the so-called Itˆo process (Xt)t≥0. (7) Xt=X0+Zt 0 µsds +Zt 0 σsdWs where µsis the drift function and σsis the instantaneous or sport volatility. The Itˆo process is also written in differential form for simplicity and it is also called Diffusion Process (8) dXt=µtdt +σtdWt. Itˆo processes are commonly used in finance to model asset log-prices. If we take into account the Quadratic variation and the Wiener process properties, it can be proven that given any Itˆo process Xtthe quadratic variation of Xis
16 2. MATHEMATICAL MODELING (9) [X, X]t=Zt 0 σ2 sds This quantity is also known as Integrated Volatility (IV) and is what we aim to estimate. Then a consistent estimator for IV can be naturally derived from Equation (5), the quadratic variation of (Xt)t≥0: (10) c IV t=RVn= n X i=1 (Xi−Xi−1)2≃Zt 0 σ2 sds. This estimator is commonly known in financial-econometric world as Realized Volatility (RV) and instead of taking log-price increments as in Equation (10) it is often used the returns for notation RVn=Pn i=1(ri)2, where ri=Xi−Xi−1. Moreover, it can be proved that RVnis an unbiased estimator of [X, X]t, namely E[RVn|Ft]=[X, X]t.
4. AVERAGE REALIZED VOLATILITY 23 G= K [ k=1 G(k). We can now define our Sub-sampling RV estimator in terms of the k-th subgrids as (19) g RV (k) n= n(k) X i=1 ˜r2 i. and if averages of estimator g RV (k) nacross K grids of average size ¯nare taken, one can construct g RV average ¯nwhich is distributed as (20) g RV average ¯n L ≈IVT+ 2¯nE[ϵ2] | {z } bias due to noise +h4¯n KE[ϵ4] + 4T ¯nZT 0 σ4 tdti1/2Ztotal | {z } total variance . Although bias has been reduced from of the initial g RV nestimator, and variance is now lower then it was with g RV sparse ¯n, Zang et al. 2005 showed that g RV (average) ¯n still was not able to estimate consistently the IVT.g RV (average) nestimator is also known in the econometrics world as Average Realized Volatility (ARV).
24 3. MICROSTRUCTURE NOISE AND RV ESTIMATION Fig. 3. IV estimators comparison over different frequencies. The blue line is the realized volatility sub-sampled for the latent Bitcoin log-prices. ARV is a very well known method in finance to filter out the noise and estimate the true signal. Although it can be seen in Figure 3 it is still a biased estimator for IVT. Many work has been done to fix this bias due to noise and many estimators have also been proposed to filter it out and allow to estimate the true variance of the price process.
Chapter 4 Estimation Methods The initial approach to the microstructure problem sets the basis behind more complex estimation methods. The philosophy behind sparse sampling is that the size of noise ϵis really small, and so if we have a small data set, then the effect of the noise will be limited. While true, we have seen that this method uses the data inefficiently. Moreover, the sub-sampling is more efficient in terms of using the data, but it is still biased estimator for the Integrated Volatility, as can be seen in Figure 3. Both methods do not estimate accurately the IVTof the process, but they do provide for some guidance on how to proceed to more complex schemes. In this chapter we present the Two Time Scales estimator,Kernel-Based estimator and Pre-Averaging estimator for the Integrated Volatility and we give evidences of their efficiency dealing with microstructure problem. 1. Two Time Scales Since sparse sampling and sub-sampling are not yet efficient estimators for the IVT, as can be seen in the preceding section, Zhang et al. (2005) proposed an estimator which successfully corrects the bias g RV (average) nhas. This is the well known Two Time Scales estimator and is defined as the combination of g RV (average) nand g RV n as follows: (21) g RV (T T S) n=g RV (average) ¯n−¯n ng RV n where g RV (average) ¯nis constructed by averaging sparse sampled RV estimators across Ksubgrids of average size ¯n≈n K. Furthermore, since the bias of g RV (average) ¯nis 25
26 4. ESTIMATION METHODS Ehg RV (average) ¯n−IVTi= 2¯nE[ϵ2] from Equations (15) and (20) respectively, it can be shown that, the Estimator (21) is unbiased under Assumption 3.1 as follows: Ehg RV (T T S) ni=Ehg RV (average) ¯n−¯n ng RV ni =IVT+ 2¯nE[ϵ2]−¯n n2nE[ϵ2]−IVT =IVT+ 2¯nE[ϵ2]−2¯nE[ϵ2]−¯n nIVT =IVT−¯n nIVT→IVT. (22) as n, ¯n→ ∞ such that ¯n n→0. 1.1. Optimal Parameters. The Two Time Scale estimator is a function that ultimately depends only on parameter K. Finding an expression for the optimal number of non-overlaping subgrids Kis of relevance for the estimator to be efficient. One can derive the asymptotic properties of the estimator which will lead on finding the number of non-overlaping subgrids that minimizes its variance. Zhang determines the optimal choice of Kas (23) K=cn2/3. At a rate of n1/6we have a convergence in law in terms of unknown parameter c (24) n1/6g RV T T S n−IVTL = 8c−2E(ϵ2)2+nξ2T!1/2 N(0,1). cis a constant and ξ2=4 3RT 0σ4 tdt is the realized quarticity. One can find such c∗ that minimizes the variance of Equation (24)
2. KERNEL-BASED ESTIMATOR 27 (25) c∗= 16E(ϵ2)2 Tξ2!1/3 . The second moment of the noise can be consistently estimated as E[ϵ2] = 1 2ng RV n. Regarding the rescaled quarticity ξ2, from Bandi-Russell 2003 and Zhang et al. (2005) we estimate it with the following equation (26) ξ2=4 3ZT 0 σ4 tdt ≈4 3m 3T m X i=1 ˜ri4 estimated by summing up mreturns at a lower frequency. Hence, one can choose Kbased on past data since c∗can be consistently estimated from data in past time periods. Alternatively, Zhang et al. (2005) also proposed a re-scaling or adjustment for g RV TTS n. When the data set is large but one wants to work or is only available to have a certain small amount of data, such a representative sample, they proved that a bias correction has to be done. The final estimator becomes (27) g RV T T S,adj. n=1−¯n n−1IVT. Both estimators are derived from Assumption 3.1 and have the same asymptotic properties. 2. Kernel-Based Estimator Estimating the quadratic variation when prices have the underlying effect of microstructure noise has, in a sense, similarities with studying the long term variance of a time series (see Andersen (1997) and Newey and West (1986)). RV is analogous to the sum of square variances, as seen in the first chapters, and kernel −based estimators rely on this work. The Two Scales Estimator form Zhang was the first to consistently estimate the quadratic variation of the prices. In this section we present the next estimator of Realized Kernel type
28 4. ESTIMATION METHODS The first to consider on using Kernel methods was Zhou (1996) to deal with the problem of microstructure noise on high frequency data. Zhou proposed for integrated volatility estimation the linear combination of auto-covariances between observed log-prices with a certain leads and lags. The estimator he proposed is known as Asymmetric Kernel Estimator (28) g RV KB n=w0bγ0+ 2 H X h=1 whbγh. where bγhare the realized auto-covariances for the returns γh= n X j=|h|+1 ˜rj˜rj−|h|. Zhou (1996) took H=1 and so the first order auto-covariance to estimate the expost variance. It can be proven that this estimator successfully eliminates the bias caused by the microstructure frictions. However, Hansen and Lunde (2006) proved the estimator was inconsistent at H=1. In their work they also discuss the properties of the estimator (28) for a more general framework with unrestricted H, w0= 1 and w1=m m−sbut it still was inconsistent. Given these inconsistencies, Barndorff-Nielsen et al. (2008) show that sutably weighting auto-covariance of the returns, the estimator g RV KB ncan be consistent at a certain rate. Using kernel weight functions k(x), Barndorff-Nielsen et al. (2008) showed that the error of the estimator was a mixed normal Gaussian converging at a rate of n1/6when H=cn2/3. Moreover, they advocate that an estimator with a kernel weight function which satisfies k′(0)2=k′(1)2= 0 have the fastest convergence rate until now at n1/4. After testing eight of this smoother-kernel based estimators, Barndorff-Nielsen et al. (2008) stated that the one with the lowest asymptotic variance was the Parzen kernel. For our tests we will use the following Parzen Kernel-based estimator (29) g RV P arzenKB n= H X h=−H kh H+ 1γh. The Parzen kernel function is given by
3. PRE-AVERAGING ESTIMATOR 29 k(x) = 1−6x2+ 6x30≤x≤1/2 2(1 −x)31/2≤x≤1 0x > 1. 2.1. Optimal Parameters. From Barndorff-Nielsen et al. (2008) we know that H=cξ4/5n3/5offers the best trade off between bias and asymptotic variance. Although true, this is only achieved if cmakes Hminimize the variance. BarndorffNielsen et al. (2009) set c∗to be considered (30) c∗= ((12)2/0.269)1/5 for the Parzen kernel. Moreover, following Barndorff-Nielsen et al. (2009), ξ2is now the integrated quarticity RT 0σ4 tdt and can be estimated as in Equation (26) without the correction, as follows: (31) ξ2=m 3T m X i=1 ˜r4 i, for m being a sample in lower frequency to avoid microstructure problem. This allows one to estimate the optimal H∗=c∗ξ4/5n3/5and make the estimator asymptotically consistent. 3. Pre-Averaging Estimator Until now we have described two of the main approaches for estimating the integrated variance, i.e. the two scale estimator from Zhang (2005) and the Kernel approach from Barndorff-Nielsen (2008). The first one is based on the linear combination of Realized volatilities obtained by sub-sampling while the Kernel estimator is composed from a weighted auto-covariances of the returns. Christensen, Kinnebrock, and Podolskij (2010) propose a pre-averaging technique in order to reduce the microstructure effects. The idea is not far from the previous methods we have described. It is supposed that observed prices Ytare as in Equation (11) and noise follows Assumption 3.1. In addition, observations are recorded at times ti=i∆nequally spaced. To build the estimator, they considered that by averaging knobserved log-returns ˜rt,i, one is closer to the latent process Xt, because with that they reduce the variance of the error by kn. This can be shown as follows
30 4. ESTIMATION METHODS 1 kn kn X i=1 Yi=1 kn kn X i=1 Xi+1 kn kn X i=1 ϵi; V ar(1 kn kn X i=1 ϵi) = 1 k2 n kn X i=1 V ar(ϵ) = 1 kn V ar(ϵ). (32) The estimator of the integrated variance is called Pre-Averaging estimator. Like sub-sampling and auto-covariance methods, the pre-averaging approach gives rise to optimal estimators when parameters are well implemented. As it will be seen later on, the pre-averaging estimator is a function of a parameter knwhich, when well implemented, it will lead to the lower bound for the variance. The estimator can be written in terms of parameter knas follows (33) g RV P A n=√∆n kn∆tψ2 [t/∆n]−kn+1 X i=0 (Yi)2−ψ1∆n 2kn∆2 tψ2 [t/∆n] X i=0 (˜ri)2 | {z } Bias correction due to noise with the weighted average for the returns, Yiare defined as Yi= kn−1 X j=1 gj kn˜ri+ji= 0, . . . , n −kn+ 1. Consideration 1. The Pre-averaging method differs form the ARV in the sense that it does not only average over log-price differences from non-overlaping intervals, but rather uses a moving window kn. knis a sequence of integers satisfying (34) kn=θ √∆n . Moreover, to define the constant θone will have to decide the moving window length. In our high frequency framework kngoes to ∞as n→ ∞. Consideration 2. As can be seen from Equation (33), a correction bias due to noise has to be considered. Although, it is worth mentioning that it plays no role
3. PRE-AVERAGING ESTIMATOR 31 in the central limit theorem when we want to know the asymptotic properties of the estimator. Consideration 3. As stated before, pre-average estimator is a weighted function for the moving window averages, not far from the kernels intuition. The estimator can be generalized by the use of a general weight function gon [0,1] and g: [0,1] → Rwith the following weight function gi=g(i/kn). This allows us to consider g(0) = g(1) = 0 and makes the estimator much easier with an straightforward construction. In addition the simplest weight function gis g(x) = min(x, 1−x), also known as triangular function, it will allow us to reach optimal parameters specified later. g(x) in this case is a symmetric function with a maximum when x= 1/2 and minimum when x= 0 and x= 1. It gives symmetric relevance to auto-correlations equally spaced. Given this weight g(x) function, the weighted returns average form Equation (33) can be defined as follows: (35) Yi=1 knkn−1 X j=kn/2 Yi+j− kn/2−1 X j=0 Yi+j. Yiequality can be proven when knis an even integer as Yi=g1 kn˜ri+1 +g2 kn˜ri+2 +g3 kn˜ri+3 +. . . +g(kn/2) −1 kn˜ri+kn 2−1+gkn/2 kn˜ri+kn 2+. . . +gkn−1 kn˜ri+kn−1=1 kn (Yi+1 −Yi) + 2 kn (Yi+2 −Yi+1)+ +3 kn (Yi+3 −Yi+2) + ···+(kn/2) −1 kn (Yi+kn 2−1−Yi+kn 2−2)+ +(kn/2) kn (Yi+kn 2−Yi+kn 2−1) + ···+1 kn (Yi+kn−1−Yi+kn−2) = =1 kn (−Yi−Yi+1 −Yi+2 −···−Yi+kn 2−1+Yi+kn 2+Yi+kn 2+1 +···+Yi+kn−1). (36)
32 4. ESTIMATION METHODS 3.1. Optimal Parameters. The Pre-averaging estimator depends on the bandwidth kn. Our weight function g(x) = min(x, 1−x) allows us to reach the optimal bandwidth that minimizes the asymptotic variance of the estimation error. For any other weight function than a triangular one, the decision for optimal parameter will be different from the ones we have considered (see Jacod and Mykland, 2015). Essentially, pre-averaging estimator is a kernel based estimator. One can take the pre-average estimator (33) and adapt it in terms of a weight function k(x) instead of g(x) as follows (37) g RV P A n=Kt(1 + O(p∆t)) −ψ1∆n 2θ2ψ2 n X i=1 (˜ri)2 where Kt= n−kn+1 X i=kn (˜ri)2 +X kn≤i≤n−kn+1,1≤j≤kn kj−1 kn(˜ri˜ri+j+ ˜ri˜ri−j) (38) where the weight function kis defined between [0,1] having k(0) = k(1) = 0 and also k’(0) = k’(1) = 0 as with our kernel based proposal in past sections. Furthermore, Equations (33) and (37) will have the same asymptotic distributions. It can be proven that our estimator g RV P A nreaches the convergence with the Integrated Volatility at a rate of n−1/4 (39) 1 n1/4d RV P A n−IVT. Since asymptotic distributions are the same as for the Kernel estimator, also one can take the same optimal values to reach the lowest bound of the variance. From Barndorff-Nielsen et al. (2010), the bandwidth parameter which will minimize the variance is kn=θn3/5with θ=c∗ξ4/5. Note that H∗and knare equivalent, namely H∗=kn=c∗ξ4/5n3/5.c∗is the same constant as for kernel estimator and ξ2can be estimated as in the previous sections.
3. ESTIMATORS IMPLEMENTATION 39 Fig. 4. K number of subgrids that minimizes the asymptotic variance of the error on every day in the sample. Fig. 5. TTS Estimation for optimal k with minimum variance criteria on every day of the sample. 3.2. Kernel-Based Estimator. In the following section we apply the same test performance for the Kernel based Estimator (29). Similarly with the Two Scale estimator, we first derive the optimal parameter, number of auto-covariances Hfor Kernel estimators, taken into account to be weighted for the kernel function. In this case the kernel function we will use is the Parzen kernel as can be seen in Chapter 4. The optimal estimated Hfor every day in the sample is shown in Figure 6
40 5. BITCOIN REALIZED VOLATILITY ANALYSIS Fig. 6. Number of realized auto-covariance Hwhich minimizes the asymptotic variance of the estimation error on every day of the sample. The estimations for the Integrated Volatility using asymptotic optimal parameter Hfrom Equations (30) and (31) for Kernel Estimator are shown in Figure 7 Fig. 7. KB Integrated Volatility estimation using minimum variance criteria on every day of the sample. Similarly to the previous estimator, the Kernel Estimator also produces a reliable estimations for the IV, as can be seen in Figure 7. However, comparisons cannot
3. ESTIMATORS IMPLEMENTATION 41 be made up to this point just based on the plots. Later on we will provide for a quantitative comparison method. 3.3. Pre-Averaging Estimator. This is our last proposed estimator. The PreAveraging estimator in Equation (33) is defined in terms of parameter kn. In this section, and similar to what has been done with previous estimator parameters, we will compute the optimal parameter that minimizes the asymptotic variance in the CLT according to Equation (32). In Figure 8 we can see the different bandwidth estimations for parameter knand for each trading day. Fig. 8. Optimal bandwidth parameter knfor every Trading day. It has to be a even integer due to the estimator construction. The integrated volatility estimation with asymptotic optimal parameter knis shown in Figure 9.
42 5. BITCOIN REALIZED VOLATILITY ANALYSIS Fig. 9. PA Integrated Volatility estimation using minimum variance criteria on every trading day. 4. Comparison It is very common to compare estimators performance using the Mean Squared Error (MSE) method. This provides for a measure of how accurate are the estimations with respect to the true underlying process. Moreover, MSE will also allow us to compare between models since it calculates the estimation error, a quantitative comparable measure. To apply MSE one has to know the actual process or value that aims to estimate to be able to make comparisons on performances. This methodology is also known as unfeasible since one has to know the actual values, however, it applies since we know the actual IV from our RVnnoise-free estimation. Therefore, we can use this estimation as a proxy of the true IV to compute the MSE of the different estimates built on the noise contaminated data set. The MSE for the Integrated Volatility, given estimators, is: (41) MSE =Eh(c IV n−IVT)2i. Up to now, and from the Figures 5, 7 and 9, we have only been able to reproduce every estimation, but not to perform any comparison. Since we have introduced the quantitative MSE, we can calculate the error every estimation has with the parameters found in the limit. Starting from Equation (41), we calculate the MSE for every estimator given their optimal bandwidth parameter in previous sections. The results are provided in Table 3 and indicate the PA as the most efficient estimator in terms of MSE.
4. COMPARISON 43 MSETTS∗MSEP arzen,KB∗MSEP A∗ 4.166906e-06 5.391602e-06 3.06716e-06 Table 3. MSE calculated over optimal parameters K∗,H∗and k∗ n, for every estimator, respectively. In bold the estimation with smaller error. Since we are working with a finite sample from a high frequency data, the optimal parameters we are able to estimate minimizing the asymptotic variance could turn out to be sub-optimal ones. In our framework we don’t achieve the CLT rate, namely our n→ ∞, therefore, one cannot rely just on these asymptotic results (see. Bandi and Russell (2003)). We will base the choice of the bandwidth parameters for the different estimators on the minimization of the MSE over the whole sample. In the following sections we give minimum MSE parameter estimations for each estimator. 4.1. MSE for the Two Time Scale Estimator. Calculus for the Two Time Scale estimation error is as follows (42) MSET T S =Eh(g RV T T S n−IVT)2i KMSET T S 2 3.083455e-05 3 1.204682e-05 4 6.791061e-06 5 4.777700e-06 6 4.015265e-06 7 3.647966e-06 8 3.512836e-06 9 3.477572e-06 10 3.721111e-06 12 4.310056e-06 15 5.308646e-06 18 6.002404e-06 Table 4. MSE for the Two Time Scale estimator over different K. In bold the estimation with less error. In Table 4 MSE for g RV T T S nis evaluated at a different constant parameter K. The best trade-off between them both is K= 9 constant for every trading day. It can also be seen that, in one hand, the error for lower K, i.e high sub-sampling
44 5. BITCOIN REALIZED VOLATILITY ANALYSIS frequency, is due to the fact that the estimator is not able to get rid of all the bias, and on the other hand, the error for high K, i.e lower sub-sampling frequency, is due to the variance addition due to discretization. 4.2. MSE for the Kernel Based Estimator. Calculus for the Parzen Kernel Estimator error is as follows (43) MSEP arzen,KB =Eh(g RV P arzen,KB n−IVT)2i HMSEP arzen,KB 2 1.144340e-04 3 4.773067e-05 4 2.373312e-05 5 1.385073e-05 10 5.424316e-06 11 5.370415e-06 12 5.411094e-06 15 5.806344e-06 18 6.339024e-06 20 6.677436e-06 25 7.413086e-06 30 8.098493e-06 35 8.835234e-06 Table 5. MSE for the Parzen Kernel based estimator over different H realized autocovariances. In bold the estimation with less error. The performance of the Parzen Kernel estimator through different His captured in Table 5. The MSE is minimized when H= 11 and very close to the MSE of the variance-based estimator in Table 3. Moreover, we notice the kernel estimator is less effective than the TTS in terms of MSE.
4. COMPARISON 45 4.3. MSE for the Pre-Averging Estimator. Calculus for the Parzen Kernel Estimator error is as follows (44) MSEP A =Eh(g RV P A n−IVT)2i knMSEP A 2 1.311846e-04 4 7.050681e-06 6 2.256767e-06 8 2.711014e-06 10 3.462410e-06 12 4.125213e-06 14 4.729814e-06 16 5.306066e-06 18 5.850685e-06 20 6.344426e-06 22 6.770901e-06 24 7.152013e-06 26 7.507414e-06 28 7.845780e-06 Table 6. MSE for the pre-average estimator over different kn bandwidth. In bold the estimation with less error. Similarly with the previous estimators, MSE is computed for Pre-averaging estimator over different knin Table 6. As can be seen the bandwidth that minimize the MSE for the estimator is kn= 6. Furthermore, we can see when comparing all errors in Tables 4, 5 and 6 that Pre-Averaging is the one providing the best estimations.
Chapter 6 Conclusions In this work we have applied and tested three different methods to quantify and correct the effect of noise when estimating the Integrated Volatility of the returns from high frequency data. We have argued that due to the presence of noise, the Realized Volatility from observed prices Yt, in a fixed period of time, does no longer converge to the Integrated Volatility. We have used sub-sampling, realized auto-covariance and averaging techniques to construct our estimations. The first method based on sub-sampling is g RV TTS n, the Zhang et al. (2005) estimator build non-overlaping sub-grids to filter out the noise still using all the data available. The motivation behind building this type of estimator was to try to use all high frequency data available without throwing away any as in sparse sampling. Regarding the Kernel estimators, they are built on realized auto-covariances. We have seen that there are many proposals for kernel estimators, but following BarndorffNielsen et al. (2008, 2009) we have used the Parzen Kernel estimator g RV P arzen,KB n. This type of estimators use a weighted averaging based on auto-covariances to filter out the noise. Finally, Pre-Averaging estimator, g RV P A nfrom Podolskij et al. (2009) has been proposed. This estimator uses the averaging idea as the previous ones, but instead of creating sub-grids or taking into account auto-covariances, it is based on local weighted averages of returns. Also it has to be decided which kernel function to use. We have tested these estimators on a sample of 1-minute Bitcoin data form February 7th, 2022 to April 27th, 2022. From the results we are able to say that every estimator gets rid of the bias due to the microstructure noise successfully. The choice of the bandwidth parameters for the defined estimators can be based both on minimization of the asymptotic variance of the error in the CLT and on the minimum of the MSE. Through our analysis, it turns out that in both cases the one with less error is the Pre-averaging estimator. These evidences leads us to conclude that the Pre-averaging estimator from Podolskij et al. (2008) is the one that showed the best performance in terms of getting 47
48 6. CONCLUSIONS rid of the microstructure noise when estimating the continuous Integrated Volatility. Even though the Pre-averaging method showed the best results, the Two Time Scales estimator has an estimation error close to the best one. Further on the results, all the estimators showed a fair IV estimations, even the kernel Based one that has the greatest MSE of all them.