scieee AI-readable full text Open interactive document viewer

The adoption of Big Data by National Statistics Offices - an exploratory research

Cardoso, Fabio dos Santos

Abstract

Official Statistics play a relevant role in modern societies. Their indices, rates, values, and time series are used to support decisions at the individual and collective levels. Behind this information are the data generators that provide Official Statistics to society. These organizations are called National Statistical Offices (NSOs). As part of this public function, NSOs work through processes of collecting, storing, and analysing data from society to update Official Statistics. For more than two centuries, NSOs have performed the function of generating official statistics in a linear fashion, reporting results, and informing society. However, in the last two decades, the role of NSOs has been shaken by disruptive innovations related to digital transformation. One impact is the advent of Big Data. Innovative techniques and methodologies have enabled companies to manage large volumes and distinct types of data in a timely and permanent manner. This change has affected NSOs because the risk of obsolescence has impacted not only internal processes but also the response time required by society. In this challenging context, NSOs are faced with a dilemma: either adopting Big Data technologies and methodologies to maintain their social relevance or face fatal obsolescence. The current research therefore intends to study how and why are NSOs adopting or avoiding Big Data technologies and methodologies. To understand the adoption of Big Data (ABD), three NSOs in different realities were investigated. In South America, the selected case was the Brazilian NSO, IBGE. In Europe, Portugal's INE and the UK's ONS were chosen. These NSOs represent different geographical contexts, occupy a different the position in Global Statistical System and present different degrees of autonomy relative to the Central National Government. The Technology-Organization-Environment (TOE) Framework was used as the analytical model. A rigorous and broad review of the literature allowed us to extract and consolidate a set forty-three factors previously used to study the adoption of Big Data. Data collection methods included a documentary collection and application of twenty semi-structured interviews with members of the three NSOs: ONS, INE, and IBGE. Data analysis led us to identify ten additional new factors related to ABD in NSOs, including six technological, three organisational and one environmental. During the cases comparison, findings showed that the adoption of Big Data technologies and methodologies have similarities and common characteristics in the Official Statistics sector, despite each institution following a distinct trajectory. Based on this, we propose a matrix that depicts different positions in the evolution from avoidance to adoption of Big Data by NSOs. Moreover, ABD in NSOs advances on a cumulative trajectory where associations of specific factors define the path of adoption or avoidance. In conclusion, NSOs must not see ABD as a matter of choice, but as inevitable if they are to survive.

Full text

Fabio dos Santos Cardoso THE ADOPTION OF BIG DATA BY NATIONAL STATISTICS OFFICES – AN EXPLORATORY RESEARCH July 2023 THE ADOPTION OF BIG DATA BY NATIONAL STATISTICS OFFICES – AN EXPLORATORY RESEARCH Fabio dos Santos Cardoso UMinho|2023 Universidade do Minho Escola de Economia e Gestão January 2024 2024 Fabio dos Santos Cardoso THE ADOPTION OF BIG DATA BY NATIONAL STATISTICS OFFICES – AN EXPLORATORY RESEARCH July 2023 Doctoral Thesis Ph.D. Thesis in Business Administration Work developed under supervision of Professor Ana Cristina Almeida Carvalho University of Minho - Portugal Professor João Eduardo Quintela Alves Sousa Varajão University of Minho - Portugal Professor Dimária Silva e Meirelles Presbyterian University Mackenzie - Brazil Universidade do Minho Escola de Economia e Gestão January 2024 ii DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/ iii Acknowledgements My first acknowledgement is for God to the permanent presence in my whole life, including in the darkest moments of this PhD journey. In sequence, I offer my acknowledgment to my wife, Analice Araújo, for her permanent and incontestable support to me during those years. She had to abnegate from some projects, and besides those personal renouncing, she stayed though by my side. Analice this achievement is ours. My mother Claudete Santos, my daughter Ana Clara Cardoso, my brother Luíz Cardoso and my sister Simone Rodrigues, thank you so much for your supports. During the journey, in each time one of you stayed by my side, bringing solutions or motivating me to go beyond. Still in the family, I offer my acknowledgment to my aunts Zenaide and Zélia, my cousins Iris, Tânia, Sérgio and Júlio. When it was necessary, they were there and did not disappoint me. Regarding the closer friends, my acknowledgment has been offered to Ricardo Oliveira, Bartira Rignel, and Robson Alves. After decades of friendship, I am happy and proud to keep calling all of you as my brothers and sister. A great acknowledgment to my supervisors Ana Carvalho, João Varajão, and Dimária Meirelles. This marathon is in the final stage because I received more than supervision from you. Actually, you prepared an academic professional. Thank you. My friends from University of Minho, you were permanent fellows during this stony trip. Bassem, Núbio, Talita, João, Mahamoud, Heshem, Loaloa, Yasmin, Mona, Nadine, Nada, Tilmar, Maria, Luís, Samuel, Sandrine, Alene, Enrikson, Dilson, Ziad, Tarek, Zakarias, Hussein, Maher, Samer, Kevian, thank you so much for your company. Still about friendship, my acknowledgment to Rafael Lopes, Solange Oliveira, and Iara Baldin. My grateful to the IBGE’s employees who supported me with this research during several times. It is impossible to cite all of them. So, I say my thanks to all of you, specially to Marcelo Holanda, Georgia Assumpção, Ana Cristina von Calabach, Luíz Arbex, Bianca Walsh, Paulo Tostes, Hugo Baptista, Fabiane Lucena, Pamela Calazans, Aline Bezerra, Eduardo Derbli, Flávia Pinto, and Diana Baptista. Throughout the development of this research, I have received financial support from the Instituto Brasileiro de Geografia e Estatística (IBGE), who sponsored partially this study. iv STATEMENT OF INTEGRITY I hereby declare having conducted my Thesis with integrity. I confirm that I have not used plagiarism or any form of falsification of results in the process of the Thesis elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. v THE ADOPTION OF BIG DATA BY NATIONAL STATISTICS OFFICES – AN EXPLORATORY RESEARCH Abstract Official Statistics play a relevant role in modern societies. Their indices, rates, values, and time series are used to support decisions at the individual and collective levels. Behind this information are the data generators that provide Official Statistics to society. These organizations are called National Statistical Offices (NSOs). As part of this public function, NSOs work through processes of collecting, storing, and analysing data from society to update Official Statistics. For more than two centuries, NSOs have performed the function of generating official statistics in a linear fashion, reporting results, and informing society. However, in the last two decades, the role of NSOs has been shaken by disruptive innovations related to digital transformation. One impact is the advent of Big Data. Innovative techniques and methodologies have enabled companies to manage large volumes and distinct types of data in a timely and permanent manner. This change has affected NSOs because the risk of obsolescence has impacted not only internal processes but also the response time required by society. In this challenging context, NSOs are faced with a dilemma: either adopting Big Data technologies and methodologies to maintain their social relevance or face fatal obsolescence. The current research therefore intends to study how and why are NSOs adopting or avoiding Big Data technologies and methodologies. To understand the adoption of Big Data (ABD), three NSOs in different realities were investigated. In South America, the selected case was the Brazilian NSO, IBGE. In Europe, Portugal's INE and the UK's ONS were chosen. These NSOs represent different geographical contexts, occupy a different the position in Global Statistical System and present different degrees of autonomy relative to the Central National Government. The Technology-Organization-Environment (TOE) Framework was used as the analytical model. A rigorous and broad review of the literature allowed us to extract and consolidate a set forty-three factors previously used to study the adoption of Big Data. Data collection methods included a documentary collection and application of twenty semi-structured interviews with members of the three NSOs: ONS, INE, and IBGE. Data analysis led us to identify ten additional new factors related to ABD in NSOs, including six technological, three organisational and one environmental. During the cases comparison, findings showed that the adoption of Big Data technologies and methodologies have similarities and common characteristics in the Official Statistics sector, despite each institution following a distinct trajectory. Based on this, we propose a matrix that depicts different positions in the evolution from avoidance to adoption of Big Data by NSOs. Moreover, ABD in NSOs advances on a cumulative trajectory where associations of specific factors define the path of adoption or avoidance. In conclusion, NSOs must not see ABD as a matter of choice, but as inevitable if they are to survive. Key-words: adoption, Big Data, National Statistics Offices, Official Statistics, TOE framework. vi A ADOÇÃO DE BIG DATA POR INSTITUTOS NACIONAIS DE ESTATÍSTICAS – UM ESTUDO EXPLORATÓRIO Resumo As Estatísticas Oficiais desempenham um papel importante nas sociedades modernas. Seus índices, taxas, valores e séries temporais são usados para apoiar decisões em nível individual e coletivo. Por trás dessas informações estão os geradores de dados que fornecem estatísticas oficiais à sociedade. Essas organizações são chamadas de Institutos Oficiais de Estatísticas (INEs). Como parte dessa função pública, os INEs trabalham por meio de processos de coleta, armazenamento e análise de dados da sociedade para atualizar as estatísticas oficiais. Por mais de dois séculos, os INEs desempenharam a função de gerar estatísticas oficiais de forma linear, relatando os resultados e informando a sociedade. Entretanto, nas últimas duas décadas, o papel das INEs foi abalado por inovações disruptivas relacionadas à transformação digital. Um dos impactos é o advento do Big Data. Técnicas e metodologias inovadoras permitiram que as organizações gerenciassem grandes volumes e tipos distintos de dados de maneira oportuna e permanente. Essa mudança afetou os INEs porque o risco de obsolescência atingiu não apenas os processos internos, mas também o tempo de resposta exigido pela sociedade. Nesse contexto desafiador, os INEs se deparam com um dilema: adotar tecnologias e metodologias de Big Data para manter sua relevância social ou enfrentar a obsolescência fatal. Esta investigação pretende então estudar como e porque é que os INEs adotam ou evitam tecnologias e metodologias de Big Data. Para compreender a adoção de Big Data (ABD), foram estudados três INEs em diferentes realidades. Na América do Sul, o caso selecionado foi o IBGE, o INE brasileiro. Na Europa, foram escolhidos o INE de Portugal e o ONS do Reino Unido. Estes INEs representam diferentes contextos geográficos, diferentes posições no Sistema Estatístico Global e apresentam diferentes graus de autonomia em relação ao governo nacional central. O modelo Tecnologia-Organização-Ambiente (TOE) foi definido como modelo analítico. Uma revisão rigorosa e ampla da literatura permitu extrair e consolidar um conjunto de quarenta e três fatores usados anteriormente para estudar a ABD. Os métodos de recolha de dados incluíram uma coleta documental e a aplicação de vinte entrevistas semiestruturadas com membros dos três INEs: ONS, INE e IBGE. A análise dos dados levou-nos a identificar dez novos fatores relacionados com a ABD nos INEs, incluindo seis tecnológicos, três organizacionais e um ambiental. Durante a comparação dos casos, os resultados mostraram que a adoção de tecnologias e metodologias de Big Data possuem semelhanças e características comuns no setor de Estatísticas Oficiais, apesar de cada instituição seguir uma trajetória distinta nesta adoção. Com base nisso, propomos uma matriz que representa diferentes as posições na evolução dos INEs na adopção de Big Data desde o evitamento à adopção. Além disso, a ABD nos INEs progride em uma trajetória cumulativa em que associações de fatores específicos definem o caminho da adoção ou da evitamento. Em conclusão, os INEs não podem encarar a ABD como uma questão de escolha, mas sim como uma inevitabilidade para garantir a sobrevivência. Palavras-chaves: adoção, Big Data, Estatísticas Oficiais, Institutos Nacionais de Estatísticas, modelo TOE. ix Table of Contents Chapter 1 – Context and research roadmap ....................................................................... 1 1.1. Official Statistics – The history and features of a business base on knowledge and technology. 4 1.1.1. A summary of the Official Statistics History .................................................................. 5 1.1.2. Attributions and Uses of National Statistics Offices ....................................................... 7 1.1.3. The challenge of the adoption of Big Data .................................................................. 13 1.2. Thesis structure ................................................................................................................. 14 Chapter 2 – Adoption models for Big Data - a literature review ......................................... 16 2.1. Literature Review approach ....................................................................................................... 17 2.2. Adoption of Big Data – definitions and frameworks .................................................................... 22 2.2.1. Conceptualizing Big Data ................................................................................................... 23 2.2.2. Theoretical frameworks for the adoption of Big Data ........................................................... 36 2.3. The Technology-Organization-Environment model (TOE) ............................................................. 42 Chapter 3 – TOE as a framework for the adoption of Big Data .......................................... 47 3.1. Factors of the adoption of Big Data ............................................................................................ 49 3.1.1. Technology factors ............................................................................................................. 56 3.1.2. Organisation factors ........................................................................................................... 58 3.1.3. Environment factors ........................................................................................................... 60 Chapter 4 – Methodology: how the research has been built .............................................. 63 4.1. Addressing the Research problem and the research question ..................................................... 65 4.2. Why the Case studies method? .................................................................................................. 68 4.3. Designing the research .............................................................................................................. 70 4.3.1. Selecting the cases ............................................................................................................ 70 4.4. Data collection .......................................................................................................................... 73 4.4.1. Documental analysis .......................................................................................................... 73 4.4.2. Semi-structured Interviews ................................................................................................. 75 4.5. Data Analysis Techniques .......................................................................................................... 80 Chapter 5 – Understanding the National Statistics Offices ................................................ 85 5.1. Features and Summarised trajectory ......................................................................................... 86 5.2. Outputs and deliveries ............................................................................................................... 88 5.3. Innovation cycles ...................................................................................................................... 89 5.4. The Census operation ............................................................................................................... 91 5.4.1. The Census in the United Kingdom .................................................................................... 92 5.4.2. The Census in Portugal ...................................................................................................... 95 vii 2 Source: Google trends. Retrieved 2023, January 30th. Figure 1 – Searches in Google, from 2008-Jan to 2022-Dec, using the themes "Big Data", "Web Scraping", "IoT", "Artificial Intelligence", and "Cloud computing". Beyond the definitions, Big Data is recognised by Schwab (2016) as a disruptive technology with revolutionary capacity when applied in organisational contexts, especially in decision-making processes. This author lists the benefits obtained by the adoption of Big Data (ABD): (i) " faster and better decision-making process ", (ii) " decision-making process in real-time ", (ii) " open data to support and foster innovation ", (iv) " more efficient public services ", (v) " new job opportunities creation ", (vi) " new types of public and private services ." All of them, by his predictions, might result in more quality services deployed to citizens and customers, profound change in legal and commercial norms, and improvement of global integration in areas such as public services, as the Official Statistics. The impacts can positively improve citizens, companies, and States. Despite those possible opportunities, Big Data is also considered a disruptive technology. This disruptive feature may bring risks when adopted by governments and companies. Schwab (2016) draws attention to the lack of governmental accountability, loss of individual privacy, and individual profile featuring, "profiling", to use for civil manipulation by private groups, such as in the novel “ 1984” written by Orwell (2013). As a matter of fact, Harari (2017, 2019) also reinforces the risks associated with the use of Big Data, drawing attention to the civil control made by private groups. Exemplifying, the author cites two examples: the Brexit referendum and the United States presidential elections, both in 2016. Harari emphasised the risk that a "digital dictatorship" will guide humanity. In the same way, O’Neil (2016) goes beyond and calls the erroneous use of Big Data a tool to control citizens' life. 0 10 20 30 40 50 60 70 80 90 100 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 Big Data Web Scraping IoT Artificial Intelligence Cloud computing 3 This dystopia is approached in Amer and Noujaim's (2019) documentary "The Great Hack." The movie showed how Big Data had been used by Cambridge Analytica company, associated with Facebook, to influence electors in Argentina, Trinidad and Tobago, the United Kingdom, the USA, and Brazil. In the same film, the statement done by the Royal Statistical Society about data is confirmed, and the sentence is clear: today, data has more value than oil. Notwithstanding the internal risks embedded in Big Data technologies and methodologies, this technological-organisational asset calls the attention of companies, States, and academia. One motive for this interest is the association between Big Data and decision-making processes. Sheng et al.'s (2017) research shows the link between those two themes. In their article, the authors showed the link between decision-making and Big Data: "(…) also point out that data-driven decision-making is a promising trend and the primary factor attributed to successful Big Data utilisation is the organisational alignment in every aspect in the companies. (…) From the information management perspective, Big Data research primarily concerns data acquisition and process effectiveness. The availability and feasibility of information are critical to organisational success in terms of strategic decision-making. " (p. 102). Some social actors are responsible for the relevance of data. One of them is a group of public data generators named National Statistical Offices (NSOs). They are in charge of delivering Official Statistics to the respective societies. Regarding the relevance of the NSOs, this research focuses on their adoption of Big Data technologies and methodologies. As official data generators, the NSOs collect, store, and analyse data to deliver index prices, rates of inflation and unemployment, among others (Radermacher, 2020). In fact, NSOs are critically relevant in this period of History because they are public institutes, accountable, subject to laws, and transparent in their data production (United Nations, 2014). Different from private companies, the Official Statistics Offices, as a business sector, has to work to support transparency, accuracy, and social trust (United Nations, 2014). In this context, this research presents a strong motivation that can be summarised as the demand for innovation by NSOs in their processes to produce Official Statistics, considering the adoption of Big Data (ABD henceforth) solutions (technologies and methodologies). So, this research seeks to identify and analyse factors that inhibit or boost ABD. If the NSOs suffer obsolescence, this circumstance can represent risks to societies and decision-makers. The reason is that using outdated data can guide decision-makers and governments to fail. 4 Furthermore, studies about the ABD in NSOs are scarce, despite the relevance of those institutes. This gap provides an opportunity to address this research, contributing to the body of knowledge on ABD, management, and Official Statistics as a business sector. Indeed, this business sector presents a plethora of specifications, and these details and elements are explained in the next section. 1.1. Official Statistics – The history and features of a business base on knowledge and technology. It is necessary to start by defining and differentiating the terms related to Official Statistics. Radermacher (2020) approaches Official Statistics as a business sector in his research. This author brings definitions for three main concepts to the Official Statistics field: “(i) statistics”, “(ii) statistical results”, and “(iii) statistical institutions.” According to this author, statistics is “the science of learning from data.” This science produces statistical results that “are used for all conceivable information and decision-making processes.” In the same path, Radermacher (2020, p.2) remarks that “the producers of statistics” are statistical institutions. The production and application of statistics can be routinely used by a commercial company, a non-governmental organisation and a soccer team. When the statistics are generated and broadcast by a governmental authority, it becomes official information. This official information refers to Official Statistics. Official Statistics means “any set of statistics produced by an organisation named under secondary legislation and described by that organisation as an official statistic or part of a set of official statistics" (UK Statistics Authority, 2022). Statistical statistics offices are the producers of Official Statistics. This type of statistics can be differentiated by mainly two general groups, according to the UK Statistics Authority (2022): (a) “National Statistics, which have been assessed by the statistics office as fully compliant with the regulation ; ” (b) “experimental statistics, which are newly developed or innovative statistics . ” Related to the mentioned regulation, it refers to a set of rules to guide the planning, collection, analysis, and dissemination of the data processed by the Statistics Offices. Those rules can be laws, acts, or administrative codes to ensure that the Official Statistics have been produced with ethics and high standards. Statistical Offices are, therefore, public organisations working at three different geographic levels: local, national and international. In this research, we consider the organisations which work at the 5 national and international levels. This distinction is relevant because of the comparative approach related to the case study method adopted in this study. Official Statistics and their operators are not a 21st Century phenomenon. All those statistics offices, their outputs and outcomes, the methodologies and techniques result from a long historical trajectory. This trajectory of knowledge adjudication and constant innovation is described in the following subsection. 1.1.1. A summary of the Official Statistics History Kotz (2005) and Stigler (1986) highlight the evolution of statistics as a science and Official Statistics as a subfield of this science, both part of the European Modern History. However, Kotz and Stigler diverge about the timeline. Kotz (2005) has adopted a longer trajectory since the 16th Century. The use of counts and numbers to measure military capabilities and taxation dates remotes from Antiquity (Kotz 2005), with registers in the Mediterranean zone. In the 16th Century, seminal studies about Official Statistics started to be published in three European countries: Germany, the United Kingdom, and France. As remarked by Kotz (2005), during the 17th Century, in Germany, the publication from Johann Peter Süssmilch (1707-1767) represented the initial event for the Official Statistics in that European Zone. As a Protestant Pastor, Süssmilch recorded religious events, such as marriages, baptisms, and funerals. Indeed, these registers generated the initial tables and data about immigration, emigration, nationality, age of marriage, and birth rates. Many years later, in the 19th Century, the Royal Prussian Statistical Bureau adopted Süssmilch’s methods in its processes to produce Official Statistics. The United Kingdom is another country where statistics and Official Statistics methods arose early. Kotz (2005) remarks that the adoption of the word “statistics” is an appropriation of the similar German word “statistik.” John Gaunt (1620-1674) was one of the first involved in social and demographic measurements in the United Kingdom (UK). As a London citizen, he published a book on the city's mortality metrics based on official registers. Another relevant British author for statistics and Official Statistics was John Sinclair (1754-1835). Sinclair’s twenty-one volumes about statistical accounts in Scotland adopted an innovative method of data collection: the questionnaire. The compilation from questionaries applied in 938 Scottish Parishes and responded to by the respective church ministers generated the data explored by 6 Sinclair (Kotz (2005). Thus, Sinclair’s effort represents a social statistical survey developed over two hundred years ago! Still in the UK, during the first half of the 19th Century, the Royal Statistical Society was founded in London (Royal Statistical Society, 2022). With 188 years of history, this institution is the active think tank for Official Statistics in the British islands. The third country is France, as remarked by Kotz (2005). The oldest register of Official Statistics in France is a public financial publication from 1581. Surveys followed this historical register to collect administrative data in the second half of the 16th Century. In fact, the French focus on Official Statistics crossed the centuries and kept steady progress until the French Revolution. Chaptal, Neufchateau, Laplace and Duvillare are some names of intellectuals who contributed to developing French Official Statistics in the 18th and 19th Centuries. In the same period, specifically in 1843, a Belgian scientist, Adolphe Quetelet, presented his research to the French academia. His studies concentrated on demographic themes. The target of his studies was addressing three questions: (i) “what are the laws of human reproduction, growth, and physical force” that influence a country or society? (ii) “what influence has nature over man?”, (iii) “can human forces compromise the stability of the social system?” (Quetelet, 2013). Attempting to answer those questions, Quetelet developed a theoretical construct still used nowadays in Official Statistics: the average man. The proposal was to achieve the possibility of understanding a great set of data, such as a nation’s society. So, Quetelet understood citizens as units of this society. In this way, he tried to reduce the extremes in matters as behaviours, physical features and natural actions. The reduction proposed by him adopted a statistic tool: normalisation of the society features in a average, the average man. This concept was created to concentrate in a symbolic-model-citizen, predominant elements from each society. He pointed out that his intention was not to create a flat view of mankind. France should have its own average man, such as England, Germany and Scotland should have their own (Quetelet, 2013). Until the 21st Century, Quetelet’s (2013) legacy still holds. In 2010, the British Office for National Statistics did a research to describe the average man and the average woman in England (BBC, 2019). The relevance of Official Statistics is also apparent in other countries, including Portugal and Brazil, which feature in the current study along with the UK. In 1864, Portugal ran its first demographic Census, adopting modern statistical methodologies (Instituto Nacional de Estatística, 2022). Before 7 this year, other governmental accountings with military proposals took place in the 13th, 15th, 16th, 17th, and 18th centuries. Since the end of the Middle Ages, national accountings were present in the routines of the Portuguese government (Instituto Nacional de Estatística, 2021a). In the case of Brazil, the reference year is 1870, when the “Diretoria Geral de Estatísticas” (General Statistics Directory) was founded. This organisation was the first entity to centralise the Official Statistics in this South American country. Two years later, in 1872, that organisation planned, organised and ran the first Brazilian demographic Census (Diretoria Geral de Estatística, 1874). Considering the extensive history of Official Statistics, a reflection emerges in this context: in the current days, the Official Statistics shall be up to date with the current data technology applications. However, this is a huge challenge to be addressed by the Official Statistics operators in the execution of their attributions. Those operators are named in this study as National Statistics Offices (NSOs), considering the geographic approach: national and international levels. 1.1.2. Attributions and Uses of National Statistics Offices According to Radermacher (2020), Official Statistics are essential to societies because they can be used to support public policies, business and investment decisions. Official Statistics also offer citizens the capacity to scrutinise the performance of governments. In this way, NSOs play a central role in the business of Official Statistics. Radermacher (2020) highlights those offices as keepers of the uniform production of Official Statistics, following effective and efficient patterns. As shown in Table 1, this perspective can be observed in the mission statements of several National Statistics Offices (NSOs), that claim a clear commitment with society and its players. The selection of those three examples is defined by the choice of the three case studies to be addressed in this research: IBGE, from Brazil; INE, the Portuguese NSO; and ONS, from United Kingdom. Chapter 5 describes each case in detail. Table 1 – Examples of National Statistics Offices' Mission Statements “High quality data and analysis to inform the UK, improve lives and build the future.” ONS - Office for National Statistics “The Mission of Statistics Portugal is to produce, in an independent manner, high-quality official statistical information, relevant for the society, while promoting the coordination, 8 the analysis, the innovation and the dissemination of the national statistical activity and ensuring integrated data storage.” INE - Instituto Nacional de Estatística “To portray Brazil by providing the information required to the understanding of its reality and the exercise of citizenship.” IBGE - Instituto Brasileiro de Geografia e Estatísticas It is possible to verify the relevance of some concepts presented in the mission statements using a tag cloud chart (Figure 2). The relevance of the concepts: “information”, “quality” and “analysis” indicates a common basis for these NSOs to define their missions. Those concepts can be translated to values and targets to be achieved by those organisations. In this way, it is necessary to explain that the NSOs follow a kind of common “constitutional bill”, with the principles and guidelines to define the Official Statistics procedures. Legend: Tag cloud for the NSOs strategic mission statements. The prevalence of three words (information, quality, and analysis) shows the commitment of those organisations with the duty to deliver Official Statistics to the society. Figure 2 – Relevant terms mentioned in the NSOs strategic mission statements In fact, this “constitutional bill” works as a list of limits and protocols that NSOs must observe. Everything out of this frame is considered misconduct by NSOs’ peers. The United Nations defines this document as the “Fundamental Principles of Official Statistics” (United Nations, 2014). Ten principles guide the operations of Official Statistics by NSOs. Table 2 shows the principles as they are presented by United Nations Statistics Division (United Nations, 2014). 9 Table 2 - Fundamental Principles of Official Statistics Principle 1 Official statistics provide an indispensable element in the information system of a democratic society, serving the Government, the economy, and the public with data about the economic, demographic, social and environmental situation. To this end, official statistics that meet the test of practical utility are to be compiled and made available on an impartial basis by official statistical agencies to honour citizens’ entitlement to public information. Principle 2 To retain trust in official statistics, the statistical agencies need to decide according to strictly professional considerations, including scientific principles and professional ethics, on the methods and procedures for the collection, processing, storage, and presentation of statistical data. Principle 3 To facilitate a correct interpretation of the data, the statistical agencies are to present information according to scientific standards on the sources, methods and procedures of the statistics. Principle 4 The statistical agencies are entitled to comment on erroneous interpretation and misuse of statistics. Principle 5 Data for statistical purposes may be drawn from all types of sources, be they statistical surveys or administrative records. Statistical agencies are to choose the source with regard to quality, timeliness, costs and the burden on respondents. Principle 6 Individual data collected by statistical agencies for statistical compilation, whether they refer to natural or legal persons, are to be strictly confidential and used exclusively for statistical purposes. Principle 7 The laws, regulations and measures under which the statistical systems operate are to be made public. Principle 8 Coordination among statistical agencies within countries is essential to achieve consistency and efficiency in the statistical system. Principle 9 The use by statistical agencies in each country of international concepts, classifications and methods promotes the consistency and efficiency of statistical systems at all official levels. Principle 10 Bilateral and multilateral cooperation in statistics contributes to the improvement of systems of official statistics in all countries. Source: United Nations. Retrieved from “https://unstats.un.org/unsd/dnss/gp/FP-Rev2013-E.pdf”. Accessed in 01/06/2022. These ten principles cover ten specific themes about de execution of Official Statistics. The first principle announces the relevance and the use of Official Statistics. Next, the second principle approaches the preservation of trust from Official Statistics based on scientific methods and professional ethics. Data interpretation and reinterpretation are the themes of principles three and four. Principle five covers the data sources used, including surveys and administrative records. 10 Individual privacy, regulations and laws are principles six and seven themes. Regarding cooperation at the international level, principles eight, nine and ten not only propose but also encourage partnerships among NSOs. It is necessary to highlight that those principles result from a consensus between United Nations Statistics Division (UNSD) and the NSOs. This set of organisations integrates the Global Statistical Division. The exchange of information and methodologies among NSOs, and from NSOs to USND characterises the Global Statistical System, a global organisational network, as remarked by Cardoso et al. (2021, p.3): “As nodes of the NSOs network, these organisations have similarities regarding the organisational mission and operations. They respond to their countries’ respective governments as well as the United Nations Statistics Division - UNSD (…). The joint operation of NSOs within UNSD forms the Global Statistical System (GSS). In other words, the GSS natural configuration is a global network organisation (…). On top of that worldwide structure, some NSOs and UNSD play central roles in operational standardisation and technology innovation processes. In contrast, others GSS members act under a follower perspective. The GSS represents a Core-Periphery network standard (…) Its members keep working as a network (…) related to Big Data Adoption. “ In order to guide and advance the use of Official Statistics, those principles work as a starting point for NSOs to develop their research and targets. Indeed, the principles are not a self-centred set of rules. Radermacher (2020) remarks that Official Statistics are also being used to inform about the United Nations’ Millennium Development Goals (MDGs). In this case, surveys and administrative data support measuring the countries’ progress towards achieving the goals. So, it is possible to define Official Statistics as a type of business (Radermacher, 2020). As a business, it has its production cycle; in this case, a data production cycle. There is also a business architecture to support the Official Statistics Business (Radermacher, 2018). Both are structured and aligned with the Fundamental Principles of Official Statistics (Radermacher, 2020). Figure 3 presents the data generation life cycle of Official Statistics. It is possible to observe statistics users as the origin of the demands for Official Statistics outcomes. Development, production, and communication are stages in the Official Statistics life cycle. 11 Legend: There are two groups involved during the life cycle of the Official Statistics: producers and users. The first group develops the work system, manage the statistical data and its metadata, generate predictions, and communicate the outcomes by the statistical information system. Users share information needs and their data. Source: Based on Radermacher (2020). Figure 3 – Life cycle of the Official Statistics It is relevant to remark that the second green dot represents the raw material of Official Statistics, data and metadata. The statistical information can be disseminated only after the processing and analysis of those elements. So, the Official Statistical life cycle, together with the Ten Fundamental Principles, allows the development of the business architecture followed by the National Statistics Offices (Radermacher, 2018). In Figure 4, this extensive scheme shows the value chain of the Official Statistics in the rectangle named “Production”. Surveys and metadata are at the core of this value chain. The methodology and the technology are highlighted in the domain “Develop”. Training, legal, organisation, and ICT in all perspectives (capabilities, complexities, readiness, infrastructure) are among the elements that constitute the support line for NSOs to run official statistics businesses. Communication is the final stage of this architecture. It is understood to cover a wide range of dissemination channels, from technological capacity to offer user access (data warehouses and microdata access), to publishing (publications and press documents). 18 There are different ways to structure the LR. Ferfolja and Burnett (2002, p. 3) list four types of LR organization: (i) chronological – research sources organized by historical context; (ii) “classic” studies – the focus of the LR is on the prime publications and authors from a specific topic; (iii) topical or thematic – the LR is structured on conceptual categories, also named sections and subsections; and (iv) inverted pyramid – from a large thematic perspective to a well-focused issue, the LR advances towards the core issue of the research question. Subjects organized by thematic sections can be used as the foundation of an LR’s framework (Webster and Watson, 2002). Additionally, an LR can be structured in concept-centric and author-centric logics (Webster and Watson 2002). The first is the research that adopts contributions based on a central concept, such as “Adoption of Big Data”. An author-centric LR is organized in function of the specific publications of an author or a group of authors. Our LR adopted a concept-centric framework (Webster and Watson, 2002), organized in topical or thematic sections (Ferfolja and Burnett, 2002). The concept-centric framework enabled us to deploy a systematic LR on a specific subject, in this case, the adoption of Big Data (ABD). Concerning topical or thematic structures, Ferfolja and Burnett (1998) argue that they organize the discussion. As parts of those structures, sections and subsections develop a “research space” (Ferfolja and Burnett, 2002) to further results, including those from empirical studies, such as ABD technologies and methodologies by the National Statistics Offices. Our LR, therefore, operates with the use of key concepts. Webster and Watson (2002) consider that a robust LR must be based on a set of key concepts. Additionally, the authors reinforce the argument that concepts confer organization and boundaries to the LR. About the selected concepts for this research, they are presented in the research question: “ How and why do National Statistical Offices adopt or avoid Big Data technologies and methodologies? ” The primary concepts are National Statistics Offices, Adoption, and Big Data. Derived from these primary concepts are the terms: acceptance, and the names of each National Statistics Office (IBGE, INE, ONS) under study. Additionally, the database search used those terms, organized in sets of keywords or strings. The principal sources used are two databases: Scopus (SCP) and Web of Science (WOS). Despite the existence of other databases, SCP and WOS gathered the relevant journals related to business, management and innovation adoption. Harzing and Alakangas (2016) reinforce the reliability of SCP and WOS, and additionally highlight that both databases return more accurate outputs than 19 Google Scholar. The search included three types of published materials: journal articles, conference papers, and book chapters. Summaries, websites, and magazine articles were not considered in this LR. Ten relevant scientific fields were identified as relevant for this study: Computer Science; Engineering; Mathematics; Decision Sciences; Social Sciences; Business, Management, and Accounting; Multidisciplinary; Economics, Econometrics, and Finance. The keywords used for the searches were: 1. National Statistics Office (or NSO); 2. National Statistics Offices (or NSOs); 3. Big Data; 4. Adoption; 5. Acceptance; 6. Innovation. Those six keywords were used in different combinations (strings) during twenty-seven search rounds in SCP and WOS. The Boolean connector “and” linked the keywords to compose the strings. During the current search string process, English as the primary language was used. Portuguese, as the mother language from Portugal (INE) and Brazil (IBGE) was applied in this string search. However, the results in Portuguese were null. Table 3 summarizes the adopted set of strings and the search results. Table 3 - Results from the Strings Search Results from the Strings Search References obtained Strings No. of rounds Results “National Statistics Office” and “adoption” and “[country name]” 3 0 “National Statistics Office” and “acceptance” and “[country name]” 3 0 “National Statistics Office” and “innovation” and “[country name]” 3 0 “National Statistics Office” and “Big Data” and “[country name]” 3 10 “[NSO name]” and “adoption” 3 4 “[NSO name]” and “innovation” 3 5 “[NSO name]” and “Big Data” 3 1 “[NSO name]” and “acceptance” 3 0 “Big Data” and “adoption” 1 93 “Big Data” and “acceptance” 2 219 TOTAL 27 322 Considering the low number of results (only ten) in the initial rounds, other string searches were executed. In these additional searches, the reference to NSOs was dropped, and strings combining 20 “Big Data” with “adoption” or “acceptance” were used. This procedure followed the example of Ram et al. (2019, p. 566). Firstly, ninety-three (93) results from “Big Data” and “adoption” were found. Secondly, the string “Big Data” and “acceptance” generated three hundred forty-nine (349) references. After eliminating 130 duplicated or failed registers, 219 documents were considered for screening with the string “Big Data” and “acceptance.” Finally, three hundred and twenty-two (322) references were obtained after twenty-seven string searches, considering adoption (1 round), acceptance (2 rounds), and NSOs (24 rounds). Summarizing, the aggregated results (WOS+SCP) are presented also in Table 1. Figure 5 shows the steps taken in this process that followed the stages proposed by Santos et al. (2020): (i) identification, (ii) screening, (iii) eligibility, (iv) selection. Identification is the initial stage, described in the previous paragraph and Table 1. In the screening stage, the title, abstract, keywords, and graphic elements were read. The exclusion criteria followed these patterns: (a) studies about technology per se , without link to organizational elements (e.g., coding, mathematical model to algorithm development); (b) papers concerned about other issues (e.g. medicine, nursing, astronomy, etc.); (c) research in progress papers (RiPs) without academic findings, or just proposing an investigation agenda; (d) references with register-errors in the databases, or noaccess by the principal repository websites (ScienceDirect, Springer Link, JStor, EBSCO, etc.). Two hundred and fifteen (215) references were thus rejected, and the number of references was reduced from 322 to 110. In the eligibility stage, four studies were removed because the study design and the results were unclear. So, the total references reviewed was 106, which were read in full. Finally, at the selection stage, thirteen references were rejected because of low convergence with the research theme. The ten articles about NSOs remained in the selected studies. Our LR has a set of 93 references. 21 Legend: The four stages of the literature review process: (i) Identification, (ii) Screening; (iii) Eligibility; (iv) Selection. The initial 322 references were reduced to a final 93. Figure 5 – The systematic search process 22 Considering the results of the Literature Review, the following section proposes to establish standard concepts about the central terms of this research. So, the strings “Big Data”, “Adoption”, and “Acceptance” are discussed based on the 93 references. This set of references constitutes the basis for achieving four objectives: (i) develop consolidated definitions used in this thesis, such as Big Data and Big Data applied in Official Statistics; (ii) identify the theoretical frameworks used in the literature to study the adoption of Big Data; (iii) focus specifically on studies using the Technology-Organization-Environment framework (TOE), and (iv) devise a consolidated list of factors affecting ABD in the 3 dimensions of the TOE framework: technological, organizational, and environmental. 2.2. Adoption of Big Data – definitions and frameworks As remarked at the beginning of the current chapter, the initial discussion about Big Data positioned it as an asset within Information Technology (IT) with impacts to the organization. In this perspective, IT adoption constitutes a fundamental pilar on which to base the study related to ABD. IT adoption has been a relevant research topic in Information Systems (IS) and Management for the last 30 years (Iacovou et al., 1995; Chau and Tam, 1997; Karahanna et al. , 1999; Brown and Venkatesh, 2005; Ahuja and Thatcher, 2005; Komiak and Benbasat, 2006; Pavlou and Fygenson, 2006; Sarker and Valacich, 2010; Trigo et al. , 2013; Angst and Agarwal, 2016; Bayerl et al. , 2016; Duan et al. , 2017; Kim and Son, 2017; Retana et al. , 2018; Wunderlich and Veit, 2019; Bose and Leung, 2019; Jung et al. , 2019; Ho et al. , 2020). Most studies are concentrated in the last ten years (e.g., Al-Rahmi et al. 2019; Angwar 2018; Baig et al. 2019; Bilal et al. 2016, 2016, 2018; Bonner 2013; Cabrera-Sánchez and Villarejo-Ramos 2019; Cabrera-Sanchez and Villarejo-Ramos 2019; Félix et al. 2018; Kwon et al. 2014; Lai et al. 2018; Martinez-Mosquera and Luján-Mora 2019; Mcmahon et al. 2020; Memon et al. 2016; Park and Kim 2019; Ramadoss and Elango 2015; Shahbaz et al. 2020; Silva et al. 2019; Sun et al. 2018; Verma and Bhattacharyya 2017; Yadegaridehkordi et al. 2018). 23 To establish the state of the art related to ABD we therefore review models applied to IT adoption, and then more specifically to ABD. The TOE model demanded particular attention as it was followed in this research. So, we start by a discussion of the concepts and definitions of Big Data and ABD. 2.2.1. Conceptualizing Big Data Big Data is a concept with high diversity in its definition. It is possible to affirm this situation based on the several different definitions presented in academic articles (Ylijoki and Porras 2016; Chen et al. 2017; Brock and Khan 2017; Urbinati et al . 2019; Mcmahon et al. 2020; Bakici et al. 2022). Looking just at the articles included in this literature review (see section 2.1), we find many different conceptualizations of Big Data. Big Data is therefore a polysemic concept. Polysemy is the “ multiplicity of meanings ” (Ravin and Leacock, 2000) from different words which express a concept. Considering this polysemous feature of Big Data, the current study proposes to build a conceptualization of this term in an adjudicative (Cronin and George, 2020) sequence. A clear conceptualization of the term Big Data will improve the analytical capacity of this research, from the theoretical review to the interview analysis. In this way, conceptualization or concept formation is a complementary process of theory formation, as remarked by Kaplan (1964). In the same perspective, Welch et al. (2016) define concepts as elements of a theory or, in their own words, “building blocks”. Additionally, Podsakoff et al. (2016) stated that a concept is “ a cognitive symbol that has meaning” , adopted and use by researchers. In addition to their definition, Podsakoff et al. (2016) highlighted that concept formation is a core part of social science research. Based on the relevance to social research, Welch et al. (2016) assert that the concept has a dynamic nature. Those authors understand that concepts must be “ tested, challenged and revised” to verify their validity in a specific study. These procedures check if a concept present high intensity and low extension, or the opposite status. Looking at the plethora of Big Data definitions, there is an indication that this concept is very extensive and poor in intensity. Welch et al. (2016) also indicates that parts of a concept – its attributes – must be compared as part of the conceptual test. This comparison process is relevant to identify if the attribute presents more coverage or accuracy. In the following pages and tables, the concept of Big Data is analysed and rebuilt with a specific purpose: to develop an updated conceptualization to be used in the context of National Statistical Offices and Official Statistics. 24 There are distinct methodologies to test, challenge and revise a concept. Morakanyane et al. (2017) adopt a concept-centred matrix to evaluate a concept, checking its core fundaments present in the literature state of the art. Another method of conceptualization is proposed by Podsakoff et al (2016) with the use of the literature review. This procedure defines ten steps/questions to check if the concept is reliable and robust. Although the previous two methods use literature review as the background, there are other possibilities to conceptualize a research term. For example, Welch et al. (2016) evaluated a concept with the use of a research case. The case therefore works as the basis of concept building. In this doctoral investigation, Morakanyane et al. 's (2017) method is the selected one because the applied technique – a concept-centric matrix – is objective with a clear orientation of each step, and presents the stages in sequential tables. Indeed, those features are compatible with the structure of the LR adopted in this investigation (see section 2.1). The next paragraphs present the steps of building the concept of Big Data in the National Statistics Offices contexts following this method. Firstly, Morakanyane et al. 's (2017) method starts with a structured literature review. The references extracted from the literature review define the landscape and coverage of the analysed concept. In the case of Big Data, our systematic literature review confirmed a polysemic status of its definition. Considering the ninety-three references, our LR found that most of the articles, conference papers, and book chapters mention, cite but do not define the term “Big Data”. This is an indicator that the polysemy is present and acts as an inhibitor to the adjudicative research. Probably, many authors avoid conceptualizing the term Big Data considering the limited time and space available in a published manuscript. Also, the lack of consensus about the concept is an additional reason to avoid defining it. The next stage in the Morakanyane et al. 's (2017) method is selection. In the sample of ninetythree references, just twenty presented a clear, objective, and specified conceptualization. Nineteen references are extracted from journals and reviews in Management, Business, and Information Systems. Those publications are top ranked by the academic portal SCIMAGOJR (https://www.scimagojr.com/) with high scores Q1 and Q2. Additionally, an international peerreviewed conference paper was included in this set. In total, twenty references were applied to develop a Big Data adjudicative concept. 25 Table 4 shows the authors, year of publication, article title and journal. The term Big Data is part of the titles in nineteen references. Only one does not mention the term in the title; notwithstanding, it defines it. Another relevant find is that reputable reviews and journals are more receptive to publish manuscripts with Big Data conceptualization. This is not a subject frequent in conference papers and periodicals Q3 or Q4, as identified in the ninety-three references of this literature review. Additionally, the journal titles bring an indicator of the multidisciplinary related to Big Data. Nine of the nineteen journals have the term “management” in their titles, showing Big Data is a concept present not only in “information systems” studies but also in “management” research. Table 4 – References for developing the Big Data adjudicative concept Authors Title Journals and Conference Titles Gangwar, 2018 Understanding the Determinants of Big Data Adoption in India: An Analysis of the Manufacturing and Services Sectors Information Resources Management Journal Sun et al., 2018b Understanding the Factors Affecting the Organizational Adoption of Big Data Journal of Computer Information Systems Baig et al., 2019 Big Data adoption: State of the art and research challenges Information Processing and Management Yadegaridehkordi et al., 2018b Influence of big data adoption on manufacturing companies' performance: An integrated DEMATELANFIS approach Technological Forecasting and Social Change Salleh et al., 2015 SEC-TOE Framework: Exploring Security Determinants in Big Data Solutions Adoption PACIS 2015 Proceedings - Pacific Asia Conference on Information Systems (PACIS) Sun et al., 2020 Organizational intention to adopt big data in the B2B context: An integrated view Industrial Marketing Management Ram et al., 2019 Adoption of Big Data analytics in construction: development of a conceptual model. Built Environment Project and Asset Management Bakici et al., 2022 Big Data Adoption in Project Management: Insights from French Organizations IEEE Transactions on Engineering Management Paradza and Daramola, 2021 Business Intelligence and Business Value in Organizations: A Systematic Literature Review Sustainability Maroufkhani et al., 2022 Determinants of big data analytics adoption in small and medium-sized enterprises (SMEs) Industrial Management & Data Systems, Baig et al., 2021 A Model for Decision-Makers’ Adoption of Big Data in the Education Sector Sustainability Lai et al., 2018b Understanding the determinants of big data analytics (BDA) adoption in logistics and supply chain management: An empirical investigation International Journal of Logistics and Management Al-Rahmi et al., 2019 Big Data Adoption and Knowledge Management Sharing: An Empirical Investigation on Their Adoption and Sustainability as a Purpose of Education IEEE ACCESS 26 Mcmahon et al., 2020 Requirements for Big Data Adoption for Railway Asset Management IEEE ACCESS Müller et al., 2018 The Effect of Big Data and Analytics on Firm Performance: An Econometric Analysis Considering Industry Characteristics Journal of Management Information Systems Urbinati et al., 2019 Creating and capturing value from Big Data: A multiplecase study analysis of provider companies Technovation Günther et al., 2017 Debating Big Data: A literature review on realizing value from big data. The Journal of Strategic Information Systems Brock and Khan, 2017. Big data analytics: does organizational factor matters impact technology acceptance? Journal of Big Data Ylijoki and Porras, 2016 Conceptualizing Big Data: Analysis of case studies. Intelligent systems in accounting, finance and management Chen et al., 2017 How Lufthansa Capitalized on Big Data for Business Model Renovation. MIS Quarterly Executive In sequence, an in-depth analysis of the references (Morakanyane et al., 2017) took place. During this stage, it was possible to identify how researchers understand and define the concept of Big Data. The analysis was based on the definition, characteristics, and the attributes of the concept. Attributes are a central lens of this analysis, considering their relevance to conceptualization (Welch et al. , 2016). So, the development of a concept centric matrix was deployed, showing the different concepts and sources, presented in Table 5. It is necessary to remark that the twenty different definitions represent a polysemic phenomenon. As an example, there is a distinct understanding of Big Data in the same scientific periodical, such as the “IEEE Access” journal, where Al-Rahmi et al. (2019) and Mcmahon et al. (2020) conceptualized Big Data with different meanings and approaches. Table 5 – Big Data: the current conceptualizations Authors Big Data Definitions Gangwar, 2018 “ It is a collection of huge and complex amalgamation of data sets which make it difficult to process using traditional data processing platforms. ” Sun et al. , 2018b “In this research, from a business solution perspective, we define big data as a new technology that is primarily characterized by its advanced business intelligence and analytics (BI&A) function ." 27 Authors Big Data Definitions Baig et al. , 2019 “ Big Data is defined as a large dataset that could not be controlled or examined through conventional databases .” Yadegaridehkordi et al. , 2018b “ Big Data is currently considered as a fast-emergence amount of data from different processes that gradually poses a challenge to industrial organizations all around the world. ” Salleh et al. , 2015 “ The term (Big Data) is often associated to three unique characteristics of data: Volume, Variety, and Velocity or more widely known as the 3Vs. ” Sun et al. , 2020 (…) “ describe the data sets and analytical techniques in applications that are so large (from terabytes to exabytes) and complex (from sensor to social media data) that they require advanced and unique data storage, management, analysis, and visualization technologies. ” Ram et al. , 2019 For the purpose of clarity, BD (Big Data) is defined here as “high-volume, high-velocity and/or high-variety information assets that demand cost-effective, innovative forms of information processing that enable enhanced insight, decision making and process automation.” Paradza and Daramola, 2021 “The latest term that evolved from BI and BA is Big Data or Big Data Analytics which represents larger volumes and complex sets of data requiring dedicated tools to synthesize information while still upholding the same emphasis on reporting and predictive analytics.” Bakici et al. , 2022 “In this article, Big Data are defined as datasets in such a large scale that typical tools and software cannot store, manage, and analyse it as before. The Big Data come in various forms (structured and unstructured) from various sources and requires a fast-realtime analysis. Broadly, Big Data technologies include distributed file systems (Hadoop MapReduce framework), NoSQL, MPP systems, cloud computing platforms (storage and computing resources), in-memory database processing, and data mining tools.” Maroufkhani et al ., 2022 “The term “Big Data” refers to the massive amounts of data, including both unstructured and structured, which is accessible in real-time.” Baig et al. , 2021 “ Big Data refers to huge and multifaceted data sets that need potent storage systems and tools [2]. The characteristics of Big Data make it distinguishable from conventional data. Most of the research characterized Big Data into seven V’s [3], namely: Volume (amount of data), Velocity (speed of data), Variety (structure of data), Veracity (accuracy of data), Variability (meaning of data), Volatility (validity of data), and Value (use of data).” 34 This trend indicates that Big Data conceptualization is suffering a refinement owing to its multidisciplinary nature, not only circumscribed to the information systems and information technology domains. Like the categories, the attributes support the conceptualization analysis, indicating a possibility to develop a phrase-template or a “pattern that suggests” (Morakanyane et al., 2017) a concept for Big Data: It is something… (+) with/that [a special feature] added to a [complexity level] (+) to deliver/deploy solutions form a specific/general proposal or exemplification. Based on this phrase-template, it is possible to evolve to an adjudicated conceptualization of Big Data. To build this concept it is necessary to select which attributes can be adopted from the references in the concept centric matrix. Table 7 presents the selected references and the extracted attributes to build a new Big Data concept. All four categories (use, data features, strategy and coverage) were selected to conceptualize the Big Data. There were the uses of five attributes from the first part, “essence”, three from “special feature”, and six from “uses/purpose”. In the end, the developed conceptualization of Big Data balances the trade-off between large coverage and strong accuracy, as remarked by Welch et al. (2016). Considering the selected fourteen attributes, it is possible to remark that Big Data, as a concept, presents the prevalence of two categories: data features and uses. The conceptualization to this term must approach the types of data that this technology/methodology are based on. In the same context, the uses and applications for Big Data have to be considered as part of its meaning. If Big Data is a disruptive innovation (Ylijoki and Porras 2018; Maroufkhani et al. 2020), this disruption demands an applicability to take place. Then, the adjudicative concept of Big Data can be present in two sentences, as shown below: As a multidisciplinary phenomenon, Big Data are used to define the huge data sets and complex analytical techniques in applications with various data formats (structured and unstructured) from various data sources, requiring a fast-real-time capability to generate, capture, store, process, analyse and visualize data. Big data adopters are able to enhance insight, decision-making and process automation, by extracting value from data.” 35 Table 7 - Selected Attributes Matrix Authors and Year Categories Essence Special feature Uses/purpos es Ylijoki, O., & P orras , J. (2016) Big data is a multidisciplinary phenomenon. Sun, S., Hall, D. J., & Cegielski, C. G (2020) Big data and big data analytics are used to “describe the data sets and analytical techniques in applications that are so large (from terabytes to exabytes) and complex (from sensor to social media data) that they require advanced and unique data storage, management, analysis, and vis u a liz a t io n t e c h n o lo g ie s . ” BAKICI, Tuba; NEMEH, André; HAZIR, Öncü. (2021) that typical tools and software cannot store, manage, and analyze it as before. The big data come in various forms (structured and unstructured) from various sources and requires a fast-real-time analysis Broadly, big data technologies include distributed file systems (Hadoop MapReduce framework), NoS QL, MP P systems, cloud computing platforms (s torage and computing res ources ), inmemory database processing, and data mining tools . Günther, W. A., Mehrizi, M. H. R., Huys man, M., & Feldberg, F. (2017) that are generated, captured, and processed at high velocity Baig M.I., Shuib L., Yadegaridehkordi E. (2021) Big data refers to huge and multifaceted data sets Maroufkhani, P ., Ira n m a n e s h , M. a n d Ghobakhloo, M. (2022) which is acces s ible in real-time. La i, YY; S un, HF; Re n, JF (2018) Strategy which us ually require organizations to improve proces s ing abilities and analyzing capabilities Sun S., Cegielski C.G., Jia L., Hall D.J. (2018) that is primarily characterized by its advanced business intelligence and analytics (BI&A) function." Ram, J., Afridi, N. K., & Khan, K. A. (2019) that enable enhanced insight, decision making and process automation” Al-Rahmi W.M., Yahaya N., Aldraiwees h A.A., Alturki U., Alamri M., Bin Saud M.S., Kamin Y.B., Aljeraiwi A.A., Alhamed O.A. (2019) Big Data are Information asset Paradza, D., & Daramola, O. (2021) The latest term that evolved from BI and BA is Big Data or Big Data Analytics which represents larger volumes and complex sets of data Converage Sources Us es 36 This Big Data conceptualization is a reference that can be used in following studies about the theme. It is composed of a wide range of attributes, requirements and applications which makes this as a feasible and updated concept. Focusing on the Official Statistics domain, this concept can also be adopted. However, this context presents singularities described in the first chapter, in the section “Global Statistical System”. It is possible to summarize those singularities in the ten principles of the Official Statistics. Then, the developed conceptualization of Big Data, presented above, can be applied to Official Statistics with some adjustments. In view of this demand, the Big Data concept for Official Statistics can be defined as: Big data are used to define the huge data sets and complex analytical techniques in applications with various data formats (structured and unstructured), from various data sources, requiring a fast-real-time capability to generate, capture, store, process, analyse and visualize data. Following the principles of Official Statistics, those datasets and techniques can be adopted to produce reliable, accurate and timely public data. Concluding this section, the conceptualization of Big Data for Official Statistics is the current meaning to be understood when the term Big Data will be cited in the following pages. This adjusted concept fills a remaining gap in the literature involving Big Data and Official Statistics. The adoption models used to study Big Data adoption are reviewed in the next subsection. Considering the number of frameworks used in Big Data studies, the models are analysed under the applicability to this research focus. 2.2.2. Theoretical frameworks for the adoption of Big Data The diversity of frameworks can be an indicator of the complexity in studying and understanding the technological adoption process. Figure 6 presents those frameworks organized by the year of the original publication. This timeline represents the evolution of frameworks applied in Technology adoption studies. The more representative are briefly described next. 37 Legend: Innovation adoption frameworks organized by the year of the original publication. Source: adapted from Varajão et al. (2021). Figure 6 – Timeline of the innovation adoption frameworks evolution Varajão et al. (2021) listed eleven models and analytical structures as the most used in the literature to understand technological adoption processes or the adoption project to deploy a new information technology asset. This will be used to analyse the 93 articles selected in the systematic literature review described in section 2.1 in order to identify which IT adoption models have been applied to the study of ABD. We start by reviewing eight of those models. Diffusion Of Innovation (DOI) theory . Developed by Rogers (1962, 1983, 1995, 2004), this framework was originally adopted in rural sociology studies. The analytical level of this model is the individual, and the influence of adopter/user’s decision in the organizations (Gallivan, 2001). About the analytical dimensions, the DOI framework considers the adopter/user, the organization, and the technology. Its set of factors are composed by innovation, communication, social system, time, individual education and cultural level, knowledge transfer, barriers to knowledge transfer, individual learning capacity, and assimilation gap (Fichman and Kemerer, 1999; Gilbert and Cordey-hayes, 1996; Lundblad, 2003; Miranda et al., 2016). Theory of Reasoned Action (TRA). This framework was applied to study behavioural processes, and emerged from psychological research (Fishbein and Ajzen, 1977). Although its initial application did not approach information system management and organizational behaviour, the TRA is relevant as a framework-base. It supported the development of other 38 frameworks, like TAM and UTAT. At the individual level, TRA applies two factors (attitude towards behaviour and subjective norms) to identify individuals' behaviour intention (Varajão et al ., 2021). Additionally, TRA proposes that user/adopter behaviour can be measured by beliefs and expectations related to the outcomes and attributes from specific actions (Mital et al. , 2018). Currently, TRA is applied in distinct research fields studies, from the Adoption of Internet of Things (Mital et al. , 2018) to “ the consumer attitude towards green consumption” (Paul et al. 2016, p. 124). Theory of Planed Behaviour (TPB). Another individual-level framework, the Theory of Planed Behaviour (or TPB), is an evolution of TRA (Varajão et al. , 2021). It was built by one of the TRA’s authors, Ajzen (1991). This framework proposes that an individual's behaviour results from three factors: subjective norms, attitude towards behaviour, and the individual-self-perception of the attitudes control (Paul et al. , 2016). Like TRA, the TPB also supported the development of the Technology Acceptance Model (TAM). Mital et al. (2018, p. 341) classify TPB and TAM in function of the dimensions of analysis. TPB analyses reduced parts of the user perceptions. In its turn, TAM has a wider analytical focus, considering the mix of perception and behaviour. Technology of Acceptance Model (TAM). As an evolution of TRA and TPB, the Technology of Acceptance Model (or TAM) is one of the most powerful frameworks to study how users decide to adopt or avoid vanguard technologies (Al-Rahmi et al. , 2019). Davis (1986)and Davis et al. (1989) developed TAM considering that the user’s attitude toward using technologies receives influence from two specific factors, perceived usefulness and perceived ease of use (Varajão et al. , 2021). Following Davis’s original proposal, other studies added new factors to this framework, reinforcing its adaptability and feasibility (Mital et al., 2018). Technology-Organization-Environment Model (TOE). Technology-OrganizationEnvironment model (or TOE) is an adaptative framework to study technology adoption (Varajão et al. , 2021). DePietro et al. , (1990) developed it as an evolution of the DOI model (Angwar 2018, p.4). Based on its three dimensions, the technological, organizational, and environmental, TOE’s unit of analysis is the organization, and not the individual. This feature supports the model to be used in studies about ABD by companies, as remarked by Baig et al. (2019). Task-Technology Fit (TTF) - This framework complements a gap in the Technology of Acceptance Model (TAM) concerned about IT users task motivation (Varajão et al. , 2021). The authors of TTF, Goodhue and Thompson (1995), developed Technology-to-Performance Chain 39 structure, composed by three factors: (i) IT characteristics, (ii) tasks, and (iii) individual users. Those three factors can support the understanding of why a person uses an information system and improves or reduces his/her performance. DeLone and McLean IS Success Model (DMISM). Distinct from some previous frameworks, DeLone and McLean IS Success Model (or DMISM), has been developed specifically to Information System studies (Varajão et al., 2021). The authors, DeLone and McLean (2003), proposed the use of this model in e-commerce systems success adoption. One important feature of DMISM is the link between the individual and the organizational dimensions, offering a bottomup perspective of the adoption process inside an organization. Wang and Liao (2008) tested DMISM to measure the success of e-Government systems. Petter et al. (2013) operated an extensive literature review and identified twenty success factors to be achieved in a well-deployed adoption process, with the use of DMISM. Unified Theory of Acceptance and Use of Technology (UTAUT). Venkatesh et al. (2003, p. 428-432) developed this framework with the aggregation of features from other previous models: (i) TRA, (ii) TAM, (iii) MM, (iv) TPB, (v) DOI, (vi) SCT; (vii) model of Personal Computer Utilization (Thompson et al., 1991); (viii) hybrid model using TAM and TPB (Taylor and Todd 1995). In this model, a set of individual factors associated with environmental elements influence the behavioural intention to adopt technology. Respectively, the individual components are age, gender, experience, and voluntariness of use, and the environmental elements are facilitating conditions, performance expectancy, effort expectancy, and social influence (Venkatesh et al . 2003, p. 447). However, the technological adoption field of study is not restricted to those frameworks. There are theoretical constructs adapted from the other areas, such as project management. 2.2.2.1. Frameworks used to study the adoption of Big Data. We now focus specifically on the models used in the 93 references included in the literature review. In total, twenty-two frameworks or theoretical constructs are mentioned. Twenty-six of the 93 references do not apply any specific framework, especially literature review articles. In some papers, two or more frameworks are adopted (e.g., Pillai and Sivathanu, 2020; Queiroz and Farias Pereira, 2019). Figure 7 presents the frequency of frameworks in this literature review. 40 Figure 7 – The relevant frameworks present in this literature review Specifically, about ABD, five frameworks were identified: Diffusion of Innovation (DOI), Technology-Organization-Environment model (TOE), Technology Acceptance Model (TAM), TaskTechnology Fit (TTF), Unified Theory of Acceptance and Use of Technology (UTAUT). The frameworks’ features considered in this discussion are those relevant to study ABD processes. The review will focus mainly on the following elements: levels of analysis (individual or organizational), dimensions (internal or external), and influencing factors (antecedent variables). Diffusion of Innovation theory (DOI) Since its development by Rogers (1983, 1995), the DOI model is widespread in technological adoption studies (Daradkeh, 2019). ABD is also studied with the support of DOI. Baig et al. (2019) applied a combination of DOI with the TOE framework to understand the adoption decision. The authors remarked that five factors influence ABD: relative advantage, complexity, compatibility, trialability, and observation. Daradkeh (2019) also associated DOI with TOE to understand which factors are relevant to the adoption of Big Data Analytics and identified thirty-four. In both studies (Baig et al. 2019 and Daradkeh 2019), the researchers applied the combination of frameworks to understand ABD from the individual to the organizational level. In a different approach, Saheb (2020) used DOI to explain the impact of digital services of social connectivity and personalization on the adoption behaviour. 26 19 14 14 10 531 1 APLICATIONS No model TAM TOE Other models UTAUT DOI TPB TRA DeLone and McLean 41 Technology of Acceptance Model (TAM) Al-Rahmi et al. (2019) used TAM to understand how users adopt or avoid Big Data technologies in the educational sector, at the individual level. In this study, the authors added eight factors to the TAM framework: (i) perceived risk; (ii) users age diversity; (iii) cultural diversity; (iv) motivators; (v) behaviour intention; (vi) knowledge management sharing; (vii) ABD readiness; (viii) sustainability for education. Other research about the adoption of Big Data analytics in healthcare organizations (Shahbaz et al., 2020) included two factors – gender and resistance to change – to the original framework, considering the internal dimension of the organization. Task-Technology Fit (TTF) Sinha et al. (2019) applied the TTF framework to verify the effectiveness of the Internet of Things (IOT) and Big Data in disaster management. Gangwar (2020) studied the integration of TAM and TTF to understand how the usage of Big Data Analytics affects business performance. Considering the research developed by Gangwar (2020), TIF has the potential to evolve from the individual to the organizational analytical level. Unified Theory of Acceptance and Use of Technology (UTAUT) The UTAUT model is used in ABD studies, considering both the individual and the organizational levels. Cabrera-Sanchez and Villarejo-Ramos (2019), based on the UTAUT model, investigated the adoption of Big Data Analytics in companies from different sectors: human resources, finance, marketing, and retail. Like other frameworks, UTAUT is adaptable, and in this study, it worked with two added new factors: perceived risk, and resistance to use. Expanding UTAUT's applicability, Cabrera-Sánchez and Villarejo-Ramos (2020) added a new adoption factor to the original framework: opportunity cost. In that study, the investigation core was the adoption process of Big Data in 199 Spanish service companies. Silva et al. (2019), applying UTAUT, concluded that the intention to adopt Big Data in Peruvian companies is influenced by Performance Expectancy, Social Influence, and Facilitating Conditions. This framework represents a resilient and adaptable model, and it is well-explored to understand the ABD process at the individual level. TRA, TPB, and DeLone and McLean IS Success Model (DMISM) These three frameworks are present in two articles included in the LR. Firstly, Kohli and Tan (2016) studied the adoption of electronic health records by healthcare companies. In this 42 research they verified that the knowledge developed with the use of DeLone and McLean IS Success Model supported the innovation process. Secondly, Theory of Reasonable Action and Theory of Planned Behaviour are two frameworks relevant to be used in studies involving data applications such as blockchain adoption (Wong et al., 2020). There is no strict link between those three frameworks and ABD. Technology-Organization-Environment Model (TOE) In our literature review (LR), 14 references (16%) use TOE as the principal framework to study ABD. Salleh et al. (2015) remarked the power of TOE is clustered in its factor analysis using three analytical dimensions: technological, organizational, and environmental. Additionally, Salleh and Janczewski (2016) highlighted TOE’s capacity to explain an adoption element, such as IT security, under internal and external perspectives, including factors such as top management support, organisational culture, and IT complexity. Considering the status of the National Statistics Offices, and their organizational network structure, it is critical to adopt a model that considers the organizational analytical level. Following this logic, and considering the characteristics of the reviewed frameworks, the TOE model is the selected framework to be used in this doctoral research. In the following subsection, TOE model’s origins, features, and factors are discussed with the focus on ABD. 2.3. The Technology-Organization-Environment model (TOE) The TOE model encompasses three dimensions: technology, organization and environment (DePietro et al., 1990), as presented in Figure 8. This triple analytical dimensions represent an advantage when TOE is compared with other frameworks, such as DOI, with only two analytical dimensions (Park and Kim 2019; Lai et al. 2018b). Organizational elements, such as culture, skills, company size, or internal resources (slack) are considered in this framework. In this perspective, TOE analyses also the economic and business environment and evaluates how those external factors affect the adoption process. About the Environment dimension, industry characteristics, the external stakeholders, national economic scenario, national business environment, and public regulation are considered relevant factors to influence the adoption process (Baker, 2012). 43 Legend: Original version of TOE framework, developed by DePietro, et al (1990). Source: Tornatzky and Fleischer (1990); Baker (2012). Figure 8 – TOE Framework: original graphic structure The seminal publication about the TOE framework took place thirty-two years ago, by DePietro et al. (1990). Previously, Rogers (1962), Fishbein and Ajzen (1977), Ajzen (1991), Davis (1986), and Bandura (1985) approached the theme of technological innovation by considering the individual and the technological dimensions. In order to bring the external influence to technological adoption, DePietro et al. (1990) included the environment as the external dimension of the technological innovation process, thus completing its present configuration of the three delimiting contexts of innovation decisions: technology dimension, organization dimension, environment dimension. Baker (2012) remarks the three dimensions offer a real picture of the organization during the adoption process. Firstly, there is the Technology dimension, which comprises an external and an internal component. The external component is where the IT adoption is analysed under the perspective of technology availability, demanded methods, risks and costs. The internal component looks at decision-making capacities, and current assets (DePietro et al. , 1990). This availability means the technological assets that need to be acquired by the organization from external suppliers. It includes licenses and copywriting, number of suppliers able to offer IT solutions, and the costs of the acquisition. Inside the firm, there is another set of adoption factors related to the internal technological conjuncture (the central box in Figure 8). In other words, an organization must consider the current state of its technological equipment. Additionally, the complexity of the IT adoption process must be examined. Complexity refers to the incremental or discontinuous 50 Authors Title Journal /Conference Title Verma and Bhattacharyya, 2017b Perceived strategic value-based adoption of Big Data Analytics in emerging economy A qualitative approach for Indian firms Journal of Enterprising Information Management Gangwar, 2018 Understanding the Determinants of Big Data Adoption in India: An Analysis of the Manufacturing and Services Sectors Information Resources Management Journal Sun et al., 2018b Understanding the Factors Affecting the Organizational Adoption of Big Data Journal of Computer Information Systems Yadegaridehkordi et al., 2018 Influence of Big Data adoption on manufacturing companies' performance: An integrated DEMATEL-ANFIS approach Technological Forecasting and Social Change Lai et al., 2018a Understanding the determinants of Big Data analytics (BDA) adoption in logistics and supply chain management: An empirical investigation International Journal of Logistics Management Ram et al., 2019 Adoption of Big Data analytics in construction: development of a conceptual model. Built Environment Project and Asset Management Park and Kim, 2019 Factors Activating Big Data Adoption by Korean Journal of Computer Information Systems Baig et al., 2019 Big Data adoption: State of the art and research challenges Information Processing and Management Sun et al., 2020 Organisational intention to adopt big data in the B2B context: An integrated view Industrial Marketing Management Picoto et al., 2021 The influence of the technology-organisationenvironment framework and strategic orientation on cloud computing use, enterprise mobility, and performance. Revista Brasileira de Gestão de Negócios Bakici et al., 2022 Big Data Adoption in Project Management: Insights from French Organisations IEEE Transactions on Engineering Management Al Hadwer et al., 2021 A systematic review of organisational factors impacting cloud-based technology adoption Internet of Things 51 Authors Title Journal /Conference Title using Technology-organization-environment framework Paradza and Daramola, 2021 Business Intelligence and Business Value in Organisations: A Systematic Literature Review Sustainability Baig et al., 2021 A Model for Decision-Makers’ Adoption of Big Data in the Education Sector Sustainability El-Haddadeh et al., 2021 Value creation for realising the sustainable development goals: Fostering organisational adoption of big data analytics. Journal of Business Research Ghaleb et al., 2021 The Assessment of Big Data Adoption Readiness with a Technology–Organisation–Environment Framework: A Perspective towards Healthcare Employees Sustainability Maroufkhani et al., 2022 Determinants of Big Data analytics adoption in small and medium-sized enterprises (SMEs) Industrial Management & Data Systems As presented in Table 9, the number of factors in the references reaches a total of 234 items! This tremendous and diverse number of factors constitutes a barrier to research advance in this area. Considering the epistemological proposal of Cronin and George (2020), 234 factors were submitted to check for similarity. Table 9 – Factors per dimensions in Adoption of Big Data studies Dimensions Number of Factors Technology 99 Organisation 77 Environment 58 Total 234 In an attempt to reduce the semantic noise, the factor description of each factor was checked and compared using a convergence criterion. If the factor’s descriptions present similarities and the text approaches the same theme, this is considered a factor with high convergence. In the opposite direction, a factor with discrepant descriptions is considered a low convergence factor. Following this thought, the research proposed rearranging the factors based on the respective 52 descriptions. In the end, high convergence factors kept the original name. Low convergence factors were joined under a wide common name to allow the reader a better understanding of the meaning. There are factors with close nomenclature. An example is the organisational factor, named “top management support”, which presented similar names but similar descriptions, as shown in Table 10. This factor counted fourteen definitions and three non-definitions. Much more than an accounting issue, different names for the same factor indicate the necessity to check all descriptions and verify whether they are part of a thematic group. Table 10 – Polysemy in TOE factors nomenclature Authors Authors’ Factor Description Same factor, different labels Sun et al. (2018, p. 198) Managers are willing to allocate sufficient resources and encourage the initiative adoption of Big Data (e.g., top executives responsible for data management, CIOs’ willingness to adopt Big Data). Management Support Yadegaridehkordi et al . (2018, p. 201) The authors gave no definition. Management Support Park and Kim (2019, p. 3) Management support refers to the degree to which management recognises the importance and relevance of Big Data. Management Support for Big Data El-Haddadeh et al. (2021) Top management support explains the extent to which top management understands the importance of innovation and the extent to which it is involved in new innovative activities, such as ABD Top Management Support Gangwar (2018, p. 7) Top management support refers to the degree to which top management understands the strategic importance of IS innovation and the extent to which it is involved in IS activities. Top Management Support Verma and Bhattacharya (2017, p. 361) Referred to devoting time to the (IS) program in proportion to its cost and potential, reviewing plans, following up on results and facilitating the management problems involved Top Management Support 53 Authors Authors’ Factor Description Same factor, different labels with integrating ICT with the management process of the business (p. 361) Baig et al. (2019) Management support is a degree to which it perceives the significance and applicability of Big Data adoption (Tornatzky et al., 1990). Usually, top management does not support the ‘change’ factor. In big data, adoption process changes are required when top management becomes reluctant to change for improvement, then the whole organisation starts following higher management decisions, which will linger the adoption process. Management support is necessary to encompass the organisational policies, rules, handling of data storage issues, technological and financial capability. Therefore, management support is a significant factor that contributes to the adoption of Big Data. Management Support for Big Data Bakici et al. (2021, p. 12) Commitment, availability of data, skilled human Resources, training. Top Management and Strategy Lai et al. (2018, p. 361) The degree to which top management understands the importance of the IS function and the extent to which it is involved in IS activities. Technology Management Support Ghaleb et al. (2021, p.12) Refers to the extent to which managers grasp and embrace a new technology system’s technological potential. Top Management Support Paradza and Daramola (2021) No definition by the authors Top Management Support Al Hadwer et al. (2021) No definition by the authors Top Management Support On the opposite side, an example of a low convergence name is the environmental factor “Security and privacy regulations”, as presented in Table 11. In this case, four different definitions, two non-definitions and six different names integrate the factor’s conceptualisation. “Security and privacy regulations” is a factor related to the plethora of laws and acts to regulate 54 one specific issue: security and privacy. Clients’ data, collected and processed by a company’s Big Data application, must have privacy protection from any leak or irregular dissemination. Table 11 – Low convergence factor: the example of “Security and privacy regulations” Authors Authors’ Factor Description The same factor, different names Sun; Cegielski; Jia; and Hall. (2018, p. 198) Data collection from individuals causes individuals’ security, privacy concerns. (e.g., legal implications of collecting customers’ private information). Security, Privacy and ethical concerns in collecting data. Salleh; Janczewski; and Beltran. (2015) Security and privacy regulatory concerns refer to organisational concerns in ensuring compliance to security and data privacy regulations Security and Privacy Regulatory Concerns Baig; Shuib; and Yadegaridehkordi. (2019, p. 10) Security, privacy and risk factors will not only affect the big data adoption process but will also destroy the company reputations. Security, Privacy and Risk Baig; Shuib; and Yadegaridehkordi. (2021, p.7) Security and privacy concerns are associated with individual safety and moral values. Security and privacy concerns were the obstructing factors of ABD. Security and Privacy Concerns Salleh and Janczewski. (2016) No definition by the authors Privacy Regulatory Concerns Al Hadwer; Tavana; Gillis; and Rezania. (2021) No definition by the authors Cloud Locality The factors' polysemy is also present in the Technology Dimension but to a lower extent. For example, the factors “Complexity” and “Compatibility” have nineteen and fourteen references. Besides the vast number of citations, it is possible to assert that both names define specific concepts. Table 12 brings three definitions of “Compatibility”. They all consider how a Big Data solution is compatible or incompatible with the current organisational processes and systems. In this factors analysis, an association to a familiar name took place regarding the redundancies and the convergences of each factor’s descriptions. The reduction of factors was drastic, and it was possible to gather several factors in a shorter-term/definition group. Beyond the polysemic reduction, the other outcome for future TOE studies is the possibility of using a set of references proposed by this literature review. 55 Table 12 – Technological factors description: the example of “Compatibility.” Compatibility means… " The characteristics of big data are perceived as being consistent with the existing IT architecture in an organisation (e.g., scalability, integration into the existing information systems). (p. 198) " Sun, Cegielski, Jia & Hall (2018) “ The concept of compatibility refers to the level at which innovation is perceived as consistent with internal organisational processes and information systems. (p. 283)” Picoto, Crespo & Carvalho (2021) “ Compatibility examines “the degree to which a new system is consistent with the current system within the company. ” Maroufkhani, Iranmanesh, & Ghobakhloo (2020) Summarising, Figure 9 shows the procedure of the assembling factors. From the original number, 234 (99 Technological, 77 Organisational, and 58 Environmental), to the final, the reduced set presented 43 factors. This adjudication in the factors promoted the capability to analyse ABD with a consistent set of factors. This consistency comes from (i) the sources, as relevant articles, (ii) the description associated with the noise reduction/elimination, and (iii) the evaluation of which factor describe ABD in organisations. Legend: The reduction of the TOE factors presented in the literature about Adoption of Big Data. The original 99 technological, 77 organisational, and 58 environmental factors were refined to 14 technological, 17 organisational, and 12 environmental factors, after eliminating redundant meanings. Figure 9 – Aggregating and simplifying the TOE factors. 14 (T) + 17 (O) + 12 (E) = 43 E (58) O (77) T (99) 56 This section presents an extensive list of the factors and their descriptions. Those descriptions and the factors’ names are detailed in Tables 11, 12 and 13. It is necessary to highlight that those tables result from a synthesis, as described before, including names and definitions. Hence, the final factors set is the base to analyse ABD under the TOE’s perspective. Following this interpretation, the research's analytical lens refers to the 43 factors. In this way, the three dimensions are considered subsets of factors, ordering groups of Technological, Organisational, and Environmental factors. Regarding those selection criteria, the result is detailed in the following three subsections regarding the TOE dimensions. 3.1.1. Technology factors The Technology dimension kept the totality of the fourteen reviewed factors. There is a reason to support all factors: Big Data is not a single technological application. As cited at the beginning of this chapter, it presents a combination of technologies, such as AI, IoT, Cloud, and Web Scraping. In this perspective, an extensive list of factors allows the content analysis to capture details of ABD in NSOs. So, the technological factors used in this research are those presented and defined in Table 13. Table 13 – Technology factors Name Authors Factor Description Compatibility Picoto et al. (2021, p. 283) “The concept of compatibility refers to the level at which innovation is perceived as consistent with internal organisational processes and information systems.” Data quality Lai et al. (2018a, p. 682) Data quality means the degree to which the data needed for Big Data are accessible, consistent, and complete. Data integration Park and Kim (2021) Data integration refers to the degree to which the data is analysed by the Big Data systems, considering the relevance of Big Data applications and the integration of those with the data collected. Complexity Sun et al. (2018, p. 198) Complexity refers to “the characteristics of Big Data that are perceived as being difficult to understand and use (e.g., the difficulty of learning related knowledge for employees who will use Big Data applications),” Perceived benefits Park and Kim (2019, p.2) Perceived benefits refers to "the degree to which firms perceive benefits (such as cost reduction, operation improvement, and marketing performance) deriving from Big Data." 57 Name Authors Factor Description Technology resources El-Haddadeh et al. (2021) Technology resources represent tangible and intangible resources such as hardware, software, human resources, skills, and experience that an organisation acquires for implementing Big Data. Security and privacy Park and Kim (2021, p. 287) Security and privacy are associated with privacy invasions and security risks posed by Big Data. Relative advantage Sun et al. (2018, p.198) Relative advantage means “the characteristics of Big Data are perceived as being better than those of the idea it supersedes (e.g., its unique role for innovation, competition, productivity, customer value creation, and good business problem solution).” Trialability Baig et al. (2019, p. 10) "Trialability is the degree to which companies can experiment with Big Data before fully implementing or committing to Big Data adoption." Technology competence Sun et al. (2018, p.198) Technology competence means the "organization has sufficient internal IT expertise and technological infrastructure to adopt Big Data (e.g., IT knowledge and skills within the organisation)." Cost of adoption Baig et al. (2019), Sun et al. (2018b), Bakici et al. (2022) Cost of adoption refers to the financial resources demanded and executed expenses related to Big Data adoption (e.g., the costs of using Big Data technology, the significant initial investment required to embrace ABD). Accuracy Based on Baig et al. (2019), Yadegaridehkordi et al. (2018a), and Bakici et al. (2022). Accuracy refers to ability to overcome the data collection and evaluation limitations due to sampling, allowing the organisation to deploy faster, predictable, accountable and truthful outcomes. Internal versus external technologies Baig et al. (2019, p. 9) Internal versus external technologies refers to “the hardware and software technologies used for Big Data adoption. Technologies revealed by retailers are internal technologies, whereas those provided by vendors are called external technologies.” Vendor support Baig et al. (2019) Vendor support can be defined as the services and activities to support the ABD in NSOs when the solutions are bought with commercial suppliers. 58 3.1.2. Organisation factors Orlikowski (2000) highlighted that technology influences the organisation and vice-versa, limiting or spreading the adoption into the departments and areas. In this way, Organisational the selection of factors was less straightforward than technological ones. Firstly, two reviewed factors were discarded. “System perspective” did not present a description. “Optimism” presented a description indicating this is not a factor but an output from ABD. Secondly, the most cited factor, “top management support,” indicated a divergence point. Some authors indicate that management support is widespread, from line management to the top level. Others consider this factor exclusively from the strategic layer of a company. Despite this double approach, after examining the descriptions, it was concluded that this is the same factor. However, this indication will be checked during the interview analysis and define if this is indeed a single factor, or two conflated under one name. Thirdly, some factor labels are misleading and could lead to a misinterpretation of their real meaning. “IS strategy orientation” and “Technological capabilities” are organisational factors. However, they could be understood as technological factors. As a strategic orientation, “IS strategy orientation” integrates the strategic planning of each company. “Technological capabilities” is closer to employees than computers. It means the skills of the whole organisation and the employees to deal with the adoption of Big Data. This challenge includes the organisational impacts, such as processes and methodologies. Another issue of the Organisational factors is the word “readiness”. To avoid noise, this word, in this study, means “ willingness or a state of being prepared for something 1 ”. For that matter, the list of organisational factors is thus presented in Table 14. Table 14 – Organisation factors Name Authors Factor Description Top Management support Park and Kim (2019) Top Management support is the degree to which high-level management recognises the importance of Big Data, sponsoring and supporting the adoption process. Firm Size Sun et al. (2018, p.198) Firm size refers to "the firm’s annual revenue and number of employees that could support the adoption of Big Data (e.g., leading companies with more revenue)." 1 Cambridge Dictionary. Retrieved from https://dictionary.cambridge.org/dictionary/english/readiness. Accessed in 01/02/2021. 59 Name Authors Factor Description Financial capacity Baig et al. (2019) Financial capacity means the skills of the staff to manage budget and financial resources in order to make available the monetary conditions for ABD. Change efficiency and efficacy Baig et al. (2019) Change efficiency and efficacy refers to the idea that the employees of NSOs should be capable of handling changes related to ABD and obtaining better results from this adoption. Business strategy orientation Baig et al. (2019, p.9) Business strategy orientation "refers to a strategy that is used to learn about business analytics and making strategic decisions for organisations in adopting Big Data." Human resources and Skills Baig et al. (2019, p.9) "Human resources play a significant role in employing and maintaining technological innovation. Sufficient skilled human resources are required for Big Data adoption. Strong programming, statistical and analytical skills help in implementing Big Data in retail organisations." Decision-making culture Baig et al. (2019, p.9) Decision-making culture "refers to as a belief that the adoption of Big Data is a key to success in enhancing organisational productivity. At the company level, executive or managerial staffs make a decision." Technological capabilities Park and Kim (2019, p.3) Technological capabilities refers to "the degree to which firms have the ability to operate Big Data systems in terms of technological readiness." Organisational structure (2018, p. 198) Organisational structure implies "a well-organised structure is well-suited to the adoption of Big Data (e.g., crossorganisational collaboration structure, IT departments, and staff configurations)." Organisational learning culture Salleh et al. (2015) Organisational learning culture is “the learning characteristics and orientation of an organization, especially in the process of complex technology adoption” Organisational data environment Verma and Bhattacharyya (2017, p.361) Organisational data environment "refers to the extent to which data resources are managed in an organisation” Organisational readiness Gangwar (2018, p.6) Organisational readiness is "the extent to which financial, technological, and skilled human resources are accessible to an organisation wishing to embrace a technology." IS strategy orientation Sun et al. (2018, p. 198) IS strategy orientation occurs when "the firm’s IS strategy prioritises Big Data usage (e.g., information strategy, information governance policy)." 66 too vague problems should be avoided. (…)”. In the face of those criteria, it is necessary to define a research problem with technique. One of this technique’s elements is the possibility to develop ideas through discussions with peers and individuals, searching for “ a position to enlighten the researcher on different aspects of the proposed project (p. 28)”. Additionally, Kothari (2004) proposed a validation test to check whether a research problem proposal is reliable. This validation is composed of 5-checking-elements: (a) problem owner – an individual, a group, or an organisation; (b) objective(s) to be attained; (c) alternative means/channels to attend the objective(s); (d) a portion of uncertainty about the access of the sources; (e) an environmental conjuncture where the problem takes place/occurs. Firstly, the problem owner has been identified as the National Statistics Offices (NSOs). Next, the objectives to be attained are understanding why and how NSOs adopt or avoid Big Data solutions. Thirdly, alternative information sources will be official documents, structured interviews, and direct observations. The information access policy from each NSO represents the fourth item. One NSO, e.g., INE (Portugal), may be closer to external researchers compared to a similar IBGE (Brazil). Lastly, the environment where the problem takes place is the environment of disruptive changes, where NSOs have to adopt or avoid Big Data solutions. Based on the analysis of Kothari's (2004) elements, the research problem has to demonstrate consistency and resilience. Besides this validation, the research problem shall be verified in a double-check review under another approach. So, for Cresswell and Cresswell (2018, p. 57), “(…) if a concept or phenomenon needs to be explored and understood because little research has been done on it or because it involves an understudied sample, then it merits a qualitative approach. Qualitative research is especially useful when the researcher does not know the important variables to examine. This type of approach may be needed because the topic is new, the subject has never been addressed with a certain sample or group of people, and existing theories do not apply with the particular sample or group under study. ” In this case, the research problem is examined relative to those four questions: (I) Has some research been done? (II) Is the topic new? (III) Has the subject never been addressed with a specific sample? (IV) Have no existing theories been applied to the particular sample or group under study? The checklist proposed by Cresswell and Cresswell (2018) indicated that the current research problem is correctly addressed. The answer to question (I) is objective: Yes, but on a reduced 67 scale. There is no expressive number of academic works in Management and Information Systems to understand how NSOs are adopting Big Data solutions. Next, the answer to question (II) reports that the adoption of Big Data applications by NSOs is a research issue quite innovative and not widely explored. In sequence, the answer to question (III) is: no, not yet. Specifically, the TOE framework has not been commonly applied in adoption studies related to organisational networks, such as Global Statistical System. Finally, the answer to question (IV) is negative. The adoption theory has not been applied to ABD in NSOs. The theories of adoption models and pervasiveness may, but have not been applied, in an innovative perspective, to cases like NSOs. Thus, after the double validation checking, the research problem is able to be explored and detailed into a research question. This central question is addressed to understand the actual dilemma lived by NSOs. As described in Chapter 1, until the last decade, those organisations were the primary official data generators for societies. After the information technology revolution and the spread of social media, mobile phones, ubiquitous internet, 5G speed and others, NSOs are losing prominence, as remarked by Allin (2021). Eisenhardt (1989) places the research question definition in the first stage of a case study journey. According to the author, this choice allows the student to focus her/his efforts and develop solid constructs and measures. So, the research question has to summarise this scenario and the challenges to NSOs. Additionally, considering the case study method, the research question must be composed of “how” and “why” questions (Yin, 2014). In this way, the final version of the central question is this: How and why do National Statistical Offices adopt or avoid Big Data technologies and methodologies? Following this perspective, the central objective of this research is to answer the research question and explain how and why NSOs adopt or avoid Big Data solutions in their processes and business cycle. Then, if Big Data is an innovation for NSOs and it might be adopted to offer better processes and outcomes, this is probably a context of innovation adoption and acceptance. After this initial step, it is essential, also, to verify if, when adopted, Big Data solutions become pervasive inside the NSO or remain localised as an asset used only by one or a couple of departments. In this way, multiple case studies arise as an adequate research strategy (Yin, 2014). 68 4.2. Why the Case studies method? Eisenhardt (1989) remarks that the case study method is suitable for unexplored research topics. Regarding multiple case studies, initially, each of the three NSOs will be evaluated as a specific case. This specification was followed by a search for patterns in the adoption or avoidance of Big Data. The proposal is to identify convergences or divergences in the adoption processes for each NSO. Considering the classification proposed by Yin (2014), there are six different possibilities to describe multiple case studies: (i) linear-analytic structure; (ii) comparative structure; (iii) chronological structures; (iv) theory-building structures; (v) suspense structure; (vi) un-sequence structure. It is possible to propose that linear-analytic and comparative structures are more pertinent to this study. The reason for this selection considers the exploratory nature of this research. As reported in Chapter 2, the issue of ABD by NSOs has not been explored in-depth and large scale. So, the three case studies, and the conclusions extracted from them, allow the development of comparisons and standards. Specifically, in this research, two types of analyses can be adopted related to multiple case studies: pattern matching and cross-case synthesis (Yin, 2014). Furthermore, the literature review explains ABD as a non-standard process in different business sectors. The current research attempts to identify if ABD is similar to organisations in the same business sector, the Official Statistics. In this context, the multiple case studies with three different NSOs (IBGE, INE, ONS) can offer precise and objective understandings. Not only the analytical possibilities but also the features sustain the selection of the method. Yin (2014) remarks case study is a method to be applied in empirical inquiries investigating a realworld phenomenon. This phenomenon occurs in a grey zone where the limits are unclear between it and the context. It is possible to infer that ABD is part of the context. At the same time, it is also a phenomenon, considering the dimensions of the TOE framework. In the same way, there is another strength of case study methods. It has the capacity to offer insights related to a (i) new or (ii) not-well-explored topic. ABD in NSOs fits in the topic (ii). Regarding the method’s typology, it is necessary to highlight the selected case study type. Yin (2014) defined four types of case study design. Types 1 and 2 are exclusively dedicated to single- 69 case designs. This research adopts a multiple-case design, precisely Yin’s fourth type, the embedded multiple-case design. Explaining the selection, this type looks at a phenomenon replication, such as ABD. The proposal of type 4 is to analyse an initial proposition and verify how it takes place in each case. In this way, the current research investigates how and why Big Data is adopted or avoided in the three NSOs (IBGE, INE, ONS). Another feature of the adopted type in this research is delimitation. Regarding the focus on ABD, the study did not spill over to all routines from the NSOs. When supporting processes and systems were studied, they were restricted to the adoption of Big Data process interface. In summary, this research follows the same sequence proposed by Yin (2014, p. 60), as presented in Figure 12. Related to the multiple case studies, the flow of activities contains eight stages: (1) Develop theory, (2) Select cases, (3) Design data collection protocol, (4) Conduct the case studies, (5) Writing of the three individual case reports, (6) Draw cross-case conclusions, (7) Modify theory, (8) Develop policy implications. Legend: The eight stages of multiple case studies activities flow: (1) Developing the theory, (2) Selecting the cases, (3) Designing the data collection protocol, (4) Carrying out the case studies, (5) Writing the reports on the three individual cases, (6) Drawing cross-cutting conclusions, (7) Modifying the theory, (8) Developing policy implications. Source: Yin (2014). Figure 12 – Flow of multiple-case study designs Those stages converge with the roadmap proposed by Eisenhardt (1989). Stake (1995) exemplifies this generalisation, citing a case study about a child's behaviour. In this example, a particular child had been studied, and the research conclusions could be extended to other 70 children in similar contexts. In a similar way, Yin (2014) considers the case study able to generalise findings. No statistical findings were to be tested, but theoretical and conceptual conclusions were. 4.3. Designing the research As defined by Yin (2014), the research design includes an explanation of the selection criteria for each case, how the data collection will run, and the logical link between the research question and the data collection instrument. Each stage is detailed in the following subsections. 4.3.1. Selecting the cases The selection of the cases follows the logic attributed by Yin (2014, p. 57): “ The logic underlying the use of multiple-case studies is the same. Each case must be carefully selected so that it either (a) predicts similar results (a literal replication) or (b) predicts contrasting results but for anticipatable reasons (a theoretical replication).” This research attempts to achieve similar or contrasting results, looking for possible generalisations. Stake (1995) highlights that the diversity of the attributes shall not be the sole criterion for selecting a case. Variety and balance are other features to be realised in a case study sample/selected group. As the author remarks, case selection has to offer the researcher the possibility to learn about the differences. In the same way, Eisenhardt (1989) remarks that case selection is not random but based on theory and literature review, considering a specific population. In this study, the population is the members of the Global Statistical System (GSS), the NSOs. Under this perspective, the selection of the cases followed three criteria: organisational nature, geographical location, and position in the Global Statistical System. Those three criteria are analysed under the TOE model perspective. In fact, the differentiation from these criteria provides information that allows us to examine TOE factors in the Technology, Organisational and Environment dimensions. 1st criterion: Governmental Nature – Autonomous or Independent This criterion is analysed to identify differences among NSOs with autonomous and independent natures. The first means a strictly formal hierarchy from a country’s central political power. In sequence, the second refers to administrative independence, free from ministerial control. Additionally, the difference between organisational perspectives can support the analysis of each Private 71 case. In this way, the TOE framework remarks that, within Organisational factors, the governmental nature can be classified as centralised/formal (Baker, 2012). On the other hand, a non-governmental statute can be read as a decentralised/organic one. Those variations in the organisational nature allow the research related to factors in the three TOE dimensions. In the Environment dimension, factors such as national/institutional context , partners pressures , and Security and privacy regulations can be differentiated when an NSO works under an independent orientation or in a limited one. Under the Organisational dimension, business strategy orientation , decision-making culture , and financial capacity can be better supported or constrained under a specific type of governmental nature. Related to the Technology dimension, factors such as technology competence and technology resources may advance or restrict the ABD. Considering that an independent NSO will probably have a more robust financial capacity to improve its competences and infrastructures, those two technological factors can influence the adoption. 2nd criterion: Geographical location It is relevant to consider that ABD can be influenced by the business sector (Gangwar, 2018; McMahon et al., 2020) or gender (Shahbaz et al., 2019). Perhaps, the national and international context can interfere with the ABD process too. Then, the selected NSOs are organisations from three nations with asymmetrical realities. This allows us to explore the environmental factors in TOE, including the national/institutional context . The national context may explain or inform patterns to adopt or avoid Big Data. Competitive pressure and partner pressure are other two factors related to national asymmetries. A country’s regulation system, or the barriers to importing technology, determine the access opportunities for companies to adopt national or foreign solutions. This criterion can influence factors such as Government support/ laws and policy , Human resources and Skills , and Vendor support . Laws and policies vary from one country to each other. This environmental factor may support or constrain the development of the labour force and its skill. Additionally, if a country is attractive to technological companies, vendor support will probably be superior to a country with a closed economy. In this way, an open economy, such as the Portuguese, can offer access to buy national or international Information and Systems Technologies (IST). In contrast, in Brazil, the import tariffs are high, and, in many cases, it is easier and cheaper to develop internal solutions. 72 Complementarily, when a national economy is in recession, NSOs may suffer from budget restrictions, impacting their financial capacity . In fact, asymmetric geographic contrasts are potential sources for this comparative research. Illustrating those differences, Table 16 shows basic figures to highlight the geographic contrast. Table 16 – Main figures Indexes / Numbers Brazil Portugal United Kingdom Human Development Index (2017) 0,76 0,86 0,92 Gross Domestic Product (2020-US$ trillion) 1,75 0,20 2,89 Population in 2018 (millions) 211 10 67 *Source: Our World in Data. Retrieved from https://ourworldindata.org/, 2022 November 1st. 3rd criterion – Position in the Global Statistical System NSOs respond to their countries' respective governments and the United Nations (UN) through the United Nations Statistics Division - UNSD (United Nations Statistics Division, 2019). As remarked in Chapter 1, the joint operation of NSOs with UNSD forms the Global Statistical System (GSS). In other words, the position in the GSS offers the possibility to compare the three NSOs and understand if a member of this network accesses the technology and resources more simply than others. Even so, an NSO in a central position in GSS can participate in knowledge exchanges more intensely than a peripheral partner. This exchange flow can improve an NSO, making it able to modernise its outcomes and reflecting them in technological factors such as compatibility, complexity, cost of adoption, data integration, and data quality . Indeed, those five technological factors present significative levels of deployed knowledge. On the other hand, the influence of the GSS can differ in each NSO. This potential singularity can allow comparisons in the environmental factor industry features . If an NSO receives more support from the GSS and, in this way, adopts Big Data faster, the industry features can represent a critical factor. These organisations work as official data providers, and the results of their work processes - surveys, studies and projections - guide and support the stakeholders’ decision-making processes from public and private organisations across the globe. Each selected NSO for this research plays a significant role in local societies. 73 Then, it is possible to use TOE to analyse innovation in organisations embedded in an organisational network. The three NSOs, INE-PT, IBGE-BR, and ONS-UK, are connected by a central node, UNSD. This is the core/brain of the Global Statistical System. On the regional scale, European NSOs work integrated into Eurostat, the statistical office of the European Union. In summary, these three organisations, IBGE, INE and ONS, serve around 276 million citizens. The results of their work supply supranational institutions, UNSD and EUROSTAT, and the governments of 6 nations: Brazil, Portugal, England, Scotland, Wales, and Northern Ireland. 4.4. Data collection As a stage of the research project, data collection includes deciding the collection instrument(s) regarding the selected method. Quantitative research uses specific methods, like questionaries, while qualitative studies apply another type, such as structured interviews (Cresswell and Cresswell, 2018). Besides the type of research, data collection seeks to answer the research question (Yin, 2014). To achieve this target, the data collection methods shall be “flexible and opportunistic” (Eisenhardt, 1989). The design of the data collection instrument follows the literature review and the case selection under Yin's (2014) perspective. Following this understanding, Cresswell and Cresswell (2018, p. 262) remark, “ the data collection steps include setting the boundaries for the study through sampling and recruitment, collecting information through unstructured or semi-structured observations and interviews, documents, and visual materials (…).” In fact, not only the interview but also the selection and collection of documents are included in this critical research stage. Yin (2014) and Cresswell and Cresswell (2018) identify the documental collection as a rich source of research information. This understanding is discussed in the following subsection. 4.4.1. Documental analysis How and why do National Statistical Offices adopt or avoid Big Data technologies and methodologies? To achieve the objective of answering this question, data from different sources need to be collected using triangulation procedures (Yin, 2014; Eisenhardt, 1989). In the first place, it is worth explaining why to use documental analysis to collect data about a dynamic issue such as ABD in NSOs. Bowen (2009) identifies positive and negative points about 74 using this collection data technique. The positives are (i) efficient method, (ii) availability, (iii) costeffectiveness; (iv) lack of obtrusiveness and reactivity; (v) stability; (vi) exactness; (vii) coverage. The negatives are (i) insufficient detail; (ii) low retrievability; (iii) biased selectivity. The balance is positive and proves this is a well-disseminated research technique. As an initial tool for this research, the documental analysis will be used to understand and explore the three case studies and obtain support to elaborate the case study analysis. Following this perspective, after the selection of the cases, documents were collected, exclusively in digital format, considering this research was advanced during the COVID-19 pandemic. So, the documental search is the first data collection in this methodology. Bardin (2016) considers document collection as an autonomous research source. Her content analysis technique attributes a similar relevance to documents and interviews as data sources. Documents extracted from the three organisations, including website publications, digital repositories and social media channels, were obtained. INE, IBGE, and ONS offer an amount of content in their virtual addresses. Additionally, the site from UNSD was also consulted. This source considers the directives and proposals for adopting Big Data by NSOs. The documents were classified into four types: reports, descriptions, publications and outcomes. In the business architecture of Official Statistics (Radermacher, 2020), those four research data sources are related to the phase called “communication”, specifically in “user support”, “press office”, and “publications” segments. Related to the first type, reports inform the state of ABD in NSOs, such as the Big Data report from The United Nations (The United Nations Statistical Commission, 2020). Descriptions involve information about the features, operations, organisational structure, history and achievements of the NSOs. For example, the history of the first Brazilian Census is available on IBGE’s website (IBGE, 2022b). The third documental source, publications, refers to the contents in the NSOs' websites about Big Data for Official Statistics and the operations deployed by ONS, IBGE, and INE. Finally, outcomes are the descriptive material resulting from the NSOs operations. In this context of operations, the Census is the most relevant one to the studied NSOs. Considering the magnitude of the Census and its impact on NSOs, this was the more relevant operation/outcome to be observed by this researcher in the three cases. Between 2021 and 2022, those three institutions launched nationwide census operations to portray the three populations. This common period was a unique window of opportunity to realise how Big Data technologies were 75 adopted and deployed or, on the other hand, avoided by the NSOs. Another feature of the Census operations is the fact that it is a well-standardised operation for GSS members. This feature allows comparisons between IBGE, ONS and INE. In studies about new technologies, Yin (2014) proposed to observe this new technology at work. So, observations can be done by taking photographs during fieldwork or collecting physical artefacts, such as an electronic device used in a technological journey. Van Campenhoudt et al. (2017) go beyond and consider the electronic register in the field as a variation of this data collection procedure. So, photos, videos, informal conversations with NSOs members, and screenshots of the three Censuses were elements of this documental collection. Van Campenhoudt et al. (2017) remark that field findings shall be organised in a grid of evidence. It is necessary to remark that this thesis does not focus only on ABD during Census operations. Other processes and surveys are also considered, such as price index surveys, unemployment rate surveys, inflation rate, and industrial production indexes. Thus, the documental analysis associated with the literature review supported the development of the interview script in two complementary ways. Using the documents from the NSOs, it was possible to elaborate the questions with the proposal to photograph the stages where each organisation was placed in the ABD journey: avoidance, initial adopter, medium adopter, and advanced adopter. Another relevant contribution from the documental analysis was understanding the relevant organisational roles to be interviewed. This diagnostic represented time-saving and focus gain to the research during the interviews stage. 4.4.2. Semi-structured Interviews The semi-structured interview is a main evidence source in social research (Van Campenhoudt et al. , 2017). The reason for this heavy use is the instrument's flexibility, which attributes some degree of freedom to the researcher in interacting with the respondent (Van Campenhoudt et al. , 2017). Based on Kvale (1996), those interviews tried to take a photograph of the world of these organisations and to understand the researched phenomena, ABD. Yin (2014) asserts that multiple case studies shall do a cross-check between each case conclusion to offer a generalisable result. In this study, beyond this, another triangulation took place. The cross-check between documents, and interviews was performed. 82 resources, suffered budget cut-offs, and was procrastinating the use of Big Data, this code/factor was verified as reinforcing the avoidance of Big Data. Legend: Saldaña's coding process. From the real/particular to the abstract/general perspective, promoting affirmations and theoretical advances based on codes and categories. Source: Saldaña (2013). Figure 13 – Coding Process by Saldaña (1993) Despite the extensive effort to analyse the interviews under the lens of each factor, the result was rewarding. A robust outcome was elaborated from those analyses with the evidence of elements strictly linked with ABD for each code/factor. In fact, those analyses brought summaries from each code/factor. Again, in the example of financial capacity , the result of analysis understood this code/factor impacts ABD in NSO because it is (i) a requirement for the adoption of Big Data; demands (ii) budget, cost, and stakeholder management; suffer pressures for (iii) considerable investments in technology and training; and forces the NSO to keep (iv) proactively seeking funding. The adoption effort probably fails if those elements are not achieved/executed. At the end of the analysis, a relevant result was a matrix developed based on the three TOE dimension, their respective factors and elements. This matrix allowed the research to verify if codes/factors were more critical to ABD than others. A code/factor might be irrelevant, relevant, or critically relevant. Just one factor/code presented irrelevant contributions to ABD in NSOs and was discarded, the IS fashion . Others were relevant but not critical, like “ improved design and 83 execution efficiencies ” and “ vendor support ”. Those codes/factors were identified in the answers of the interviewees, but their impacts were not significant for ABD. On the other side, if the NSO members remarked that a code/factor was massively relevant, as the financial capacity , it was considered critical. Following this analytical path, the matrix added an expressive value for this research: a summary of the findings related to ABD elements in the context of NSOs, as presented in Table 19. Table 19 – Design of the code/factors and elements Matrix Categories Factors/codes (FC) Elements (EL) Technology (T) T_001; T_002; T_003; … T_FC_01 (EL01, EL02; EL03,); T_FC_02 (EL01, EL02; EL03,); T_FC_03 (EL01, EL02; EL03,); … Organisation (O) O_101; O_102; O_103; O_104 Financial capacity; … O_FC_01 (EL01, EL02; EL03,); O_FC_02 (EL01, EL02; EL03,); O_FC_03 (EL01, EL02; EL03,); (Requirement for ABD; budget, cost, and stakeholder management; huge investments in technology and training; proactively seeking funding.) … Environment (E) E_FC_01; E_FC_02; E_FC_03; … E_FC_01 (EL01, EL02; EL03,); E_FC_02 (EL01, EL02; EL03,); E_FC_03 (EL01, EL02; EL03,); … This matrix of findings in the case study was based on the Eisenhardt (1989) proposal. According to the author, the researcher has to shape hypotheses and enfold the literature after the data analysis. Internal validity and measurability are necessary to sharpen the construction definition, also named the shaping hypothesis (Eisenhardt, 1989). About enfolding literature, comparing the findings with the literature reinforced the internal validity of the construct, the TOE framework. An expected outcome for this thesis is the resultant TOE version for ABD in NSOs. At this moment, the last step of the case study journey (Eisenhardt, 1989) occurs: “ reaching closure ”, when theoretical saturation is realisable. After the TOE’s adapted version for ABD was built, there were improvements in the theory. Those theoretical improvements are discussed in Chapter 7, “ Results ”. 84 Another finding from the analysis was the influences and linkages a code/factor deploys on or receives from each other. This result showed that TOE’s factors work under a web of interactions, where each relevant factor acts like a node of influence and trend. One example is the technological factor named “ technology resources ”. It is influenced by factors from the three dimensions and represents the results of added actions and practices executed by the respective NSO. Finally, it is necessary to remark that the multiple-case study method demands a generation of a repository with the research's records, files, prints, and videos (Yin, 2014). In this study, the adopted software, NVivo, followed the techniques proposed by Jackson and Bazeley (2019). Documental, interviews and Direct Observation analysis are processed by NVivo, version 12 for Microsoft Windows. Considering the described methodology, Table 20 summarises the features of the research design adopted in this thesis. It is a proposal to figure out how ABD takes place in NSOs under an exploratory perspective, with openness for future studies. Table 20 – Summarized Methodological Design Approaches Definitions Philosophical Realistic Research features Empirical and Exploratory Method Multiple Case Studies Data collections Documental research, structured interviews, direct observation Sampling INE-PT; IBGE-BR; ONS-UK. Target Group Presidents, Census managers, IT managers, and Regional Local Offices. Data analysis Content analysis, speech analysis and text mining Geographic horizon Transnational After the detailed description of the methodology, the next chapter presents the cases descriptions based on the documental analysis. The cases presentations supported the following interview analysis. 85 Chapter 5 – Understanding the National Statistics Offices This research studied the adoption of Big Data (ABD) in three National Statistics Offices (NSOs): the Office for National Statistics (ONS) in the United Kingdom, the Instituto Nacional de Estatística (INE) – Statistics Portugal - in Portugal, and Instituto Brasileiro de Geografia e Estatística (IBGE) – the Brazilian Institute of Geography and Statistics - in Brazil. Each organisation advanced in a singular path for adopting technologies and innovations. In the case of Big Data technologies and methodologies, the same happens. This chapter briefly presents the three NSOs' trajectories and features. After this, these organisations’ primary statistical operation, the demographic Census, is discussed under the comparison of the three NSOs. All of them had launched Census operations in the last three years, achieving different results. Next, about the operations deployed by the NSOs, the data generated by them is presented under the structure proposed by the United Nations Statistics Division (UNSD). Official Statistics include different topics beyond economics indices. Those data, surveys, indices, and rates represent the outcomes and deliveries that NSOs present to society. The last subsection approaches the studies about the association between Big Data and Official Statistics. Additional methods are necessary to advance ABD in NSOs. The following sections explain the description of ONS, INE, and IBGE based on a mix of documental references and testimonials drawn from interviews with members of the NSOs. A couple of questions about professional trajectory were applied in the interviews' scripts. The result is a rich view of the NSOs by their employees, with clarifying notes. Innovation cycle is one of the mentioned notes. It includes the perceptions from the NSOs’ employees about the tipping points where technological changes took place. Another important aspect of the cases is the Census operations management. In this section, the British, the Portuguese, and the Brazilian NSOs had their demographic Census scrutinized. The justification for this detailed examination is the relevance Census has to NSOs, as their major survey. Finally, the outputs and deliveries addressed by NSOs are presented in this chapter. The relevance of the outputs and their standardised structure signalises how the Global Statistical System influences the NSOs routines. 86 5.1. Features and Summarised trajectory The NSOs as the current organisations we know were created in different periods of the last one hundred years. In chronological order, INE is the oldest institute, followed by IBGE, and ONS is the most recent. However, the activity of Official Statistics in the three countries had previous operators before those three. The “ Instituto Nacional de Estatística ” of Portugal (INE) was created in 1935 (Valério and Bastien, 2010). It collects and disseminates Official Statistics to the Portuguese society, which totalises approximately 10 million inhabitants. This organisation generates economic and demographic data collected in the 18 districts of the mainland and the autonomous regions of Madeira and the Azores. Regarding the regions of Madeira and the Azores, both have dedicated structures that work with INE. Those governmental bodies are not independent NSOs, but members of the Portuguese Statistical System. " Diretoria Regional de Estatística da Madeira " (Diretoria Regional de Estatística da Madeira, 2023) and " Serviço Regional Estatístico dos Açores " (SREA, 2023) work as partners of INE in the Portuguese archipelagos of Madeira and Azores. INE is a public agency subordinate to the Minister of the Presidency (Instituto Nacional de Estatística, 2023). In the Portuguese case, this link means that the NSO has limited budget autonomy. Portugal is a member of the European Union. The 27 countries of the European Union, through their respective NSOs, are part of the EUROSTAT group (Eurostat, 2023). This set of NSOs composes a regional body tasked with maintaining and improving the quality of the Official Statistics of the member-states of the European Union. In 1938, IBGE was created by the fusion of two previous organisations: the Brazilian National Statistics Office and the National Geographic Council (Penha, 1993; Gonçalves, 1995). Since that genesis, IBGE has presented a singularity compared to other NSOs. Its operations included not only Official Statistics but also Official Geographic data (IBGE, 2023a). During those eightyfive years, IBGE has achieved a double challenge: deploying Official Statistics and managing all geographic information for the citizens of the fifth-biggest country in the world. 87 Related to the link with the Brazilian Federal Government, IBGE is a national agency subordinated to the Ministry of Planning (Ministério do Planejamento, 2023). This feature is remarkable because it justifies the limited autonomy that NSO has. Nowadays, IBGE provides information for approximately 210 million Brazilian citizens. It operates in 26 states and the Federal District, collecting data from 5570 Brazilian cities. In addition, this organisation disseminates demographic, geographic, social and economic data from Official Statistics. One remarkable feature in the IBGE trajectory is the capacity to develop outcomes about the complex Brazilian reality. This competence is proven by constructing experimental statistics to count the indigenous populations (IBGE, 2023b) and verify their conditions. Another example of the IBGE response capacity was a survey created to measure the impacts of the COVID-19 pandemic on Brazilians, named PNAD COVID-19 (Penna et al. , 2020). In the case of the British Office for National Statistics (ONS), it delivers Official Statistics for a group of countries integrated by England, Scotland, Wales and Northern Ireland (Office for National Statistics, 2023a). This organisation presents a singularity in comparison with its counterparts in Brazil and Portugal: it is a public-independent research institute with the status of a non-ministerial organisation (Office for National Statistics, 2023b). ONS has another distinctive feature, its shorter lifetime. Unlike INE and IBGE, the British NSO was created only twenty-seven years ago, in 1996. Pullinger (1997, p. 291) remarked that the " merge of the Central Statistical Office (CSO) and the Office of Population " created the ONS. Despite its relative youth, ONS absorbed the two hundred years of heritage on Official Statistics in the United Kingdom. Beyond the historical register, ONS works for different nations and cultures. This diversity of nationalities represents a context different from INE and IBGE, which work only for people in a respective country. As a non-ministerial public body, ONS reports only to the British Parliament and the UK Statistic Authority (UK Statistics Authority, 2023). This last organisation acts like a regulatory agency with the " objective of promoting and safeguarding the production and publication of official statistics that 'serve the public good " (UK Statistics Authority, 2023). In this way, it is relevant to verify in further analysis (Chapter 6) if this institutional design has some influence on the ABD process. Speaking of institutional structures, it is possible to inquire if an NSO, with more autonomy than a similar one, can adopt Big Data faster and more broadly. An unfolding result of the linking 88 structures may be the other TOE's element, government regulation . In a formal and centralised relationship, the regulations are based on laws and governmental acts. The Brazilian case can confirm this, and a quick view of IBGE's regulatory bases shows more than ten pages referencing official rules (IBGE, 2021). On the other hand, ONS presents fewer formal norms, as shown on the National Census webpage (Office for National Statistics, 2022a). In the same path, data regulation laws affect those organisations, and they are under distinct legal norms. INE-PT and ONS-UK are covered by The General Data Protection Regulation – GDPR (European Commision, 2022), and IBGE-BR is working under a domestic governmental norm, Personal Data Protection General Law - PDPGL (República Federativa do Brasil, 2018). 5.2. Outputs and deliveries During a search on the UNSD webpage (United Nations, 2023), it is possible to verify, in the menu "Topics", five Official Statistics main categories. In Figure 14 there is an illustrative scheme of those categories. The proposal presented an illustrative list of Official Statistics, not an extensive one. Legend: A summarized portfolio of the Official Statistics standards outputs/deliveries following United Nations Statistics Division guidance. Based on United Nations Statistics Division – Department of Economics and Social Affairs. Available at: https://unstats.un.org/UNSDWebsite/ Figure 14 – Official Statistics by topics – an exemplificative scheme. As presented above, four topics summarise the main Official Statistics developed by NSOs. Economics statistics support decision-making with data about business sectors, national accounts and trade statistics. In the opposite circle, Population and Society include demographic statistics such as Census data and migration. Environmental statistics is an emergent topic, 89 reporting data from water, waste, forests, and associated issues. Finally, the Development Indicators represent two sets of Official Statistics, including the "Sustainable Development Goals" (United Nations, 2022) and the "Millennium Development Goals" (United Nations Statistics Division, 2023d). Those two sets work as monitors to check how countries cooperate to promote a sustainable feature. All those surveys, indices and rates are generated by NSOs, including ONS, INE, and IBGE, as presented on their websites. A simple comparison between the NSOs' websites makes it possible to realise how standardised the Official Statistics portfolio is. This feature allows the exchange of practices and proceedings in the GSS network. One possibility is to spread the ABD among the NSOs. Those outcomes are part of the business architecture presented in Figure 4 (Chapter 1). Big Data technologies and methodologies can be adopted or avoided in this business architecture element of surveys. 5.3. Innovation cycles During the interviews, the members of NSOs were asked two questions about their professional trajectory. The first inquired about their trajectory in the respective institute, regarding their achievements and professional experience. Secondly, the members answered about their perceptions of the relevant innovations they had witnessed during their journey in the NSO. After collecting those questions, they are compared with the documental sources. The proposal was to verify if the described facts are also relevant to the respective NSO. In fact, in the three cases, the convergence was strong. Accordingly, the interviews and documents showed that ONS, INE, and IBGE work on innovation cycles. IBGE and INE presented similar innovation cycles to the demographic Census operation, occurring every ten years. This information was relevant to this research because it shows when those NSOs receive from the Governments the window of opportunity to update their technologies and processes. Members from both institutes subscribe to this fact. (In this research, the citations are presented with the code of the "interviewees" between brackets after the quotation marks.) " Another point that I would also like to draw attention to is the modernisation of data collection and verification. Related to the collection, the introduction of the CAPI was 90 fundamental because it made the collection quicker. CAPI gave more time for the interviewer to work on the approach process. " [06_IBGE] In 2010, during the demographic Census, IBGE innovated by adopting Computer Assisted Personal Interviewing - CAPI (Instituto Brasileiro de Geografia e Estatísticas, 2023a). Those devices substituted the sheet-form questionnaires for Census data collection. For IBGE members, this innovation represented an icon of the institute adopting new technologies. As a remarkable event, the members cite the prize won by IBGE from UNESCO for adopting CAPIs in Census 2010 (UNESCO, 2023). Notwithstanding, the interviews did not mention breakthrough innovations after the CAPIs adoption. " IBGE won the technology adoption prize (by UNESCO). So, I would say that this was the greatest technological advance that I have witnessed. Production of technologies in the collection of the Demographic Census. In collection in all senses: in collection management and in the collection itself, with the CAPIs. " [09_IBG_HUB] Regarding the advances of INE, there is no mention of international prizes. However, the members highlighted the use of electronic forms since the first decade of this century. In sequence, the Portuguese professionals pointed out a continuation of technological adoption. The data collected in digital form through INE’s website (Instituto Nacional de Estatística, 2011) started in the CENSUS 2011. " Here in Portugal, Census operations are always marked by some component of innovation. In 2001, the great innovation in the Census was optical reading and character recognition (...), which was a very big leap in what was done until then. (...) Then, after ten years, in the 2011 Census, one of the very innovative components that were done in Portugal quite successfully - and let me say that it is a pride to work in an Institute like this - refers to the introduction of the internet response in the Census. In 2011 we already introduced the internet response, and I can tell you that we achieved at the time a 50% response rate. " [04_INE_COC] The realised difference between IBGE and INE refers to the innovation cycle continuation. While IBGE's cycle appears to be interrupted, the Portuguese one maintained the effort to transform processes after the Census operations in 2011 (Instituto Nacional de Estatística, 2011). A particular innovation cycle was verified in the ONS case. Since its foundation in 1996, this institute deployed two Census, 2001 and 2011 (Office for National Statistics, 2022b). In 2016, 91 with twenty years of executed services, ONS suffered a structural reform. An act from the British Parliament determined a review of the ONS from a data supplier to a changing driver for other governmental bodies. This ambitious proposition was registered in the study developed by Bean (2016). This initiative resulted in relevant innovations. The significant output was the creation of the "Data Science Campus", also launched in 2017 (Data Science Campus, 2023). Interviewees reported this as a tipping point in the ONS trajectory. " They decided to establish a so-called Data Science Campus. So that not only their employees could be trained to understand the (data science) concepts, basic techniques, but also they want to attract some kind of external collaboration within the institution. " [01_ONS_PRO] In summary, the data collection revealed two types of innovation cycles for NSOs. The cycle associated with the Census operations is verified in IBGE and INE. Conversely, the innovation cycle provoked by external agents refers to ONS. In this case, the external agent is the major stakeholder, the British Parliament. 5.4. The Census operation The previous section identified the Census operation as a window for innovation for IBGE and INE. This attributed relevance is consensual among NSOs and supranational agencies, such as the United Nations Statistics Division (UNSD). Best practices and challenges are discussed in forums organised by the UNSD (United Nations Statistics Division, 2023b). The International Committee on Census Coordination (ICCC) coordinates those forums and the exchange of information. According to the UNSD (2008), the demographic/populational Census means: "(…) the total process of collecting, compiling, evaluating, analysing and publishing or otherwise disseminating demographic, economic and social data pertaining, at a specified time, to all persons in a country or in a well-delimited part of a country." ICCC runs recommendations and best practices for NSOs to deploy the demographic Census. Those prescriptions are necessary, as this is the major Official Statistics operation performed by NSOs. UNSD (2008) presents best practices for Census operations. Some examples of those guidelines are (i) definition of the topics to be investigated, as immigration and fertility; (ii) 98 Related to Big Data technologies and methodologies, INE adopts Artificial Intelligence with the use of machine learning algorithms to support citizens in filling out forms. As the INE representative cited, this adoption represented an innovation for the Portuguese institute. " Now, in 2021, there was a reinforcement of this response component. There was the introduction of autocomplete, which, as the person writes, we suggest entries for the text in the open questions. We had the use of machine learning in the coding of the descriptives, and one of the other innovations was the incorporation and the possibility of getting information from the administrative records. " [04_INE_COC] 5.4.3. The Census in Brazil Since 2020, Brazil has been one of the countries most affected by the COVID-19 pandemic. Behind the United States, with 1,111,342 deaths, Brazil is the second country in total lost lives (WHO, 2023), with 699.276 deaths, during the period from 2020 to 2022, as presented in Figure 19. This externality has impacted different Brazilian business sectors, including the Official Statistics. Legend: Global map categorizing, by collors, the total of deaths by Coronavirus (COVID-19) in the World (2020-2023). Source: World Health Organisation (WHO). In: WHO Coronavirus (COVID-19) Dashboard. Extracted from: https://covid19.who.int/ Figure 19 - Total of Deaths by Coronavirus (COVID-19) in the World (2020-2023) Initially, the Census had the original plan to begin in 2020. Because of COVID-19, it was postponed to 2021 (IBGE, 2020b). A new delay occurred in 2021 (IBGE, 2022a). The justification was based on the Brazilian Government's allegation that there were insufficient funds to deploy the Census. Finally, IBGE launched the Census on 2022 August 1st (Gandra, 2022). 99 Unlike ONS and INE, the Census deployed by IBGE focuses mainly on direct data collection based on the work of hired data collectors. To advance this strategy, IBGE disseminated visual content to promote the engagement of Brazilians. Two examples are the posters presented in Figure 20. Legend: Informative posters to disseminate the Census to the Brazilian citizens. Source: IBGE (2022). Dissemination material. Extracted from: https://censo2022.ibge.gov.br/pecas-dedivulgacao/pecas-de-divulgacao.html, on 09/01/2023. Figure 20 – Dissemination posters for the Brazilian Census 2022 Related to technological adoptions, no reference was mentioned on the IBGE website. Just the use of the CAPIs is presented in the dissemination materials. Figure 21 shows the set of materials used for data collection by Census professionals. This set includes one cap, one waistcoat, and a CAPI. The data collection suffered delays because of the reduced number of Census professionals (Almeida, 2022). Originally, IBGE had planned to hire 183,000 professionals. During the peak of the collection, 120,000 data collectors were working. This deficit impacted the conclusion of the Census. In this scenario, the dissemination of the final results. scheduled for April 2023 (Silveira, 2023), was yet to occur. On 2023, June 28th, finally, IBGE broadcast the initial results from the last demographic Census. This new populational counting presented around 203 million Brazilian citizens. For this doctoral 100 research two numbers are relevant: the percentage of non-responses and the number of surveys deployed by Internet, in an experimental mode. Five percent of the Brazilians refused to answer the Census survey. Although the proportion of 95% of responses can be understood as a strong value, the non-responses are significative. On the other side, the number of responses by Internet was reduced when compared with the door-to-door interviews. Just 0,56% of the Census questionnaires were answered using the internet. The majority of the data collection, 98,88%, took place using face-to-face interviews. Another residual percentage, 0,56%, was executed by telephone call interviews. Legend: Data collectors’ uniform on the left. CAPI model used in the Brazilian Census on the right. Source: photos kindly provided by IBGE staff. Rio de Janeiro, 2022 August 1st. Figure 21 – Data collectors' work kit. In summary, the delays suffered by the Brazilian Census are a strong warrant for others NSOs. When the three cases are compared, INE and ONS advanced with their Census, and IBGE just concluded it in 2023, after three years of attempts. The focus on data collection by the internet associated with ABD allowed INE and ONS to overcome the impacts of the pandemic. On the opposite side, Brazil insisted on the same formula of the Census 2011, despite the adverse context. Maybe if IBGE had adopted a hybrid strategy as INE, the Census would be concluded on time, consuming less resources from the taxpayers. 101 Chapter 6 - Data analysis The adoption of Big Data (ABD) by NSOs can be verified in the operational process changes to deploy Official Statistics. Additionally, ABD can also be verified in initiatives to promote the use of Big Data technologies and methodologies in and out of the NSO. One example of the effort to associate Big Data with Official Statistics is the creation of Hubs in the Global Statistical System (United Nations Statistics Division, 2023e). Those entities have the mission to promote and spread the use of Big Data for NSOs in specific regions. Currently, there are four Big Data Hubs in the world, located in four regions on four continents: Brazil (America), Rwanda (Africa), China (Asia), United Arab Emirates (Middle East). Each Hub is supported and managed by the NSO in the respective country. In this way, the Brazilian Hub receives support from IBGE (Instituto Brasileiro de Geografia e Estatísticas, 2023b). The opened catalogue of courses (United Nations Statistics Division, 2023a) to teach Big Data techniques or the creation of different Hubs are examples of adoption acts. In fact, those acts impact the TOE model on its Organisation and Environment dimensions. Adding a new business unit into the IBGE structure, as the Hub is, represents a change in the organisational processes and interactions between departments. On the other hand, the environmental dimensions are affected by new interactions between the Hub host and the other regional NSOs. Related to the statistics and ABD, this interaction figures in the Technology dimension. According to Allende-Alonso et al. (2019), a Big Data set delivers unnormalised data. This feature demands that Statisticians and Data Scientists use additional statistical methods to interpret and extract value from massive datasets. Indeed, this journey to develop Big Data methodologies for Official Statistics is complex and represents an ongoing challenge. Bollineni-Balabay et al. (2015), van den Brakel and Kreig (2016), Van Den Brakel et al. (2017), and Schiavoni et al. (2021) guided researches to develop or adapt statistics tools and methods to incorporate Big Data in the Official Statistics productions. In this way, for NSOs, it is necessary to use the expression Big Data technologies and methodologies. A mandatory link between the statistical methodology and the technology allows to collect, store, process and analyse the data. 102 Next, the sections and subsections present the analyses of the interviews with ABD in perspective. This adoption refers to technologies and methodologies for Big Data applied in the Official Statistics processes. 6.1. How National Statistics Offices adopt Big Data - TOE factor analysis During one semester, the interviews with ONS, IBGE, and INE members ran in a systematic way. All interviewees demonstrated a thorough understanding of Official Statistics as well as a genuine perception from the adoption of Big Data (ABD). Aside from the fact that the interviews took around forty minutes, each piece of work used by them was precise and carefully selected. The result was relevant raw data to be analysed as proposed in the Chapter 4. Interviewees did not assign the same importance and attention to the different TOE factors. The most cited factors presented more quotations and descriptions than the others, with fewer mentions. In fact, the factors that drive ABD in NSOs do not run in a flat and balanced way. So, the analysis advances in a systematic mode, following the lists of factors per dimensions, with the proposal to bring a basic order in the sets of factors and simplify the localization of each factor in this chapter, like a catalogue, as presented in Figure 22. According to the original proposal of DePietro et al. (1990), TOE should be, initially, approached from the Technology dimension. Hence, the Technology dimension is the first to be analysed, followed by the Organisation and Environment dimensions. 103 Legend: Continuous Rectangles depict TOE factors found in the Literature review. Dashed Rectangles depict emergent TOE factors resulting from this research. The colours represent the three TOE dimensions: Environmental factors are brown; Organisational factors are green; Technological factors are orange. Arrows represent influences from one factor to each other. Figure 22 – TOE factors for the Interview analysis. 104 6.1.1 Technology factors The literature review presented fourteen TOE factors in the Technology dimension related to the adoption of Big Data (ABD). As presented in Chapter 4, this set of factors was used as a code book to analyse the content of the interviews (Saldaña, 2013). During the interview analysis, those fourteen factors were validated by members of the three NSOs (ONS, INE, IBGE). However, some factors presented validation and high relevance, while others were validated, but showed less importance. Additionally, six original factors emerged from the content analysis specifically related to ABD in NSOs: " obsolescence ," " Big Data technologies and methodologies ," " data restrictions ," " polysemy and noise ," " adoption design for Big Data ," and " Big Data types ". Clarifying, Figure 22 presented those emergent factors in the yellow dashed rectangle, and in the continuous one, the factors from the literature review. Each new factor was defined to support the respective content analysis. The first technological emergent factor, “ Obsolescence ”, means the incapacity of the traditional methodologies and technologies to address current issues in Official Statistics (OS), such as large volumes of data, randomness, and imputation. Second, “ Big Data Technologies and Methodologies” include descriptive characteristics of Big Data technologies and methodologies used by NSOs. Thirdly, " Data restrictions " refers to every condition, norm, economic feature, or technological issue that restricts the access, processing, use, and broadcast of a data source or data set required by an organisation engaged in BD adoption. In regard to “ Polysemy and noise ”, it is related to the different interpretations about what Big Data for Official Statistics means. The emergent factor named “ Adoption design for Big Data ” refers to the plans and actions designed for a successful adoption of Big Data. Last, but not least, “ Big Data types ” define the data sources and types used in ABD processes in NSOs. During the analysis, each Technology factor presented relevant elements about ABD in the three NSOs – ONS, INE, and IBGE. Further, the analysis of each factor, related to ABD in NSOs, includes the respective definition, the elements/items that integrate the factor, illustrating quotes from the interviews and respective comments. As the anonymity of the interviewees was guaranteed, each quote is identified only with the institute of the interviewee, between brackets. Finally, in the end of each factor analysis, there is a summarized list of integrating elements/items. This list informs the findings of this research 105 related to the respective factor. Like with the quotes, each item has the names of the respective NSO between brackets. In this way, it is possible to compare and verify in which case each item was identified, allowing the reader to compare the three cases. For this purpose, Table 21 has the objective to show where the respective item was more salient. In this context, the first column informs the factor name. In sequence, column two shows the consolidated findings. Inside this column, each element has an order number and brackets with the respective NSO where it is manifest. Some elements were observed in the three NSOs. Others just in one or two cases. Finally, the last three columns inform the presence of each element in the respective case, following the order number from column two. So, there is the synthesis of the results from the Technology dimension in Table 21. After Table 21, the detailed analysis of each Technology factor is presented. 106 Table 21 – Technology factors: Summarised findings Technology Consolidated Findings per case Factors Findings ONS-UK INE-PT IBGE-BR 001.Compatibility 1. Why is compatibility important for ABD in NSOs? Because the adoption must integrate Big Data to the internal organisational processes and information systems. [ONS, INE, IBGE] 2. How does compatibility influences ABD? By meeting specific requirements: a. Managing long-term projects. [INE] b. Alignment of technological and methodological procedures to deploy new outcomes for Official Statistics. [ONS, INE, IBGE] c. First milestone: compatibility of traditional and new data sources. [ONS] d. Following NSOs calendar: Big Data outcomes inserted in the Official Statistics (OS) schedule. [ONS] 3. Benefit of the compatibility to ABD: the comparability between NSOs, improving the benchmark. [IBGE] 1; 2b; 2c; 2d. 1; 2a; 2b. 1; 2b; 3. 002.Complexity 1. Why complexity for ABD: avoidance of a potential barrier for the adoption. [ONS; INE; IBGE] 2. How does complexity influence ABD? It affects ABD initiatives when: a. The replacement between Big Data methods and survey methods is too complex. [ONS] b. Coverage issues arise: “Just because it’s big, it doesn’t mean that its coverage is right.” [ONS] c. Methodological challenges include the imputation and missingness. [ONS] 3. Benefit of reducing complexity: Merging of the Census operation and administrative Census. [INE] 1; 2a; 2b; 2c. 1; 3. 1. 003.Cost of adoption 1. Why and how does cost of adoption impact ABD? Because it includes cost elements necessary to ABD, such as: a. ABD is expensive, complex, and risky. [IBGE] b. Cost drivers: technology and digital tools, infrastructure, training, and datasets access. [ONS, INE, IBGE] c. Specific budget for ABD. [ONS, INE, IBGE] d. Training costs: a critical item associated with the lack of skills. [ONS; IBGE] 2. Requirement: dealing with high costs in earlier stage of adoption. [ONS, INE, IBGE] 1b, 1c, 1d, 2. 1b, 1c, 2. 1a, 1b, 1c, 1d, 2. 004.Data integration 1. Why is data integration relevant for ABD? Because it analyses the complementarity between Big Data and traditional surveys. [ONS, INE, IBGE] 2. How does data integration influence ABD? Big Data and surveys work in a complementary way. Exemplifying: administrative records used as complements to the OS. [ONS; INE; IBGE] 3.Benefits: a. More precise measurements with real-time data collection. [ONS] b. Ability to use different types of data, including semi-structured and unstructured. [ONS; IBGE] 4. Outcome: a nationwide administrative database. [INE] 1; 2; 3a; 3b. 1; 2; 4. 1; 2; 3b. 107 Technology Consolidated Findings per case Factors Findings ONS-UK INE-PT IBGE-BR 005.Data quality 1. Why is data quality relevant for ABD in NSOs? Because OS demand reliable and consistent data to deliver the outcomes. [ONS, INE, IBGE] 2. How does data quality influence ABD? The influence advances by the quality requirements demanded to adopt Big Data as inputs to the OS. [ONS; INE; IBGE] 3. Requirements: a. Rigorous quality for ABD projects. [ONS; INE; IBGE] b. Achievement of quality standards, as harmonisation, before the integrations with OS. [ONS; INE] c. Validation of administrative records before integration with traditional data sets. [ONS; INE] d. Quality tests to verify the quality of big datasets. [ONS; INE; IBGE] 1; 2; 3a; 3b; 3c; 3d. 1; 2; 3a; 3b; 3c; 3d. 1; 2; 3a; 3d. 006.Internal versus external technologies 1. Why is the internal vs external technologies factor relevant for ABD in NSOs? Because the origin of the technology can promote cost and time savings. [ONS, INE, IBGE] 2. How does the internal vs external technologies factor influence ABD? The technological strategy defines the origin of the technology and the type of adoption. [ONS; INE; IBGE] 3. Results: combination amongst internal, external, and hybrid solutions via consortia. [ONS] 4.Requirements: a. First stage: attempt to use internal solutions. [ONS; INE; IBGE] b. Second stage: call for outsourcing. [ONS; INE; IBGE] c. Suppliers: Universities, Research Centres, private companies. [ONS; INE; IBGE] 1; 2; 3; 4a, 4b, 4c. 1; 2; 4a, 4b, 4c. 1; 2; 4a, 4b, 4c. 007.Perceived benefits 1. Why is the perceived benefits factor relevant for ABD in NSOs? Because NSOs' top managers base their decisions on the perceived benefits of the Big Data for OS. [ONS, INE, IBGE] 2. How does the perceived benefits factor influence ABD? This factor works as a promoter for ABD through: a. The potential capacity to manage sorts of data storage as a service (cloud). [ONS] b. The possibility to model and deploy innovative data collection processes, as using mobile devices. [ONS] c. The option for transitioning to use on time data streaming (data sources) from retail chains. [ONS] 3. Requirements: a. Plan and execute the adequate Big Data project for the NSO reality. [ONS] b. Dedicated budget as a support for ABD projects. [ONS] c. Organisational redesign with a dedicated Big Data unit. [ONS; INE] 4. Benefits: a. The improvement in the reporting time. [ONS; INE] b. Better, faster, and more precise OS outcomes. [ONS; INE] c. Data quality improvement: accuracy and granularity gains. [ONS] 1; 2a; 2b; 2c; 3a; 3b; 3c; 4a; 4b; 4c. 1; 3c; 4a; 4b. 1 114 d. Following NSOs calendar: Big Data outcomes inserted in the Official Statistics (OS) schedule. [ONS] 3. Benefit of the compatibility to ABD: the comparability between NSOs, improving the benchmark. [IBGE] 002. Complexity Complexity is “the characteristics of Big Data that are perceived as being difficult to understand and use (e.g., the difficulty of learning related knowledge for employees who will use Big Data applications),” as defined by Sun et al. (2018a, p. 198). Employees from an NSO must deal with complexity because this factor, if not addressed, can be a barrier to ABD. If complexity becomes too great, ABD can be abandoned by the NSO. This dropout includes incompatibility between Big Data methods and survey methods, or methodological challenges, as imputation and missingness. In this way, the factors of compatibility and complexity work together, as presented in citations, to ensure that ABD in NSOs is not a simple phenomenon. In the interviewees’ words, ABD is associated with technological and methodological procedures to deploy new outcomes from NSOs. The difficult mentioned in the definition of complexity is part of ABD in NSOs context. In fact, there is a demand to make a replacement between Big Data methods (innovative) and survey methods (traditional). Another challenge for ABD in NSOs includes dealing with imputation and missingness during the data collection. It represents a methodological frontier to be expanded when an NSO is working with large and non-normalized amounts of data. As highlighted by a member of ONS, Big Data is important and necessary. However, it does not have a complete coverage of the country and society. "Just because it’s big (data), it doesn’t mean that its coverage is right. (…) Nobody yet, including Stats Nederland, has got to a position where the perfect replacement between Big Data methods and the survey methods is achieved. We have not gotten in place at this moment.” [02_ONS] The adoption of data lakes and data streaming to generate Official Statistics demands the convergence between administrative records with surveys, in a mandatory way. An example is the Portuguese Census operation, where the use of administrative records filled up gaps in the traditional Census data collection procedures. 115 So, the complexity of ABD is an obstacle to be overcome. It is complex because some results from Big Data are not originally adequate to Official Statistics, requiring imputation and missingness treatment. Below, the interviewee cites this methodological challenge. “It is the big methodological challenge. We’ve also started thinking now about imputation. So, one of the big challenges that we go with Big Data is that the target population from which we got the data that isn’t the same as the target population we’re interested. Our challenge is to understand the missingness, and how to address it..” [02_ONS] By the way, the adoption of Big Data technologies and methodologies involves distinct complexity levels. From a wider perspective, complexity appears in the Portuguese Census operations regarding the association between Census data collection and the administrative records. This association focuses on the creation of a Census based on the administrative data, an administrative Census, with regular updates in the demographic database, such as in emigration and permanent resident population counting. "In addition, these administrative bases, which is not only for the administrative Census, but it's the backbone to expand our production of statistics related to households, social statistics..." [01_INE] “We have confirmation that it is feasible to adopt an administrative Census methodology.” [04_INE] For ABD in NSOs, the complexity factor is summarised as: 1. Why complexity for ABD: avoidance of a potential barrier for the adoption. [ONS; INE; IBGE] 2. How does complexity influence ABD? It affects ABD initiatives when: a. The replacement between Big Data methods and survey methods is too complex. [ONS] b. Coverage issues arise: “Just because it’s big, it doesn’t mean that its coverage is right.” [ONS] c. Methodological challenges include the imputation and missingness. [ONS] 3. Benefit of reducing complexity: Merging of the Census operation and administrative Census. [INE] 003. Cost of adoption 116 Considering Baig et al. (2019), Sun et al. (2018b), Bakici et al. (2022), the cost of adoption refers to the financial resources demanded and executed expenses related to ABD (e.g., the costs of using Big Data technologies, the large initial investment required to embrace Big Data). According to Bakici et al. (2021, p. 11), " when Big Data tools are expensive, they would be more difficult to adopt by companies that do not have sufficient resources. " So, cost of adoption impacts ABD, imposing high costs for earlier stages of adoption. In fact, cost of adoption is critical to ABD because this factor includes the cost drivers of this innovation and remarks the necessity to a dedicated budget and training costs management. Those elements are mandatory for an NSO to adopt Big Data technologies and methodologies. Without those cost elements, the NSO does not accesses knowledge and technological resources for ABD. ABD is expensive, complex, and risky. It is expensive because many technological applications and infrastructures are not free, such as cloud storage or streaming. Another cost driver is the access to external datasets developed by other companies different from the NSOs. In this example, the cost is a restrictor because innovative methods for delivering experimental statistics can demand those external data sources. As a restriction, the limitation of datasets access means complexity and risk. " And just to conclude, Big Data is awfully expensive. You know, it's not something that is going to be cheaper in the next few years. It will cheapen in the medium term.” [02_IBG] The cost of Big Data is so significant that ONS is considering allocating a specific budget for its implementation. So, the cost of adoption includes the budget dedicated for acquiring Big Data technologies. The progress of the adoption process is determined by budgetary capacity, which is linked to the financial capacity of the respective NSO. “ I think it is important to think the importance of budgetary spending. Big Data sources are hard to find. The technologies, if we want to develop from scratch, will take a lot of effort and time, and money as well. So, at the current stage, I think ONS wants a budget solution for Big Data. If we opt for a small budget, quality will sometimes be sacrificed, which will cause some problems. As a governmental body, ONS deals with a background of tight budgets. I think it is always all over the world. I think this kind of question probably is universal.” [01_ONS] 117 “So, ONS wanted us to do the research so that they can improve the quality of their statistics because they are different layers of quality based on their budget." [01_ONS] “Budget is always an onerous. We never have enough money to do everything that we would like to do and that is a good thing.” [03_ONS] Another relevant cost driver for ABD in NSO is training. The training cost is a requirement to overcome the lack of Big Data skills in the staff of an NSO. This item is a cost-intensive and critical part of the adoption process. It is necessary to train a large number of employees to initiate the adoption process inside NSO departments. Cost management is an obligation to advance in ABD. This type of adoption demands financial resources on a large scale. Indeed, ONS is an example of costs of training reverted in Big Data skills. “The training effort deployed by ONS was something quite strong. The ONS launched and funded the master’s in data Analytics of Government. I was the manager of that training program. There was also a trainee program. Everything was funded. This ONS’ effort is something I consider brilliant. (…) I contracted various universities according to the strategic goals that we wanted to achieve.” [05_ONS_PRO] For ABD in NSOs, the cost of adoption factor is summarised as: 1. Why and how does cost of adoption impact ABD? Because it includes cost elements necessary to ABD, such as: a. ABD is expensive, complex, and risky. [IBGE] b. Cost drivers: technology and digital tools, infrastructure, training, and datasets access. [ONS, INE, IBGE] c. Specific budget for ABD. [ONS, INE, IBGE] d. Training costs: a critical item associated with the lack of skills. [ONS; IBGE] 2. Requirement: dealing with high costs in earlier stage of adoption. [ONS, INE, IBGE] 004. Data integration Based on Park and Kim (2021, p.287), data integration refers to the degree to which the data is analysed by the Big Data systems, considering the relevance of Big Data applications and the integration of those with the data collected. Under this concept, data integration for ABD in NSOs refers to accessing and processing administrative records and aggregating them for 118 complementing Official Statistics outcomes. A successful data integration allows the NSO to delivery more precise measurements, and deal with different types of data. Data integration is critical to ABD because this factor analyses the complementarity between Big Data and traditional surveys. Indeed, ABD becomes a reality when it is integrated into the NSO's current processes and deliveries. "It is a harmonization. Actually, it is a complementation of registers. Like administrative records that you can use in association with the survey, and you can join other data to update the Official Statistics." [07_IBG_COC] Data integration refers to the complementarity between Big Data and surveys, associating methodologies and technologies. In the initial stages, Big Data was understood as a disruptive hype. However, when the advancement of ABD took place, it became clear that this innovation came to improve data quality for the current Official Statistics, not substitute them. “(…) When I joined ONS, their “mantra” was that surveys were in the last year. That sonly we wouldn’t need surveys anymore. That all we would need would be Big Data, transactional data, administrative data… and that we would be able to produce even better statistics than we were producing before. Since then, I think we’ve accepted that it was a very naïve approach…kind of a data strategy. Now, our understanding is that we need integrated data assets and administrative records. Basically, to pull together administrative records, transactional data, survey data…or, also called Big Data whatever. “ [02_ONS] Data integration is critical to ABD because a new developed dataset has to be compatible with the traditional data structure in the NSO's operations. In this process, measurements become more precise using real-time data collection. This approach represents an improved quality for NSOs’ data collection. Therefore, a more precise and faster data collection process means an improvement in the service level of NSOs. This improvement positively impacts organisational performance. “I think that they are starting to use alternative data sources to see whether they can measure the economic activity more precisely, more in a real-time sense.” [01_ONS_PRO] 119 The changes embraced by NSOs have delivered recognised benefits, including the ability to use different types of data and data sources. Related to the types, unstructured and semi structured data can be added to complement the Official Statistics outcomes. “The ability to use any type of data whether it is survey data or large volume, administrative records, or unstructured text data, such as documents, or unstructured data, such as images and video. “ [03_ONS] “Now, I can see another part of the Big Data, the real-time processing, referring to unstructured statistics, using the famous NoSQL." [01_IBG] "IBGE works essentially with structured data. So, there is a very interesting and very important movement in IBGE to interact with government agencies and use the administrative records made by other government agencies" [04_IBG] Referring to data sources, administrative records, videos and images represent the target source to be explored with focus on the Official Statistics improvement. Data integration and the organisational data environment are benefited from the new internalised knowledge. One example is the gain of data quality when INE adopted administrative records in the Portuguese Census. Actually, a Portuguese administrative database has been developed with a nationwide coverage. “We have the big project to change from the traditional Census to the Census based on an administrative base.” [01_INE] For ABD in NSOs, the data integration factor is summarised as: 1. Why is data integration relevant for ABD? Because it analyses the complementarity between Big Data and traditional surveys. [ONS, INE, IBGE] 2. How does data integration influence ABD? Big Data and surveys work in a complementary way. Exemplifying: administrative records used as complements to the OS. [ONS; INE; IBGE] 3. Benefits: a. More precise measurements with real-time data collection. [ONS] b. Ability to use different types of data, including semi-structured and unstructured. [ONS; IBGE] 4. Outcome: a nationwide administrative database. [INE] 120 005. Data quality Based on Lai et al. (2018a, p. 682), data quality means the degree to which the data needed for Big Data applications are accessible, consistent, and complete. This factor is critical for ABD because Official Statistics demand reliable and consistent data to deliver their outcomes. So, data quality includes the requirements for the adoption of datasets related to Big Data. In fact, the quality of data is a fundamental element in keeping the Official Statistics trustworthy and reliable. In the case of NSOs, ABD follows this assumption. All the three cases studied – ONS, INE, and IBGE – presented a solid commitment with data quality for their current Official Statistics outcomes. When ABD is the focus, firstly ONS, secondly INE and IBGE, demonstrated the same rigour in the data quality criteria as used in Official Statistics. Regarding the process to verify the quality of data derived from Big Data applications, the NSOs deploy a quality test applied in a current Official Statistics outcome. If the Big Data outcome is approved in this control test, it is integrated into the Official Statistics portfolio of the respective NSO. An example is the administrative records and the harmonisation methods they are submitted to in order to achieve the quality patterns of Official Statistics. This is the reason why there is a segmented run by NSOs named “experimental statistics”, where outcomes (data) derived from ABD are tested and validated. Those results can be evolved to Official Statistics just after the quality criteria are achieved by the Big Data results. “(…) I’m talking about quality of data already. But it is very common the statistical challenge around understanding these data sources. And the last challenge is how do we ensure, publish, disseminate, explain why we use Big Data. How do we create a high-quality result from a source that can be organic? Maybe collecting for administrative purposes. Maybe from sources that we have no control over like satellite imagery. It is a particularly important part of our work to ensure that the user trusts our work until it is concluded. So, it is just thinking about public trust and what we need to do to keep people informed.” [ONS_02] In this context, the new methods used by NSOs have to maintain the data quality of the Official Statistics. Consequently, by advancing ABD in NSOs, it is mandatory to achieve established quality standards. When the Big Data outcome is working on those standards, NSOs understand Official 121 Statistics acquired an improvement in their accuracy and updating capacity. This is one of the more realised benefits of ABD in NSOs. This benefit comes with “ the use of better-quality data sources ” [ONS_02] associated with other elements like: (i) “the ability to bring administrative data into our production technologies for routine statistical production;” [ONS_02] (ii) “to incorporate these sorts of administrative data into our systems;” [ONS_02] (iii) “ to collect and use to being a clean data set that is nationally representative" [ONS_02] (iv) "(…) to validate the entire methodology of that data, how that data was collected and worked on, and sometimes this is not very clear. So, I need to validate that these data are really reliable and useful to complement the Official Statistics from IBGE.” [04_IBG] Those data quality advances are associated with the validation of administrative records. Just when the data quality requirements are achieved, are the administrative records able to be integrated into datasets for Official Statistics generation. "It is a mode. You can call it “collecting”. But the source is not a survey. The source is an administrative record. (...) Once again, the example of the Tax Authority: it extracts information and sends this information to us, many times without metainformation. There is no metadata; there is nothing. We have to create that. We have to clean that up, put those bases in a way that is workable." [01_INE] An example of Big Data technology submitted to the data quality rigour, is the Web scraping used for prices index calculation. This technology adoption involves complexity, considering the level of noise in the extracted data. It is relevant but, from the interviewee’s perspective, not very useful yet. This technology is in an ongoing process to be better applied with significant data quality patterns for the Official Statistics. “(…) alternative price data is quite messy. We all know Big Data is good, but the messiness is very difficult to tackle, and this is another area that we try to build the capacity. So, two areas (AI and web scraping), as the technological front, that I am personally evolved into the technology sets. “ [01_ONS] For ABD in NSOs, the data quality factor is summarised as: 1. Why is data quality relevant for ABD in NSOs? Because OS demand reliable and consistent data to deliver the outcomes. [ONS, INE, IBGE] 122 2. How does data quality influence ABD? The influence advances by the quality requirements demanded to adopt Big Data as inputs to the OS. [ONS; INE; IBGE] 3. Requirements: a. Rigorous quality for ABD projects. [ONS; INE; IBGE] b. Achievement of quality standards, as harmonisation, before the integrations with OS. [ONS; INE] c. Validation of administrative records before integration with traditional data sets. [ONS; INE] d. Quality tests to verify the quality of big datasets. [ONS; INE; IBGE] 006. Internal vs External technologies Baig et al. (2019, p. 9) define this factor as “the hardware and software technologies used for Big Data adoption. Technologies revealed by retailers are internal technologies, whereas those provided by vendors are called external technologies.” The origin of technology for ABD is relevant because it can represent cost and time savings to the respective NSO. Internal, external, and hybrid sources are selected by the NSOs based on their technological strategy and the suppliers’ availability. Moreover, there are stages to be followed. The initial work is to attempt an internal solution. If it is not possible, calls for outsourcing projects take place. During the research, it was possible to verify that ONS demands more external and hybrid solutions than INE and IBGE. The justification is that British NSO has advanced in the initiatives and projects related to ABD more than the others. “Yeah, Both! So, I think you must have a mixture, a hybrid. The cloud providers … you know… Amazon, Google, Microsoft, and others… they are the sorts of groups that we would be talking about for deploying and supporting the work of specific applications for us. Many of those applications are Statistical Office depending specialised. So, they tend to be developed in-house. And lot of other things, like survey tools, that we develop in partnership with all the other Statistical Offices. So, a mixture of them all.” [03_ONS] Additionally, other interviewees reinforced this mix of internal, external, and hybrid solutions. The IT strategy emerges as a driver of the ADB process, indicating what needs to be internal and what 123 can be external or hybrid. Regarding this strategic guidance, it is noteworthy to highlight the partners and suppliers, such as universities, research centres and private companies. Differently from the ONS and IBGE, INE presented a degree of openness for external solutions mainly when the supplier is a Portuguese University. In fact, this strategy signals an effort to promote the development of Portuguese technology. "As for machine learning in specific, it had an in-house part but with collaboration from the University of Porto. In the development of methodologies, the academia has been requested. But it doesn't mean that there isn't also a collaboration with companies." [04_INE] One example of the hybrid effort with partners is the ESCoE, advanced by ONS. The Economic Statistics Centre of Excellence (ESCoE) is a consortium led by ONS, where universities and research centres not only pursue technology but methodology topics as well. "So, part of the solution might be developed from the internal research team, especially the capacity built by the Big Data campus. (…) But from my understanding, the innovative methodology is not developed by them. Basically, they try to borrow the technology, try to learn the language developed by other experts and areas to see whether they can use the algorithm, and test the algorithm do themselves to see the impact. So, that's one way of doing things. It's very, very expensive to develop something new. And secondly, for the external collaboration with their internal projects that to be supported by the so-called ESCOE. They want to see the external researchers' output as they want external researchers to devise algorithms to develop the methodology for them to use their data. So, these are two ways for the technology and for the algorithm, and for the methodology.” [01_ONS] For ABD in NSOs, the Internal vs External technologies factor is summarised as: 1. Why is the internal vs external technologies factor relevant for ABD in NSOs? Because the origin of the technology can promote cost and time savings. [ONS, INE, IBGE] 2. How does the internal vs external technologies factor influence ABD? The technological strategy defines the origin of the technology and the type of adoption. [ONS; INE; IBGE] 3. Results: combination amongst internal, external, and hybrid solutions via consortia. [ONS] 4. Requirements: a. First stage: attempt to use internal solutions. [ONS; INE; IBGE] 130 “Brazilians need to understand IBGE is the coordinator of the National Statistical System, and IBGE has a policy focused on privacy and security.” [03_IBG] Security and privacy concerns are present in innovation initiatives, like with ABD. ONS welcomes proposals since they prioritise security and privacy. “But we also allow people to highlight to us where their ideas are. Obviously, those ideas need to meet the security requirements. Nothing gets happened if it doesn’t meet security requirements.” [04_ONS] For ABD in NSOs, the Security and Privacy factor is summarised as: 1. Why is the security and privacy factor relevant for ABD in NSOs? Because it can be considered a hindrance for ABD. [IBGE] 2. How does security and privacy influence ABD? This factor is perceived as a daily, expensive, and inglorious struggle, considering the risk of cyberattacks. [IBGE] 3. Requirement for security and privacy: corporate policy and programs focusing on prevention. [ONS; INE; IBGE] 011. Technology competence According to Sun et al. (2018a, p. 198), technology readiness or technology competence is verified when the “organization has sufficient internal Information Technology expertise and technological infrastructure to adopt Big Data (e.g., IT knowledge and skills within the organization).” In the context of ABD in NSOs, technology competence is essential because this factor enables the NSO staff to require/develop Big Data technologies and methodologies, and at last, to design the adoption process. This factor can be understood as an outcome of the employees’ skills and process improvements. ABD demands from an NSO the hiring of top-skilled professionals or internal skills development. Those skilled employees spread their knowledge into the NSOs departments through the exchange of expertise about Big Data. As a result of this spill over, the NSO staff acquires technology competence for ABD. Not only IT professionals, but also specialists in economics, statistics, and business administration are necessary to take part in this staff. 131 During this research, ONS demonstrated a well-established set of technological competences, adopting practices such as hiring highly-qualified professionals in Big Data. In sequence, the British organisation provoked the cooperation between traditional statisticians and data scientists, data analysts and business analysts. An established data-driven culture was built based on those interactions. So, this mixed strategy, hiring policy and internal skills development, brought a cultural transformation to ONS. The change allowed ONS to adopt Big Data technologies and methodologies in a gradual expansion. This managed change reduced the level of disturbance and resistance to ABD. Related to skills development, the key for the success was in the continuous investment on longterm training programs. Instead of short trainings, ONS staff underwent a master’s degree in data science to learn topics such as Big Data design, process design and standardization, digital data collection, data visualization, business models, and project management techniques. In sequence, ONS deployed and established experimental statistics journeys to promote the engagement of the skilled staff. The rule is “learning by doing.” Finally, the staff must be tasked to use Big Data technologies and methodologies for core activities, after the data quality criteria is achieved. Thus, the raw material required to begin the ABD process is a highly qualified workforce drawn from the labour market. Big Data knowledge is delivered to NSOs through skills and top-tier human resources. Following those requirements, technology competence is forged to prepare the NSOs for ABD. The time and money required to hire a Big Data expert may be less than those required to prepare someone from inside the organization. “My background: I have spent time in Industry as a start-up chief executive, and for companies and those are in academia using a background in physics (undergraduate). But then, I used what we would now describe as deep learning (PhD in artificial intelligence and neuroscience) to develop control brains for robots, and there is a link with Big Data. (…) So, I joined ONS at the same time as joining the public sector civil service in the UK government, when I applied for the role of Campus Director for the Data Science Campus.” [03_ONS] ONS members remarked that individual productivity is part of technology competence . Based on this perspective, the change from MS-Excel to Stata and from Stata to “R” and “Python” represents 132 a technology competence gain. This upgrade flow showed that technology competence emerges as an asset to deal with ABD. This asset is associated with the human resources and skills and the organisational learning culture , demonstrated the capacity to embrace change and innovation. “Yep. So, let’s start with the technological side. So, there is a big push for the organization to move some old and heavy systems. We have quite a large and varied technology estate, and that’s been quite a journey. I think one of our biggest challenges is knowing which options to choose because technology moves so quickly that what looks sensible at one point in time can very quickly become the sort of new legacy system tomorrow." [04_ONS] In summary, Technology competence for ABD demands actions from the NSO. Indeed, those actions can be summarised in four points for the NSOs’ reality: 1. Why is technology competence essential for ABD in NSOs? Because it enables the staff to require/develop BD technologies and methodologies, and to design the adoption process. [ONS; INE; IBGE] 2. How does technology competence allow ABD? It allows ABD through three pillars: a. Mixed strategy: hiring policy and skills development. [ONS] b. Internal skills development: long-term training projects associated with experimental statistics construction. [ONS] c. Moving the staff to a data-driven culture. [ONS] 3. Requirement for technology competence: a. Continuous investment in training. [ONS] b. Execution of core actions: to propose, adopt, plan, deploy, and establish Big Data technologies and methodologies for experimental statistics. [ONS]. 4. Benefit of technology competence: individual productivity gain with “R” and “Python” applications. [ONS] 012. Technology resources Based on El-Haddadeh et al. (2021), technology resources represent tangible and intangible resources such as hardware, software, human resources, skills, and experience that an organisation acquires for implementing Big Data. 133 During this research, specifically ONS presented four additional elements of technological resources for ABD in NSOs: computing, storage and processing power, data integration, and software licenses. Those elements are associated with the capacity of the respective NSO staff to realize and demand uncommon and untraditional technological resources necessary to ABD. One component of those technology resources is the coding-controlled applications, as remarked by ONS employees. The adoption of those tools allows employees to manage and control the Big Data deployment in the Official Statistics processes. However, the interviewees from ONS and INE stressed that this is a work in progress rather than a finished process. This statement brings the understanding that the technology resources for ABD in NSOs include other assets. One is the access to new data sources, allowing the development of experimental statistics. Another is the presence of top qualified professionals to manage the infrastructure dedicated for Big Data. “On the one hand, an increase in technological capacity. A big investment not only in hardware - servers, for example, and processing power, for the other side. (…) As I said, it was a couple of decades. So, in that sense, there were many, many differences. And I can say that the biggest of all differences refers, exactly, to the data-warehouse area.” [01_INE] “Now, in 2021, there was a reinforcement of this response component. There was the introduction of autocomplete function, which, as the person writes, we suggest entries for the text in the open questions. We had the use of machine learning in the coding of the descriptive; and one of the other innovations was the incorporation and the possibility of yielding information from the administrative registers.” [04_INE] Moreover, the ABD in NSOs means transforming the traditional technological infrastructure and the medium data process into Big Data outcomes. This infrastructure is part of the ABD effort. “ Statistical Offices are on a very big journey to modernize our infrastructure for technology and for data used to manipulation. Every National Statistics Office is the same, and it always will be. This is never a journey that will be completed. (…) Over the period since then, their investment in infrastructure that ONS is doing really in a very large-scale. Now, we have system-wide organizational platforms and tools that support analysts, statisticians, data scientists, domain experts, such as 134 demographers, Model S, Health analysts, and so on. I am using a very rich range of data. “ [03_ONS] Namely, the technologies to arrange Big Data include Cloud storage, peer-review version-controlled (Git Hub and Git Lab), and Python and “R” applications. Those solutions are always controlled by code and represent technological resources . “(…) So, data infrastructure, which clouds systems, and tools with all of the data science stack that I would like to see from you know, free per review versioncontrolled Git Hub, Git Lab type approaches. Throw well-understood software development in documented practices. (…) the end applications (…) in Python and “R” packages and applications for text analysis, image analysis, for time series (…) right throw the publication using things like (R-)Markdown. And all of them are controlled by a code. ” [03_ONS] In summary, the technology resources factor for ABD in NSOs means: 1. Why is the technology resources factor determinant for ABD in NSOs? Because technology resources allow the NSO to run Big Data, that implies the use of uncommon and untraditional technologies, such as coding-controlled applications. [ONS; INE] 2. How does the technology resources factor run ABD in NSOs? Technology resources run ABD through: a. Increasing computing, storage, and processing power, and data integration. [ONS] b. Accessing data sources. [ONS] c. Managing technological infrastructure by skilled professionals. [ONS] 3. Examples of technology resources for Big Data in NSOs: cloud storage, free per-review version-controlled (Git Hub and Git Lab), Python and “R” applications. [ONS] 013. Trialability Baig et al. (2019, p. 10) define trialability as “Trialability is the degree to which companies can experiment with Big Data before fully implementing or committing to Big Data adoption.” This factor is necessary because it implies running experimental statistics, regarding tests with low-sensitive datasets. Testing and verifications are mandatory to verify if the proposed BD solution is adequate for the NSO need. 135 In fact, trialability is relevant for ABD through experimental statistics operating a pipeline of trial projects. Additionally, trialability works by the adoption of peer-reviewable Big Data applications. In the context of Official Statistics, trialability arises as a dependent factor of security and privacy, financial capacity, human resources, skills, and time. Indeed, the three cases – ONS, INE, and IBGE – showed that trialability demands a professional's willingness to try and experiment innovations, combined with a risk-taking capacity. For ONS, a trailable mentality must come before trialability. In the operational domain of NSOs, trialability is represented by experimental statistics, as remarked for the three cases. This type of statistics refers to a pipeline of trial projects related to ABD for improving or creating innovative Official Statistics outcomes. Our finding was the fact that ONS is more advanced in projects to develop trial/experimental statistics and evaluate then to the Official Statistics standard. On the other side, INE and IBGE are early adopters of experimental statistics. Those trial projects require customized and peer-reviewable procedures with the use of new methodological approaches, using specific types of data. In this context, low-security and lowsensitivity datasets are mandatory for experimental statistics, to avoid security and privacy damages. Additionally, the supporting requirements for trialability/experimental statistics include budget, skilled professionals, IT infrastructure, and time (schedule). There is a consensus in the three cases that those requirements are mandatory for any Big Data project for Official Statistics. “So, orderable, a peer reviewable, and then reusable for other applications. (…) But data scientists can practice, trying innovative approaches, new tools, application packages, etc. Not being loaded or set up on the corporate systems. (…) And we can use those for low security, low sensitivity datasets. So, we have much more control of we’ve run on those systems. Then, we have much less access to secure data.” [03_ONS] Trialability can also be understood as the capacity to take risks for innovation. In other words, adopters of experimental statistics are potential adopters of Big Data technologies and methodologies. Those departments, as risk takers, represent the entering gate for Big Data in an NSO. 136 “ It depends on the dataset; it can depend on your risk appetite… So, price statistics, by legislation, we can’t revise our price statistics. Probably, Brazil is very similar. And therefore, you got to get it right the first time. There’s no go back and revise it. So, our risk appetite is a lot lower in that area than in some other statistics that where we might still be in an experimental state. And actually, we’re actively exploring how to use this data in more ways. We’re still looking to use this scan of data in prices. It just means the process. It’s probably longer and there’s more checks along the way. So, I think we’re very actively looking to deploy Big Data technologies.” [04_ONS_DEX] In summary, the trialability of ABD in NSOs means: 1. Why is trialability necessary for ABD in NSOs? Because trialability involves experimental statistics, regarding tests with low-sensitive datasets. Testing and verifications are mandatory to verify if the proposed BD solution is adequate for the NSO need. [ONS; INE] 2. How does trialability support ABD in NSOs? Trialability supports ABD through: a. Experimental statistics operating a pipeline of trial projects. [ONS; INE; IBGE] b. Adoption of orderable and peer-reviewable Big Data applications. [ONS] 3. Requirements for trialability: a. Development of a trailable mentality before the trialability. [ONS] b. Dedicated budget, professionals, IT infrastructure, and time (schedule). [ONS; INE; IBGE] c. Capacity to take risks through innovative methodological approaches. [ONS; INE] d. Use of data for trials: low-security, low-sensitivity datasets. [ONS; INE; IBGE] 014. Vendor support Baig et al. (2019, p. 9) consider this factor as “helpful in implementing Big Data technologies. Vendors can fix information technology related issues more quickly and efficiently.” So, it can be defined as the services and activities to support the ABD in NSOs when the solutions are bought from commercial suppliers. This factor presented low relevance for this research. Despite this reduced influence, vendor support avoids the lock-in dependence from a specific supplier. If an NSO has only one vendor for the Big Data solution, this may pose a risk due to over-reliance on that vendor. 137 Related to the support offered by vendors, it works as services to support the adoption of demand programming hours and Business Intelligence services. However, it is possible to try to avoid becoming reliant on a single supplier. Only INE considered this as an important factor for ABD, specifically for projects where business intelligence support and programming hours are involved. In fact, the Portuguese NSO showed a prevention attitude to avoid a lock-in dependence from a supplier/vendor. "In other words, when INE contracts with a company-whether it's Microsoft or Amazon or the big players that are out there in the market - it is not just contracting for cloud storage. It is also contracting for intelligence." [01_INE] "Our relationship with companies related to information systems, technologies and refers to software development. Typically, the development is in-house. We can hire companies. But our relationship with companies is almost for programming hours." [01_INE] In summary, the vendor support factor of ABD in NSOs means: 1. Why is vendor support necessary for ABD in NSOs? Because vendor support avoids the lock-in dependence from a specific supplier. [ONS; INE] 2. How does vendor support help ABD in NSOs? Vendor supports works as services to support the adoption of demand programming hours and Business Intelligence services. [INE] 0X2. Obsolescence Obsolescence means the incapacity of the traditional methods and technologies to address current issues in Official Statistics, such as large volumes of data, randomness, and on-time data collection. It works under a timing perspective, from the past (traditional methods) to the present (Big Data methods). Obsolescence is a concern for NSOs. The risk of Official Statistics being considered outdated information in the age of Big Data is well known by the staffs in Brazil, Portugal, and United Kingdom. “ And our old type of technology and old type of methods weren’t really suitable for that. But that also means that what you call Big Data is a moving target. So, what we 138 called Big Data five years ago probably isn’t big anymore, given the technology that we’ve got at the moment. So, you know? Years ago, the Census of the population in the UK, which was rounding about sixty-five million, would be considered Big Data. That is not Big Data anymore. Now, you can analyse it on your PC." [02_ONS] In the context of ABD, obsolescence is critical because this factor presents the limited capacity to launch and manage Big Data long-term project by some NSOs. So, obsolescence is manifested as over-centralised technological and departmental structures, and the maintenance of traditional methodologies against the Big Data alternatives. In the case of IBGE, the over-centralization in an organisational unit to manage technological subjects can be understood as obsolete and a constrain for the ABD. Another effect of obsolescence is the response time or outcomes delay derived from that over-centralization. "Precisely, because of the survey that we did. We can't respond at the speed that citizens need. We have a lot of difficulty, because our time for data collection and processing is very long in relation to citizens necessities.” [10_IBG] The Brazilian NSO is outdated in its deliveries, as it owns members remarked. One example of delay, as presented in Chapter 5, is the extension of the Brazilian demographic Census with estimated conclusion in 2023, May, after several postponements. In this scenario, a feature that sustains obsolescence was found. It refers to theorizing more than executing proposals, specially about ABD projects. This statement was remarked by IBGE members. " So, I said, ‘what is the use of theorizing and not doing anything?’. But this is the IBGE culture (...) There was not a strong conception of Big Data in any directorate." [01_IBG]. Different from UK and Portugal, Brazil demonstrated a limited capacity to launch and manage Big Data long-term projects. Regarding the previous citation, obsolescence dominates where there are more discussions than actions. In fact, ABD is a hard challenge to be addressed, and it becomes critical in cases where the NSO is suffering obsolescence. An ONS member remarked that “ Big Data is a moving target ” with 139 changing requirements constantly. So, the insistence to focus on traditional methodologies while Big Data projects are avoided represents limited capacity to improve and update Official Statistics. Despite the damages and risks of obsolescence, there is an escape route for this failing. An ONS member pointed the way to overcome obsolescence and adopt Big Data. First, it is mandatory to map the flow of the data and all the processes involved. In sequence, the potential improvements must be listed. Finally, the technology resources must be defined to be deployed in a future ABD project. Obsolescence in NSOs for ABD means: 1. Why is the obsolescence factor critical for ABD in NSOs? Because obsolescence presents the limited capacity to launch and manage Big Data long-term project by some NSOs. [IBGE] 2. How does the obsolescence factor support ABD in NSOs? Obsolescence constrains ABD through: a. An outdated organisational structure: IT over-centralisation. [IBGE] b. Maintenance of traditional methodologies that are insufficient for the ABD Adoption challenge. [ONS; INE; IBGE] 3. As a consequence of obsolescence, there are delays in the extended response time to achieve outcomes. [IBGE] 4. NSOs can react to obsolescence by: a. Talking about ABD more than actually doing anything about it - “Much theorization, less action.” [IBGE] b. Trying to overcome it by an escape route: mapping the flow of the data and all the processes being applied. [ONS] 0X3. Big Data Technologies and Methodologies (BDTM) Big Data technologies and methodologies (BDTM) can be described as the features of the Big Data technologies and methodologies adopted by NSOs. In this context, the case studies indicate that NSOs work mainly with four Big Data technologies: Cloud, Artificial Intelligence (AI), web scraping, and Big Data analytics (BDA). Related to the methodologies, the cited examples are clusteringbased alternative method for AI to detect outliers (TCBAN), natural language processing (NLP), sentiment-related information, and real-time measurement. 242 process influenced by the Technology, Organisation, and Environment dimensions. As a result, the ABD fails or advances not because the NSO decided this at one time. In fact, the internal and external institutional conditions allow or inhibit the NSO from adopting and advancing in the ABD journey. Without those conditions, the decision for ABD can be taken, but the deployment will remain an organisational illusion. 7.6. Big Data for Official Statistics: beyond disruption According to Ghaleb et al. (2021), Big Data technologies, like the Cloud or IoT, can be considered disruptive technologies when adopted by a novice organisation. In the same way, Gangwar (2018) considered Big Data technologies disruptive in the manufacturing and service sectors. However, this doctoral thesis, unlike the current literature, finds that ABD in the NSO context is not disruptive. After the data analysis, this doctoral research proposes a different interpretation of ABD for National Statistics Offices. In the three case studies, IoT, Cloud, Web Scraping, and AI/Machine Learning, as Big Data technologies, were considered powerful and useful complements for Official Statistics production. An ONS interviewee mentioned this transition from the disruptive to the complementary interpretation, explaining that, initially, Big Data was considered a new beginning for Official Statistics. After a while, ONS members realised that Big Data was a useful supplement to, not a replacement for, current Official Statistics. So, the main change that the adoption of Big Data brings is the improvement in the production of the Official Statistics outcomes. This perspective can change with the progress of the adoption. Under the current view of the NSOs, Big Data does not apparently mean changes in the types of Official Statistics outcomes for the short term. Even with the use of Big Data, the price index, unemployment rate, industrial production rate, and demographic variation are still relevant Official Statistics. According to this contemporary understanding, ABD brings changes to "how" (processes) Official Statistics are produced, not "what" (types) Official Statistics are. In fact, the three cases have confirmed this understanding. From the use of AI algorithms to support data collection in the demographic Portuguese Census to the expanded use of cloud computing in ONS and the limited application of web scrapping in IBGE, Big Data acts as a booster to support NSO members in their day-to-day processes. 243 In summary, ABD is welcomed by Official Statistics operators as a valuable asset to improve their delivery to society with more accuracy in real-time and with less demand on material and human resources. Indeed, Big Data is understood as an incremental innovation, not a breakthrough one. This positive perception encourages the adoption of Big Data as long as it does not jeopardise institutional based trust . On the contrary, Big Data can reinforce it. 244 Chapter 8 - Contributions, Limitations and Propositions In this conclusion stage, it is necessary to remember how this research advanced and the limitations it suffered. Not only those aspects but also the contributions are summarised in this last chapter. Finally, the possibilities for future studies are proposed under the perspective that this exploratory result is just a starting point. Adopting or avoiding Big Data technologies and methodologies represents a critical step for National Statistics Offices (NSOs). The seven chapters presented the relevance of Big Data and how it can contribute to Official Statistics. In Chapter 4 this question was proposed: "How and why do National Statistical Offices adopt or avoid Big Data technologies and methodologies?" The answer to the research question was obtained following a specific methodological design. Base on the adjudicative perspective proposed by Cronin and George (2020), the methodology advanced a systematic literature review (LR). This LR was developed using a knowledge-synthesis orientation, integrating studies about Big Data adoption, including several business sectors, organisational types, and methodologies. As an outcome of the LR, the TOE factors were reviewed and, in some cases, redefined to add juxtapositions. This application of the knowledge-synthesis orientation contributes to reinforcing the approach developed Cronin and George (2020). A multiple case-study approach was followed to pursue the empirical study. The case selection for this research followed three different criteria: governmental nature, geographic location, and position in the Global Statistical System (GSS). As a result, ONS (United Kingdom), INE (Portugal), and IBGE (Brazil) were selected. In fact, the differences between the cases offered a rich opportunity for comparisons and a broad understanding about ABD. Using the study about TOE, a semi-structured interview guide was developed. It was applied to twenty representants from the three NSOs. In sequence, the interviews and the collected documents were analysed using codification as proposed by Saldaña (1993). This methodological strategy brought relevant findings, such as the specific features of ABD in the NSO context, and the stage each NSO is at in the ABD process. Conclusions emerged from those findings and are detailed in the next pages. The overall answers found are, firstly, that NSOs adopt Big Data to avoid obsolescence and keep the relevance of Official Statistics for citizens and societies. Related to the mode of adoption, NSOs 245 adopt Big Data in incremental stages, initializing with selected deliveries, and in sequence, spreading this adoption for the other surveys and Census operations. Related to avoidance, NSOs’ resistance to ABD involves the attempt to keep the status quo, focusing on the preservation of internal outdated power structures. A combination of environmental and organisational factors affects the technological context, blocking the development and acquisition of Big Data technologies and methodologies. In the end, this research brings a clear message for the guardians of Official Statistics: the adoption of Big Data is not a choice but a matter of survival for those organisations and their stakeholders. 8.1. Summarizing the journey The initial idea for this research occurred seven years ago, at the beginning of 2016. Since that period, the demand for Big Data in the NSOs has been a common thought among those organisations. Over the years, this demand has remained high, and the aspiring future for Official Statistics using Big Data is a moving target. Regarding this constant search for Big Data adoption, NSOs realised that, as a technological target, Big Data is not a push-button solution. Beyond this, Big Data is a rich asset for Official Statistics in the 21st century, and as an innovation, it demands adaptation and flexibility. Unlike previous studies, this research revealed an incremental orientation to the adoption process, not a disruptive one. After forty-eight months, three international technical visits, three international conferences, twenty interviews, three NSOs accessed, and more than three hundred papers analysed, the result was satisfactory. Those numbers and efforts supported the documentary and content analysis. During this trajectory, it was hard to identify how to compare an NSO with continental coverage with another located in an archipelago and another in a small European country. Even so, it was possible to advance with the comparability using the TOE framework. Additionally, the participation of the PhD student in several seminars and workshops allowed a broad comprehension of how Big Data is powerful and valuable for Official Statistics. This understanding signalled to the researcher that ABD must overcome obsolescence risks. So, it is mandatory that every NSO adopt Big Data instead of avoiding it. 246 8.2. Contributions This study incorporates both theoretical and practical contributions. On the first branch, the results propose an advance in the current theory. In the context of practice, the findings are helpful for NSOs to develop or update their strategies and processes. Firstly, at the theoretical level, this research verified a novel feature about TOE framework, the dimensional precedence. In the context of NSOs, the Environmental dimension is the starting point influencing the adoption of Big Data. Key-factors like g overnment support/ laws and policy , National context, and Industry features trigger the respective NSO in the direction of Big Data adoption or avoidance. In fact, this research showed that the decision to include IoT, Cloud, Web Scraping, and AI in the Official Statistics activities is a sequential and cumulative process. This process is drastically affected by the external environment, the Environmental dimension, and, in sequence, by the Organisational dimension. In the final process stage, the Technological dimension demonstrates a conditioned influence on the adoption of Big Data. So, the clockwise direction is necessary to explain the ABD in the NSOs reality. Also at the theoretical level, this study contributes by asserting the relevance of some factors over others. In the three dimensions – Environmental, Organisational, and Technological –, there are subsets of factors that influence ABD more than others. Institutional based trust and International - Institutional Context are examples of the critical Environmental factors for ABD in NSOs. Top Management Support, Cultural Blockers, and Potential Pervasiveness are part of the determining Organisational factors subset. Last, but not least, Polysemy and noise , Adoption design for Big Data , and Obsolescence refers to the group of Technological factors with relevant influence on ABD in NSOs. As a qualitative research, this doctoral thesis complements the research about TOE framework for ABD, demonstrating that there are two levels of factor influence: some more and some less relevant. Considering this relevance of factors from the three dimensions, one more finding emerged during the research: the interdimensional association. ABD in NSOs showed that factors interact with each other in chain of ties, inducing inputs to adopt or avoid Big Data technologies. This finding shows the factors work together, and not isolated in the respective dimension. Regarding practical contributions, three specific findings were obtained: (i) idiosyncrasies of Big Data adoption processes in NSOs; (ii) adoption as a cumulative process over time, not a single decision at a certain point; and (iii) ABD as a complementary/incremental process. Beyond the 247 academic results, those three findings show solid orientations for NSO managers and staff to advance in their ABD journey. Idiosyncrasies of adoption of Big Data processes in NSOs include the critical role of the procurement departments, as part of the business architecture. This finding is embedded in the Organisational factor Human resources and Skills because a highly skilled procurement staff is able to drive the procurement process to acquire Big Data technologies. In the same path, the hiring process is also another particularity of NSOs as public organisations. Submitted to the Central Government control, those organisations are restricted in hiring employees or consultants, different from private companies. Those specifications strengthen the ABD specific landscape in NSOs. In this particular reality, the adoption of Big Data, or its avoidance, results from a cumulative process of interactions between the Top Management Team and the staff, influenced by the TOE factors remarked in Chapter Seven. This aggregation of interactions and decisions showed that ABD presents distinct intensities of avoidance or adoption. NSOs can adopt Big Data in a strong or weak mode. Strong adoption means spreading Big Data technologies in several departments and processes. A weak adoption signifies the use of Big Data is restricted to a few core processes, or in some experimental efforts. In the opposite direction, strong avoidance means the adoption of no Big Data technology whatsoever, whereas weak avoidance one refers to isolated adoptions without the support from the Top Management team. The last finding refers to the adoption as an incremental/complementary process, not a disruptive one. This is the way the three NSOs understand and deploy ABD in their business architecture and processes. Progressive efforts to adopt the Big Data, technology by technology, reinforce this incremental orientation. When an NSO adopts Big Data, it adopts one technology, such as Web Scraping, or at most, two technologies in specific processes under an experimental mode. In this effort, the technological factor Trialability is critical. If an NSO supports and reinforces Trialability , it increased the possibility to advance ABD successfully. So, firstly, the NSO adopts Big Data experimentally. Based on the results, if satisfactory, the new technology is added to the Official Statistics processes. So, there is no evidence in this research that ABD disrupts the Business Architecture or dispels the original pipeline of each NSO in a dramatically mode. A managed transition appears as the rule of ABD in NSOs. 248 Beyond those six main findings, this doctoral research brought additional contributions. Firstly, the resilience of TOE was proven in this research, guiding the analysis with its logic of three dimensions and allowing findings about the specific context of Official Statistics. It is necessary to highlight that this context of Official Statistics had not been previously studied using the TOE framework. Second, the TOE framework demonstrated flexibility to be adopted in this research as it was adopted in previous articles. One example is the proposition offered by Baker (2012) about applying TOE in adoption studies involving specific business sectors. In this doctoral research, TOE was feasible and intensively valuable for understanding how NSOs have adopted Big Data. TOE is thus an aggregative framework that can be applied to research agendas over time. Regarding originality, the Literature Review (LR) followed a critical perspective, allowing the construction of an integrative compilation of the TOE factors. This compilation advanced to a review of the TOE framework through the consolidation of terms and concept for each factor. Indeed, the consolidation was necessary in the face of the verified polysemy related to the factors. In summary, a set of reviewed factors was identified and defined, and this set was adopted to understand how and why NSOs adopt or avoid Big Data. The concept of Big Data also presented polysemy and was therefore revised, in general, and specifically for the context of NSOs. Based on the empirical study, this doctoral research advanced ten new TOE factors never mentioned before in the literature. This group of factors – including six new technological, three organisational, and one environmental – brought new possibilities to the evolution of the studies about ABD under the lens of the TOE framework. Likewise, the contribution of new factors showed the explanatory capacity of TOE in a complex qualitative study beyond the traditional quantitative papers where it is used to support structural equations. As a result of this four-and-a-half-year research, the contributions of this research advanced on specific topics related to the adopters, the NSOs. One finding concerns the vital and urgent demand for adopting Big Data in the Official Statistics, considering this as a business sector. It is a matter of survival, not a choice. Avoidance is harmful for the NSOs persisting on it. Related to additional findings, there is the concerns about ABD stages. Each NSO advances in the adoption effort and related organisational changes along specific stages. This process takes time — in some cases, over five years — as demonstrated in the Portuguese and British cases. As identified by this research, there are four phases for the ABD in NSOs: (i) strong avoidance, (ii) weak avoidance, (iii) weak adoption, and (iv) strong adoption. Each phase of adoption seems to be 249 related to a specific maturity level. Some NSOs are at a high maturity level, others in the middle, and some at a low one. However, there are NSOs that are a rejection level, which is avoidance. Finally, the last finding confirms that the adoption of Big Data need not have a disruptive impact on NSOs. When organizations adopt Big Data, the pace is incremental, gradually aggregating the new technologies and methodologies into the existing portfolio of outcomes. In some previous studies, Big Data and disruption were considered synonyms. 8.3. Limitations This doctoral research was subjected to structural and situational limitations. In the first category, structural limitations, the methodological restrictions are explained in the following paragraphs. In addition, adverse events impacted the planned evolution of the study. Regarding the methodological limitations, the multiple-case study focused on three specific realities: British, Portuguese, and Brazilian. As defined in the methodological chapter, the selection of the cases attempted to offer generalisations for the findings. However, this generalisation must be tested in a future qualitative study with other NSOs. So, the impossibility of checking the generalisation of those conclusions is a structural limitation of this research. Although the use of the TOE framework to address the adoption of Big Data proved highly valuable in this research, other frameworks may be used and tested to understand how ABD advances in the context of Official Statistics and NSOs. Given the temporal dimension identified in this study, a Process Perspective (Langley et al., 2013) and Maturity models (Chen and Nath, 2018) are examples of analytical tools to be used in future studies. As for practical constraints, the research plan was initially based on a schedule of forty-eight months, with the conclusion estimated for the 2022 second semester. However, the COVID-19 pandemic had a significant impact on this planning. The lockdowns in 2020 and 2021; the flight reductions between Brazil, Portugal, and the United Kingdom in 2021; and the sanitary barriersimposed limitations on the research. Almost all interviews advanced remotely, contrasting with the original intention to conduct them in person. Because of the restrictions, ONS, INE, and IBGE had to postpone this researcher's access because those organisations were dealing with their reviews during the three Census calendars. Actually, it is possible to measure that the pandemic caused a delay of one semester in this research. 250 8.4. Future studies An exploratory and qualitative study can be a starting point for future investigations like this doctoral research. A research agenda is proposed in this last section; understanding ABD in National Statistics Offices is a relevant theme for those organisations, governments, citizens, and also for the academy. This research agenda can be advanced with a quantitative study applied to other NSOs, analysing ABD in their processes and outcomes. In this scenario, the universe of 198 NSOs should be considered. A clear objective can be proposed for this future research: identify if there are specific patterns for ABD in NSOs. In another study, ABD can be verified in the complete portfolio of surveys or only in Census operations. However, if the second proposal is considered, only NSOs where the Census took place in the last three years should be included, considering that Big Data became part of the Census recently. All surveys or just the Census operation can allow the verification of best practices for ABD, generating rules to advance this adoption in the current decade for all GSS members. Another potential research topic is examining the role of networks on ABD in NSOs where Big Data Hubs were deployed. In this line, NSOs from Brazil, China, Rwanda, and the United Arab Emirates could be studied qualitatively, following a similar orientation to this doctoral research. As a central objective, the verification of a spill-over effect to other regional partners can advance. The role of networks could also be analysed by studying the impact of the Global Statistical System (GSS) on the ABD process. In this future study, it would be possible to investigate how and why some NSOs are more influenced than others by this global network. In this proposal, quantitative research can be adequate to discover if GSS influence works as a supportive environmental factor for the ABD. 8.5. Final Consideration: NSOs, an organisational species at risk of extinction? Besides this doctoral research focus on the adoption of Big Data, it is necessary to understand the emergent findings about the adopters, the National Statistics Offices (NSOs). This group of organisations is dealing with threats, as Cardoso (2022) remarked. One is obsolescence, and the other is loss of social relevance, as highlighted by Allin (2021). 251 In this multiple-case study, it became clear that NSOs can face those threats when ABD is at a middle or high maturity level. So, those adopters’ control and administer the changes and demands presented by governments and citizens. A demonstration of this capability is the perception of INE and ONS that Official Statistics production has to be faster and cheaper than it is nowadays. With reduced expenses to deliver an outcome, those NSOs plan to keep their relevance for the population, in spite of the escalation of fake news. On the other side, IBGE is still operating from an outdated perspective, concluding the Demographic Census with three years of delay and at a cost of around half a billion euros 3 . Contrary to this, the IBGE members realise the Brazilian Official Statistics have to advance and be modernised, but this perception is not reflected in corporate initiatives. In cases similar to the Brazilian, NSOs are at considerable risk of losing social relevance. If this scenario is fulfilled, NSOs can be substituted by private data generation, compromising not only Official Statistics in some countries but also the credibility of governments and nations. Another risk refers to the Global Statistical System. In a scenario where some members of this network will go extinct or are unable to keep the exchange flow of knowledge going, the GSS will suffer a disruption, and this system can start to collapse. The GSS members must struggle to avoid the extinction of any NSO. This proposition can come by influencing traditional organisations such as the OECD and UN to have their members' input on the mandatory preservation of the National Statistics Offices in each country. Concerning preservation, this can be examined by peers. In this context, inspectors from the GSS would verify whether an NSO was achieving the Global patterns and principles of the Official Statistics during an in-field inspection. A similar approach has been taken by the International Atomic Energy Agency (IAEA) 4 . Those inspections would attribute a quality certification to an NSO, provided the organisation runs the Official Statistics regularly, just like the IAEA does with atomic energy uses in many countries. 3 Reference value of the exchange rate between Brazilian Real (BRL) and Euro (EUR) in the launching day of the demographic Census in Brazil, August 1st 2022. Retrieved from https://www.bcb.gov.br/conversao, in January 10th 2023 4 Retrieved from https://www.iaea.org/topics/verification-and-other-safeguards-activities, in January 10th 2023. 258 https://doi.org/10.3390/su13158379 Gilbert, M., & Cordey-hayes, M. (1996). Understanding the process of knowledge transfer to achieve successful technologicla innovation. Technovation , 16 (6), 301–312. Goodhue, D. L., & Thompson, R. L. (1995). Task-Technology Fit and Individual Performance. MIS Quarterly , 19 (2), 213-236. https://doi.org/10.2307/249689 Günther, W. A., Rezazade Mehrizi, M. H., Huysman, M., & Feldberg, F. (2017). Debating big data: A literature review on realizing value from big data. The Journal of Strategic Information Systems , 26 (3), 191–209. https://doi.org/10.1016/j.jsis.2017.07.003 Harari, Y. (2017). Homo Deus - A Brief History of Tomorrow . Vintage Publishing. https://www.bookdepository.com/Homo-Deus-Yuval-NoahHarari/9781784703936?ref=pd_detail_1_sims_b_p2p_1 Harari, Y. (2019). 21 Lessons for the 21st Century . Vintage Publishing. https://www.bookdepository.com/21-Lessons-for-the-21st-Century/9781784708283 Harzing, A.-W., & Alakangas, S. (2016). Google Scholar, Scopus and the Web of Science: a longitudinal and cross-disciplinary comparison. Scientometrics , 106 (2), 787–804. https://doi.org/10.1007/s11192-015-1798-9 Hatch, M. J., & Cunliffe, A. L. (2013). Organization Theory - Modern, Symbolic and Postmodern Perspectives . Oxford University Press. Hiebl, M. R. W. (2021). Sample Selection in Systematic Literature Reviews of Management Research. Organizational Research Methods , 109442812098685. https://doi.org/10.1177/1094428120986851 Ho, C. K. Y., Ke, W., Liu, H., & Chau, P. Y. K. (2020). Separate versus joint evaluation: The roles of evaluation mode and construal level in technology adoption. MIS Quarterly: Management Information Systems , 44 (2), 725–746. https://doi.org/10.25300/MISQ/2020/14246 Iacovou, C. L., Benbasat, I., & Dexter, A. S. (1995). Electronic data interchange and small organizations: Adoption and impact of technology. MIS Quarterly: Management Information Systems , 19 (4), 465–485. https://doi.org/10.2307/249629 Instituto Brasileiro de Geografia e Estatísticas. (2020a). Censo Agropecuário 2017 . Sobre. https://censoagro2017.ibge.gov.br/sobre-censo-agro-2017.html Instituto Brasileiro de Geografia e Estatísticas. (2020b). Postponing the Brazilian Census form 2020 to 2021 . Highlights. https://www.ibge.ge.gov.br/en/highlights/27164-censo-2020adiado-para-2022.html 259 Instituto Brasileiro de Geografia e Estatísticas. (2021). Coletânea Legislação . Memória. https://memoria.ibge.gov.br/images/pdf/memoria/coletanea_legislacao_ibge.pd Instituto Brasileiro de Geografia e Estatísticas. (2022a). Brazilian Census postponement form 2021 to 2022 . Highlights. https://www.ibge.gov.br/en/highlights/30572-postponement-of-thepopulation-census.html Instituto Brasileiro de Geografia e Estatísticas. (2022b). Recenseamento do Brazil em 1872 (G. Leuzinge (ed.)). https://biblioteca.ibge.gov.br/visualizacao/livros/liv25477_v1_br.pdf Instituto Brasileiro de Geografia e Estatísticas. (2023a). Linha do tempo . Memória IBGE. https://memoria.ibge.gov.br/linha-do-tempo.html Instituto Brasileiro de Geografia e Estatísticas. (2023b). O que é . População. https://www.ibge.gov.br/estatisticas/sociais/populacao/31876-dimensionamentoemergencial-de-populacao-residente-em-areas-indigenas-e-quilombolas-para-acoes-deenfrentamento-a-pandemia-provocada-pelo-coronavirus.html?edicao=31877&t=o-que-e Instituto Brasileiro de Geografia e Estatísticas. (2023c). Recenseamentos Gerais e Estatísticas Populacionais no Brasil . Memória IBGE. https://memoria.ibge.gov.br/historia-doibge/historico-dos-censos/censos-demograficos.html Instituto Brasileiro de Geografia e Estatísticas. (2023d). Regional Hub of Big Data in Brazil . UN Big Data Regional Hub in Brazil. https://hub.ibge.gov.br/ Instituto Nacional de Estatística (2011). E-Censos . Censos 2011. https://censos.ine.pt/xportal/xmain?xpid=CENSOS&xpgid=censos2011_e-censos Instituto Nacional de Estatística (2022). Mensagem do Presidente do INE . Censos 2021. https://censos.ine.pt/xportal/xmain?xpgid=censos21_menpresidente&xpid=CENSOS21&xl ang=pt Instituto Nacional de Estatística (2023). Presentation . Statistic Council. https://www.ine.pt/xportal/xmain?xpid=CSE&xpgid=cse_main&cont_cse=51354&cse_sme nu.boui=3109257&xlang=en INE. (2021a). Edital para Recrutamento Recenseadores . Instituto Nacional de Estatísticas. https://www.ine.pt/xportal/xmain?xpid=INE&xpgid=ine_rh_reccensos&contexto=rc&selTab =tab4&xlang=pt Instituto Nacional de Estatísticas (2021b). História dos Censos em Portugal . Censos Em Portugal. https://censos.ine.pt/xportal/xmain?xpid=INE&xpgid=censos_historia_portugal Instituto Nacional de Estatísticas (2021c). Recenseamento Agrícola - Análise dos principais 260 resultados - 2019 (Instituto Nacional de Estatísticas (ed.)). Instituto Nacional de Estatísticas. https://censoagro2017.ibge.gov.br/sobre-censo-agro-2017.html Instituto Nacional de Estatísticas (2022). História dos Censos em Portugal . Censos Em Portugal. https://censos.ine.pt/xportal/xmain?xpid=INE&xpgid=censos_historia_portugal Jackson, K., & Bazeley, P. (2019). Qualitative Data Analysis with NVIVO (First). SAGE Publications Ltd. Jung, J. H., Bapna, R., Ramaprasad, J., & Umyarov, A. (2019). Love unshackled: Identifying the effect of mobile app adoption in online dating. MIS Quarterly: Management Information Systems , 43 (1), 47–72. https://doi.org/10.25300/MISQ/2019/14289 Kaplan, A. (1964). The conducto fo inquiry: Methodology for behavioral science . Chandler. Karahanna, E., Straub, D. W., & Chervany, N. L. (1999). Information technology adoption across time: A cross-sectional comparison of pre-adoption and post-adoption beliefs. MIS Quarterly: Management Information Systems , 23 (2), 183–213. https://doi.org/10.2307/249751 Kim, S. S., & Son, J.-Y. (2017). Out of Dedication or Constraint? A Dual Model of Post-Adoption Phenomena and its Empirical Test in the Context of Online Services. MIS Quarterly , 110 (9), 1689–1699. Köhler, T. (2016). From the Editors: On Writing Up Qualitative Research in Management Learning and Education. Academy of Management Learning & Education , 15 (3), 400–418. https://doi.org/10.5465/amle.2016.0275 Kohli, R., & Tan, S. S. L. (2016). Electronic health records: How can is researchers contribute to transforming healthcare? MIS Quarterly: Management Information Systems , 40 (3), 553–573. https://doi.org/10.25300/MISQ/2016/40.3.02 Komiak, S. Y. X., & Benbasat, I. (2006). The effects of personalization and familiarity on trust and adoption of recommendation agents. MIS Quarterly: Management Information Systems , 30 (4), 941–960. https://doi.org/10.2307/25148760 Kothari, C. R. (2004). Research Methodology: Methods and Techniques (Second). New Age International Publishers. https://www.researchgate.net/publication/269107473_What_is_governance/link/54817 3090cf22525dcb61443/download%0Ahttp://www.econ.upf.edu/~reynal/Civil wars_12December2010.pdf%0Ahttps://thinkasia.org/handle/11540/8282%0Ahttps://www.jstor.org/stable/41857625 Kotz, S. (2005). Reflections on Early History of Official Statistics and a Modest Proposal for Global 261 Coordination. Journal of Official Statistics , 21 (2), 139–144. Kunisch, S., Menz, M., Bartunek, J. M., Cardinal, L. B., & Denyer, D. (2018). Feature topic at organizational research methods: How to conduct rigorous and impactful literature reviews? Organizational Research Methods , 21 (3), 519–523. https://doi.org/10.1177/1094428118770750 Kvale, S. (1996). InterViews - An Introduction to Qualitative Research Interviewing . SAGE Publications, Inc. Kwon, O., Lee, N., & Shin, B. (2014). Data quality management, data usage experience and acquisition intention of big data analytics. International Journal of Information Management , 34 (3), 387–394. https://doi.org/10.1016/j.ijinfomgt.2014.02.002 Lai, Y., Sun, H., & Ren, J. (2018a). Understanding the determinants of big data analytics (BDA) adoption in logistics and supply chain management: An empirical investigation. International Journal of Logistics Management , 29 (2), 676–703. https://doi.org/10.1108/IJLM-06-20170153 Lai, Y., Sun, H., & Ren, J. (2018b). Understanding the determinants of big data analytics (BDA) adoption in logistics and supply chain management. The International Journal of Logistics Management , 29 (2), 676–703. https://doi.org/10.1108/IJLM-06-2017-0153 Langley, A., Smallman, C., Tsoukas, H., & Van de Ven, A. H. (2013). Process Studies of Change in Organization and Management: Unveiling Temporality, Activity, and Flow. Academy of Management Journal , 56 (1), 1–13. https://doi.org/10.5465/amj.2013.4001 Lundblad, J. (2003). A Review and Critique of Rogers’ Diffusion of Innovation Theory as It Applies to Organizations. Organization Development Journal , 21 (4), 50–64. Maroufkhani, P., Iranmanesh, M., & Ghobakhloo, M. (2022). Determinants of big data analytics adoption in small and medium-sized enterprises (SMEs). Industrial Management & Data Systems. 123(1). 278-301. https://doi.org/10.1108/IMDS-11-2021-0695 Maroufkhani, P., Wan Ismail, W. K., & Ghobakhloo, M. (2020). Big data analytics adoption model for small and medium enterprises. Journal of Science and Technology Policy Management , 11 (4), 483–513. https://doi.org/10.1108/JSTPM-02-2020-0018 Martinez-Mosquera, D., & Luján-Mora, S. (2019). Framework for big data integration in egovernment. DYNA , 86 (209), 215–224. https://doi.org/10.15446/dyna.v86n209.77902 McGuire, A. (2016). Researching Big Data Skills Using a Mixed-Methods Approach. SAGE Research Methods Cases , 1-16. https://doi.org/10.4135/9781526419620 262 Mcmahon, P., Zhang, T., & Dwight, R. (2020). Requirements for Big Data Adoption for Railway Asset Management. IEEE Access , 8 , 15543–15564. https://doi.org/10.1109/ACCESS.2020.2967436 Memon, S., Changfeng, W., Rasheed, S., Pathan, Z. H., Saddozai, S. K., Yixin, Q., & Yanping, L. (2016). Adoption of Big Data Technologies for Communication Management in Large Projects. International Journal of Future Generation Communication and Networking , 9 (10), 73–82. https://doi.org/10.14257/ijfgcn.2016.9.10.07 Ministério do Planejamento. (2023). Unidades do Ministério do Planejamento . Institutional. https://www.gov.br/economia/pt-br/acesso-a-informacao/institucional/planejamento Miranda, M. Q., Farias, J. S., Schwartz, C. D. A., Pascualote, J., & Almeida, L. De. (2016). Technology adoption in diffusion of innovations perspective: introduction of an ERP system in a non-profit organization. RAI Revista de Administração e Inovação , 13 (1), 48–57. https://doi.org/10.1016/j.rai.2016.02.002 Mital, M., Chang, V., Choudhary, P., Papa, A., & Pani, A. K. (2018). Adoption of Internet of Things in India: A test of competing models using a structured equation modeling approach. Technological Forecasting and Social Change , 136 , 339–346. https://doi.org/10.1016/j.techfore.2017.03.001 Morakanyane, R., Grace, A., & O’Reilly, P. (2017). Conceptualizing Digital Transformation in Business Organizations: A Systematic Review of Literature. Bled Proceedings , 427–443. https://doi.org/10.18690/978-961-286-043-1.30 Moss, L. (2023). Census data reveals LGBT+ populations for first time. BBC News . https://www.bbc.com/news/uk-64184736 Moura, I., Dominguez, C., & Varajão, J. (2021). Information systems project team members: factors for high performance. The TQM Journal , 33 (6), 1426–1446. https://doi.org/10.1108/TQM-07-2020-0170 Müller, O., Fay, M., & vom Brocke, J. (2018). The Effect of Big Data and Analytics on Firm Performance: An Econometric Analysis Considering Industry Characteristics. Journal of Management Information Systems , 35 (2), 488–509. https://doi.org/10.1080/07421222.2018.1451955 Noor, K. B. M. (2008). Case Study: A Strategic Research Methodology. American Journal of Applied Sciences , 5 (11), 1602–1604. O’Neil, C. (2016). Weapons of Math Destruction . Penguin Books. 263 Office for National Statistics. (2022a). Data and analysis from Census 2021 . Legislation and Policy. Office for National Statistics. (2022b). Story of Census . Office for National Statistics. https://www.ons.gov.uk/visualisation/storyofthecensus Office for National Statistics. (2023a). About Us . What We Do. https://www.ons.gov.uk/aboutus Office for National Statistics. (2023b). Home . Organisations. https://www.gov.uk/government/organisations/office-for-national-statistics Office for National Statistics. (2023c). Release Plans . Release Plans. https://www.who.int/emergencies/diseases/novel-coronavirus-2019/mediaresources/science-in-5/episode-63---omicronvariant?gclid=Cj0KCQiAgaGgBhC8ARIsAAAyLfHI9hHoI1v4EQWHLKaKB8d1p36GoxgZGbeH wzS4lH_2Bozu5KualtYaAlBjEALw_wcB Olszak, C. M., & Zurada, J. (2020). Big Data in Capturing Business Value. Information Systems Management , 37 (3), 240–254. https://doi.org/10.1080/10580530.2020.1696551 Orlikowski, W. J. (2000). Using Technology and Constituting Structures: A Practice Lens for Studying Technology in Organizations. Organization Science , 11 (4), 404–428. https://doi.org/10.1287/orsc.11.4.404.14600 Orwell, G. (2013). 1984 . Penguin Books. https://www.bookdepository.com/Nineteen-Eighty-FourGeorgeOrwell/9780141393049?redirected=true&utm_medium=Google&utm_campaign=Base1&u tm_source=PT&utm_content=Nineteen-EightyFour&selectCurrency=EUR&w=AF7DAU99ZT93J0A8YC13 Paradza, D., & Daramola, O. (2021). Business intelligence and business value in organisations: A systematic literature review. Sustainability , 13 (20), 3–27. https://doi.org/10.3390/su132011382 Park, J. H., & Kim, Y. B. (2021). Factors Activating Big Data Adoption by Korean Firms. Journal of Computer Information Systems , 61 (3), 285–293. https://doi.org/10.1080/08874417.2019.1631133 Paul, J., Modi, A., & Patel, J. (2016). Predicting green product consumption using theory of planned behavior and reasoned action. Journal of Retailing and Consumer Services , 29 , 123–134. https://doi.org/10.1016/j.jretconser.2015.11.006 Pavlou, P., & Fygenson, M. (2006). Understanding and Predicting Electronic Commerce Adoption: 264 An Extension of the Theory of Planned Behavior. MIS Quarterly , 30 (1), 115–143. Penna, G. O., Silva, J. A. A. da, Neto, J. C., Temporão, J. G., & Pinto, L. F. (2020). PNAD COVID19: um novo e poderoso instrumento para Vigilância em Saúde no Brasil. Ciência & Saúde Coletiva , 25 (9), 3567–3571. https://doi.org/10.1590/1413-81232020259.24002020 Petter, S., Delone, W., & McLean, E. R. (2013). Information systems success: The quest for the independent variables. Journal of Management Information Systems , 29 (4), 7–62. https://doi.org/10.2753/MIS0742-1222290401 Phillips-Wren, G., Iyer, L. S., Kulkarni, U., & Ariyachandra, T. (2015). Business Analytics in the Context of Big Data: A Roadmap for Research. Communications of the Association for Information Systems , 37 , 448–472. https://doi.org/10.17705/1CAIS.03723 Picoto, W. N., Crespo, N. F., & Carvalho, F. K. (2021). The influence of the technology-organizationenvironment framework and strategic orientation on cloud computing use, enterprise mobility, and performance. Review of Business Management , 23 (2), 278–300. https://doi.org/10.7819/rbgn.v23i2.4105 Pillai, R., & Sivathanu, B. (2020). Adoption of internet of things (IoT) in the agriculture industry deploying the BRT framework. Benchmarking , 27 (4), 1341–1368. https://doi.org/10.1108/BIJ-08-2019-0361 Podsakoff, P. M., MacKenzie, S. B., & Podsakoff, N. P. (2016). Recommendations for Creating Better Concept Definitions in the Organizational, Behavioral, and Social Sciences. Organizational Research Methods , 19 (2), 159–203. https://doi.org/10.1177/1094428115624965 Pramana, S., Mariyah, S., & Takdir. (2021). Big data implementation for price statistics in Indonesia: Past, current, and future developments. Statistical Journal of the IAOS , 37 (1), 415–427. https://doi.org/10.3233/SJI-200740 Pullinger, J. (1997). The Creation of the Office for National Statistics. International Statistical Review , 65 (3), 291–308. https://doi.org/10.1111/j.1751-5823.1997.tb00310.x Queiroz, M. M., & Farias Pereira, S. C. (2019). Intention to adopt big data in supply chain management: A Brazilian perspective. Revista de Administração de Empresas , 59 (6), 389– 401. https://doi.org/10.1590/S0034-759020190605 Quetelet, A. (2013). A Treatise on Man and the development of this faculties . Cambridge University Press. http://www.cambridge.org/9781108064422 Radermacher, W. (2018). ESS vision and ways for cooperation Workshop on strategic 265 developments in business. Issue: January. https://doi.org/10.13140/RG.2.2.32812.77448 Radermacher, W. J. (2020). Official Statistics 4.0 . Springer International Publishing. https://doi.org/10.1007/978-3-030-31492-7 Ram, J., Afridi, N. K., & Khan, K. A. (2019). Adoption of Big Data analytics in construction: development of a conceptual model. Built Environment Project and Asset Management , 9 (4), 564–579. https://doi.org/10.1108/BEPAM-05-2018-0077 Ramadoss, R., & Elango, N. M. (2015). Proactive exploratory testing methodology during enterprise application modernization. International Journal of Engineering and Technology , 7 (2), 673– 681. Ravin, Y., & Leacock, C. (2000). Polysemy: an overview. In Y. Ravin & C. Leacock (Eds.), Polysemy: Theoretical and computational approaches (pp. 1-29.). Oxford Press Inc. República Federativa do Brasil. (2018). Lei Geral de Proteção de Dados Pessoais (LGPD) . http://www.planalto.gov.br/ccivil_03/_Ato2015-2018/2018/Lei/L13709.htm Retana, G. F., Forman, C., Narasimhan, S., Niculescu, M. F., & Wu, D. J. (2018). Technology support and post-adoption IT service use: Evidence from the cloud. MIS Quarterly: Management Information Systems , 42 (3), 961–978. https://doi.org/10.25300/MISQ/2018/13064 Rogers, E. M. (1962). Diffusion of innovations (1st ed.). Free Press of Glencoe. Rogers, E. M. (1983). Diffusion of innovations (3rd ed.). Free Press of Glencoe. Rogers, E. M. (1995). Diffusion of innovations (4th ed.). Free Press. Rogers, E. M. (2004). A prospective and retrospective look at the diffusion model. Journal of Health Communication , 9 , 13–19. https://doi.org/10.1080/10810730490271449 Royal Statistical Society. (2014). The Data Manifesto. https://rss.org.uk. Royal Statistical Society. (2022). History . About. https://rss.org.uk/about/history/ Saheb, T. (2020). An empirical investigation of the adoption of mobile health applications: integrating big data and social media services. Health and Technology , 10 (5), 1063–1077. https://doi.org/10.1007/s12553-020-00422-9 Saldaña, J. (2013). The Coding Manual for Qualitative Researche. In The Coding Manual For Qualitative Researhers (Secpond). SAGE Publications Ltd. Salleh, K. A., & Janczewski, L. (2016). Technological, Organizational and Environmental Security and Privacy Issues of Big Data: A Literature Review. Procedia Computer Science , 100 , 19– 28. https://doi.org/10.1016/j.procs.2016.09.119 266 Salleh, K. A., Lech, J., & Beltran, F. (2015). SEC-TOE Framework : Exploring Security Determinants in Big Data Solutions Adoption. Pacific Asia Conference on Information Systems , 1-11. https://aisel.aisnet.org/pacis2015/203/ Santos, C., Santos, V., Tavares, A., & Varajão, J. (2020). Project Management in Public Health : A Systematic Literature Review on Success Criteria and Factors. Portugues Journal of Public Health . https://doi.org/10.1159/000509531 Sarker, S., Sarker, S., & Sidorova, A. (2006). Understanding Business Process Change Failure: An Actor-Network Perspective. Journal of Management Information Systems , 23 (1), 51–86. https://doi.org/10.2753/MIS0742-1222230102 Sarker, S., & Valacich, J. (2010). An Alternative to Methodological Individualism: A NonReductionist Approach to Studying Techology Adoption by Groups. MIS Quarterly , 34 (4), 779–808. Saunders, M., Lewis, P., & Thornhill, A. (2009). Reserach Method for Business Students (5th). Schiavoni, C., Palm, F., Smeekes, S., & van den Brakel, J. (2021). A dynamic factor model approach to incorporate Big Data in state space models for official statistics. Journal of the Royal Statistical Society. Series A: Statistics in Society , 184 (1), 324–353. https://doi.org/10.1111/rssa.12626 Schwab, K. (2016). The Fourth Industrial Revolution (1st ed.). Crown Business. Scott, S. V., & Wagner, E. L. (2003). Networks, negotiations, and new times: the implementation of enterprise resource planning into an academic administration. Information and Organization , 13 (4), 285–313. https://doi.org/10.1016/S1471-7727(03)00012-5 Shahbaz, M., Gao, C., Zhai, L., Shahzad, F., & Hu, Y. (2019). Investigating the adoption of big data analytics in healthcare: the moderating role of resistance to change. Journal of Big Data , 6 (1), 6. https://doi.org/10.1186/s40537-019-0170-y Shahbaz, M., Gao, C., Zhai, L., Shahzad, F., & Arshad, M. R. (2020). Moderating Effects of Gender and Resistance to Change on the Adoption of Big Data Analytics in Healthcare. Complexity , V. 2020, 1–13. https://doi.org/10.1155/2020/2173765 Sheng, J., Amankwah-Amoah, J., & Wang, X. (2017). A multidisciplinary perspective of big data in management research. International Journal of Production Economics , 191 (November 2016), 97–112. https://doi.org/10.1016/j.ijpe.2017.06.006 267 Silva, J., Hernández-Fernández, L., Torres Cuadrado, E., Mercado-Caruso, N., Rengifo Espinosa, C., Acosta Ortega, F., Hernández P, H., & Jiménez Delgado, G. (2019). Factors affecting the big data adoption as a marketing tool in SMEs. Communications in Computer and Information Science , 1071 , 34–43. https://doi.org/10.1007/978-981-32-9563-6_4 Silveira, D. (2023, January 25). IBGE prevê para abril divulgação dos resultados definitivos do Censo 2022. Portal G1, Economia . https://g1.globo.com/economia/noticia/2023/01/25/ibge-preve-para-abril-divulgacaodos-resultados-definitivos-do-censo-2022.ghtml Sinha, A., Kumar, P., Rana, N. P., Islam, R., & Dwivedi, Y. K. (2019). Impact of internet of things (IoT) in disaster management: a task-technology fit perspective. Annals of Operations Research , 283 (1–2), 759–794. https://doi.org/10.1007/s10479-017-2658-1 Soto-Acosta, P. (2020). COVID-19 Pandemic: Shifting Digital Transformation to a High-Speed Gear. Information Systems Management , 37 (4), 260–266. https://doi.org/10.1080/10580530.2020.1814461 Serviço Regional de Estatística dos Açores. (2023). O que somos O que fazemos Como fazemos . About Us. https://srea.azores.gov.pt/Conteudos/Relatorios/lista_relatorios.aspx?idc=307 &idsc=6209&lang_id=1” Stake, R. (1995). The Art of Case Study Research . Sage Publications, Inc. Statistics Canada. (2020). The Integration of Web-Scraped Data into the Clothing and Footwear Component of the Consumer Price Index (Issue 62). Stigler, S. (1986). The history of statistics: The measurement of uncertainty before 1900 . Harvard University Press. Sun, S., Cegielski, C. G., Jia, L., & Hall, D. J. (2018a). Understanding the Factors Affecting the Organizational Adoption of Big Data. Journal of Computer Information Systems , 58 (3), 193– 203. https://doi.org/10.1080/08874417.2016.1222891 Sun, S., Hall, D. J., & Cegielski, C. G. (2020). Organizational intention to adopt big data in the B2B context: An integrated view. Industrial Marketing Management , 86 (September 2019), 109– 121. https://doi.org/10.1016/j.indmarman.2019.09.003 Sun, Z., & Huo, Y. (2021). The Spectrum of Big Data Analytics. Journal of Computer Information Systems , 61 (2), 154–162. https://doi.org/10.1080/08874417.2019.1571456 Taylor, S., & Todd, P. A. (1995). Understanding information technology usage: A Test of Competing Models. In Information Systems Research, 6 (2), 144–176. 274 19. [OPENED QUATION 1] 20. [OPENED QUATION 2] 21. Are there other external factors to support or constrain Big Data technologies in the [NSO_NAME] ? [Closing sentence: "I will stop the record now. Thank you so much for your time and answers.”]