Full text
Rui Pedro Oliveira Machado março de 2020 UMinho | 2020 Readout Circuit for Time-Based Automotive Sensors Universidade do Minho Escola de Engenharia Rui Pedro Oliveira Machado Readout Circuit for Time-Based Automotive Sensors
março de 2020 Tese de Doutoramento Programa Doutoral em Sistemas Avançados de Engenharia para a Indústria Trabalho efetuado sobre a orientação de Professor Doutor Jorge Miguel Nunes dos Santos Cabral Professor Doutor Luís Alexandre Machado Rocha (à memória de) Rui Pedro Oliveira Machado Readout Circuit for Time-Based Automotive Sensors Universidade do Minho Escola de Engenharia
Readout Circuit for Time-Based Automotive Sensors ii DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/
Readout Circuit for Time-Based Automotive Sensors iii ACKNOWLEDGMENTS Accepting the PhD challenge was one of the hardest decisions that I have ever made. Not giving up was even harder. This journey has brought up the best and the worst of me, allowing me to know myself better and to become a better person. Throughout this emotional rollercoaster, I was only able to maintain my perseverance due to some people, to whom I want to thank for the time, support and valuable teachings. To Dr. Jorge Cabral, my supervisor which I consider as a friend, thanks for guiding and caring about me throughout my entire academic path. To Dr. Luís Alexandre Rocha, may he rest in peace, thanks for everything. I will always remember you as an example of what a true scholar should be. Special thanks to Filipe Serra Alves, for being my supervisor without any obligation. I have learnt so much from you, both professionally and personally. I hope one day I can repay you for everything you did for me. Thanks to INL for receiving me. Thanks to ESRG for being an amazing group. Thanks to Fundação para a Ciência e Tecnologia (FCT) and Bosch Car Multimedia for financing my PhD (grant PD/BDE/114562/2016). Thanks to all my friends, Pedro (Pimba), Luís Novais, Tiago Vasconcelos, Vasco Lima, Eurico Moreira, Miguel Azevedo, Tânia Moreira and everyone from LAR group. Special thanks to Juca, for the Saturday chill out coffees and gaming sessions. A Special thanks to my best friend, Antero, for the snooker games on Saturday afternoons remembering the times when we were kids, and for the trust and strength. To my mother and father, a heartful thanks for the love, education, sacrifice and all the rest. I am who I am thanks to you. For all it is worth, you are the best parents one could have wished for. To my beloved sister, thank you for making me laugh so many times, you may not see it, but you are truly important. I am proud of who you have become. To my grandfather and role model, Manuel Machado, my source of strength and guardian angel, to whom I dedicate this thesis. Wherever you are, I hope you are proud. Finally, to the love of my life, Juliana Martins. Thank you for everything you did for me the last 10 years. Whatever the future holds for us, please know that you are definitely the best thing that ever happened to me. You changed me in so many ways, and always for the better. Thank you for never doubting me, for seeing in me what nobody else sees, for truly knowing how and who I am. Thanks for making me believe that everything is possible and for encouraging me to always challenge myself. A friend, a girlfriend, a safe haven. You are my everything and only with your caress was I able to surpass this challenge. Rui Machado, Guimarães, March 9th, 2020.
Readout Circuit for Time-Based Automotive Sensors iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho.
Circuito de Leitura para Sensores Automóveis baseados em Tempo de Voo v RESUMO A pesquisa pelo veículo autónomo (AV) iniciou-se há já algumas décadas, com a introdução de vários sistemas inteligentes nos veículos do nosso quotidiano. O melhor exemplo deste tipo de sistemas são os Advanced Driver-Assistance Systems (ADAS). As grandes marcas da indústria automóvel e principais Original Equipment Manufacturers (OEMs) estão focados no desenvolvimento do primeiro AV. O Light Detection And Ranging (LiDAR) é considerado como uma tecnologia chave para implementação do AV. O sistema de leitura e medição do tempo de voo ( Time-of-Flight - ToF) é um dos subsistemas constituintes do sensor LiDAR, e assume extrema importância. Os sistemas de medição de ToF de alto desempenho são normalmente implementados recorrendo ao desenho de células lógicas específicas e customizadas, o que leva a um aumento do tempo de desenvolvimento do sistema e, consequentemente, do custo. Estes tipos de sistemas apresentam desempenhos superiores aos necessários e o seu nível de integração é reduzido. O desenvolvimento de sistemas de medição de ToF capazes de serem completamente desenhados por linguagens de descrição de hardware (HDL) e implementados através de um fluxo de desenvolvimento totalmente automatizado permitirá alcançar maior portabilidade e níveis de integração. O propósito desta tese é o desenvolvimento e implementação de uma arquitetura para um sistema de medição de ToF, capaz de facilitar o processo de migração destes sistemas entre tecnologias e plataformas. As arquiteturas existentes foram analisadas e foram implementadas e avaliadas múltiplas arquiteturas recorrendo a plataformas de prototipagem. Para assegurar um processo de migração fluído, as ferramentas de desenho de Application Specific Integrated Circuit (ASIC) foram estudadas. Como resultado, foi desenvolvido um sistema de medição de ToF para aplicações automóveis LiDAR e estabelecido um fluxo de desenvolvimento que suporta a migração automatizada de arquiteturas ToF. O contributo da presente tese baseia-se no estudo sobre como devem ser desenhados e implementados os sistemas de medição de ToF para permitirem um fluxo de desenvolvimento automatizado e aumentar a sua portabilidade e integração, mantendo o desempenho necessário em aplicações automóveis LiDAR. A investigação iniciou-se com uma revisão do estado da arte em sistemas de medição de ToF, que culminou no desenvolvimento de duas arquiteturas em Field-Programmable Gate Array (FPGA) e na fabricação de um ASIC utilizando fluxo de desenvolvimento proposto baseado em Structured Data Path (SDP) para migrar arquiteturas de medição de ToF baseados em FPGA para tecnologia ASIC. Palavras-Chave: ASIC, FPGA, LiDAR, Tempo de voo, Time-to-Digital Converter (TDC).
Readout Circuit for Time-Based Automotive Sensors vi ABSTRACT The pursue for autonomous car started long ago with multiple smaller and smarter systems being researched and introduced gradually in our daily vehicles. The so-called Advanced Driver Assistance Systems (ADAS) are the best example of these stepwise process towards a full-autonomous vehicle (AV). Nowadays, all the major automotive groups and Original Equipment Manufacturers (OEMs) are pursuing the goal of launching an AV on the market. Light Detection And Ranging (LiDAR) is considered the key enabling technology to implement this AV. LiDAR sensor are composed by a multitude of systems and components, being the time-of-flight (ToF) readout system of extreme importance. State-of-the-art high-performance ToF readouts are implemented using highly customized cells, which increases both development time and costs. The performance achieved with these solutions usually highly exceeds the required for LiDAR and their level of integration is also reduced. The development of ToF measurement architectures capable of being completely described using hardware description languages (HDL) and implemented using a full automatized design flow, will help to attain reduced development time, increased portability and a high level of integration. This Thesis aims to develop and implement a ToF readout architecture to simplify the migration effort between platforms and technologies. Existing architectures are analyzed and, based on the acquired knowledge, multiple architectures developed and assessed using a fast prototype platform. To ensure a seamless migration process, the tools used on Application Specific Integrated Circuits (ASIC) development are studied, and the design flow steps that can be automated or supported by scripting are identified. The accomplishment of these activities enabled the development of a ToF measurement system for automotive LiDAR sensors and a design flow guideline and respective supporting scripts. The present Thesis contribute by reasoning about how should a ToF measurement system be designed and implemented to enable a full automated design flow process, increase portability and integration, while maintaining the required performance for automotive LiDAR based systems. The research started with an exhaustive literature review on FPGA ToF measurement systems, which lead to the implementation of two FPGA-based architectures, and to the fabrication of an ASIC TDC using the proposed Structured Data Path (SDP) design flow for migrating FPGA-based ToF systems to ASIC technology. Keywords: ASIC, FPGA, LiDAR, Time-of-Flight (ToF), Time-to-Digital Converters (TDC).
Readout Circuit for Time-Based Automotive Sensors vii INDEX Acknowledgments ............................................................................................................................... iii Resumo............................................................................................................................................... v Abstract.............................................................................................................................................. vi Index ................................................................................................................................................. vii Index of Figures .................................................................................................................................. xi Index of Tables .................................................................................................................................. xv List of Abbreviations and Acronyms .................................................................................................... xvi 1. Introduction ............................................................................................................................... 1 1.1. Contextualization and Problem statement ............................................................................ 2 1.2. Motivation, scope and Research Questions ........................................................................ 11 1.2.1. Research Questions and Objectives ............................................................................ 12 1.2.2. Research methodology ............................................................................................... 13 1.2.3. Research Development Timeline ................................................................................ 16 1.3. Contributions .................................................................................................................... 17 1.4. Thesis Organization ........................................................................................................... 18 References ................................................................................................................................... 21 2. Time-based Readout Circuits .................................................................................................... 23 2.1. Performance Metrics ......................................................................................................... 25 2.1.1. Dynamic range .......................................................................................................... 26 2.1.2. Resolution and Precision ............................................................................................ 26 2.1.3. Non-linearity .............................................................................................................. 27 2.1.4. Dead Time................................................................................................................. 28 2.1.5. Power Consumption, Area and Resource Usage ......................................................... 28 2.2. TDC Architectures ............................................................................................................. 29 2.2.1. Coarse Counter Architectural Group ........................................................................... 29 2.2.2. Analog TDC ............................................................................................................... 30 2.2.3. Phased Clocks ........................................................................................................... 31 2.2.4. Tapped-Delay Lines (TDL) and Delay-Locked-Loops (DLL) ........................................... 35 2.2.5. Time-Amplifier (TA) TDCs ........................................................................................... 43 2.2.6. Differential Delay Lines .............................................................................................. 44 2.2.7. Pulse Shrinking Architecture ...................................................................................... 48 2.2.8. Summary .................................................................................................................. 49
Readout Circuit for Time-Based Automotive Sensors xiv Figure 5.12Code Density Test Results for Temperature Corner Cases ............................................ 172 Figure 5.13Single-Shot Precision variation with Temperature ......................................................... 173 Figure 5.14Corrected Single-Shot Precision Variation with Temperature using Equation 5.3 ........... 174 Figure 6.1Block Diagram of Possible Integration Scenarios ............................................................ 186 Figure 6.2Block Diagram of the Proposed TDC and Processing Unit integration ............................. 186
Readout Circuit for Time-Based Automotive Sensors xv INDEX OF TABLES Table 1.1Vision Technologies Comparison ......................................................................................... 5 Table 1.2 -LiDAR Commercial Devices ................................................................................................ 8 Table 2.1 - TDC Literature Review Summary ..................................................................................... 54 Table 2.2TDCs Commercial Devices Summary ................................................................................ 60 Table 3.1Synchronizer Correction Factors ....................................................................................... 89 Table 3.2Gray-Code Datapath Delay Analysis ................................................................................. 102 Table 3.3Routing Delays for the Gray-Code TDC ............................................................................ 109 Table 4.1 - TSMC clock digital cells propagation delay analysis ........................................................ 129 Table 4.2Gray-Code Path Propagation Delay Analysis .................................................................... 133 Table 4.3I/O PAD Configuration .................................................................................................... 153 Table 5.1Example of the Sampled Values of Figure 5.9 ................................................................. 169 Table 5.2Single-Shot Precision Values along the Studied Temperature Range ................................. 173 Table 6.1State-of-the-art Comparison ............................................................................................ 181
Readout Circuit for Time-Based Automotive Sensors xvi LIST OF ABBREVIATIONS AND ACRONYMS ADAS Advanced Driver-Assistance Systems ADC Analog-to-Digital Converter ADPLL All-Digital Phased Locked Loop ASIC Application Specific Integrated Circuit AV Autonomous Vehicle AXI Advanced eXtensible Interface BMW Bayerische Motoren Werke BSP Board Support Package CCOpt Clock Concurrent Optimization CCS Composite Current Source CLB Configurable Logic Block CTS Clock Tree Synthesis DAC Digital-to-Analog Converter DLL Delay Locked Loop DNL Differential Non-Linearity DR Dynamic Range DRC Design Rules Check DSP Digital Signal Processor ECU Engine Control Unit EOF End of File
Readout Circuit for Time-Based Automotive Sensors xvii FIFO First-In First-Out FPGA Field-Programmable Gate Array GM General Motors HDL Hardware Description Language IC Integrated Circuit INL Integral Non-Linearity IP Intellectual Property LEF Library Exchange Format LiDAR Light Detection And Ranging LSB Least Significant Bit LUT Look-Up Table LVS Layout Versus Schematic MCMM Multi Corner Multi Mode MEMS Microelectromechanical Systems MSB Most Significant Bit NDR Non-Design Rules NLDM Non-Linear Delay Model OCV On-Chip variation OEM Original Equipment Manufacturer PET Positron Emission Tomography PL Programmable Logic
Readout Circuit for Time-Based Automotive Sensors xviii PLL Phased Locked Loop PS Processing System PVT Process Voltage Temperature RADAR Radio Azimuth Direction And Ranging RAM Random Access Memory RTL Register Transfer Level SAE Society of Automotive Engineers SDC Synopsys Design Constraints SDF Standard Delay Format SDP Structured Data Path SoC System on Chip SPI Serial Peripheral Interface TA Time Amplifier TCL Tool Command Language TDC Time-to-Digital Converter TDL Tapped Delay Line ToF Time-of-Flight TSMC Taiwan Semiconductor Manufacturing Company VCO Voltage Controlled Oscillator WU Wave-Union XDC Xilinx Design Constraints
Readout Circuit for Time-Based Automotive Sensors 1 1. Introduction The world, and the understanding humanity has of it, is constantly changing. New technologies are developed daily to address people’s needs and ever-increasing demands, changing the reality in which we live. Automotive industry always played an important role on this technology advancement. Over the last few years, a collection of systems has been introduced in vehicles, aiming to improve the overall driving experience, increase drivers’ safety and reduce driving hazards due to human error. The current trend fueling automotive research is the autonomous vehicle (AV) concept, a technology that, when achieved, will completely shift the driving paradigm. The realization of such concept relies on the ability to endow a vehicle with capabilities to percept its surrounding environment. Thus, vison systems such as Radar Cameras and LiDAR are currently in the spotlight of research, being LiDAR pointed out as one of the core technologies, propelling numerous research works worldwide. In this chapter, an introductory concept of this Thesis is presented. First, this work’s motivation is explained, contextualizing its pertinence, formulating the problem statement and defining the Thesis scope. Then, the research questions are devised, and the objectives and methodologies defined. Finally, the main contributions to the scientific state-of-the-art achieved during this Thesis are listed and the structure of the remainder of this Thesis is presented. The Chapter is organized as follows: Section 1.1 formalizes the problem statement and presents the project’s dependent requirements and constraints; Section 1.2 motivates this Thesis, describes the Thesis’ scope and presents its research questions, targeted objectives and proposed methodologies to
1. Introduction 2 attain them, while answering the formulated questions. Section 1.3 presents this Thesis contributions and Section 1.4 describes the structure and organization of this document. 1.1. Contextualization and Problem statement Some decades ago, if someone talked about a car, it would be describing an almost full mechanical system. It was in 1882 that Karl Benz first patented the Benz Patent-Motorwagen. Then, in 1908, the first automobiles became accessible to the masses with the famous Model T, commercialized by Ford. Over the years, futuristic insights about driverless cars, pushed the automotive industry further, resulting in impressive technological improvements. This pursue lead to the incremental appearance of various functionalities, supported by various sensors and electronics, which reduced the driver’s workload and increased safety. Nowadays, a variety of sensors can be found in cars, among them, angular sensors used in the pedals to measure the throttle position, pressure sensors used to measure the fuel and boost pressure, temperature sensors, imaging sensors, etc. Recently, sensors like LiDAR, RADAR and Cameras, are gaining relevance due to its role on sensing the car surrounding environment, which is mandatory when implementing an Automated Vehicle (AV). For instance, in 2018, there were already 53 companies working on LiDAR technology for automotive in California [1.1]. Also, in 2016, the demand for all kind of sensor solutions was increasing [1.2]. Electronics have always played an important role on automotive industry, improving safety, automatizing some driving tasks in a controlled environment, and generally improving the driving experience with infotainment systems. Although the acceptance and introduction of Advanced Driving Assistance Systems (ADAS) was slow [1.3], nowadays they are part of everyday driving experience and are already considered as a must have system depending on the tier of the car. According to [1.3], an increase of 50% on the number of ADAS systems included in cars was verified in just two years, from 2014 to 2016, with the inclusion of surrounding view systems having an increase of more than 150%, in the same time span. ADAS are playing an important role on teaching people about autonomous driving and its revenues may be used to finance AV research [1.4], [1.5]. Nevertheless, this will only be possible if the adoption rate of ADAS on mass production vehicles increases [1.4], [1.5]. This highlights the need for cheaper, but still high reliable systems [1.3], [1.6]. The ever-increasing adoption of ADAS, together with the recent search for a full automated driving experience is continuously pushing electronic sensors and systems forwards. With sensors technologies holding the key for the future of automobile industry [1.7], research on low
Readout Circuit for Time-Based Automotive Sensors 3 cost and more reliable sensors and readout mechanisms used in ADAS systems and AV is required to tackle the new challenges ahead in the competitive automotive industry. According to the Society of Automotive Engineers (SAE), a full autonomous vehicle, level 5 in Figure 1.1, must be capable of controlling every aspect of the driving task in the entire imaginable scenarios. This means that electronic sensors and systems must be capable of performing even in the most severe conditions [1.8]. When no user interaction is allowed, as in the case of a level 5 vehicle, the electronics requirements drastically change compared to a typical ADAS. According to Goel [1.9], solving the challenges introduced by AV will require better hardware, to collect more data and with higher precision, and better software to analyze and make decisions based on the gathered data. Moreover, this new hardware will also have to be able to cope with the typical automotive requirements, i.e. low power, low area, resistance to harsh environment, etc. Fulfilling all these requirements is not a trivial task, and multiple OEMs have already invested millions of dollars and made partnerships trying to be the leaders on this new and emerging AV market. Companies like Honda, Ford, GM, Toyota, Volvo, Hyundai, BMW and Tesla have already announced their intention of having fully autonomous vehicles on the road between 2020 and 2030 [1.1], [1.8]. Figure 1.1Levels of Automation (according to SAE) Although being relatively new in automotive markets, the vehicle surroundings mapping advantages offered by LiDAR, when compared to the more established technologies (i.e. cameras, ultrasonic sensors and radar), have propelled massive innovations. Thus, LiDAR is already considered as a key enabling technology for achieving AVs [1.7], [1.10]. This popularity increase has also enabled a dizzying drop on LiDAR’s cost, from around US$50000 to US$10000, with forecasts predicting a lower than US$200 cost per LiDAR module in 2022 [1.10].
1. Introduction 4 The current scenario on ambient mapping solutions make use of Radar for longand short-range measurements, video cameras for medium range and ultrasound for short-range measurements, mainly used on parking assistance ADAS [1.11]. The introduction of LiDAR may replace and/or complement some of these technologies, resulting in a schema like the one depicted on Figure 1.2. Figure 1.2Car Vision Technologies (based on [1.3]) Ultrasonic sensors are only viable for short range measurements since the effects of attenuation are strong beyond a few meters distance. Furthermore, although ultrasonic sensors resolution may be suitable for object detection, it is not for objected identification. Although LiDAR can also perform in short range measurements and give detailed data on the object shape, facilitating the object identification, its cost is still way above of an ultrasonic sensor. Therefore, LiDAR sensors research mainly targets medium to long range measurement applications. Other technologies targeting medium to long range measurements are Cameras and Radar [1.10]. Although Cameras offer a cost-effective solution due to its availability, the data processing power required, in order to extract useful information from the captured data, is a drawback of this technology. Furthermore, cameras are highly sensible to ambient light conditions. The only parameter in which LiDAR cannot outperform cameras is road signs and color detection. Thus, time critical detection tasks are better performed by LiDAR and cameras can be used to complement the information acquired by it, using a sensor fusion approach. Most of LiDAR current applications could also be addressed by Radar. Since Radar solutions are available at lower cost and are easy to integrate, due to smaller size, this technology is the one used on modernday vehicles. However, LiDAR research enabled LiDAR solutions to be shrunk over the years and the recent industry shift to solid-state LiDAR will further enhance integration and reduce costs [1.10]. Therefore, LiDAR is now capable of compete with Radar since it offers a set of performance improvements
Readout Circuit for Time-Based Automotive Sensors 5 like larger distance, better angular resolution and larger field-of-view. This enables an improved object classification, with higher resolution in a broader scene/frame, without significant backend processing. The possibility to cover short, medium, and long range with a single sensor is attractive. Radar relies on the technology used to cover different ranges, meaning that a combination of technologies per each range must be made. Although LiDAR performance degrades with adverse weather conditions, when compared with Radar which offers a robust performance even under heavy rain, snow and fog, the use of 1550nm wavelength enables LiDAR to reach acceptable performance values [1.10]. An overview on the performance comparison between LiDAR, Radar and Cameras, based on the data reported by Mizuho Securities USA, and presented at the AutoSens 2017 conference held in Brussels [1.12], is given in Table 1.1. Table 1.1Vision Technologies Comparison LiDAR Radar Camera Range Best Best Worst Field of View Best Better Worst Width and Height Best Worst Worst 3D Shape Best Worst Worst Object recognition at long range Best Worst Worst Rain, Snow, Dust Best Best Worst Night Best Best Worst Signs and Color Worst Worst Best Source: AutoSens 2017, “LiDAR systems for automotive: Benefits and the challenges for OEMs” [1.12]. LiDAR working principle is based on transmitting a pulsed or continuous light signal (generated by a laser beam) which will be reflected by the different objects at the scene being scanned (Figure 1.3). By measuring the characteristics of the reflected signal, a high-fidelity picture of the scene being illuminated can be reconstructed. The usual parameters used in LiDAR measurements are the pulse’s power, timeof-flight (ToF), and phase shift of the received signal [1.7], [1.10], [1.13]. The maximum achievable range for LiDAR measurement is presented in [1.13], and can be calculated according to the following equation: 𝑅𝑎𝑛𝑔𝑒= √𝑃∗𝐴∗𝑇𝑎∗𝑇𝑜 𝐷𝑠∗𝑃𝑖∗𝐵 , (1.1)
1. Introduction 12 concluded in a three-year time window, and there must be enough time for system integration and testing, it was defined that the ToF measurement unit research and implementation had to be done in less than three years. Moreover, apart from the time constraint, the ToF measurement unit must be easily integrated within the LiDAR sensor and easily ported between prototyping platforms (used in the initial stages of the project) and the final product platform (the final system is intended to be implemented in ASIC). Nowadays, solutions for high performance time-to-digital conversion usually imply a custom-made process which increases both project’s costs and development time. Thus, the aim of this Thesis is to develop a time interval measurement readout system to perform the Time-to-Digital Conversion (TDC) in a LiDAR sensor, capable of achieving high performance, with reduced customization and highly automated design flow processes, to accelerate development and reduce the system’s cost. The main principle is based on being capable of directly migrating a fully digital, synthesizable TDC architecture, implemented in a lowcost prototype platform, i.e. FPGA, to an ASIC with minimum intervention. The main motivation for this Thesis consists on the possibility to work with a large set of tools for hardware and software development which will allow the author to improve its knowledge. Furthermore, prior to the start of this work, there was no knowledge regarding TDCs development in the author’s research group or Bosch Car Multimedia, which rises the challenge even further. Finally, the tools used to design digital systems are developed to be efficient in optimizing and analyzing synchronous designs, and therefore, the design of a system in which the relevant information is the one hidden between clock cycles, and that requires a precise characterization of the real circuit timings and not only the worst and best case scenarios is, by itself, a large and interesting challenge. 1.2.1. Research Questions and Objectives The following research questions were formulated in order to guide this research, making possible to reach the aforementioned objective: • RQ1: What are the current research trends on ToF measurement systems for LiDAR sensors? • RQ2: Which architectures are simultaneously suitable for FPGA deployment (fast prototyping) and ASIC implementation, while maintaining the required performance for LiDAR sensors?
Readout Circuit for Time-Based Automotive Sensors 13 • RQ3: During the porting process (of the selected ToF measurement architecture) from FPGAbased to ASIC technology, how to minimize the development effort and time? • RQ4: How does the developed ToF measurement solutions (FPGA and ASIC) compare with the current state-of-the-art for LiDAR sensors? The following sub-objectives were defined in order to gradually pursue the main goal of this Thesis, while answering the Thesis’ research questions: • O1: Review the state-of-the-art in time interval measurement systems; • O2: Develop different architecture working prototypes to gain insight on the main challenges and technical limitations when developing high performance time interval measurement systems; • O3: Evaluate the developed architectures performance to understand the scenarios in which they are viable, and how should the ToF measurement systems characterization be performed; • O4: Study the digital design flow for ASIC technology; • O5: Create scripts to configure and automatize the ASIC digital design flow; • O6: Implement a time interval measurement peripheral, addressing the issues identified during the state-of-the-art review and prototype development; • O7: Evaluate the proposed architecture, development flow and implemented peripheral and position it within the existing state-of-the-art. 1.2.2. Research methodology To focus and guide the activities involved in the research process that would enable the attainment of the proposed objectives (O1 to O7) and answer the formulated research questions (RQ1 to RQ4), several research methodologies, from RM1 to RM8, were adopted during this PhD Thesis. RM1 – State-of-the-art of TDC: A review on the state-of-the-art of time interval measurement systems was performed on both academia and commercial fields to understand which type of architectures are being implemented, what are its typical performances, and which applications are being targeted. The review focused on architectures implemented in FPGA since these are the ones that can easily be ported to ASIC due to its intrinsic digital nature. The results of this study were published in article J1 [1.14], which proposes a taxonomy for FPGA-based TDCs classification, and identifies research gaps and new approaches that are not yet explored and may be an important contribution to build TDC systems. This
1. Introduction 14 study was performed to address RQ1 and RQ2, targeting O1 and providing a solid field knowledge, crucial to target the remaining research questions. Afterwards, a study on the available ASIC architectures was made to understand the typical design flow used on ASIC TDC development and the performance values achieved. RM2 – Study the FPGA development framework: Developing time interval measurement systems, with resolution under the clock system frequency, requires a deep understanding on how the development frameworks are configured, as well as advanced knowledge of the FPGA platform being used. Therefore, a preliminary study of the Xilinx Vivado framework was made in order to: test multiple optimization configurations; assess how to avoid automatic optimizations on parts of a design; force the framework to generate specific hardware directly mapped to the FPGA configurable logic blocks (CLB); and understand how manual layout (placement and routing) could be applied to parts of the design, to improve the overall time interval measurement system’s performance. The FPGA platforms used were studied to learn which type of resources were available and how should them be configured. This enabled to establish a solid expertise required to target O2 and O3. RM3 – Development of FPGA-based TDC prototypes: To understand the challenges and issues involved on TDCs development, two different architectures were implemented in a FPGA device, namely the Xilinx Z7010. The first architecture was implemented with the objective of achieving the highest possible resolution, while the second one was designed targeting low resource utilization. Both architectures were designed to ensure portability and scalability. Details of the first architecture were presented in a conference proceeding C1 [1.15], with special focus on the synchronization block. Later, the same synchronization block was updated and a design methodology for synchronizer blocks was developed and presented in another conference proceeding C2 [1.16]. The details of the second architecture are partially presented in publication P1. The architecture consists on a modified version of the TDC presented by Wu and Xu in [1.17], which sacrifices the maximum achievable resolution in order to obtain improved linearity and homogeneity when multiple time interval measuring channels are implemented, and low resource and power consumption is required. Using this methodology, RQ2 was addressed and objective O2 achieved. RM4 – Evaluate FPGA-based TDCs: The evaluation and characterization of the developed architectures was done to address RQ2 and target O3. The characterization process enabled a better understanding of the main metrics that need to be considered for proper TDC assessment, namely, which tests need to be performed and how to perform them. Two tests were performed on both architectures. The first test was
Readout Circuit for Time-Based Automotive Sensors 15 a code density test with 100 thousand samples, made to retrieve the TDC’s mean resolution and differential and integral non-linearity. After, a performance measurement with 100 thousand samples was executed to obtain the TDC’s precision based on its standard deviation. A third test could be done to analyze the TDCs performance variation with temperature. However, since the objective of the FPGA implementations is to study the viable solutions for ASIC porting, and because the performance temperature variation is highly dependent on the technology used, there is no significant advantage in performing such test. Furthermore, according to [1.18], the Xilinx Zynq 7000 FPGA family does not have a significant performance variation with temperature. Therefore, the temperature tests were only made with the final ASIC TDC implementation. The results of the tests performed for the first and second architecture are presented in the conference proceeding C1 [1.15] and in publication P1, respectively. RM5 – Study the ASIC design flow tools: The development tools for ASIC design are not the same as the ones for FPGA. Moreover, while using FPGA, a single framework with all the included tools was available. Although there are frameworks comprising every tool needed for an ASIC digital design flow, it is a de facto standard in industry that for synthesis, Synopsys’ DesignCompiler offer the better results while Cadence’s Innovus performs better during system layout. These tools are from different vendors, therefore the data transferred from one stage to the other must be handled by the ASIC designer. Moreover, as in the FPGA case, the tools are optimized to handle synchronous designs and to perform hardware intensive optimizations to reduce area, power consumption and design complexity. Those kinds of functionalities can only be applied to parts of the design (the synchronous part), while the module responsible for measuring sub-clock timings must be left out by the tool. Consequently, a thoughtful study on the tools used during digital ASIC design flow was performed, to understand how to manage the data exchange between tools, how to correctly configure the tools, how to manually manipulate the layout, and how to automate the different processes involved. With this methodology O4 was completed while addressing RQ3 and building the required knowledge to target O5 and O6. RM6 – Migrate and Evaluate the FPGA-based TDCs to ASIC: Although being designed to be easily ported between platforms, the results from a TDC architecture porting could lead to unbearable performance drops. Therefore, the implemented FPGA TDC architectures were ported to ASIC and a pre-evaluation on its performance was made, to understand which of the architectures would be able to produce better results. The migration process implies the creation of multiple design flow scripts, to configure Synopsys and Cadence IC design Tools and to generate the final architecture layout, used during fabrication and final test simulations. From the obtained results, the first architecture was chosen to be fabricated. The
1. Introduction 16 porting results of the two architectures are presented in this Thesis on Chapter 4. These results address RQ3 and RQ4, while enabling to achieve O5 and start progressing into O6 and O7. The details on the ASIC design flow methodology adopted are reported in article J2 [1.19], while in article J3 [1.], the migrated architecture is presented in detail, together with the description of the scripts used to configure the tools during the synthesis and layout process. RM7 – Tape-out the ASIC TDC peripheral: the final process before ASIC fabrication is known as tape-out. This process consists on the final chip layout configurations, i.e. creation of the ASIC’s pad-ring, DRC (Design Rules Check) and LVS (Layout Vs Schematic) checks, and final timing simulation tests. During this process, a Printed Circuit Board (PCB), used to integrate and assess the ASIC TDC was designed and developed. Details on the entire ASIC design flow and PCB development are presented in article J3. These activities address RQ3 and enabled to complete O6, producing all the required resources to address RQ4 and achieve O7. RM8 – Characterize and integrate the fabricated TDC: In order to be able to compare the implemented TDC with available implementations, a set of tests was performed to characterize the device. The main parameters are the TDC resolution, non-linearity, precision and temperature performance drifts. The first two parameters can be obtained by a code density test, similar to the ones performed on the FPGA-based TDCs prototypes. The precision and temperature drift tests were done in a thermal chamber, in a range from 0 to 50 Celsius degrees. A set of Matlab scripts were developed to analyze the data from the performed tests. The obtained results allow the comparison of the developed TDC with the existing solutions. The architectures comparison, described in Chapter 6, were done using the main TDc performance metrics, described in Chapter 2. This comparison addresses RQ4 while enabling to achieve O7. 1.2.3. Research Development Timeline In order to better understand which research methodologies were applied to address the formulated research questions and how were the proposed objectives targeted, the research timeline is depicted in Figure 1.5. The RQs are represented as circles and used has the starting point to achieve a single or group of objectives, also represented as a circle. The connection is done by a rectangle specifying which methodology or set of methodologies were used to answer the research questions, accomplishing the objectives.
Readout Circuit for Time-Based Automotive Sensors 17 Figure 1.5Thesis Timeline 1.3. Contributions To support the development of this Thesis and to validate its scientific contribute, in the field of TDCs and ASIC design flow, the following publications were submitted to peer-reviewed international indexed conferences and journals: Journal Papers: • R. Machado, J. Cabral and F. S. Alves, "All-Digital Time-to-Digital Converter Design Methodology Based on Structured Data Paths," in IEEE Access, vol. 7, pp. 108447-108457, 2019. doi: 10.1109/ACCESS.2019.2933496 • R. Machado, J. Cabral and F. S. Alves, "Recent Developments and Challenges in FPGA-Based Time-to-Digital Converters," in IEEE Transactions on Instrumentation and Measurement, vol. 68, no. 11, pp. 4205-4221, Nov. 2019. doi: 10.1109/TIM.2019.2938436 Conference Papers: • R. Machado, L. A. Rocha and J. Cabral, "A novel synchronizer for a 17.9ps Nutt Time-to-Digital Converter implemented on FPGA," 2018 AEIT International Annual Conference, Bari, 2018, pp. 1-6. doi: 10.23919/AEIT.2018.8577365 • R. Machado, J. Cabral and F. Alves, "Designing Synchronizers for Nutt-TDCs," 2019 5th International Conference on Event-Based Control, Communication, and Signal Processing (EBCCSP), Vienna, Austria, 2019, pp. 1-6. doi: 10.1109/EBCCSP.2019.8836914 RM1 Review Apr-17 Jul-17 Oct-17 Jan-18 Apr-18 Jul-18 Oct-18 Jan-19 Apr-19 Jul-19 Oct-19 Jan-20 FPGA-Based TDC Prototypes Evaluation ASIC TDC Migration Characterization RQ1 System Results Publications O1 Priority Level RM1 RM2 RM3 RM4RQ2 O2 O3 TDC Prototype 1 TDC Prototype 2 Literature Review Design Flow Methodology ASIC TDC Readout System J1 J2C1 C2 RM5 RM6 RM7RQ3 O4 O5 O6 J3P1 RM8RQ4 O7
1. Introduction 18 Under Revision: • R. Machado, F. Alves, A. Geraldes, J. Cabral, “Technology Independent ASIC based Time to Digital Converter”, submitted IEEE Transactions on Circuits and Systems I: Regular Papers • R. Machado, F. Alves, J. Cabral, “Gray-Code TDC Architecture with Improved Linearity and Scalability”, submitted 2020 6th International Conference on Event-Based Control, Communication, and Signal Processing (EBCCSP) 1.4. Thesis Organization The remainder of this Thesis is structured as follows (Figure 1.6): Chapter 2 introduces the basic concept regarding TDCs, providing a theoretical background to understand the design decisions made throughout this Thesis work. First, the main performance metrics are explained. After, the most popular architectures, implemented in FPGA and ASIC platforms, are described and discussed in detail. For each architecture, the most relevant research works present on the literature are mentioned. Since this Thesis targets a system with improved integration and portability, more emphasis is given to architectures that can be implemented in a fully digital system. A description of the FPGA prototype platform and respective development environment is presented in Chapter 3, followed by the description of the selection process regarding which TDC architectures should be explored. Two architectures were selected considering the requirements for automotive LiDAR applications. These architectures modifications and support modules are described in detail, highlighting the benefits and identifying its limitations. A discussion on the performance assessment results, for each architecture, is presented at the end of the chapter. The main conclusions of this chapter support the decisions made throughout the development process of the final ASIC TDC. Chapter 4 introduces the development tools for digital ASIC design. Afterwards, the complete system implementation process is presented, describing the required architectural changes, the details of the scripted development design flow that enables an almost seamless migration from FPGA to ASIC platform, and the expected TDC performance, inferred from the simulation results and extracted timing information.
Readout Circuit for Time-Based Automotive Sensors 19 The main experimental results of the developed TDC are presented in Chapter 5. Chapter 5 starts by describing the performed tests and experimental setup used. Then, the performance of the TDC is assessed, identifying its limitations and discussing the needed improvements. Chapter 6 concludes this Thesis and summarizes the acquired knowledge and future research work for the time interval measurement device is proposed either to enhance its performance or its integration level.
1. Introduction 20 Figure 1.6Thesis Structure Chapter 2: Time-Based Readout Circuits Chapter 6: Conclusions & Future Work Chapter 5: Experiemtal Results Chapter 4: ASIC-based TDC Design and Developement Chapter 3: FPGA-based TDC Prototypes Chapter 1: Introduction Why are time based readout systems important to Vision Sensors? ASICFPGA Prototypes of the choosen architectures What are the interfaces? How do they work? How to address the issues? What are the main issues? Prototypes Characterization How was the prototype migrated? Layout changes Synthesis changes Interface Changes Results Does it work? Discusion Conclusion Future Improvements MicroProcessor Integration ROM-based TDC case of study Reduced Synchronizer TDC version Automotive Vision Technologies Innovation Scenario What is the Current Scenario on Time-Based Readout systems? How are they characterized? Survey of the State-of-the-Art Which is the most suitable archtecture for ASIC? What is the expected performance?
Readout Circuit for Time-Based Automotive Sensors 21 References [1.1] M. Avary, “3 autonomous vehicle trends to follow in 2019,” 2019. [Online]. Available: https://www.weforum.org/agenda/2019/01/3-autonomous-vehicle-trends-to-follow-in-2019/. [Accessed: 22-Jul-2019]. [1.2] N. Tyler, “Demand for automotive sensors is booming,” 2016. [Online]. Available: http://www.newelectronics.co.uk/electronics-technology/automotive-sensors-market-isbooming/149323/. [Accessed: 16-Oct-2019]. [1.3] K. Heineke, P. Kampshoff, A. Mkrtchyan, and E. Shao, “Self-driving car technology: When will the robots hit the road?,” 2017. [Online]. Available: https://www.mckinsey.com/industries/automotive-and-assembly/our-insights/self-driving-cartechnology-when-will-the-robots-hit-the-road. [Accessed: 22-Jul-2019]. [1.4] A. Padhi and P. Kampshoff, “Autonomous-driving disruption: Technology, use cases, and opportunities,” 2017. [Online]. Available: https://www.mckinsey.com/industries/automotiveand-assembly/our-insights/autonomous-driving-disruption-technology-use-cases-andopportunities. [Accessed: 22-Jul-2019]. [1.5] P. Gao, H.-W. Kaas, D. Mohr, and D. Wee, “Disruptive trends that will transform the auto industry,” 2016. [Online]. Available: https://www.mckinsey.com/industries/automotive-andassembly/our-insights/disruptive-trends-that-will-transform-the-auto-industry. [Accessed: 22-Jul2019]. [1.6] S. Choi, F. Hansson, H.-W. Kaas, and J. Newman, “Capturing the advanced driver-assistance systems opportunity,” 2016. [Online]. Available: https://www.mckinsey.com/industries/automotive-and-assembly/our-insights/capturing-theadvanced-driver-assistance-systems-opportunity. [Accessed: 22-Jul-2019]. [1.7] V. Hiligsmann, “How sensor technology will shape the cars of the future,” 2017. [Online]. Available: https://www.melexis.com/en/insights/knowhow/how-sensor-technology-shape-carsfuture. [Accessed: 22-Jul-2019]. [1.8] J. Walker, “The Self-Driving Car Timeline – Predictions from the Top 11 Global Automakers,” 2019. [Online]. Available: https://emerj.com/ai-adoption-timelines/self-driving-car-timelinethemselves-top-11-automakers/. [Accessed: 22-Jul-2019]. [1.9] A. Goel, “What Tech Will it Take to Put Self-Driving Cars on the Road?,” 2016. [Online]. Available: https://www.engineering.com/DesignerEdge/DesignerEdgeArticles/ArticleID/13270/WhatTech-Will-it-Take-to-Put-Self-Driving-Cars-on-the-Road.aspx. [Accessed: 22-Jul-2019]. [1.10] M. Khader and S. Cherian, “An Introduction to Automotive LIDAR.” Texas Instruments, 2018. [Online]. Available: http://www.ti.com/lit/wp/slyy150/slyy150.pdf. [1.11] D. Bronzi, Y. Zou, F. Villa, S. Tisa, A. Tosi, and F. Zappa, “Automotive Three-Dimensional Vision Through a Single-Photon Counting SPAD Camera,” IEEE Trans. Intell. Transp. Syst., vol. 17, no. 3, pp. 782–795, Mar. 2016. [1.12] AutoSens. LIDAR Systems for Automotive: Benefits and the Challenges for OEMs. (02-Mar-2018). Accessed: 16-Oct-2019. [Online Video]. Available: https://www.youtube.com/watch?v=lnyXQ3IiTBA. [1.13] P. McCormack, “LIDAR System design for Automotive/Industrial/Military Applications.” Texas Instruments, p. 10, 2011. [1.14] R. Machado, J. Cabral, and F. S. Alves, “Recent Developments and Challenges in FPGA-Based Time-to-Digital Converters,” IEEE Trans. Instrum. Meas., vol. 68, no. 11, pp. 4205–4221, Nov. 2019.
2.Time-based Readout Circuits 28 Figure 2.3Linearity metrics 2.1.4. Dead Time The dead time of a TDC is the time interval required, from the arrival of the stop signal, until the TDC is ready to perform a new measure. This metric is highly dependent on the TDC architecture. For instance, TDCs based on tapped delay lines usually report dead times equal to one reference clock period [2.14], while pulse shrinking [2.15] or ring oscillators [2.16] architectures can have a dead time dependent on the time interval to be measured, which can reach several hundreds of nanoseconds. Since high sample rate is required for modern applications, architectures capable of achieving low dead times are becoming popular. A common practice to reduce dead time and increase sample rate is to use multiple TDC channels multiplexed, measuring the same input signal in an interleaved schema [2.7]. 2.1.5. Power Consumption, Area and Resource Usage Power and area play an important role on modern applications since the current mobile trend requirements focus on low power and high levels of integration. In digital systems, the power is usually characterized as dynamic (switching) or static (leakage). The first is directly related to the operation frequency of the system, while the second one is technology dependent. Another technology dependent characteristic is the system’s area or used resources, depending on whether the system is being implemented in ASIC or FPGA, respectively. Smaller ASIC technologies does not necessarily mean that Digital Code Time Interval Ideal Transfer Function 0.5 LSB 11111 00000 First Transition Last Transition Ideal Code Center Ideal Transition Point (50%) 2LSB code Missing Code (Bubble) Ideal Digital Transfer Function Real Transfer Function INL
Readout Circuit for Time-Based Automotive Sensors 29 the same TDC architecture can be implemented in a smaller area, since the delay elements need to be redesigned and it is not always possible to shrink the cells. However, smaller technologies can definitely enable higher performances to be achieved, at the expense of higher fabrication costs. Modern FPGA technologies also enable higher performance values at higher platform costs. FPGA architectures size are measured in resources utilization count rather than on area dimensions as in the ASIC scenario. In the case of FPGAs, the selected platform can also constrain the type of TDC architectures that can be implemented. Regarding the FPGA platforms used during this Thesis, when the term resources is used, it will be referring to: the FPGA’s Configurable Logic Blocks (CLB) elements, namely, registers, Carry4 and Look-Up tables (LUTs); the FPGA’s BRAM blocks; the FPGA’s PLL blocks; and the FPGA DSP blocks. The type of resource being used will be enumerated whenever pertinent. 2.2. TDC Architectures TDCs are highly dependent on the available resources and/or technology in use. While in ASIC platforms there is theoretically no constraint to the implementation of any TDC architecture, FPGA platforms limit the range of implementable architectures. Therefore, contrarily to FPGA where architectures share a lot of similarities, in ASIC platforms it is often possible to find completely new approaches. In the following sections, the TDC architectures are presented and analyzed, divided accordingly to their principle of operation. Nevertheless, the differences and nuances between the FPGA-based and ASIC-based TDC architectures are presented, whenever the implementation is possible on both platforms. The most relevant and distinct FPGA and ASIC architectures are presented from section 2.2.1 to section 2.2.7. 2.2.1. Coarse Counter Architectural Group Course counters are TDC implemented using binary-, gray-code or ripple counters, which are incremented by a reference clock. The works in [2.17]–[2.19] are examples of coarse counters’ implementation. The main advantage of such architectures is the simplicity of the design, its portability and the low resources utilization in FPGAs or small area in ASIC. Nevertheless, the achievable resolution is bounded to the frequency of the reference clock used. For instance, in order to achieve a resolution equal to 1 nanosecond, a 1 GHz clock is required. The course counters can be implemented using a free running schema, in which the clock is always enabled, and the time event to be measured, usually denoted as hit, is responsible for sampling the value in the counter registers. Coarse counters can also be
2.Time-based Readout Circuits 30 implemented using an enable schema, in which the hit signal is responsible for enabling and disabling the counter. The range and resolution of this architecture is given by: 𝑅𝑎𝑛𝑔𝑒=2𝑛, (2.7) 𝜏= 1 𝑓𝐶𝐿𝐾, (2.8) where n is the number of bits of the counter register and fCLK is the system clock frequency. When multiple bit coarse counters are implemented, the routing of the hit signal for the counter’s registers must be carefully planned, regardless of its use as enable or sampling signal. Otherwise errors greater than 1 LSB may occur, especially when a binary code schema is used. Although the operating frequencies of nowadays FPGAs and ASICs technologies are higher, for example, the works in [2.20] and [2.21] reported the use of 500 MHz and in [2.22] a 710 MHz reference clock, for resolutions under the nanosecond scale, these architectures are still not suitable. Furthermore, even if it was possible to use a high frequency reference clock to achieve picosecond resolutions (a clock higher than 10 GHz would be required), it would be harder to secure low skew values between the signals for the counter’s registers. The largely enhanced skew effects increase the risk for metastability and counting errors. Therefore, the Coarse Counter architecture should only be employed when resolutions of a few nanoseconds and high measurement ranges are required. Recently, Wu and Xu proposed a Gray code counter without sampling stage [2.23]. This enables the operating frequency of the counter to be approximately equal to the cells’ plus routings’ propagation delays. Since the counting schema used is the Gray code, only one bit is changing at a time, which eliminates the possibility for incorrect counting sequence due to delays mismatch between the counting cells. 2.2.2. Analog TDC Analog TDCs are usually built using a Time-to-Analog converter (TAC), followed by an Analog-to-Digital Converter (ADC). In these types of TDCs, a capacitor is charged by a fixed current source. The basic architecture is depicted in Figure 2.4. The amount of charge stored in the capacitor is proportional to the time the capacitor was charged. The final digital value is obtained using an ADC to convert the charge value into a digital value. The dynamic range (DR) of this architecture is given by equation (2.9):
Readout Circuit for Time-Based Automotive Sensors 31 𝐷𝑅=2𝑛∗𝑇𝐿𝑆𝐵, (2.9) where n is the number of bits of the ADC and TLSB is the resolution of the ADC. Although resolutions under 50 ps are possible, the use of a capacitor and a controlled current source, greatly increase the area and power consumption [2.4]. Moreover, this type of TDCs are highly susceptible to temperature drifts. More details on this architecture can be found in literature [2.4], [2.24]. Recently, a 50 ps resolution and precision analogue TDC has been implemented in a 110 nm process technology by Cossio [2.25]. Figure 2.4Analog TDC overview 2.2.3. Phased Clocks An alternative to achieve resolutions in the range of a few hundreds of picoseconds is the use of TDCs based on phased clocks. These architectures are simple to implement and require very low resources when implemented in FPGA, since FPGAs platforms have PLL blocks integrated [2.26]–[2.30]. In ASIC, if a PLL is needed to generate the different clock phases, then the complexity of the architecture increases. Therefore, this type of architecture is especially advantageous in FPGA platforms. Besides being relatively easy to implement, phased clocks architectures have good linearity and can achieve resolutions better than 300 picoseconds [2.29]. However, when compared to other high performance TDCs, the resolution of phased clocks continues to be one of its main drawbacks. Phased clocks architectures are based on two main techniques: oversampling and phase detection. GND C ADC Start/Stop Reset Δ t = (C.Vc)/I1 I1 I2 Start Stop Δ t Digital Time Interval Output
2.Time-based Readout Circuits 32 1) Oversampling: this architecture is based on using phased clocks as reference clocks to independent counters. The time event to measure is used as the counters’ enable, just like in the case of coarse counters architecture. In fact, this approach is identical to the previous one, replicated m times, where m is the number of phased clocks used, and consequently the number of independent counters. It is important to mention that these phases have to be generated from the same reference clock and ideally, equally spaced. To determine the measurement value, equation (2.10) can be used: 𝑡𝑂𝑣𝑒𝑟𝑠𝑎𝑚𝑝𝑙𝑖𝑛𝑔 =(𝑛0+𝑛1+⋯+𝑛𝑛)∗𝑇𝐶𝐿𝐾 𝑚, (2.10) where n0 to nm represent the number of counts in each counter. These numbers are added and then multiplied by the TDC’s resolution (given by the reference clock period TCLK , divided by the number of phases used). Resolutions equal to 1 nanosecond have been reported when using this architecture in [2.26]–[2.28], [2.31]. The main implementation challenge is related to the routing of the signal to be measured. The phase difference between the generated clocks is responsible for defining the clocks resolution. Therefore, the signal to be measured must be routed with the minimum skew possible between counters to avoid degrading the measurement. Also, as the number of used phases increases, the effect of jitter accumulates, degrading the TDC’s performance. Therefore, to avoid phase overlap, special attention should be given to the design of the phase generation mechanism. 2) Phase detection: phase detection architectures sample the event to be timed with multiple phases (see Figure 2.5). The output of the sample process is a unique code dependent on which was the phase that first detected the event. The resolution (LSB) of these architectures is given by the phase difference between the multiple phases. As in the previous case, with the increase of the number of generated phases, the performance of the TDC tends to degrade due to the errors associated with the phase generation. With the increase of the frequency used as reference clock, and the routing skews, jitter and uncertainty of the phase generation, scenarios where the phase m arrives before phase m+1 can occur, resulting in bubble errors, which lead to missing codes, jeopardizing the TDC’s performance. Again, the phase generator mechanism assumes high relevance and its design must be carefully done when targeting sub-nanosecond resolution. A synchronization stage, like the one depicted in Figure 2.5, is also needed to assure the creation of a common clock domain allowing the correct sampling of the code pattern by the reference clock, before it can be used to determine the instant of arrival of the time event.
Readout Circuit for Time-Based Automotive Sensors 33 The synchronization module number of stages and resource utilization increases with the number of phases implemented. Opposed to what happens in the coarse counters and oversampling architectures, phase clocks based on phase detection offer great resolution and linearity but are not suitable for large dynamic ranges. To address this issue, it is common to extend the dynamic range of this architecture, complementing it with a coarse counter. In this way, the phase detection module just needs to cover the time equivalent to a full reference clock period while the coarse counter covers time intervals greater than the reference clock. Using the setup depicted in Figure 2.5, this TDC architecture can be described by the following equations: 𝜏= 1 𝑁𝑝ℎ𝑎𝑠𝑒𝑠, (2.11) 𝑇𝑓𝑖𝑛𝑒 =𝑇𝐶𝐿𝐾 −(𝑝ℎ𝑎𝑠𝑒∗𝜏), (2.12) 𝑇𝐷𝐶𝑚𝑒𝑎𝑠𝑢𝑟𝑒 =𝑛∗𝑇𝐶𝐿𝐾 +𝑇𝑓𝑖𝑛𝑒, (2.13) where phase is the clock phase’s number that sampled the input signal (from 0, for the 0˚ phase clock, to m , to the m˚ phase clock), and τ is the resolution given by the phase difference between clocks. Figure 2.5Phased Clocks based architecture (adapted from [2.29]) Research works, that use this architecture in FPGA platforms, report great linearity values without implementing calibration mechanisms [2.29], [2.32], [2.33]. A resolution of 625 picoseconds was Phase Capture Synchronizer D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK D Q CLK Fine Time Channel Buffer D Q CLK 270° 180° 90° 0° Coarse time 440MHz 440MHz 880MHz Quad Phase Clocks Measured value hit
2.Time-based Readout Circuits 34 reported in [2.32], the recent research work by Sano et al. [2.29] reports a 280 picoseconds (LSB) resolution with high linearity, being the DNL errors below 0.5 LSB. In [2.33], a precision of 56 picoseconds and resolution below 156 picoseconds was achieved, proving the potential of phase detection architectures. These results corroborate that high performance TDCs with lower design complexity and resource utilization can be achieved using phased detection architectures. Nevertheless, it is important to understand that the high linearity is also related to the relatively large LSB size. As the number of phases generated increase, the size of the LSB gets smaller and the errors associated with the clock’s phase generation and routing paths get more pronounced, thus deteriorating the TDC’s linearity. Phase detection interpolation is not a common architecture in ASIC technology. Although it has the advantage of being a pure digital architecture which simplifies the design and implementation process. The performance achieve by such architectures cannot compete with more sophisticated ones, like DLLs and pulse shrinking. Nevertheless, the linearity values reported by phased clocks are usually in the range of less than 0.2 LSB, making this architecture very attractive for applications where resolution in the range of a few hundreds of picoseconds is required. The resource consumption or area utilization per TDC channel is also reduced when compared to other TDC architectures. Since the architecture can be fully implemented in a digital flow, this architecture is a good candidate for ASIC migration. For the same reason, FPGA and ASIC platforms usually share the same phase detection architectures and issues. The main difference between FPGA and ASIC phased clocks architectures is the PLL block, which is already included in modern FPGA but, in ASIC platforms, must be designed and implemented, increasing the overall TDC architecture complexity. Recently, the work by Wang et al. [2.34] proposed a phased detection architecture where, instead of sampling the time event with the phased clocks, it was the phased clocks which were sampled by the time event (Figure 2.6). This enables for power savings since the sampling process will only occur once per time event. Furthermore, it removes the need of the synchronization stage, as presented in Figure 2.5, since the sampling is done in the same clock domain, enabling area savings. The TDC was implemented in a 130 nm process technology, reporting a 780 ps resolution, with a maximum bin variation of +/- 40 ps, corresponding to a DNL of +/-0.05 LSB.
Readout Circuit for Time-Based Automotive Sensors 35 Figure 2.6ASIC Phased Clocks (Adapted from [2.34]) 2.2.4. Tapped-Delay Lines (TDL) and Delay-Locked-Loops (DLL) Even though in ASIC a multitude of architectures exist, DLL architectures and variants are one of the most popular since they enable high resolution, which can be further improved using n-stage interpolation schemas, for FPGA platforms, the most researched and adopted architecture to achieve high performance TDCs is based on tapped-delay-lines (TDL) [2.7], [2.8], [2.21], [2.31], [2.35]-[2.104]. The design process of these architectures may seem simple however, when high linearity and high performance is required, special precautions must be taken during their implementation phase. Tapped-Delay Lines A typical TDL architecture is composed by an input stage, a delay line paired with a sample stage, a decoding block, and a calibration block, as depicted in Figure 2.7. The first and last blocks, the input stage and the calibration stage, are not mandatory. However, with technology scaling down and therefore, a higher impact of PVT variations in the propagation delays of the cells, if high precision is required, these blocks must be implemented. Otherwise, the non-linearity errors will deteriorate the TDC performance. The basic principle of operation consists in delaying an event signal throughout a chain of buffers (delay cell or interpolation step). The state of the delay chain is sampled by a reference clock every cycle. This results in a thermometer code that has the information of the number of delay cells that the event signal was capable of traverse in between reference clock cycles. This code is then passed to a decoder that will convert this thermometer code to a binary value corresponding to number of delay cells traversed. This value can be directly used to recover the event timing information, or it can be passed to a calibration table as an index to obtain the calibrated time information. The input stage can be used to manipulate Q Q SET CLR D 0° Q Q SET CLR D 90° Q Q SET CLR D 180° Q Q SET CLR D 270° hit Q Q SET CLR D Q Q SET CLR D Q Q SET CLR D Q Q SET CLR D q0 q1 q2 q3 Clk 0°
2.Time-based Readout Circuits 36 the time event in order to generate multiple pulses, allowing multiple transitions to be sampled in the delay line, which, using of statistical methods, contributes to an increase of the TDCs precision. The critical blocks are the pair composed by the delay line and the sample line (Figure 2.8), since these define the base resolution of the TDC. The components presented in Figure 2.7 are usually implemented per TDL channel. In multiple chains TDCs, the base blocks are replicated for each channel. For a single TDL channel, the time interval measurement value can be obtained according to: 𝑡𝑓𝑖𝑛𝑒 =𝑛∗𝜏, (2.14) where τ is the propagation delay of each cell element and n is the number of cells traversed by the delayed signal. Figure 2.7TDL TDC Block Diagram Figure 2.8TDL architecture In TDL architectures, the maximum achievable resolution is always dependent on the propagation delay of the cell used as the basic delay element (TDL step) and the clock skew to the sampling registers [2.59]. Input Stage Delay Line Sample Line Thermometer-to-Binary Decoder Calibration hit ... ... Measurement Value CLK Delay Line Sample Line D Q CLK τ τ τ D Q CLK D Q CLK hit CLK ... ... ... TDL step/bin (delay cell) T0 T1 Tn ...
Readout Circuit for Time-Based Automotive Sensors 37 In ASIC-base implementations, these basic delay cells are usually custom designed so that a specific propagation delay time is achieved. In FPGA-based implementations, one of the available cell resources must be selected. In modern FPGA there are two cells that can potentially be used to implement delay lines, the look-up tables (LUTs) or the Carry4 cells. Carry4 cells offer lower propagation delay and have dedicated routing paths, which are desired characteristics when designing a TDL TDC and therefore, a vast number of reported implementations used these resources as the TDL base step. Nevertheless, there are still some implementations of TDLs using LUTs [2.44], [2.52], which were able to achieve interesting resolutions. In [2.105], flip-flops were used to implement the TDL steps, using the output of a previous flip-flop to clock the next one. Although good linearity has been achieved, the resolution was not as good as the one achieved when using Carry resources to implement the TDL step. Regardless of the cell used to build the delay chain steps (Carries, LUTs, flip-flops, or custom designed cells), when implementing a TDL architecture, the following issues, divided by variation, must be addressed. 1) Single TDL: TDLs are usually designed to cover a time interval equal to the period of a reference clock. When larger dynamic range are required, the TDL architecture is paired with coarse counters, otherwise a very large number of steps would be required. Since the arrival of the time event to be measured by the TDC is asynchronous to the coarse counting mechanism, the TDL and coarse counter must be synchronized for proper operation and to avoid metastability. Otherwise, the performed measurement may have an error of several coarse counter clock periods. For this reason, a second coarse counter, with a clock signal delayed by 180º is usually implemented and the value outputted by the TDL (sampled value) is used to identify which counter has the correct value, i.e. the value that is not metastable. A metastability error occurs when the hit signal arrives close to the rising edge of the reference clock (used to increment the coarse counter). Therefore, if the TDL value sampled is close to zero or to the maximum TDL value, there is a chance for metastability on the coarse counter. In these scenarios, the value on the second coarse counter, incremented by the 180º delayed clock, is used. When multiple transitions per time event are to be measured by the same TDL, the described method does not work [2.106]. In those scenarios, a methodology like the one proposed in [2.107] should be adopted. Since the resolution of TDLs is attached to the delay element used, statistical methods are usually employed to overcome this limitation. In the case of single TDLs, this is done using a Wave-Union (WU) launcher, which was first proposed by Wu [2.108]. Research works reporting a 10 ps resolution and 38 ps precision, using a Lattice FPGA to implement a WU TDL can be found in [2.80] and [2.109]. The
2.Time-based Readout Circuits 44 FPGA platforms. This enables to achieve lower resolutions since the effective resolution will be the LSB of the TDC divided by the amplification factor [2.120]. A resolution of 980 fs in a 65 nm technology have been reported in [2.120]. Molaei and Hajsadeghi [2.121] also implemented a TA TDC, achieving a 5.3 picoseconds precision in a 180 nm process technology. 2.2.6. Differential Delay Lines An alternative to improve TDC resolution under the intrinsic propagation delay of a cell is to adopt a differential approach. These types of architectures have a resolution equal to the difference between the delay step of two elements. This is obtained by, for example, delaying both the time event to be measured using a TDL and delaying the clock signal for the registers that are sampling the TDL (usually called 2D TDL or Vernier TDL). This way, the resolution is given by the difference between the cells used in the TDL and the cells used to delay the clock signal (see Figure 2.13 where the resolution of the TDC is given by τ1-τ2). TDCs based on the Vernier principle can also be considered as an interpolative TDC, similar to the TDL and DLL architectures. Nevertheless, since the resolution of the TDC is given not by a single delay element, but rather by the difference between two interpolative stages, these architectures were considered as differential, according to the taxonomy proposed in [2.122]. Another approach using two ring oscillators with slightly different frequencies could also be adopted, as presented in [2.123] and depicted in Figure 2.16. Figure 2.13Differential Delay line Architecture Slow Delay Line Sample Line D Q CLK τ1 τ1 τ1 D Q CLK D Q CLK hit CLK ... ... Fast Delay Line τ2 τ2 τ2 ... T0 T1 Tn ... τ 1>τ2 Different Delay cells
Readout Circuit for Time-Based Automotive Sensors 45 1) 2D TDL: On a 2D TDL, depicted in Figure 2.11, the resolution of the TDC is calculated according to equations (2.17) and (2.18). During implementation it must be guaranteed that the propagation delay of the cells used to build the clock delay chain is lower than the propagation delay of the cells used in the hit delay chain. Otherwise, the clock signal will never be able to catch up with the hit signal and the TDC will always return the maximum value. Another issue with this architecture is the length of the delay chains. As in the case of TDLs, the longer the chain is, the higher will be the error due to non-linearities. Furthermore, because it must also be guaranteed that an entire reference clock is covered by the delay line, and because the size of the step in 2D TDLs is lower than in simple TDLs, the delay chains will be longer. This requires more resources and decreases the TDCs precision due to the accumulation of nonlinearity errors across the chains. 𝜏=𝜏1−𝜏2, (2.17) 𝑡𝑓𝑖𝑛𝑒 =𝑛∗𝜏, (2.18) For FPGA implementation, the choice for different cells is limited and therefore, when implementing 2D TDLs, the routing is used to obtain chains with slightly different delays steps. Therefore, getting a uniform delay difference across the two delay chains is a demanding process, hard to replicate. For these reasons, regarding differential TDCs in FPGA platforms, the ring oscillators are often more popular due to its simpler implementation. On the other hand, in ASIC implementation, it is possible to design different delay cells, achieving better resolutions. The research works in [2.124] and [2.125] reported a 30 ps and 5 ps LSB resolution, respectively. 2) Ring Oscillators: Ring oscillator architectures are usually implemented using TDLs in a loop schema [2.112]. Because the parameter responsible for defining the TDC resolution is the difference between the two ring oscillators, the cells’ delay mismatch issue is solved [2.112] (see Figure 2.16 and equations (2.19) to (2.22)). The main challenge is to obtain a high accurate and stable oscillation period in order to achieve high performance [2.123]. Two different implementations for ring oscillators TDC were proposed over the last few years for FPGA platforms [2.123], [2.126]. The first one is based on two counters incremented each by one oscillator and a phase detector which samples the counters when the phases of both oscillators align [2.123] (see Figure 2.14). The values on the counters is then added and multiplied by the TDC resolution. The measurement value can be calculated according to equation (2.20): 𝜏=𝑇1−𝑇2, (2.19)
2.Time-based Readout Circuits 46 𝑇𝑟𝑖𝑛𝑔𝑂𝑠𝑐𝑖𝑙𝑙𝑎𝑡𝑜𝑟𝑇𝐷𝐶 =(𝑛1−1)∗𝑇1−(𝑛2−1)∗𝑇2, (2.20) 𝑡𝑚𝑎𝑥𝐶𝑜𝑛𝑣 =𝑇1∗𝑇2 𝜏, (2.21) being n1 the slow and n2 the fast counters’ value respectively. Accordingly, T1 and T2 are the slow and fast oscillators’ periods. The waveform diagram on Figure 2.15 depicts a typical measurement process for this type of architecture. The main drawback of the architecture is its long conversion time that can reach several hundreds of nanoseconds. Figure 2.14Ring Oscillator with two independent counters Figure 2.15Different frequency oscillators Vernier Architecture waveform The other architecture is based on a single counter and a phase detector. Once the phase of the two oscillators align, the value on the counter is sampled and the measurement value is obtained by multiplying the counter value for the TDCs resolution as demonstrated in equation (2.22): 𝑡𝑟𝑖𝑛𝑔𝑂𝑠𝑐𝑖𝑙𝑙𝑎𝑡𝑜𝑟𝑇𝐷𝐶 =𝑛𝑓𝑖𝑛𝑒 ∗𝜏, (2.22) Startable Oscillator 1 Startable Oscillator 1 Coincidence Detector Counter 1 Counter 2 Q Q SET CLR D Q Q SET CLR D Start Stop Enable f1 f2 n1 n2 CoincidenceStart Stop T1=1/f1 T2=1/f2 n2T2 n1T1 tDifferentialTDC
Readout Circuit for Time-Based Automotive Sensors 47 being nfine the number of counts in the counter until the fast oscillator is able to overtake the slow one, and τ the resolution of the TDC (given by equation (2.19)) [2.126] (see Figure 2.16). Figure 2.16Ring Oscillator with single counter The main drawback of these architectures is the occurrence of pulse shrinking/stretching phenomenon that may happen when implementing ring oscillators that propagate a pulse. If this is not addressed the oscillation behavior will cease. Therefore, a pulse reshaping mechanism, like the one presented in [2.16] and [2.112], needs to be implemented. Ring oscillators architectures are also known for long measurement dead times. The time required to finalize a conversion can be calculated using equation (2.21). Based on it, depending on the time interval to be measured, the conversion time can reach several microseconds, which limits both the throughput of the TDC and the acceptable input rate. It is possible to decrease the conversion time by reducing the resolution of the TDC or by increasing the oscillation frequency of the ring oscillators. Neither are good solutions since high resolution is usually desirable and maintaining a stable oscillation behavior at high frequencies is not a trivial task. Ring oscillators topologies in ASIC share the same structure as the ones implemented in FPGA platforms. As most of the ASIC architectures that can be implemented in FPGA, the difference resides on the type of cell used to build the module responsible for defining the TDC resolution, in the case of the ring oscillators. Nguyen et al. [2.127] presented a ring oscillator TDC with 377 ps LSB resolution and 0.8 LSB precision in a 0.18 µm technology. Slow Ring Oscillator hit_sync Pulse width reshaping Pulse width reshaping D Q CLK Fast Ring Oscillator Cascaded Carry Chain Pulse width reshaping Pulse width reshaping τ τ ... τ τ ... clk clk_sync clear Counter en clr Ctrl2 Generator nfine ctrl2
2.Time-based Readout Circuits 48 2.2.7. Pulse Shrinking Architecture The pulse shrinking phenomenon, cause by the mismatch between the rise and fall times for the cells used in a delay chain [2.128], although being an undesirable effect when trying to propagate a pulse in a ring oscillator, can be used to implement a high resolution TDC. Basically, if a counter is being incremented at each ring oscillator cycle, and if a pulse is propagating and being shrunk every cycle, the number of counted cycles until the oscillation behavior ceases will be proportional to the size of the pulse. Therefore, the resolution is given by the so-called shrinking factor, i.e. the amount the pulse that is being propagated is shrunk every cycle. This shrinking factor is given by the sum of the difference between the rise and fall times of every cell in the looped delay line that is implementing the ring oscillator. These TDCs, although achieving high resolution values, have a high measurement deadtime that is proportional to the pulse size that is being measured. Furthermore, there is an offset associated to the measurement since near the end of the measured, although the pulse is still circulating in the ring, it no longer has the capability of triggering the loop counters clock [2.128]. The measured value can be calculated according to equation (2.23): 𝑡𝑖𝑛 =𝑛∗𝑅−𝑡𝑜𝑓𝑓𝑠𝑒𝑡, (2.23) being R the pulse shrinking factor, i.e. the resolution, and toffset the offset size of the pulse circulating on the loop that can no longer trigger the counter’s clock (this value has to be obtained by a time consuming experimental procedure or estimated by exhaustive simulations). The main advantage of this architecture is its non-linearity which is bellow half of the LSB. The shrinking factor of a delay cell can be adjusted by controlling its power supply. Therefore, in FPGA platforms, implementing pulse shrinking architectures is hard since there is no direct mean to control the shrink factor of the ring oscillator. Nevertheless, the work from Chen et al. [2.128] reported a resolution in the range of 110-115 picoseconds with ±1 LSB INL, using a Xilinx XC3S200An FPGA. The authors also proposed a schema to address the offset issue, eliminating the time-consuming process of determining its value through experimental measurements. The proposals to the pulse shrinking architecture changes are depicted in Figure 2.17. Figure 2.18 depicts the waveform diagram of the architecture with the offset canceler mechanism. Although offering good linearity and relatively high resolutions, the extra complexity of implementation makes this architecture less popular when developing for FPGA platforms since TDLs and Phased clocks can offer the same or even better performances with lower design complexity. In fact, FPGA TDLs have
Readout Circuit for Time-Based Automotive Sensors 49 proven to be capable of reaching performance levels similar to the ones achieved by some ASIC pulse shrinking solutions [2.129]–[2.131]. Figure 2.17Offset canceller pulse-shrinking architecture Figure 2.18Offset canceller pulse-shrinking waveforms Though pulse shrinking architectures are not suitable for FPGA implementation, these are good solutions when developing for ASIC due to the finer grain control that a custom design cell can has over the shrinking factor, precisely controlling the TDC’s resolution. Moreover, the linearity of the TDC is just related with the shrinking factor and not to the individual cells’ propagation delays, enabling high performances to be attained. 2.2.8. Summary Table 2.1 presents a summary on the recent FPGA-based and ASIC-based TDCs, grouped according to the Taxonomy proposed on the journal article J1 [2.122]. Note that a direct comparison between ASIC Cyclic Delay Line Time Subtractor Time Adder Pulse-Shrinking Unit ... reset Delay Line t1 t2 PWD Counter EOC tout tout’n tp t1 t2 tin t in tp-nR tp-R . . . tp-(n-1)R td tout tp tp-2R tcycle toffset . . . EOC tout’. . . . . . . . . . . . . . . td
2.Time-based Readout Circuits 50 and FPGA based TDCs cannot be made. Although they may share some applications, usually the goal of these two implementations are not the same. Nevertheless, by the analysis of Table 2.1 it is possible to verify that FPGA-based TDCs’ performances are closing the gap to the ASIC-based ones. 2.3. Commercial Devices Apart from the research that has been done in TDC, it is also important to analyze the available commercial devices. The performance values reported help on understanding the current state of the industry. Table 2.2 presents the most relevant TDC devices available. 2.4. Conclusion By analyzing the state-of-the-art and the TDC’s architectures evolution throughout recent years a set of conclusions can be drawn: 1) First, as technology scales down and FPGA platforms improve, architectures with higher resolutions are expected due to reduction on the cells’ propagation delays. However, this will also contribute to enhance the negative effect of PVT variations on the linearity of the TDC. Therefore, calibration mechanisms and methods to reduce PVT variations will assume higher relevance. 2) Second, nowadays ToF applications are requiring multiple TDC channels. For example, in LiDAR applications, capturing larger parts of the scene in a single shot process by having multiple receivers, each with a TDC channel associated, has the advantage of saving time that can be used by the scanning and image processing algorithms. The alternative solution is to sweep a scene point-by-point, which is a much slower process and puts hard time constraints on the image processing algorithms, if a rate of 10-20 frames per second is required. Thus, lower hardware resources TDC architectures are desirable in order to keep both costs and area utilization low, in order to increase systems integration. In FPGA platforms, some research works have already reported 128-channels [2.35] and 256-channels [2.37], [2.38] implementations. However, these works used large FPGA platforms and the resource utilization was on its limits. Recently the research work by Wu and Xu [2.23] proposed an interesting low resource architecture with good linearity and stability. This architecture seems promising for high channel count using low area, low resources platforms and states the need for more architectures focusing on other aspects of a TDC rather than its resolution. Therefore, architectures capable of reducing the length of the
Readout Circuit for Time-Based Automotive Sensors 51 delay chains without compromising system resolution will become more popular and will be the focus of future research works [2.61], [2.97]. 3) Third, nowadays market is rapidly changing, and time-to-market constraints are tighter than ever. As stated in [2.4], porting an architecture from one application to another is not an easy task, and requires lengthy manual customizations. Therefore, research regarding reconfigurable and customizable TDCs architectures, is required. Also, tools that can help the designer generate the base core of the TDC would accelerate the development process and contribute for system portability and reutilization. A recent research work [2.55] has addressed this issue with promising results. Due to its characteristics, phased clocks architectures are promising candidates for exploring automated generation of TDCs. Automatically generated TDLs offer an extra challenge due to the non-linearity of the delay-chain. Nevertheless, some research works [2.14], [2.61] have explored a multichain architecture which improves the overall TDL’s linearity before calibration. These works are good use cases to test automatic TDC generation tools. With nowadays integration of microprocessors on most FPGA, systems that allow for programmable logic reconfiguration could also be used to achieve automatically generated TDCs. An algorithm to analyze the automatically generated TDL’s non-linearity could be implemented on the microprocessor. Depending on the results from the histograms, the chain’s hardware could be rearranged automatically in order to improve its linearity and reduce missing codes. In [2.162], a framework is proposed to dynamically control the FPGAs routing with precision. The use of such framework in TDC designs could contribute to achieve high linearity architectures in a simplified way. 4) Fourth, in order to comply with modern application requirements, higher sampling rates architectures are required. Usually TDC architectures based on TDLs or DLL report sampling rates equal to the frequency of the reference clock being used. An alternative to reduce other architectures dead time, such as pulse shrinking and ring oscillators, is to have multiple TDC channels operating in an interleaved schema. However, this solution has a negative impact on resources utilization. Therefore, for applications with high input event rates, these architectures may not present the best solution, being phased clocks and delay lines more appropriated. 5) Fifth, building on the literature analysis, it is expected that applications’ requirements will continue to drive TDCs evolution. FPGA-based TDCs have recently reported performance values that can compete with the ones achieved by ASIC TDCs. Therefore, it is expected that FPGA-based TDC would grow in popularity and start to be included on commercial products, rather than just being used in research field experiments or as prototype platforms. A full automated implementation and migration process for FPGA-
2.Time-based Readout Circuits 52 based and ASIC-based TDC will certainly increase its popularity, reduce the production cost and ease the use of these systems in a broader range of applications. The state-of-the-art clearly shows that TDL are the most explored architecture when FPGA platforms are used. On the other hand, in ASIC-based TDCs there is no dominant architecture. Nevertheless, research works reporting the used of DLL, with or without a second stage interpolator to achieve resolutions under the values of the cells’ propagation delay, are the most popular architectures. In FPGA-based TDCs, linearity issues are often addressed by introducing a calibration stage to the TDC architecture, either by doing decimation or bin-by-bin calibration. In ASIC-based TDCs, the shielding against PVT variations is achieved through carefully design the delay element and by a DLL schema, locked to a fixed frequency that dynamically adjusts the power supply of the cells. With the evolution on automotive technology and the appearance of LiDAR sensor for autonomous driving scenarios, it is predictable that TDCs will become more popular, due to their performance on ToF measurements. Although requiring high performance systems, automotive applications have other tighter requirements such as area and power consumption. Therefore, architectures that can combine all these requirements will have the upper hand. In FPGA implementations, this trend can already be identified. The research work by Dinh et al. [2.41] proposes a mixture between a ring oscillator and a second stage interpolator based on a TDL architecture, that enables to reduce hardware resources utilization while maintaining high resolutions. A phased clock architecture with a TDL to cover the time interval in-between phases, proposed in [2.9], is also an indicative to this trend in FPGA platforms. Although FPGA constraint the TDC design by having the available resources pre-defined, the portability between different FPGAs is greatly enhanced when compare to the ASIC scenario where, if a TDC architecture needs to be ported to another technology, the delay cells must be completely redesigned to agree with the new technology rules. In FPGA-based implementation, it is only necessary to change the name of the cell used to create the delay chain in the HDL (Hardware Description Language) file. Furthermore, being capable of testing an architecture in FPGA and directly migrate it to ASIC improves system testability and reduce the risks associated to the porting. Therefore, research on TDC architectures and technology migration processes would definitely prove advantageous on the design of TDCs for new ToF sensors. The research on migration of digital FPGA-based TDC architectures to ASIC platforms is scarce, being the research work by Wang et al. [2.34] one of the few architecture proposals that could be directly migrated
Readout Circuit for Time-Based Automotive Sensors 53 from FPGA, since the TDC channel was developed using HDL. Thus, digital TDC architectures were studied during this Thesis’ research to understand which architecture could be directly migrated to ASIC. Maximum resolution, high sampling rate, power consumption and size were also considered when selecting the TDC architectures, since these requirements are extremely important for LiDAR sensors, which require multiple TDC channels. The selected architectures and design decisions are introduced and discussed in Chapter 3.
2.Time-based Readout Circuits 60 Pulse Shrinking TDCs [2.10]-10 Spartan-3 42 56 [-0.98:0.5] [-4.17:3.5] 11.5 710 - X - [2.160]-14 Virtex-5 4.56 - - - - - - X - [2.128]-16 Spartan-3AN - 115 - ±1 15 - - X [2.161]-18 Actel SmartFusion 8.5 42.4 0.36 0.91 10 1042 - √ - *-number of slices (for a TDL each slice usually has 8 logic units used in Xilinx FPGAs); **-units in mm2 Table 2.2TDCs Commercial Devices Summary Parameter Product Name Resolution (ps) Precision (ps) Range #channels System Clock (MHz) Readout rate Interface Reference vendor Texas Instruments TDC7201 55 35 Mode1:12 ns - 2 us Mode2:250 ns - 8 ms 2 16 - SPI http://www.ti.com/product/TDC7 201#features AMS AS6500 - 20 0 s - 16 s 4 2-12.5 1.5 MSamples/s SPI https://ams.com/as6500 AS6501 - 10 0 s - 16 s 2 2-12.5 70 MSamples/s LVDS and SPI https://ams.com/as6501 TDC-GPX - 10 9.8 us 8 40 40 MSamples/s (200 M peak) 28-bit parallel https://ams.com/tdcgpx#tab/features
Readout Circuit for Time-Based Automotive Sensors 61 TDC-GPX2 - 10 0 s - 16 s 4 2-12.5 35 MSamples/s (70 M peak) Serial LVDS and SPI https://ams.com/tdc-gpx2 Maxim integrated* MAX35101/ MAX35102/ MAX35103 - 20 8 ms 2 - - SPI https://www.maximintegrated.com /en/products/industries/meteringenergymeasurement/MAX35101.html/ Cronologic TimeTagger 4 500 - 8 ms 4 - 48 MHits/s PCIe 1.1 https://www.cronologic.de/time_ measurement/timetag/timetagger 42g/ HPTDC 25 to 12,800 419 us 8 78.125 4 MHits/s PCIe 2.2 https://www.cronologic.de/time_ measurement/tdc/hptdc/ xTDC4 13 - 218 us (14 ms extended) 4 - 48 MHits/s PCIe 1.1 https://www.cronologic.de/time_ measurement/tdc/xtdc4/ *for ultrasonic heat meters and flow meters markets
2.Time-based Readout Circuits 62 References [2.1] J. Mauricio, D. Gascón, D. Ciaglia, S. Gómez, G. Fernández, and A. Sanuy, “MATRIX: a 15 ps resistive interpolation TDC ASIC based on a novel regular structure,” J. Instrum., vol. 11, no. 12, pp. C12047–C12047, Dec. 2016. [2.2] R. A. Dias et al., “Real-Time Operation and Characterization of a High-Performance Time-Based Accelerometer,” J. Microelectromechanical Syst., vol. 24, no. 6, pp. 1703–1711, Dec. 2015. [2.3] J.-P. Jansson, “A stabilized multi-channel CMOS time-to-digital converter based on a low frequency reference,” University of Oulu, 2012. [2.4] Z. Cheng, X. Zheng, M. J. Deen, and H. Peng, “Recent Developments and Design Challenges of High-Performance Ring Oscillator CMOS Time-to-Digital Converters,” IEEE Trans. Electron Devices, vol. 63, no. 1, pp. 235–251, Jan. 2016. [2.5] R. Szplet, R. Szymanowski, and D. Sondej, “Measurement Uncertainty of Precise Interpolating Time Counters,” IEEE Trans. Instrum. Meas., pp. 1–9, 2019. [2.6] J.-P. Jansson, A. Mantyniemi, and J. Kostamovaara, “Multiplying delay locked loop (MDLL) in time-to-digital conversion,” in 2009 IEEE Intrumentation and Measurement Technology Conference, 2009, pp. 1226–1231. [2.7] Z. Jachna, R. Szplet, P. Kwiatkowski, and K. Różyc, “Permanently calibrated interpolating time counter,” Meas. Sci. Technol., vol. 26, no. 1, p. 015006, Jan. 2015. [2.8] R. Szplet, P. Kwiatkowski, Z. Jachna, and K. Rozyc, “Precise three-channel integrated time counter,” in 2015 Joint Conference of the IEEE International Frequency Control Symposium & the European Frequency and Time Forum, 2015, pp. 575–578. [2.9] R. Szplet, P. Kwiatkowski, Z. Jachna, and K. Rozyc, “An eight-channel 4.5-ps precision timestamps-based time interval counter in FPGA chip,” IEEE Trans. Instrum. Meas., vol. 65, no. 9, pp. 2088–2100, 2016. [2.10] R. Szplet and K. Klepacki, “An FPGA-Integrated Time-to-Digital Converter Based on Two-Stage Pulse Shrinking,” IEEE Trans. Instrum. Meas., vol. 59, no. 6, pp. 1663–1670, Jun. 2010. [2.11] P. Kwiatkowski and R. Szplet, “Time-to-Digital Converter with Pseudo-Segmented Delay Line,” in 2019 IEEE International Instrumentation and Measurement Technology Conference (I2MTC), 2019, pp. 1–6. [2.12] R. Plassche, “Specifications of converters,” in CMOS Integrated Analog-to-Digital and Digital-toAnalog Converters, Boston, MA: Springer US, 2003, pp. 57–64. [2.13] S. Cova and M. Bertolaccini, “Differential linearity testing and precision calibration of multichannel time sorters,” Nucl. Instruments Methods, vol. 77, no. 2, pp. 269–276, Jan. 1970. [2.14] J. Y. Won and J. S. Lee, “Time-to-Digital Converter Using a Tuned-Delay Line Evaluated in 28-, 40-, and 45-nm FPGAs,” IEEE Trans. Instrum. Meas., vol. 65, no. 7, pp. 1678–1689, Jul. 2016. [2.15] J. Zhang and D. Zhou, “A new delay line loops shrinking time-to-digital converter in low-cost FPGA,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 771, pp. 10–16, 2015. [2.16] K. Cui, Z. Ren, X. Li, Z. Liu, and R. Zhu, “A High-Linearity, Ring-Oscillator-Based, Vernier Time-toDigital Converter Utilizing Carry Chains in FPGAs,” IEEE Trans. Nucl. Sci., vol. 64, no. 1, pp. 697–704, 2017. [2.17] M. Arkani, “A High Performance Digital Time Interval Spectrometer: An Embedded, FPGA-Based System With Reduced Dead Time Behaviour,” Metrol. Meas. Syst., vol. 22, no. 4, pp. 601–619, Dec. 2015.
Readout Circuit for Time-Based Automotive Sensors 63 [2.18] Q. Guo, R. Feng, Y. Wu, and N. Yu, “Measurement of the AFDX switch latency based on FPGA,” in 2016 IEEE International Conference on Aircraft Utility Systems (AUS), 2016, pp. 45–49. [2.19] D. N. Grigoriev, P. V. Kasyanenko, E. A. Kravchenko, A. G. Shamov, and A. A. Talyshev, “A 32channel 840Msps TDC based on Altera Cyclone III FPGA,” J. Instrum., vol. 12, no. 8, 2017. [2.20] C. Liu, Y. Wang, P. Kuang, D. Li, and X. Cheng, “A 3.9 ps RMS resolution time-To-digital converter using dual-sampling method on Kintex UltraScale FPGA,” 2016 IEEE-NPSS Real Time Conf. RT 2016, pp. 1–3, 2016. [2.21] Y. Wang and C. Liu, “A 4.2 ps Time-Interval RMS Resolution Time-to-Digital Converter Using a Bin Decimation Method in an UltraScale FPGA,” IEEE Trans. Nucl. Sci., vol. 63, no. 5, pp. 2632– 2638, 2016. [2.22] R. Szplet, “Time-to-Digital Converters,” in Design, Modeling and Testing of Data Converters, P. Carbone, S. Kiaei, and F. Xu, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2014, pp. 211– 246. [2.23] J. Wu and J. Xu, “A Novel TDC Scheme: Combinatorial Gray Code Oscillator Based TDC for Low Power and Low Resource Usage Applications,” in 2019 5th International Conference on EventBased Control, Communication, and Signal Processing (EBCCSP), 2019, pp. 1–7. [2.24] J. Kalisz, “Review of methods for time interval measurements with picosecond resolution,” Metrologia, vol. 41, no. 1, pp. 17–32, Feb. 2004. [2.25] F. Cossio, “A mixed-signal ASIC for the readout of Gas Electron Multiplier detectors design review and characterization results,” in 2017 13th Conference on Ph.D. Research in Microelectronics and Electronics (PRIME), 2017, pp. 33–36. [2.26] D. Calvo, “1 ns time to digital converters for the KM3NeT data readout system,” in AIP Conference Proceedings, 2014, vol. 1630, no. 2014, pp. 98–101. [2.27] Z. Li et al., “Development of an integrated four-channel fast avalanche-photodiode detector system with nanosecond time resolution,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 870, no. November 2016, pp. 43–49, 2017. [2.28] A. Balla et al., “The characterization and application of a low resource FPGA-based time to digital converter,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 739, pp. 75–82, 2014. [2.29] Y. Sano, Y. Horii, M. Ikeno, O. Sasaki, M. Tomoto, and T. Uchida, “Subnanosecond time-to-digital converter implemented in a Kintex-7 FPGA,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 874, no. February, pp. 50–56, 2017. [2.30] H. Huang and W. Chou, “Hysteresis Switch Adaptive Velocity Evaluation and High-Resolution Position Subdivision Detection Based on FPGA,” IEEE Trans. Instrum. Meas., vol. 64, no. 12, pp. 3387–3395, Dec. 2015. [2.31] H. H. Fan, P. Cao, S. Bin Liu, and Q. An, “TOT measurement implemented in FPGA TDC,” Chinese Phys. C, vol. 39, no. 11, 2015. [2.32] W. Yonggang, C. Xinyi, L. Deng, Z. Wensong, and L. Chong, “A linear time-over-threshold digitizing scheme and its 64-channel DAQ prototype design on FPGA for a continuous crystal PET detector,” IEEE Trans. Nucl. Sci., vol. 61, no. 1, pp. 99–106, 2014. [2.33] T. Xiang et al., “A 56-ps multi-phase clock time-to-digital convertor based on Artix-7 FPGA,” in 2014 19th IEEE-NPSS Real Time Conference, 2014, pp. 1–4. [2.34] J. Wang et al., “Development of a time-to-digital converter ASIC for the upgrade of the ATLAS Monitored Drift Tube detector,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 880, pp. 174–180, Feb. 2018.
2.Time-based Readout Circuits 64 [2.35] C. Liu and Y. Wang, “A 128-Channel, 710 M Samples/Second, and Less Than 10 ps RMS Resolution Time-to-Digital Converter Implemented in a Kintex-7 FPGA,” IEEE Trans. Nucl. Sci., vol. 62, no. 3, pp. 773–783, 2015. [2.36] W. Pan, G. Gong, and J. Li, “A 20-ps time-to-digital converter (TDC) implemented in fieldprogrammable gate array (FPGA) with automatic temperature correction,” IEEE Trans. Nucl. Sci., vol. 61, no. 3, pp. 1468–1473, 2014. [2.37] Z. Song, Y. Wang, and J. Kuang, “A 256-channel, high throughput and precision time-to-digital converter with a decomposition encoding scheme in a Kintex-7 FPGA,” J. Instrum., vol. 13, no. 05, pp. P05012–P05012, May 2018. [2.38] Y. Wang, P. Kuang, and C. Liu, “A 256-channel multi-phase clock sampling-based time-to-digital converter implemented in a Kintex-7 FPGA,” in Conference Record - IEEE Instrumentation and Measurement Technology Conference, 2016, vol. 2016-July. [2.39] B. Qi et al., “A compact readout electronics for the ground station of a quantum communication satellite,” IEEE Trans. Nucl. Sci., vol. 62, no. 3, pp. 883–888, 2015. [2.40] W.-S. Choong, F. Abu-Nimeh, W. W. Moses, Q. Peng, C. Q. Vu, and J.-Y. Wu, “A front-end readout Detector Board for the OpenPET electronics system,” J. Instrum., vol. 10, no. 08, pp. T08002– T08002, Aug. 2015. [2.41] V. L. Dinh, X. T. Nguyen, and H.-J. Lee, “A New FPGA Implementation of a Time-to-Digital Converter Supporting Run-Time Estimation of Operating Condition Variation,” in 2018 IEEE International Symposium on Circuits and Systems (ISCAS), 2018, pp. 1–4. [2.42] N. Franch, O. Alonso, J. Canals, A. Vila, A. Herms, and A. Dieguez, “A low cost fluorescence lifetime measurement system based on SPAD detectors and FPGA processing,” in 2016 Conference on Design of Circuits and Integrated Systems (DCIS), 2016, pp. 1–6. [2.43] L.-Y. Hsu and J.-L. Huang, “A multi-channel FPGA-based time-to-digital converter,” in 2016 IEEE 21st International Mixed-Signal Testing Workshop (IMSTW), 2016, pp. 1–4. [2.44] H. Y. T. To et al., “A Novel Programmable On-chip Voltage Droop Detector for FPGA Applications,” in 2016 IEEE 66th Electronic Components and Technology Conference (ECTC), 2016, pp. 2009– 2015. [2.45] H. Chen, Y. Zhang, and D. D.-U. Li, “A Low Nonlinearity, Missing-Code Free Time-to-Digital Converter Based on 28-nm FPGAs With Embedded Bin-Width Calibrations,” IEEE Trans. Instrum. Meas., vol. 66, no. 7, pp. 1912–1921, Jul. 2017. [2.46] K. Katoh et al., “A Small Chip Area Stochastic Calibration for TDC Using Ring Oscillator,” J. Electron. Test., vol. 30, no. 6, pp. 653–663, Dec. 2014. [2.47] D. R. E. Gnad, F. Oboril, S. Kiamehr, and M. B. Tahoori, “An Experimental Evaluation and Analysis of Transient Voltage Fluctuations in FPGAs,” IEEE Trans. Very Large Scale Integr. Syst., pp. 1– 14, 2018. [2.48] G. Cao, H. Xia, and N. Dong, “An 18-ps TDC using timing adjustment and bin realignment methods in a Cyclone-IV FPGA,” Rev. Sci. Instrum., vol. 89, no. 5, p. 054707, 2018. [2.49] F. Nogrette et al., “Characterization of a detector chain using a FPGA-based time-to-digital converter to reconstruct the three-dimensional coordinates of single particles at high flux,” Rev. Sci. Instrum., vol. 86, no. 11, p. 113105, Nov. 2015. [2.50] J. Jung, Y. Choi, K. bom Kim, S. Lee, and H. Choe, “An improved time over threshold method using bipolar signals,” Phys. Med. Biol., vol. 63, no. 13, p. 135002, Jun. 2018.
Readout Circuit for Time-Based Automotive Sensors 65 [2.51] E. Venialgo et al., “An order-statistics-inspired, fully-digital readout approach for analog SiPM arrays,” in 2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and RoomTemperature Semiconductor Detector Workshop (NSS/MIC/RTSD), 2016, pp. 1–5. [2.52] J. Michel et al., “Electronics for the RICH detectors of the HADES and CBM experiments,” J. Instrum., vol. 12, no. 01, pp. C01072–C01072, Jan. 2017. [2.53] E. Arabul, J. Rarity, and NaimDahnoun, “FPGA based fast integrated real-time multi coincidence counter using a time-to-digital converter,” in 2018 7th Mediterranean Conference on Embedded Computing (MECO), 2018, pp. 1–4. [2.54] A. T. Eshghi, S. Lee, M. K. Sadoughi, C. Hu, Y.-C. Kim, and J.-H. Seo, “Generic high resolution PET detector block using 12×12 SiPM array,” Smart Mater. Struct., vol. 26, no. 10, p. 105037, Oct. 2017. [2.55] N. Lusardi, A. Palmucci, and A. Geraci, “Fully-migratable TDC architecture for FPGA devices,” in 2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and Room-Temperature Semiconductor Detector Workshop (NSS/MIC/RTSD), 2016, pp. 1–3. [2.56] S. Grzelak, L. Wydzgowski, J. Czokow, D. Chaberski, and M. Zielinski, “High precision ΔE effect measurement with the use of ultrasonic-wave-time-of-flight method,” Prz. Elektrotechniczny, vol. 1, no. 11, pp. 85–88, Nov. 2016. [2.57] W. Pan, G. Gong, Q. Du, H. Li, and J. Li, “High resolution distributed time-to-digital converter (TDC) in a White Rabbit network,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 738, pp. 13–19, Feb. 2014. [2.58] N. Lusardi, A. Geraci, J. Marjanovic, and M. Gustin, “High-resolution TDL-TDC system for MTCA.4 standard,” in 2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and RoomTemperature Semiconductor Detector Workshop (NSS/MIC/RTSD), 2016, pp. 1–4. [2.59] S. Grzelak, M. Kowalski, J. Czoków, and M. Zieliński, “High Resolution Time-Interval Measurement Systems Applied To Flow Measurement,” Metrol. Meas. Syst., vol. 21, no. 1, pp. 77–84, Mar. 2014. [2.60] B. Neumeier and D. Schmitt-Landsiedel, “Online Condition Measurement of High Power Solid State Laser Cutting Optics using Ultrasound Signals,” Phys. Procedia, vol. 56, pp. 1252–1260, 2014. [2.61] H. Chen and D. D.-U. Li, “Multichannel, Low Nonlinearity Time-to-Digital Converters Based on 20 and 28 nm FPGAs,” IEEE Trans. Ind. Electron., vol. 66, no. 4, pp. 3265–3274, Apr. 2019. [2.62] T. Polzer, F. Huemer, and A. Steininger, “Measuring metastability using a time-to-digital converter,” in 2017 IEEE 20th International Symposium on Design and Diagnostics of Electronic Circuits & Systems (DDECS), 2017, pp. 116–121. [2.63] R. Szplet, P. Kwiatkowski, K. Rozyc, M. Sawicki, and Z. Jachna, “Modular time interval counter,” in 2014 European Frequency and Time Forum (EFTF), 2014, pp. 494–497. [2.64] M. Pałka et al., “Multichannel FPGA based MVT system for high precision time (20 ps RMS) and charge measurement,” J. Instrum., vol. 12, no. 08, pp. P08001–P08001, Aug. 2017. [2.65] Y. Wang, Q. Cao, and C. Liu, “A Multi-Chain Merged Tapped Delay Line for High Precision Timeto-Digital Converters in FPGAs,” IEEE Trans. Circuits Syst. II Express Briefs, vol. 65, no. 1, pp. 96–100, Jan. 2018. [2.66] P. Deng et al., “Readout Electronics of T0 Detector in the External Target Experiment of CSR in HIRFL,” IEEE Trans. Nucl. Sci., vol. 65, no. 6, pp. 1315–1323, 2018. [2.67] D. Yang et al., “Readout electronics of a prototype time-of-flight ion composition analyzer for space plasma,” Nucl. Sci. Tech., vol. 29, no. 4, p. 60, Apr. 2018.
2.Time-based Readout Circuits 66 [2.68] T. Polzer, F. Huemer, and A. Steininger, “Refined metastability characterization using a time-todigital converter,” Microelectron. Reliab., vol. 80, pp. 91–99, Jan. 2018. [2.69] E. Arabul, A. Girach, J. Rarity, and N. Dahnoun, “Precise multi-channel timing analysis system for multi-stop LIDAR correlation,” in 2017 IEEE International Conference on Imaging Systems and Techniques (IST), 2017, pp. 1–6. [2.70] H. Li, T. Xue, G. Gong, and J. Li, “The integration of FPGA TDC inside White Rabbit node,” J. Instrum., vol. 12, no. 04, pp. P04020–P04020, Apr. 2017. [2.71] S. Grzelak, J. Czoków, M. Kowalski, and M. Zieliński, “Ultrasonic Flow Measurement with High Resolution,” Metrol. Meas. Syst., vol. 21, no. 2, pp. 305–316, Jun. 2014. [2.72] F. Huemer, T. Polzer, and A. Steininger, “Using a Duplex Time-to-Digital Converter for Metastability Characterization of an FPGA,” in 2018 IEEE 21st International Symposium on Design and Diagnostics of Electronic Circuits & Systems (DDECS), 2018, pp. 141–146. [2.73] A. Aguilar et al., “Timing Results Using an FPGA-Based TDC with Large Arrays of 144 SiPMs,” IEEE Trans. Nucl. Sci., vol. 62, no. 1, pp. 12–18, Feb. 2015. [2.74] Y. Wang, J. Kuang, C. Liu, and Q. Cao, “A 3.9-ps RMS Precision Time-to-Digital Converter Using Ones-Counter Encoding Scheme in a Kintex-7 FPGA,” IEEE Trans. Nucl. Sci., vol. 64, no. 10, pp. 2713–2718, Oct. 2017. [2.75] Y. Wang and C. Liu, “A nonlinearity minimization-oriented resource-saving time-to-digital converter implemented in a 28 nm Xilinx FPGA,” IEEE Trans. Nucl. Sci., vol. 62, no. 5, pp. 2003–2009, 2015. [2.76] Q. Shen et al., “A multi-chain measurements averaging TDC implemented in a 40 nm FPGA,” 2014 19th IEEE-NPSS Real Time Conf. RT 2014 - Conf. Rec., pp. 6–8, 2015. [2.77] Y. Wang, J. Kuang, C. Liu, Q. Cao, and D. Li, “A flexible 32-channel time-to-digital converter implemented in a Xilinx Zynq-7000 field programmable gate array,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 847, no. September, pp. 61–66, 2017. [2.78] Q. Cao, Y. Wang, and C. Liu, “A Combination of Multiple Channels of FPGA Based Time-to-Digital Converter for High Time Precision,” in Nuclear Science Symposium, Medical Imaging Conference and Room-Temperature Semiconductor Detector Workshop (NSS/MIC/RTSD) 2016, 2016. [2.79] P. Chen, Y. Y. Hsiao, and Y. S. Chung, “A high resolution FPGA TDC converter with 2.5 ps bin size and -3.79~6.53 LSB integral nonlinearity,” in Proceedings of the 2nd International Conference on Intelligent Green Building and Smart Grid, IGBSG 2016, 2016, pp. 2–6. [2.80] C. Ugur, S. Linev, J. Michel, T. Schweitzer, and M. Traxler, “A novel approach for pulse width measurements with a high precision (8 ps RMS) TDC in an FPGA,” J. Instrum., vol. 11, no. 1, 2016. [2.81] A. Tontini, L. Gasparini, L. Pancheri, and R. Passerone, “Design and Characterization of a LowCost FPGA-Based TDC,” IEEE Trans. Nucl. Sci., vol. 65, no. 2, pp. 680–690, Feb. 2018. [2.82] J. Y. Won, S. Il Kwon, H. S. Yoon, G. B. Ko, J. W. Son, and J. S. Lee, “Dual-Phase Tapped-DelayLine Time-to-Digital Converter with On-the-Fly Calibration Implemented in 40 nm FPGA,” IEEE Trans. Biomed. Circuits Syst., vol. 10, no. 1, pp. 231–242, 2016. [2.83] S. Burri, H. Homulle, C. Bruschini, and E. Charbon, “LinoSPAD: a time-resolved 256×1 CMOS SPAD line sensor system featuring 64 FPGA-based TDC channels running at up to 8.5 giga-events per second,” Opt. Sens. Detect. IV, vol. 9899, p. 98990D, 2016. [2.84] J. Zheng, P. Cao, D. Jiang, and Q. An, “Low-Cost FPGA TDC With High Resolution and Density,” IEEE Trans. Nucl. Sci., vol. 64, no. 6, pp. 1401–1408, 2017.
Readout Circuit for Time-Based Automotive Sensors 67 [2.85] S. Guo, Y. Wang, N. Li, J. Diao, and L. Chen, “Multi-chain time interval measurement method utilizing the dedicated carry chain of FPGA,” in 2017 7th IEEE International Conference on Electronics Information and Emergency Communication (ICEIEC), 2017, no. 1, pp. 489–492. [2.86] D. Chaberski, R. Frankowski, M. Zieliński, and Ł. Zaworski, “Multiple-tapped-delay-line hardwarelinearisation technique based on wire load regulation,” Meas. J. Int. Meas. Confed., vol. 92, pp. 103–113, 2016. [2.87] Q. Shen et al., “A 1.7 ps equivalent bin size and 4.2 ps RMS FPGA TDC based on multichain measurements averaging method,” IEEE Trans. Nucl. Sci., vol. 62, no. 3, pp. 947–954, 2015. [2.88] N. Lusardi, J. W. N. Los, R. B. M. Gourgues, G. Bulgarini, and A. Geraci, “Photon counting with photon number resolution through superconducting nanowires coupled to a multi-channel TDC in FPGA,” Rev. Sci. Instrum., vol. 88, no. 3, 2017. [2.89] N. Lusardi, A. Abba, F. Caponio, and A. Geraci, “Quantization noise in non-homogeneous calibration table of a TCD implemented in FPGA,” in 2014 IEEE Nuclear Science Symposium and Medical Imaging Conference (NSS/MIC), 2014, no. i, pp. 1–5. [2.90] Y. Wang, C. Liu, X. Cheng, and D. Li, “Spartan-6 FPGA based 8-channel time-to-digital converters for TOF-PET systems,” in 2015 IEEE Nuclear Science Symposium and Medical Imaging Conference, NSS/MIC 2015, 2016, pp. 1–3. [2.91] J. Torres et al., “Time-to-Digital Converter Based on FPGA With Multiple Channel Capability,” Nucl. Sci. IEEE Trans., vol. 61, no. 1, pp. 107–114, 2014. [2.92] D. Chaberski, “Time-to-digital-converter based on multiple-tapped-delay-line,” Measurement, vol. 89, pp. 87–96, Jul. 2016. [2.93] H. Homulle, F. Regazzoni, and E. Charbon, “200 MS/s ADC implemented in a FPGA employing TDCs,” Proc. 2015 ACM/SIGDA Int. Symp. Field-Programmable Gate Arrays, pp. 228–235, 2015. [2.94] S. Y. Wang, J. Wu, S. H. Yao, and W. C. Chang, “A Field-Programmable Gate Array (FPGA) TDC for the Fermilab SeaQuest (E906) experiment and its test with a novel external wave union launcher,” IEEE Trans. Nucl. Sci., vol. 61, no. 6, pp. 3592–3598, 2014. [2.95] M. Pałka et al., “A novel method based solely on field programmable gate array (FPGA) units enabling measurement of time and charge of analog signals in positron emission tomography (PET),” Bio-Algorithms and Med-Systems, vol. 10, no. 1, pp. 41–45, 2014. [2.96] T. Chujo et al., “Experimental verification of timing measurement circuit with self-calibration,” in 19th Annual International Mixed-Signals, Sensors, and Systems Test Workshop Proceedings, 2014, vol. 1, pp. 1–6. [2.97] R. Narasimman, A. Prabhakar, and N. Chandrachoodan, “Implementation of a 30 ps resolution time to digital converter in FPGA,” in 2015 International Conference on Electronic Design, Computer Networks & Automated Verification (EDCAV), 2015, pp. 12–17. [2.98] X. Qin, L. Wang, D. Liu, Y. Zhao, X. Rong, and J. Du, “A 1.15-ps Bin Size and 3.5-ps Single-Shot Precision Time-to-Digital Converter With On-Board Offset Correction in an FPGA,” IEEE Trans. Nucl. Sci., vol. 64, no. 12, pp. 2951–2957, Dec. 2017. [2.99] A. Aguilar et al., “Optimization of a Time-to-Digital Converter and a coincidence map algorithm for TOF-PET applications,” J. Syst. Archit., vol. 61, no. 1, pp. 40–48, 2015. [2.100] A. Aguilar et al., “Time of flight measurements based on FPGA using a breast dedicated PET,” J. Instrum., vol. 9, no. 5, 2014.
2.Time-based Readout Circuits 68 [2.101] A. Aguilar et al., “Time of flight measurements based on FPGA and SiPMs for PET-MR,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 734, no. PART B, pp. 127–131, 2014. [2.102] N. Lusardi, F. Garzetti, G. Bulgarini, R. B. M. Gourgues, J. W. N. Los, and A. Geraci, “Single photon counting through multi-channel TDC in programmable logic,” in 2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and Room-Temperature Semiconductor Detector Workshop (NSS/MIC/RTSD), 2016, pp. 1–4. [2.103] Y. Wang and C. Liu, “A 3.9 ps Time-Interval RMS Precision Time-to-Digital Converter Using a DualSampling Method in an UltraScale FPGA,” IEEE Trans. Nucl. Sci., vol. 63, no. 5, pp. 2617–2621, 2016. [2.104] N. Lusardi and A. Geraci, “8-Channels high-resolution TDC in FPGA,” in 2015 IEEE Nuclear Science Symposium and Medical Imaging Conference, NSS/MIC 2015, 2016, pp. 1–2. [2.105] Y.-C. Chen, H.-C. Chang, and H. Chen, “Two-Dimensional Multiply-Accumulator for Classification of Neural Signals,” IEEE Access, vol. 6, pp. 19714–19725, 2018. [2.106] R. Machado, L. A. Rocha, and J. Cabral, “A novel synchronizer for a 17.9ps Nutt Time-to-Digital Converter implemented on FPGA,” in 2018 AEIT International Annual Conference, 2018, pp. 1– 6. [2.107] R. Machado, J. Cabral, and F. Alves, “Designing Synchronizers for Nutt-TDCs,” in 2019 5th International Conference on Event-Based Control, Communication, and Signal Processing (EBCCSP), 2019, pp. 1–6. [2.108] J. Wu and Z. Shi, “The 10-ps wave union TDC: Improving FPGA TDC resolution beyond its cell delay,” in 2008 IEEE Nuclear Science Symposium Conference Record, 2008, pp. 3440–3446. [2.109] C. Ugur, G. Korcyl, J. Michel, M. Penschuk, and M. Traxler, “264 Channel TDC Platform applying 65 channel high precision (7.2 psRMS) FPGA based TDCs,” in 2013 IEEE Nordic-Mediterranean Workshop on Time-to-Digital Converters (NoMe TDC), 2013, pp. 1–5. [2.110] N. Lusardi, M. Luciani, and A. Geraci, “Single-chain 4-channels high-resolution multi-hit TDC in FPGA,” in 2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and RoomTemperature Semiconductor Detector Workshop (NSS/MIC/RTSD), 2016, pp. 1–4. [2.111] J. Kuang, Y. Wang, Q. Cao, and C. Liu, “Implementation of a high precision multi-measurement time-to-digital convertor on a Kintex-7 FPGA,” Nucl. Instruments Methods Phys. Res. Sect. A Accel. Spectrometers, Detect. Assoc. Equip., vol. 891, no. February, pp. 37–41, May 2018. [2.112] K. Cui, X. Li, Z. Liu, and R. Zhu, “Toward Implementing Multichannels, Ring-Oscillator-Based, Vernier Time-to-Digital Converter in FPGAs: Key Design Points and Construction Method,” IEEE Trans. Radiat. Plasma Med. Sci., vol. 1, no. 5, pp. 391–399, Sep. 2017. [2.113] P. Chen, Y. Hsiao, Y. Chung, W. X. Tsai, and J. Lin, “A 2.5-ps Bin Size and 6.7-ps Resolution FPGA Time-to-Digital Converter Based on Delay Wrapping and Averaging,” IEEE Trans. Very Large Scale Integr. Syst., vol. 25, no. 1, pp. 114–124, Jan. 2017. [2.114] R. Szplet, D. Sondej, and G. Grzeda, “Subpicosecond-resolution time-to-digital converter with multi-edge coding in independent coding lines,” in 2014 IEEE International Instrumentation and Measurement Technology Conference (I2MTC) Proceedings, 2014, pp. 747–751. [2.115] R. Szplet, Z. Jachna, P. Kwiatkowski, and K. Rozyc, “A 2.9 ps equivalent resolution interpolating time counter based on multiple independent coding lines,” Meas. Sci. Technol., vol. 24, no. 3, p. 035904, Mar. 2013. [2.116] G. Grzęda and R. Szplet, “Time interval measurement module implemented in SoC FPGA device,” Int. J. Electron. Telecommun., vol. 62, no. 3, pp. 237–246, Sep. 2016.
Readout Circuit for Time-Based Automotive Sensors 69 [2.117] I. Diehl et al., “Readout ASIC for fast digital imaging using SiPM sensors: Concept study,” in 2015 IEEE Nuclear Science Symposium and Medical Imaging Conference (NSS/MIC), 2015, pp. 1–3. [2.118] L. Perktold and J. Christiansen, “A multichannel time-to-digital converter ASIC with better than 3 ps RMS time resolution,” J. Instrum., vol. 9, no. 01, pp. C01060–C01060, Jan. 2014. [2.119] T. Watanabe and H. Isomura, “All-digital ADC/TDC using TAD architecture for highly-durable timemeasurement ASIC,” in 2014 IEEE International Symposium on Circuits and Systems (ISCAS), 2014, pp. 674–677. [2.120] J.-C. Lai and T.-Y. Hsu, “Cost-Effective Time-to-Digital Converter Using Time-Residue Feedback,” IEEE Trans. Ind. Electron., vol. 64, no. 6, pp. 4690–4700, Jun. 2017. [2.121] H. Molaei and K. Hajsadeghi, “A 5.3-ps, 8-b Time to Digital Converter Using a New GainReconfigurable Time Amplifier,” IEEE Trans. Circuits Syst. II Express Briefs, vol. 66, no. 3, pp. 352–356, Mar. 2019. [2.122] R. Machado, J. Cabral, and F. S. Alves, “Recent Developments and Challenges in FPGA-Based Time-to-Digital Converters,” IEEE Trans. Instrum. Meas., vol. 68, no. 11, pp. 4205–4221, Nov. 2019. [2.123] M. Maamoun, I. S. Arami, R. Beguenane, A. Benbelkacem, and A. Meraghni, “A 3ps Resolution Time-to-digital Converter in Low-cost FPGA for Laser Rangefinder,” in Proceedings of the World Congress on Engineering, 2017, vol. I, no. figure 2, pp. 7–11. [2.124] P. Dudek, S. Szczepanski, and J. V. Hatfield, “A high-resolution CMOS time-to-digital converter utilizing a Vernier delay line,” IEEE J. Solid-State Circuits, vol. 35, no. 2, pp. 240–247, Feb. 2000. [2.125] C. T. Ko, K. P. Pun, and A. Gothenberg, “A 5-ps Vernier sub-ranging time-to-digital converter with DNL calibration,” Microelectronics J., vol. 46, no. 12, pp. 1469–1480, 2015. [2.126] K. Cui, Z. Liu, R. Zhu, and X. Li, “FPGA-based high-performance time-to-digital converters by utilizing multi-channels looped carry chains,” in 2017 International Conference on Field Programmable Technology (ICFPT), 2017, pp. 223–226. [2.127] V. Nguyen, D. Duong, Y. Chung, and J.-W. Lee, “A Cyclic Vernier Two-Step TDC for High Input Range Time-of-Flight Sensor Using Startup Time Correction Technique,” Sensors, vol. 18, no. 11, p. 3948, Nov. 2018. [2.128] C.-C. Chen, C. Hwang, Y. Lin, and G. Chen, “Note: All-digital pulse-shrinking time-to-digital converter with improved dynamic range,” Rev. Sci. Instrum., vol. 87, no. 4, p. 046104, Apr. 2016. [2.129] C.-C. Chen, S.-H. Lin, and C.-S. Hwang, “An Area-Efficient CMOS Time-to-Digital Converter Based on a Pulse-Shrinking Scheme,” IEEE Trans. Circuits Syst. II Express Briefs, vol. 61, no. 3, pp. 163–167, Mar. 2014. [2.130] Yue Liu et al., “A 6ps resolution pulse shrinking Time-to-Digital Converter as phase detector in multi-mode transceiver,” in 2008 IEEE Radio and Wireless Symposium, 2008, pp. 163–166. [2.131] R. Enomoto, T. Iizuka, T. Koga, T. Nakura, and K. Asada, “A 16-bit 2.0-ps Resolution Two-Step TDC in 0.18-um CMOS Utilizing Pulse-Shrinking Fine Stage With Built-In Coarse Gain Calibration,” IEEE Trans. Very Large Scale Integr. Syst., vol. 27, no. 1, pp. 11–19, Jan. 2019. [2.132] P. Fischer, I. Peric, M. Ritzert, and T. Solf, “Multi-Channel Readout ASIC for ToF-PET,” in 2006 IEEE Nuclear Science Symposium Conference Record, 2006, pp. 2523–2527. [2.133] P. Fischer, I. Peric, M. Ritzert, and M. Koniczek, “Fast Self Triggered Multi Channel Readout ASIC for Timeand Energy Measurement,” IEEE Trans. Nucl. Sci., vol. 56, no. 3, pp. 1153–1158, Jun. 2009.
3.FPGA-based TDC Development 76 one output, or as two 5-input LUTs with one output each. Furthermore, in the case of SLICEM, the LUTs can also be used to implement 32-bit distributed RAMs and shift registers. Four of the Slice’s storing elements must be edge -triggered D-type flip-flops. The remaining four Slice storing elements can be configured as edge-triggered D-type flip-flops or latches, with a set or reset signal. However, if these storing elements are configured as latches, the first four flip-flops of the Slice cannot be used. Moreover, the storing elements must keep the same configuration inside the same Slice, i.e. if one of the storing elements is configured as a flip-flop with asynchronous reset, then the other storing elements can only implement the same type of flip-flop, since the control signals (clock, enable, set/reset) are shared inside a Slice. The Carry4 cell is provided to enable fast arithmetic operations (addition and subtraction). Each CLB as two identical 4-bit carry chains, one per slice. In order to increase the number of inputs supported by the carry element, multiple carries from different slices can be cascaded using the COUT and CIN ports (See Figure 3.3). The CYINIT input is used as the CIN bit in the first carry of a carry chain or to select between the add operation (0) and the subtract operation (1). The carry element outputs the result of the addition/subtraction on O0 to O3 while the carry out of each bit can be accessed through CO0 to CO3 outputs, the last one, the most significant bit, is also connected to COUT to be used to cascade the carry chain. 3.2. TDC Design Flow and General Architecture The first step when designing any digital system is to analyze the application requirements to understand which constraints and limitations must be addressed. According to the problem description presented in Chapter 1, a typical LiDAR sensor application requires resolutions below 7 cm and a range near 180 m. These requirements represent a time resolution for the ToF measurement better than 467 ps and a measurement range of approximately 1.34 µs. In order to address all the aforementioned constraints, the block diagram presented in Figure 3.4 was developed to guide the implementation of the TDC. The system’s architecture was designed to guarantee a modular and flexible implementation, thus, all the blocks represented in Figure 3.4 can operate in standalone and be reused in other digital designs. This enabled the use of the same design structure to implement both FPGA-based TDC architectures, by only changing the fine measurement module.
Readout Circuit for Time-Based Automotive Sensors 77 Figure 3.3CLB disposition overview (top) and Slice detailed view (bottom) SliceX CLB SliceL0 (XnYm+1) SliceL1 (Xn+1Ym+1) CLB SliceL0 (XnYm) SliceL1 (Xn+1Ym) Switch BOXSwitch Box CLB SliceM0 (XnYm+1) SliceL1 (Xn+1Ym+1) CLB SliceM0 (XnYm) SliceL1 (Xn+1Ym) CLB Switch Box CLB Switch Box CLB Switch Box CLB Switch Box Switch BOXSwitch Box CIN COUTCOUT COUTCOUT CIN CIN CIN CINCIN CINCIN COUTCOUT COUTCOUT Dx D[6:1] Q Q SET CLR D E LUT6 O6 O5 MUXCY COUT DMUX DQ DMUX/DQ Carry Chain Block (CARRY4) Cx C[6:1] Q Q SET CLR D E LUT6 O6 O5 MUXCY CMUX CQ CMUX/CQ Bx B[6:1] Q Q SET CLR D E LUT6 O6 O5 MUXCY BMUX BQ BMUX/BQ Ax A[6:1] Q Q SET CLR D E LUT6 O6 O5 MUXCY AMUX AQ AMUX/AQ CIN 01 O0 O1 CO0 CO1 O2 CO2 O3 CO3 S3 DI3 S2 DI2 S1 DI1 S0 DI0 CYINIT
3.FPGA-based TDC Development 78 Figure 3.4TDL TDC Architecture Overview The proposed TDC architecture is composed by two measurement units, one for fine time counting, which enables higher resolution to be achieved, and the other for coarse time counting, which secures large dynamic range. Because the two measurement units operate asynchronously, a synchronization module is required to ensure that the values used from the fine and coarse measurement modules are correct. Two different synchronizers were developed during this Thesis. Further details on these modules and on its implementation will be presented in Section 3.3. A merge block was also implemented to combine the two measurement values and generate a set of control signals. The interface to the TDC is done through a storage module implemented using a FIFO (dual-port RAM) which is also used as the clock domain crossing mechanism between the TDC and the implemented interface. Two different interfaces were developed. One to interface the TDC with the PS Arm processor on the FPGA platform used, which is based on the AXI-Lite protocol. A SPI slave interface was also developed to enable the TDC system to be used by other external microcontrollers. The two communication protocols are mutually exclusive, either the TDC is implemented using the AXI-Lite or using the SPI slave interface. With the general TDC architecture defined, it is now important to understand the multiple steps and resources involved in its implementation. These steps are closely coupled with the development tools used, in this case, the Vivado Design Suite framework. An overview of the design flow used and required Interface FPGA TDL Edge Detector Decoder x2 Coarse Counter FIFO Interface Synchronizer Merge Block Input Filter Time Interval Reference Clock ... wEn wdata rdatarEn Full Empty SPI SS SCK Mode MOSI MISO AXI Write Address Channel OR Write Data Channel Read Data Channel Read Address Channel aclk aresetn
Readout Circuit for Time-Based Automotive Sensors 79 resources can be seen in Figure 3.5. The digital flow adopted by Vivado has four main phases: System design entry, RTL Synthesis, Place & Route, and bitstream generation. System simulation can be done in-between each of the design steps. Figure 3.5FPGA Design Flow The description of the system can be done using VHDL, Verilog or using both HDLs. In this Thesis Verilog was used to implement the digital design of the TDC. Once the description is completed, it is possible to generate an RTL (Register Transfer Level) view of the implemented system and perform the behavioral simulation. The generated RTL view represents a generic implementation of the described system. However, this is a technology independent view of the system and thus, the subsequent steps may introduce multiple changes due to technology mapping and optimization algorithms. Moreover, this step allows for non-synthesizable code to be used which, although helpful for debugging, must be used only for simulation purposes and never to describe a functionality that is intended to be implemented in the final system. Nevertheless, the generation of the RTL view can be considered equivalent to the analysis phase during synthesis in the ASIC design flow. After, a behavioral simulation can be performed to assess our design and guarantee that it is functioning as intended. This is particularly useful to test the system’s state-machines and sequential behavior. Vivado can be configured to directly interface different simulators. During this Thesis the default Vivado simulator was used. Once the RTL is validated through behavioral simulation, the synthesis step is responsible to map the RTL code to the technology available on the selected platform. Since the Zybo Z7 board is mainly composed by LUTs, Storage elements and Carry arithmetic, all the combinatorial circuitry is converted to a logic expression and implemented using LUTs. Analogously, all registers triggered by a clock are mapped to storing elements. If any non-synthesizable code exists in the hardware description, a warning Design Closure HDL Code (VHDL, Verilog, ...) Test Bech RTL Simulation (Minimal Test) RTL Synthesis Place & Route Netlist Simulation Placed Netlist Simulation Bitstream Generation & Programming Constraints Hardware System Debug (Exhaustive Functional Testing)
3.FPGA-based TDC Development 80 or error will occur stopping the synthesis process. At the end of the synthesis step a new RTL description is obtained. This RTL description is no longer technology independent. Thus, functional and timing simulations are now possible. However, timing simulations prior the implementation step do not have information regarding placement and routing of the logic elements used. Therefore, these simulations only take into account the propagation delays on sequential storing elements and consider the propagation delay throughout combinatorial logic as ideal. Nevertheless, functional simulation in this step is important to guarantee that the optimization and technology mapping process did not change the intended behavior of the digital design. With the design synthesized and its functional behavior validated, the cells instantiated by the technology dependent RTL must be placed and routed. This step is done during the Place & Route phase (also known as implementation). Timing constraints are important throughout the entire design flow since the tool’s optimization algorithms make use of them to decide if a part of the design should be replicated or if buffers should be added to a given output. However, in the implementation phase, these constraints are very important, since these are the major drivers when deciding the optimal spot to place a logic element and which routing box and path must be selected. The lack of timing constraints may lead to design malfunction, solely due to arbitrary placement and routing. Upon completion of the implementation phase, functional and timing simulation of the placed and routed digital design can be performed. The main difference between the post-synthesis timing simulation and the post-implementation timing simulation is that, in the later, the propagation delay of the combinatorial circuits and the routing delays are also considered during simulation. This timing information is represented and described in a Standard Delay Format (SDF) file that is generated by Vivado during the implementation phase. All timing simulations are made considering the worst-case timing scenarios by default. Apart from timing constraints, physical constraints must also be defined to map the input and output ports of the design to the FPGA physical locations. At the end of the implementation phase, power, timing and resources utilization reports are made available to further analyze the final design result. The last step on the FPGA-based digital design flow is the generation of the bitstream used to program the FPGA. During this step a set of design rules, Design Rules Check (DRC), are made to validate the design. The output of this step is a .bit file that is used to configure the FPGA device according to the developed design.
Readout Circuit for Time-Based Automotive Sensors 81 Vivado also offers an IP integrator tool that enables multiple IPs to be instantiated, connected and validate. This functionality is especially useful to integrate the created designs with the processing system with minimal effort, since all the intermediary modules required are automatically generated by the framework. To make use of this functionality, the design must be encapsulated. More details regarding the created IP from the implemented design and the use of the IP integrator functionality will be presented in Section 3.3.4. 3.3. TDL TDC The TDL was one of the selected architectures to be implemented in FPGA due to its intrinsic digital nature, attractiveness for attaining a full autonomous migration for ASIC platforms in a later stage of this Thesis, and achievable high resolutions, due to the fast carry blocks available on Xilinx FPGAs. Although being a pure digital system, the design of a TDC demands for additional steps and considerations during implementation, mainly to avoid unwanted optimizations, automatically done by the framework in the various steps of the design flow. This section will describe in detail the design and implementation of the multiple modules presented on Figure 3.4. 3.3.1. Architecture Design The main block of the TDC system is the fine measurement module, since the major performance metrics of the TDC are defined or highly influenced by it. It is also the module that distinguishes the design from any other digital system design. Although a TDL is used as the base architecture for the fine measurement module, some changes were made to the typical approach in order to minimize resources utilization. The applications that this Thesis targets has both start and stop signals asynchronous to the reference clock. Usually, in such scenarios, two fine TDL measurement channels are implemented, one for measuring the start signal arrival time and another for the stop signal. In this Thesis, a fine measurement module designed with a single TDL for capturing both start and stop time interval is proposed. Using this approach, only the second stage sampling block and the decoder block must be replicated. Figure 3.6 depicts an overview of the RTL of the implemented fine measurement module. The decision of using a single TDL adds a timing constraint to the time interval to be measured. There must be a minimum time interval, equal to at least one reference clock cycle, between the start and stop
3.FPGA-based TDC Development 82 Figure 3.6TDL RTL Overview Decode Stop Decode Start Store Stop Stage Store Start Stage Delay Line Sample Stage CARRY4 CO[3:0] O[3:0] DI[3:0] S[3:0] CI CYINIT hit clk Store_start Store_stop FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE CO[3] GND VDD VDD GND FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE GND Q[3:0] Thermometer_stop_val_o[3:0] CARRY4 CO[3:0] O[3:0] DI[3:0] S[3:0] CI CYINIT FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE CO[3] VDD VDD GND FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE GND Q[3:0] GND CARRY4 CO[3:0] O[3:0] DI[3:0] S[3:0] CI CYINIT FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE VDD VDD GND FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE FDCE QCLR D C CE GND Q[3:0] GND ... ... ... ... Thermometer_start_val_o[3:0] Thermometer_stop_val_o[7:4] Thermometer_start_val_o[7:4] Thermometer_stop_val_o[255:252] Thermometer_start_val_o[255:252] Start_binary_val[7:0] Stop_binary_val[7:0]
Readout Circuit for Time-Based Automotive Sensors 83 signal in order to properly capture both start and stop time of arrival. Otherwise, multiple transitions will appear in the sampled TDL thermometer code, jeopardizing the measurement result. The main drawback of TDL architectures implemented in FPGA platforms is its poor linearity, due to process variation, which results in multiple steps having zero propagation delay, as mentioned before, this linearity issue is usually solved using decimation or bin-by-bin calibration. As already explained in Chapter 2, decimation sacrifices the TDC maximum achievable resolution, thus bin-by-bin calibration was the selected approach in this Thesis. This calibration technique can either be implemented directly in hardware or by software. Since one of the goals of this Thesis is to understand how to efficiently port a TDC design from a prototype platform to ASIC, and because the calibration tables required to implement bin-by-bin calibration require large chip areas to implement in ASIC platforms, it was decided that any sort of post-measurement calibration would be done by software and not directly implemented on hardware. Thus, bin-by-bin calibration tables, for the start and stop signals propagation delays, was built in software using the results obtained from a code density test with 100 thousand samples. The values received from the TDC are then used as index to address the calibration tables and the values returned are used in the final ToF measurement calculation. The principle of operation of the proposed TDC is as follows: the arrival of the start signal generates a rising edge that starts propagating throughout the delay chain, creating a 1-to-0 pattern on the last propagated step; on the following reference clock rise edge, the first sampling stage stores the state of the delay chain. Simultaneously, an edge detector module generates a start signal event. The first sampling stage is always enabled, updating the TDL state at each reference clock cycle; in order to secure a stabilized value for the decoding stage, a second sampling stage, enabled by the start signal event during one clock cycle, is also implemented. Furthermore, this double sampling method reduces the probability of metastability. A second sampling stage for the stop signal was implemented in an analogous way; two reference clock cycles after a start or stop signal, the second sampling stage has a stable value that can be used by the decoder to obtain the equivalent binary state of the TDL from the sampled thermometer code; the decoder module is a purely combinatorial priority encoder that converts the thermometer code sampled from the TDL to a binary value. This value corresponds to the position of the last step of the delay chain at logic level ‘1’. This approach has good performance and shields the TDC against bubbles since only the last ‘1’-to’0’ transition is considered. Further details regarding the implementation process to reduce bubble occurrence are discussed in Section 3.3.2; the start signal event generated by the edge detector module also enables the coarse counter module, which starts
3.FPGA-based TDC Development 84 incrementing at every reference clock cycle until the stop signal event is generated, disabling it and storing the value in a second set of registers; the process of obtaining the stop time interval is analogous to the start, but with the detection of a ‘0’-to-‘1’ pattern being propagated in the delay chain. Because the pattern to search is different, the decoder module must be different to the one used for the start signal. A waveform diagram exemplifying a typical measurement is shown in Figure 3.7. Figure 3.7Typical TDL TDC Operation Waveform The time interval to be measured, hereafter denoted as hit signal, is not directly connected to the TDL. Instead, an input stage is implemented to guarantee that no time interval measurement is accepted until the conversion of the previous one is finished. After reset, the input stage follows the hit signal. When a hit rise edge is detected, the input stage propagates it to the TDL and waits a falling edge transition of the hit signal. This event signalizes the stop signal. Upon the arrival of the stop signal, the input stage keeps the signal propagating to the TDL at a low logic level until receiving an end-of-conversion (EOF) signal, generated by the merge block. This EOF signal indicates that the time interval has successfully been measured and stored in the FIFO memory. To avoid situations where the hit signal might already be at a high logic level when the EOF signal is generated, the input stage checks if the input hit signal is at ‘0’. If this is the case, the input stage starts bypassing the hit signal to the TDL, otherwise, the input stage waits for the hit signal to return to ‘0’, while keeping ‘0’ at the input of the TDL. Only after does the circuit bypass the hit signal to the TDL again. 0xFFF...FFFFF clk hit 0x0 0x000...00000 0x000...00000 0xFFF...FFE00 0x00 0x09 0x01 0x02 0x03 0x04 0x05 0x7D 0x7E 0x7F Sample Stage Registers Store Start Stage Registers Start Binary Value Store Stop Stage Registers Stop Binary Value Coarse Counter 0x00000000 0x000...011111 0x0080 0x00 0x05 0x000...00000 0x000...00000 Delay Line 0x0..11 0x0000 0x0000 Sampled Coarse Counter 0x80 0x0000 0x0000 0xFFF...FFF 0xF..00 Start Event Stop Event TDC Value 0x00000000 0x00800509 End of Conversion
Readout Circuit for Time-Based Automotive Sensors 85 Figure 3.8 depicts the RTL view of the input stage and the waveform diagram demonstrating its normal operation and the scenario in which partial time measurement could have been done if the hit signal check was not performed. Figure 3.8Input Stage Schematic (top) and Operation Waveforms (bottom) The input stage also increases the flexibility of the TDL stage. Throughout the design of the TDC, it was assumed that the time interval to measure (hit signal) would have a pulse shape, being the rising edge of the hit signal used as the start event and the falling edge as the stop event. However, multiple applications require the generation of a pulse for the start event and another pulse for the stop event, being the relevant information the time between pulses and not the pulse’s duration. To target such applications while using the same architecture, it is enough to add a D-type flip-flop with asynchronous reset at the beginning of the input stage, maintaining the remaining modules unaltered. The start event could be used to set the flip-flop while the stop event would be responsible for resetting it. Thus, a pulse would be generated with a width proportional to the time interval between the start and stop pulses. This would introduce an error on the measurement due to different propagation of the clock and reset signals on the flip-flop. However, this would be just an offset error correctable by software. The coarse counter module is a 16-bit binary counter with enable, which is incremented by one at every reference clock cycle. The module has a second set of registers to store the counting value when a stop event is detected. The synchronizer block is mandatory to guarantee the proper operation between the asynchronous and synchronous part of the TDC system. As already mentioned, since the hit signal is asynchronous to the reference clock, scenarios where the time of arrival of the hit signal violates the setup or hold time of the nrst Input Stage hit D Q clk clr D Q clk clr 1 1 End of Convertion Filtered_hit O1 O2 O3 O4 hit O1 O2 End of Conversion nrst O3 O4 Filtered_hit Ignored hit due to arrival before end of conversion
3.FPGA-based TDC Development 92 the carry delay chain can only be propagated upwards, each CLB has four carry elements grouped up in a Carry4 block, and the Zybo Z7-10 platform has 50 rows per clock region, it is not possible to keep the TDL inside the same clock region. For this reason, an ultra-wide bin is expected around the 200th step. This could be avoided if LUT elements were used to implement the TDL, since they do not have the restriction of only enable a propagating chain upwards. However, LUTs have higher propagation delay (approximately 123 ps in the Zybo Z7-10 according to the worst-case timing simulations), which would greatly reduce the TDCs resolution. Moreover, LUTs do not have a dedicated routing like the one connecting all the carry blocks in a column, instead LUTs use the route boxes. This would ultimately result in longer propagation delays due to longer routing paths and worse linearity performance as result of the non-uniform routing across the multiple TDL steps. Thus, having some ultra-wide bins was considered preferable, since the issue can be easily targeted by a calibration mechanism. The routing between the output of the carry elements and the sampling stage registers is also an important factor to achieve higher linearity across the delay chain. While it is enough to constraint the propagation delay between the first and second sampling stages to one reference clock cycle, between the delay chain and the first sampling stage, the routing must be uniform across all steps and as small as possible. Therefore, the first sampling stage registers must be placed inside the same Slice as the carry elements they are sampling. FPGAs have a highly optimized clock tree structure. Thus, inside the same clock region the routing skew has typical values under 30 ps. However, the same cannot be said when multiple clock regions are used since the insertion delays from the clock input pin to the distribution buffers of each region varies, thus increasing the skew between clock regions. Controlling the clock signal in the TDC architecture presented is not as critical as controlling the hit signal. As mentioned during the presentation of the synchronization module, the hit signal insertion delay to the TDL and enable pins of the coarse counter registers must be closely match. Otherwise, the synchronization window will suffer a shift equal to the hit signal skew. The insertion delay of the hit signal can be controlled by adding LUTs configured as buffers to delay the signal on the fastest paths or by rerouting it.
Readout Circuit for Time-Based Automotive Sensors 93 3.3.4. Interface To interface the TDC and store the measurement values, a FIFO module was implemented. Apart from performing a clock domain crossing mechanism, the FIFO module is also useful to temporarily store some measurement values when the processor’s reading rate is lower than the acquisition rate of the TDC. The signals from the FIFO module were designed to enable the simple exchange of communication protocols, increasing the system modularity. In fact, the FIFO can also be addressed directly, reading the 32-bit output in parallel, with no communication protocol in-between the TDC and the processor reading from it. The FIFO module is composed by a dual port RAM memory and two smaller modules to generate the write and read address pointers and the full and empty flags. Because the empty and full flags are generated by comparing the write and read pointers and these are generated in two different clock regions, a clock domain crossing module was implemented, using a double register method. Figure 3.14 depicts an overview of the FIFO implementation (top) and the write pointer and full flag module (bottom). Figure 3.14FIFO Module Overview and FIFO write pointer and Full Flag Generation module FIFO wRst wInc D Q clk clr D Q clk clr rInc wClk rClk rData wData rRst full empty Dual-Port RAM Wptr & Full Gen Rptr & Empty Gen rptr_sync wptr D Q clk clr D Q clk clr waddr raddr inc wEn DataIn DataOut wptr_sync inc raddr waddr rptr wrst rrst x9 x9 x9 x9 Wptr & Full Gen D Q clk clr x9 + wInc wClk Binary to Gray logic [8:0] waddr D Q clk clr x9 wRst wptr == rptr_sync full
3.FPGA-based TDC Development 94 The write and read address pointers generators are similar. The structure of the implemented write pointer generator is depicted at the bottom of Figure 3.14. The system is based on a (n+1)-bit binary counter, to address the 2n FIFO memory positions. The extra bit is used to calculate the empty and full flags. If the n least significant bits of the read and write pointers are the same, then the logic XOR of the n+1 bit from those pointers define whether the FIFO is full or empty (empty when the MSB are the same and full when they differ). Before crossing the write pointer to the read clock domain, the value is converted from binary to Gray-code. The same is done when passing the read pointer to the write clock domain. This is done to avoid multiple bit state changes when the pointers’ values are updated, increasing the system’s robustness. The FPGA Processing System’s Arm Cortex-A9 uses an AXI bus to communicate with internally mapped peripherals. Using the Vivado framework it is possible to automatically generate a 32or 64-bit AXI4 slave or master interface, with a default state-machine implemented, to encapsulate custom made IPs and automatically map them into the processor’s peripheral address space. This functionality was used to automatically generate a 32-bit AXI-Lite slave interface. By default, four registers were instantiated in the AXI-Lite state machine to communicate with the processor. Only two of those registers are used, one to read the next value from the TDC’s FIFO and another used by the processor to send commands to the implemented TDC IP. Since the AXI interface presented in the Arm Cortex-A9 processor is still a legacy AXI3 version, it was necessary to implement a bridge between it and the AXI4 interface encapsulating the TDC peripheral. This bridge may be generated by the Vivado framework or manually instantiated by the user when using the IP Integrator tool. With the TDC IP implemented and mapped into the processor’s memory, the last step was to develop a software application to read the values from the TDC peripheral. The Xilinx Software Development Kit (XSDK) is integrated in Vivado framework and enables the development of embedded application, grants access to the automatically generated Board Support Package (BSP) of the implemented system and provides a full debugging environment. The algorithm for the software application to develop is as follows: first, the processor sends a read command to the TDC FIFO by writing to the memory position of register 1 of the TDC’s AXI interface. Then the processor reads from the address of TDC’s register 0, which has the updated value from the TDC FIFO. The 32-bit value received is then decoded to obtain the coarse, fine start and fine stop measurement values (see Figure 3.15 and Figure 3.16).
Readout Circuit for Time-Based Automotive Sensors 95 The TDC AXI peripheral was mapped in memory from address 0x43C00_0000 to 0x43C00_FFFF, being the base address represented by the macro XPAR_TDC_0_S00_AXI_BASEADDR in Figure 3.16. TDC_S00_AXI_SLV_REG0_OFFSET and TDC_S00_AXI_SLV_REG1_OFFSET are macros representing the offset address of the TDC AXI registers (0 in the case of register 0 and 1 for the register 1). Figure 3.15TDC Read Application Flow Figure 3.16TDC Read Application Send Read Request to AXI register 1 Read Measure Application Init Device While true Read from AXI register 0 to result variable Clear Read Request on AXI register 1 Decode result (stop = result[7:0]) (start = result[15:8]) (coarse = result[31:16]) Calculate and print time interval measured true false Clear Device Exit Read Measure Application
3.FPGA-based TDC Development 96 The coarse value was multiplied by the reference clock frequency and the values from the fine start and stop measurements are used to address the calibration table. The final measurement result can be obtained using equation (3.1): 𝑡=𝑐𝑜𝑎𝑟𝑠𝑒𝑐𝑜𝑢𝑛𝑡𝑠 ∗1 𝑇𝐶𝐿𝐾 +(𝑠𝑡𝑎𝑟𝑡𝑐𝑎𝑙𝑖𝑏𝑟𝑎𝑡𝑒𝑑 −𝑠𝑡𝑜𝑝𝑐𝑎𝑙𝑖𝑏𝑟𝑎𝑡𝑒𝑑), (3.1) 3.4. Gray-Code TDC The Gray-code oscillator architecture was first proposed in [3.4]. The architecture reported a mean bin size of 256 ps and 271 ps, for the two TDC channels implemented, using 8-LUTs and 8 flip-flops per channel, being its main components a gray counter and a 2-bit oscillations counter. This architecture was developed in order to achieve high resolution for a low-power and low-resource system. Typically, a counter is composed by a combinatorial stage, which calculates the next value in the counting schema, and by a sampling stage, responsible for latching the value of the counter at each clock cycle, assuring a stable value for the combinatorial stage, so that the next value can be correctly calculated and latched in the next clock cycle. This is of extreme importance for binary counters in which multiple bits can change from one counting value to the next, activating multiple combinatorial paths at the same time. For example, in a 4-bit value, during the increment from seven (0111) to eight (1000), all the bits change. If no latching stage is present, i.e. the output of the combinatorial stage is directly connected to its inputs, there is the risk of having random values, and therefore a random counting sequence, due to different propagation delays in the counter’s datapath. However, if a Gray-code counting schema is used, this problem is avoided. In the original Gray-code, only one-bit changes from one state to the next one. Thus, the Gray-code can be configured in a loop, without the latching stage, since there is no risk of missing codes or making a random counting sequence. This enables the implementation of a counter with a resolution that is no longer limited by the system clock used to sample the state of the counter. Instead, the maximum achievable resolution is given by the propagation delay of the counter’s signals trough the Datapath (cells’ propagation delay plus routing delays). The Gray-code counter implemented in [3.4] is based on the a 5-bit reflected binary code (RBC) schema. In such schema, the first bit of the Gray-code does not have a dependence on itself. As the Gray-code has only 5-bits, a single 5-input LUT is enough to calculate each bit next state. However, because the counter also needs a bit to enable the counter, 6-input LUTs are used to implement it. Apart from the Gray-code
Readout Circuit for Time-Based Automotive Sensors 97 counter, a 3-bit cycle counter was also included, which counts the number of full counts performed by the Gray-code counter. This mechanism was implemented to extend the range of the TDC, enabling a lower clock frequency to be used in the final system. From the three bits, only two are used. The decision of which bits should be used is done depending on the output of the Gray-code TDC. This guarantees that the selected bits are not metastable. The architecture was implemented in a Kintex-7, a Xilinx 28 nm technology FPGA. In this platform, as seen in Section 3.1, each CLB has two Slices with four 6-input LUTs, a 4-bit Carry block and eight flip-flops each. Therefore, a single TDC channel can be implemented in a single CLB, allowing for multiple channel implementation even on small FPGAs. The research reports step delays in-between 100 ps and 500 ps, which, for a mean step delay of 256 ps corresponds to a DNL in the range of -0.61 LSB to +0.95 LSB. To obtain these results, the building cells of the TDC were manually placed inside the same CLB, in a specific order. By the analysis of the steps’ delays for the two TDC channels presented in [3.4], it is possible to notice that, for some steps, the same TDC step in different TDC channels has delay differences greater than 200 ps, leading to the conclusion that placement influences TDC channel performance. Finally, to improve linearity, a four measurement per input pulse, followed by averaging, is proposed in [3.4]. This method, although enabling better linearity with no extra resource usage, decreases the system throughput. The Gray-code architecture was studied during this Thesis research, since it could be interesting for Flash LiDAR applications due to its very low resource utilization. In the following sections, a proposal for improving the base Gray-code TDC architecture linearity and scalability based on controlling the routing propagation delay is presented. 3.4.1. Architecture Design Based on the wire load regulation principle presented in [3.5], and the base TDC architecture presented by Wu and Xu in [3.4], an improved linearity Gray-code TDC architecture was designed during this Thesis research. Furthermore, the followed approach enabled the improvement of the TDC architecture scalability, by reducing the channels mismatch when implementing multiple TDC channels. Thus, improved performance is obtained reducing the need for calibration circuitry or post-measurement calibration software routines. The base block diagram of the proposed architecture is depicted in Figure 3.17, and Figure 3.18 present the logic equations to calculate each Gray-code bit. The core of the TDC channel, presented in Figure 3.19, is a pure combinatorial 5-bit Gray-code counter that, when enabled,
3.FPGA-based TDC Development 98 starts looping through the 32 possible values. The Gray counter is enabled on the rise edge of the hit signal. To avoid that the counter continues to oscillate undefinably, the input stage presented in Figure 3.17 was used to guarantee that the enable for the Gray counter has a maximum duration of one reference clock cycle. After the arrival of a hit signal, on the next clock rise edge, the value of the Graycode counter is sampled and simultaneously, hit_r (the output of the Input Stage register) is cleared thus stopping the Gray counter. When the value sampled is different from zero, a store signal is generated to sample the value of a free-running coarse counter, used to increase the measurement range of the TDC. The Gray-Code value sampled is also stored in a second set of registers on the second rising edge of the reference clock. Like in the TDL case, this is done to reduce the probability of metastability on the sampling of measurement value from the Gray counter and to secure a stable value for the next operations (the first set of registers samples the Gray-code at every clock period and therefore, if the value was not stored, the measurement value would be lost at the second rise edge of the clock after the hit’s signal arrival). Figure 3.17Gray-Code Architecture Overview Figure 3.185-bit Gray-Code Logic Equations FPGA TDC Channel (Start) TDC Channel (Stop) Input Stage (Start) Input Stage (Stop) Q Q SET CLR D hit nrst clk VDD Coarse Counter Merge FIFO Interface store_start store_stop
Readout Circuit for Time-Based Automotive Sensors 99 Figure 3.19Gray-Code channel RTL View A second TDC channel was implemented to measure the arrival time of the hit’s signal falling edge. This way, it was possible to measure the time of an impulse. A waveform, with the typical TDC’s channel signals, during a measurement procedure, is depicted in Figure 3.20, and in Figure 3.21 the state machine of a measurement process is presented. Differently to the architecture presented by Wu and Xu [3.4], no bit extension was used in this implementation of the TDC channel. Therefore, the maximum counting value, assuming the 250 ps average delay per stage reported for the 7-Series Xilinx FPGA, is 8 ns. This constrains the minimum clock of the system to 125 MHz, which, for modern FPGAs, is a constraint easy to meet. Although this decision forces the use of a faster clock, which can lead to higher power consumption, it also enables to save three LUTs, three flip-flops and one Carry4 element per TDC channel. CLB_XnYm SLICEM_Xn+1Ym CLB_XnYm SLICEL_Xn-1Ym SLICEL_XnYm I0 I1 I2 I3 I4 I5 O6 bit4 bit3 bit2 bit1 bit0 hit D C Q bit4 Gray4 clk D C QGray4_stored store E LUT FDCE FDCE I0 I1 I2 I3 I4 I5 O6 bit4 bit3 bit2 bit1 bit0 hit D C Q bit3 Gray3 D C QGray3_stored E LUT FDCE FDCE I0 I1 I2 I3 I4 I5 O6 bit4 bit3 bit2 bit1 bit0 hit D C Q bit2 Gray2 D C QGray2_stored E LUT FDCE FDCE I0 I1 I2 I3 I4 I5 O6 bit4 bit3 bit2 bit1 bit0 hit D C Q bit1 Gray1 D C QGray1_stored E LUT FDCE FDCE I0 I1 I2 I3 I4 I5 O6 bit4 bit3 bit2 bit1 bit0 hit D C Q bit0 Gray0 D C QGray0_stored E LUT FDCE FDCE SLICEL_Xn+2Ym Gray4 Gray3 Gray2 Gray1 Gray0 store
3.FPGA-based TDC Development 100 Figure 3.20Gray-Code TDC Typical Measurement Waveform Figure 3.21Gray-Code TDC State-Machine Since all the modules developed for the TDL architecture were designed targeting portability and modularity, the new Gray-code TDC channel was integrated with the synchronizer, merge, FIFO and AXI-Lite Interface modules automatically. It was only necessary to substitute the TDL Channel module by the Gray-code TDC module. clk hit 2 hit_start hit_stop 1 start_gray_cnt 0 3 4 5 6 07 stop_gray_cnt 210 3 4 5 6 0 start_sampled stop_sampled 0 6 0 0 5 0 start_store registers 0 6 stop_store registers 0 5 start_store stop_store Measuring Measuring (gray code) Sample Store (start_store) Wait hit_start 1 1 1 !count_reset Idle Measuring Merge Store Count Reset !hit_start hit_start !stop_store stop_store !(end_of_conversion) (end_of_conversion) 1 hit !hit/ count_reset hit_stop Measuring (gray code) Sample Store (stop_store) Wait 1 1 1 !count_reset Idle Idle count_reset count_reset
Readout Circuit for Time-Based Automotive Sensors 101 3.4.2. Implementation notes and Layout Considerations Gray’s Counter oscillator Datapath Analysis Based on the results reported in [3.5], the propagation delay of a cell in FPGA platforms can be partially controlled by controlling the load of the cell. This can be done by adding dummy buffers to the output of a cell, to increase its load, or by increasing/decreasing the size of the routing, which will increase/decrease the parasitic capacitance of the net, changing the load at the output of the cell. The second option is more suitable for smaller adjustments, either to increase or reduce the propagation delay, but demands for a great knowledge on the FPGA routing resources, making its implementation more complex. Furthermore, the propagation delays obtained from the tool, which are dependent on the wire loads, are always the worst-case scenario. Nevertheless, the second option does not require extra resource usage, apart from the routing resources which are required for both methods. By assuring a close match between the propagation delays in the worst-case scenario, it is possible to assume that the typical conditions would be similar as well. Moreover, as the TDC channel can be confined to a single CLB, the voltage and temperature conditions should be similar in all the five LUTs used to build the Gray-code oscillator. Therefore, exploring the routing possibilities during the layout of the Gray-code TDC may lead to an implementation with higher linearity with no extra hardware cost, improving the overall system’s performance. By analyzing the pattern on the 5-bit Gray-code it is possible to conclude that only 8 out of the 24 datapath connections affect the size of the steps. Namely the paths from the output of LUT0 to the input of all the other LUTs (four connections) and the output of all the other LUTs to the inputs of LUT0 (another four connections, see Table 3.2 and Figure 3.19). This greatly reduces the effort of implementing the linearity correction through routing. The remaining 16 datapath will not affect the delay of the steps (as long as the propagation delay of these paths do not exceed two clock cycles), and therefore can be automatically routed by the framework. Manual Routing to Control Datapath’s delay The manual routing process can be done in Vivado design tool using the implemented design graphical interface, or by creating a file with the set of physical constraints annotated in a Xilinx Design Constraints (XDC) format. If the graphical interface is used, the manual routing is done by entering the implementation design view and selecting the routing resources option of the Device tab. Then the routings on the nets
3.FPGA-based TDC Development 108 pattern identified before (in Section 3.4). Thus, if the routing timings is subtracted to the times obtained from the simulation, the worst-case scenario of the LUT’s propagation delay can be calculated according to equation (3.2). Figure 3.27Detailed View of the Gray-Code Sequence Generation and step size 𝑡𝐿𝑈𝑇𝑖 =𝑡𝑃𝐷 −𝑡𝑅𝑂𝑈𝑇𝐸𝑖, (3.2) where tLUTi is the LUT propagation delay, tROUTEi is the propagation delay of the LUT’s output wire to the LUT which will change state for the next count, and tPD is the total propagation delay, extracted from the simulation. Taking the LUT responsible for generating the least significant Gray-code bit and the 1 (00001) to 2 (00011) transition as an example, the LUT propagation delay would be equal to 123 ps (999 ps - 876 ps, see Figure 3.27 and Table 3.3). As in the TDL scenario, 123 ps is the worst-case propagation delay for all the LUTs. These results conform with the information on the Xilinx datasheet stating that the LUT’s propagation delay is independent of the truth table being implemented. Since the first LUT has one less output connection than the remaining LUTs, but all of them have the same propagation delay, one can also conclude that the LUT’s propagation delay is independent of its output load (at least for the timing simulation). Furthermore, the routing resources appear as the main delay source, further supporting the statement made previously regarding the impact of controlling it to achieve better linearity. The experimental results presented in Section 3.7 study the differences between manual and automatic routing and its impact on the TDC channel linearity and scalability. 3.7. TDC Performance Assessment The TDC architectures were deployed in Xilinx Zybo Z7 development board (depicted in Figure 3.1). The Tektronix AFG1022 arbitrary waveform generator was used to generate the time intervals to assess the developed TDCs. Figure 3.28 depicts the test setup used during all the tests performed. It is composed clk_i 0x00 0x0b hit_i start_val_w[4:0] 0x01 0x03 0x02 0x06 0x07 0x05 0x04 0x0c 0x0d 0x0f 0x0e 0x0a 999ps 600ps 809ps 704ps 999ps 600ps814ps 832ps 999ps 600ps809ps704ps
Readout Circuit for Time-Based Automotive Sensors 109 Table 3.3Routing Delays for the Gray-Code TDC STOP CHANNEL MANUAL ROUTING AUTOMATIC ROUTING LUT4 LUT3 LUT2 LUT1 LUT0 LUT4 LUT3 LUT2 LUT1 LUT0 BIT0 665 691 686 876 - 664 701 696 295 - BIT1 700 198 193 475 477 700 198 193 475 477 BIT2 911 732 737 162 580 909 730 735 1080 165 BIT3 394 306 307 711 709 394 306 307 711 709 BIT4 296 916 914 518 513 297 612 609 330 514 by the development board, the waveform generator and a host PC running a MATLAB script to analyze the data read from the board. The TDC IP with the AXI interface and the integrated Arm processor was used (the SPI interface was built to target the ASIC solution). For each architecture, a code density test was performed to extract the real delay of each step of the fine interpolation stage. The results from these tests were used to calculate the non-linearity of the fine measurement module, namely, the DNL and INL. Figure 3.28FPGA Test Setup A total of 100 thousand measurements were made to reduce probabilistic errors. A single-shot precision test was also performed, with 100 thousand samples. Then, to reduce the influence of the errors introduced by the arbitrary waveform generator, a 10 measurements average was made, and the precision recalculated. To understand the impact of a calibration mechanism in the TDCs’ performance, a post-measurement software bin-by-bin calibration was applied to the 100 thousand samples collected during the single-shot precision test. The calibration tables were created based on the results from the code density test. Then, the single-shot precision and average precision were recalculated after applying the calibration. All the tests were performed at ambient temperature of 25°C and with a power supply of 3.3V. The following assessment results are presented by TDC architecture.
3.FPGA-based TDC Development 110 3.7.1. TDL TDC Code Density Test In order to evaluate the real delay distribution across the implemented TDL, a code density test was performed. The waveform generator was configured to output a square wave signal at a frequency unrelated with the 250 MHz reference clock used, thus creating a sliding window effect on the sampling steps of the TDC, which, in an ideal scenario, would have the same probability to be sampled. The selected frequency was 999133 Hz. The code density test for the start and stop events propagation is presented in Figure 3.29. Regarding the start signal, it is possible to notice that no step was captured prior to the 22nd step. The last step through which the start event was able to propagate was the 254th. Thus, the start signal can propagate through a total of 232 steps in one clock period. Since the reference clock used is operating at 250 MHz, the average delay of each step when propagating the start signal is 17.2 ps. Notice that almost half of the steps have zero delay. This is mainly caused due to process variation and it is the main reason for the mandatory implementation of a calibration mechanism in FPGA-based TDL TDCs. Since the proposed architecture uses the same TDL to measure both time events, a similar behavior was expected for the stop signal propagation. This can be observed when analyzing the results from the stop signal code density test. The ultra-wide steps are the same (for example, the 200th step) and the number of steps with zero propagation delay is also similar. The main difference is regarding the first bin to be sampled that, in the stop propagation scenario is the 10th step. Since a rising edge (start signal) and a falling edge (stop signal) have different propagation behaviors, this difference was expected. Thus, the average delay per step when propagating a stop signal is 16.4 ps. The values obtained from the code density test were used to create two calibration tables. Each of the rows in filled with the cumulative sum of the delays of the TDL steps until that row position, i.e. at the 70th row, the value would be equal to the sum of the delays of the first 70 steps of the delay line and so on. 3.7.2. TDL TDC Linearity The DNL and INL of the TDL were calculated using the data obtained from the code density test, according to the equations (2.5) and (2.6) presented on Chapter 2 (considering 17.2 ps and 16.4 ps as the LSB for start and stop respectively). Figure 3.30 depicts the non-linearity for the start and stop events propagation. For the start event propagation, the maximum DNL is equal to 3.3 LSB (56.16 ps), while for the stop
Readout Circuit for Time-Based Automotive Sensors 111 event, a maximum DNL of 3.7 LSB (60.84 ps) was obtained. The INL ranges between -3.8 and 1.8 LSB when propagating the start event, and -3.7 and 2 LSB for the stop event. Figure 3.29TDL TDC Code Density Test for Start (top) and Stop (bottom) event propagation Figure 3.30TDL TDC Linearity results for start (left) and stop (right) signals propagation Time (ps) Bin Number Bin Number Time (ps)
3.FPGA-based TDC Development 112 3.7.3. TDL TDC Precision To assess the TDC precision, the waveform generator was configured to output a square-wave with 999.133 kHz frequency with 50% duty-cycle. The output frequency was verified using an oscilloscope to check the duration of the pulse between the rising and falling edge of the signal (portion of the signal measured by the TDC). A duration between 480.242 ns and 480.434 ns was observed. From the code density test performed and the linearity results obtained, it is possible to conclude that the TDC precision will be considerably affected if equation (2.14) from Chapter 2 (which uses an average cell delay value) is used to calculate the time interval measurement. This is highlighted in Figure 3.31, which presents the results from the single-shot measurement without calibration (top) and with calibration (bottom). A precision improvement of 40.3 ps (2.4 LSB) was obtained when calibration is applied. The raw precision test shows a 481.007 ns average measurement (573 ps offset regarding the expected value) and a precision of 211 ps. After calibration, the precision is improved to 179 ps, with an average time interval measurement of 480.891 ns (457 ps offset regarding the expected value). Figure 3.31TDL TDC Single-Shot Precision before (top) and after (bottom) calibration However, even with calibration, the obtained precision is still far from the ideal LSB size (approximately 17 ps). The reason for such results may be explained by two factors: first, the existence of multiple ultra-
Readout Circuit for Time-Based Automotive Sensors 113 wide bins along the TDL deteriorates the TDC’s precision, even when the real cell delays are used to calculate the time interval measured; second, the arbitrary waveform generator jitter and noise introduces errors to the time interval generated. To reduce the influence of the errors introduced by these factors, an average of 10 measurements was performed for both raw and calibrated data. The results are depicted in Figure 3.32 (being the non-calibrated measurement precision depicted in the top graph and the calibrated measurement precision on the bottom graph). A precision of 59 ps and 56.7 ps was attained for the raw and calibrated data, respectively. Figure 3.32TDL TDC 10 Measurement Average Precision before and after Calibration 3.7.4. The synchronizer contribution In order to understand the necessity of a synchronizer block, a set of measurements with the synchronizer disabled were performed. The results obtained for 1,000 measurements with and without the synchronizer implemented are presented in Figure 3.33, being the results without synchronizer displayed at the top of the figure while the results with the synchronizer enabled displayed at the bottom. As can be observed in Figure 3.33, multiple errors in the range of ±1 coarse counter LSB appear in the
3.FPGA-based TDC Development 114 measurements. Moreover, there are even measurements where the error is equal to multiple clock cycles. This is because the coarse counter is binary, thus multiple bits can change simultaneously. When this happens, if only some bits are properly updated, errors equal to multiple clock periods can appear. With the synchronizer implemented, no measurement deviations equal or greater than the system clock period were recorded. Thus, it is possible to conclude that, in the FPGA implementation, the synchronizer module is working properly. Figure 3.33Synchronizer Effect on TDC Measurement Value Output. 3.7.5. Gray-code TDC Code Density Test The same setup used to test the TDL TDC was used during the assessment of the Gray TDC. However, since the proposed Gray TDC architecture was targeting an improvement on the TDC linearity, two different implementations were deployed to FPGA and tested. The first implementation followed the strategy adopted by Wu and Xu [3.4], constraining just the placement of the TDC’s channel LUTs and storing elements, while the routing was performed automatically, using the Vivado framework default implementation run. Time(µs) Sample Number Time(µs) Sample Number Time(µs) Sample Number 4ns Error 130ns Error 4ns Error 200ps Error
Readout Circuit for Time-Based Automotive Sensors 115 The second implementation followed the design flow proposed in this Thesis, first placing the TDC’s channel cells, then letting the tool perform automatic routing using a low_net_latency implementation run, and finally manually routing the nets identified to try to have equal parasitic capacitances. The TDC's start and stop channels' code density test results, for both implementations (default and manual routing), are presented in Figure 3.34 (being the default results are presented on the left while the manual routing implementation results are presented on the right). Since reducing the delays of the slowest nets is usually harder (most of the times impossible when using the low_net_latency option), the strategy adopted was to increase the delay of the fastest nets. This resulted in a higher average step delay. However, as can be seen in Figure 3.34, it reduces the delay differences across the TDC channel. Figure 3.34Gray-Code TDC Code Density Test for Default and Manual Routings The previously presented Table 3.3 depicts the preand post-manual routing net delays for the stop channel. A 125 MHz reference clock was used in the Gray-code TDC. Thus, the average step delay is 380.9 ps, a 33.1 ps increase regarding the solution where only the placement is constrained. Another important factor to notice is the delay uniformity across channels. When manual routing is performed, the start and stop channels steps’ delays are closely matched with a maximum difference of 95 ps for the same step (for the worst-case scenario), while the automatic routing implementation presented a Time(ps) Time(ps) Time(ps) Time(ps) Bin Number Bin Number Bin Number Bin Number Stop Propagation Stop Propagation Start Propagation Start Propagation 340ps 400ps 510ps 215ps 240ps 320ps 225ps τavgstart= 347.8ps τavgstop= 307.7ps Δmaxstart= 400ps Δmaxstop= 340ps Δmaxstartstop= 295ps τavgstart= 380.9ps τavgstop= 380.9ps Δmaxstart= 250ps Δmaxstop= 240ps Δmaxstartstop= 95ps 250ps
3.FPGA-based TDC Development 116 maximum difference of 295 ps. When analyzing the work by Wu and Xu, a maximum difference of 230 ps can be observed. Thus, the proposed implementation improves the TDC scalability, since it will secure a uniform performance when multiple channels are implemented. Furthermore, the same calibration mechanism, for instance a single bin-by-bin calibration table, can be deployed to calibrate multiple TDC channels due to its similar delays’ distribution, enabling resource and power savings. 3.7.6. Gray-code TDC Linearity The DNL and INL calculated from the code density test results is presented in Figure 3.35 (being the default results are presented on the left while the manual routing implementation results are presented on the right). As expected, the manual routing scenario shows higher linearity with a maximum DNL of 0.38 LSB and an INL in the range of 0.01 and 0.7 LSB for the start channel (worst case) (see Figure 3.35). Figure 3.35Gray-Code TDC Linearity Results for Default and Manual Routing
Readout Circuit for Time-Based Automotive Sensors 117 3.7.7. Gray-code TDC Precision Although calibration can be applied to the TDC channel, given the obtained linearity, high performance without calibration is expected. The single-shot precision test results are presented in Figure 3.36 (being the default results are presented on the left while the manual routing implementation results are presented on the right). A difference of 2.2 ps precision can be seen when bin-by-bin calibration is applied to the manual routing implementation. On the automatic routing implementation, the precision difference is about 109 ps. Thus, the proposed design flow proved to be efficient in improving the performance of the TDC channel without extra resource costs. The trade-off is purely in the design stage which require an additional manual step. Figure 3.36Gray-Code TDC Single-Shot Precision for Default and Manual Routing Again, to reduce the influence of the errors introduced be the waveform generator, another test, in which each measure was a 10-measurement average was performed. The precision results are depicted in Figure 3.37 (being the default results are presented on the left while the manual routing implementation results are presented on the right). Here is important to highlight that the non-calibrated average precision of the manual routed TDC was able to surpass the average precision of the calibrated default routed TDC. Precision Before Calibration Precision After Calibration Precision Before Calibration Precision After Calibration 40000 30000 20000 10000 60000 40000 20000 Counts Counts
4.ASIC-based TDC Development 124 4.1.1. Development tools A multivendor development environment was adopted, following the current industry trend. The Verilog analysis and RTL synthesis were done using Synopsys’ Design Compiler. The Layout of the obtained Netlist was performed using Cadence’s Innovus, and SimVision was used during behavioral and functional simulations. The pad ring definition and final DRC and LVS checks were done using Cadence’s Virtuoso. Design Compiler (DC) According to Synopsys Design Compiler datasheet [4.1], DC Ultra is a RTL synthesis tool which performs concurrent timing, area and power optimization to guarantee better time Quality of Results (QoR). Apart from analysis, elaboration, compilation and static timing analysis of the developed design, DC Ultra offers a set of graphical interfaces and tools to help study the design’s critical paths and possible congestion areas. On later versions of the software, cross-probing between the RTL source code and other design views (like the Netlist schematic view) was introduced, enabling designers to have better control and understanding regarding the optimizations and design changes that are being performed by the tool, and efficiently identify possible design issues in early stages of the development [4.1], [4.2]. The main optimization operations performed by the tool are based on arithmetic changes, logic duplication (to reduce the high fan-out loads on critical paths), design ungrouping (to reduce silicon area and achieve better timing performance), buffer insertion on high fan-out nets (to improve the total negative slack), and register retiming (which can also add pipeline registers on long combination paths) [4.1], [4.3]. The DC Ultra also enables the automatic addition of scan registers for test and debugging purposes. Innovus Genus is Cadence’s framework for digital IC design. From the set of tools included, Innovus is used in the layout step (place and route). One of the main distinctive features of this tool, when compared to its competitors is the placement engine, GigaPlace. The GigaPlace follows a slack driven approach, as opposed to the traditional “time-aware” approach [4.4], [4.5]. According to Cadence [4.4], this enables a concurrent convergent optimization of both electrical and physical metrics. The new placement engine builds slack models considering the design floorplan, routings topologies, congestion and other electrical constraints, and optimizes the design placement. Another distinctive feature is the clock tree synthesis (CTS) engine. In order to improve useful skew and merge physical optimization with clock tree synthesis, Innovus introduces a new CTS engine named Clock
Readout Circuit for Time-Based Automotive Sensors 125 Concurrent Optimization (CCOpt). The optimizations are based on true propagated clocks and consider on-chip variations (OCV) [4.4]. Apart from placing and routing the design, the tool also offers verification and validation options in order to test the implemented layout against the technology’s DRCs. Similar to Synopsys’ tools, Innovus has a TCL interface and a graphical interface, in which all the design characteristics can be inspected in detail, including, among others, power, electromigration, routing congestion analysis, clock tree debugging, design hierarchical view search. SimVision SimVision, also known as NCSim, is a Cadence software tool for debugging digital, analog, and mixed-signal designs [4.6]. It supports testbenches written in Verilog, SystemVerilog, VHDL and SystemC languages (or a combination of these). This tool offers multiple graphical functionalities that simplify the debugging process and can be used for behavioral and functional simulation. It includes support for functional timing simulations using standard delay format (SDF) files, ideal to validate a post-layout design. Virtuoso Virtuoso is an analog and mixed-signal design environment that offers multiple capabilities for electrical analysis and verification [4.7]. A graphical user interface and a TCL-based command line supports the development, when using this tool. In this Thesis, Virtuoso was used solely to create the pad ring for the developed chip, to perform the last DRC and LVS verifications and validations, and to tape-out the design. 4.1.2. Technology adopted Depending on the technology adopted, some files required by the IC design tools to perform various verifications might not be available. This limits the depth of the analysis that can be performed. For instance, if only the worst-case capacitance tables ( captables ) are available on a technology pack, then, a best-case scenario analysis cannot be performed. Furthermore, if the layout of the digital cells is not available, it is not possible to perform a complete Design Rules Check (DRC) and Layout Versus Schematic (LVS) verification. The set of available technologies for ASIC design was limited to the AMS 0.35 µm technology package and the TSMC 0.18 µm technology package (the ones available at the International Iberian
4.ASIC-based TDC Development 126 Nanotechnology Laboratory). The development package selected was the TSMC 0.18 µm, since it was expected that the lower node technology would result in higher performances to be achieved on synthesizable TDCs. There are three different available libraries with different routing resources namely, 4-, 5-, and 6-metal layers, and multiple combinational and sequential logic cells. In the adopted TSMC package, the core library cells are always designed using metal layer 1. The 6-metal layers library was used to implement the proposed ASIC-based TDC. The technology package also includes Engineering Change Order (ECO) cells (designed using metal layers 1 and 2) to enable minor changes to the design after tape-out. According to TSMC documentation, the digital standard cell library is compatible with a vast set of design tools. The development package is described in multiple formats, like the .db and .lib format required by Synopsys Design Compiler, and LEF files, required by layout tools. Moreover, these libraries are described in multiple timing models (like Non-Linear Delay Model - NLDMand Composite Current Source Model - CCS), which allow the user to reach a compromise between the tool’s run time and static timing analysis precision. Each timing model can be described based on the best, worst, or typical case scenarios. A timing model is composed by tree models: the driver; the receiver; and the net. Driver and receiver models are characterized using a circuit simulation, like SPICE simulator LTSPICE. The wire model can be extracted from the layout using the various metal, via and contact parameters (among others), or it can be estimated. Only NLDM and CCS timing models are available in TSMC 0.18 µm technology package used. Thus, the other timing models will not be described in this Thesis. NLDM estimates cells’ delays and transition times for the driver model based on six points (three for the input and three for the output). These points are: in/out slope lower threshold; in/out slope higher threshold; and in/out delay threshold. The receiver model is characterized using a single capacitor (load). For technology nodes above 65 nm, this model suffices for proper static timing analysis. However, when targeting 65 nm technology nodes or lower, the 3-point schema used by NLDM is not sufficient to properly reflect the circuits’ non-linearity during static timing analysis. Furthermore, the miller effect on the receiver side, which in small impedance nets dominates the delay calculation, is not captured by NLDM. CCS models provide better accuracy when the net impedance is high, when compared to the driver resistance, since it models the driver as a current source. Regarding the receiver model, CCS is very similar to NLDM. However, the capacitance is divided in two, giving the model more granularity and enabling it to account for the miller’s effect. Since 0.18 µm technology node is being used, the NDLM models were used in this Thesis.
Readout Circuit for Time-Based Automotive Sensors 127 During synthesis only the .db files are required for static timing analysis, since placement information is still not available, and routings are considered ideal. However, during layout, in order to perform proper time closure, there are additional files that must be provided. These files comprise: .lef (which contains the physical characteristics of the library being used); worst and best cases timing model libraries, for multi-corner multi-mode (MCMM) analysis; capacitance tables ( .captable ), used to model the interconnect parasitics of the design. Ideally, a worstand best-case capacitance tables should be provided to the layout tool. However, the TSMC kit only provides the typical-case capacitance table file. Therefore, the same file had to be used when creating the time analysis corners, used during layout. 4.1.3. Design flow When implementing digital systems, the traditional design flow can usually be automated by the design tools. As can be concluded by the aforementioned description of the design tools, the focus is on area, power and timing optimizations, which more often than not lead to changes on the generated Netlist when compared to the inputted RTL design. Although this might be advantageous to several designs, there are scenarios where it can lead to erroneous circuit behavior. Thus, when implementing a synthesizable TDC in ASIC, some changes must be performed to the typical design flow (see Figure 4.1). Typical HDL designs are technology independent, however this is not the case of a HDL TDC design. Thus, a previous study of the technology to be used is required to select the logic elements that will be selected to build the fine stage measurement. Afterwards, depending on the TDC architecture, a set of additional constraints must be loaded during synthesis to avoid the optimization of certain parts of the design (specifically if delay lines are used). These optimization constraints must also be included in the synthesis exported constraints file. As verified during the implementation of the Gray-code architecture in FPGA, routing has great influence on the TDC linearity. Therefore, the placement and routing of the design during layout phase must also be constrained, in contrast to the traditional time-driven placement and routing, implemented by the IC design tools. The remaining of the design flow is similar to the typical one. Figure 4.1 presents the digital design flow adopted during the migration of the FPGA-based TDC architectures. The additional synthesis’ and layout’s constraints used are explained in detail in Sections 4.3 and 4.4.
4.ASIC-based TDC Development 128 Figure 4.1 - Adopted Design Flow 4.2. TDC architecture migration A preliminary study aiming to understand which one of the implemented FPGA-based TDC architectures would attain better performance when migrated to ASIC was made. This study was based on the obtained synthesized Netlists and the information available on TSMC 0.18 µm technology datasheet for the typical operation scenario. During this analysis, routing and placement was treated as ideal, i.e., not influencing the steps’ propagation delays. 4.2.1. TDL preliminary results The FPGA-based TDL TDC was implemented in a Hardware independent language, however it has a direct reference to a technology dependent cell responsible for creating the delay line. Thus, this HDL code must be changed to instantiate a cell available on the technology library being used. The TSMC digital standard cells library was analyzed to select which cell should be used in the TDL implementation. Since Synthesis Place & Route Signoff SDC TCL Script SDP TCL Script SDC GDSII Post-Synthesis Functional Simulation Post-Layout Timing Simulation Pre-Synthesis Behavioral Simulation RTL Description SDF Manually CreatedAutomatically Generated Added Steps Optimization constraints (SDC) Technology Dependent Verilog Code RTL Description Cadence’s Virtuoso CAD Tools Cadence’s NcVerilog Cadence’s Innovus Cadence’s NcVerilog Cadence’s NcVerilog Verilog Code Synopsys’ DesignCompiler
Readout Circuit for Time-Based Automotive Sensors 129 the same TDL is propagating the start and the stop signals, ideally the cell should present a similar lowto-high and high-to-low transition time from input to output. This type of characteristic is common on cells used to implement clock trees. Therefore, the set of clock buffers, inverters, and gates were analyzed. The cell that theoretically would offer better resolution would be a clock inverter. However, since the developed decoders expect a thermometer code with a sequence of 1s followed by a sequence of 0s (or vice-versa), if inverters were to be used, each TDL step would have to be comprised by two inverters, otherwise the decoder blocks would have to be changed. Another solution would be to use a clock AND gate with short-circuited inputs. Nevertheless, both solutions have higher propagation delay than the one obtained when a single clock buffer per step is used (see Table 4.1). So, the clock buffer with lower propagation delay and sufficient fan-out to supply the sampling flip-flop and the next step clock buffer was selected, since it was able to comply with the resolution requirements established for this Thesis application. According to the TSMC typical case datasheet, the CKBD0BWP7T has the lowest low-to-high and high-to-low propagation delays and fulfils the fan-out requirement. A comparison between the generate block used to implement the delay line in FPGA and the ASIC one is presented in Figure 4.2. Table 4.1 - TSMC clock digital cells propagation delay analysis Digital Cell Connection Schema Total Fanout Capacitance (pF) High-to-Low Propagation Time (ps) Low-to-high Propagation Time (ps) Inverters (CKND0BWP7T) 0.007966 (one stage) 0.003597 + 0.007966 (two stages) 62.4 110.3 68.8 121.9 AND gates (CKAN2D0BWP7T) 0.009842 141.8 126.8 Buffers (CKBD0BWP7T) 0.004137+0.004369 0.008506 107.1 105.1 Q Q S ET CLR D Q Q S ET CLR D Q Q S ET CLR D Q Q S ET CLR D
4.ASIC-based TDC Development 130 Figure 4.2FPGA vs ASIC HDL Comparison The modified code was synthesized, avoiding the delay line optimization, according to the process described in Section 4.3. Analyzing the fine measurement module in detail, it is possible to verify the correct delay line generation (see Figure 4.3). Accordingly, a good estimation for the expected TDC performance can be obtained from the TSMC datasheet, using equation (4.1) and (4.2). Figure 4.3TDL Synthesized Netlist Overview
Readout Circuit for Time-Based Automotive Sensors 131 𝑡𝑃𝐿𝐻 =0.0743(𝑛𝑠)+3.6185(𝑛𝑠/𝑝𝐹)∗𝐶𝑙𝑜𝑎𝑑(𝑝𝐹), (4.1) 𝑡𝑃𝐻𝐿 =0.0785(𝑛𝑠)+3.3704(𝑛𝑠/𝑝𝐹)∗𝐶𝑙𝑜𝑎𝑑(𝑝𝐹), (4.2) where tPLH and tPHL are the low-to-high and high-to-low propagation delays (in nanoseconds), respectively, and Cload is the total load being driven by the cell (in picofarad). Considering that the clock buffer used has an input capacitance equal to 0.004137 pF and that the generated flip-flop D pin has a 0.004369 pF capacitance, the typical propagation delay of each step is approximately 105 ps and 107 ps for low-to-high and high-to-low transitions, respectively. These results comply with the defined LiDAR application’s requirements, making the migration of this architecture feasible. The migration process can also be fully automated, requiring only a change in the instantiation of the cell used to build the delay line, if a different technology is desired. Regarding linearity, it is known that longer delay chains tend to have its performance degraded. This is mainly caused by process mismatch (something that cannot be controlled by design), temperature and voltage variations (which can be minimized by layout). Thus, if this architecture is to be ported to ASIC, the layout will have a significant impact on the system’s linearity. The layout consideration will be further discussed in Section 4.4. 4.2.2. Gray-code TDC architecture preliminary results The Gray-code TDC architecture does not need any Verilog change before being synthetized by the ASIC tools. It also does not require any special attention regarding optimization constraints when being synthesized. However, contrarily to what happens in FPGA-based implementations, where each Gray-code bit was generated by a single LUT, the resultant Netlist in ASIC is composed by multiple logic gates, arranged in a combinatorial loop, being the different Gray-code bits extracted at different points of the circuit (see Figure 4.4). Thus, while in FPGA every LUT has the same propagation delay independently of the truth table implemented, in ASIC development the various paths of the Gray-code combinatorial logic Netlist must be analyzed.
4.ASIC-based TDC Development 132 Figure 4.4Gray-Code TDC Channel Synthesized Netlist Overview When analyzing the Netlist to determine the estimated propagation delay of each step, it must be considered that, at any time, from one step to the next one, only paths that does not change multiple Gray-code bits must be considered as valid. Using the detailed Netlist schematic of the start channel fine measurement module, presented on Figure 4.4, and the transition from the code 00001 to 00011 as an example, only the paths “AOI221D0-INVD0-OAI31D2” and “AOI221D0-INVD0-ND3D1” are valid. The other ones like, for example, “AOI221D0-NR2D0-CKND2D1-IAO22D1-NR2D0-XOR2D0-MAOI22D0OAI31D2”, does not corresponds to a valid system behavior, since it would generate multiple bit changes on the outputted Gray-code. For brevity, the equations used to calculate the propagation delay of each logic cell (depicted in grey color at the top of each cell in Figure 4.4) are not presented. However, the process is analogous to the one used in the TDL architecture. First, the input and output capacitance for each cell pin is annotated. Then, for each cell, the output load is calculated and the value is used in the propagation delay equation of the cell (obtained from the TSMC standard digital library datasheet). With the propagation delay of each cell calculated (using the typical-case datasheet values), the Netlist paths are analyzed for each step, and the propagation delay of the cells contained in the path are added. It is important to notice that the time of a step is equal to the propagation delay of the bit that changes to the following bit to change, i.e., considering the sequence 00001-00011-00010, the step size of the code 00011 will be equal to the time required for the 4th bit (the one that changed from 00001 to 00011) to generate a change in the 5th bit (the one that changes from 00011 to 00010). Table 4.2 presents all the Gray-code steps delays and corresponding path according to the presented Netlist. CKND2D1 AO211D1 OAI31D2 INVD0 OAI21D2 IAO22D1 A1 A2 B1 B2 C Zn EN U1 INVD0 U2 NR2D0 U3 gray_w[3] U4 NR2D0 XOR2D0 U5 U6 MAOI22D0 gray_w[2] gray_w[0] U7 U8 U9 AOI221D0 U11 U10 INVD0 U12 NR2D0 U17 U16 ND2D0 U15 INVD0 U18 INVD0 U13 ND3D1 U14 gray_w[1] AOI22D2 U19 U20 INVD0 gray_w[4] 177ps 313ps 123ps 360ps 264ps 122ps 312ps 383ps 62ps 155ps 234ps 154ps 59ps 131ps 81ps 277ps 110ps 94ps 114ps 58ps
Readout Circuit for Time-Based Automotive Sensors 133 Since the goal is not to obtain a detailed characterization of the TDC channel, but rather an overview of the performance and architecture limitations, although some steps change at the high-to-low transition, the delay calculations were done considering the propagation delay of a low-to-high transition. Table 4.2Gray-Code Path Propagation Delay Analysis Gray-code Step Path Propagation Delay (ps) 0 U1-U9 560 1/9/17/25/5/13/21/29 U10-U11 411 2/10/18/26 U9 383 3/19 U10-U14-U6 554 4/12/20/28 U7-U8-U9 757 6/14/22/30 U20-U9 441 7/23 U12-U17-U3-U4 718 8/24/16 U5-U7-U8-U9 1021 11/27 U10-U13-U6 387 15 U12-U16-U18-U19-U2 806 31 U12-U16-U19-U2 712 The analysis of the results presented on Table 4.2 demonstrate that eleven different paths are responsible for Gray-code changes (three more than in the FPGA case). Moreover, there is a considerable step delay variation, with a maximum variation equal to 638 ps. This value is even higher than the worst case scenario obtained in the FPGA simulation, which already included routings. Another important aspect to highlight is the complexity of the routing that will be required to be managed during the layout step. There are multiple insertion points per step and no routing pattern (when compared to the TDL case where the routing from one step to the next one followed a well-defined pattern). In conclusion, although performance loss was expected due to the process technology in which the TDC is being fabricated (the same happened in the TDL architecture), the migration of the Gray-code TDC architecture to ASIC introduces multiple other issues that do not exist in FPGA-based designs. The following section compares the preliminary results from the TDL and Gray-code architectures. 4.2.3. Discussion When analyzing the preliminary results of the implemented TDCs, it is clear that the performance of the Gray-code architecture is deteriorated when migrated to ASIC. This is mainly due to the dissimilarity of the combination logic for each Gray-code bit. While in FPGA-based implementation, a LUT cell with a fixed propagation delay (regardless of the implemented truth table) is available, in ASIC-based implementation logic cells must be used which, depending on the logic equation to implement, may have different levels