scieee AI-readable full text Open interactive document viewer

Reliability and availability in reconfigurable computing: A basis for a common solution

Manuel G. Gericota,Gustavo R. Alves,Miguel L. Silva,José M. Ferreira

Abstract

Dynamically reconfigurable SRAM-based field-programmable gate arrays (FPGAs) enable the implementation of reconfigurable computing systems where several applications may be run simultaneously, sharing the available resources according to their own immediate functional requirements. To exclude malfunctioning due to faulty elements, the reliability of all FPGA resources must be guaranteed. Since resource allocation takes place asynchronously, an online structural test scheme is the only way of ensuring reliable system operation. On the other hand, this test scheme should not disturb the operation of the circuit, otherwise availability would be compromised. System performance is also influenced by the efficiency of the management strategies that must be able to dynamically allocate enough resources when requested by each application. As those resources are allocated and later released, many small free resource blocks are created, which are left unused due to performance and routing restrictions. To avoid wasting logic resources, the FPGA logic space must be defragmented regularly. This paper presents a non-intrusive active replication procedure that supports the proposed test methodology and the implementation of defragmentation strategies, assuring both the availability of resources and their perfect working condition, without disturbing system operation.

Full text

IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS, VOL. 16, NO. 11, NOVEMBER 2008 1545 Reliability and Availability in Reconfigurable Computing: A Basis for a Common Solution Manuel G. Gericota, Member, IEEE, Gustavo R. Alves, Miguel L. Silva, and José M. Ferreira Abstract—Dynamically reconfigurable SRAM-based field-programmable gate arrays (FPGAs) enable the implementation of reconfigurable computing systems where several applications may be run simultaneously, sharing the available resources according to their own immediate functional requirements. To exclude malfunctioning due to faulty elements, the reliability of all FPGA resources must be guaranteed. Since resource allocation takes place asynchronously, an online structural test scheme is the only way of ensuring reliable system operation. On the other hand, this test scheme should not disturb the operation of the circuit, otherwise availability would be compromised. System performance is also influenced by the efficiency of the management strategies that must be able to dynamically allocate enough resources when requested by each application. As those resources are allocated and later released, many small free resource blocks are created, which are left unused due to performance and routing restrictions. To avoid wasting logic resources, the FPGA logic space must be defragmented regularly. This paper presents a non-intrusive active replication procedure that supports the proposed test methodology and the implementation of defragmentation strategies, assuring both the availability of resources and their perfect working condition, without disturbing system operation. Index Terms—Active replication, availability, field-programmable gate array (FPGA), online structural testing, reliability. I. INTRODUCTION RECONFIGURABLE logic devices, namely field-programmable gate arrays (FPGAs), experienced a considerable expansion in the last few years due in part to an increase in their size and complexity, with advantages in terms of board space and flexibility. The availability of SRAM-based FPGAs supporting fast runtime partial reconfiguration (e.g., the Virtex family from Xilinx used to validate this work, the only family of FPGAs that supports dynamic reconfiguration so far) considerably reinforced these advantages, wide-spreading their usage as a base for reconfigurable computing platforms. Manuscript received May 10, 2006; revised October 10, 2007. Current version published October 22, 2008. This work was supported by an FCT Program under Contract POSC/EEA-ESE/55680/2004. M. G. Gericota and G. R. Alves are with the Department of Electrical Engineering, Polytechnic Institute of Porto, 4200-072 Porto, Portugal (e-mail: [email protected]; [email protected]). M. L. Silva is with the Department of Electrical and Computer Engineering, Faculty of Engineering, University of Porto, 4200-465 Porto, Portugal (e-mail: [email protected]). J. M. Ferreira is with the Department of Electrical and Computer Engineering, Faculty of Engineering, University of Porto, 4200-465 Porto, Portugal, and also with the Buskerud College of Engineering, 3603 Kongsberg, Norway (e-mail: [email protected]). Digital Object Identifier 10.1109/TVLSI.2008.2001141 Dynamically reconfigurable FPGAs enable the implementation of virtual hardware as defined in [1] in the beginning of the 1990’s, by using temporal partitioning to implement those applications whose area requirements exceed the available logic space (as if there were unlimited hardware resources). This approach is viable because each application comprises a set of functions, predominantly executed sequentially or with a low degree of parallelism, making simultaneous availability hardly ever required. The static implementation of an application is separated in two or more independent hardware contexts, which may be swapped during runtime [2]. Extensive work was done to improve the multi-context handling capability of these devices, by storing several configurations and enabling quick context switching [3], [4]. The main goal was to improve the execution time by minimizing external memory transfers, assuming that some amount of on-chip data storage was available in the reconfigurable architecture. However, this solution was only feasible if the functions implemented on hardware were mutually exclusive on the temporal domain, e.g., context-switching between coding/decoding schemes in communication, video or audio systems; otherwise, the length of the reconfiguration intervals would lead to unacceptable delays in most applications. The reduction of manufacturing scales contributed significantly to eliminate these restrictions, by enabling higher levels of integration and higher frequencies of operation. The increasing amount of logic available in FPGAs and the smaller reconfiguration times, partly owing to the possibility of partial reconfiguration [5], extended the concept of virtual hardware to the implementation of multiple applications sharing the same logic resources in the spatial and temporal domains. However, if the functions required by the different applications cannot be scheduled in advance, all resource allocation decisions will have to be made at runtime, preventing a long term resource allocation strategy. In this case, fragmentation of the logic space is inevitable, leading to poor resource utilization. Since the type and amount of resources required by different functions varies greatly, as resources are allocated to functions and later released, many “islands” of free resources are created. These areas tend to become so small that they are left unused due to performance and routing restrictions. To avoid wasting resources and degrading system availability, the defragmentation of the FPGA logic space must be performed systematically. Smaller submicrometer scales also have disadvantages, such as the higher electronic current densities in metal traces—which increase the threat of electromigration—and the lower threshold voltages. As scale goes down the number of defects related to small manufacturing imperfections that are not detected by production testing goes up. These defects are especially prone to 1063-8210/$25.00 © 2008 IEEE Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. 1546 IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS, VOL. 16, NO. 11, NOVEMBER 2008 electromigration phenomena and lead to permanent faults after long operation periods [6]. On the other hand, the exponential growth in the number of configuration memory cells, and their lower threshold voltages, make these components more susceptible to gamma particle radiation and to the appearance of transient faults, such as single event upsets (SEU) and multi-bit upsets (MBU) [7]–[9]. These faults do not physically damage the chip, but their effects are permanent, since the functionality of the circuits mapped into the device is modified [10], [11]. Yet, and since the cause of the failure is actually transient, online reconfiguration is sufficient to restore the original functionality. In the case of permanent faults, and after the faulty elements are located—either configurable logic blocks (CLBs) or routing resources—they must be excluded and replaced by previously unused fault-free resources. A non-intrusive online concurrent test strategy, performing a structural test of all FPGA resources is required to detect and diagnose any emerging permanent fault. This solution involves the periodic release of all resources from their active functions for the system to be able to perform their structural test. Previous considerations about the advantages and disadvantages of new FPGA features show that to increase the availability and reliability of reconfigurable computing systems, transparent onlineFPGAlogic spacedefragmentation andtransparentonline test operations must be run simultaneously [12], [13]. Both procedures require resource reallocation strategies that do not interfere with any currently running functions. This article presents a non-intrusive procedure for concurrent replication of active logic blocks and resource interconnections (i.e., logic resources and interconnections that are currently being used to implement running functions from one or more applications). This procedure is able to support the implementation of the proposed structural test methodology and also to serve as a basis for the implementation of defragmentation strategies. This paper is organized as follows. Section II presents a survey of FPGA test strategies proposed in the literature. A detailed account of the proposed procedure for the active replication of logic resources is presented on Section III, while Section IV describes the online structural concurrent test methodology. A brief introduction to the defragmentation issue and an overview of how the implementation of defragmentation strategies may benefit from the proposed active replication procedure are provided in Section V. Section VI presents the software tool developed to support the active replication procedure and the test methodology. Finally, conclusions are drawn in Section VII. II. TESTING DYNAMICALLY RECONFIGURABLE FPGAS To achieve higher reliability in reconfigurable computing systems, the structure of all FPGA resources has to be continuously tested and error correction/fault tolerance has to be introduced. These requirements will ensure that any function will perform correctly independently of its type. For SRAM-based FPGAs, they translate into the following features: 1) to be non-intrusive; 2) to be able to detect any permanent structural fault emerging during system lifetime; 3) to be able to correct transient faults affecting function functionality. Several offline and online strategies have been proposed to test and diagnose FPGA faults. An offline built-in self-test (BIST) technique that uses reprogrammability to set up the BIST logic is presented in [14]–[16]. Some of the logic blocks are configured as test pattern generators or response analyzers, while testing the other blocks, and vice versa. Since the test sequences are a function of the FPGA architecture and independent of its functionality, this approach is applicable at all levels (wafer, packaged device, board, and system). This technique requires a fixed number of reconfiguration sessions and presents no area overhead or performance penalty, since the BIST logic is eliminated when the circuit is reconfigured for normal operation. A slightly different BIST technique, involving a structural modification of the original configuration memory, is proposed in [17]. This technique enables the automation of the test process while reducing test time and off-chip memory. However, the modification required to the FPGA hardware is a major disadvantage, implying the non-universality of the solution. An offline test methodology based on a non-BIST approach, targeting the FPGA CLBs, is presented in [18] and [19]. After setting up a specific test configuration, the FPGA input/output blocks (IOBs) are used to support the external application of test vectors and to capture the test responses. In order to achieve 100% fault coverage at CLB level, different test configurations must be programmed and specific sets of test vectors applied in each case. Based on the same principles, a fault diagnosis method is presented in [20]. Extensive work on the structural testing of FPGA lookup tables (LUT) and interconnections is also presented in [21] and [22]. The previous approaches require the device to be offline, increasing fault-detection latency, and as such are not admissible in highly fault-sensitive, mission-critical applications. In order to overcome these limitations, online testing and diagnosis methods based on a scanning strategy were presented in [6] and [12]. The idea underlying these methods is to have a relatively small portion of the chip being tested offline (instead of the whole chip, as considered in previous proposals), while the rest continues its normal online operation. Testing is accomplished by sweeping the test functions across the entire FPGA. If the functionality of a small number of FPGA elements can be relocated on another portion of the device, then those elements may be taken offline and tested in a completely transparent way. This fault scanning procedure moves on to copy and test another set of elements, systematically testing the whole FPGA. However, the replication procedure of the first approach [6] requires a modified cell structured, while the second approach [12] halts the whole system to relocate an entire CLB column. Since reconfiguration is performed through the boundary scan (BS) infrastructure (IEEE 1149.1 Standard) [23], reconfiguration time is long, and it seems likely that repeatedly halting the system will severely disturb its operation. The design for test features proposed in [24] are essentially concerned with fault detection, instead of carrying out any structural or functional test functions (the FPGA logic structure is not taken into account). Their main goal consists of detecting the presence of faults in the current application, and therefore, a physical defect may escape detection if that particular application is not using the damaged resource. Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. GERICOTA et al.: RELIABILITY AND AVAILABILITY IN RECONFIGURABLE COMPUTING 1547 A new application-oriented method that generates a functional test for the configured circuit, taking into account the logic structure of the FPGA where it is implemented, was proposed in [25]. However, this method corresponds to an offline field-oriented test to be used with a given application, thus presenting the same drawbacks of the previous method. The online test approach proposed in this article reuses some of the previous ideas, but eliminates their disadvantages by using a novel concept herein referred as active replication, which enables the functionality of a given set of resources to be relocated without halting the system. This approach is feasible even when the resources are active, i.e., when they are being used by a function that is currently running [26], [27]. Conceptually, an FPGA may be visualized as an array of uncommitted CLBs, surrounded by a periphery of IOBs, which are interconnectable by configurable routing resources. A set of memory cells that lies beneath controls the configuration of the whole structure. Complete (100%) usage of the FPGA resources is hardly ever achieved, even when independent hardware blocks, from different applications, dynamically share the same device. The dynamic and partially reconfigurable features that are offered by some FPGAs make it possible to test all free CLBs and interconnection resources, without disturbing system operation. After being tested, defect-free CLBs and interconnection resources remain available as spare blocks, ready to replace others that are found defective. The CLBs and routing resources currently being used are released for test following a dynamic rotation mechanism, after having their current functionality relocated into other areas already tested. That dynamic rotation mechanism ensures that all FPGA resources are released and tested within a given latency. The active replication of the FPGA resources is therefore at the core of this proposed non-intrusive online structural test approach, which can be carried out concurrently with system operation. Since all FPGA resources are released and tested using the BS test infrastructure, there is no overhead at board level. Being application-independent, and oriented to test the FPGA structure, the proposed strategy guarantees FPGA reliability after many reconfigurations, and helps to ensure correct operation throughout the system lifetime. III. CONCURRENT REPLICATION OF ACTIVE LOGIC BLOCKS The replication of CLBs and interconnections is required to release any active resources for testing. However, it is not trivial to do it non-intrusively due to two major issues: 1) configuration memory organization and 2) internal state information. The configuration memory may be visualized as a rectangular array of bits, which are grouped into one-bit wide vertical frames, extending from the top to the bottom of the array. The atomic unit of configuration is one frame—it is the smallest portion of the configuration memory that can be written to or read from. These frames are grouped together into larger units called columns. Each CLB column has an associated configuration column, with multiple frames, which mixes internal CLB configuration and state information, and column routing and interconnecting information. The organization of the entire conFig. 1. Two-phase CLB replication process. figuration memory into frames enables the online concurrent partial reconfiguration of the FPGA. The configuration process is a sequential mechanism that spans through some (or eventually all) CLB configuration columns. More than one column may be affected during the replication of an active CLB, since its input and output signals (as well as those in its replica) may cross several columns before reaching its source or destination. Any partial reconfiguration procedure must ensure that the signals from the replicated CLB are not broken before being totally reestablished from its replica. It is also important to ensure that the functionality of the CLB replica is perfectly stable before its outputs are connected to the system, to avoid output glitches. A set of experiments performed with Virtex FPGAs from Xilinx demonstrated that the replication process has to be divided into two phases, as illustrated in Fig. 1. In the first phase, the internal configuration of the CLB is copied and the inputs of both CLBs are placed in parallel. Due to the low-speed characteristics of the configuration interface used (the BS infrastructure), the reconfiguration time is relatively long when compared to the system speed. Therefore, the outputs of the CLB replica will be perfectly stable before being connected to the circuit, in the second phase. Both CLBs must remain in parallel for at least one system clock cycle, to avoid output glitches. Notice that rewriting the same configuration data does not generate any transient signals. Therefore, the remaining resources covered during this process by the rewritten configuration frames are not affected, even if in an active state. The correct transference of state information is another major requirement for the success of the replication process. If the current CLB function is purely combinational, a simple readmodify-write procedure will suffice to accomplish a successful replication. However, in the case of a sequential function, the internal state information must be preserved and no write-operations may be lost while this process goes on. In the Virtex FPGA family, each CLB slice comprises two storage elements, which can be individually configured as a latch or flip-flop (FF). Although a read back operation of the configuration memory may be performed to read the value of a storage element, it is not possible to perform a direct write operation. In addition, when dealing with active CLBs during a replication procedure, if state information changes between read and write operations, a coherency problem will occur. For this reason, no time gap is allowed between the two operations. The solution to this problem depends on the type of implementation. The following three cases are considered: 1) synchronous free-running clock circuits; 2) synchronous gated-clock circuits, and; Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. 1548 IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS, VOL. 16, NO. 11, NOVEMBER 2008 Fig. 2. Implementation of the synchronous gated-clock FF replication scheme. 3) asynchronous circuits. When dealing with synchronous free-running clock circuits, the two-phase replication process that was previously described solves the state transfer problem. Between the first and the second phase, the CLB replica has the same inputs as the replicated CLB, and all its storage elements acquire the state information, even if the system clock frequency is an order of magnitude lower than the clock frequency (of the BS infrastructure) used for reconfiguration purposes. Several experiments were carried out and showed the effectiveness of this method to replicate active CLBs. No loss of state information and no output glitches were observed. Notice that this procedure is valid even when dealing with asynchronous circuits. If the longest interval between consecutive update operations of asynchronous latches is lower than the interval between the first and the second phases, the replicated latch always acquires the correct state information. Despite the effectiveness of this solution, its usefulness is very restricted. A broad range of applications use synchronous gated-clock circuits, where input acquisition is controlled by a clock enable signal. In such cases, it is not possible to ensure that this signal will be active during the replication process, and that the value at the input of the replica FFs will be captured. On the other hand, it is not feasible to set this signal as part of the replication process because the value present at the input of the replica FFs may differ from the one captured by the replicated FFs, resulting in a coherency problem. Furthermore, the FFs could be updated during the replication process, since this procedure is asynchronous in relation to system operation. A replication aid block is used to solve this problem. This block manages the transfer of state information from the replicated FFs to their replicas. State information may also be updated by the circuit at any moment, without losing information or delaying the replication process. The replication scheme is represented in Fig. 2 for a single CLB logic cell (for this purpose each CLB logic cell in the Virtex FPGA family can be considered individually). Fig. 3 represents the flow diagram of the replication process. One input of the 2:1 multiplexer in the replication aid block is connected to one temporary transfer path from the output of the replicated FF (FF OUT). The other one is connected to the output of the combinational logic block in the replica cell (LOGIC OUT), which is normally applied to the input of the FF. If the clock enable (CE) signal—controlling the Fig. 3. Replication process flow. multiplexer—is not active, the output of the replicated FF (FF OUT) is applied to the input of the replica FF. A clock enable signal, coming from the replication aid block (capture control signal—CC), forces the replica FF to store the transferred value. If the CE signal is active or is activated during this process, the multiplexer selects the LOGIC OUT signal and applies it to the input of the replica FF. This FF is, therefore, updated simultaneously with the replicated FF, and captures the same value, guaranteeing state coherency. Neither simulations, nor the ensuing practical experiments, have shown any loss of information. The control signals CC and BY C are driven by configuration memory bits. BY C directs the state signal to the input of the replica FF, while CC enables its acquisition. It is, therefore, possible to control the whole replication process through the BS infrastructure, and as such no extra pins are required. Fig. 4 Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. GERICOTA et al.: RELIABILITY AND AVAILABILITY IN RECONFIGURABLE COMPUTING 1549 Fig. 4. Simplified representation of the replication aid block implemented on a CLB slice. shows a schematic implementation of the replication aid block in a CLB slice. To enable all signals to be controlled through the configuration memory, the CC net includes the FF shown in Figs. 2 and 4. However, it is there simply as a consequence of the structure of the CLB slice, and does not play any role in this process. After the transference of state information BY C is driven low, disconnecting the replica FF from the replication aid block. The state transfer ends and the replica FF may now be directly updated by the circuit. The CE signal of both CLBs is placed in parallel, all the signals to/from the replication aid block are disconnected, and the outputs are also placed in parallel. After at least one clock cycle, the replicated block is disconnected, and the resources used in its implementation are released. Each of these steps (corresponding to a rectangle in the flow diagram shown in Fig. 3) requires a new reconfiguration file. A total of nine files are therefore needed to complete the replication process, instead of four, as would be necessary when dealing with synchronous free-running clock circuits. However, in most cases, their size is much smaller (to change the value of CC and BY C only one configuration frame is needed). Table I details the average sizes of the partial reconfiguration files and their respective reconfiguration times, when using a 20-MHz test clock [the TestClock (TCK) signal of the BS infrastructure] [23]. Replication of synchronous free-running clock circuits takes roughly 18 ms, as steps 2–6 are not necessary. Practical experiments performed using a Virtex device to implement the ITC’99 Benchmark Circuits from the Politecnico di Torino [28], demonstrated the effectiveness of the proposed approach. These circuits are purely synchronous with only one single-phase clock. However, the procedures presented are also applicable to multiple clock/multiple phase circuits, since only one clock signal is involved in the replication process at a time. Still, the slowest “clock” period must be shorter than the duration of the replication process, thus enabling the FFs to be updated meanwhile. The proposed method is also effective when dealing with asynchronous circuits, where storage elements are configured TABLE I COST OF EACH PARTIAL RECONFIGURATION FILE DURING REPLICATION Fig. 5. Relocation of routing resources. Fig. 6. Propagation delay during the relocation of routing resources. as latches instead of FFs. In this case, the CE signal is replaced by an input control signal. Data present in the D input is stored in the gated D latch when the control input signal changes from “1” to “0”. The same replication aid block and the same replication sequence are used. The register present in the replication aid block may be configured as a latch, instead of as a FF, if this is preferred or if no adequate clock signal is available. The replication of routing resources does not pose any special problems, since the same two-phase replication procedure is also effective to relocate local and global interconnections. The interconnections involved are first duplicated in order to establish an alternative path, and then disconnected, becoming available to be reused, as illustrated in Fig. 5. A last remark must be made about the replication of routing resources. The different paths used while paralleling the original and replica interconnections will likely have different propagation delays. This means that if the logic level at the output of the source CLB changes, there will be an interval of fuzziness at the input of the destination CLB, as shown in Fig. 6. However, the impedance of the routing switches will limit the current flow in Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. 1550 IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS, VOL. 16, NO. 11, NOVEMBER 2008 the interconnection, and hence this behavior does not damage the FPGA. Nevertheless, and for transient analysis, the propagation delay associated to parallel interconnections shall be the longer of the two paths [29]. The LUTs in the CLB can also be configured as memory modules (RAMs) for user applications. However, the extension of this concept to the replication of LUT/RAMs is not feasible. The content of the LUT/RAMs may be read and written through the configuration memory, but there is no mechanism, other than to stop the system, capable of ensuring the coherency of the values if there is a write attempt during the replication interval [30]. Furthermore, since frames span an entire column of CLB slices, a given bit in all slices is written with the same command. Therefore, it is necessary to ensure that all the remaining data in the slice is constant, or else it must also be changed externally through partial reconfiguration. Even if not being replicated, LUT/RAMs should not lie in any column that could be affected by the replication process. According to the overall test (and/or defragmentation) strategy, this method could be used to replicate more than one CLB simultaneously, improving scalability aspects. Considering that the smallest configuration unity is a frame, and that frames span the FPGA from top to bottom, it takes exactly the same time to replicate one CLB or the whole CLB column (the number of reconfiguration frames involved is the same in both situations). Scalability is therefore not an issue. The time required increases proportionally to the number of CLB columns in the FPGA and is independent of the number of CLB rows. IV. ONLINE STRUCTURAL CONCURRENT TEST Our online structural concurrent test method is divided in the following three parts: 1) the replication procedure; 2) the test strategy; 3) the dynamic rotation mechanism. The replication procedure has already been presented. A detailed presentation of the proposed test strategy and of the dynamic rotation mechanism used to release resources for test will follow. A. Fault Detection and Error Recovery The replication procedure used with synchronous free-running clock circuits did not perform a true state transfer operation, but rather an acquisition of the values present at the inputs of the replica CLB FFs. For this reason, the acquired state information is correct, despite any permanent or transient fault that may affect the content of the replicated CLB FFs. As a consequence, and after the replication process, the outputs of the CLB replica always display the correct values, automatically correcting any faulty behavior. On the other hand, when replicating synchronous gated-clock circuits (or asynchronous circuits), a true state transfer operation is performed. In this case, the replica CLB FFs (or latches) will acquire exactly the same value held by the replicated FFs (or latches). Erroneous state information may therefore be propagated to the replica CLB, and will survive until an update occurs. A permanent fault in the replicated CLB will be detected during the subsequent test phase and the CLB will be flagged as defective, meaning that it will not be used again in a later reconfiguration. Depending on the method used to create the reconfiguration files, the replication procedure can also recover from errors caused by transient faults in the on-chip configuration memory. Typical examples of such errors are SEUs, which modify the logic function originally implemented in the FPGA. Until now, they used to be a major concern only for space applications. Yet for designs manufactured at advanced technology nodes—such as 90, 65 nm, and downward—system-level soft errors become an issue also at ground level. They are now much more frequent than in previous generations [31]. Since Virtex FPGAs enable read back operations, a completely automatic read-modify-write procedure may be implemented to replicate the CLBs using local processing resources. In this case, any transient fault in the configuration memory is propagated and will affect the functionality of the CLB replica. On the other hand, if the reconfiguration files are generated from the initial configuration file stored in an external memory, any error due to SEUs is corrected when the affected blocks are replicated. B. Interconnection Resources and I/O Blocks Successful structural testing of the CLB replica ensures its good functionality, but the replicated CLB may be faulty. When the inputs and outputs of both CLBs are placed in parallel, nodes with different voltage levels may be interconnected. Due to the impedance of the routing switches, this apparent “short-circuit” behaves as a voltage divider, limiting the current flow in the interconnection. Therefore, no damage results to the FPGA, as proven by extensive experimental essays. Since we are dealing with digital circuits, the analog value resulting from the voltage divider leads to a well defined value (logic “0” or logic “1”) when it propagates through a routing buffer, or at the input of the next CLB or IOB. No logic value instability was observed in our experiments [26]. In the FPGA, signals are routed using the global routing resources, which are located in horizontal and vertical routing channels between each routing array. The routing resources may be unidirectional or bidirectional. Besides a pair of dedicated paths providing high-speed connections between vertically adjacent CLBs (to propagate carry signals), few routing resources are available to establish direct interconnections with other CLBs. As such, the majority of interconnections required by the replication process can only be done through global routing resources. To place the inputs in parallel, the interconnection segments to be used between routing arrays may be unidirectional (from the replicated CLB inputs towards the CLB replica inputs), or bidirectional. Concerning the outputs, interconnection segments between routing arrays may also be unidirectional (from the CLB replica outputs towards the replicated CLB output), or bidirectional, as illustrated in Fig. 7. Since signals do not propagate backwards, if propagation direction is not taken into account, no signals would exist at the inputs of the CLB replica, and the outputs of both CLBs would not be placed in parallel. As a result, when the outputs of the replicated CLB were disconnected, no signals would be propagated to the rest of the circuit. Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. GERICOTA et al.: RELIABILITY AND AVAILABILITY IN RECONFIGURABLE COMPUTING 1551 Fig. 7. Replication CLB interconnection. Fig. 8. Internal architecture of an IOB and associated BS cells. Local routing, at the inputs (and outputs) of the CLB, is unidirectional and therefore the logic values present at the inputs of the replica CLB will not be affected by the interconnection, even if the replicated CLB is faulty. As such, all CLB replica inputs will always reflect the correct values, because no fault at any of the replicated CLB inputs may propagate backward. This is also true when replicating active interconnections, with faults in the replicated net being automatically corrected when the replication takes place. Depending on the location of the fault in the replicated interconnection, it may be corrected while the path is duplicated, or only after it is disconnected. Any FPGA pin could be used as an input, an output, a tristate output, or a bidirectional pin. The output and tristate signals may or may not be registered. The IOB circuitry provides an FF for each of these signals and two multiplexers, controlled through the configuration memory. The input signal is available to the internal logic both in registered and nonregistered form. A generic implementation of an IOB is illustrated in Fig. 8. In spite of the configuration of each IOB or of its use (or not) to implement a system function, the number of BS cells of the BS register remains constant. All IOBs are considered as independent tristate bidirectional pins, placed in a single BS chain. For this reason, BS cells are provided on the input, output, and tristate signal paths, as required by the IEEE 1149.1 Standard [23]. Notice that, even when a bidirectional pin is used only as an input, its tristate and output BS cells are still part of the BS register, as well as the three BS cells of unused bidirectional pins. All IOBs have a pad, as seen in Fig. 8, but not all of them have an associated output pin. IOBs without a bond wire connecting the pad to a pin on the package are called unbonded IOBs. These IOBs may be used on register intensive applications or as tristate buffers in internal bus implementations, with the bus signals being returned to the internal logic through the input path. Usually, design tools offer an option that enables the user to pack registers into IOBs. Despite not being true input/outputs, these IOBs have BS cells and, therefore, are part of the BS register. Test vector application to the IOBs and response capturing should take account the following factors: 1) BS register enables controllability of the input signal path and observability of the output and tristate signal paths; 2) observability of the input signal paths and controllability of the output and tristate signal paths are not possible through the BS register; 3) observability and controllability of the control and clock signals are not possible through the BS register; 4) not all IOBs have an attached pin; therefore, external access to improve the controllability/observability of the IOB can not be assumed; however, since they all have BS cells, this limitation is not a problem. These remarks lead to the conclusion that a feasible and reliable online test of the IOBs is not possible. The observability and controllability of all the paths in the IOB implies the direct access through the external pin (if it exists), or the execution of intrusive operations through the BS register. An offline test method for the IOB structure and its interconnections at board level, which presents no area overhead or performance penalty (since the logic functionality required to implement it is eliminated when the circuit is reconfigured for its normal operation), is presented in [32]. C. Test Configurations The configurable structure of the CLB requires the use of a minimum number of test configurations to perform a complete test of its structure, with a specific set of test vectors applied to each test configuration. Since the implementation structure of the CLB primitives (LUTs, multiplexers, FFs) is not known, a hybrid fault model was considered [18] (see also [21] and [22] for an extensive study concerning FPGA fault models). To test the SRAM elements of the LUT, each bit is set to both “0” and “1”. By programming the LUTs to implement XOR and XNOR functions—which requires at least two test phases—it is easy to propagate any activated faults to a primary CLB output. Due to the XOR/XNOR functions, all LUT input stuck-at faults, together with their respective addressing faults, are also detected. For test purposes, Virtex CLB multiplexers have to be divided in two types: conventional and programmable multiplexers (those where selection lines are controlled through configuration memory bits). Since the existing maximum number of selection lines is two (in both cases), at least four test configurations are needed to exhaustively test each programmable multiplexer. The CLB structure presents a chain of three configurable primitives, which requires at least six test configurations to completely test its combinational part. Notice that the test of primitives in a chain could not take place simultaneously, because the controllability and observability of a primitive under test depends on the configuration of its immediate neighbours Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply. 1552 IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS, VOL. 16, NO. 11, NOVEMBER 2008 TABLE II NUMBER OF TEST VECTORS PER TEST PHASE TABLE III COST OF EACH PARTIAL RECONFIGURATION in the propagation path, except in the case of primitives with primary inputs and/or outputs. All FFs are tested during these six phases for data input and hold, clock enable, initialize and reverse, and stuck-at faults. Since reconfiguration is slower than test vector application, the small number of test phases is a good measure of our reduced test time. Notice also that test reconfiguration time is not constant through all six phases. In the first test phase the initial test configuration has to be set up. In the five subsequent test phases, only a few configuration bits related to the LUT function, to the programmable multiplexers and to the FFs configuration, are changed. Therefore, test reconfiguration time is smaller. Table II details the content of each CLB structural test session. The average values for the partial reconfiguration file sizes and reconfiguration times (using a 20-MHz TCK) are shown in Table III. D. Test Application The BS infrastructure is also reused to apply the test vectors and to capture the test responses, with the outputs of the CLB(s) under test being routed to unused BS register cells associated to the IOBs. However, the application of test vectors by means of the BS register would be intrusive, so an alternative User Test Register is needed (the Virtex family enables the definition of two user registers controlled through the BS infrastructure). The User Test Register created for this purpose comprises 13 cells, corresponding to the maximum number of CLB inputs in the Virtex family, and is fully compliant with the IEEE 1149.1 Standard [23]. The schematic representation of a User Test Register cell is illustrated in Fig. 9. The seven CLBs occupied by this register and the two CLBs occupied by the replication aid block, associated to the CLB needed to perform the replication, are the only FPGA hardware overhead that is implied by the proposed test methodology. In total, it accounts for less than 1% of the CLB resources of a Xilinx Virtex XCV200 device (array size CLBs), one of the FPGAs used to validate this work. Fig. 10 illustrates the Fig. 9. User Test Register cell. Fig. 10. Test of CLBs through the BS infrastructure. TABLE IV SHIFTING TIME FOR TEST VECTOR APPLICATION test infrastructure setup that is required to implement this procedure. Notice that more than one CLB may be under test at the same time, provided that enough routing resources and unused BS register cells are available. Since the same set of test vectors are applied simultaneously to all CLBs under test, the length of the User Test Register (13 bits) is fixed. Therefore, scalability of the test procedure is also possible, although dependent on the usage of the FPGA resources. Each Virtex CLB comprises two slices that are exactly equal. In total, each CLB has 13 inputs (test vectors are applied to both slices of all CLBs under test simultaneously) and 12 outputs (6 from each slice). Since the outputs of each slice are captured independently, fault location can be resolved to a single slice. The same principles apply to Virtex-II CLBs. Experimental results, obtained using a Virtex XCV200 with a TCK of 20 MHz, are shown in Table IV—the shifting time for each test vector application—and in Table V—the shifting time for the test vector responses from a CLB under test. The test of global interconnections is achieved using the same principles, with the CLB under test being replaced by a set of Authorized licensed use limited to: IEEE Xplore. Downloaded on November 27, 2008 at 13:28 from IEEE Xplore. Restrictions apply.