scieee AI-readable full text Open interactive document viewer

An 0.5-μm CMOS analog random access memory chip for TeraOPS speed multimedia video processing

Carmona Galán, Ricardo; Espejo Meana, Servando Carlos; Domínguez Castro, Rafael; Rodríguez Vázquez, Ángel Benito; Roska, Tamás; Kozek, Tibor; Chua, Leon O.

Abstract

Data compressing, data coding, and communications in object-oriented multimedia applications like telepresence, computer-aided medical diagnosis, or telesurgery require an enormous computing power - in the order of trillions of operations per second (TeraOPS). Compared with conventional digital technology, cellular neural/nonlinear network (CNN)-based computing is capable of realizing these TeraOPS-range image processing tasks in a cost-effective implementation. To exploit the computing power of the CNN Universal Machine (CNN-UM), the CNN chipset architecture has been developed a mixed-signal hardware platform for CNN-based image processing. One of the nonstandard components of the chipset is the cache memory of the analog array processor, the analog random access memory (ARAM). This paper reports on an ARAM chip that has been designed and fabricated in a 0.5-μm CMOS technology. This chip consists of a fully addressable array of 32×256 analog memory registers and has a packing density of 637 analog-memory-cells/mm2. Random and nondestructive access of the memory contents is available. Bottom-plate sampling techniques have been employed to eliminate harmonic distortion introduced by signal-dependent feedthrough. Signal coupling and interaction have been minimized by proper layout measures, including the use of protection rings and separate power supplies for the analog and the digital circuitry. This prototype features an equivalent resolution of up to 7 bits-measured by comparing the reconstructed waveform with the original input signal. Measured access times for writing/reading to/from the memory registers are of 200 ns. I/O rates via the l6-line-wide I/O bus exceed 10 Msamples/s. Storage time at room temperature is in the 80 to 100 ms range, without accuracy loss

Full text

A 0.5µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing *Ricardo Carmona1, Servando Espejo1, Rafael Domínguez-Castro1, Ángel RodríguezVázquez1, Tamás Roska2, Tibor Kozek3, Leon O. Chua3 1 Instituto de Microelectrónica de Sevilla-CNM-CSIC-Universidad de Sevilla. Edificio CICA, Avda. Reina Mercedes s/n, 41012-Sevilla, Spain. Ph. No.: 34+ 954 239923, Fax: 34+ 954 231832 E-mail: rcar[email protected] 2 MTA-SZTAKI, Analogic & Neural Computing Laboratory, Computer and Automation Institute of the Hungarian Academy of Science, Budapest, H-1111, Hungary. 3 Electronics Research Laboratory, University of California, Berkeley 258M Cory Hall, Berkeley, CA 94720, USA. Submitted for revision to the IEEE Transactions on Multimedia September 7, 1998 ABSTRACT Data compressing and coding and communications in object oriented multimedia applications like telepresence, computer-aided medical diagnosis or telesurgery require an enormous computing power − in the order of Trillion Operations per Second (TeraOPS). Compared with conventional digital technology, Cellular Neural/Nonlinear Network (CNN) based computing is capable of realizing these TeraOPS-range image processing tasks in a cost-effective implementation. To exploit the computing power of the CNN Universal Machine (CNN-UM), the CNN Chipset architecture has been developed − a mixed-signal hardware platform for CNN-based image processing. One of the non-standard components of the chipset is the cache memory of the analog array processor, the Analog Random Access Memory (ARAM). This paper reports an ARAM chip that has been designed and fabricated in a 0.5µm CMOS technology. This chip consists of a fully addressable array of analog memory registers and has a packing density of 637 analog-memory-cells/mm2. Random and non-destructive access of the memory contents is available. Bottom-plate sampling techniques have been employed to eliminate harmonic distortion introduced by signal-dependent feedthrough. Signal coupling and interaction have been minimized by proper layout measures, including the use of protection rings and separated power supplies for the analog and the digital circuitry. The prototype features an equivalent resolution of up to 7 bits −measured by comparing the reconstructed waveform with the original input signal. Measured access times for writing /reading to/from the memory registers are 200ns and 800ns, respectively. I/O rates via the 16-line wide I/O bus exceed 10Msamples/s. Storage time at room temperature is in the 80 to 100ms range, without accuracy loss. EDICS: 2-CIRC, 2-EXTN Front page footnotes12 1. This work is supported by the JSEP Grant No. FDF49620-97-1-0220-03/98 and by the ONR Grant No. N00014-98-1-0052 2. Research of the authors from IMSE-CNM (CSIC) has been supported by the spanish CICYT (Project TIC96-1392-C0202 SIVA) and the EU (Project ESPRIT IV 27077-DICTAM). 32 256× A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 3 I. INTRODUCTION Cellular Neural Networks (CNNs) are analog nonlinear dynamic processor arrays in which direct interconnections among the basic processing units are restricted to a finite local neighborhood [1]. Their potential for image processing applications was advanced shortly after their invention [2] and is based on the fact that many image processing tasks can be realized by means of weighted local interactions between neighbouring pixels [1][3]. Because of their inherently parallel processing architecture, CNNs achieve a high computation speed in the realization of these tasks. Besides, their uniformity and local connectivity make them especially suited for VLSI implementation [4][5][6][7][8]. The CNN paradigm provides the framework for the definition of an algorithmically programmable analog array computer with supercomputer power on a chip: the CNN Universal Machine (CNN-UM) [9]. Its dual-computing property enables the realization of highly complex image processing tasks by means of an on-chip analogic − analog and logic − stored program, and renders it a highly competitive alternative to the conventional digital approach to parallel image processing [3]. For example, almost 104 Pentium® are required for the TeraFLOPS array computer shipped by Intel® in 1997 [10]. Whenever accuracy in the computation is not a critical issue, as it actually happens in early-vision tasks [11], CNN-UM analogic chips are advantageous in terms of power consumption and computation speed as compared to these digital counterparts [12]. The working CNN-UM chips reported to date, with up to [5], [6] and [7] cells, respectively, contain a much smaller number of pixels than practical image sizes. For instance, conventional television applications require pixels per frame − not including the necessary scanning overhead involved in any display system [13]. Although larger chips will be available in the near future − [8] − processing of practical size images requires the adoption of system-level solutions to overcome technology limitations on the number of parallel processing cells [14]. Particularly, multiplexing the CNN-UM processors, i.e. making them operate onto a fraction of the complete input image at a time, appears sometimes the only way of operation. One possible strategy is using space-multiplexed, or multichip, CNN hardware [15]. In a multichip CNN, large arrays are built by interconnecting chips with a smaller number of cells. Each module operates simultaneously onto a fraction of the input image which is, in this way, processed in parallel. One drawback of this approach are the random fluctuations of the process 20 22×16 16× 48 48× 644 483× 64 64× A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 4 parameters among the different processors. This may cause incorrect or inaccurate operation and, thus, requires the incorporation of different correction strategies; for instance, using tuning to correct parameter deviations during the generation of the analog weights [16]. However, the major drawback of multichip CNNs is the very large number of chip modules and, specially, offchip interconnections needed. For instance, around chips and connections are required to process a pixels video frame using the CNN module reported in [17]. And around chips and connections are needed using the last generation processor reported in [8]. A different approach to using small size CNN chips for large images is time-multiplexing. By taking advantage of the computing power of the CNN-UM, a single chip can be used to process a complete video frame by operating on a fraction of the image at a time. A frame rate of − adequate for high quality video applications [13] − represents a data flow of pixels per second. Real-time processing of such rate demands processing time per pixel. Thus, by allowing for a 2-pixel wide overlap between image subsets in each scan direction − required for correct processing of the border pixels [18] −, a CNN chip should be capable to process each subimage in about ; and for a chip. Because the time constant of CNN-UM chips is in the range of [4][8] we can conclude that the time-multiplexed approach is feasible and, hence, constitutes a more cost-effective solution than the multichip one. The time-multiplexed approach requires the definition and development of an appropriate hardware platform for the CNN processor: the CNN chipset [19]. It is designed to support high speed data transmission and interfacing of the analogic processor to the sensory devices and the digital host circuitry. The Analog RAM (ARAM) is one of the non-standard parts of this chipset. It is a high-speed short-term memory buffer that operates as the cache memory [20] of the CNN processor. A straightforward realization of the required functionality would be the use of a conventional digital RAM interfaced with A/D and D/A converters. However, the resulting I/O rates between the memory and the processor would render this solution impractical. In order to realize a direct data interchange between the memory and the processor, avoiding data conversion, the implementation of a truly analog RAM chip is proposed. For full compatibility with the digital host environment and reduced fabrication cost, this ARAM should be designed using standard CMOS. The problem of on-chip analog signal storage has been faced by different authors in connection to quite diverse applications. Particularly, CMOS realizations of scanning delay-lines 8E3 4.1E5 644 483×66× 75 3.8E4 64 64× 40Hz 12.3E6 81ns 32 32× 73.8µs 320µs6464× 1µs A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 5 for video processing are presented in [21] and a high-speed SC sampling circuit is reported in [22] to capture analog waveforms from an array of sensory devices. However, no random access or non-destructive reading of the memory contents can be done. An ARAM for early vision applications was reported in [23]. However, its accuracy relays on mismatch compensation and no switching error reduction strategies are adopted. In this paper an improved version of a wellknown Sample-and-Hold (S/H) circuit is proposed to implement a fully addressable analog memory chip. It is realized in a CMOS single-poly triple-metal technology and allows non-destructive reading and random access to memory locations with a cell density of 637 cells/mm2. It features around 7 bits equivalent resolution with writing/reading access times of 200ns/800ns, respectively, and storage time at room temperature in the 80 to 100ms range. Besides, its power consumption is of only 73mW from a 3.3V power supply − achieved through multiplexing of the active S/H circuitry. In the next section, a brief review of video signal processing with CNNs is given together with the specifications of the ARAM in the CNN chipset. The, Sect. III reports the details of the ARAM prototype chip architecture and circuit design. Test results are displayed and discussed in Sect. IV. And finally, a summary of concluding remarks is given. II. VIDEO SIGNAL PROCESSING WITH CNNs A. CNN based image processing and ARAM chip specifications In the CNN Universal Machine − which has been demonstrated to be universal in the Turing sense [24] − programmable nonlinear analog dynamics are combined with programmable logic operations and analog and logic distributed memories. Complex image processing tasks are described by an analogic program [25], consisting of a sequence of analog and logic operations. This analogic program has to be compiled into a platform-dependent machine code to be executed by a particular hardware implementation. Fig. 1 depicts a diagram of the CNN-UM and its principal building blocks: the basic processing units (cells), and the Global Analogic Programming Unit (GAPU). The GAPU stores the analogic program and controls its execution. For this purpose, it is divided into two main functional blocks. First, the storage unit consisting of the Analog Program Register (APR), the Logic Program Register (LPR) and the Switch Configuration Register (SCR). They contain the machine code instructions for the analog and logic operations and the switch configuration, respectively. Second, the Global Analogic Control Unit (GACU) that decodes these instructions into a microcode that is transmitted to the cells. Inside 0.5µm 32 256× A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 6 the basic cell, three parts can be distinguished which are responsible for signal processing, storage and control of the operation − Fig. 1. For the implementation of the programmable analog dynamics, the CNN core contains the integrator and the limiter blocks. Synaptic operators can be considered as a part of the analog processing unit. A Local Logic Unit (LLU) realizes programmable logic operations between stored binary magnitudes. Short-term storage of intermediate signals is realized by Local Analog and Logic Memories (LAMs and LLMs). Signal transference and operation control is performed by the Local Communication and Control Unit (LCCU). And, finally, data exchange between the cell array and the external circuitry is realized via the Local Analog Output Unit (LAOU). In order to exploit the computing power of this architecture, The CNN chipset of Fig. 2 has been developed to interface the CNN-UM processor to the sensors and the digital environment. Data transmission is supported by three different buses. A high-speed analog bus connects the processor, the ARAM and the video signal sources. The width of this analog bus is determined by the I/O bus of the CNN-UM chip, otherwise it will limit the total throughput of the GAPU Fig. 1: CNN Universal Machine architecture, basic processing cell and global analog-and-logic programing unit. CNN Universal Machine LAM LLM LLU CNN core LCCU LAOU GAPU APR LPR SCR GACU processing storage control control A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 7 system. Digital data are transmitted via the digital bus, which is interfaced to the analog bus through A/D and D/A converters. In addition, there is a digital instruction bus. The required storage capacities and local throughput values have to be evaluated to determine the specifications for the non-standard parts, i. e. the CNN-UM and the ARAM. Assume an input image composed of -pixels (Fig. 2). It has to be decomposed into -pixel subsets that are temporarily stored one-by-one in the analog RAM chip for their processing. However, pixels in the border of this window will not be properly processed unless a certain overlap between the image fractions is allowed. Therefore, and pixel overlaps in the vertical and the horizontal direction, respectively, are considered. Taking this into account, a straightforward calculation shows that, (1) subimages are needed to cover the whole image. Each of these subimages has to be captured, processed and downloaded, thus resulting into the following total processing time for the input frame, Fig. 2: Diagram of the CNN chipset architecture. Analog RAM 1 Analog bus Digital bus Instruction bus A/DD/A DRAM VRAM µprocessor CCD Imager CNN-UC MiNi ×MaNa ×MpNp × Bai Bao Bpi Bpo MiNi × MaNa × MaNa × mono kMimo –()Nino –()× Mamo –()Nano –()× -------------------------------------------------------= Ti MiNi × A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 8 (2) where , and are the times required to acquire, download and process each subimage, respectively. For the former two times, and assuming that and are the widths of the input and output buses of the ARAM, the following is obtained, (3) where and are the times required for writing and reading, respectively, an analog register of the ARAM chip. With regards to the processing time in (2) we have to take into account that, in the more general case, the processor size is smaller than the ARAM size. Hence, the necessity arises for another multiplexation. Assume the size of the processor is and that each analogic program contain data acquisition steps, analog processing steps, logic processing operations, and data downloads. Thus, the time needed to perform the analogic algorithm on each subset is given by, (4) where, , (5) and and are the times required for the analog and the digital circuitry of the CNNUM to settle and complete the logic operation, respectively. These parameters are part of the timing specs of the CNN-UM chip. and in the expression above represents I/O times which are given by, (6) where and are the widths of the input and output buses of the CNN-UM, respectively, and and are the times required for updating and downloading analog data from one cell Ti Mimo –()Nino –()× Mamo –()Nano –()× -------------------------------------------------------Tai Tap +Tao +()⋅= Tai Tao Tap Bai Bao Tai MaNa × Bai ---------------------τai ⋅= Tao MaNa × Bao ---------------------τao ⋅= τai τao Tap MpNp × ninap nlp nd MaNa × Tap Mamo –()Nano –()× Mpmo –()Npno –()× ------------------------------------------------------- Tpp ⋅= Tpp niTpi napTpap +nlpTplp ndTpo ++= Tpap Tplp Tpi Tpo Tpi MpNp × Bpi --------------------- τpi ⋅= Tpo MpNp × Bpo --------------------- τpo ⋅= Bpi Bpo τpi τpo A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 9 of the CNN array − also defined as temporal specs of the processing chip. Assume a frame rate of frames per second. The following must be accomplished in order to process the whole input image ( ) in real-time: (7) Thus, from the mathematics above, the following design equation can be obtained, (8) We find convenient to illustrate this design equation using typical values. For instance, consider a frame rate of 40 frames per second, an input image of pixels, an analog RAM buffer of registers and a CNN array of cells. Consider as well a 2-pixel wide overlap in both, vertical and horizontal, scan directions and 16-line wide I/O buses. Then, for a typical I/O time of 500ns per memory cell, the CNN-UM chip should be capable to complete the analogic algorithm over each subimage in less than 26µs− well within the specs of CMOS CNN-UM chips [4][8]. The larger the CNN processor size the faster the system is. Besides, pipelined architectures and some interleaving of the memory blocks can be used for a more relaxed constrain on the processing time. Let us now derive the specifications for the ARAM block. It must exhibit the following features for proper usage within the CNN chipset architecture, •Non-volatility. The analog information contained in the memory registers should be maintained for a sufficiently long time. In this case, and because of the high-speed of the computation, a storage time of 100-200ms should be enough. Being a cache memory, power-off non-volatility is not necessary. •Resolution. Accuracy levels for a wide range of early-vision tasks are in the 0.8-1.5% range. It represents an equivalent resolution of 6-7 bits. Cooperative phenomena derived from the parallel processing nature of CNNs, like hyperacuity [26], allow for a moderate resolution requirement. •Random access. Some analogic algorithms designed for the CNN Universal Machine [27] require repeated reading and writing to a specific location of the memory. Thus, random access to any memory register should be provided. •Non-destructive reading. For the same reason, reading any memory location should not affect the contents, because access to then might be required several times in an anaNf MiNi × Ti1 Nf ------- ≤ 1 Nf -------MaNaMimo –()Nino –() Mamo –()Nano –() -------------------------------------------------------------- τai Bai ------- τao Bao --------+   Mimo –()Nino –() Mpmo –()Npno –() ------------------------------------------------- Tpp +≥ 512 512× 32 256×32 32× MpNp × A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 10 logic program. •High-speed. Narrow access times to the memory allow a faster operation. Although difficult to achieve, access times smaller than 100ns will be required to realize complex image processing tasks in real-time. •Input/Output. On the one hand, a serial analog input channel is needed to interface the image acquisition devices − CCD imager, composite-video signal source, ... On the other, the communication with the CNN-UM processor is accelerated by the use of parallel analog channels of width and − see Fig. 2. Obviously, the memory cell should be the smallest possible to allow obtaining the larger possible memory arrays without important yield problems. Besides, compatibility with digital CMOS voltage levels is implicitly assumed for integration with a digital environment at the system level via the instruction and digital data buses. B. Video signal interface to the CNN chipset A standard composite-video signal has a limited bandwidth of 5MHz and must, hence, be sampled at a minimum rate of 10Msamples/s. The maximum time interval between consecutivesamples is hence 100ns. In addition, the composite-video signal carries information on the luminance and chrominance of each pixel, and a synchronization pulse generated by the raster scanning of the object picture. Fig. 3 displays the envelope spectrum of a NTSC coded signal and the waveform of a scan line. Although NTSC is a color encoding standard, it is also commonly used to refer to its associated scanning standard 525/59.94. A simple implementation of a video-signal interface to the CNN chipset is portrayed in Fig. 4. It can be built up by using offthe-shelf components. Here, the incoming video signal (NTSC coded in this case) is fed into a video decoder chip. It is decomposed into its luminance (Y) and chrominance (C) components plus the recovered timing signals. By now, only the luminance component will be of interest as we are not considering color information processing. After some amplification and level shifting, if required, the ARAM chip take samples of the input via the serial input channel. Control signals and memory address codes are generated by some programmable logic device from the synchronization pulses extracted from the raw input by the NTSC decoder. Time requirements for the ARAM in this video interface can be easily derived. Using a square pixel grid -- equal horizontal and vertical sample pitch, each frame in the 525/59.94 scanning standard is composed of pixels, this includes the required blanking intervals. It means that each line of the image, containing 780 pixels, will be transmitted in 64µs approximately. Acquisition of this Bpi Bpo 780 525× A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 17 at the end of the sampling phase. It is controlled by the signal that falls slightly before . In this way, the feedthrough error is introduced via the bottom-plate of , which is maintained at a constant voltage by the opamp. Now, is independent of the input and, therefore, its derivatives with respect to are equal to zero. Consequently, no harmonic distortion due to clock feedthrough will be present at the output. The stored voltage is only affected by an additional voltage offset. A small pedestal error of magnitude (18) If the finite DC gain and the parasitic capacitor are accounted for, the output voltage is an attenuated copy of the input and an offset term appears, (19) Fig. 10 shows the opamp schematics, which has been realized through a folded cascode architecture to better fit the 3.3V power supply voltage. For 7 bits equivalent resolution of the S/ H crcuit, and assuming that a 16mV error is allowed for each sample, the opamp output swing has to be larger than 2V. Other opamp specifications are: of 20MHz − required to follow the input during the tracking phase; and Slew-Rate (SR) of 8V/µs− required to sample 4MHz band limited signals with up to 2V amplitude (peak-to-peak). Let be the small-signal transconductance of the transistors in the input differentialpair of the opamp, and the tail-current. A relation between the transistors aspect ratio and can be derived from the specifications. Because and assuming φ1 *φ1 Cmem VREF ε f Vi ε f Cgds Cmem Cgds + ------------------------------–VREF VTVREF VSS –()VSS –+[ ] ⋅= Vo11 A0 ------1Cp Ck ------+   +1– Vi1 A0 ------1Cp Ck ------+   Vos +≈ Fig. 9: Opamp schematic. IB Vi+ ViVO VB2 VB1 M1M2 M3 M4M5 M6 M7 M10 M9 M8 VBp VBn IB IB Table I: Transistor sizes M1−M224/1.2 M316/1.2 M4−M548/2.4 M6−M748/1.2 M8−M924/2.4 M10 24/0.6 GBW gm1 IBIB GBW GBW gm12πCL ()⁄= A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 18 operation within saturation region in strong inversion, one obtains, (20) where is the intrinsic transconductance of the MOS transistor. On the other hand, the necessary tail-current is fixed by the slew-rate, (21) This current determines the appropriate aspect ratio of the input differential-pair for a constant of 20MHz. The folded-cascode output stage is specified by the DC gain. By providing at least 60dB for the DC gain −− the error introduced by the parasitic capacitance is reduced to 0.1%. As is now fixed, the output stage has to be designed so as to achieve the necessary output impedance. Final compromises are resolved by phase margin and matching considerations. C. Leakage currents and storage time During the hold period, several leakage currents attempt to discharge the storage capacitor, contributing to degrade the sampled voltage value. In the first place, the reverse-biased junction formed by the n-diffusion area, corresponding to the source terminal of the pass transistor and the substrate pumps out of the upper plate of the capacitor a current that can be approximated by the reverse-biased saturation current of the parasitic diode. Another leakage is due to the subthreshold drain-to-source current of the pass MOS transistor. These effects add up resulting in a total current in the range of the pA. In this occasion, capacitors are implemented by a poly-overdiffusion structure lying on top of a weakly-doped n-well (Fig. 10). Then, the n-well/p-substrate junction is reverse-biased and the current that flows out of the bottom plate of the capacitor correspond to the associated reverse-bias saturation current. Since it is in the fA range, it limits the effect of the upper plate leakage. Stored voltage degradation in time during the hold period is now given by (22) where is the capacitance per unit area of the poly-over-diffusion structure. In these conditions, a self-discharge rate, independent of the capacitor size, is defined: W L -----2πCLGBW() 2 2knIB ------------------------------------= kn IBSR CL ⋅= GBW A 0gm1Ro = gm1 td dVc1 Cmem -------------– td dq- ⋅Iself CaA ----------–≈= Ca A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 19 , (23) where is the charge of an electron, and are the diffusion coefficients for holes and electrons, and their diffusion lenghts and and are the minority-carrier concentrations in each side of the junction. In this technology is 50mV/s. Then, the voltage at the capacitor decays linearly in time during the hold period. A maximum storage time can be defined in terms of the accuracy requirements. For an equivalent resolution of bits and a full scale range of the input signal given by , the maximum storage time ( ) is the period in which the difference between and the initially stored voltage does not exceed , that is 1/2 LSB. That is (24) which is in the 200ms range for a 10mV error. These figures, however, must be understood only as orientative because of the strong sensitivity of the leakage currents to the operating temperature. Also, incidence of light on the circuit surface can seriously degrade the contents of the memory because of the light induced generation of an extra amount of carriers. rself q Ca ------ Dppn0 Lp ----------------Dnnp0 Ln ---------------+   = qD pDn LpLnpn0np0 rself N At sto VcA2N1+ ⁄ Fig. 10:Polysilicon over n-diffusion capacitor. A’A B B’ AA’ B’ B p-substrate n-well n-diffusion polysilicon metal-1 oxide VC+ VCtsto A rself 2N1+ ⋅ ----------------------------= A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 20 D. ARAM chip floorplan This CMOS ARAM chip is composed of an array of analog memory cells. Each one contains a capacitor, a pass transistor and some local logic for address decoding. The system includes as well some digital control circuitry and an I/O interface consisting in an analog MUX/ DEMUX and 16 output buffers. Fig. 11 shows a picture of the ARAM chip floorplan. The memory matrix is arranged into 32 S/H lines with 256 capacitors each. Random access to any memory location is available with the help of two binary-to-one-hot address decoders. A code of 5 bits activates one out of the 32 row selection lines, by means of the row address decoder. Similarly, each one of the 256 columns is selected by an 8-bit code. Different access schedules can be implemented by an adequate programming of the address codes. In order to avoid the selection of more than one capacitor per row at a time, what would seriously degrade the operation, a global clock controls the duty cycle of the access signals leaving a tunable guard time interval for address codes to change. Now, with respect to the I/O interface, the 32 data lines of the array are multiplexed either to the 16-line wide I/O bus or the serial I/O channel. A digital control signal sets the serial or parallel I/O mode. Row selection signals are employed to scan the 32 data lines with either the I/O serial channel or the 16-line wide I/O bus. Some test pads have been added to characterize the output buffers for a better analysis of the test results. 32 256× bias stage Fig. 11:System architecture of the ARAM chip. column address decoder (8:256) I/O mux/demux (32:1/16) to output output buffers Analog memory cells array pads from input pads (32 x 256) row decoder (5:32) A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 21 Guidelines concerning signal interaction prevention in mixed-signal IC’s have been followed in the development of the prototype. It is a well-known fact that the integration of a significant amount of digital circuitry along with analog signal processing in the same substrate can potentially degrade system performance. A conservative layout style, with an extensive use of grounded guard rings, reduces signal coupling by opening alternative return paths to the currents induced into the substrate [30]. This is reinforced by the implementation of separated power supply and ground connections for the analog and digital circuitry and guard rings [31]. Digital lines switching at higher rates have been routed over insensitive areas and critical crossings have been shielded with a grounded metal intermediate layer. Also, analog bus lines are made wider and are separated to a larger distance than recommended by technology rules, in order to reduce cross-talk at higher frequencies. IV. EXPERIMENTAL RESULTS The first prototype of this ARAM chip has been integrated in the Hewlett-Packard 0.5µm CMOS process offered by the MOSIS service. The 24 available samples of the chip has been tested and proved to be functional. No major discrepancies have been found during the test of the different samples. First of all, a functional characterization test has been developed. Several input sine waves of different frequencies have been sampled at different rates. Fig. 12 shows a plot of the measured root-mean-square error during the reconstruction of the input waveform. It has been computed by taking the square root of the average of the squared difference between the input signal and the recovered waveform over the samples of the input wave: Fig. 12:Measured RMS error in the reconstructed waveform Chip sample No. RMSE (mV) 100Hz @10Ks/s 1Kz @10Ks/s 1Kz @100Ks/s 10Kz @100Ks/s 2 4 6 8 10 12 14 0 10 20 30 40 50 60 115 N A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 22 (25) It is important to mention that no correction of the output buffer offset or the feedthrough induced pedestal error has been made. Fig. 13 displays a reconstructed triangular wave sampled at 10KHz and a recovered sine wave sampled at 100KHz. The computed absolute RMSE is in the 13-25mV range, which means a relative error of 0.7-1.4% for a 1.8V output swing. A revealing picture of the test results is obtained by computing the FFT of the output signal. In this case, a 10KHz sine wave has been sampled at 250Ksamples/s. It has been fed to the ARAM chip through the serial input channel, therefore, 8192 samples of the input waveform Fig. 13:Recovered triangular and sine waveforms 0.5 1 1.5 2 2.5 x 10-3 0.5 1 1.5 2 2.5 3 Input Signal Freq. 100Hz Sampling Freq. 10KHz RMSE abs: 11.9mV Output swing 1.668V Time (seconds) Output waveform (volts) 0.5 1 1.5 2 2.5 x 10-3 0.5 1 1.5 2 2.5 3 Time (seconds) Output waveform (volts) rel: 0.71% Input Signal Freq. 1KHz Sampling Freq. 100KHz RMSE abs: 22.3mV Output swing 1.725V rel: 1.29% RMSE 1 N ----VikVok –() 2 k1= N ∑ ⋅= A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 23 have been taken. Fig. 14 shows the spectrum of the output signal, directly measured from the output of the chip without eliminating irrelevant information or filtering of the digitizer readings. It means that not only the stored voltage samples but also the voltage peaks occurring during address changes are captured. The magnitude of the single-tone at 10KHz is nearly 80dB above the background level. The following peak in magnitude, that takes place at the sampling rate, is approximately 30dB below the sine wave tone. Fig. 15 displays the input and the output signals as -pixel images using a linear 256-levels grayscale (8 bits deep). Each pixel in the image represents the voltage at a memory capacitor in the array. The absolute value of the difference between the input and output images is represented in the same grayscale. Besides, some real images have been loaded to the chip at 200ns per pixel and downloaded at 800ns. Fig. 16 displays the input and output pictures together with a grayscale representation of the absolute difference between them. The first two examples are -pixel pictures in a 256-level grayscale. The last one is a color picture. They have been processed in -pixel pieces because of test equipment requirements. Some spatial noise can be detected in the output picture. It is partly due to image partitioning and, on the other side, due to an improper tracking of the input at the beginning of each pixel group -- vertical lines at the 1st, 129th, 257th and 385th pixels. Because of the clocking scheme adopted to avoid the selection of more than one memory register at a time, the feedback loop of the opamp in the S/H stage is left open for a certain period. Consequently, the voltage of the output node goes up to the power supply voltage or down to the negative rail. In these conditions, the slew-rate of the opamp is insufficient to catch up with the input in the required acquisition time. Fig. 14:Spectrum of the output sinewave at 10KHz (no filtering of the readings) 0 1 2 3 4 5 x 105 -120 -100 -80 -60 -40 -20 0 Frequency (Hz) Magnitude (dB) 32 256× 512 512× 256 256× 32 128× A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 24 Finally, storage time has been measured for randomly selected cells of the array. Fig. 17 shows the difference between the initially stored voltage and the instant value through time. These data represent 24 cells in the 24 different samples of the chip. Stored voltage degradation exceeds the required accuracy levels after 80-100ms. Recursive reading of the same memory spot does not have a noticeable influence on the stored voltage. Finally, Fig. 18 shows a photograph of the prototype circuit and Table II provides a survey of data extracted from the tests results. V. CONCLUSIONS The only missing part of the CNN chipset architecture has been implemented. A random access analog memory chip has been designed and integrated in a standard 0.5 µm CMOS single-poly triple-metal technology. Measured equivalent resolution is around 7 bits. Storage time is larger than 80ms. DC power dissipation remains 73mW for a 3.3V power supply. Access times of 200ns have been obtained, while reading time is 800ns. Higher sampling and output rates can be achieved using the 16-line wide analog I/O bus. In future generations of the CNN Fig. 15:Input and output images (256 gray levels) input output abs(difference) A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 25 Fig. 16:Test input and output images input output abs(diff) A 0.5 µm CMOS Random Access Analog Memory Chip for TeraOPS Speed Multimedia Video Processing 26 input output abs(diff) Figure 16: (Continued)