CNN universal chip in CMOS technology
Abstract
This paper describes the design of a CNN universal chip in a standard CMOS technology. The core of the chip consists of an array of 32×32 completely programmable CNN cells. Input image can be loaded in optical or electrical form. Accuracy is in the range of 7-8 bit, and cell density is of 33 cells/mm2.
Full text
CNNA-94 Third IEEE International Workshop on Cdlular Neural Ne4wodcs and their Applications Rome, Italy, December 1821,1994 Fig. 1 shows the conceptual block diagram of the implcmcntation described in this paper. The system has the same fundamental capabilities as the original version proposed by Roska and Chua [I], along with the possibility of optical initialization. During the system design process, anention was dedicated A CNN Unived Chip in CMOS Technology PROORAMc” APR + MABLE Figure 1 - Schematic cell architeclurc of (I CNN Universal Chip. R. Dominguez-Castro, S. Espejo, A. Rodriguez-VaZqucz, and R. Carm~na Centro Nacional de Microelectr6nica-Universidad de Sevilla Edificio CICA, Wda dn, 41012-Sevilla, SPAIN Phone: 34 - 5 - 423 99 23. Fax: 34 - 5 - 462 45 06. E-mail: [email protected] Absltact - This paper describes the design of a CNN universal chip in a standani CMOS technology. The core of the chip consists of an away of 32 x 32 conpletely pmgrammabk CNN cells. Input image can be loaded in optical or ciectrical form. Accuracy is in the range of 7-8 bit, and cell density is of 33 cel1s/mm2. 1. Introduction CNN Universal Chips are the main components of CNN Universal Machines [ 11. Their universality [2], together with their ability to implement any CNN application, makes their electronic implementation extremely anractivc. the original CNN model has been uscd-[3]. This-model has properties which are very similar to those of the original one, results in higher area and power efficiency. and is more tolerant to process parameter variations. It can be described by the following equation, where the sum extends to the neighborhood of cell c, and every symbol is used with its traditional meaning [3] except for the term Ex,, which here represents a programmable (by factor kc) offset term. 2. A synergy of analog and digital prog”b1lity A hybrid strategy which combiaes the advantages of analog and digital programmability [4] has been used to control the programmed values of CNN coefficients. This approach is based 0-7803-2070-OB4/$4.00 6 1994 IEEE 91
on a combined use of analog-programmable multipliers within the cells and of digital control signals from the outside of the cell array. Clearly, some interface circuitry is needed to generate the internal analog weight-signals from their digitally coded values. Due to the large number of multipliers in the network, even small reductions in multiplier-area justify the area dedicated to the interface circuitry. In our case, the cell-area is substantially reduced, thus system-area reduction is extremely high. The interface circuitry consists of several identical blocks, one for each programmable parameter in the network. The functionality of each interface blocks is that of a nonlinear D/A converter, and its implementation follows the adaptive architecture shown in Fig. 2, in which the analog weight signal adapts its value until the scaling factor of the analogand digitally-con- - - RAM memory. trolled multipliers coincide [4]. This approach merges most of the advantages of digitallyand analog-programmed multipliers, achieving low areas and a reduced number of control lines, simplifying the control of the weight values, and eliminating their sensitivity to global process paraxbeter variations. In addition, it allows the realization of the APR 111 using a digital required for the hybrid-control approach. . Adaptive stage+Cell array Figure 2 - Architerrum of the D/A interface circuirm 92
by two analog buffers with very low output impedance, whose output are also the weight signals transmitted to every cell in the system. During a precalibration step, controlled by +op these buffers are disconnected from the multiplier, and the differential weight signals are shorted at the input of the multiplier. This has the effect of setting the common mode voltage V,, to the voltage VL at the input of the low impedance loading stages. The error current flowing out of the p-channel current mirror is stored in the p-channel transistors with gates driven by switches controlled by +op which constitute the analog memories. After this calibration step, the output of the cmnt memories ace added to the difference of the output signals of the analog multiplier and the DIA converter, and the resulting current is integrated at the input nodes of the analog buffers. The output of these buffers control the weight signals of the multiplier, thus closing a feedback loop which settles after the analog weight signals have the correct, adapted value. The sign of the propmmed weight is introduced by swapping the connections of the input and output nodes of the buffers, with respect to the adaptive core of the circuitry. The signals driving the cells do not have any switch in their path, in order to avoid output impedance degradation. The D/A converter is implemented by a binary weighted array of n-channel transistors in saturation. A final important issue is the power dissipation of the buffers, which is extremely high. The amount of current required from the output of one buffer depends on the value of the programmed weight. For this reason, the buffers include a weight-dependent bias current, controlled by the most significative bits of the digital word encoding the weight value. 3. Error Sources and Evaluation Criteria The design of area-efficient analog integrated circuits requires a careful analysis directed towards the identification, characterization, and reduction of all possible errors arising in the practical realization. A first subdivision of these errors can be into dynamic and static non-idealities. In addition, errors can be classified as deterministic and random. Deterministic errors are originated by any non ideality derived from the circuit implementation assuming that identically designed devices are actually identical, while random errors are originated by mismatch effects. Mismatch errors are strongly dependent on the area of the devices, which is of extreme relevance in our application. Mismatch errors are the dominant error source whenever low area devices are used, and for CNN implementations, these errors are specially relevant on the multipliers. 3.l.Mismatch errors on the multipliers We assume that every multiplier has a weight deviation and an output offset. That is, while the ideal characteristic of the multiplier with a generic weight p is, xo = pxi, in practice we have x0 = (p+6p)x,+SxOff, whereboth6p(weigbterror)andGx,~(offseterror) arestochasticvariables, assumed statistically independent, with zero mean. Using this model for all the multipliers in Eq. 1, and considering the cell to be in its linear region, we have, Tdrew = { =>‘(I) + b:Ud) + kcXSat + (2) where the first line of the equation contains the nominal terms, and the second the error terms. The right-hand side of equation Eq. 2 constitutes the integrand F(t) of the integral equation governing every cell (sec J2q. l). Assuming that the state variables and the input values are dr + 2 { SO>d(t) + shy} + c { 6*& + S&} + 6kCxat +qat 93
limited to the same interval [-x,, xsat], the maximum nominal value of f(t) is given, at any time instant by, If for every multiplier, the relative weight ermr and the relative offset error are bounded by max{f(r)) =xsaI' t~{laj+Ib~l+lkll (3) then, the error of integrand f(r) relative to its maximum value is bounded by E, i,q = 5kc.xa+5x~ + {S.~d(r)+66;ud+6x+6u~}l s IC\ I dc N(c) After a detailed analysis of a large number of multipliers [5], a multiplier based on four MOS transistors operating in the triode region (Fig. 4b) was selected. This multiplier exhibits low mismatch errors and a wide range of linearity. In addition, its parasitic dynamic behavior is negligible. I Figure 4 - a) Schematic of rhe analog core of the cell. b) Multiplier 4. General characteristics and functiodity The external management of the chip is completely digital. Analog weights are specified and internally stored in digital form. For each template value, an adaptive stage transforms the digital code into an analog voltage, which is then transmitted to the network. This methodology results in weight insensitivity to process-parameter variations. as well as on accurate external control. Every cell incorporates a photosensitive device, which allows the system to be optically initialized. Electrical initialization is also possible, while output image is always downloaded in electrical form. Input and output images are assumed to be binary in every case. Electrical image uploading and downloading is realized through 32 YO bonding pads, on a row by row basis. The digital circuitry at each cell includes a four-bit static memory (LLM), a completely programmable two-input digital gate (LLU), and initialization and control circuitry (LCCU) for many different operations. Memory contents can be moved from one location to another. The four-bit memory at each cell allows the network to store four complete images. Two additional "read-only-memories" with fixed +1 and -1 values are also available. Any memory can be used as input U or as initial conditions X(0) of the network. 94
Data-transference processes (to or from the exterior of the chip, the CNN, the photosenson, or the LLU) are centralized on the LLM, as shown in Fig. 1. Microinstructions (templates, offset terms and local logic operations) are digitally stored in an on-chip static RAM memory, which implements the analog and logic program registers (APR and LPR) and has a capacity of 8 instructions. The information contained in each of these instructions is described in Tab. 1. After an instruction set has been loaded into the digital memory, the individual CNN and logic operations can be used any number of times in any order. Minimum allowed increment of the analog parameters is smaller than the expected standard deviation in the weight of the multipliers dw to mismatch effects. Therefore, coefficient discretization does not have a significant effect on the Linal precision. The allowed range of the coefficients can be scaled (at the expense of similar change in the time constant of the network) by an arbitrary value W,-. This is a consequence of using the FSR model [3]. Data description Feedback coefficients Number of Bits per AllOWtd Minimum coefficients coefficient ValUCS increment 4 9 8 [-W,,,,. Wm,l W,,,,/128 Control coefficients Offset term Boundarycells state variable Bouadarycells input value LLU truth table 95 - 9 8 I-W,,,,, Wm,l Wm,/128 d 1 8 [-Wm,. W,,I Wm,/128 xs 1 2 {-LO, 11 us 1 2 {-LO, 11 'IT 1 4 bY G ________ ---_-_--
Bias and tuning stages, which generate the analog reference voltages required for the analog core of the cells, are located at every comer of the chip area, and connected among them. A digital decoder, placed at the left side of the cell array, is used to generate the 32 control signals required for the row by row U0 protocol. The 32 YO cells located at the top of the chip include input and output digital buffers, as well as the circuitry required to multiplex the input and output signals through the same 32 lines. The circuitry located at the bottom of the ceIl array can be divided into two large sections. The first, located below, is the SRAM block implementing the APR and LPR, which contains 8 words (for 8 microinstructions) of 160 bits. The second, located above, contains 10 adaptive stages (nine weights plus the offset term). Finally, the blocks on the right side are used to generate some control signals and for miscellaneous purposes. Fig. 5b shows the layout of one cell. The size of the cell is 180 x 170 pn2, and the size of the complete chip is 7,7 x 6,8 mm2. The design has been sent for fabrication in a standard lpm CMOS technology, and is due from the foundry in the following weeks. Test results are expected to be available for presentation at the time of the conference. I Figure 5 - a) Schematic of the complete chip b) Layout of one cell of the FSR CNN universal chip References [ 11 T. Roska and L.O. Chua: “The CNN Universal Machine: An Analogic Array Computer”, IEEE Transactions on Circuits and Systems-11: Analog and Digital Signal Processing, [2] L.O. Chua, T. Roska and P.L. Venetianer: “The CNN is Universal as the Turing Machine“. IEEE Trans. Circuits and Systems I: Fundamental Theory and Applications, Vol. 40, pp [3] S. Espejo. R. Dom’nguez-Castro, A. Rodriguez-Vizquez and R. Carmona: “Convergence and Stability of the FSR CNN Model”. In this proceeding. [4] S. Espejo, R. Dom’nguez-Casuo, A. Rodriguez-Vbquez and R. Carmona: “Weight-Control Strategy for Programmable Ch\I Chips”. Ln this proceeding. [5] S. Espejo: “VU1 Design and Modeling of CNNs” Ph. Dissertation, University of Sevilla, March 1994. Vol., 40, NO.-3, March 1993. 289-29 1, April 1993. 96