Efficient bit-level design of an on-board digital TV demultiplexer
Abstract
A bit-level description of the signal processing stage of an on-board integrated VLSI multi-carrier demodulator is presented in this paper, along with a description of the optimization procedure that has been developed for the signal processing functions1. The demultiplexer is capable of handling a varying number of carriers in a 36 MHz bandwidth on the satellite up-link. Its architecture has been optimized at bit-level in a way dependent on the known input signal statistics and carrier distributions allowed by the frequency plan.
Full text
EFFICIENT BITLEVEL DESIGN OF AN ONBOARD DIGITAL TV DEMULTIPLEXER S. Calvo J.Sala A. Pages G. Vazquez Dept. of Signal Theory and Communications Universitat Politecnica de Catalunya c Jordi Girona 13 Modul D5114b 08034 Barcelona SPAIN Tel 3434015894 Fax 3434016447 email f sergioalvarez g gps.tsc.upc.es ABSTRACT A bitlevel description of the signal pro cessing stage of an onb oard integrated VLSI multicarrier demo dulator is presented in this pap er along with a description of the optimization pro cedure that has b een develop ed for the signal pro cessing functions 1 .The demultiplexer is capa ble of handling a varying numb er of carriers in a 36 MHz bandwidth on the satellite uplink. Its architecture has been optimized at bitlevel in away dep endent on the known input signal statistics and carrier distributions allowed by the frequency plan. 1INTRODUCTION Space Digital Video Broadcasting Systems are evolving toward the DVB Digital Video Broadcasting Standard based on MPEG2. An increasingly larger amount of pro cessing is b eing moved toward the space segment so that complex regenerativepayloads shall haveto becar ried by the forthcoming satellite generation. This pap er describ es the architecture that has b een developed for an OB Multicarrier Demultiplexer ASIC prototype to provide services for digital television and multimedia in the frame of the HISPANET network pro ject. HISPA NET is aimed at providing broadcasting of digital multi programme television to Spanishsp eaking communities in Europ e and America. The basic concept is to provide access to individual broadcasters and service providers through sp ecic transp onders carried by the HISPASAT satellite. One carrier conveying all programmes recei ved on the individual uplinks Multifrequency TDMA is transmitted on the downlink. Therefore demultiple xing and demo dulation not considered in this pap er must b e carried out onb oard. The design of digital onb oard systems and speci cally of the ltering stages of the digital demultiplexer havetokeep p ower consumption gate count and imple mentation losses to a minimum while maintaining ac ceptable system performance. A sp ecial criterion that 1 This researchwork has b een partially supported by the Nati onal Research Plan of Spain CYCIT TIC951022C0501 and TIC960500C1001 and the Catalonian Regional Government CIRIT 1996SGR00096. takes into account the structure of the interfering ad jacent carriers has b een develop ed 4 to derive suitable decimation lters for the demultiplexing function. The criterion optimizes jointly the lter response in the pass transition and stopbands for a given number of coe cients as the complexity of ltering is exp onential in the lter length. In this followup paper we consider the bitlevel design of the architecture therein describ ed. A system overview is presented in section 2 System Des cription. Section 3 presents the approach followed in bitlevel design for VLSI integration. Results and Con clusions are shown in Section 4. 2 SYSTEM DESCRIPTION The architecture of the digital onb oard demultiplexer shall have to deliver any carrier combination of those allowed see Fig. 2.1 and 2.2 of the following signa ling rates R s 2 R s 3 R s and 4 R s with R s the lowest signaling rate. Each carrier is QPSK mo dulated with a square ro ot raised cosine pulse rollo 0.35. Two p ossible frequency plans have been tailored to facilitate the demultiplexing scheme where the separation with adjacent carriers is 1.5 R s . In the nal architecture b oth frequency plans depicted in Fig. 2.1 and Fig. 2.2 are processed by two independent demultiplexers that can b e internally congured to deal with either of them. The overall bandwidth 36 MHz can contain up to 18 small carriers at the R s signaling rate. The sampling scheme is IF sampling at f s 36 R s 44.64 MHz. Both frequency plans have been devised to contain the four p ossible mentioned signaling rates with two constraints a that very simple frequency shifting operations should b e carried out and b that the output sampling rate of each carrier should be the same in samples per sym b ol for all rates. These two constraints have led to the construction of two frequency plans and the design of the demo dulators at 3 samples p er symb ol. The inner architecture of the demultiplexer consists of intercom municating p olyphase processors inatree scheme. In principle it would have b een p ossible to dene a com mon architecture capable of processing both frequency plans. Note in gures 2 and 2 that it should only be
neccessary to intro duce the input signal at a decimation bytwo or at a decimationbythree block. Nevertheless the impact this approachhas at bitlevel is considera ble as each pro cessing blo ck must be dimensioned to handle complex signals at sucient rate. In the end it was opted for implementing each trees hardware sepa rately. Then hardware optimization is more straight forward and can b e handled more eectively using the techniques describ ed in the section on BitLevel Design. A A A A A A A A A A A A A A A A A A A A A A A A A A H H H H H H H H H A A A A A A A A A A A A A A A A A A A A A A A A A A H H H H H H H H . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 0 222222222 666 0 4 4 4 4 2 8 8 Figure 1 Frequency Plan No. 1 and No. 2 ab ove 2.1 Mbps and 6.3 Mbps carriers b elow 2.1 Mbps 4.2 Mbps and 8.4 Mbps carriers. Demultiplexing is carried out according to the scheme dened in gures 2 and 2 1/2 band 21/2 band 2 1/2 band 2 1/2 band 2 1/3 band 3 1/3 band 3 1/3 band 3 h 1 n [] j −n h 1 n [] j n h 0 n [] j n e −j π n h 1 n [] k 6 =0 k 6 =1 k 6 =− 1 e −j π k 6 n h 2 n [] e −jn π /3 h 2 n [] e jn π /3 h 2 n [] ′ k 2 =0 ′ k 2 =1 ′ k 2 =− 1 e −j π ′ k 2 n 6 Mbps 2 Mbps Figure 2 Filtering and Decimation structure for Tree no. 1 or 223. Each blo ck is a p olyphase processor consisting of a decimated lter bank and a IDFT op eration. Bitlevel optimization is carried out separately at b oth blo cks The architecture for the 322 conguration is dis played in gure 2. Each blockis implemented with a p olyphase pro cessor that performs demultiplexing of 6 and 4 carriers with a decimation ratio of 3 and 2 res p ectively. Note that the working rate of each pro cessor is twice as fast if compared to that of aconventional p olyphase processor the decimation ratio is only half the number of carriers. This only aects the IDFT part of the p olyphase as the same outputs of the lter bank maybeusedtwice at dierent input p ositions to the IDFT to evaluate odd output samples. Only some 1/2 band 2 1/2 band 2 1/2 band 2 1/2 band 2 1/3 band 3 1/3 band 3 1/3 band 3 h 3 n []− 1 () n h 3 n [] e jn π 3 h 3 n []− 1 () n e −jn π 3 k 8 =1 k 8 =0 k 8 =− 1 e −j π k 8 n h 4 n [] e −jn π /4 h 4 n [] e jn π /4 ′ k 4 =− 1 ′ k 4 =1 h 5 n [] e −jn π /4 h 5 n [] e jn π /4 ′ k 2 =1 ′ k 2 =− 1 e −j π ′ k 4 /2⋅n e −j π ′ k 2 /2⋅n 8 Mbps 4 Mbps 2 Mbps Figure 3 Filtering and Decimation structure for Tree no. 2 or 322 of the outputs of each processor contain useful data de p ending on the frequency plan so that inhibited out puts will not be synthesized in the nal hardware. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 4 4 4 T T T h2 h1 h3 h2 h0 h3 h1 T T h0 z1 z3 z0 z2 IDFT xn x0 x1 x2 x3 Figure 4 Architecture of the 2 decimationbytwo pro cessor lter bank. The output vector y is fed into a IDFT pro cessor. The four dierent sublters constitute the de cimated versions of the carrier demultiplexing lter h n hi h 4 n i . For this particular application we have h 4 n 1 n L and h 4 n 3 0. Upp erlower branc hes of the four logic multiplexers are selected according to eveno dd IDFT output samples vector z . T he3processor is designed in a similar way. 3 BITLEVEL DESIGN Bitlevel designs of signal pro cessing algorithms often require careful wordlength dimensioning in eachnodeof the architecture. The overall performance of the system in terms of a distortion criterion dep ends on two factors a the wordlength assigned to eachinternal variable and b the numb er of quantization levels at the input. The usual criterion is thus to optimize p erformance for a re asonable complexity. It is useful to consider several op timization levels in the design of a digital architecture a top level where algorithmical optimization is carried out choice of the algorithm b bit level or register transfer level RTL where the optimization pro cedure
describ ed in this pap er applies and c gate level where the RTL sp ecication is synthesized into an interconnec tion of gatesregisters. Tradeos must be met during the design phase. Therefore it is a usual pro cedure in the provision of the system sp ecications to state the maximum tolerated level of implementation loss evalu ated in dB. We have chosen this therefore as the dis tortion criterion under which the architecture must b e optimized. The implementation loss is dened according to L I 10log 1 0 E b N o arch E b N o true dB 1 with E b N o arch the bitenergy to noise sp ectral p ower density ratio obtained at the output of the architecture nally implemented and E b N o true the true ratio of the ideal signal mo del. Previouslywehave describ ed the highlevel architec ture for the digital TV demultiplexer in terms of a p olyp hase tree. The next step in the design pro cess is to translate this algorithmics to RTL VHDL primitives. Sp ecications of the present system in terms of the number of carriers to be pro cessed the four admissi ble symb ol rates for each carrier and the varying input dynamics must b e approached with a suitable strategy at the arithmetical and logic levels. The complexityof the overall system dep ends on a large degree on the number of bits assigned to each node in the architec ture. It is therefore imp ortanttoidentify those p oints in the architecture that are more critical in terms of the implementation loss intro duced. This ob jective is usually achieved after recurrentsi mulations using probing sequences previously dened for anumb er of scenarii we refer here to those critical ca ses dened in the system sp ecications. Tables are pro duced showing the implementation loss asso ciated with each node in terms of the implicated complexityanda nal decision is reached for the its prop er dimensioning. A close understanding of the b ehaviour of bit dynamics in terms of the data statistics can provide shortcuts to this pro cedure. In our setting data statistics can be intuitively related to the spectrum or frequency plan present at the input to each p olyphase processor. The working margin of each processor can b e dened as that range in terms of signal p ower that is presented to it from the preceding stage in the architecture. Provi ded that this condition is met that particular pro cessor will work according to sp ecications. In this particular case the whole demultiplexing tree is implemented as a cascade of polyphase pro cesors so that careful monito ring of signal dynamics is crucial to guarantee that each p olyphase pro cessor is near its optimum working p oint. Bitlevel dynamics dep end heavily on the input data statistics. In particular cascaded architectures are ex tremely sensitive to this eect. Bitlevel primitives are implemented as xedp oint op erations so that recurrent pro cessing on the input data vector results in signal at tenuation along the cascade. This attenuation eect being detrimental to system p erformance in terms of the quantization SNR or implementation loss L I is heavily dep endent on the input data sp ectrum. The approach that has b een taken for this design is to monitor the data histogram along the cascade as pro bing sequences are fed through. This attenuation can b e eliminated if xedp oint amplication is implemented at key points in the architecture. The choice of the ampli cation factors is critical in the sense that only a precise understading of the signal statistics can provide a sui table value for the whole range of carrier distributions contemplated in the frequency plan. Asuitable value must b e chosen to prevent either saturation of arithme tics or an excessively low signal to quantization noise ratio for the sp ecied carrier p owers. Therefore in order to determine the b ehaviour of one target architecture it is only necessary to determine those input data distributions or sp ectra that deliver at the output the maximally and minimally attenuated signal power for constantinput signal power. It can be shown that the attenuation induced by the archi tecture dep ends on the randomness of the input signal sp ectrum. All op erations involved in demultiplexing the carrier set are linear op erations. It is straightforward to provide an intuitive justication let f x i i 2Ig be a set of input correlated and bounded random variables and let us p erform a linear op eration L on these vari ables y L x 1 x N . Then the probability density function of the random variable Y 0 p Y 0 y 0 y 0 def y max y 2 is atter the more those input random variables are cor related this can be justied from the Central Limit Theorem. In other words absence of correlation at the input can be interpreted as the output taking its maximum values with vanishing probability. Therefore wordlength dimensioning is critical in terms of the data statistics to guarantee minimum im plementation loss. The most signicant bits at several points in the architecture can be dropp ed as they will only activate with negligible probability dep ending on the data statistics. Thus the necessary logic to evalu ate those bits can be obviated in the synthesis pro cess leading to area and gate delay reductions in the nal implementation. That is true a tradeo shall haveto be established between the clipping probability those MSB bits that would activate and the logic complexity. In the nal hardware simple scaling operations with factors 1 shift bits to the upper p ositions. Rounding is p erformed afterwards to limit the wordlength passed on to the next pro cessor. In conclusion the use of histograms and characteris tic probing sequences has provided the necessary means
Gates stage 1 stage 2 stage 3 Tree 223 5229 12477 13698 Tree 322 31185 25252 27235 Table 1 Numb er of gates used bytheintegrated circuit for each tree. Trees 223 and 322 total 31404 and 83672 gates resp ectively. Tree 322 displaysaheavier computational load. to reduce the digital architecture complexity of the de multiplexer to an acceptable level. 4RESULTS Histograms and sp ectra are shown at key p oints in the architecture. We have considered two dierent scena rii a one carrier conguration containing all nine 2.1 Mbps carriers and b one carrier conguration contai ning two 8.4 Mbps carriers and one 2.1 Mbps carrier. In this way we can showtheeect of data statistics and the sp ectrum at those p oints of interest in the architec ture. Particularlywehavechosen the input to eachof the p olyphase pro cessors. The dynamic range is always sp ecied as 1 = 2 ; 1 = 2. −0.5 −0.4 −0.3 −0.2 −0.1 0 0.1 0.2 0.3 0.4 0.5 10 −4 10 −3 10 −2 10 −1 (1) −0.5 −0.4 −0.3 −0.2 −0.1 0 0.1 0.2 0.3 0.4 0.5 10 −4 10 −3 10 −2 10 −1 (2) (b) (b) (a) (a) Figure 5 Histograms a and b at the output of the rst 1 and second 2 stages. Note that histogram b is more spread than a as the numb er of indep endent carriers the rein contained is four times lower than a. It is advisable to monitor the dynamics asso ciated with b to keep a suf ciently low clipping probabilitywhileawillbepronerto granular noise. Note also than in 2 histogram b is already departing from the Gaussian shap e. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 10 −3 10 −2 10 −1 10 0 10 1 10 2 10 3 (b) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 10 −3 10 −2 10 −1 10 0 10 1 10 2 10 3 (a) 2.1 Mbps Figure 6 ab ove sp ectrum of the b conguration and b elow sp ectrum of the a conguration. The frequency axis is normalized to the sampling frequency of 44.64 MHz. Extensivesimulations have been run to obtain the wordlength assignment for each p olyphase processor capable of meeting the sp ecications. A 8bit analog to digital converter ADC is used to IFsample the input signal. Wordlengths passed b etween p olyphase pro ces sors are limited to 8 bits. The complexityofeach stage is presentedin table4asevaluated from the VHDL synt hesis pro cess. Note that Tree 322 is more complex due to more demanding computational requirements. For comparison see gures 2 and 2 5Conclusions It has been shown that linear signal pro cessing op era tions can be eciently synthesized onto an integrated circuit when the statistical dep endence between data is taken to advantage in the design pro cess. The reduction in complexitymust be traded o against clipping pro bability. Therefore saturating arithmetics is necessary to avoid excessive distortion. References 1 J.Sala A. Pages J.Riba S.Calvo G. Vazquez M.A.Rey. Algorithms Study and Simulation Re sults. Rep ort AEO001887 300697 submitted to the Europ ean Space Agency under contract ESAESTEC1209296NLUS. 2 J.Prat A. Ro drguez F.Ortega M.A.Rey. Digital Architectural Design. Report AEO 0018487 300697 submitted to the Europ ean Space Agency under contract ESAESTEC 1209296NLUS. 3 F. Ortega A. Ro drguez et al. An advanced MultiCarrier Demo dulator for the ESA OBP Sys tem. Pro ceedings of the Fifth ESA Internatio nal Workshop on Digital Signal Pro cessing Tech niques Applied to Space Communications. 4 J. Sala A. Pages S. Calvo J. Prat. Design and Implementation of aDVB OnBoard Multi Carrier Demo dulator. Pro ceedings of ICASSP98 COMM10.8.