scieee AI-readable full text Open interactive document viewer

Efficient bit-level design of an on-board digital TV demultiplexer

Sala Álvarez, José,Pagès Zamora, Alba Maria,Vázquez Grau, Gregorio

Abstract

A bit-level description of the signal processing stage of an on-board integrated VLSI multi-carrier demodulator is presented in this paper, along with a description of the optimization procedure that has been developed for the signal processing functions1. The demultiplexer is capable of handling a varying number of carriers in a 36 MHz bandwidth on the satellite up-link. Its architecture has been optimized at bit-level in a way dependent on the known input signal statistics and carrier distributions allowed by the frequency plan.

Full text

EFFICIENT BITLEVEL DESIGN OF AN ONBOARD DIGITAL TV DEMULTIPLEXER S. Calvo J.Sala A. Pages G. Vazquez Dept. of Signal Theory and Communications Universitat Politecnica de Catalunya c Jordi Girona 13 Modul D5114b 08034 Barcelona SPAIN Tel 3434015894 Fax 3434016447 email f sergioalvarez g gps.tsc.upc.es ABSTRACT A bitlevel description of the signal pro cessing stage of an onb oard integrated VLSI multicarrier demo dulator is presented in this pap er along with a description of the optimization pro cedure that has b een develop ed for the signal pro cessing functions 1 .The demultiplexer is capa ble of handling a varying numb er of carriers in a 36 MHz bandwidth on the satellite uplink. Its architecture has been optimized at bitlevel in away dep endent on the known input signal statistics and carrier distributions allowed by the frequency plan. 1INTRODUCTION Space Digital Video Broadcasting Systems are evolving toward the DVB Digital Video Broadcasting Standard based on MPEG2. An increasingly larger amount of pro cessing is b eing moved toward the space segment so that complex regenerativepayloads shall haveto becar ried by the forthcoming satellite generation. This pap er describ es the architecture that has b een developed for an OB Multicarrier Demultiplexer ASIC prototype to provide services for digital television and multimedia in the frame of the HISPANET network pro ject. HISPA NET is aimed at providing broadcasting of digital multi programme television to Spanishsp eaking communities in Europ e and America. The basic concept is to provide access to individual broadcasters and service providers through sp ecic transp onders carried by the HISPASAT satellite. One carrier conveying all programmes recei ved on the individual uplinks Multifrequency TDMA is transmitted on the downlink. Therefore demultiple xing and demo dulation not considered in this pap er must b e carried out onb oard. The design of digital onb oard systems and speci cally of the ltering stages of the digital demultiplexer havetokeep p ower consumption gate count and imple mentation losses to a minimum while maintaining ac ceptable system performance. A sp ecial criterion that 1 This researchwork has b een partially supported by the Nati onal Research Plan of Spain CYCIT TIC951022C0501 and TIC960500C1001 and the Catalonian Regional Government CIRIT 1996SGR00096. takes into account the structure of the interfering ad jacent carriers has b een develop ed 4 to derive suitable decimation lters for the demultiplexing function. The criterion optimizes jointly the lter response in the pass  transition and stopbands for a given number of coe cients as the complexity of ltering is exp onential in the lter length. In this followup paper we consider the bitlevel design of the architecture therein describ ed. A system overview is presented in section 2 System Des cription. Section 3 presents the approach followed in bitlevel design for VLSI integration. Results and Con clusions are shown in Section 4. 2 SYSTEM DESCRIPTION The architecture of the digital onb oard demultiplexer shall have to deliver any carrier combination of those allowed see Fig. 2.1 and 2.2 of the following signa ling rates R s  2 R s  3 R s and 4 R s with R s the lowest signaling rate. Each carrier is QPSK mo dulated with a square ro ot raised cosine pulse rollo 0.35. Two p ossible frequency plans have been tailored to facilitate the demultiplexing scheme where the separation with adjacent carriers is 1.5 R s . In the nal architecture b oth frequency plans depicted in Fig. 2.1 and Fig. 2.2 are processed by two independent demultiplexers that can b e internally congured to deal with either of them. The overall bandwidth 36 MHz can contain up to 18 small carriers at the R s signaling rate. The sampling scheme is IF sampling at f s 36 R s 44.64 MHz. Both frequency plans have been devised to contain the four p ossible mentioned signaling rates with two constraints a that very simple frequency shifting operations should b e carried out and b that the output sampling rate of each carrier should be the same in samples per sym b ol for all rates. These two constraints have led to the construction of two frequency plans and the design of the demo dulators at 3 samples p er symb ol. The inner architecture of the demultiplexer consists of intercom municating p olyphase processors inatree scheme. In principle it would have b een p ossible to dene a com mon architecture capable of processing both frequency plans. Note in gures 2 and 2 that it should only be neccessary to intro duce the input signal at a decimation bytwo or at a decimationbythree block. Nevertheless the impact this approachhas at bitlevel is considera ble as each pro cessing blo ck must be dimensioned to handle complex signals at sucient rate. In the end it was opted for implementing each trees hardware sepa rately. Then hardware optimization is more straight forward and can b e handled more eectively using the techniques describ ed in the section on BitLevel Design. A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  H H H  H H H  H H H   A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A  A A                     H H H H   H H H H . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .  0  222222222 666 0 4 4 4 4 2 8 8 Figure 1 Frequency Plan No. 1 and No. 2 ab ove 2.1 Mbps and 6.3 Mbps carriers b elow 2.1 Mbps 4.2 Mbps and 8.4 Mbps carriers. Demultiplexing is carried out according to the scheme dened in gures 2 and 2 1/2 band 21/2 band 2 1/2 band 2 1/2 band 2 1/3 band 3 1/3 band 3 1/3 band 3 h 1 n [] j −n h 1 n [] j n h 0 n [] j n e −j π n h 1 n [] k 6 =0 k 6 =1 k 6 =− 1 e −j π k 6 n h 2 n [] e −jn π /3 h 2 n [] e jn π /3 h 2 n [] ′ k 2 =0 ′ k 2 =1 ′ k 2 =− 1 e −j π ′ k 2 n 6 Mbps 2 Mbps Figure 2 Filtering and Decimation structure for Tree no. 1 or 223. Each blo ck is a p olyphase processor consisting of a decimated lter bank and a IDFT op eration. Bitlevel optimization is carried out separately at b oth blo cks The architecture for the 322 conguration is dis played in gure 2. Each blockis implemented with a p olyphase pro cessor that performs demultiplexing of 6 and 4 carriers with a decimation ratio of 3 and 2 res p ectively. Note that the working rate of each pro cessor is twice as fast if compared to that of aconventional p olyphase processor the decimation ratio is only half the number of carriers. This only aects the IDFT part of the p olyphase as the same outputs of the lter bank maybeusedtwice at dierent input p ositions to the IDFT to evaluate odd output samples. Only some 1/2 band 2 1/2 band 2 1/2 band 2 1/2 band 2 1/3 band 3 1/3 band 3 1/3 band 3 h 3 n []− 1 () n h 3 n [] e jn π 3 h 3 n []− 1 () n e −jn π 3 k 8 =1 k 8 =0 k 8 =− 1 e −j π k 8 n h 4 n [] e −jn π /4 h 4 n [] e jn π /4 ′ k 4 =− 1 ′ k 4 =1 h 5 n [] e −jn π /4 h 5 n [] e jn π /4 ′ k 2 =1 ′ k 2 =− 1 e −j π ′ k 4 /2⋅n e −j π ′ k 2 /2⋅n 8 Mbps 4 Mbps 2 Mbps Figure 3 Filtering and Decimation structure for Tree no. 2 or 322 of the outputs of each processor contain useful data de p ending on the frequency plan so that inhibited out puts will not be synthesized in the nal hardware. .  . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .             4 4 4 4 T T T h2 h1 h3 h2 h0 h3 h1 T T h0 z1 z3 z0 z2 IDFT xn x0 x1 x2 x3 Figure 4 Architecture of the 2 decimationbytwo pro cessor lter bank. The output vector y is fed into a IDFT pro cessor. The four dierent sublters constitute the de cimated versions of the carrier demultiplexing lter h  n  hi  h 4 n  i . For this particular application we have h 4 n 1    n  L  and h 4 n 3  0. Upp erlower branc hes of the four logic multiplexers are selected according to eveno dd IDFT output samples vector z . T he3processor is designed in a similar way. 3 BITLEVEL DESIGN Bitlevel designs of signal pro cessing algorithms often require careful wordlength dimensioning in eachnodeof the architecture. The overall performance of the system in terms of a distortion criterion dep ends on two factors a the wordlength assigned to eachinternal variable and b the numb er of quantization levels at the input. The usual criterion is thus to optimize p erformance for a re asonable complexity. It is useful to consider several op timization levels in the design of a digital architecture a top level where algorithmical optimization is carried out choice of the algorithm b bit level or register transfer level RTL where the optimization pro cedure describ ed in this pap er applies and c gate level where the RTL sp ecication is synthesized into an interconnec tion of gatesregisters. Tradeos must be met during the design phase. Therefore it is a usual pro cedure in the provision of the system sp ecications to state the maximum tolerated level of implementation loss evalu ated in dB. We have chosen this therefore as the dis tortion criterion under which the architecture must b e optimized. The implementation loss is dened according to L I 10log 1 0  E b N o  arch  E b N o  true dB 1 with  E b N o  arch the bitenergy to noise sp ectral p ower density ratio obtained at the output of the architecture nally implemented and  E b N o  true the true ratio of the ideal signal mo del. Previouslywehave describ ed the highlevel architec ture for the digital TV demultiplexer in terms of a p olyp hase tree. The next step in the design pro cess is to translate this algorithmics to RTL VHDL primitives. Sp ecications of the present system in terms of the number of carriers to be pro cessed the four admissi ble symb ol rates for each carrier and the varying input dynamics must b e approached with a suitable strategy at the arithmetical and logic levels. The complexityof the overall system dep ends on a large degree on the number of bits assigned to each node in the architec ture. It is therefore imp ortanttoidentify those p oints in the architecture that are more critical in terms of the implementation loss intro duced. This ob jective is usually achieved after recurrentsi mulations using probing sequences previously dened for anumb er of scenarii we refer here to those critical ca ses dened in the system sp ecications. Tables are pro duced showing the implementation loss asso ciated with each node in terms of the implicated complexityanda nal decision is reached for the its prop er dimensioning. A close understanding of the b ehaviour of bit dynamics in terms of the data statistics can provide shortcuts to this pro cedure. In our setting data statistics can be intuitively related to the spectrum or frequency plan present at the input to each p olyphase processor. The working margin of each processor can b e dened as that range in terms of signal p ower that is presented to it from the preceding stage in the architecture. Provi ded that this condition is met that particular pro cessor will work according to sp ecications. In this particular case the whole demultiplexing tree is implemented as a cascade of polyphase pro cesors so that careful monito ring of signal dynamics is crucial to guarantee that each p olyphase pro cessor is near its optimum working p oint. Bitlevel dynamics dep end heavily on the input data statistics. In particular cascaded architectures are ex tremely sensitive to this eect. Bitlevel primitives are implemented as xedp oint op erations so that recurrent pro cessing on the input data vector results in signal at tenuation along the cascade. This attenuation eect being detrimental to system p erformance in terms of the quantization SNR or implementation loss L I  is heavily dep endent on the input data sp ectrum. The approach that has b een taken for this design is to monitor the data histogram along the cascade as pro bing sequences are fed through. This attenuation can b e eliminated if xedp oint amplication is implemented at key points in the architecture. The choice of the ampli cation factors is critical in the sense that only a precise understading of the signal statistics can provide a sui table value for the whole range of carrier distributions contemplated in the frequency plan. Asuitable value must b e chosen to prevent either saturation of arithme tics or an excessively low signal to quantization noise ratio for the sp ecied carrier p owers. Therefore in order to determine the b ehaviour of one target architecture it is only necessary to determine those input data distributions or sp ectra that deliver at the output the maximally and minimally attenuated signal power for constantinput signal power. It can be shown that the attenuation induced by the archi tecture dep ends on the randomness of the input signal sp ectrum. All op erations involved in demultiplexing the carrier set are linear op erations. It is straightforward to provide an intuitive justication let f x i i 2Ig be a set of input correlated and bounded random variables and let us p erform a linear op eration L    on these vari ables y  L  x 1    x N . Then the probability density function of the random variable Y 0  p Y 0  y 0   y 0 def  y max y  2 is atter the more those input random variables are cor related this can be justied from the Central Limit Theorem. In other words absence of correlation at the input can be interpreted as the output taking its maximum values with vanishing probability. Therefore wordlength dimensioning is critical in terms of the data statistics to guarantee minimum im plementation loss. The most signicant bits at several points in the architecture can be dropp ed as they will only activate with negligible probability dep ending on the data statistics. Thus the necessary logic to evalu ate those bits can be obviated in the synthesis pro cess leading to area and gate delay reductions in the nal implementation. That is true a tradeo shall haveto be established between the clipping probability those MSB bits that would activate and the logic complexity. In the nal hardware simple scaling operations with factors  1 shift bits to the upper p ositions. Rounding is p erformed afterwards to limit the wordlength passed on to the next pro cessor. In conclusion the use of histograms and characteris tic probing sequences has provided the necessary means Gates stage 1 stage 2 stage 3 Tree 223 5229 12477 13698 Tree 322 31185 25252 27235 Table 1 Numb er of gates used bytheintegrated circuit for each tree. Trees 223 and 322 total 31404 and 83672 gates resp ectively. Tree 322 displaysaheavier computational load. to reduce the digital architecture complexity of the de multiplexer to an acceptable level. 4RESULTS Histograms and sp ectra are shown at key p oints in the architecture. We have considered two dierent scena rii a one carrier conguration containing all nine 2.1 Mbps carriers and b one carrier conguration contai ning two 8.4 Mbps carriers and one 2.1 Mbps carrier. In this way we can showtheeect of data statistics and the sp ectrum at those p oints of interest in the architec ture. Particularlywehavechosen the input to eachof the p olyphase pro cessors. The dynamic range is always sp ecied as   1 = 2 ; 1 = 2. −0.5 −0.4 −0.3 −0.2 −0.1 0 0.1 0.2 0.3 0.4 0.5 10 −4 10 −3 10 −2 10 −1 (1) −0.5 −0.4 −0.3 −0.2 −0.1 0 0.1 0.2 0.3 0.4 0.5 10 −4 10 −3 10 −2 10 −1 (2) (b) (b) (a) (a) Figure 5 Histograms a and b at the output of the rst 1 and second 2 stages. Note that histogram b is more spread than a as the numb er of indep endent carriers the rein contained is four times lower than a. It is advisable to monitor the dynamics asso ciated with b to keep a suf ciently low clipping probabilitywhileawillbepronerto granular noise. Note also than in 2 histogram b is already departing from the Gaussian shap e. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 10 −3 10 −2 10 −1 10 0 10 1 10 2 10 3 (b) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 10 −3 10 −2 10 −1 10 0 10 1 10 2 10 3 (a) 2.1 Mbps Figure 6 ab ove sp ectrum of the b conguration and b elow sp ectrum of the a conguration. The frequency axis is normalized to the sampling frequency of 44.64 MHz. Extensivesimulations have been run to obtain the wordlength assignment for each p olyphase processor capable of meeting the sp ecications. A 8bit analog to digital converter ADC is used to IFsample the input signal. Wordlengths passed b etween p olyphase pro ces sors are limited to 8 bits. The complexityofeach stage is presentedin table4asevaluated from the VHDL synt hesis pro cess. Note that Tree 322 is more complex due to more demanding computational requirements. For comparison see gures 2 and 2 5Conclusions It has been shown that linear signal pro cessing op era tions can be eciently synthesized onto an integrated circuit when the statistical dep endence between data is taken to advantage in the design pro cess. The reduction in complexitymust be traded o against clipping pro bability. Therefore saturating arithmetics is necessary to avoid excessive distortion. References 1 J.Sala A. Pages J.Riba S.Calvo G. Vazquez M.A.Rey. Algorithms Study and Simulation Re sults. Rep ort AEO001887 300697 submitted to the Europ ean Space Agency under contract ESAESTEC1209296NLUS. 2 J.Prat A. Ro drguez F.Ortega M.A.Rey. Digital Architectural Design. Report AEO 0018487 300697 submitted to the Europ ean Space Agency under contract ESAESTEC 1209296NLUS. 3 F. Ortega A. Ro drguez et al. An advanced MultiCarrier Demo dulator for the ESA OBP Sys tem. Pro ceedings of the Fifth ESA Internatio nal Workshop on Digital Signal Pro cessing Tech niques Applied to Space Communications. 4 J. Sala A. Pages S. Calvo J. Prat. Design and Implementation of aDVB OnBoard Multi Carrier Demo dulator. Pro ceedings of ICASSP98 COMM10.8.