Full text
1 INTRODUCTION 1 Fuzzy logic-based embedded system for video de-interlacing Piedad Brox1, Iluminada Baturone1,2, Santiago S´ anchez-Solano1 1Instituto de Microelectr´onica de Sevilla. Am´erico Vespucio s/n. 41092 Seville (Spain) Tel.: +34 95446666, Fax: +34 95446600 email contact: [email protected] 2Departamento de Electr´onica y Electromagnetismo, Universidad de Sevilla Av. Reina Mercedes s/n. 41012 Seville (Spain) Abstract Video de-interlacing algorithms perform a crucial task in video processing. Despite these algorithms are developed using software implementations, their implementations in hardware are required to achieve real-time operation. This paper describes the development of an embedded system for video de-interlacing. The algorithm for video de-interlacing uses three fuzzy logic-based systems to tackle three relevant features in video sequences: motion, edges, and picture repetition. The proposed strategy implements the algorithm as a hardware IP core on a FPGA-based embedded system. The paper details the proposed architecture and the design methodology to develop it. The resulting embedded system is verified on a FPGA development board and it is able to de-interlace in real-time. Keywords: Embedded system, video de-interlacing, Fuzzy system, IP core 1. Introduction Video de-interlacing is a key task in digital video processing. Some current digital transmission standards use interlaced scan format to halve video bandwidth. However, modern display devices require a progressive scan. Therefore, algorithms to interpolate the missing lines during the transmission have to be implemented at the receiver side. This kind of algorithms are called de-interlacing since they perform the reverse operation of interlacing [1]. Digital video processing chain usually includes the four stages shown in the block diagram of Figure 1: •Reception: video TV is received by a TV decoder, which is in charge of decoding the analog or digital signal. •Removal of artifacts: analog signal suffers white gaussian noise whereas digital signal is affected by video compression artifacts, which introduces two kinds of artifacts: ‘block’ and ‘mosquito’ noise. •Conversion of resolution: signal is firstly converted from interlaced to progressive. After performing deinterlacing, an algorithm for down or up scaling is applied accordingly to the format required by the device. •Picture improvement: the quality of video signal is enhanced by applying several picture enhancement algorithms such as color improvement, sharpness or edge enhancement, and improvements of contrast, details and textures. Despite video processing algorithms are developed using software implementations, their implementations in hardware are required to achieve real-time operation. Although primitive consumer equipments had several ASICs
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 2 !"#$%&'()*+%' *+,%*)"&' -)&).#$'()*+%' *+,%*)"&' !"#"$%&'() /%)0+' 1+*2,.)%"' 3+*2,.)%"'%4' #1.)4),#.0'*2+' .%' ,%561+00)%"' !"*'+,-)'.) ,!%&.,#%/) -+7)".+1$#,)"&' 869-%:"' ;,#$)"&' #'(+"!/&'()'.) !"/'-0%&'() <)*+%' +"=#",+5+".' $&#%0!") &*$!'+"*"(%) <)*+%')"62.' 0%21,+' <)*+%' %2.62.' Figure 1. The four typical stages of digital video processing to perform each of these four tasks, the current high-end products in the market contain a highly integrated chip to perform the last three tasks and even all the tasks in Figure 1. Video processing chips are included in numerous consumer devices, such as Liquid Crystal Display (LCD) TVs, plasma TVs, Audio-Video (AV) receivers, DVD players, High Definition (HD)-DVDs, Blu-ray players/recorders, and projectors. Independently of the video source (during the last years multiple video sources are proliferating), these chips normally perform intelligent and costly de-interlacing methods to obtain a high quality output progressive signal. Concerning video de-interlacing, several choices for hardware implementations are currently available in the market, such as Application-Specific Integrated Circuits (ASICs) [2]-[9], programmable solutions on Digital Signal Processing (DSPs) [10]-[14] and/or Field Programmable Gate Arrays (FPGAs) [15]-[18], and Intellectual Property (IP) cores [19]-[23] that can be used as building blocks within ASICs or FPGA designs. Recently, application specific instruction-set processors (ASIPs) for high speed computation of intra-field de-interlacing have been also proposed [24]. Each of these alternatives offers advantages and disadvantages regarding high performance, flexibility and easy upgrade, and low development cost. In commercial equipments, ASICs offer the most efficient solution justified by the huge demand of video products. In the other side, FPGAs are a good solution when there is a low volume of consumer products or in the case that the aim is to develop a prototype. Furthermore, FPGAs address with the requirement of flexibility and easiness to upgrade. Taking into account these considerations, FPGA implementation is the option chosen herein to implement the de-interlacing algorithm on an embedded system. There are many proposals of fuzzy logic-based embedded systems for control applications [25]-[26]. Recent advances in de-interlacing algorithms propose the inclusion of soft computing techniques to increase the picture quality [27]-[31]. An embedded system that implements the algorithm of [27] in real-time is proposed in this work. To the best of our knowledge this is the first hardware implementation of a fuzzy logic-based video de-interlacing algorithm. This paper details the implementation strategy that consists of an IP core on a FPGA-based embedded system. The paper is organized as follows. Section 2 summarizes a description of the algorithm and the proposed architecture to develop its hardware implementation. Section 3 explains the design methodology to build the IP core for video de-interlacing. The development of the embedded system and its validation on a development board from Xilinx is detailed in Section 4. Finally, the conclusions of this work are expounded in Section 5. 2. Architecture of fuzzy IP core for video de-interlacing The algorithm for video de-interlacing is the result of combining three fuzzy logic-based systems, each of them tackling a relevant feature: motion, edges, and possible repetition of areas in fields. The edge-adaptive system is called spatial interpolator since it performs a non-linear interpolation among pixels in the spatial neighborhood. The system that is capable of detecting repeated areas in the fields is called temporal interpolator since it interpolates pixels in the temporal neighborhood. The third system combines the outputs of the spatial and temporal interpolators according to a motion measurement in the current pixel. A block diagram of the complete system is shown in Figure 2. The three fuzzy systems are simple since they contain a low number of inputs and a low number of rules in their rulebases. The spatial interpolator (denoted as FS1 in Figure 2) is an edge-adaptive algorithm that uses five potential
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 3 !"#$"%& !"#$ !"%$ !"&$ '$(($!$)*+$#,& -'.-& /"++-)*#$"%(& 0%#-+1")*#-'&1$2-)&345& I S IT Figure 2. Block diagram of the video de-interlacing algorithm that employs the three fuzzy logic-based system Table 1. Analysis of storage resources and POs of several state-of-the-art de-interlacers De-interlacing Algorithm No. of field memories No. of Primitive Operations (POs) MC Field Insertion [32] 1 529 MC VT Filtering [32] 1 536 MC TBP [32] 2 535 MC TR [32] 2 529 MC AR [32] 2 544 GST [32] 1 543 Robust GST [32] 1 555 GST-2D [32] 2 559 Robust GST-2D [32] 2 571 Algorithm in [27] 3 153 !"##$%&'()*$+' ,-' !-' .-' /-'0-' 1-' 2-'3-' 4-' 5' ,' !' .' 6' 7' 2' 3' 4' 8' /%'0%'1%' 0' 9#:%;<)&&$='+)%$' 7)$+='>&?@A'7)$+='>&A' 7)$+='>&B@A' 4%&$#(C+:&$='+)%$' Figure 3. Pixels used for de-interlacing edge directions and six fuzzy rules to adapt the interpolation to the presence of edges. The second interpolator (FS2) is a fuzzy area-repetition-dependent temporal interpolator that uses a simple convolution to measure the dissimilarity between consecutive fields. It employs two simple fuzzy rules to adapt interpolation to repetition of areas in fields. Despite its simplicity, it reduces considerably annoying artifacts such as feathering. The feature of motion is considered by a fuzzy motion-adaptive interpolator (FS3) that uses a simple convolution to measure the motion at each pixel. This third system evaluates how this motion measurement should influence on the interpolation decisions. The algorithm for video de-interlacing employs an off-line tuning process to obtain the values of the parameters in the fuzzy systems as detailed in [27]. The tuning stage has been successfully performed by using a supervised learning algorithm that minimizes the mean square error between a set of data corresponding to progressive and de-interlaced results of different standard sequences. A detailed description of the complete system and its comparison (in terms of PSNR and visual inspection) with other state-of-the-art de-interlacers are presented in [27]. The algorithm is able to improve the results obtained by several Motion-Compensated (MC) algorithms in areas of the images with small and large motion, with clear and unclear edges, and with film and video material mixed. Concerning embedded system solutions, the hardware implementation of this algorithm is advantageous over the considered MC algorithms, as depicted in Table 1. The number of Primitive Operations (OPs) required is quite low. The number of field memories is three since it is necessary to store the values of the interpolated pixels in the previous field (as shown in Figure 3).
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 4 Table 2. Rule base of the fuzzy system for motion-adaptive interpolator Rule Antecedents Consequent 1. motion(x,y,t) is S MALL IT(x,y,t) 2. motion(x,y,t) is S MALL −MEDIUM γ1IT(x,y,t)+λ1IS(x,y,t) 3. motion(x,y,t) is MEDIUM γIT(x,y,t)+λIS(x,y,t) 4. motion(x,y,t) is MEDIUM −LARGE γ2IT(x,y,t)+λ2IS(x,y,t) 5. motion(x,y,t) is LARGE IS(x,y,t) 2.1. Description of fuzzy logic-based systems for video de-interlacing 2.1.1. Fuzzy logic-based system for motion adaptation (FS3) Motion-adaptive approaches are based on the fact that linear temporal interpolators are perfect in the absence of motion, whereas linear spatial methods offer a most adequate solution in case that motion is detected [1]. Motion detection can be implicit, as in median-based techniques, or explicit, using a motion detector. Explicit motion-adaptive algorithms calculate a new pixel value by interpolating between a spatial and a temporal de-interlacing method: Ip=(1 −α)IT+αIS(1) where ISis the output of a spatial interpolator, and ITis the output of a temporal interpolator. The variable α, which is the output of a motion detector, ranges from 0 to 1 and determines the level of motion. The performance of explicit motion-adaptive algorithms relies on the quality of the motion detector, since it is strongly dependent on the combination of both de-interlacing algorithms. The fuzzy system (FS3) performs the non-linear interpolation between a spatial and a temporal interpolator according to the presence of motion. This fuzzy system for motion adaptation has one input, which is a motion measurement based on the use of bi-dimensional convolution between a matrix of picture differences, M, and a matrix of weights, C. For a current pixel (X), the input of this fuzzy system is as follows: motion =Σ3 a=1(Σ3 b=1Ma,bCa,b) Σ3 a=1Σ3 b=1Ca,b (2) where M(i,j)are the elements of the following matrix, M, of difference values among three consecutive fields or pictures (see Figure 3): M= M(−1,−1) M(0,−1) M(1,−1) M(−1,0) M(0,0) M(1,0) M(−1,1) M(0,1) M(1,1) = |B−B0| 2 |C−C0| 2 |D−D0| 2 |Wn−W0| 2 |Xn−X0| 2 |Yn−Y0| 2 |G−G0| 2 |H−H0| 2 |I−I0| 2 (3) And Ca,bare the elements of the following weight matrix (their values have been adjusted to ease hardware implementation): C= 010 121 010 (4) The first step of the fuzzy inference process is called fuzzification. It consists of determining the degree to which the input belongs to each of the appropriate fuzzy sets via the chosen membership functions. The shape of the membership functions is piece-wise linear since this selection eases the hardware implementation of the algorithm (see Figure 4). Because of linguistic coherence, the degree of overlapping between two consecutive sets is two. The rule base of the inference system is strongly related to the selection of the number of membership functions (see Table 2). The first rule states that ‘if motion(x,y,t) is SMALL then the interpolated pixel is calculated by applying a temporal interpolator (IT)’. On the contrary, the rule number five asserts that ‘if motion(x,y,t) is LARGE the interpolation is performed by applying a spatial interpolator (IS)’. The other three rules consider intermediate situations that, when they are activated, perform different linear combinations of the spatial and the temporal interpolators. The final step to calculate the value of the interpolated pixel is the defuzzification process. Among the defuzzification methods, the Fuzzy Mean (FM) has been chosen since it is a simplified method that allows hardware simplicity. It
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 5 !"# $# !"#"$# %&%'&()*+,# -./01+2/#3&4(&&# µ! 5# "# !6# !7# !8# 9"# 96# 97# )%:;;# )%:;;<%&3+.%# %&3+.%# ;:(4&# %&3+.%<;:(4&# 98# !=# Figure 4. Normalized piecewise linear functions used for the membership functions. The parameters that define the Z, triangular, and S functions are break points (ai) and slopes (mi) Table 3. Rule base of the fuzzy system for edge-adaptive interpolation Rule Antecedents Consequent 1. bis S MALL and cis LARGE and dis LARGE X =B+I 2 2. bis LARGE and cis LARGE and dis S MALL X =D+G 2 3. bis VERY S MALL and cis LARGE and dis VERY S MALL X =B+D+G+I 4 4. ais S MALL and bis LARGE and cis LARGE and dis VERY LARGE and eis VERY LARGE X =A+J 2 5. ais VERY LARGE and bis VERY LARGE and cis LARGE and dis LARGE and eis S MALL X =E+F 2 6. Otherwise X =C+H 2 consists of a weighted average of the rule consequents, cr(see Table 2), where the weights are the activation degrees, αr, of the corresponding rules: FM =Prαr·cr Prαr(5) 2.1.2. Fuzzy logic-based system for edge adaptation (FS1) The fuzzy system (FS1) for spatial interpolation takes inspiration from well-known conventional edge-adaptive de-interlacing algorithms. This kind of techniques explores a neighborhood of the current pixel to extract information about the edge orientation [33]. Among them, Edge-based Line Average (ELA) algorithm interpolates the new pixel value by analyzing the luminance differences in the upper and lower lines. ELA looks for the most possible edge direction and then applies ‘line average’ along the selected direction. The pseudo-code of the ELA algorithm with 5+5 taps is as follows (see Figure 3): i f min(a,b,c,d,e)=b→X=(B+I)/2 elsei f min(a,b,c,d,e)=d→X=(D+G)/2 elsei f min(a,b,c,d,e)=a→X=(A+J)/2 (6) elsei f min(a,b,c,d,e)=e→X=(E+F)/2 else →X=(C+H)/2 where :a=|A−J|,b=|B−I|,c=|C−H|,d=|D−G|,e=|E−F| ELA performs well when the edge direction agrees with the maximum correlation, but otherwise introduces mistakes and degrades the image quality. The fuzzy inference system proposed in [27] is inspired by the ELA scheme but it uses a rule-based inference system to overcome ELA limitations. The inputs of this fuzzy system are the edge correlations in the five directions (a,b,c,d,e). As the starting point to design the rules of the fuzzy system heuristic knowledge expressed linguistically has been applied (see Table 3): 1. An edge is clear in direction bnot only if bis small but also if cand dare large. 2. An edge is clear in direction dnot only if dis small but also if band care large.
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 6 Table 4. Rule base for the temporal interpolator Rule Antecedents Consequent 1. dissimilarity(x,y,t) is S MALL X0 2. dissimilarity(x,y,t) is LARGE Xn 3. If band dare very small and cis large, neither there is an edge nor vertical linear interpolation performs well; the best option is a linear interpolation between the neighbors with small differences: B, D, G, I. 4. If three antecedents are large, that is, no edge is found in the three directions (b,c, and d). An edge is clear in direction a, if ais very small and eis very large. 5. If three antecedents are large, that is, no edge is found in the three directions (b,c, and d). An edge is clear in direction e, if eis very small and ais very large. 6. Otherwise, a vertical linear interpolation would be the most adequate. This heuristic knowledge is fuzzy since the concepts of ‘small’, ‘large’, and ‘very small’ are not understood as threshold values but as fuzzy ones. Fuzzy sets SMALL,LARGE,VERY SMALL, and VERY LARGE are again represented by fuzzy sets with a piecewise linear functions. The overlapping of the membership functions has been selected to ensure that no more than two rules are simultaneously activated. This strategy allows the use of the minimum operator as connective ‘and’ since a positive value for the activation degree of the sixth rule is always obtained. FM defuzzification method is used to calculate the output value. The output of the FS1 is the spatial interpolator, (IS), which is used in the complete algorithm (see Figure 2). 2.1.3. Fuzzy logic-based system for picture-repetition adaptation (FS2) The simplest temporal interpolator is named ‘field insertion’, and it consists of repeating the pixel with same spatial coordinates in the previous picture. This technique introduces annoying artifacts such as feathering when the previous pixel is not a good option. This is especially noticeable in film sequences when the previous field does not correspond to the same but to a different frame, which has motivated an active research field on film-mode detectors. Since de-interlacing should be able to cope with any kind of material: video, film and hybrid (currently very popular due to the explosion of multimedia market), the idea developed in [27] is to adapt locally temporal interpolator to the presence of repeated areas in the fields as spatial interpolator was adapted locally to the presence of edges. In this sense, area repetition, as edges and motion, is considered, in general, a local (as happens to video and hybrid sequences) and fuzzy feature of the image and, hence, two simple fuzzy rules are proposed to implement a pixel-bypixel fuzzy selection between pixels in the previous (t-1) and posterior (t+1) pictures. The input of these rules is a measure of dissimilarity between consecutive fields. Taking advantage of previous experience, dissimilarity between two consecutive fields, which is calculated by using a bi-dimensional convolution (as in the case of motion): dissimilarity = ΣΣM(i,j)C0 (i,j) ΣΣC0 (i,j) (7) where M(i,j)are the elements of the matrix, M, defined in (3) and C0 (i,j)are the coefficients of the convolution mask: C0= 1 0 1 (8) The influence of dissimilarity in selecting the kind of temporal interpolation is evaluated by considering the following fuzzy rules (see Table 4): 1. If dissimilarity between the fields (t-1) and (t) is SMALL, the most adequate interpolated value is obtained by selecting the pixel value in the previous field at the same spatial position (I(x,y,t−1)). 2. On the contrary, if dissimilarity is LARGE, the pixel value in the previous field is not a good choice and is better to bet on the pixel in the next field (I(x,y,t+1)). The shape of membership functions to model the fuzzy concepts SMALL and LARGE are also piece-wise linear functions. The output of this fuzzy system is given by applying the Fuzzy Mean defuzzification method.
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 7 . !"##$%% &'()% *+))'*,-.'% *+))'*,-.'% !"##$!$%&'$()* +'&,-* ."/-*0.(%-++$),* +'&,-* 1-2!"##$!$%&'$()* +'&,-* (),'*'/'),0% (),'*'/'),0% Figure 5. Parallel rule processing architecture 2.2. Hardware implementation of the fuzzy inference systems The proposed architecture to implement the fuzzy inference systems provides a data path for each rule. This strategy is very adequate when the rule bases contain a low number of rules and it provides a high inference speed since the rules are processed in parallel. The modules that compose the proposed architecture are a memory, an inference unit, a control circuit, and a defuzzification block. The memory stores the values to define the membership functions of the antecedents and the consequents. Since the fuzzy systems have a reduced number of rules, the memory size required is not large. The connective AND in the antecedents of the rules is represented by minimum operator. The inference can be performed in parallel since there is not a high number of the rule antecedents. The architecture has multi-input operators dedicated to aggregate the antecedents of the rules. These operators are designed by the cascade connection of two-input operators. The architecture introduces pipeline stages, which separate independent operations, and allow an overall improvement of the inference speed. The general structure of the architecture consists of three fundamental stages: the fuzzification stage, which calculates the membership degree for each of the inputs; the rule processing stage, in which the activation degrees of the rules are computed, and the defuzzification stage, which provides the system output. A block diagram of the proposed rule processing architecture is illustrated in Figure 5. 2.2.1. Fuzzification stage There are two well-known strategies to implement the membership functions: memory-based approach and arithmetic-based approach. Memory-based approaches store the membership degrees of every possible input value into a memory. The input value acts as the address of the memory to retrieve the fuzzy labels that are activated and the corresponding membership functions. The main advantage of this kind of approaches is that there is no restriction in the shape of the functions. The main drawback is the memory size, which can be very large to implement high-resolution systems, since the size of memory increases exponentially with the bit number of the inputs and the number of the resolution bits to code the membership functions. The second approach is particularly appropriated if the shapes of the membership functions are piecewise linear. The antecedent memory only stores the parameters required to carry out the calculation of the functions: the break points and the slopes. The proposed architecture implements the membership functions by using the arithmetic-based approach in parallel. From a hardware point of view, this strategy requires more resources but it achieves a high inference speed for real-time purposes. Three shape of membership functions are implemented: triangular, Z, and S functions (see Figure 4). From a hardware point of view, this selection has two advantages. The first one is that the parameters stored for each function are a point and a slope. The second one is that only one Membership Function Circuit (MFC) is required to calculate the two membership degrees per input because the maximum overlapping degree is two and the functions are normalized, so that calculating one of the degrees, µn, the other is obtained as 1µn. The flow chart to implement
2 ARCHITECTURE OF FUZZY IP CORE FOR VIDEO DE-INTERLACING 8 µ ! !"#$$%% &'%% !"#$$(")*+,"%% &-%% µ ! ")*+,"%% &-%% &-%% µ ! ")*+,"($#./)%% µ ! 012'%%3)!% 45% !"#$$%% &627(089:'%% !"#$$(")*+,"%%&'(%% µ ! µ ! 0127%%3)!% 45% !"#$$% µ ! 012;%%3)!% 45% 012<%%3)!% 45% !=#.=% +4>,=%608% 5,=>,=%6%%%%!"#$$?%%%%%!"#$$(")*+,"?%%%%")*+,"?%%%%")*+,"($#./)?%%%%$#./)8% µ !µ !µ !µ ! $#./)%% µ !&-%% 012@%% ")*+,"%% &-%% µ ! ")*+,"($#./)%% $#./)%% µ !&-%% µ !&-%% !"#$$%% &62;(089:7%% !"#$$(")*+,"%% &'(%% µ ! µ ! !"#$$(")*+,"% µ ! ")*+,"%% &-%% µ ! ")*+,"($#./)%% $#./)%% µ !&-%% µ ! &-%% !"#$$%% &62<(089:;%% !"#$$(")*+,"%% µ ! µ ! ")*+,"% µ ! ")*+,"%% &'(%% ")*+,"($#./)%% $#./)%% &-%% &-%% µ ! µ ! µ ! &-%% !"#$$%% &-%% !"#$$(")*+,"%% µ ! µ ! ")*+,"%% ")*+,"($#./)%% $#./)%% &'(%% &-%% µ ! µ ! µ ! &-%% µ ! &62@(089:<%% ")*+,"($#./)% 3)!% µ ! Figure 6. Flow chart used to implement the membership functions that are shown in Figure 4 the membership functions in Figure 4 is illustrated in Figure 6. The hardware resources required to implement these membership functions are a multiplier, an adder, and registers to store intermediate results. 2.2.2. Rule processing stage A parallel architecture, which computes the activation degree of the rules, is chosen since the rule bases contain a low number of rules and there is a low number of the rule antecedents. This means that the implementation cost in terms of hardware resources is affordable. Furthermore, this solution provides the best performance in terms of timing since the proposed activation degrees of the rules are evaluated with the smallest number of clock cycles.
3 DESIGN METHODOLOGY FOR FPGA-BASED FUZZY IP CORE 9 !"# $"# %# %# ) ) ) ) ) ) ) ) %# %# ) ) %# ) ) ) ) %# ) ) ) ) %# ) ) &'(# &)(# !*# $*# !+# $+# !,# $,# $-# !-# !"# $"# !*# $*# !+# $+# !,# $,# !-# $-# ./01/0# ./01/0# Figure 7. (a) Block diagram to implement the Fuzzy Mean defuzzification method in FS3. (b) The divider is suppressed since the sum of the activation degrees of the rules is equal to one 2.2.3. Defuzzification stage Figure 7(a) shows the block diagram of the parallel architecture to implement the Fuzzy Mean method that is employed by the fuzzy systems. It requires as many multipliers as the number of rules in the rulebase, adders and one divider. The divider in equation (5) can be suppressed if the sum of the activation degrees of the rules is equal to one. The shapes of all membership functions are chosen to ensure that no more than two rules are simultaneously activated and the sum of the activation degrees is always normalized. Hence, the expression in (5) is simplified as follows: FM →output =X i αi·ci(9) The block diagram to implement the method in (9) for the case of five rules (as is the case in FS3) is depicted in Figure 7(b). 3. Design methodology for FPGA-based fuzzy IP core The hardware implementation of the three fuzzy systems that compose the algorithm for video de-interlacing has been done following the architecture detailed in Section 2.2. A methodology has been proven to implement the algorithm on Xilinx FPGAs. A design flow, which connects the abstract algorithmic level with its physical FPGA implementation, is depicted in Figure 8. At algorithmic level, Matlab and its Image and Video Processing Toolbox have been employed to develop the algorithms. Xfuzzy 3 [34] and its integrated tool called xfsl have been used to tune the parameters of the fuzzy rule bases. System Generator from Xilinx (XSG) has been employed to develop efficient hardware-description-language (HDL) definitions from the schematic diagram described by the designer. XSG consists of a specific blockset for Simulink, called Xilinx Blockset, and the necessary software to translate the mathematical model described in Simulink into a hardware model described in HDL language. XSG maps the system parameters (defined like mask variables in Xilinx Blockset blocks) into entities, architectures, ports, signals, and attributes in the hardware realization with VHDL. The hardware description developed with XSG has been verified with Simulink simulations of test video sequences stored in the Matlab workspace. Finally, CAD environments to easily develop FPGA designs from HDL descriptions (in particular, ISE from Xilinx) have been employed to develop the implementations that perform experimental validation. Furthermore, XSG allows the verification of the implementation on the FPGA introducing the inputs through the Matlab environment. This verification is also known as hardware-in-the-loop testing. A picture of the top level description in the XSG design is shown in Figure 9. Since the design in XSG is integrated into Simulink, other elements from several blocksets can be used to verify the correct behavior of the design. For