scieee AI-readable full text Open interactive document viewer

Linking Bitstream Information to QoE: A Study on Still Images Using HEVC Intra Coding

Miždoš, Tomáš

Abstract

The coding tools used in image and video encoders aim at high perceptual quality for low bi-trates. Analyzing the results of the encoders in terms of quantization parameter, image partitioning, prediction modes or residuals may provide important insight into the link between those tools and the human perception. As a first step, this contribution analyzes the possibility to transcode reference images of three well-known image databases, i.e. IRCCyN/IVC, LIVE and TID2013, from their original, older formats to HEVC; thus creating a homogeneous database of 327 HEVC encoded images accompanied with bitstream parameters and values obtained from objective and subjective as-sessments. Secondly, it analyzes some of the HEVC intra coding parameters regarding their influence on the image quality by using machine learning, namely Support Vector Machine - Regression

Full text

INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER Linking Bitstream Information to QoE: A Study on Still Images Using HEVC Intra Coding Tomas MIZDOS1, Marcus BARKOWSKY 2, Miroslav UHRINA1, Peter POCTA1 1Department of Multimedia and Information-Communication Technologies, Faculty of Electrical Engineering and Information Technology, University of Zilina, Univerzitna 1, 010 26 Zilina, Slovak Republic 2Department of Interactive Systems and Internet of Things, Faculty of Applied Computer Science, Deggendorf Institute of Technology, Dieter-Gorlitz-Platz 1, 94469 Deggendorf, Germany [email protected], marcus.barko[email protected], mirosla[email protected], peter.po[email protected] DOI: 10.15598/aeee.v17i4.3625 Abstract. The coding tools used in image and video encoders aim at high perceptual quality for low bitrates. Analyzing the results of the encoders in terms of quantization parameter, image partitioning, prediction modes or residuals may provide important insight into the link between those tools and the human perception. As a first step, this contribution analyzes the possibility to transcode reference images of three wellknown image databases, i.e. IRCCyN/IVC, LIVE and TID2013, from their original, older formats to HEVC; thus creating a homogeneous database of 327 HEVC encoded images accompanied with bitstream parameters and values obtained from objective and subjective assessments. Secondly, it analyzes some of the HEVC intra coding parameters regarding their influence on the image quality by using machine learning, namely Support Vector Machine - Regression. Keywords HEVC, HEVC intra coding parameters, image quality. 1. Introduction Since the introduction of first standardized video coding algorithm ITU-T H.261, the video coding experts have improved the performance by halving the lower bitrate at the same perceptual quality every few years. Most of this gain is due to improved algorithms for intra and inter frame prediction and further refinement of residual coding. The results also show that the human visual system is as satisfied with an HEVC bitstream as it was with H.261 of more than ten times the size. The question for the QoE community is then: What can we learn from the information reduction process in the video encoder for the analysis of the image and video quality? This paper may serve as a first step towards that question by creating a dataset of HEVC intra coded images from older still image coding standards such as JPEG. This step provides us with sufficiently large group of source images, roughly associated to score obtained from subjective tests. Then, some first properties of the HEVC bitstream are analyzed towards identifying the importance of the highly flexible partitioning process which is one of the strengths of HEVC. While the motivation in this endeavor is different, the work is also closely related to No-Reference quality measurement and also can serves as a base for development of new No-Reference quality measurements. For this reason, we provide a brief overview of them. In recent years many approaches to link bitstream parameters to QoE were introduced [1], [2], [3], [4], [5], [6], [7], [8], [9], [10], [11], [12] and [13]. The quality prediction approaches developed for H.264/AVC coded videos, especially the bitstream based ones, are not applicable to HEVC coded videos [2]. A few HEVC bitstream based quality estimation models have been already proposed in the literature. The model proposed by Lei et al. in [3] is based on a linear regression and extracts Quantization Parameter Average and Skip Coding Unit Percent from HEVC bitstream to estimate a video quality predicted by PSNR. In [4], Anegekuh et al. developed the regression model, based on a moution amount metric (a metric described in this paper determining video content type) and Quantization Parameter (QP), being able to estimate video c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 436 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER quality predicted again by PSNR. An extended version of this model taking into account also a complexity of video sequences was proposed in [5]. Similarly as in [4] and [5], Anegekuh et al. proposed the regression model based on the QP and content type in this work characterized by the content type classification metric, estimating quality values predicted by PSNR. Shahid et al. designed in [6] the model, based on a two-layer feedforward artificial neural network and involving 43 HEVC bitstream parameters, being able to estimate a video quality predicted by PSNR, VQM, VIF, and PVQM. In [7] Izumi et al. proposed an objective perceptual video-quality-measurement for HEVC. They introduced two parametric NR methods to estimate a perceptual picture quality for HEVC. In [8], He et al. proposed No-Reference model which considers bitstream and display parameters, to be more generalized for different terminal scenarios. The following parameters were considered as inputs for the model: Quantisation Parameter, motion vectors, video complexity and key-frame indicator and display parameters (resolution, PPI, . . . ). The authors tried to make general model for both H.264/AVC and H.265/HEVC standards. In [9] Fazliani et al. proposed a near realtime No-Reference video quality assessment method. They trained a fully connected neural network with features extracted from both bitstream and pixel domains along with their respective subjective quality scores. In [10], Alizadeh et al. presented a novel No-Reference Video Quality Assessment (NR-VQA) based on Convolutional Neural Network (CNN) for the HEVC. In [11], Ren et al. presented a No-Reference quality assessment algorithm for UHD HEVC encoded videos. The algorithm directly extracts the specific video characteristics from the HEVC bitstream and establishes the mathematical model between the video characteristics and the final video quality through a linear regression. The algorithm extracts three video features, i.e. quantization parameters of the measurement compression damage, number of CTUs of different sizes (8×8 & 32×32), and numbers of the SAO blocks. Huang et al. proposed in [12] a No-Reference (NR) Video Quality Assessment (VQA) method for videos distorted by the HEVC. The assessment was performed without an access to a bitstream. The proposed analysis was based on the transform coefficients estimated by the decoded video pixels, which are used to estimate a level of quantization. In [13], Nawada et al. presented and described more quality indicators that can be used in a No-Reference QoE calculation, since some of them detect specific errors. Such errors are difficult to include in a global QoE model but are important from the operation point of view. In this paper, we first map HEVC Quantization Parameters (QPs) using three well-known image quality databases, i.e. the IRCCyN/IVC [14] database, the LIVE image database [15], and the TID2013 database [16], to MOS. As a result of the mapping process, a new merged dataset containing 327 HEVC encoded images accompanied with the corresponding bitstreams and second order MOS values is created. Second order MOS means indicative quality values derived from the MOSs included in the mentioned datasets by linear alignment. In the second step, the new dataset is used to analyze an impact of some HEVC intra coding parameters, i.e. QPs, distribution of Coding Unit (CU), Prediction Unit (PU) and Transform Unit (TU), on image quality experienced by the end user by deploying Support Vector Machine - Regression (SVM). The remainder of the paper is organized as follows. Section 2. describes the experiment dealing with the mapping of the HEVC QPs to MOS. In Sec. 3. , the analysis of the selected HEVC intra coding parameters impact on image quality using SVM is presented. Section 4. provides the final conclusions and suggests future work. 2. Mapping of HEVC Images to Subjective Quality of JPEG Images 2.1. Description of the Used Datasets We would like to find out whether there exists a relationship between bitstream parameters of intra coded H.265/HEVC images and their subjective quality. Unfortunately, there is no sufficiently large dataset of HEVC intra coded images annotated by subjective quality tests. We have decided to reuse older annotated datasets with similar type of distortions like the HEVC coding produces. Suitable datasets should be those which contain images with JPEG degradation because it introduces similar types of distortions as HEVC coding. In our experiments, three well-known image quality databases, namely IRCCyN/IVC [14] database, LIVE image database [15], and TID2013 database [16], containing non-degraded and also JPEG-degraded images, besides other degradations, were used. All of the above mentioned datasets involve results from subjective tests as well. In the case of the IRCCyN/IVC database, the Double Stimulus Impairment Scale (DSIS) [17] method was used to obtain quality scores. This method uses a five-grade impairment scale where one means very annoying and five imperceptible difference between the degraded and reference image. The Absolute Category Rating (ACR) [18] method was employed for a subjective evaluation in the case of the LIVE image database. The grading scale was divided into five linear equal regions ranging from low c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 437 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER IVC/IRCYN LIVE TID 2013 HM reference software QP: 0-51 Well known data Objective measuremets (PSNR, SSIM, VIF) Comparison of objective values min(VQMHEVC-VQMJPEG)2 Median of VQM values 327 HEVC encoded images INLSA MOS scaling to 1-5 Objective measuremets (PSNR, SSIM, VIF) New ILT-HEVC dataset 327 images with second order MOS 1-5 Reference images (Non-degraded) JPEG distorted images PSNR, SSIM, VIF values SSIM values Subjective scores (IVC/IRCyN MOS 1-5, LIVE MOS 1-100, TID2013 0-9) HEVC distorted images PSNR, SSIM, VIF values PSNR minial difference SSIM minial difference VIF minial difference Fig. 1: Basic schema of data preparation process. to excellent quality. The scale was then linearly converted into 1–100. A pairwise sorting methodology described in [16] was used in the case of the TID2013 database to obtain a visual quality. The subjective tests were conducted in five countries, i.e. Finland, Ukraine, France, Italy, and USA. The MOS values obtained by this methodology vary from 0 to 9 where the larger values correspond to better visual quality. More details, like resolution, number of the used reference and degraded images are presented in the Tab. 1. Tab. 1: Parameters of the used image datasets. Dataset IRCCyN/IVC LIVE TID2013 Name Number 10 29 25 of images Number 50 175 125 of JPEG distorted images Resolution 512×512 768×512 512×384 (typically) 2.2. Data Preparation for Experiment Applying machine learning algorithms require a large amount of input data. Due to this fact, we have merged the three above mentioned datasets into larger one. Source (reference) images without any distortions serve as a base for the new bigger dataset. Whole process of data preparation and creation of new dataset suitable for machine learning is depicted in Fig. 1. Firstly, a colour space of each reference image was converted from RGB to YUV420p. Secondly, all the reference images from the datasets were encoded by the HEVC compression standard using the HM reference software (version 16.20) [19]. All the encoding settings were kept default, except Quantisation Parameter. The Quantisation Parameter (QP) has been continuously changed from 1 to 51 during the encoding process. As a result, a new dataset containing 3 264 HEVC encoded images accompanied with the corresponding bitstreams was created. This dataset involves 51 levels of compression degradation of each reference image. To investigate the relationship between HEVC stream parameters and quality of image, it is necessary to have quality assessments for the corresponding images at hand. The best way to get them is to execute subjective tests but this is also a very expensive and time consuming approach. For our purpose, an approximate value of subjective score is sufficient just to see whether it is possible to map HEVC intra coding parameters to image quality. So, instead of assessing the quality of all the images subjectively, an equivalent quality to the degraded images from the original databases was sought in order to assign their subjective MOS values to the HEVC intracoded images. In other words, we mapped available quality scores of the JPEG image to the counterpart HEVC coded image. In order to find out a relationship between the HEVC and JPEG degradations, three different objective measures, namely Peak Signal to Noise Ratio (PSNR), Structural SIMilarity index, and Visual Information Fidelity (VIF), were computed. According to [20], there is a relationship between objective measure values and subjective score. Mentioned objective measurements were done on both the new HEVC dataset and JPEG degraded images. For calculating the PSNR, SSIM, and VIF values, a Video Quality Measurement Tool (VQMT) developed by Multimedia Signal Processing Group (version 1.1) was used in [21]. c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 438 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER Dispesion 4; (Selected QP 23) 23 20 24 PSNR SSIM VIF 0 5 10 15 20 25 (a) Dispesion 4; (Selected QP 48) 49 48 45 PSNR SSIM VIF 0 10 20 30 40 50 (b) Dispesion 3; (Selected QP 42) 43 42 40 PSNR SSIM VIF 0 10 20 30 40 50 (c) Dispesion 3; (Selected QP 42) 43 42 40 PSNR SSIM VIF 0 10 20 30 40 50 (d) Fig. 2: Distributions of the QP values obtained by the measures for the dispersion worst cases. 2.3. Mapping HEVC Degradation to JPEG Subjective Score Residues between the HEVC and JPEG metrics results were computed according to following Eq. (1): HEV C_QP = min(((V QM(JP EGimage)+ −V QM(HEV Cimage))2),(1) where HEV C_QP is the HEVC image with the closest objective quality to the JPEG image and V QM ={P SNR, SSIM, V IF }. Each result from the HEVC dataset was compared to the results from the JPEG dataset. This comparison was done for each objective method separately. For the HEVC encoded image with the QP value where the objective measures matched best (minimum residual), the MOS value of the JPEG degraded image was assigned. As three different measures (PSNR, SSIM, VIF) were used, sometimes happened that three different candidates for the single JPEG score were denoted. It was caused by different approaches and sensitivity of the used objective measurements. The different candidates correspond to HEVC images with slightly different values of QP. Therefore, we have calculated dispersion between images selected by each objective measurement according to the following Eq. (2): QP _Dispersionpvs = max(selectedQP (V QM))+ −min(selectedQP (V QM)), (2) where QP _Dispersionpvs is a dispersion for each image and selectedQP (V QM)is the QP of the images selected by the corrensponding VQM. A maximum dispersion obtained by the three objective metrics was 4 QP units and have appeared only twice. A difference of 4 QP units is visually hardly noticeable. It is worth noting that the difference is rather small considering the fact that PSNR, SSIM, and VIF are based on quite different approaches. The dispersion and frequency of occurrence of different QP values are depicted in Fig. 3. Distribution of QP dispersion 69 188 58 10 2 01234 QP dispersion 0 20 40 60 80 100 120 140 160 180 200 Number of occurance Fig. 3: Size of QP dispersion vs number of occurrence. We also analysed distributions of all the objective measures deployed in our experiment. Figure 2 depicts four worst cases in terms of the QP dispersion. As you can see from this picture, the values obtained from c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 439 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER the objective measures slightly differ. This is caused by different approaches used by the measures and also by a content of the test images. In order to get only one image associated to the subjective quality, a median value of QPs provided by different objective measurements was chosen and the corresponding HEVC image was tagged with the MOS score from the original JPEG dataset. Figure 4 depicts a percentage of the cases when the value obtained by the measure differs from the selected median value. It is obvious that the results obtained by PSNR mostly differ from the selected median value. The reason can be that PSNR is a rather simple measure which does not consider properties of human visual system. As the MOS scores were not directly obtained by subjective testing but by the alignment using the objective measures, we will refer to them as a second order MOS in this work. 45% 34% 16% PSNR SSIM VIF 0 5 10 15 20 25 30 35 40 45 Fig. 4: Percentage of the cases when the measures provided the different QP value from the selected median value. 012345 MOS 25 30 35 40 45 50 Median QP IVC MOS to median QP Avion Barba Boats Clown Fruit House Isabe Lena Mandril Pimen Fig. 5: Relationship between the HEVC QP and JPEG MOS for the IRCCyN/IVC database with denoted different content of images. By the above specified approach, we map the Quantisation Parameter of the images to values describing their quality, i.e. JPEG MOS. As the reused older image quality datasets have included different MOS scales, we have executed a mapping of QP to quality score on each dataset separately. A relationship between QP and JPEG MOS is depicted in Fig. 5, Fig. 6 and Fig. 7. 0 10 20 30 40 50 60 70 80 90 100 MOS 0 10 20 30 40 50 60 Median QP LIVE MOS to median QP Lighthouse Sailing2 Sailing3 Statue Woman Womanhat Building2 Coinsinfountain Flowersonih35 Bikes Buildings Caps house lighthouse2 monarch ocean paintedhouse parrots plane rapids sailing1 sailing4 stream carnivaldolls cemetry churchandcapitol dancers manfishing studentsculpture Fig. 6: Relationship between the HEVC QP and JPEG MOS for the LIVE image database with denoted different content of images. 0123456789 MOS 0 10 20 30 40 50 60 Median QP TID MOS to median QP I01 I02 I03 I04 I05 I06 I07 I08 I09 I10 I11 I12 I13 I14 I15 I16 I17 I18 I19 I20 I21 I22 I23 I24 I25 Fig. 7: Relationship between the HEVC QP and JPEG MOS for the TID2013 database with denoted different content of images. Figure 5, Fig. 6 and Fig. 7 show a typical behaviour of JPEG MOS versus QP, showing that the mapping may be considered reasonable. However, the three distinct datasets are still considered small in terms of machine learning needs. c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 440 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER 2.4. Datasets Merging Due to the requirement of larger dataset for more reliable results, the MOS values coming from the three datasets were aligned and merged. To merge the datasets, an Iterated Nested Least-Squares Algorithm (INLSA) proposed in [22] was deployed. Inputs to the algorithm were objective values of SSIM and subjective JPEG MOSs from all the three datasets. The LIVE dataset was used as the reference one when it comes to the merging process. It means that remaining datasets were aligned according to the LIVE database. The reason why the LIVE as reference was chosen is that the LIVE dataset is the biggest one and contains a higher range of subjective MOS. As the INLSA utilizes the functional relationships between an objective quality metric (extracted from the images) and the corresponding subjective MOS, SSIM was used in the merging process as the input objective quality metric. In our case, the INLSA algorithm calculated weights and scaling parameters on the basis of SSIM values and then scaled MOS linearly to 1–5 scale. Firstly, the original MOS values were linearly transformed to the interval [0, 1] in each dataset separately. It is so called distortion domain, where 0 represents no-impairment and 1 represents severe impairment. Then the scaling process follows a typical linear Eq. (3): scaledMOS =a·MOS0+b, (3) where amean gain, bshift and MOS0is MOS in the distortion domain. As it was already mentioned above, the LIVE database was used as the reference dataset. It means that a= 1 and b= 0. For the IRCCyN/IVC and the TID2013, we have obtained a= 0.7696 and b= 0.2160 and a= 1.0475 and b=−0.1168 by the INLSA respectively. As a result of the merging process, the new dataset entitled ILTHEVC containing 327 HEVC encoded images accompanied with the corresponding bitstreams and second order MOS values was made. 2.5. Possible Limitations of the Approach It is worth noting here that the proposed approach might have been positively/negatively influenced by a couple of effects. The merging process of the datasets can be one of them. We have used the INLSA to merge the values of the MOS by using the objective values but an understanding of the task by the observers can be, at the end, different. Secondly, different subjective test methodologies were deployed to get the MOS scores when it comes to the datasets. The different methodologies led to the MOS scores with diverse meaning and in different scales. Thirdly, the mapping based on the objective values can be seen as one of the prospective sources of noise in this context. Despite the fact that we have deployed the full reference objective methods, which are considered rather reliable, they are still far from being perfect. It is worth to reiterate here that we have used three different full reference measures and selected the final quality on the basis of the median value. Finally, a different appearance of the JPEG and HEVC distortions can also have some influence in this case. It is worth noting that the JPEG compression uses fixed size of blocks (8×8 pixel) unlike the HEVC, which uses adaptive size. Higher size of blocks can lead to blurring important lines and can subsequently cause worse subjective assessment. 3. Analysis of the Selected HEVC Intra Coding Parameters Impact on Image Quality by SVM-Regression 3.1. Experiment Description First, the ILT-HEVC dataset containing 327 HEVC encoded images accompanied with the corresponding bitstreams and second order MOS values was divided into two parts by the 80/20 ratio, typically deployed when it comes to the machine learning approaches. The dataset division was random while considering some restrictions to avoid a model overfitting. Thirteen different contents covering all the quality levels, ranging from 4 to 7, were randomly selected for a validation subset. It means that 68 images, out of the 327 distorted images, were selected for a validation and the remaining images were used for a training of SVM model. It should be mentioned that the content deployed in the training dataset was not replicated in the validation dataset. This fact should avoid overfitting. The SVM was selected as a machine learning technique to be used for this analysis. For a SVM training, we have used a function included in MATLAB, entitled “fitrsvm”, which trains a Support Vector Machine (SVM) regression model on a low-through moderate-dimensional predictor data set. Regarding a kernel function, as second order polynomial function has led to the best results, this function was used in this analysis as the kernel function. The epsilon parameter was set to 0.75. Other SVM settings remain unchanged (default). Input to the SVM were the QP, the Coding, Transform, and Prediction Unit Size (CU, TU, PU). The target value was the second order MOS. The used function has returned a full SVM c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 441 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER regression model trained by the selected HEVC intra coding parameters and the corresponding reference to quality values the second order MOS. The training process was repeated 9 times with different combinations of the HEVC intra coding parameters in order to find out how different HEVC intra coding parameters influence image quality. Pearson Linear Correlation Coefficient (PLCC), Spearman Rank-Order Correlation Coefficient (SROCC), and Root Mean Square Error (RMSE) were used as performance indicators in this analysis. 3.2. Experimental Results Figure 8, Fig. 9 and Fig. 10 compare the second order MOS values with the predictions provided by the SVM regression model for the selected combinations of the HEVC intra coding parameters involved in the SVM-based regression process. It can be observed from Fig. 8 that the correlation results are rather good when all the selected HEVC intra coding parameters are used. Moreover, the reported RMSE value is also rather good. 1 1.5 2 2.5 3 3.5 4 4.5 5 Real MOS 1 1.5 2 2.5 3 3.5 4 4.5 5 SVR predicted MOS Real MOS to SVR predicted MOS (Polynomial kernel q=2) PLCC = 0.9029 SROCC = 0.9226 RMSE = 0.4167 Fig. 8: Correlation between the second order MOS values and predictions provided by the SVM regression model for all the selected HEVC intra coding parameters. The results for all the investigated combinations are summarized in Tab. 2. As it can be clearly seen from this table, the best results in terms of RMSE were achieved for the combination involving the QP, CU, and PU, followed by the QP, CU, and TU combination as well as the QP and CU combination. It is worth noting here that the reported RMSE value is even smaller than that obtained for the combination involving all the parameters. Figure 9 depicts the correlation between the second order MOS values and predictions provided by the SVM regression model for the best performing parameters combination in terms of RMSE. When it comes to the combination involving the CU, TU and PU parameters and a number of the support Tab. 2: Results obtained for the different combinations of the investigated HEVC intra coding parameters. QP CU TU PU Performance indicators X PLCC = 0.90 SROCC = 0.92 RMSE = 0.45 Number of SVs = 28 X X PLCC = 0.92 SROCC = 0.93 RMSE = 0.37 Number of SVs = 26 X X PLCC = 0.92 SROCC = 0.93 RMSE = 0.38 Number of SVs = 26 X X PLCC = 0.90 SROCC = 0.93 RMSE = 0.38 Number of SVs = 24 X X X PLCC = 0.92 SROCC = 0.93 RMSE = 0.37 Number of SVs = 26 X X X PLCC = 0.92 SROCC = 0.93 RMSE = 0.37 Number of SVs = 24 X X X PLCC = 0.91 SROCC = 0.93 RMSE = 0.38 Number of SVs = 24 X X X PLCC = 0.85 SROCC = 0.86 RMSE = 0.52 Number of SVs = 48 X X X X PLCC = 0.90 SROCC = 0.92 RMSE = 0.42 Number of SVs = 27 1 1.5 2 2.5 3 3.5 4 4.5 5 Real MOS 1 1.5 2 2.5 3 3.5 4 4.5 5 SVR predicted MOS Real MOS to SVR predicted MOS (Polynomial kernel q=2) Parameters without TU size PLCC = 0.9152 SROCC = 0.9284 RMSE = 0.3688 Fig. 9: Correlation between the second order MOS values and predictions provided by the SVM regression model for the QP, CU and PU distributions. vectors, the number doubled in comparison to the other investigated combinations. We were interested in the dependency of the resulting performance on the input values. Even only by usc 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 442 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER ing QP, we discovered that the ranking order was still reasonable as it can be seen in Fig. 10. It should be noted here that the results obtained in this study were achieved by a rather small dataset. So, a statistical significance of the results is rather limited. 1 1.5 2 2.5 3 3.5 4 4.5 5 Real MOS 1 1.5 2 2.5 3 3.5 4 4.5 5 SVR predicted MOS Real MOS to SVR predicted MOS (Polynomial kernel q=2) QP only PLCC = 0.8971 SROCC = 0.9218 RMSE = 0.4533 Fig. 10: Correlation between the second order MOS values and predictions provided by the SVM regression model for the QP only. 4. Conclusions and Future Work In this contribution, we have analyzed the possibility to create a database with a more recent coding standard (HEVC) from the existing annotated databases of older coding standard, i.e. JPEG, by using different objective measures. While the results could not yet be validated by subjective testing, some evidence was presented that the approach is reasonable. In addition, machine learning, in particular SVM, was used to relate the bitstream information of HEVC to subjective quality. While this is a common approach for training NoReference bitstream models, our work is focused on understanding the link between the encoding process and the human perception. First results concerning the image partitioning used by the HEVC intra codec were presented in this paper. Further analysis of the encoding process and the resulting bitstream information is required, also extending from still image to video in future research. Acknowledgment This publication is the result of the project implementation: Centre of excellence for systems and services of intelligent transport II, ITMS 26220120050 supported by the Research & Development Operational Programme funded by the ERDF. References [1] SEVCIK, L., M. VOZNAK and J. FRNDA. QoE prediction model for multimedia services in IP network applying queuing policy. In: International Symposium on Performance Evaluation of Computer and Telecommunication Systems (SPECTS 2014). Monterey: IEEE, 2014, pp. 593–598. ISBN 978-1-4799-5745-3. DOI: 10.1109/SPECTS.2014.6879998. [2] VAN WALLENDAEL, G., N. STAELENS, L. JANOWSKI, J. DE COCK, P. DEMEESTER and R. VAN DE WALLE. No-reference bitstreambased impairment detection for high efficiency video coding. In: 2012 Fourth International Workshop on Quality of Multimedia Experience. Yarra Valley: IEEE, 2012, pp. 7–12. ISBN 978-1-46730725-3. DOI: 10.1109/QoMEX.2012.6263845. [3] LEI, X., X. JIANG and X. MA. A no-reference video quality evaluation method based on HEVC bitstream. In: Eighth International Conference on Digital Image Processing (ICDIP 2016). Chengdu: SPIE, 2016, pp. 1068–1074. ISBN 978-1-51060503-9. DOI: 10.1117/12.2243822. [4] ANEGEKUH, L., L. SUN and E. IFEACHOR. Encoded bitstream based video content type definition for HEVC video quality prediction. In: 2014 IEEE International Conference on Communications (ICC). Sydney: IEEE, 2014, pp. 1296–1301. ISBN 978-1-4799-2003-7. DOI: 10.1109/ICC.2014.6883500. [5] ANEGEKUH, L., L. SUN and E. IFEACHOR. Encoding and video content based HEVC video quality prediction. Multimedia Tools and Applications. 2015, vol. 74, iss. 11, pp. 3715–3738. ISSN 1380-7501. DOI: 10.1007/s11042-013-1795-z. [6] SHAHID, M., J. PANASIUK, G. VAN WALLENDAEL, M. BARKOWSKY and B. LOVSTROM. Predicting full-reference video quality measures using HEVC bitstream-based no-reference features. In: 2015 Seventh International Workshop on Quality of Multimedia Experience (QoMEX). Pylos-Nestoras: IEEE, 2015, pp. 1–2. ISBN 978-1-4799-8958-4. DOI: 10.1109/QoMEX.2015.7148118. [7] IZUMI, K., K. KAWAMURA, T. YOSHINO and S. NAITO. No reference video quality assessment based on parametric analysis of HEVC bitstrea. In: 2014 Sixth International Workshop on Quality c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 443 INFORMATION AND COMMUNICATION TECHNOLOGIES AND SERVICES VOLUME: 17 |NUMBER: 4 |2019 |DECEMBER of Multimedia Experience (QoMEX). Singapore: IEEE, 2014, pp. 49–50. ISBN 978-1-4799-6536-6. DOI: 10.1109/QoMEX.2014.6982287. [8] HE, T., R. XIE, J. SU, X. TANG and L. SONG. A no reference bitstream-based video quality assessment model for h.265/hevc and h.264/avc. In: 2018 IEEE International Symposium on Broadband Multimedia Systems and Broadcasting (BMSB). Valencia: IEEE, 2018. ISBN 978-1-53864729-5. DOI: 10.1109/BMSB.2018.8436879. [9] FAZLIANI, Y., E. ANDRADE and S. SHIRANI. Learning based hybrid no-reference video quality assessment of compressed videos. In: 2019 IEEE International Symposium on Circuits and Systems (ISCAS). Sapporo: IEEE, 2019, pp. 1– 5. ISBN 978-1-7281-0397-6. DOI: 10.1109/ISCAS.2019.8702584. [10] ALIZADEH, M., A. MOHAMMADI and M. SHARIFKHANI. No-reference deep compressed-based video quality assessment. In: 2018 8th International Conference on Computer and Knowledge Engineering (ICCKE). Mashhad: IEEE, 2018, pp. 130–134. ISBN 978-15386-9569-2. DOI: 10.1109/ICCKE.2018.8566395. [11] REN, J., X. JIANG, L. XU and X. GUO. No-reference quality assessment for UHD videos based on HEVC encoded bitstream. In: 2017 4th International Conference on Systems and Informatics (ICSAI). Hangzhou: IEEE, 2017, pp. 1319–1323. ISBN 978-1-5386-1107-4. DOI: 10.1109/ICSAI.2017.8248490. [12] HUANG, X., J. SOGAARD and S. FORCHHAMMER. No-reference pixel based video quality assessment for HEVC decoded video. Journal of Visual Communication and Image Representation. 2017, vol. 43, iss. 1, pp. 173–184. ISSN 1047-3203. DOI: 10.1016/j.jvcir.2017.01.002. [13] NAWADA, J., L. JANOWSKI and M. LESZCZUK. Modeling of quality of experience in no-reference mode. Journal of Telecommunications and Information Technology. 2017, vol. 2017, iss. 2, pp. 11–17. ISSN 1509-4553. DOI: 10.26636/jtit.2017.114517. [14] NINASSI, A., P. LE CALLET and F. AUTRUSSEAU. Pseudo no reference image quality metric using perceptual data hiding. In: Human vision and electronic imaging XI. San Jose: SPIE, 2006, pp. 146–157. DOI: 10.1117/12.650780. [15] SHEIKH, H. R., Z. WANG, L. CORMACK and A. C. BOVIK. Live image quality assessment database 2003. In: Laboratory for Image &Video Engineering [online]. 2006. Available at: http://live.ece.utexas.edu/ research/Quality/subjective.htm. [16] POMONARENKO, N., L. JIN, O. IEREMEIEV, V. LUKIN, K. EGIAZARIAN, J. ASTOLA, B. VOZEL, K. CHEHDI, M. CARLI, F. BATTISTI and C.-C. J. KUO. Image database TID2013: Peculiarities, results and perspectives. Signal Processing: Image Communication. 2015, vol. 30, iss. 1, pp. 57–77. ISSN 0923-5965. DOI: 10.1016/j.image.2014.10.009. [17] ITU-R BT.500-13. Methodology for the subjective assessment of the quality of television pictures. Geneva: ITU, 2012. [18] ITU-T P.910. Subjective video quality assessment methods for multimedia applications. Geneva: ITU, 2008. [19] HEVC. High Efficiency Video Coding. Berlin: HHI, 2015. [20] SEVCIK, L., L. BEHAN, J. FRNDA, M. UHRINA, J. BIENIK and M. VOZNAK. Prediction of subjective video quality based on objective assessment. In: 2018 26th Telecommunications Forum (TELFOR). Belgrade: IEEE, 2018, pp. 1–4. ISBN 978-1-5386-7171-9. DOI: 10.1109/TELFOR.2018.8612127. [21] VQMT: Video Quality Measurement Tool. In: L’Ecole polytechnique federale de Lausanne (EPFL) [online]. 2019. Available at: https://mmspg.epfl.ch/vqmt. [22] PINSON, M. H. and S. WOLF. An objective method for combining multiple subjective data sets. In: Visual Communications and Image Processing 2003. Lugano: SPIE, 2003, pp. 583–593. ISBN 0-8194-5023-5. DOI: 10.1117/12.509909. About Authors Tomas MIZDOS was born in 1993 in Poprad, Slovak Republic. He received his B.Sc. degree in Multimedia technologies at the of Multimedia and Information-Communication Technologies, Faculty of Electrical Engineering and Information Technology, University of Zilina in 2015. Currently he is a Ph.D. student at the same department. His main area of interest is Functionality and Quality of Multimedia Services, Quality of Experience, Video Coding and Digital Signal Processing. Marcus BARKOWSKY received his Dr.-Ing. degree from the University of Erlangen-Nuremberg c 2019 ADVANCES IN ELECTRICAL AND ELECTRONIC ENGINEERING 444