scieee AI-readable full text Open interactive document viewer

Gesture recognition using histograms of optical flow

Wen, Ruochen

Abstract

In the field of Computer Vision, Gesture Recognition is kind of crucial problem. What differ video classification from normal image classification is that the enormous amount number of video data can not be ignored, because those complex data could lead to significant decline of computational efficiency, Therefore, this article mainly focus on how to obtain a video classification with both accuracy and efficiency in the meanwhile. In order to create a n ideal video classification system, the article use an improved speed-up Bag-of-Words model as basic pipeline. In each part of the pipeline, we apply and evaluate various strategies. In particular, in the step of feature extraction, we create a type of fast information feature descriptor for video, which is called Histogram of Optical Flow. Besides, we try to modify frame sampling rate of video, aiming to reduce calculation. In the process of creating features, we use sampling rate which is same to the size of a block. In this way, each block could be used repeatedly and the calculation will be reduced. When building visual word vocabulary and using SVM for classification, we use different methods to find a best performance of our system. As a final result, we get a trade-off between efficiency and accuracy of our gesture recognition system.

Full text

1 Gesture Recognition Using Histogram of Optical Flow  Ruochen Wen 2017 June 2 Abstract In the field of Computer Vision, Gesture Recognition is kind of crucial problem. What differ video classification from normal image classification is that the enormous amount number of video data can not be ignored, because those complex data could lead to significant decline of computational efficiency, Therefore, this article mainly focus on how to obtain a video classification with both accuracy and efficiency in the meanwhile. In order to create a n ideal video classification system, the article use an improved speed-up Bag-of-Words model as basic pipeline. In each part of the pipeline, we apply and evaluate various strategies. In particular, in the step of feature extraction, we create a type of fast information feature descriptor for video, which is called Histogram of Optical Flow. Besides, we try to modify frame sampling rate of video, aiming to reduce calculation. In the process of creating features, we use sampling rate which is same to the size of a block. In this way, each block could be used repeatedly and the calculation will be reduced. When building visual word vocabulary and using SVM for classification, we use different methods to find a best performance of our system. As a final result, we get a trade-off between efficiency and accuracy of our gesture recognition system. Key words: Optical Flow, Feature Extraction, Bag-of-Words, Gesture Recognition 3 Contents Gesture Recognition Using Histogram of Optical Flow 1 Contents 3 1.Introduction 5 1.1 The main challenge in gesture recognition 5 1.2 Research status at home and abroad 5 1.2 The main content of this article 7 2 Bag-of-Word Model 8 2.1 Introduction of Bag - of - Word Model 8 2.2 Innovation and design 11 3.Feature Extraction 12 3.1 Optical flow method 13 3.2 Horn-Shrunk dense optical flow 14 3.3 Lucas-Kanade sparse optical flow 14 3.4 Establish optical flow histogram 14 3.4 Fast optical flow histogram descriptor 16 4. Building Visual Vocabulary 19 4.1 Generate visual words 19 4.2 k-means mean clustering 19 4.3 Hierarchical k-means hierarchical clustering 20 4.4 Fisher Vector Fisher 21 5.Video Classification 24 5.1 Introduction of Support Vector Machines 25 6. Experiment 25 "6.1 Experimental setup 26 4 6.2 Experimental results 27 6.2.1 The establishment of visual dictionary and the impact of video classification27 6.2.2 Impact of video sampling rate 28 6.2.3 Influence of optical flow algorithm 28 6.2.4 Selection of the best solution 29 7. Project Management 30 7.1.Considerations 30 7.2 Budget monitoring 30 7.3 Hardware budget 30 7.4 Software budget 30 7.5 Total budget 31 Conclusion 32 Acknowledgement 33 Reference 34 Code 34 5 1.Introduction 1.1 The main challenge in gesture recognition Today, one of the main challenges of gesture recognition is the processing of large numbers of data sets. Because with the rapid development of the Internet, more and more data will appear in the form of video, which led to the already just for the image data, some applications have been unable to solve some problems. Because, although some theoretical methods have a good performance in the field of static images, since the video data contains complex information and the information is in the process of changing, resulting in the calculation of the processing of video data is quite large and is quite Therefore, when these methods are applied to the posture recognition problem composed of video, the result is not satisfactory, and there are some specific and new problems. Therefore, in order to solve the special problems in the attitude recognition, the face of the system needs to deal with a large number and its number is still growing in the video data, to explore and find those who do not enter the accuracy of the performance of the method, There is also a significant increase in the speed of the method is very necessary. 1.2 Research status at home and abroad In this section, we mainly introduce the characteristics of the description, introduce their domestic and international development. Today, the most commonly used local descriptors (on time and space spans) are descriptive features based on scale-invariant feature transformations. The establishment process is as follows: Each local video sequence will be divided into different partitions, in each partition we will partition all the pixels in the calculation of the optical flow or the direction of the gradient results together, get a total of this partition A (flow of light or direction) to calculate the results. And the final descriptor 6 is the adjacent, a specific number of partition results through a certain way to get together. Respectively in 2006 and 2008, the gradient direction histogram descriptor and the optical flow histogram descriptor on the two-dimensional plane are proposed. In addition, Dalal also presents a descriptor obtained by calculating the optical flow variation: the motion boundary histogram feature. Scovanner and Kläser, respectively, in 2007 and 2008, have proposed on the basis of two-dimensional plane, while the time dimension also established the direction gradient descriptor theory, which has been three-dimensional The descriptor. In 2013, the three-dimensional spatial description sub-theory was extended in terms of image channels. But then found that the fact that the proposed descriptor is not better than the direction of the histogram description sub-effect is better. We can see that some points of interest selection theory and some time-space feature descriptors are evaluated. He found that intensive sampling theory in general would be better than the way of selecting points of interest, especially for some of the more complex and difficult to identify the data set. Recently, the theory of video-intensive trajectory is proposed. In this theory he suggests that local video information moves over time, and the result is that, even if the time changes, the above-mentioned local information is still trying to remain in the same part of the target object. In addition, they also use the flow of light rather than light flow as the basis for characterization. Their work in the basic gradient direction histogram, optical flow histogram and action boundary histogram characterization descriptor made some improvements. However, combining their new tracing descriptive characteristics with traditional gradient directions, optical flow histograms, and motion boundary histogram characterization descriptors will actually result in a better effect than using a tracking feature descriptor alone. It is proposed to use the integral graph to effectively calculate the characteristics of the accelerated file in the static image. And put forward some in the image classification problem, to accelerate the classification of word bag classification of some of the theory and their performance made a very detailed assessment. In recent years, in the application of the word pocket model, the Fisher vector has 7 shown a better performance than the standard classification theory (such as k-means mean clustering). And the recently proposed VLAD descriptor local descriptor aggregation is seen as a non-probable version of the Fisher's vector. 1.2 The main content of this article The main objective of this paper is to obtain a posture recognition system that achieves a relative optimal result (or a balance between the two) with the efficiency of the calculation and the accuracy of the results. Around this purpose, we used a fast word bag model as the basis of the pipeline, at different stages of the program, the use of image feature extraction, visual vocabulary classification and classification, video data classification and identification of a variety of methods, and They were evaluated and compared separately. First of all, according to some of the characteristics of video data, we have adopted a number of methods, we hope to improve the speed of the system: For example, we have how to focus on the video data sampling method to do a certain study: We try to change the sampling rate for video frames Improve the efficiency of the method. Secondly, since the optical flow histogram feature is used as the descriptor, we will also pay attention to the selection of the calculation method of the optical flow. We use and evaluate the different optical flow algorithm. At the same time, we try to build a dense, accelerated optical flow histogram descriptor on the basis of the underlying optical flow histogram descriptor. Furthermore, on the basis of the existing phonetic model, we try to find and evaluate a combination of the most suitable for video classification to achieve the above-mentioned automatic recognition system with both accuracy and computational efficiency. 8 In addition, we used and evaluated the effect of using Fisher's vector theory and kmean, hierarchical k-means hierarchical clustering in video classification. Finally, we use different classifiers to match the previous steps to achieve the best video classification results. 2 Bag-of-Word Model In the field of attitude recognition based on visual features, the emergence of a model occupies the mainstream: the word pocket model. The word pocket model has been proven in a number of experiments that he is the most effective strategy for general classification strategies. We can derive this conclusion from the excellent performance of the various types of mainstream evaluation systems over the past few years, such as the TRECVID advanced feature extraction task in video data. In these tasks, conceptual descriptive features can be detected and assigned to different categories such as chairs, cats, cars, and the like. Based on the excellent performance of the word pocket model, this text is based on this classic model as our basic design pattern. But the basis of the performance of the word pocket model still can not achieve our goal, its huge amount of calculation can not be ignored defects. So on this basis, we have improved the classic word bag model, so that it is not only in the efficiency and accuracy are reached a higher level, this strategy we will be detailed after the introduction. 2.1 Introduction of Bag - of - Word Model The name of the phonetic model is actually very vivid, that is, with a group of "words" to express our extracted features. In the beginning, there is actually a question about the text retrieval problem. The specific process is as follows: First, a dictionary containing all the words in the training set text is constructed according to 9 the training set text. From another perspective, As a fixed-length vector, because the words in the text will be repeated, we will be able to each text in which the frequency of the word that, so that each text can be used to express the fixed-length vector, through training these data we will You can sort the text to be tested. Similarly, in the word pocket model, we use a similar strategy: in accordance with the data itself will be divided into training set and test set, we will train all the characteristics of the data set together, and apply a certain way to them The automatic classification (the total number of categories is determined by the user), the resulting classification center points represent each category, so that we get an array of size-specific information that contains the center point information to represent the visual word in our word pocket model. Then we will focus on the video of the video with the visual word in the dictionary, generate each video fixed-length vector used to train the support vector machine model. After completing the training, and then take the test set video for video classification. The basic word pocket model contains the four main steps: 1. Video extraction features. 2. Create a visual dictionary. 3. Generate the video histogram. 4. Video classification. The process is as follows: ! 16 ! figure 3.1 3.4 Fast optical flow histogram descriptor In this chapter, we will focus on a fast optical flow histogram descriptor we built. Unlike the feature information in the image, we made the corresponding improvement according to the characteristics of the video. In order to completely extract the video stream histogram features, the main need to complete the following steps: Before we start feature extraction, we have a problem to consider: There are many frames in the video. To calculate the flow of each frame, it is conceivable that the computation is huge. At the same time, however, the information contained in the adjacent video sequence is mostly duplicated. Thus, some frames in the video can be extracted by frame sampling instead of calculating the number of frames, such as every two, three, or six frames, and the sample rate is According to the video itself, determined by experiment. By taking this approach, the computational efficiency of the feature extraction phase system is significantly increased. After completing the preprocessing step above, you can start the feature extraction. First, it is necessary to calculate the optical flow of each frame in the frame in units of frames. In the first step, it is necessary to calculate the optical flow 17 displacement vector in the horizontal direction and the vertical direction in the twodimensional image of each frame. Then, on this basis, in order to establish the optical flow histogram, we need to quantize the amplitude of the optical flow obtained for each pixel, and generally o = 8. In this way, the optical flow of each pixel can be represented by a vector whose direction represents the instantaneous motion of the pixel and its absolute value represents the instantaneous velocity of the pixel. The next step is to combine the results of each pixel, according to the time and space set in the block together, so we get each piece of light flow results; this process we will be detailed in the back , Which uses a coefficient matrix to calculate. The last step is to associate the adjacent blocks according to the pre-set size, and the resulting optical flow histogram results together to obtain the optical flow histogram characterization of the video. Among them, the optical flow calculation part, the default Horn-Shrunk method used to calculate the optical flow of each pixel, in addition to the Lucas-Kanade method used to calculate the optical flow. In this paper, the optical flow histogram characterization descriptor is constructed by using block as the basic unit. So, how to select the video in the block is a feature extraction in a key point. In this paper, we choose to construct the feature descriptor with the same step size as a single block size, that is, we can block these characters can be reused by different feature descriptors, so that each block Information has been calculated in advance, then the calculation efficiency will be significantly improved. Once the result of each block is calculated, the descriptor can be obtained by concatenating a certain number of adjacent blocks of the optical flow histogram. In this paper, we use each frame to occupy 3 * 3 blocks in space, and occupy two blocks in time to form an optical flow histogram characterization descriptor. Of course, the size of this feature descriptor is not necessary, it should be based on the specific data set to make the appropriate changes to get the best results. According to the above method, in addition to the video at the edge of the block, each block will be repeated by the different descriptors 18 times. 18 19 4. Building Visual Vocabulary This chapter will focus on how to build the corresponding visual dictionary with the light stream histogram descriptor after getting the video. The establishment of a visual dictionary consists of two phases: 1. Generating visual words. 2. Using visual words to represent video information features, and further use the image feature histogram to visually reflect this information. 4.1 Generate visual words Generating visual words is a more abstract concept. It is the same dimension of each descriptor in each of the extracted videos, and in fact a video can produce a lot of descriptors and thus can not work concisely and concisely in the subsequent matching work, resulting in the use of a video descriptor After clustering, a number of feature vectors with a number less than the total number of descriptors are generated, but the dimensions are invariant, to effectively respond to the video information. As can be seen from the above description, the choice of clustering is a key step in determining the visual word, and it also determines the speed of the process and the accuracy of the results. Therefore, it is important to choose a suitable word classification method. In this paper, we mainly introduce three methods: 1. k-means clustering 2. Hierarchical k-means hierarchical clustering 3. Fisher Vector. 4.2 k-means mean clustering The goal of k-means mean clustering is to classify the data to be classified into k 20 clusters. First, the first step randomly selected k center points. The next step is to calculate the distance between each point to be classified and all the center points, where the center point closest to the result of the distance to be measured is the classification to be classified. And then calculate the distance between all points to be measured and the center point, to complete the first classification. On the basis of the completion of the first round of classification, according to the class has been divided into another class according to the algorithm to recalculate and select the center of each category and update, and then repeat the second step is once again classified. The above steps are repeated until the center point in the last category does not change, that is, the classification is completed and the final classification result is obtained. Among them, k-means mean clustering mainly faces the following problems: 1. The initial center point selection is random, leading to the final result due to the initial center point has a huge gap. 2. The number of selection k is not certain, the need for users based on data sets and experimental goals to determine. In general, the larger the value of k, the greater the amount of computation, but the results are more accurate (especially for large sample datasets). 4.3 Hierarchical k-means hierarchical clustering As mentioned above, the main problem with k-means mean clustering is that the number of initial center numbers needs to be artificially determined, and finding a suitable k value is more difficult and requires multiple attempts; The initial center of the random selection of the results of the change is difficult to control. In view of these cases, Hierarchical k-means hierarchical clustering shows a better solution. Hierarchical k-means Hierarchical clustering algorithm can be divided into two categories: "bottom-up" and "top-down". Their processes are mainly used in the iterative algorithm. First introduce the "bottom-up" principle: Suppose we have a data set, which contains a point. First, we need to calculate the distance between all the points, get the two points with the smallest distance, divide them into groups, and denote them with a new point. In this way, the number of our data points is reduced to one. Then, repeat the above process 21 until the last only one classification. On the contrary, "top-down" is the opposite process. It will eventually get a tree model results, which contain data depth, node, classification and other information. 4.4 Fisher Vector Fisher As mentioned above, the main problem with k-means mean clustering is that the number of initial center numbers needs to be artificially determined, and finding a suitable k value is more difficult and requires multiple attempts; The initial center of the random selection of the results of the change is difficult to control. In view of these cases, Hierarchical k-means hierarchical clustering shows a better solution. Hierarchical k-means There are two types of hierarchical clustering: "bottomup" and "top-down". Their processes are mainly used in the iterative algorithm. First introduce the "bottom-up" principle: Suppose we have a data set, which contains a point. First, we need to calculate the distance between all the points, get the two points with the smallest distance, divide them into groups, and denote them with a new point. In this way, the number of our data points is reduced to one. Then, repeat the above process until the last only one classification. On the contrary, "top-down" is the opposite process. Hierarchical k-means Hierarchical clustering The final result is a tree model that contains information such as depth, node, classification, and so on. Fischer vector method is mainly based on the principle of Gaussian mixture distribution. First, we use the training set to get all the parameters of the Gaussian mixture distribution, that is to say, through the Gaussian mixture distribution to reflect the training set of video model, this step can be solved by EM method. This process can be further explained by setting the Gaussian mixture distribution model that has been known to have known parameters, substituting the video descriptor to be trained in step E, and obtaining a vector in the M step F, which is a correction value, indicating how it should be updated to apply to this data set. And this vector F is the Fisher Vector calculation results. The next step is to use the square root of the 22 absolute value to normalize the Fisher Vector. Although the Fisher Vector results are much more dimensionally than other methods, the linear classifier is more suitable for the Fisher Vector for subsequent use of the classifier, thus making up for the larger dimension in this step due to multidimensional The amount of calculation. 23 24 5.Video Classification In this chapter, we mainly introduce the process of image classification. In the previous chapter we have obtained a histogram representation of all the images: that is, each image is finally represented by the corresponding fixed length vector. In the classification process, we need to randomly divide all the images into three categories: training set, test set and verification set. However, in the actual situation, the verification set is used less, so we mainly divide the image into training set and test set in this paper. Among them, the number of their own data is determined by the user's own, but in general, the proposed training set and test set data ratio of 3: 2. Of course, there are other special methods such as cross validation, it will turn each data in turn as a test set to experiment. After the above steps are completed, the support vector machine will be used in the classification. First, we need to enter the training set data, support the vector opportunity to learn the characteristics of each image in the training set, and then training the model. After the training model is obtained, the support vector machine will use the obtained model to predict the result of testing the concentrated image, and finally get the classification result. It should be noted that the user needs to adjust the parameters of the model to achieve the optimal results based on each training result in this process. 25 5.1 Introduction of Support Vector Machines In many cases, we will need to attribute it according to the characteristics of a specific thing to classify it as one or more types. In computer science, we call this case a classification problem. SUPPORT VECTOR MEACHINE or support vector machine is a supervised learning model, which is a method of analyzing data and performing pattern recognition. The purpose is to automatically establish a series of rules that do not exist before the data to be classified Classify it. By supervising learning we want to get the result that the system can infer the series of rules mentioned in the previous paragraph through the data that has already been marked. And no supervised learning means that the system has the ability to classify different targets into different clusters according to the similarity without knowing any known classification. The support vector machine belongs to the former, and we need to provide the tagged data All the classification, so that the system has the ability to follow our request to the new unmarked data assigned to the above categories. 6. Experiment In this chapter, we will elaborate on our experimental process and some of the details. In addition, we will briefly introduce the database used in the experiment. The entire experimental process is done through matlab, the main reason for choosing matlab is because it is easy to use the programming language, easy to implement the experimental process, and for image processing, computer vision has a rich toolbox. At the same time, because our program involves a large number of arrays and vector calculations and use, matlab is a better choice. 32 Conclusion This paper mainly uses a kind of excellent performance and robust video characteristics: optical flow histogram characterization descriptor to express video data information. At the same time, in order to make the accuracy and computational efficiency of the system reach a high level, we evaluate and To find a kind of video classification for the most suitable for the word bag model combination. Experiments show that, when the need to extract the optical flow histogram descriptor, every two frames to extract features than each frame are extracted features 1.5-1.7 times faster, but the loss of accuracy at this time can be ignored. When we are demanding the speed of the video classification system, it is possible to consider sampling the video for every six frames to obtain the optical flow histogram characterization descriptor. In addition, for the optical flow histogram descriptor, we find that the selection of the optical flow calculation method has a significant effect on the performance of the final system. Between the best performing Horn-Schunck algorithm and the second best Lucas-Kanade algorithm, the result is that there is a 5% gap. For the construction of visual dictionary this step, the most accurate method is undoubtedly the Fisher method. Hierarchical k-means Hierarchical clustering takes up one-sixth of the time spent by the Fischer vector, but accordingly it sacrifices by a 2% reduction in accuracy. Therefore, if we want to get a more efficient video classification system, we can consider using Hierarchical k-means hierarchical clustering. For support vector machine video classification, the use of cross-histogram kernel support vector machine in the performance of accuracy than the linear support vector 33 machine excellent. But the linear support vector machine is faster than the support vector machine of the crossed histogram kernel. Acknowledgement At the last semester of my college, my student career was spent in Barcelona, Spain. I would like to thank the teacher of the domestic Beihang Quyu Yu teacher, regardless of the pre-identified topics, mid-term update to confirm the progress of my graduation design guidance, take the trouble to answer my questions, answer doubts. Thanks also to Prof. Joan at the Catalan Polytechnic University, who regularly communicates with me to guide me in reporting on the process and content of graduation design and to make improvements to my program's problems and to my interest in exotic life. 34 Reference Code ————————————————————————————————— ————— main.m ————————————————————————————————— ————— %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% %%%%%%%%%%%% % % BOW pipeline: Gesture recognition using HOF % %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% %%%%%%%% % % Part 1: extract descriptors % Part 2: Represent images by histograms of quantized features % Part 3: classification % %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% %%%%%%%%%%%%%%%%% clear; close all; run('/Applications/MATLAB_R2015b.app/toolbox/vlfeat-0.9.20/toolbox/ vl_setup'); %DATASET 35 dataset_dir = 'Gesture'; %FEATURES extraction method %'Horn-Schunck' desc_name = 'hofs'; %FLAGS do_feat_extraction = 1; do_split_sets = 1; do_form_codebook = 1; do_feat_quantization = 1; do_svm_linar_classification = 0; do_svm_intersection_classification = 1; %PATH basepath = '..'; wdir = pwd; libsvmpath = [wdir(1:end-6) fullfile('lib','libsvm-3.11','matlab')]; addpath(libsvmpath); %BOW PARAMETERS nfeat_codebook = 130000;%number of descriptors used by k-means for the codebook generation norm_bof_hist = 1; 36 % number of images selected for training num_train_img = 8; % number of images selected for test num_test_img = 2; % number of codewords (i.e. K for the k-means algorithm) nwords_codebook = 50; %image file extension file_ext = 'avi'; %Create a new dataset split file_split = 'split.mat'; if do_split_sets data = create_dataset_split_structure(fullfile(basepath, 'video', ... dataset_dir),num_train_img,num_test_img,file_ext); save(fullfile(basepath,'video',dataset_dir,file_split),'data') else load(fullfile(basepath,'video',dataset_dir,file_split)); end classes = {data.classname};%create cell array of class name string %%%%%%%%%%%%%%%%%Part 1 %%%%%%%%%%%%%%%%%%% %%%%% %% Load pre-computed HOF features for training images % The resulting structure array 'desc' will contain one % entry per images with the following fields: % desc.row = info.row; %desc.col = info.col; 37 %desc.depth = info.depth; %desc.hofs = hofs; %desc.vidname = videoName; %desc.Size = info.descSize; lasti=1; for i = 1:length(data) images_descs = get_descriptors_files(data,i,file_ext,desc_name,'train'); for j = 1:length(images_descs) f n a m e = fullfile(basepath,'video',dataset_dir,data(i).classname,images_descs{j}); %fprintf('Loading %s \n',fname); tmp = load(fname,'-mat'); tmp.desc.class=i; % tmp.desc.imgfname=regexprep(fname,['.' desc_name],'.jpg'); desc_train(lasti)=tmp.desc; desc_train(lasti).hofs = single(desc_train(lasti).hofs); lasti=lasti+1; end; end; %% Load pre-computed HOF features for test images lasti=1; for i = 1:length(data) images_descs = get_descriptors_files(data,i,file_ext,desc_name,'test'); for j = 1:length(images_descs) f n a m e = fullfile(basepath,'video',dataset_dir,data(i).classname,images_descs{j}); fprintf('Loading %s \n',fname); tmp = load(fname,'-mat'); tmp.desc.class=i; 38 %tmp.desc.imgfname=regexprep(fname,['.' desc_name],'.jpg'); desc_test(lasti)=tmp.desc; desc_test(lasti).hofs = single(desc_test(lasti).hofs); lasti=lasti+1; end; end; %% Part2: Build visual vocabulary using k-means %%%%%%%%%%%%%%% %%%%%%%%%%%%%%%%%%% if do_form_codebook fprintf('\nBuild visual vocabulary:\n'); % concatenate all descriptors from all images into a n x d matrix DESC = []; labels_train = cat(1,desc_train.class); for i=1:length(data) desc_class = desc_train(labels_train==i); randimages = randperm(num_train_img); randimages = randimages(1:5); DESC = vertcat(DESC,desc_class(randimages).hofs); end % sample random M descriptors from all training descriptors r = randperm(size(DESC,1)); %r = r(1:min(length(r),nfeat_codebook)); DESC = DESC(r,:); numClusters = 5; %[centers, assignments] = vl_kmeans(DESC', numClusters,'Initialization', 'plusplus'); 39 % x = 1:length(assignments); data = DESC * 10^4 * 5; data = uint8(data'); nleaves = 100; [tree,A] = vl_hikmeans(data,numClusters,nleaves); % run k-means VC = centers'; clear DESC; end %k-means descriptor quantization if do_feat_quantization fprintf('\nFeature quantization'); quantdist = []; for i=1:length(desc_train) dmat = eucliddist(desc_train(i).hofs,VC); [mv,visword] = min(dmat,[],2); desc_train(i).visword = visword; end for i=1:length(desc_test) dmat = eucliddist(desc_test(i).hofs, VC); [mv,visword] = min(dmat,[],2); desc_test(i).visword = visword; end end 40 %% represent images with BOF histograms %represent each image by the normalized histogram of visual N = size(VC,1); for i=1:length(desc_train) visword = desc_train(i).visword; edges = 1:1:numClusters; Hist = histcounts(visword,numClusters); Hist(1,:) = Hist(1,:)/norm(Hist(1,:),1); desc_train(i).bof = Hist; end for i=1:length(desc_test) visword = desc_test(i).visword; edges = 1:1:numClusters; Hist = histcounts(visword,numClusters); % normalize bow-hist (L1 norm) Hist(1,:) = Hist(1,:)/norm(Hist(1,:),1); % save histograms desc_test(i).bof = Hist; end %%%%%%%%%%%%%%%%% Part 3: image classification %%%%%%%% %%%%%%%%%%%%%%%%%%%% % Concatenate bof-histograms into training and test matrices bof_train=cat(1,desc_train.bof); bof_test=cat(1,desc_test.bof); 41 % Construct label Concatenate bof-histograms into training and test matrices labels_train=cat(1,desc_train.class); labels_test=cat(1,desc_test.class); % LINEAR SVM if do_svm_linar_classification % cross-validation C_vals = log2space(5,10,10); for i=1:length(C_vals); opt_string=['-t 0 -v 5 -c ' num2str(C_vals(i))]; xval_acc(i)=svmtrain(labels_train,bof_train,opt_string); end %select the best C among C_vals and test your model on the testing set. [v,ind]=max(xval_acc); % train the model and test model=svmtrain(labels_train,bof_train,['-t 0 -c ' num2str(C_vals)]); disp('*** SVM - linear ***'); svm_lab=svmpredict(labels_test,bof_test,model); method_name='SVM linear'; % Compute classification accuracy compute_accuracy(data,labels_test,svm_lab,classes,method_name,desc_test,... visualize_confmat & have_screen,... visualize_res & have_screen); end 48 % Also add descriptor sizes to info structure info.descSize = (blockSize .* numBlocks); desc.row = info.row; desc.col = info.col; desc.hofs = hofs; desc.vidname = videoName; desc.Size = info.descSize;