Full text
Efficient Deep Vision for Aerial Visual Understanding Rafael Makrigiorgis,Shahid Siddiqui,Christos Kyrkou,Panayiotis Kolios,Theocharis Theocharides KIOS Research and Innovation Center of Excellence, University of Cyprus, 1 Panepistimiou Avenue, Nicosia, Cyprus e-mails: {makrigiorgis.rafael, siddiqui.muhammad-shahid, kyrkou.christos, kolios.panayiotis,ttheocharides}@ucy.ac.cy. Abstract Unmanned Aerial Vehicles (UAVs) are becoming a growing necessity for a broad range of applications, such as emergency response, monitoring critical infrastructures, and disaster management. UAVs, due to their affordability and camera capabilities, have become a common mobile camera platform for these kinds of applications. Thus visual perception by utilizing Convolutions Neural Networks (CNNs) and Deep Learning is a key necessity for UAV based applications. The remarkable performance of deep neural networks (DNNs) for vision tasks comes at a cost of high computational demands where the problem is amplified in drone based applications due to limited energy resource. To address these drawbacks, this chapter highlights some of the key techniques of making deep vision more efficient for such resource constraint applications. The techniques include but are not limited to data selection and reduction, efficient neural network design, and hardware-oriented model optimization. Results on different use cases show that such techniques can provide improvements either when applied as standalone or in a combined manner. Keywords: Efficient Deep Learning, Object Detection, Visual Disaster Recognition, Deep Neural Networks 1 Introduction Embedded visual AI is a growing trend in applications requiring low latency, realtime decision support, increased robustness and security [14, 15]. An example of this is the increasing use of Unmanned Aerial Vehicles (UAVs) for a number of applications as a remote sensing platform, such as road traffic monitoring [19, 23], search and rescue [25], and precision agriculture [24]. Recent technological advances such as the integration of camera sensors with on-board processing provide the opportunity for new UAV applications such as i) detecting and classifying infrastructure faults during routine inspections, ii) identifying and tracking people of 1
2 Makrigiorgis et al. Fig. 1 Efficient visual understanding for UAVs requires optimizations at different levels, from data reduction, to deep neural network network architecture exploration, and hardware-driven model adjustments. interest, locating objects, and flagging unusual situations, iii) locating people who are lost etc. The aforementioned capabilities are enabled by recent advances in deep-learningbased scene understanding, and convolutional neural networks (CNNs) in particular, that provide remarkable opportunities for computer vision based applications. In such scenarios, the vision algorithms need to deployed on resource-constrained embedded devices on UAVs that support high-resolution cameras for applications beyond photography, such as environmental and infrastructure monitoring. However, deep learning algorithms are computationally very intensive, and not suitable for onboard processing on small devices such as UAVs due to battery/power constraints and are usually processed via the cloud. For certain scenarios however, where the system operates in remote areas with limited connectivity, this can result in unwanted response latency which can degrade performance, a potentially catastrophic scenario in safety-critical applications. Thus, processing information at the edge can not only eliminate unwanted lag issues but also partially handle security problems since the data is not transmitted and sensitive data cannot be intercepted. A lot of research has been conducted to address the challenges of visual perception using UAV imagery. A few notable examples are detecting vehicles for traffic monitoring scenarios in [13] and disaster management in [1]. Processing visual data information has seen significant accuracy and performance improvements due to advancements in deep learning and technologies in Graphical Processing Units (GPUs). However, relying on hardware improvements alone is not by itself sufficient to provide the optimal power/performance trade-offs necessary for edge and IoT applications. Hence, in this chapter we highlight a body of work that aims to explore different techniques that in tandem with embedded hardware improvements can lead to better efficiency for edge applications. The described techniques as depicted in Fig. 1 provide a way towards a more holistic optimization by considering all aspects from the data, to the AI model, as well as hardware aspects. These techniques are presented in the subsequent chapters where improved performance and efficiency are demonstrated compared to standard approaches.
Efficient Deep Vision for Aerial Visual Understanding 3 2 Domain-Specific Small ConvNets for UAV Applications Typically we use hardware accelerators to speedup compressed and quantized versions of well known models because the underlying operations are highly parallel. What if we also try and optimize the model directly for the application by only using the right amounts and types of each operation? In this section we will demonstrate some use-cases where a small neural network is capable of providing adequate performance. Small deep neural networks have many desirable properties and advantages that we can see across their lifetime cycle and make them more easily deployable on embedded processors where computation and even more so memory are at a premium. In addition, being able to store the model on-chip saves on power as it avoids off-chip memory access which consumes order of magnitudes more power. In addition, to inference, small neural networks facilitate faster training iterations and more easily updatable over-the-air (OTA). So whenever a remote system encounters a situation where it is not confident in its predictions or may require retraining then we can send less data even from a lower bandwidth network. Finally, when using smaller neural networks it is easier for multiple vision tasks to run on the same platform e.g., object detection and classification. But it is not only the computational and technical characteristics that need to be considered. Many of the operational characteristics also change when considering edge applications. In most cases these are narrow domain applications that need to recognize a few classes compared to the generic models trained on ImageNet [6], they have requirements for real-time processing so that a decision or action can be taken in reasonable time, and need to operate under limited energy and resources. So for many applications the full capacity of ImageNet pretrained models could be unnecessary which provides opportunities to explore smaller deep learning models in edge applications. The main principles behind the design of smaller models, as shown in Fig. 2, is the balance between the number of parameters, the downsampling rate, and the operations performed within the network. The number of parameters usually impacts how many convolutions we have in the model and this can affect the FLOP demands. Aggressive down sampling can harm the accuracy of a model even though it makes it faster. While the number of channels and type of convolution also impacts the performance. Next, two use cases for UAV applications are presented where smaller neural networks with architectural modifications are capable of providing competitive performance compared to traditional deep learning models. 2.1 Disaster Classification The first use case will demonstrate the use of UAVs for patrolling and automated recognition of disaster events. When a disaster strikes, there is limited time to act, hence, the objective is to develop an automated platform for rapid deployment and
4 Makrigiorgis et al. Fig. 2 Main approaches for producing smaller and more efficient neural networks: (a) Reducing input image resolution. (b) Using efficient operators like depthwise convolution if it is implemented efficiently on the underlying hardware. (c) Early downsampling of feature maps can improve processing speed. Fig. 3 Atrous Convolutional Feature Fusion Block large area coverage. Coupled with an automated path planning software, a UAV can navigate a monitored area and automatically recognize different events through on-board processing of its camera feed. Since connectivity is not guaranteed in such cases, this should be done in a way that is suitable for the embedded hardware of UAVs and in real-time. 2.1.1 Network Design For this application, a small neural network is designed that features a balance between the number of parameters, the downsampling rate, and the operations per-
Efficient Deep Vision for Aerial Visual Understanding 5 formed within the network. First we design a basic block that is suited for processing images where the object/area appears at varying resolutions such as in UAV imagery. Specifically this block is referred to as Atrous Convolutional Feature Fusion (ACFF) block (shown in Fig. 3) that relies on depthwise atous/dilated convolutions to aggregate context information at multiple resolutions. Each dilated convolution is factored into depth-wise convolution that performs light-weight filtering by applying a single convolutional kernel per input channel to reduce the computational complexity. The intuition is to take advantage of the different dilation rates since each path may peek up features at different object/region resolution due to changes in altitude. Another advantage of using dilated convolutions is that the same number of parameters and computations are needed regardless of the size of the kernels in the block. Prior to processing the input feature map, a 1x1 convolution filter is applied to reduce its size and then expand it after the fusion of the dilated convolutions. Other properties of the network geared towards more efficient processing on edge devices are reduced number of 16 channels at the first layer and downsampling with strided convolutions. This is particularly important in achieving better performance since the image resolution at this stage is at the highest. This first convolutional block is a standard convolution. Then the network follows a canonical architecture of ACFF blocks with a progressive reduction of spatial resolution with an increase in depth with up to 256 channels at the last layer. To further reduce the number of parameters, the fully connected layers that are usually employed in classification networks are replaced with global pooling operations. Finally the network depth is maintained at 7 blocks since deeper networks resulted in saturating performance. Additional, modifications were made to the activation function in order to make it more amicable to quantization. First, the leaky variant of the ReLU activation is used that permits for the gradient to flow for negative input values. In addition, the maximum activation is capped to the value of 255 to allow for easier quantization to 8-bits. Finally, we have two operating modes; during training, the full range of the capped ReLU is used, while during inference, the negative values are zeroed. These minor modifications allow optimizing for lower-precision execution. We refer to this architecture as EmergencyNet [18]. 2.1.2 Experiments and Results In this section, the experimental evaluation of small neural network is discussed with results from the experimental evaluation of the approach on an actual embedded platform. First the improvements over existing networks are validated on the developed dataset from [17] and results of the effectiveness are presented. Real experiments have also been conducted in two different settings: (i) On-board embedded processing, where all computations are performed on-board the resource-constrained UAV device, (ii) Remote processing, in which the UAV transmits the captured video to the controller ground station for processing on an Android tablet that controls the UAV. Herein, we focus on the former. It is worth noting that the primary interest is in single image processing speed and as such the evaluation phase is carried out with
6 Makrigiorgis et al. Table 1 Comparison with existing approaches Model Parameters Memory F1 Score FPS (MB) (%) (1/s) VGG16 [30] 14,849,349 59.39 96.4 1.1 ResNet50 [8] 24,113,541 96.4 96.1 2.9 MobileNet V1 [10] 3,492,549 13.9 95 7.9 MobileNet V2 [29] 2,587,205 10.3 95.2 9.3 MobileNet V3 [9] 3,046,037 12.1 95.3 9.0 EfficientNet (B0) [32] 4,378,785 17.5 96.0 10.4 SqueezeNet [11] 698,917 2.7 91.5 6.6 ShuffleNet [22] 4,282,425 17.1 91.1 4.5 Xception [5] 21,387,309 85.549 95.3 2.3 Fire_Net [33] 5,235,860 5.2 90.5 9.8 EmergencyNet [18] 90,892 0.368 95.7 24.3 a batch size of 1since this is common in real-time streaming applications where the camera outputs each frame sequentially. The small neural network is compared with some standard networks and results are summarized in Table 2. First, with regards to the accuracy of the pretrained models, it is observed that VGG16 outperforms all of them with a 96.4% F1 score with ResNet50 closely following with 96.1%. However, both networks have very high demands for computational and storage requirements making them unsuitable for resource constraint systems and real-time use. The latest iteration of MobileNet, i.e. V3, achieves the highest accuracy among the MobileNet family of networks with a score of 95.3%. However, it requires an order of magnitude more parameters and memory. Other MobileNet versions (V1 and V2) demonstrate similar score with the V2 version resulting in a slightly higher FPS. EfficientNet provides a high accuracy of 96.0% due to its elaborate architecture. However, the parameter count and memory requirements are higher than EmergencyNet. Other networks fail to provide adequate accuracy and require a higher number of parameters. It is clear from this analysis that it is worth investigating specifically tailored solutions for resource constraint applications in order to provide an improvement across all design aspects. Furthermore, some classification results are shown in Fig. 4. It is interesting to note that similarly appearing patterns between images do not confuse the network. For instance, the presence of cars, which is more often associated with traffic incidents, does not cause the network to fail which outputs the correct classification as flood. Overall, these results are promising since a small network manages to match or at least be very close to other more general networks while the processing speed and subsequent frame-rate improvements provide adequate trade-offs for edge applications.
Efficient Deep Vision for Aerial Visual Understanding 7 Fig. 4 Example of classification results on sample images 2.2 Vehicle Detection In the second use case, we describe how the problem of top view vehicle detection from UAV aerial imagery is tackled through the exploration of small convolutional neural networks. While there has been quite some research on reducing the complexity of well studied ConvNet models in the form of parameter compression and quantization, there has been little effort on developing specialized solutions for resource constrained embedded vision systems such as UAVs. This section briefly outlines the end-to-end investigation of different single-shot CNN detectors for drone-based vehicle detection. 2.2.1 Network Design and Approach Our approach in [16] focuses on the exploration of an efficient and lightweight network by investigating different network design considerations. The goal is to keep the accuracy as high as possible but design a faster model. We develop several models to explore the effect of different number of parameters. The models are trained to detect only 1 class, which, in our case is vehicles viewed from top. We explore the impact on performance by changing the structure of the network such as the number of filters, the number of layers, image size, the number of convolution and the pooling layers. The design approaches considered are the following: 1) Number and size of filters; Firstly, we use pooling layers sparingly and use a small number of filters in each layer in order to get a smaller network which leads to a faster detection. The largest number of filters in a network is 256 for the final convolutional layer. As the network goes deeper, we double the number of filters. In total, there are 9 convolutional and 4 max-pooling layers. 2) Input Image Size; Input size is a
8 Makrigiorgis et al. parameter that can increase accuracy but decreases the detection speed since a deeper network will be required to reach the final bounding box regression size. Moreover, a larger input image implies more convolutions per feature map and in some cases more bounding boxes to account for. Four network architectures are derived (SmallYoloV3, TinyYoloVoc, TinyYoloNet, and DroNet). These have different parameters including the layers and the type of each layer (conv, maxpool, detection) together with the configuration of the layers in terms of the number of filters, the size of filters in each layer and the input/output size of the feature maps. A common theme however, is the use of 3Γ3 and cheaper 1Γ1convolutional filters, as well as the progressive reduction of the feature maps size by a factor of 2 and use of a lower number of filters at early layers. Finally, feature shortcut connections are used to further improve accuracy for small objects. To find the CNN that optimises both accuracy and computation cost, a custom metric is employed. Given a model instance, it captures both the detection accuracy and the achieved runtime on the target hardware platform. It is defined as a composite linear combination metric that combines various performance indicators. By following this methodology, the resulting detection CNN yields the highest performing balance between detection accuracy and faster execution. πππππ(π€)=π€1ΓπΉππ +π€2ΓπΌππ +π€3Γππππ ππ‘ππ£ππ‘π¦ +π€4Γππππππ πππ (1) The score is parameterized with respect to a vector of weights which sums to one, where each weight captures the application-level importance of each metric and is normalized between [0-1]. Since real-time performance is desired, the FPS is prioritized with a weight of 0.4 over other three accuracy-related metrics which are equally weighted with 0.2. 2.2.2 Results In this section, we present a comprehensive quantitative evaluation of the four CNN architectures. The basic network models are trained and tested for various input sizes using the constructed vehicle dataset. Fig. 2.2.2 shows the performance comparison of these models. In our test set, with 386 Γ386 as input resolution, TinyYoloNet achieves 10Γhigher performance than TinyYoloVoc with decreased detection sensitivity and precision by 20% and 10% respectively, and an IoU drop of 0.11. The network SmallYoloV3, with 386 Γ386 resolution achieves the highest frame-rate of 23 FPS among all network designs. Nevertheless, the substantial reduction in the number of weights leads to a decrease in sensitivity which is 53% lower, and prohibits us from using it for robust vehicle detection. There are significant performance gains starting from TinyYoloVoc to TinyYoloNet followed by SmallYoloV3 and then to DroNet. For example, comparing these models for the same input size of 386, the performance of DroNet is 30Γfaster compared to TinyYoloVoc with a minimal drop
Efficient Deep Vision for Aerial Visual Understanding 9 Table 2 Comparison of the different models Model TinyYOLOVoc TinyYOLONet SmallYOLOv3 DroNet FPS 1.2 12 23 22 IoU 0.36 0.25 0.11 0.22 Sensitivity(%) 87.8 68.2 34.7 69.4 Presicion (%) 90.5 80.6 69.1 76.5 Size (MB) 63.1 6.4 0.115 0.283 Score 0.62 0.68 0.69 0.83 Fig. 5 Example of detection results on sample images of 0.08 on the IoU. Moreover, there is a limited drop of 2% and 6% for the detection sensitivity and precision, respectively. Comparing DroNet with different models demonstrates the effect of a composite score function. In particular, with respect to the smaller smallYoloV3 network, it provides lesser FPS while being more accurate. On the other hand, compared with the largest tinyYoloVOC, it is less accurate but much faster. The model size is also suitable for edge applications with limited resources. Overall, it provides the best trade-off between accuracy and performance. The analysis is performed on the Odroid-XU4 with an Octacore Samsung Exynos5422 CPU which is lightweight and capable of being powered by the UAV platform. Overall, the DroNet network maintains a performance of 10 FPS with the accuracy maintained around 95% for 512 Γ512 image resolution. 3 Processing Aerial Images with Tiling The previous section dealt with design of small, domain-specific deep neural networks. However, optimizing the network alone might not be enough to obtain a high performing system. In many cases, when a UAVs flies at a high altitude (e.g., 500 meters), and covers a wide field of view, the image resolution must be large enough to recognize targets. High-resolution images, however, imply an exponential increase in the amount of data processed and most of the times, there is a lot of redundant data. Hence, even a small neural network would require significant time to process such
16 Makrigiorgis et al. Fig. 9 Comparison of average processing time (CPU) and sensitivity between different EdgeNet configurations for different time frames for each stage. This figure was taken from our previous work in [27] EdgeNet-1-3-5 DroNet_Tile DroNet_V3 Tiny-YoloV3 RP3 1.5 1.7 2 2.4 ODROID 3.6 3.8 4 5 CPU 9 10 15 23 Table 4 Average Power Consumption measured in Watts which are pre-trained on massive datasets perform worst with regards to power and processing time. Moreover, upon observing the results from the sensitivity, as seen in Figure 10, we can see that π·πππππ‘_ππππ even though it comes second best in terms of power consumption and inference speed, it comes last in terms of sensitivity while ππππ¦ β πππππ3comes second. On the other hand, πΈππππππ‘ achieves a 6% higher sensitivity than ππππ¦ βπππππ3which is the largest model of all of them. πΈππππππ‘ β1β3β5 sensitivity keeps the accuracy close to 96% compared to others due to the fact that single shot models lose object resolution and features when resizing the image that leads to a decreased accuracy. Overall, by reducing the processed data, having smaller input sized CNNs and utilizing the tracker it directly impacts the processing time which leads to reduction of computation power on all platforms. Therefore, πΈππππππ‘ is a framework capable of providing real-time processing pipeline for mobile/edge devices in terms of accuracy, inference speed and power consumption.
Efficient Deep Vision for Aerial Visual Understanding 17 Fig. 10 Sensitivity of TinyYoloV3, π·πππππ‘_π3,π·πππππ‘_ππππ and EdgeNet on different platforms. This figure was taken from our work in [27] EdgeNet-1-3-5 DroNet_Tile DroNet_V3 Tiny-YoloV3 RP3 0.06 0.09 0.22 0.85 ODROID 0.05 0.1 0.3 1 CPU 0.02 0.03 0.08 0.59 Table 5 Average Processing time on different platforms 4 Combining Tiling with Quantization Embedded deployment of neural networks (NNs) are constrained by limited amount of memory, compute, power budget and real time latency requirements. As such, relying on small models, and data reduction technqiues might not be enough. Hence, a range of hardware-driven neural network optimization techniques have been proposed in the literature for efficient deployment. Among those, quantization of neural networks is of special interest since many embedded platforms support integer only arithmetic with INT8 quantization [12]. These include but are not limited to some variants of ARM Cortex-M, RISC-V GAP-8 a system-on-chip, and Googleβs Edge TPU. Quantization in general, is a method to map from input values in a large (often continuous) set to output values in a small (often finite) set e.g. rounding and truncation [7]. In the context of neural networks (NNs), quantization allows network weights and activation functions to be converted from floating point operations to either fixed point or mixed precision operations hence reducing modelβs memory
18 Makrigiorgis et al. footprint and RAM consumption and improving latency and power consumption. Using TVM quantization library [3], [20] has shown 3.89Γ, 3.32Γ, and 5.02Γspeed up for ResNet50 [8], VGG-19 [30], and inceptionV3 [31] respectively. Existing object detectors increase inference efficiency by operating on aggressively downsized, lower resolution images which causes significant information loss and hence performance degradation for small object detection. In addition, quantizing an already resized image further reduces detection accuracy. Therefore, in addition to using quantization for faster inference, an additional mechanism for maintaining high image resolution is desired for highly accurate multi-scale object detection. Our work in section 3 proposes a mechanism to intelligently select and process more relevant portions (tiles) of a high resolution image hence maintaining high detection accuracy while still improving the inference times. Thus combining the aforementioned techniques of small network design, with image data reduction, and network quantization can potentially lead to further improvements. 4.1 Quantization Techniques The way an NN is quantized can vary in many aspects. A quantization is said to be uniform if all quantization levels are equally spaced, and non-uniform otherwise. Similarly a quantization is said to be symmetric if the clipping range of the signal is centered at zero, and asymmetric otherwise. Moreover, if the clipping range of NN activations are pre-computed, the quantization is said to be static. Dynamic quantization is also possible and generally results in higher accuracy but has higher computational overhead because activation clipping range is calculated for each input during real time inference. Quantization can further be categorized into layerwise, groupwise and channelwise depending upon what set of parameters is being used to estimate the clipping range of the activations. Quantization is also classified by βwhenβ it is performed. The most wide spread trend is post training quantization where the neural network is first trained till convergence and quantized only when ready for deployment. Recently, notable accuracy gains have been reported by using quantization aware training where a network is first trained till convergence, then some of its layers are quantized and the network is fine tuned with either mixed precision or integer only weights. This is intended to make network already aware of it being quantized and finetune accordingly, hence named βquantization aware trainingβ. For a detailed review of quantization, we refer the reader to [7]. 4.2 Approach In this section we investigate a case study for object detection using quantization and selective tiling which was discussed in the previous section. In addition, the
Efficient Deep Vision for Aerial Visual Understanding 19 model under consideration is DroNetV3, which was also introduced in previous section. Hence, we study the impact of quantization and tiling techniques on detection accuracy and latency. As discussed before, aggressive image downsizing leads to poor detection performance in case of smaller objects of interest. In our previous work, [26] shows how to split a larger image into smaller tiles and process only the relevant ones. Since distributing the tiles uniformly across image leads to a drastic increase in the number of sub-images to be processed by the CNN, two statistical techniques are proposed to select and process only relevant tiles while also keeping track of non-active tiles. The first technique is making use of Intersection over Union (IoU) metric, where each newly detected objectβs bounding box (bbox) is compared to detected boxes in the previous frames and classified as either new or old object. The position of the objects in previous frames is maintained in a memory buffer. The second technique uses statistical metrics such as number of objects detected in each tile and prioritizes tiles with higher number of detected objects as well those that have not been selected recently to be processed in the subsequent frames. These techniques combined, allow an intelligent selection of image regions to be fed into a smaller and faster object detector thus keeping higher resolution intact leading to higher accuracy and real-time performance due to selective processing. A post training quantization of 8-bit is applied for both the input and the selected convolutional neural network detector. With the use of quantization, multiply-add operations are transformed into lower-precision operations which lead to large computational gains and higher performance. This implementation is based on the Darknet framework and CUDA-based Neural Network framework. The combined tiling and quantization approach can then be analyzed with respect to resulting acurracy and processing time for the application of people detection from UAV. 4.3 Experimental Results In case of YOLOV3 and Tiny-Yolov3 speed-ups up to 1.4β1.7Γare observed. However, with DroNet and DroNetV3, i.e. relatively smaller networks, the speed-up is limited to 1β1.4Γ. This is due to significant performance overhead caused by the additional processing time for input image quantization as compared to network inference time. The rest of the experiments are conducted using DroNetV3 which achieved highest speed-up using input resolution of 352 Γ352. Next, we evaluate combinations of quantization with different tiling schemes and report accuracy, IoU and average processing time (APT) metrics. All networks are trained on a UAV-based people detection dataset consisting of 1500 images and a total of 60000 people. Table 6 shows the impact of resizing, processing a single, all or only selected tiles [26] and applying quantization to all base cases. Resizing the input causes 20% accuracy drop and quantization on top drops another 8%. This result is expected due to information loss of input image and model weights respectively. However,
20 Makrigiorgis et al. Table 6 Results of quantization and tiling with resizing, single tile selection, all tiles selection and intelligent tile selection. Models DroNetV3 DroNetV3 + Resizing DroNetV3 + Single Tile DroNetV3 + All Tiles DroNetV3 + Selective Tiling Accuracy 98.929 79.643 92.143 99.464 96.071 APT 0.432 0.089 0.087 0.680 0.132 IoU 0.643 0.415 0.496 0.650 0.563 Quantized Models DroNetV3 DroNetV3 + Resizing DroNetV3 + Single Tile DroNetV3 + All Tiles DroNetV3 + Selective Tiling Accuracy 100 71.429 94.643 98.750 98.571 APT 0.315 0.071 0.068 0.574 0.098 IoU 0.648 0.312 0.494 0.619 0.562 resizing speeds up inference significantly i.e. an APT of 0.432π to 0.089π . Since tiling maintains high resolution, all techniques based on tiling improve accuracy as compared to DroNetV3 with resizing. Even when models are quantized, accuracy increases from 71% to 94% β98% i.e. an increase of 23% to 27%. Since the selective tile processing mechanism is processing 2β3per frame, the processing time is higher than either resizing or single tiling. However, quantizing this network improves latency from 0.132π to 0.098π without accuracy degradation. A similar trend has been observed for IoU metric i.e., avoiding resizing the image in general leads to higher IoU. Quantization reduces the IoU from 0.415 to 0.312 for DroNet with Resizing and its quantized version. However, combining quantization and tiling does not degrade IoU. Overall, using a combination of quantization and tiling allows processing higher resolution images hence improving latency as compared to original full precision detector and accuracy as compared to using resizing. 5 Conclusion Deep learning and computer vision are increasingly being utilized in edge applications to provide real-time intelligence. This chapter has demonstrated various techniques and approaches to apply at various stages of visual processing in order to make algorithms more suitable for the unique demands of aerial image understanding with UAVs and embedded application domains. The exploration of small neural network design considerations and architectures can be effective in deploying such models on resource constraint devices. Coupled with search reduction strategies and hardware-directed optimizations the performance of such small models can be further improved leading to real-time power efficient solutions. As a future work, we aim to further improve the accuracy-performance trade off by investigating adaptive computation schemes through networks with multiple learnable exit layers, as well as dynamically changing the computation scheme based on context and state variables.
Efficient Deep Vision for Aerial Visual Understanding 21 Acknowledgements The project is co-financed by the European Regional Development Fund and the Republic of Cyprus through the Cyprus Research Innovation Foundation (βRESTART 20162020β Program) (Grant No. INTEGRATED/0918/0056) (RONDA). This work was also supported by the European Unions Horizon 2020 research and innovation programme under grant agreement No 739551 (KIOS CoE) and from the Government of the Republic of Cyprus through the Directorate General for European Programmes, Coordination and Development. References 1. Robert Allison, Joshua Johnston, Gregory Craig, and Sion Jennings. Airborne optical and thermal remote sensing for wildfire detection and monitoring. Sensors, 16(8):1310, Aug 2016. 2. Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017. 3. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. Tvm: An automated end-to-end optimizing compiler for deep learning. In OSDI, 2018. 4. Xueyun Chen, Shiming Xiang, Cheng-Lin Liu, and Chun-Hong Pan. Vehicle detection in satellite images by parallel deep convolutional neural networks. In 2013 2nd IAPR Asian conference on pattern recognition, pages 181β185. IEEE, 2013. 5. F. Chollet. Xception: Deep learning with depthwise separable convolutions. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1800β1807, 2017. 6. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A largescale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248β255, 2009. 7. Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. CoRR, abs/2103.13630, 2021. 8. K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770β778, 2016. 9. Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam. Searching for mobilenetv3. In The IEEE International Conference on Computer Vision (ICCV), October 2019. 10. Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. CoRR, abs/1704.04861, 2017. 11. Forrest N. Iandola, Matthew W. Moskewicz, Khalid Ashraf, Song Han, William J. Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <1mb model size. CoRR, abs/1602.07360, 2016. 12. Andrey Ignatov, Grigory Malivenko, and Radu Timofte. Fast and accurate quantized camera scene detection on smartphones, mobile ai 2021 challenge: Report. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 2558β2568, June 2021. 13. Eui-Jin Kim, Ho-Chul Park, Seung-Woo Ham, Seung-Young Kho, and Dong-Kyu Kim. Extracting vehicle trajectories using unmanned aerial vehicles in congested traffic conditions. Journal of Advanced Transportation, 2019:1β16, 04 2019. 14. Christos Kyrkou. CΛ3Net: End-to-end deep learning for efficient real-time visual active camera control. Journal of Real-Time Image Processing, 18(4):1421β1433, Aug 2021. 15. Christos Kyrkou, Eftychios G. Christoforou, Stelios Timotheou, Theocharis Theocharides, Christos Panayiotou, and Marios Polycarpou. Optimizing the detection performance of smart
22 Makrigiorgis et al. camera networks through a probabilistic image-based model. IEEE Transactions on Circuits and Systems for Video Technology, 28(5):1197β1211, 2018. 16. Christos Kyrkou, George Plastiras, Theocharis Theocharides, Stylianos I. Venieris, and Christos-Savvas Bouganis. Dronet: Efficient convolutional neural network detector for realtime uav applications. In 2018 Design, Automation Test in Europe Conference Exhibition (DATE), pages 967β972, 2018. 17. Christos Kyrkou and Theocharis Theocharides. Deep-learning-based aerial image classification for emergency response applications using unmanned aerial vehicles. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 517β 525, 2019. 18. Christos Kyrkou and Theocharis Theocharides. Emergencynet: Efficient aerial image classification for drone-based emergency monitoring using atrous convolutional feature fusion. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13:1687β1699, 2020. 19. Christos Kyrkou, Stelios Timotheou, Panayiotis Kolios, Theocharis Theocharides, and Christos G. Panayiotou. Optimized vision-directed deployment of uavs for rapid traffic monitoring. In 2018 IEEE International Conference on Consumer Electronics (ICCE), pages 1β6, 2018. 20. Wuwei Lin. Automating optimization of quantized deep learning models on cuda. 2019. 21. Bruce D Lucas, Takeo Kanade, et al. An iterative image registration technique with an application to stereo vision. Vancouver, 1981. 22. Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Computer Vision β ECCV 2018, pages 122β138, Cham, 2018. Springer International Publishing. 23. Rafael Makrigiorgis, Nicolas Hadjittoouli, Christos Kyrkou, and Theocharis Theocharides. Aircamrtm: Enhancing vehicle detection for efficient aerial camera-based road traffic monitoring. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 3431β3440, 2022. 24. D. Murugan, A. Garg, and D. Singh. Development of an adaptive approach for precision agriculture monitoring with drone and satellite data. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 10(12):5322β5328, Dec 2017. 25. P. Petrides, C. Kyrkou, P. Kolios, T. Theocharides, and C. Panayiotou. Towards a holistic performance evaluation framework for drone-based object detection. In 2017 International Conference on Unmanned Aircraft Systems (ICUAS), pages 1785β1793, June 2017. 26. George Plastiras, Christos Kyrkou, and Theocharis Theocharides. Efficient convnet-based object detection for unmanned aerial vehicles by selective tile processing. In Proceedings of the 12th International Conference on Distributed Smart Cameras, ICDSC β18, New York, NY, USA, 2018. Association for Computing Machinery. 27. George Plastiras, Christos Kyrkou, and Theocharis Theocharides. Edgenet: Balancing accuracy and performance for edge-based convolutional neural network object detectors. In Proceedings of the 13th International Conference on Distributed Smart Cameras, pages 1β6, 2019. 28. Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 29. M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L. Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4510β4520, 2018. 30. Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015. 31. Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818β2826, 2016. 32. Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 6105β6114. PMLR, 2019.
Efficient Deep Vision for Aerial Visual Understanding 23 33. Yi Zhao, Jiale Ma, Xiaohui Li, and Jie Zhang. Saliency detection and deep learning-based wildfire identification in uav imagery. Sensors, 18(3), 2018.