Full text
Universidad de Málaga Escuela Técnica Superior de Ingeniería de Telecomunicación TESIS DOCTORAL Optimization of Mobility Parameters using Fuzzy Logic and Reinforcement Learning in Self-Organizing Networks Autor: Pablo Muñoz Luengo Directora: Raquel Barco Moreno
AUTOR: Pablo Muñoz Luengo http://orcid.org/0000-0002-3265-5728 EDITA: Publicaciones y Divulgación Científica. Universidad de Málaga Esta obra está bajo una licencia de Creative Commons Reconocimiento-NoComercialSinObraDerivada 4.0 Internacional: http://creativecommons.org/licenses/by-nc-nd/4.0/legalcode Cualquier parte de esta obra se puede reproducir sin autorización pero con el reconocimiento y atribución de los autores. No se puede hacer uso comercial de la obra y no se puede alterar, transformar o hacer obras derivadas. Esta Tesis Doctoral está depositada en el Repositorio Institucional de la Universidad de Málaga (RIUMA): riuma.uma.es
E.T.S.I. Telecomunicación, Campus de Teatinos, 29071-MÁLAGA, Tlf. 952 13 14 40 Fax 952 13 20 27 Departamento de Ingeniería de Comunicaciones Dra. Dª Raquel Barco Moreno, profesora doctora del Departamento de Ingeniería de Comunicaciones de la Universidad de Málaga CERTIFICA: Que D. Pablo Muñoz Luengo, Ingeniero de Telecomunicación, ha realizado en el Departamento de Ingeniería de Comunicaciones de la Universidad de Málaga, bajo su dirección el trabajo de investigación correspondiente a su TESIS DOCTORAL titulada: Optimization of Mobility Parameters using Fuzzy Logic and Reinforcement Learning in Self-Organizing Networks En dicho trabajo se han propuesto aportaciones originales para diversos problemas de optimización en redes móviles, en particular, el balance de carga mediante movilidad, la optimización del procedimiento de traspaso así como su coordinación con el balance de carga y, por último, el direccionamiento de tráfico en escenarios de redes heterogéneas. También se ha propuesto un modelo de abstracción para un simulador de la tecnología celular LTE que ha permitido la evaluación y validación de las técnicas propuestas. Los resultados expuestos han dado lugar a publicaciones en revistas y aportaciones a congresos internacionales. Por todo ello, considera que esta Tesis es apta para su presentación al Tribunal que ha de juzgarla. Y para que conste a efectos de lo establecido en el Artículo 8º del Real Decreto 778/1998, Real Decreto 56/2005 y Real Decreto 1393/2007, reguladores de los Estudios de Tercer Ciclo-Doctorado, AUTORIZA la presentación de esta Tesis en la Universidad de Málaga. Málaga a _______ de ___________________ de ________ Fdo: Dr a . D ª . Raquel Barco Moreno
UNIVERSIDAD DE MÁLAGA ESCUELA TÉCNICA SUPERIOR DE INGENIERÍA DE TELECOMUNICACIÓN Reunido el tribunal examinador en el día de la fecha, constituido por: Presidente: Dr. D. Secretario: Dr. D. Vocales: Dr. D. Dr. D. Dr. D. para juzgar la Tesis Doctoral titulada Optimization of Mobility Parameters using Fuzzy Logic and Reinforcement Learning in Self-Organizing Networks realizada por D. Pablo Muñoz Luengo y dirigida por la Dra. D.aRaquel Barco Moreno, acordó por otorgar la calificación de y para que conste, se extiende firmada por los componentes del tribunal la presente diligencia. Málaga a de del El Presidente: El Secretario: Fdo.: Fdo.: El Vocal: El Vocal: El Vocal: Fdo.: Fdo.: Fdo.:
A mis padres, mi hermano y Mari Ángeles.
“Men build too many walls and not enough bridges.” Isaac Newton.
Abstract In the last years, cellular networks have experienced a large increase in size and complexity. As a result, mobile operators have focused attention on reducing capital expenditures (CAPEX) and operational expenditures (OPEX) of their networks. This fact has stimulated strong research activity in the field of Self-Organizing Networks (SON), which is a set of principles and concepts defined for automating network management while improving network quality. Increasing automation in cellular networks means significant OPEX reduction as fewer personnel are necessary to maintain a network. In addition, network performance is enhanced as the number of human errors is diminished, also taking advantage of the analytical capabilities of computers to introduce more efficient procedures in the network. The main benefits derived from applying SON functionalities are especially valuable in the Radio Access Network (RAN) segment, since this part of the network usually forms the bottleneck due to its operational complexity and network costs. In particular, the SON functions are classified into three terms: Self-Configuration, Self-Healing and Self-Optimization. The first term, Self-Configuration, attempts to automate network deployment and parameter configuration, e.g. when a new resource is added to the existing infrastructure. Self-Healing is related to failure detection, diagnosis, compensation and recovery in order to cope with major service outages and degradations. Lastly, Self-Optimization aims to dynamically adapt network parameters to improve network quality. In the context of Self-Optimization, certain functions have been identified as key enablers in the literature, among which are Mobility Load Balancing (MLB) and Mobility Robustness Optimization (MRO). The former is an automated function where cells suffering occasional congestion can transfer load to neighboring cells, which have spare resources, by adjusting mobility parameters. The latter is a solution for automatic detection and correction of errors and suboptimal settings in the mobility configuration, which may lead to a degradation of user performance. In addition, due to the fact that these two functions can adjust the same mobility parameters, a conflict may happen if MLB and MRO tune the parameters in opposite directions. The coordination of these two functions is also an important issue in the context of SON. On the other hand, the large investments on infrastructure to cope with the growing demand of traffic has led to a heterogeneous network deployment characterized by the presence of different technologies, cell sizes, frequencies, etc. However, these solutions are expected to be insufficient, v
meaning that the existing networks will also play a key role in facing that enormous traffic demand. In this context, Traffic Steering (TS) becomes a powerful mechanism in which user distributions across the different operator’s networks are modified according to a policy or specific objectives, mainly related to cost, user satisfaction, power consumption, coverage, etc. As a result of applying those TS techniques, the overall network performance is improved. In this thesis, several optimization techniques for next-generation wireless networks are proposed to solve different problems in the field of SON and heterogeneous networks. The common basis of these problems is that network parameters are automatically tuned to deal with the specific problem. As the set of network parameters is extremely large, this work mainly focuses on parameters involved in mobility management. In addition, the proposed self-tuning schemes are based on Fuzzy Logic Controllers (FLC), whose potential lies in the capability to express the knowledge in a similar way to the human perception and reasoning. In addition, in those cases in which a mathematical approach has been required to optimize the behavior of the FLC, the selected solution has been Reinforcement Learning, since this methodology is especially appropriate for learning from interaction, which becomes essential in complex systems such as wireless networks. Taking this into account, firstly, a new MLB scheme is proposed to solve persistent congestion problems in next-generation wireless networks, in particular, due to an uneven spatial traffic distribution, which typically leads to an inefficient usage of resources. A key feature of the proposed algorithm is that not only the parameters are optimized, but also the parameter tuning strategy. Secondly, a novel MLB algorithm for enterprise femtocells scenarios is proposed. Such scenarios are characterized by the lack of a thorough deployment of these low-cost nodes, meaning that a more efficient use of radio resources can be achieved by applying effective MLB schemes. As in the previous problem, the optimization of the self-tuning process is also studied in this case. Thirdly, a new self-tuning algorithm for MRO is proposed. This study includes the impact of context factors such as the system load and user speed, as well as a proposal for coordination between the designed MLB and MRO functions. Fourthly, a novel self-tuning algorithm for TS in heterogeneous networks is proposed. The main features of the proposed algorithm are the flexibility to support different operator policies and the adaptation capability to network variations. Finally, with the aim of validating the proposed techniques, a dynamic system-level simulator for Long-Term Evolution (LTE) networks has been designed. vi
Resumen En los últimos años, las redes celulares han experimentado un gran crecimiento tanto en tamaño como en complejidad. Como resultado, los operadores móviles han centrado su atención en reducir los gastos de capital (CAPEX) y operacionales (OPEX) de sus redes. Este hecho ha suscitado una fuerte actividad investigadora en el campo de las redes auto-organizativas (SelfOrganizing Networks, SON), las cuales presentan una serie de principios y conceptos definidos para automatizar la gestión de redes a la vez que se mejora su calidad. La automatización de las redes móviles implica una reducción significativa de los gastos operacionales, debido a que se requiere menos personal para mantener y administrar la red. Además, se mejoran las prestaciones de la red puesto que se disminuye el número de errores humanos y se aprovechan las ventajas de las capacidades de cálculo de los computadores para introducir procedimientos más eficientes en la red. La aplicación de funcionalidades SON permite obtener grandes beneficios, especialmente en la red de acceso radio (Radio Access Network, RAN), debido a que esta parte de la red normalmente es el “cuello de botella” dada su complejidad desde el punto de vista operacional y los costes asociados. Las funciones SON se clasifican en tres categorías: auto-configuración, autocuración y auto-optimización. La primera categoría, auto-configuración, trata de automatizar el despliegue de la red y la configuración de parámetros, por ejemplo, cuando un nuevo elemento de red se añade a la infraestructura existente. La auto-curación está relacionada con la detección, diagnosis, compensación y recuperación de fallos con el objetivo de hacer frente a cortes y degradaciones de servicio importantes. Por último, la auto-optimización trata de adaptar dinámicamente los parámetros de red para mejorar su calidad. Dentro de la categoría de auto-optimización, ciertas funciones han adquirido gran relevancia en la bibliografía, entre las que se encuentran el balance de carga mediante movilidad (Mobility Load Balancing, MLB) y la optimización de la movilidad (Mobility Robustness Optimization, MRO). La primera de ellas, MLB, es una función automática en la que las celdas que sufren congestión pueden transferir, mediante un correcto ajuste de los parámetros de movilidad, parte de la carga a sus celdas vecinas, las cuales deben disponer de suficientes recursos libres. MRO es una función que permite detección y corrección automática de errores y ajustes sub-óptimos en la configuración de movilidad, los cuales pueden llevar a una degradación de las prestaciones de usuario. Además, debido a que estas dos funciones pueden ajustar los mismos parámetros de vii
movilidad, podría existir un conflicto si MLB y MRO ajustan dichos parámetros en direcciones opuestas. Por tanto, la coordinación de ambas funciones es también una cuestión importante en el contexto de SON. Por otro lado, la enorme inversión en infraestructura realizada por los operadores para satisfacer la creciente demanda de tráfico ha dado lugar al despliegue de redes heterogéneas, las cuales están caracterizadas por la presencia de diferentes tecnologías, tamaños de celdas, frecuencias de uso, etc. Sin embargo, es probable que estas soluciones sean insuficientes, de manera que las redes que ya han sido desplegadas con anterioridad también jueguen un papel esencial para hacer frente al fuerte crecimiento de la demanda de tráfico. En este contexto, el direccionamiento de tráfico (Traffic Steering, TS) es un potente mecanismo consistente en modificar la distribución de usuarios en las diferentes redes pertenecientes al operador con la intención de satisfacer una determinada política u objetivo, principalmente relacionado con los costes, la satisfacción de usuario, el consumo de potencia, la cobertura, etc. Como resultado de aplicar estas técnicas, se mejoran las prestaciones globales de red. En la presente tesis se proponen diversas técnicas de optimización para redes inalámbricas de próxima generación con el objetivo de resolver diferentes problemas en el campo de SON y redes heterogéneas. El fundamento básico de todas las técnicas propuestas es que los parámetros de red son ajustados automáticamente para resolver un problema específico. Puesto que el conjunto de parámetros de red existente es muy grande, este trabajo se centra fundamentalmente en aquellos parámetros que forman parte de la gestión de movilidad. Además, los esquemas de auto-ajuste propuestos están basados en controladores de lógica difusa (Fuzzy Logic Controller, FLC), cuyo potencial radica en la capacidad de poder expresar el conocimiento de una manera similar al razonamiento y la percepción humana. Por otra parte, en aquellos casos en los que se requiere una herramienta matemática para optimizar el comportamiento del FLC, en esta tesis la solución adoptada ha sido el uso de aprendizaje por refuerzo, debido a que esta metodología es especialmente apropiada para el aprendizaje mediante interacción, el cual se hace esencial en sistemas complejos como es el caso de las redes inalámbricas. Teniendo en cuenta esto, en primer lugar, se ha propuesto un esquema para MLB con la finalidad de resolver problemas de congestión persistente en redes de inalámbricas de próxima generación, en particular, debidos a una distribución espacial de tráfico desigual, la cual típicamente conlleva un uso ineficiente de los recursos. Una característica clave del algoritmo propuesto es que no solo se optimizan los parámetros, sino también la propia estrategía de ajuste de parámetros. En segundo lugar, se ha propuesto un esquema para MLB aplicable en escenarios de femtoceldas de oficina o corporativos. Estos escenarios se caracterizan por la falta de un riguroso análisis en la etapa de despliegue, de manera que es posible realizar un uso más eficiente de los recursos radio en la etapa de explotación mediante la aplicación de técnicas efectivas de MLB. Al igual que en el problema anterior, en este caso la optimización del proceso de auto-ajuste también forma parte del estudio. En tercer lugar, se ha propuesto un algoritmo de auto-ajuste para MRO. Este estudio incluye un análisis del impacto de factores contextuales tales como la carga del sistema y la velocidad de usuario, así como una propuesta para la coordinación de los esquemas de MLB y MRO diseñados. En cuarto lugar, se ha propuesto un algoritmo de auto-ajuste para TS en el contexto de redes heterogéneas. Las principales características del algoritmo propuesto son la viii
flexibilidad para soportar diferentes políticas del operador y la capacidad de adaptación a variaciones en la red. Finalmente, con el objetivo de validar las técnicas propuestas, se ha diseñado un simulador dinámico de nivel de sistema para redes Long-Term Evolution (LTE). ix
x
List of Contributions The following list presents the publications related to this thesis. Journals Arising from this thesis [I] P. Muñoz, R. Barco, D. Laselva and P. Mogensen. Mobility-based Strategies for Traffic Steering in Heterogeneous Networks. IEEE Communications Magazine, vol. 51, no. 5, pp 54-62, May 2013. [II] P. Muñoz, R. Barco and I. de la Bandera, Optimization of Load Balancing using Fuzzy Q-Learning for Next Generation Wireless Networks, Expert Systems With Applications (Elsevier), vol. 40, no. 4, pp 984-994, March 2013. [III] P. Muñoz, D. Laselva, R. Barco and P. Mogensen. Adjustment of mobility parameters for traffic steering in multi-RAT multi-layer wireless networks, EURASIP Journal on Wireless Communication and Networking, 2013:133, May 2013. [IV] P. Muñoz, R. Barco and I. de la Bandera. On the Potential of Handover Parameter Optimization for Self-Organizing Networks. IEEE Transactions on Vehicular Technology, accepted in 2013. [V] P. Muñoz, R. Barco, J. M. Ruiz, I. de la Bandera and A. Aguilar. Fuzzy Rule-based Reinforcement Learning for Load Balancing Techniques in Enterprise LTE Femtocells. IEEE Transactions on Vehicular Technology, accepted in 2012. [VI] J. M. Ruiz, S. Luna-Ramírez, M. Toril, F. Ruiz, I. de la Bandera, P. Muñoz, R. Barco, P. Lázaro and V. Buenestado, Design of a Computationally Efficient Dynamic System-Level Simulator for Enterprise LTE Femtocell Scenarios, Journal of Electrical and Computer Engineering, vol. 2012, Dec. 2012. xi
[VII] P. Muñoz, I. de la Bandera, F. Ruiz, S. Luna-Ramírez, R. Barco, M. Toril, P. Lázaro and J. Rodríguez, Computationally-Efficient Design of a Dynamic System-Level LTE Simulator, International Journal of Electronics and Telecommunications, vol. 57, no. 3, pp 347-358, Sep. 2011. Related to this thesis [VIII] R. Barco, P. Lázaro and P. Muñoz. A Unified Framework for Self-Healing in Wireless Networks. IEEE Communications Magazine, Vol.50 (12), pp.134-142. Dec. 2012. Conferences and Workshops Arising from this thesis [IX] P. Muñoz, I. de la Bandera, R. Barco, M. Toril, S. Luna-Ramírez and J.M. Ruiz, “Sensitivity Analysis and Self-Optimization of LTE Intra-Frequency Handover”, 6th Scientific Meeting, COST action IC1004, Málaga (Spain), February 2013. [X] P. Muñoz, R. Barco, I. de la Bandera, M. Toril and S. Luna-Ramírez, “Optimization of a Fuzzy Logic Controller for Handover-based Load Balancing”, Int. Workshop on SelfOrganising Networks (IWSON), IEEE Vehicular Technology Conference (VTC) Spring, Budapest, Hungary, May 2011. [XI] P. Muñoz, I. de la Bandera, R. Barco, F. Ruiz, M. Toril and S. Luna-Ramírez, “Estimation of Link-Layer Quality Parameters in a System-Level LTE Simulator”, 5th International Conference on Broadband and Biomedical Communications (IB2COM) 2010. Málaga (Spain). December, 2010. [XII] P. Muñoz, I. de la Bandera, R. Barco, M. Toril and S. Luna-Ramírez, “Optimización del balance de carga en redes LTE mediante el algoritmo de Q-Learning difuso”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (Spain), September, 2011. [XIII] I. de la Bandera, P. Muñoz, R. Barco, M. Toril and S. Luna-Ramírez, “Auto-ajuste del margen de handover en redes LTE”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (Spain), September, 2011. [XIV] P. Muñoz, I. de la Bandera, R. Barco, F. Ruiz, M. Toril and S. Luna-Ramírez, “Diseño del Nivel de Enlace para un Simulador LTE”, XXV Simposio de la Unión Científica Internacional de Radio (URSI 2010), Bilbao (Spain). September, 2010. xii
Related to this thesis [XV] J.M. Ruiz, S. Luna-Ramírez, M. Toril, F. Ruiz, I. de la Bandera and P. Muñoz, “Analysis of Load Sharing Techniques in Enterprise LTE Femtocells”, Int. Workshop on Femtocells, Wireless Advanced 2011, London, UK, June 2011. [XVI] J. Rodríguez, I. de la Bandera, P. Muñoz and R. Barco, “Load Balancing in a Realistic Urban Scenario for LTE Networks”, Int. Workshop on Self-Organising Networks (IWSON), IEEE Vehicular Technology Conference (VTC) Spring, Budapest, Hungary, May 2011. [XVII] J. Rodríguez, I. de la Bandera, P. Muñoz and R. Barco, “Balance de carga en red LTE en un entorno urbano realista”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (Spain), September, 2011. [XVIII] G. Jiménez, R. Barco, P. Muñoz and M. Toril, “Optimización del Umbral de Traspaso para Femtoceldas LTE”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (Spain), September, 2011. [VI, VII, XI, XIV] are devoted to the simulation tools that allow the performance of the proposed algorithms to be evaluated. [II, X, XII, XVI, XVII] are devoted to the load balancing problem in macrocell scenarios, while that problem in femtocell scenarios is addressed in [V, XV, XVIII]. The problem of handover optimization is tackled in [IV, IX, XIII]. [VIII] is devoted to Self-Healing in the field of Self-Organizing Networks. In [I, III], the proposed techniques for Traffic Steering in the context of Heterogeneous Networks are described. The author has been the primary author of all the contributions “arising from this thesis” except [VI, XIII], being there a key contributor. In [VI], the author wrote the section dedicated to the indoor mobility model and helped with the design of the link layer. In [XIII], the author was involved in the design, analysis and writing stages. Finally, the author helped with the writing and review of the papers “related to this thesis”. Several research projects have been involved in these contributions. In particular, the following publications were developed in the frame of these projects: •P08-TIC-4052 grant from the Junta de Andalucía: [I-VII, IX-XVIII]. •TEC2009-13413 grant from the Spanish Ministry of Science and Innovation: [II, V-VII, IX-XIV, XV, XVIII]. •IPT-2011-1272-430000 grant from the Spanish Ministry of Science and Innovation: [V]. In addition, [I, III] were also developed in collaboration with Nokia Siemens Networks in Aalborg, Denmark. xiii
PRB Physical Resource Block PUCCH Physical Uplink Control Channel PUSCH Physical Uplink Shared Channel QCI QoS Class Identifier QHLB Q-Learning Handover Margin Load Balancing QoS Quality-of-Service QPLB Q-Learning Power Load Balancing QualHO Quality Handover RACH Random Access Channel RAN Radio Access Network RAT Radio Access Technology RL Reinforcement Learning RLC Radio Link Control RLF Radio Link Failure RNC Radio Network Controller RR Round Robin RRC Radio Resource Control RRM Radio Resource Management RSCP Received Signal Code Power RSRP Reference Signal Received Power RSRQ Reference Signal Received Quality RSSI Received Signal Strength Indicator S-GW Serving Gateway SC-FDMA Single-Carrier Frequency Division Multiple Access xx
SINR Signal to Interference-plus-Noise Ratio SIR Signal-to-Interference Ratio SON Self-Organizing Network SWG Sub-Working Group TDD Time-Division Duplex TS Traffic Steering TTI Transmission Time Interval TTT Time-To-Trigger TXP Transmit Power UCI Uplink Control Information UDR User Dissatisfaction Ratio UE User Equipment UEE Unbalanced Exploration and Exploitation UL-SCH Uplink Shared Channel UMA University of Málaga UMTS Universal Mobile Telecommunications Service US Uncorrelated Scattering VH Very High VL Very Low VN Very Negative VP Very Positive VoIP Voice over Internet Protocol WCDMA Wideband Code Division Multiple Access WG Working Group xxi
WiMAX Worldwide Interoperability for Microwave Access WLAN Wireless Local Area Network WSS Wide-Sense Stationary WSSUS Wide-Sense Stationary Uncorrelated Scattering Z Zero xxii
Chapter 1 Introduction This chapter aims to explain the purpose of this thesis and its motivation, present its objectives and describe the document structure. 1.1 Motivation In the last few years, mobile networks have significantly changed the way people communicate. In particular, there has been a clear shift from fixed to mobile cellular telephony and the mobile industry has experienced an enormous growth both in terms of mobile technology and its subscribers. In 2011, global penetration of mobile-cellular technology reached 87% with almost six billion mobile-cellular subscriptions. Furthermore, mobile-broadband subscriptions have grown 45% annually over the last years and currently, there are twice as many mobile-broadband as fixed-broadband subscriptions [1]. Since the first generation of cellular networks was introduced in the 1980s, the traffic demand has grown dramatically. Due to this, mobile operators have been investing heavily in infrastructure upgrades in order to support such a traffic demand. The first generation, based on analog telecommunications standards, provided the basic mobile voice. This generation continued until being replaced by the second generation, known as Global System for Mobile communications (GSM). The second generation was characterized by introducing capacity and coverage. Later, the third generation known as Universal Mobile Telecommunications Service (UMTS) provided data services at higher speeds as a first approach to mobile-broadband experience, which will be further developed by the fourth generation. Currently, migration towards the fourth generation of mobile networks is considered one of the main goals of mobile operators. Such a migration, also known as Long-Term Evolution (LTE), will provide access to a wide range of telecommunication services, most of them supported by both mobile and fixed networks. In addition, many 1
CHAPTER 1. INTRODUCTION networking technologies will also be available to enable true ubiquitous mobile access, which means that many technologies will become available to provide different wireless solutions (e.g. wireless sensor networks, enhanced third generation) and networking protocols will connect users to the best available network. Furthermore, mobile terminals (or smartphones) will benefit from major advances in technology developments in areas such as semiconductors, nanotechnology, processing power and storage capacity, thus enabling the emergence of smaller, complex and intelligent devices. The problem of the increasing traffic demand in current cellular networks is further aggravated by considering financial constraints from the operator perspective, as higher capacity and enhanced user experience should be provided at the expense of higher capital expenditures (CAPEX) and operational expenditures (OPEX). Since users may be reluctant to pay proportionally higher bills for improved services, minimizing CAPEX and OPEX while providing better user experience and capacity becomes a crucial consideration for mobile operators. For this reason, operators attempt to achieve a trade-off between providing improved services and retaining reasonable profits to make the business model commercially viable. To cope with the enormous increase in traffic demand in a cost effective way, new technology should not only increase the speed of data transfer in the network, but also increase automation of network management. Such a requirement has triggered research to add intelligence and autonomy to future cellular networks, resulting in a new paradigm known as Self-Organizing Networks (SONs) [2]. The development of SON techniques is motivated by several factors [3]. Firstly, due the mobility of users and the unpredictable nature of the radio channel, cellular networks do not fully utilize the available resources, thus leading to potential congestion situations and suboptimal performance. Secondly, the recent deployment of smaller cells (e.g. femtocells) and outdoor relays in order to provide more capacity has brought a significant increase in the number of network elements, making configuration and maintenance tasks more tedious. In this case, classic manual and field trial based design approaches may lead to suboptimal solutions. This is of special interest in femtocell scenarios, as this type of cells is typically deployed without following a rigorous planning approach. This practice can lead to a waste of network resources or, even worse, it can cause severe interference to neighboring cells of higher sizes, thus degrading the overall performance. Thirdly, with the classic approach for periodic manual optimization, the increasing complexity of future systems (e.g. in technology, protocols) will result in higher number of human errors, which in turn will result in longer recovery and restoration times. Finally, an important benefit derived from automation of network management is that human effort can be freed from mundane and tedious tasks, so that expertise can be focused on new areas, bringing additional value to the operator. Thus, introducing automation in cellular networks means significant OPEX reduction and improved customer experience. To achieve this, the SON proposal aims to enable a set of functionalities for automated management of cellular networks, so that human intervention is minimized in the planning, deployment, optimization and maintenance tasks. In 2006, the Next Generation Mobile Networks (NGMN) Alliance identified excessive dependence on manual operational effort as a potential problem [4]. With the aim of increasing operational efficiency, the 2
1.1. MOTIVATION new paradigm SON was translated into particular functionalities, which later were grouped by the 3rd Generation Partnership Project (3GPP) into three categories, Self-Configuration, SelfOptimization and Self-Healing [5]. In particular, Self-Configuration refers to the dynamic ‘plug and play’ behavior of newly deployed base stations, which configure radio parameters (e.g. the transmission frequency and power) in an autonomous manner, leading to faster cell planning and rollout. Self-Optimization aims to dynamically adapt network parameters to improve network quality, including optimization of coverage, capacity, handover and interference. Self-Healing is related to failure detection, diagnosis, compensation and recovery in order to cope with major service outages and degradations which significantly decrease user performance. Not all of the SON functionalities considered by the 3GPP have received the same attention in the research field. Moreover, certain functions in Self-Optimization have been identified as key enablers, among which are those addressed in this thesis, Mobility Load Balancing (MLB) and Mobility Robustness Optimization (MRO) [6]. This fact has stimulated research activity in the field of parameter self-tuning for both MLB [7][8][9][10][11] and MRO [12][13][14][15][16]. MLB is a function where cells suffering occasional congestion can transfer load to neighboring cells, which have spare resources. This is achieved by adjusting mobility parameters in order to relieve the congestion in the affected area. MLB includes load reporting between base stations to exchange information such as cell load levels and available capacity. The development of MLB algorithms is motivated by some issues. Firstly, in the case of intra-system load balancing, when a user is handed over to another cell, the propagation conditions could get worse if the user is located near the cell edge because of the interference caused by neighboring cells. Thus, the degradation of signal quality for these users should be considered in the process. Secondly, in the case of inter-system load balancing, a proper interface to exchange load information in a heterogeneous scenario and additional intelligence to compare and weigh the capacities of the different radio technologies should be necessary. Thirdly, although a handover due to load balancing is carried out as a regular handover (i.e. the procedure that preserves the connection when the user moves around the network), it may be necessary to amend parameters so that the mobile terminal does not return to the congested cell. Such a correction must take place in both cells, so that the handover settings remain coherent in both sides. MRO is a solution for automatic detection and correction of errors in the mobility configuration, which may lead to a degradation of user performance. In particular, this function mainly focuses on errors causing dropped calls due to handovers that are carried out too late or early when the user moves between cells. MRO must also take into account handovers performed to a wrong cell. As LTE is being deployed with a frequency reuse of one (i.e. the same frequency is shared by all cells), the inter-cell interference becomes an important issue, especially in the cell-edge. Such interference typically results in dropped calls and failures related to handovers. Supposing that this issue is counteracted or at least reduced, another challenging issue is to cope with unnecessary handovers, among which are the so-called ping-pong handovers. A ping-pong handover occurs when a user is handed over to a cell and then it comes back to the first cell in a short time interval. Finally, another interesting challenge is the coordination with MLB, which can cooperate on the problem of a high number of dropped calls, since in the intra-frequency case, MLB may lead to an increase in the number of dropped calls. 3
CHAPTER 1. INTRODUCTION On the other hand, the large increase in both size and complexity experienced by cellular networks has led to a heterogeneous environment characterized by the presence of networks with different cell sizes, technologies, carrier frequencies, etc. In this context, the coverage area of these networks is usually overlapped. Thus, operators have some degree of freedom to steer traffic across the networks with the aim of optimizing the usage of radio resources. Some early studies in this field are [17][18][19][20][21]. The potential of traffic steering in heterogeneous networks brings new challenges to operators since the task of selecting a more suitable network can be made to achieve a goal or a combination of goals, for instance, related to cost, user satisfaction, power consumption, coverage, securing Quality-of-Service (QoS) for the requested service, etc. In practice, the operator needs to decide which performance indicators derived from network measurements must be optimized to improve some aspect of the network. Because of the lack of studies related to this new paradigm, this thesis is also devoted to traffic steering in heterogeneous networks, focusing on the development of automated techniques based on adjustment of mobility parameters. 1.2 Preliminaries The Mobile Network Optimization team [22], belonging to the Ingeniería de Comunicaciones Group (TIC-102), is a research group dedicated to the improvement of current and future mobile telecommunications networks. The group was created in 2003 by six associate professors from the Communications Engineering Department of the University of Málaga (UMA). The aim was to take advantage of the know-how acquired in a 4-year contract with Nokia Networks to develop a mobile system engineering centre at the Parque Tecnológico de Andalucía. Since then, the group has largely grown, having participated in several projects in regional, national and international calls. In addition, the group has worked in projects with the main international operators and vendors: Nokia-Siemens, Ericsson, Alcatel-Lucent, Telefónica, Orange-France Telecom, etc. The group is one of the pioneers in SON, having started to work in optimization and autodiagnosis in second generation networks even before the term SON was created. The CELTICEUREKA project called Gandalf (CP2-014 EUREKA/GANDALF: Monitoring and self-tuning of radio resource management parameters in a multi-system network, 2005-2007), in which the group participated, is a reference as one of the first international projects in the topic. In particular, the aim of Gandalf was to develop self-optimization techniques for networks with multiple radio access technologies and auto-diagnosis tools. The partners were France Telecom R+D, Ericsson R+D Ireland, Moltsen Intelligent SW, Telefónica R+D and University of Limerick. In 2009, the Junta de Andalucía funded a research project of the Communication Engineering Department named “Adaptive techniques for radio resource management in Beyond Third Generation networks”, which has been in place during five years and has supported this PhD. The main lines of this project are the development and evaluation of algorithms for self-tuning of joint radio resource management parameters, the development and evaluation of optimization algorithms for the self-tuning parameter process and the impact assessment of these proposed techniques on the global QoS. The main simulation tool used in this thesis has also been deve4
1.3. RESEARCH OBJECTIVES loped in the context of this project. Other research projects closely related with thesis are the project entitled “Self-optimizing heterogeneous mobile radio access networks”, funded by the Spanish Ministry of Science and Innovation, and the project “Localization and self-organization in femtocell environments (MONOLOC)”, funded by the Spanish Ministry of Science and Innovation and the European Regional Development Fund. One of the lines of these two projects in common with this thesis is the development of self-tuning algorithms for adjusting radio resource management parameters in LTE femtocells. Finally, among the current projects of this team, the most important ones are three contracts with Ericsson to develop SON techniques and tools for current and future mobile communication networks (2012-1015). It is foreseen that more than 40 engineers are going to be hired within this project. 1.3 Research objectives The main goal of this thesis is the design of optimization techniques for mobile networks, by self-tuning network parameters. Such a goal is applied to different problems: •A self-tuning method to solve persistent congestion problems in next-generation wireless networks, in particular, due to an uneven spatial traffic distribution, e.g. when the center of a city becomes crowded. Following the 3GPP guidelines, in which the load balancing problem is referred to as MLB, mobility parameters will take precedence for the self-tuning mechanism. In addition, the optimization of the self-tuning process must be addressed. Such an optimization will provide certain adaptation capacity to the self-tuning algorithm. Hence, not only the parameters, but also the parameter tuning strategy will be optimized. •A scheme to solve persistent congestion problems in enterprise femtocells scenarios, in which the lack of a thorough deployment of these low-cost nodes calls for algorithms that make a more efficient use of radio resources. The self-tuning process will be characterized by tuning specific parameters of femtocells, such as the transmit power or handover margins. In addition, as in the previous problem, the optimization of the self-tuning process must also be studied. •A proposal for handover optimization in next-generation wireless networks. The aim of this function, referred to as MRO in 3GPP terminology, is to reduce the number of dropped calls and avoid unnecessary handovers. In this study, the impact of context factors such as the system load and user speed must be addressed. Furthermore, the need for an optimization of this process, which depends on the complexity of the proposed self-tuning algorithm, must also be considered. •A proposal for SON coordination between the proposed MLB and MRO functions. Since these two functions adjust the same parameters, a method for conflict avoidance must be 5
CHAPTER 1. INTRODUCTION proposed. •A self-tuning algorithm for traffic steering in heterogeneous networks, which eases operators the flexibility to choose a specific policy. The scenario under study must be representative of a realistic network deployment that combines the presence of different technologies, cell sizes, frequencies, hotspots, etc. deployed in a rational manner. In addition to the flexibility to support different operator policies, the adaptation capability to network variations must also be a key feature of the proposed scheme. With the aim of validating the proposed techniques, a simulation tool for next-generation wireless networks tailored to the proposed methods and procedures will also be necessary. The design of such a simulator will therefore constitute a key part of this thesis. 1.4 Document structure This thesis is divided into those problems to be solved, and they will be treated independently. The structure of this document reflects that division in separated chapters, except for the first two problems, which share the same topic (i.e. load balancing) and thus, they have been included into the same chapter. For an easier understanding, the different problems have been treated with an unified structure. This document consists of eight chapters. The first chapter corresponds to this introduction to the thesis, which gives a general view of automated network management and main use cases. Chapters 2 and 3 include the required background to follow the rest of the report. In particular, Chapter 2 provides an introduction to the SON paradigm and a review of the state-of-the-art of the problems and techniques addressed in this thesis, while Chapter 3 provides a review of the theoretical basis of the mathematical techniques suitable for parameters self-tuning and its optimization. In Chapter 4, the simulation tools used to validate the proposed algorithms are described, focusing on the dynamic system-level LTE simulator, whose design has been part of this thesis. The following chapters, which are devoted to the design of self-tuning schemes, have a similar structure, beginning with a brief introduction of the problem to be solved, then, describing the proposed scheme as a solution to the problem in an additional section, and, finally, presenting the main results of the analysis. More specifically, Chapter 5 is dedicated to the load balancing problem in both macrocell and enterprise femtocell scenarios, for which self-optimizing schemes are proposed. In Chapter 6, a self-tuning algorithm for handover optimization and a method for coordination between load balancing and handover optimization are proposed. Chapter 7 is devoted to the development of traffic steering techniques in heterogeneous networks. It is worth noting that, although the previous chapters are mainly focused on LTE networks, the proposed techniques are also valid for other existing and future cellular networks. Finally, Chapter 8 summarizes the main conclusions of the research and future lines of action are proposed. This report also includes as appendix A a brief summary of the thesis in Spanish. 6
Chapter 2 Self-Organization in Radio Access Networks This chapter introduces the key concepts involved in the automatic management of current and future Radio Access Networks (RANs). In the last years, network operators have focused attention on reducing CAPEX and OPEX of the cellular networks. An effective manner to achieve this objective is by means of automatic management, which allows to remove several human interventions from network operation and maintenance. In this area, the concept of SONs involves different mechanisms and methods to enhance the operations of complex networks. As a result of applying SON techniques, not only the manual intervention is minimized, but also network efficiency and service quality are increased. Those benefits derived from SON functionalities are especially valuable in the RAN segment, since this part of the network usually forms the bottleneck in wireless networks due to its operational complexity and network costs. The potential of SON techniques can be further exploited if they are applied in the context of heterogeneous networks (HetNets). HetNets involve a diverse network deployment characterized by the presence of networks with different architectures, technologies, frequencies, cell sizes, elements, etc., whose objective is to cope with the increasing traffic demand which cannot be served by conventional cellular networks. In this context, since the coverage area of those networks are typically overlapped, mobile users can be steered to a particular network in order to satisfy different operator policies related to costs, user satisfaction, power consumption, coverage, securing QoS for the requested service, etc. Those actions, known as Traffic Steering (TS), can be seen as SON mechanisms if they are automatically executed with the aim of both reducing operational costs and enhancing network performance. This chapter is divided into two parts. For clarity, the first part begins by describing the RANs that are assumed in this work, focusing on LTE, which plays an important role in the 7
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS a Physical Resource Block (PRB). A PRB occupies 180 kHz in the frequency domain and 0.5 ms in the time domain [30]. As previously mentioned, downlink transmission in LTE is based on OFDM, which makes use of a large number of closely spaced orthogonal subcarriers that are transmitted in parallel. In particular, a PRB comprises 12 subcarriers at a 15 kHz spacing. Each subcarrier is modulated at a low symbol rate by using a conventional modulation scheme, e.g. QPSK, 16-QAM or 64-QAM. The combination of all the subcarriers in the time domain creates an OFDM symbol, which is then transmitted after including guard intervals to prevent inter-symbol interference at the receiver. As a result, the data-rate generated by using OFDM is similar to conventional single-carrier modulation schemes for the same bandwidth. On the one hand, OFDM provides significant benefits when compared to WCDMA systems. For example, multi-path propagation is a typical phenomenon commonly found in cellular environments. In this context, codes in WCDMA are no longer orthogonal and interfere with each other resulting in inter-user and inter-symbol interference. Such a problem is more pronounced for larger bandwidths (e.g. 10 and 20 MHz required for support of higher data rates). Conversely, OFDM provides the greatest advantage over WCDMA networks in those situations, since a cyclic extension of the OFDM signal is used to avoid this interference. More specifically, the last part of the OFDM signal is added as cyclic prefix in the beginning of the OFDM signal. Some other advantages are that OFDM can easily be scaled up to wide channels (which are more resistant to fading), OFDM is more attractive for Multiple-Input and Multiple-Output (MIMO) antenna configurations, and channel equalizers are much simpler to implement in LTE than in WCDMA, since the channel for each OFDM subcarrier experiences almost flat-fading. On the other hand, OFDM also has some disadvantages. The subcarriers are closely spaced, meaning that OFDM is more sensitive to frequency errors, phase noise and Doppler effect. In addition, as the OFDM symbol is formed by simply adding the modulated subcarrier signals, this signal has much larger signal amplitude variations than the individual subcarriers. Such a characteristic of the OFDM signal creates high peak-to-average signals, which is the main reason why SC-FDMA is used in the uplink. Thus, in the uplink, SC-FDMA is implemented to avoid the high PAPR associated with OFDM. Such a different transmission scheme offers not only the low PAPR from techniques of single-carrier transmission systems, but also the multi-path resistance and flexible frequency allocation from OFDMA. The main difference between these two schemes is that OFDMA transmits the data symbols in parallel, one per subcarrier, while SC-FDMA transmits them in series at a higher rate and also occupying more bandwidth. Therefore, visually, the OFDMA signal is clearly multi-carrier with one data symbol per subcarrier, while the SC-FDMA signal is more similar to a single-carrier in which the data symbol is represented by one wideband signal. As previously stated, in OFDMA, the transmission in parallel of multiple symbols causes the undesirable high PAPR. Conversely, in SC-FDMA systems, by transmitting Ldata symbols in series at a rate Ltimes higher, the occupied bandwidth is the same as in OFDMA but the PAPR is the same as that used for the original data symbols, regardless of the value of L. Thus, by using SC-FDMA in the LTE uplink, low signal peaks comparable to OFDM signal peaks can be achieved. It is also worth mentioning that SC-FDMA is resistant to multi-path when the data symbols are still short. In particular, to provide resistance to delay spread, it is necessary 14
2.1. PRELIMINARIES that each subcarrier is maintained constant in the frequency domain during the symbol period. In SC-FDMA, even though the symbols are not constant in the time domain over the symbol period, the associated spectrum remains constant. This is because the time-varying SC-FDMA symbol composed of Lserial data symbols is equivalent to Ltime-invariant subcarriers in the frequency domain. As a result, SC-FDMA benefits from multi-path protection despite its short data symbols. One of the main advantages of OFDM lies in its suitability to MIMO application. MIMO is a key technology to increase the capacity of wireless networks, based on the use of multiple antennas at both the transmitter and receiver. More specifically, multiple antennas allow to achieve the diversity gain, since radiated signals will take different physical paths. Such a diversity gain improves the link performance when the channel quality cannot be tracked at the transmitter, as it happens for high mobility UEs. When multiple transmission antennas and only one receiver antenna are implemented in the system (i.e. multiple-input and single-output), the data-rates cannot be increased, but the same data-rates can be supported by using less power. This MIMO configuration can also provide greater robustness of the signal to fading and better performance in low quality conditions. To achieve this, data is sent on both transmitting antennas but coded such that the receiver can identify each transmitter. In addition, higher peak data rates can be achieved by using more complex MIMO schemes in which multiple transmission antennas at the base station are combined with multiple receiver antennas at the UE. In particular, this configuration takes advantage of the MIMO spatial multiplexing, where multiple data stream are transmitted between the base station and the UE. All these characteristics highlight that MIMO is an important part of modern wireless communication standards, such as LTE, IEEE 802.11n (Wi-Fi) and Worldwide Interoperability for Microwave Access (WiMAX). The LTE network architecture is designed to support packet-switched traffic with seamless mobility, QoS and low latency. While the RAN segment is divided into the Node B and the RNC in UMTS/HSPA networks, the architecture is greatly simplified in LTE (see Fig. 2.2), resulting in a new network element, called evolved Node B (eNB), which provides the user plane and control plane protocol terminations toward the UE [5]. Thus, the RNC is removed from the network architecture and its functionality is shared between the eNB and the mobility management entity/serving gateway (MME/S-GW), located in the core network. The eNBs are connected between each other by means of the X2 interface and they can also be connected to the core network by means of the S1 interface. The following functions are covered at the eNB: •Radio resource management. •Internet Protocol (IP) header compression and encryption. •Selection of MME at UE attachment. •Routing of user plane data towards S-GW. •Scheduling and transmission of paging messages and broadcast information. •Measurement and measurement reporting configuration for mobility and scheduling. 15
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS eNBeNB eNB UE X2 X2 X2 S1 S1 MME/S-GW MME/S-GW RAN CoreNetwork S1 S1 Figure 2.2: LTE network architecture [5] The most important functions of the MME are the following: •Non-access stratum (NAS) signaling and NAS signaling security. •Access stratum security control. •Idle state mobility handling. •S-GW and MME selection (for HOs with MME change). •Bearer management control. Finally, the S-GW includes these functions: •Mobility anchor point for inter-eNB HOs and inter-3GPP mobility. •Packet routing and forwarding. •Transport level packet marking in the uplink and the downlink. The RAN segment in LTE is also called Evolved Universal Terrestrial Radio Access Network (E-UTRAN). The radio protocol architecture of E-UTRAN is represented in Fig. 2.3 for the userplane and the control-plane. Both the user-plane, which terminates in the eNB on the network side, and the control-plane include the following layers: •Physical layer. This layer offers data transport services to the next higher level through the so-called transport channels [31]. In addition, it handles coding/decoding, modulation/ demodulation, multiple antenna transmission, etc. •MAC layer. It handles the mapping between logical channels and transport channels [32]. 16
2.1. PRELIMINARIES PHY MAC RLC PDCP RRC UE PHY MAC RLC PDCP eNB MME NAS RRC NAS User-planeandcontrol-plane Control-planeonly Transportchannels Logicalchannels Layer3 Layer2 Powermeasurement andreporting Layer1 Physical channels Figure 2.3: User-plane and control-plane protocol stack and radio channels overview MAC offers services to the next higher level through the so-called logical channels. It is also responsible for error correction through HARQ and uplink and downlink scheduling. The scheduling functionality is located in the eNB. •Radio Link Control (RLC) layer. RLC is in charge of concatenation, segmentation and reassembly of RLC service data units, error correction through Automatic Repeat Request (ARQ), in sequence delivery of messages to higher layers and duplicate detection [33]. RLC offers services to the next higher level through radio bearers, which are then mapped to bearers belonging to the core network. •Packet Data Convergence Protocol (PDCP) layer. The main services and functions of this layer for the user-plane include IP header compression and decompression, ciphering and deciphering, in sequence delivery of messages to higher layers and duplicate detection [34]. The main functions for the control-plane are ciphering and integrity protection and transfer of control-plane data. •Radio Resource Control (RRC) layer. This layer is responsible for broadcast of System Information and paging, RRC connection management, radio bearer control, security and QoS management functions, UE measurement reporting and mobility functions including cell (re-)selection and HO [35]. •NAS layer. NAS protocols support the UE mobility and bearer management. This layer also defines the rules for a mapping between parameters during inter-system mobility with 3GPP and non-3GPP access networks. In addition, it provides the NAS security by integrity protection and ciphering of NAS signaling messages. By taking a closer look at the physical layer, both the signals and channels generated by this layer can be identified. On the one hand, the physical signals include the reference signals and the primary and secondary synchronization signals. The reference signals are used for channel estimation, meaning that, without the use of these signals, phase and amplitude shifts in the 17
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS received signal would make demodulation unreliable. In the uplink, the reference signals are also used for synchronization to the UE. The primary and secondary synchronization signals are used for cell search and identification by the UE. This allows the UE to identify and synchronize with the network. On the other hand, the physical channels carry data from higher layers including control, scheduling, and user payload. More specifically, in the downlink, the main physical channels and their functionalities are the following: •Physical Broadcast Channel (PBCH) carries cell-specific information. •Physical Downlink Control Channel (PDCCH) informs the UE about resource scheduling and HARQ information. •Physical Downlink Shared Channel (PDSCH) carries the user payload. •Physical HARQ Indicator Channel (PHICH) carries HARQ ACK/NACKs in response to uplink transmissions. In the uplink, the main physical channels are: •Physical Random Access Channel (PRACH) carries the random access preamble. •Physical Uplink Control Channel (PUCCH) carries CQI reports and HARQ ACK/NACKs in response to downlink transmission. •Physical Uplink Shared Channel (PUSCH) carries the user payload. Although the LTE downlink and uplink use different multiple access schemes, a common frame structure is shared. The frame structure defines the frame, slot, and symbol in the time domain. In the downlink, the primary and secondary synchronization signals, reference signals, PBCH, PDCCH and PDSCH are almost always present in a downlink radio frame. Each physical signal and channel are also associated with a specific modulation scheme. For instance, PBCH and PDCCH can only use QPSK, while PDSCH can use QPSK, 16QAM or 64QAM, depending on the AMC functionality. The next layer in the protocol stack, i.e. the MAC layer, makes use of the data transport services provided by the physical layer through transport channels and control information channels. On the one hand, the transport channels are as follows: •Downlink Shared Channel (DL-SCH) is mapped to PDSCH and characterized by supporting HARQ, dynamic link modulation and dynamic/semi-static resource allocation. •Broadcast Channel (BCH) is mapped to PBCH, it is required to be broadcast in the entire coverage area of the cell and it has pre-defined transport format. •Paging Channel (PCH) is mapped to PDSCH and it provides support for UE discontinuous reception, which is a feature to keep the UE in a sleep mode during certain periods (enabling UE power saving). •Multicast Channel (MCH) provides support for Multimedia Broadcast/Multicast service Single Frequency Network (MBSFN) and semi-static resource allocation. 18
2.1. PRELIMINARIES •Uplink Shared Channel (UL-SCH) is mapped to PUSCH and characterized by supporting HARQ, dynamic link adaptation and dynamic/semi-static resource allocation. •Random Access Channel (RACH) is mapped to PRACH and characterized by limited control information and collision risk. On the other hand, the control information channels are as follows: •Control Format Indicator (CFI) defines the number of PDCCH symbols per subframe. •HARQ Indicator (HI) represents a NACK with value ‘0’ and an ACK with value ‘1’. •Downlink Control Information (DCI) contains information including DL-SCH resource allocation. •Uplink Control Information (UCI) contains scheduling requests and acknowledgement responses or retransmission requests. The transport channels are encoded and decoded using channel coding schemes. The turbo coding is used for UL-SCH, DL-SCH, PCH and MCH, since these channels carry large data packets. A tail biting convolutional coding is used for downlink and uplink control as well as BCH. The MAC layer provides the logical channels to the next layer RLC. A logical channel is defined by the type of information which is carried. According to this, logical channels are classified into control channels and traffic channels for control-plane and user-plane information transfer, respectively. In particular, the control channels are the following: •Broadcast Control Channel (BCCCH) is a downlink channel used for broadcasting system control information. •Paging Control Channel (PCCH) is a downlink channel used for paging information and system information change notifications. Paging is the procedure that allows the network to request the establishment of a NAS signaling connection to the UE. PCCH is used for paging when the network has no knowledge about the location cell of the UE. •Common Control Channel (CCCH) transmits control information between UEs and the network. It is used for UEs having no RRC connection with the network. •Multicast Control Channel (MCCH) is a point-to-multipoint downlink channel used for the transmission of Multimedia Broadcast/Multicast Service (MBMS) control information. •Dedicated Control Channel (DCCH) is a point-to-point bi-directional channel used for the transmission of dedicated control information between a UE and the network. It is used for UEs having an RRC connection with the network. The traffic channels used by the RLC layer are: •Dedicated Traffic Channel (DTCH) is a point-to-point channel dedicated to a single UE for the transmission of user information. It is defined for both uplink and downlink. 19
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS •Multicast Traffic Channel (MTCH) is a point-to-multipoint downlink channel used for the transmission of user traffic to UEs that receive MBMS. The RRC layer handles the control-plane signaling of Layer 3 between the UEs and the EUTRAN. At this layer, information is carried by bearers, which are classified into radio bearers, S1 bearers and Evolved Packet System (EPS) bearers. Radio bearers, which are established by using RRC protocol, carry information on the radio interface. S1 bearers are defined between the eNB and the MME/S-GW, while EPS bearers are defined between the MME and the S-GW. There exists one-to-one mapping between radio, S1 and EPS bearers. An EPS bearer also has an associated QoS profile including some parameters, among which are the following: •QoS Class Identifier (QCI) is a scalar denoting the kind of service that can be supported by the bearer (e.g. bearer with/without guaranteed bit rate, priority, packet delay budget and packet error loss rate). This parameter is used as a reference to access node-specific parameters that control bearer-level packet forwarding treatment (e.g. scheduling weights, admission thresholds and queue management thresholds). •Allocation and Retention Priority (ARP) determines whether a bearer can be accepted or needs to be rejected in case of network congestion during the bearer establishment. In addition, this parameter is used for bearer modification and, occasionally, for bearer dropping. Once that the bearer has successfully been established, ARP has no impact on packet level forwarding. At Layer 3, two RRC states namely RRC_IDLE and RRC_CONNECTED are defined. In the RRC_IDLE state, the main functions are broadcast of system information, paging, cell measurement, cell selection/reselection and UE discontinuous reception. In the RRC_CONNECTED state, the UE has an E-UTRAN RRC connection, so that unicast data to/from the UE and broadcast/multicast data to the UE can be transferred. Other functions in RRC_CONNECTED are control channel monitoring to determine if data has been scheduled for the UE, CQI reporting, cell measurement, measurement reporting and system information acquisition. A UE changes from RRC_IDLE to RRC_CONNECTED state when an RRC connection is successfully established, while the UE returns to RRC_IDLE state by releasing the RRC connection. Unlike the RRC_IDLE state, in which mobility is controlled by the UE, the RRC_CONNECTED state involves network-controlled mobility. In LTE, fast and seamless HOs are of particular interest for delay-sensitive services such as VoIP. The simplified network architecture allows that HOs occur more frequently across eNBs than across core networks because the area covered by a MME/S-GW is generally much larger than the area covered by a single eNB. In addition, the signaling on X2 interface between eNBs is used for HO preparation, while the S-GW acts as anchor for inter-eNB HOs. To make proper HOs and cell reselections, decisions require knowledge of the environment. By taking and reporting measurements of the radio environment, the UE provides the information necessary to the network in order to make the correct mobility decisions. Thus, the UE and the eNB are required to make physical layer measurements of the radio characteristics. In particular, the main physical layer measurements in LTE are the following [36]: 20
2.1. PRELIMINARIES •Reference Signal Receive Power (RSRP). The RSRP is the linear average over the power contributions of the resource elements that carry cell-specific reference signals within the considered measurement frequency bandwidth. In both the downlink and the uplink there are reference signals, also known as pilot signals, which are used by the receiver to estimate, for example, the path loss. •Reference Signal Receive Quality (RSRQ). The RSRQ is defined as the ratio N ×RSRP/(EUTRA carrier RSSI), where N is the number of resource blocks of the E-UTRA carrier RSSI measurement bandwidth. The E-UTRA carrier RSSI comprises the linear average of the total received power observed only in OFDM symbols containing reference symbols, in the measurement bandwidth, over the N resource blocks by the UE from all sources, including co-channel serving and non-serving cells, adjacent channel interference, thermal noise etc. 2.1.2 Self-Organizing Networks During the last years, the growing demand of smartphones, online applications and multimedia services has led operators to upgrade their networks and use these resources more efficiently. With the extensive deployment of GSM, mobile networks are now available for a large number of people in the world. However, GSM is mainly intended for carrying voice traffic, while the capability devoted to data is rather limited. The increase in traffic demand comes from the emergence of new data services such as video and music streaming via web, social networking, locationbased services, free applications markets, etc. Furthermore, the improved data processing and storage capabilities of the new terminals encourage the development of such services and user applications. In this sense, together with the existing smartphones, the recent so-called tablets are also taking place within the market for mobile terminals. Such devices include high resolution displays, whose touch interface eases the handling, and powerful processors that allow to display high-definition video and high-quality games. To cope with such a demand for data traffic, upgrading mature networks by increasing the number of macrocells would not be sufficient. Instead, to provide the desired capacity, the introduction of incoming and future technologies such as LTE and its evolution LTE-A will be necessary. One implication of this is that, to provide seamless connectivity and mobility across the network, a more complex control-plane would be needed, since the network infrastructure will become more complex and heterogeneous. Upgrading the network infrastructure poses many challenges to operators in terms of reducing both CAPEX and OPEX [2]. CAPEX is related to the costs of adding new network resources to the existing network, while OPEX is related to the ongoing costs for network operation. To cope with the huge, necessary investment on infrastructure and taking into account that the average revenue per user has been decreasing during the last years, operators have paid special attention to cost savings, especially the OPEX. For this reason, the SON concept establishes a new paradigm for network management whose principal objective is to substantially reduce OPEX by decreasing human effort devoted to network operational tasks while at the same time optimizing network efficiency and service quality. In practice, SON means a set of functionalities for automated network planning, deployment, optimization and maintenance activities of these networks. As shown in Fig. 2.4, those functionalities are typically based on adaptive algorithms 21
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS Measurement Decision-making process Action (parametersetting) Network SONalgorithm Figure 2.4: Typical phases in a SON functionality which continuously execute a loop involving an initial measurement activity, a decision-making process and, lastly, a phase in which one or more actions are applied to the network. Firstly, the measurement phase implies a continuous activity where a lot of network measurements are collected and processed. According to the nature of the data, the information used as input for the SON algorithms can be classified into the following types [37]: •Configuration parameters. They represent the actual configuration of network elements and network resources. •Counters. This type of metric includes measurements from the network elements, which can be reported periodically or on-demand. Examples of counters are the traffic load, resource availability, the number of dropped calls, etc. •Alarms. An alarm is a message generated by a network element when there is a failure. •Mobile traces. This is information from specific UEs, which can be collected from most network elements (e.g. eNB). •Real-time monitoring. It includes online measurements of specific items, such as traffic load and HOs performed in a certain network element. •Drive tests. This refers to field measurements, e.g. related to coverage and interference, performed in a certain area by specialized equipment, which can also provide precise location information. •Key Performance Indicators (KPIs). They are derived from other measurements, e.g. through a specific formula of counters, with the aim of providing a meaningful performance measure. An example of KPI is the blocking ratio, defined as the ratio between the number of blocked calls and the number of offered calls. •Context information. This information is related to the environment, such as type of area (e.g. rural area), UE distribution (e.g. uniform), etc. 22
2.1. PRELIMINARIES Secondly, the processed measurements are used by the SON algorithm to make a decision, which is usually a variation in a network parameter. Roughly, the algorithm analyzes the incoming information and compares the actual with the desired behavior of the network. If any parameter changes are necessary, then an updated set of radio parameters is derived. Finally, the new parameter configuration is applied to the network. Examples of network parameters are antenna tilts, power settings, neighbor cell lists, HO thresholds, etc. As shown in Fig. 2.5, the implementation of the SON solution approach can be either centralized or distributed [6]. In the first case, all SON algorithms are executed in a central node (e.g. a server), which reasonably should be located close to or within the Operations, Administration, and Maintenance (OAM) system. A disadvantage of this approach is that the fault-tolerance is low, since a failure in the central node would make the system inoperative. In the second case, the SON algorithms are located on each network element, e.g. eNB. The eNBs would communicate with each other via the X2 interface in order to exchange information. The main drawback of this approach is that more communication bandwidth is needed, as well as coordination between many nodes would be also a complex task. Presently, network management consists of the processes and activities involved with planning, configuration, optimization and maintenance of networks, which are performed centrally from an OAM system. Some of these tasks are accomplished manually (e.g. adjustment of antenna tilt), while others are typically semi-automated and need to be supervised by human operators (e.g. adjustment of mobility parameters). Such a human effort is often expensive, slow and subject to errors. Thus, automation of network management allows operators to face the increasing complexity of mobile networks in order to use more efficiently resources while at the same time diminishing human involvement. In addition, SON capabilities can be further exploited in a context of multi-technology scenarios. Extending the SON concept to all RATs allows eNBeNB eNB X2X2 SON SON SON eNBeNB eNB X2X2 SON OAMsystem a)Centralizedapproach b)Distributedapproach Figure 2.5: Implementation schemes for SON functions 23
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS of the OPEX, and some environmental aspects, such as the CO2emission, have become strong arguments from the perspective of the current governments and corporations. QoS-related parameters optimization The QoS refers to several related aspects of the services calls that allow the transport of traffic with special requirements, e.g. ensuring a specific frame error rate, delay or throughput. The optimization of QoS-related parameters deals with the trade-off between the QoS and the Gradeof-Service (GoS). The latter refers to performance metrics associated with call accessibility and maintainability, which are basically related to how many calls the system can handle simultaneously. Accessibility is usually measured in terms of call blocking (it happens when a call request is rejected), while maintainability is usually measured in terms of call dropping (it happens when a ongoing call is aborted). The following use cases attempt to find a good trade-off between QoS and GoS: •Admission control parameter optimization. The basic operation of the admission control is to admit or reject new calls and incoming calls from other cells. This decision takes into account different aspects such as the current capacity, the system load, desired QoS of the newly requested and ongoing calls. In this sense, the objective of this use case is to selfoptimize the admission control thresholds in order to minimize call blocking while at the same time fulfilling the QoS requirements of the admitted calls with sufficient likelihood, which is related to uncontrollable factors, such as the user mobility and variations in radio conditions and system load. •Congestion control parameter optimization. While the admission control operates under normal conditions, the congestion control works in extreme conditions in order to cope with e.g. overload situations and prolonged bad radio conditions. In those cases, the admission control, which is of a proactive nature, cannot deal with unacceptable values of QoS requirements. In contrast, the congestion control, which is of a reactive nature, have to get the system back to a feasible load situation. Thus, the objective of this use case is to self-optimize the congestion control parameters in order to maximize resource utilization provided that certain QoS degradation due to congestion problems is allowed. •Packet scheduling parameter optimization. This use case is mainly applied to LTE networks. Since LTE is optimized for packet data transfer, the packet scheduling plays a critical role. More specifically, the packet scheduler coordinates the access to shared channel resources and operates in two dimensions, the time and the frequency domain. The objective of the packet scheduling parameter optimization is to achieve the individual QoS requirements of the ongoing calls in the most efficient way, i.e. finding a good trade-off between fairness and spectral efficiency. Self-Optimization of Home-eNBs The concept of home-eNB refers to a small base station designed to increase capacity and coverage in indoor environments, e.g. residential and small business areas. These low-power nodes allow to create cells of small size which are called femtocells. A home-eNB may be turned on and off frequently or moved to a different geographical position, it may have closed or open access and 30
2.1. PRELIMINARIES it is not physically accessible for operators. All these features transform the Self-Optimization of home-eNBs into a challenging task, which is different from the optimization related to macrocells. The home-eNBs need to be physically installed by the customer and connected to the operator network through the customer’s fixed Internet line. Since the customer usually does not have the knowledge needed to install software on home-eNBs, the configuration of these devices need to be performed in an automatic manner. Although the management of home-eNBs includes certain elements of Self-Configuration, it is more related with Self-Optimization, since these small nodes need to be constantly adapted to context variations. The objectives of this use case are various. For instance, a home-eNB should automatically detect neighboring eNBs (including other home-eNBs) in order to have seamless mobility between eNodeBs. The neighboring cell list should also be maintained and optimized throughout the time, especially because home-eNBs are switched on and off arbitrary and more frequently. In addition, the radio parameters (e.g. the transmit power) should be optimized in order to minimize coverage holes or interference. Regarding the user mobility, the decision of performing or not an HO (between macrocell and femtocell) should also be optimized in order to avoid unnecessary HOs. Mobility Robustness Optimization The MRO aims to minimize the occurrence of undesirable effects (e.g. dropped calls) due to the HO process, which is the procedure that allows the user to freely move around the network. Since recent technologies (e.g. LTE) are being deployed with a frequency reuse of 1, the interference in adjacent cells becomes an important issue, especially in the cell-edge. This typically results in dropped calls and failures related to HOs. Furthermore, sub-optimal settings of HO parameters under poor radio conditions are more sensitive to call dropping and HO signaling load, which may lead to a redundant waste of network resources. Thus, the main objective of MRO should be to reduce the number of dropped calls due to suboptimal adjustment of HO parameters, while the secondary objective would be to reduce the inefficient usage of network resources due to unnecessary HOs. The importance of this use case lies on the fact mobility management is crucial for operators, since sub-optimal settings of the mobility procedures directly affect the user performance. In particular, the call dropping is key from the operator perspective, since a reduction in this indicator can involve a significant loss of customer confidence and thus, important revenue losses for the operators. Mobility Load Balancing Optimization The MLB optimization deals with congestion situations in which some cells in the network are more heavily loaded than their neighbors e.g. due to an uneven spatial distribution of the offered traffic. In these situations, some traffic could be shifted from the heavily loaded cell towards the more lightly loaded cell by adjusting mobility parameters in order to relieve the congestion. This use case is strongly linked to the MRO use case, since the MLB aims at dynamically 31
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS adapting mobility parameters according to the current traffic load of the cell and that of neighboring cells. The potential of sending users towards non-optimal cells from the radio perspective (e.g. signal strength) may lead to a decrease in the radio conditions of those users transferred to neighboring cells, meaning that the MRO should also take care of such an issue. However, as a result of a better matching between the spatial distribution of traffic demand and network resources, it is expected that more users can be accepted in the crowded area and more capacity is provided to the ongoing connections. Traffic Steering The evolution of next-generation wireless networks is envisioned to satisfy the growing traffic demand of new services. The increasing complexity and capability of terminals (or smartphones) has encouraged the development of new applications with higher requirements of bandwidth. To deal with this traffic demand, multiple RANs will be deployed covering the same area. HetNets will be characterized by the presence of networks with different technologies, frequencies, cell sizes, etc. Future RATs must co-exist and co-operate with existing technologies to provide high data rates and good QoS. In addition, hierarchical cellular structures involving cells with different sizes (e.g. macro, micro, pico and femtocells) allow to provide more capacity in crowded areas or hotspots. In this context, since the coverage area of the deployed networks in HetNets are partially or totally overlapped, mobile operators can send users towards a specific network with the aim of utilizing resources more efficiently. This concept, known as TS, can be performed by using mechanisms of mobility management, which are standardized by the 3GPP in the case of mobile networks, providing support for interoperability with other systems [5]. In addition, such a task of selecting a more suitable network can be made to achieve a goal or a combination of goals, for instance, related to cost, user satisfaction, power consumption, coverage, etc. For this reason, the potential of TS in HetNets due to the widespread deployment of overlapping wireless networks brings many challenges to the operators. 2.2 State of the Art This section presents the state of the art in SON, focusing on the use cases studied in this thesis as well as a survey on adaptive techniques for network parameter optimization. According to the 3GPP guidelines, certain use cases have received special attention in the research field, in particular, the MRO and the MLB Optimization [6]. Moreover, enhancements for these two use cases have recently become a requirement in 3GPP Release 10. Due to this, the MRO and MLB use cases have been addressed in this thesis. In addition, the coordination of SON techniques has also become of special interest in Release 11 [6]. Since there are multiple use cases, one challenging issue arising from the application of SON techniques is that some of these use cases are strongly coupled, needing to be coordinated. For instance, the two previous example use cases (MRO and MLB) can cooperate on the problem of a high number of dropped calls. Thus, the network performance could benefit from the SON coordination in these situations. For this 32
2.2. STATE OF THE ART reason, the coordination of MRO and MLB has also been addressed in this thesis. On the other hand, to cope with the increasing mobile-broadband traffic in the last years, HetNets has become an effective solution to provide higher data rates and seamless connectivity and mobility across the network. In this sense, TS is a challenging Self-Optimizing use case from the operator perspective [45]. The concept of TS involves a powerful mechanism based on sending users to a specific subnetwork in the HetNet with the aim of optimizing the usage of radio resources. Due to this, the potential of TS techniques has also been studied in this work. After presenting a survey on international projects related to SON, the rest of the section is devoted to the bibliography related with the specific SON use cases addressed in this thesis, i.e. MLB, MRO, SON coordination and TS in HetNets. In addition, a survey on adaptive optimization techniques for these use cases is also provided. 2.2.1 International projects The introduction of new air interface technologies and network topologies due to the growing demand of smartphones and multimedia services has increased the complexity of network operation. The only feasible option for operators to cope with rapid technological changes and maintain the level of competition is by increasing the level of automation in the network [46]. In this context, there are certain areas especially attractive to offer a cost-effective implementation, especially those involving heavy calculations or signal processing. The benefit of a higher level of automation is an increase in network performance since those tasks can be executed faster and applied on a wide scale, i.e. a large number of cells. Another positive side effect is that human effort can be freed from mundane and tedious tasks and, instead, expertise can be applied in new areas, which bring additional value to the operator. As a result, the staff will be working on challenging tasks, which results in more motivation and productivity. For these reasons, the automation of network operation has gained attention in the research community, which is evident from several completed and ongoing international projects. The main public research projects and consortiums which are related with the area of automated network management and SON are the following: •The GANDALF project (2005-2007), being part of the initiative EUREKA CELTIC, was aimed at developing automatic optimization tools for wireless networks with multiple RATs, in particular, GSM, General Packet Radio Service (GPRS), UMTS and Wireless Local Area Network (WLAN). This project covered automated diagnosis for troubleshooting, auto-tuning of network parameters and joint RRM between the networks [47]. •The End-to-End Efficiency (E3) project (2008-2009) was focused on integrating cognitive wireless systems in the Beyond-3rd-generation (B3G) world, evolving current heterogeneous wireless system infrastructures into an integrated, scalable and efficiently managed B3G cognitive system framework [48]. One of the main E3objectives was to optimize the use of the radio resources and spectrum, following cognitive radio and cognitive network paradigms. To achieve this, the purpose was firstly to build a unique and generic architectural framework which was network and equipment agnostic and was mapped onto existing 33
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS and future network topologies. Then, autonomic and self-X concepts were incorporated in the system architecture to increase the efficiency of network operation. Finally, self-learning capabilities were included to improve the targets of the self-optimization processes in the management system. •COST 2100 (2006-2010) was the Action on pervasive mobile and ambient wireless communications belonging to the European, inter-governmental, scientific and technical cooperation framework COST. The Sub-Working Group (SWG) 3.1 focused on mobile wireless network optimization. As network models together with simulated environment models usually do not satisfy the operator’s requirements, SWG 3.1 aimed at substituting artificially generated data by the measured data taken from the real system. The continuation of this Action, known as COST IC1004, is currently ongoing. Its Working Group (WG) 3 is similar to the SWG 3.1 previously mentioned. •The SOCRATES project (2008-2011) developed SON algorithms to enhance the operations of wireless access networks, by integrating network planning, configuration and optimization into a single, mostly automated process requiring minimal manual intervention [2]. In this context, a functional architecture for SON coordination has also been proposed in order to avoid conflicts due to the simultaneous operation of different SON functions. The use cases considered by the SOCRATES project are listed in [44]. For each use case, the required input data, the desired output parameters, the actions to be performed, possible dependencies on further standardization and potential gain in terms of OPEX and CAPEX reduction were firstly analyzed. Later, aspects such as the implementation and interdependencies among the use cases were also studied. •The SELF-NET project (2008-2010) aimed at designing, developing and validating an innovative paradigm for cognitive self-managed elements of the future internet [49]. One of the main objectives was to incorporate new management capabilities into network elements in order to take advantage of the increasing knowledge that characterizes the daily operation of mobile users. In addition, another key objective of Self-NET was to provide a holistic architectural and validation framework that unifies networking operations and service facilities of the future internet. •The UniverSelf project (2010-2013) aims at overcoming the growing management complexity of future networking systems and reducing the barriers that complexity and ossification pose to further growth [50]. To achieve these objectives, UniverSelf presents a unified management framework for the different architectures and all required functions to achieve self-management. The impact on the industry and the push of European research into the direction of exploitation is another of the main goals. •The SEMAFOUR project (2012-2015), considered as the continuation of the SOCRATES project, will design and develop a unified self-management system, which enables the network operators to holistically manage and operate their complex heterogeneous mobile networks [51]. The first objective is to develop multi-RAT and multi-layer SON functions that provide a closed control loop for the configuration, optimization and failure recovery of the network across different RATs (e.g. UMTS, LTE, WLAN) and cell layers (e.g. 34
2.2. STATE OF THE ART macro, micro, pico, femto). The second objective is to design and develop an integrated SON management system, which acts as interface between operator policies and the SON functions. 2.2.2 Mobility load balancing Load balancing copes with congestion situations where neighboring cells with spare resources are utilized for offloading the congested cells. Load balancing is considered by the 3GPP as an important issue in SON due to its effectiveness to increase network capacity. Firstly, the concept of SON is widely addressed in next-generation networks, being part of the specification of LTE [52]. Secondly, the MLB use case is introduced in [43] together with other Self-Optimization use cases. Later, the MLB, as part of SON, is recognized as one of the key enablers in LTE-A [53], where enhancements for this use case are proposed to be investigated. Finally, the application of MLB to HetNets and inter-RAT scenarios is proposed in [6]. The load balancing problem can be solved by sharing traffic between adjacent cells. MLB is tackled here as a Self-Optimizing task that can be solved by tuning specific network parameters, in particular, those involved in the HO process. The HO process is responsible for transferring an ongoing call from one cell to another. By adjusting HO parameters settings, the service area of a cell can be modified to send users to neighboring cells. Thus, the size of the congested cell is reduced while adjacent cells increase in size taking users from the congested cell edge. As a result of a better matching between the spatial distribution of traffic demand and network resources, more users could be accepted in the crowded area so that the call blocking probability would be reduced [54]. A significant research effort has been devoted to the load balancing problem. In [54], a method for determining the best traffic share between cells in a GSM network is presented, where a closed-form expression for the optimal traffic sharing criterion is derived by solving a classical optimization problem. In [55], it was shown that the number of unnecessary HO attempts and failures can be significantly reduced by tuning the load-based HO thresholds in a GSM/WCDMA scenario, where it is assumed that a sufficient number of users are handed over to less loaded cells when the load exceeds the threshold. In [56], the decision of triggering inter-system HOs for load balancing between UMTS and GSM depends on the load in the target and the source cell, the QoS, the time elapsed from the last HO and the HO overhead. The studies in [56] show that both the capacity and the provided QoS could be notably improved in overload situations. In [7], the proposed load balancing scheme tunes HO thresholds in the soft-HO algorithm for UMTS networks, where a user can be simultaneously connected to two or more cells. In the case of LTE networks, several studies about load balancing can also be found in the literature. In [8][9][10], similar algorithms for MLB are proposed. In particular, these algorithms automatically adjust the HO margins until the load or the load difference between cells is below a threshold. In [57], load balancing is formulated as a multi-objective optimization problem. More specifically, the proposed algorithm maximizes both load balance for real-time constant bit rate users and a utility function based on QoS requirements for best effort users. In [58], the authors propose a method inspired by the flowing water algorithm that creates an HO 35
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS margin adjusting scheme from load measurements. In [59], a game theoretic scenario for load balancing is proposed. In [60], the load balancing problem is simplified to a single aggregate objective function optimization problem, which takes into account the physical resource limits and the QoS requirements. Some of the previous references are focused on inter-system load balancing, where propagation conditions are not necessarily a problem because the coverage areas of different radio technologies are typically overlapped. In the case of intra-system load balancing, when a user is handed over to another cell, the propagation conditions could get worse if the user is located near the cell edge because of the interference caused by neighboring cells. In LTE, the previous references analyze the performance of load balancing mostly in terms of achievable throughput for data services, neglecting the dropping rate, which is an important performance indicator in real-time traffic (voice, video, etc.). Also it is noted that controlling those quality indicators (e.g., dropping rate) for real-time traffic is not considered by the previous references as only performance evaluation is carried out without any further impact on the indicators. In [10], it is remarked that the trade-off between enhanced blocking and degraded dropping can be adjusted by restricting the maximum achievable HO margin. However, further study of load balancing effects on quality indicators for real-time traffic would be necessary. In this sense, an important contribution of this thesis with respect to the MLB use case is the design of an algorithm to control quality performance for real-time traffic during the load balancing process. On the other hand, the vertiginous increase in mobile-broadband traffic experienced in the last years has led operators to deploy cells of lower sizes to cope with such an increase in traffic demand. The reason for this is to keep the users closer to the base station as well as to allocate a lower number of users per cell. In this context, femtocells are envisioned to cope with such a demand of capacity in indoor environments [61]. Since those small cells are low-cost nodes, a thorough deployment is not typically performed, especially in enterprise scenarios. As a result, the matching between traffic demand and network resources is rarely the optimal. This means that the application of load balancing techniques are especially attractive in those cases. In addition, most of the tasks involved in the installation, operation and maintenance of femtocells have to be done in an automatic manner, since the customer cannot be assumed to have the knowledge necessary to manage femtocells. For this reason, SON is also a key concept for femtocell management. The challenges in femtocell self-organizing has gained attention in the research community, which is evident from several international projects (e.g. HOMESNET [62], BeFEMTO [63], FREEDOM [64]). Most efforts have been paid to the design of advanced Radio Resource Management algorithms to reduce inter-cell interference [65][66][67]. In [68], a self-tuning algorithm for selecting femtocell pilot power and antenna pattern is proposed to improve the coverage and minimize the total number of HO attempts. In [69], an adaptive algorithm for selecting the hysteresis margin based on user position is presented. In [70], an accurate model for simulating femtocells is proposed, where different levels of complexity in the prediction of femtocells and their effects on HOs as well as on call drops are also investigated. In addition, recent studies have considered networked femtocell environments, among which is the enterprise scenario [71]. 36
2.2. STATE OF THE ART However, few studies have addressed load balancing in enterprise femtocells with legacy equipment, which is of relevant interest for operators [72][11][73]. Many challenges arise from the fact that those scenarios often include three-dimensional structures, unequal distribution of traffic load is very common (e.g., crowded/uncrowded offices), the user mobility pattern is different (probably more intense) than at home and open access is usually employed instead of closed access. Thus, further study of load balancing in enterprise scenarios would be necessary. In particular, this thesis investigates the potential of different load balancing techniques to solve persistent congestion problems in enterprise LTE femtocells. 2.2.3 Mobility robustness optimization Roughly, MRO is the optimization of the HO process, which typically involves a trade-off between the amount of signaling load due to HOs and the quality of the active connections in the network. As in the case of MLB, MRO has been considered by the 3GPP as a key functionality in the field of SON. The basics of MRO are introduced in [43], while enhancements for this use case and its application to HetNets and inter-RAT scenarios are addressed in [53] and [6], respectively. Following the 3GPP guidelines [6], the main objective of MRO should be to reduce the number of dropped calls due to suboptimal adjustment of HO parameters, while the secondary objective would be to reduce the inefficient usage of network resources due to unnecessary HOs. The MRO has gained attention in the research community, which is evident from several international recently finalized projects (e.g., SOCRATES [2] and E3 [48]). Thus, several techniques for HO optimization applied to different RATs can be found in the literature. In [74][75], the optimization of the so-called soft-HO in WCDMA has been addressed. However, nextgeneration wireless networks are based on hard-HOs, meaning that the UE is attached to a single cell at a specific time. In this field, some performance analyses for intra-frequency LTE HOs are described in [76][77], but no SON algorithms are provided to find an optimal setting of HO parameters in those papers. Other works that propose a SON algorithm only consider one HO parameter in the analysis, [12][13][14]. In [15][16], techniques for modifying both the HO margin (HOM) and the Time-To-Trigger (TTT) are proposed, but no further impact on user speed is analyzed. Finally, the inter-system HO optimization has been addressed in [78][79]. Unlike previous work in this area, a significant contribution of this thesis is to study performance analysis considering both HO parameters (i.e. the HOM and the TTT) for different situations, highlighting the variation of the user speed. The impact on the GoS (e.g., call dropping and call blocking) for future services such as VoIP calls is also studied. In addition, an important contribution in the area of SON is the proposal of a novel MRO algorithm, which has several benefits over other existing approaches. For instance, in [15], the optimal values of the HO parameters are selected from a diagonal through the grid of HOM and TTT values, based on a previous sensitivity analysis derived from system performance simulations, thus limiting the results to the used, realistic, simulated scenario. Conversely, in this work, the proposed SON algorithm optimizes HOMs independently of the simulated scenario. 37
CHAPTER 2. SELF-ORGANIZATION IN RADIO ACCESS NETWORKS 2.2.4 SON coordination In Section 2.1.2, the development of stand-alone SON functions has been addressed. It is noted that, with the growing deployment of SON functions, the number of conflicts and dependencies between them increases. A conflict can happen if, for example, two different individual SON functions (e.g. the MLB and MRO) optimize the same parameter (e.g. the HOM) with different goals at a network element. The interaction between SON functions also exists if the modification of a parameter affects the operation of other SON functions. Those conflicts may have a negative impact on network performance, which is far from the operator’s objectives when the SON functions are implemented. Thus, the avoidance and resolution of conflicts and dependencies between these functions become an important issue for operators. The study of SON coordination is a recent topic addressed in the bibliography. In [80], a generic mathematical model for the interaction of multiple SON mechanisms running in parallel, along with a practically implementable coordination mechanism, are proposed. The SOCRATES project has also defined a general framework approach for SON coordination that addresses the avoidance, detection and resolution of conflicts that can occur between different individual SON functions [81][82]. However, in the SOCRATES project, minor attention has been devoted to specific SON use cases that have a strong interrelation regarding the involved control parameters. In particular, the coordination between macro and femtocell HO optimization [83], admission control and HO optimization [84], and load balancing and HO optimization [85] have been investigated. Related to this thesis, special attention has been devoted to the coordination between load balancing and HO optimization (i.e. MLB and MRO, respectively), addressed by the SOCRATES project. However, in that project, the study assumes the control parameters of the MLB and MRO algorithms to be independent of each other, i.e. the two algorithms do not tune the same parameters. In particular, the MLB function adjusts the HOM, while the MRO function adjusts the TTT and hysteresis parameters. The interactions exist because these two functions influence the same KPIs that are used as input for the optimization algorithms. Unlike this design, the MRO function proposed in this thesis do not adjust the hysteresis parameter of the HO procedure, since it cannot be adjusted per cell-basis or adjacency-basis, which is a severe limitation in mobile networks due to the variant nature of the radio environment. This consideration makes the coordination problem different from the approach studied by the SOCRATES project. The coordination between MLB and MRO has been also investigated in some other works. In [86], a constraint for the connection quality more restrictive than the one assumed in the SOCRATES project is considered. In this sense, the MLB function is restrained in favor of the HO performance optimization. However, as in the SOCRATES project, the MRO function also consider the hysteresis parameter, which is a limitation as it cannot be defined at cell level. Furthermore, the impact of the user speed on the coordination algorithm is disregarded. The user speed is an important factor as the HO signaling significantly increases with the user speed. In [87], to avoid the conflict between MLB and MRO, the HOM range of MLB is dynamically adjusted according to the TTT and the hysteresis parameters, which are first adjusted considering the effect of the user speed. However, such an optimization is related with a fast adaptation (i.e. on-line) of HO parameters, which is not the problem addressed in this thesis. In particular, the parameters in [87] are considered to be UE specific. Such UE specific 38
2.2. STATE OF THE ART properties are not available on the level of the OAM system, so this kind of optimization can only be done in the eNB. In addition, that approach requires velocity estimates which are not easy to obtain [41]. Conversely, in this thesis, the variations of HO parameters are performed to adapt to slower changes (e.g. hours, days) in the environment and they are adjusted per cell-basis or adjacency-basis. 2.2.5 Traffic steering Currently, large investments on infrastructure are needed to cope with the growing demand of traffic. In this sense, operators can choose from various alternatives, such as deploying cells of lower sizes and making use of newer technologies (e.g. LTE and LTE-A). However, these solutions are expected to be insufficient, meaning that the existing networks will also play a key role in coping with that enormous traffic demand. In this context, HetNets are shown as an effective solution to provide higher data rates and seamless connectivity and mobility across the network. To achieve these goals, HetNets involve not only deploying e.g. different RATs, cell sizes, and carrier frequencies, but also managing the hybrid network as a unified whole. In HetNets, the concept of TS involves a powerful mechanism based on sending users to the correct subnetwork with the aim of optimizing the usage of radio resources. Since the coverage area of the deployed subnetworks are partially or totally overlapped, network operators could determine toward which network a user connection should be routed to utilize resources more efficiently. The potential of steering a user connection towards a specific network layer in HetNets brings new challenges to the operators. Such a challenging task of selecting the best network due to the widespread deployment of overlapping wireless networks has been addressed in the literature. In [17], a policy-enabled HO to express policies on what is the best wireless system is proposed. These policies allow to establish different trade-offs among indicators related to cost, performance, power consumption, etc. In [18], an analytical approach to define a wide range of RAT selection policies taking into account several allocation criteria such as service type, load, etc. is proposed. Users can also use simultaneously services through different RATs. Such a problem is addressed in [19][20], where different strategies in a multi-RAT, multicellular and multiservice scenario are designed to indicate the suitability of selecting a specific RAT. Similarly, the problem of allocating multiple services onto different subsystems in multiaccess wireless systems is addressed in [21], where some principles for how this service allocation should be done to maximize the resulting combined capacity are discussed. There are also many references focused on particular cases showing the benefits of TS. For instance, a network layer suffering from a temporary traffic congestion can offload some users to a non-loaded layer (load balancing) [9][56]. High speed users can be connected to a so-called umbrella layer to avoid a frequent number of HOs [88]. Other factors contributing to the preferred choice of the access technology are addressed in [89], where aspects such as coverage, securing QoS for the requested service, minimizing the cost of delivering the service to the end user and maximizing the spectrum utilization by traffic packaging are considered to select the access technology. Potentials of dynamic TS algorithms are also addressed in [45]. 39
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES that the controller is improved by an optimization method (i.e. the optimizer) to cope with the context variations in the environment, or to refine the controller when knowledge is not (or partially) available. The latter case is of special interest in wireless networks, since the presence of many context factors makes the definition of the controller more difficult. In addition, if the reference values of the controller have not been properly selected, they can be adjusted by the optimizer, thus being more likely to find the optimal solution. Finally, as in the second approach, the main drawback is that the search space is small, as no service degradation during operation is accepted. Based on the previous statements, in this thesis, the second and third schemes shown in Fig. 3.1 are used for network parameter optimization. In particular, the control-based approach is implemented when the configuration of the controller is simple and knowledge is available. In this case, a trial-and-error strategy is typically performed to find the best configuration of the controller. Conversely, the self-optimizing control-based approach is applied in those cases in which the complexity of defining the behavior of the controller is high or knowledge is not available. In this context, Fuzzy Logic is a mathematical discipline especially appropriate to design controllers, as its potential lies in the capability to express the knowledge in a similar way to the human perception and reasoning. Such a property allows operators, for instance, to use a linguist term such as high or low instead of providing a numerical value when defining the reference values of the controller. For this reason, in this work, controllers are based on the Fuzzy Logic theory. More details about this relatively young discipline as well as FLCs are provided in the following sections. The second classification of approaches to face the problem of parameter self-tuning is made according to the frequency with which the self-tuning method is applied to the network. Such a frequency usually depends on the interval time with which measurements are collected, the delay to update the parameter settings and the computational time required to perform the calculation process. The tasks in which instantaneous performance indicators are used and a fast response of the algorithm is required (e.g. in the order of milliseconds or seconds) are referred to as on-line tuning methods. In cellular networks, this kind of tasks is known as Advanced RRM algorithms, including time and frequency domain packet schedulers, adaptive transmission bandwidth, power control, AMC, etc. In contrast, parameter changes can also be performed more slowly. For instance, the OAM system is a network entity that receives statistical performance data from the entire network, for example, every hour. After this, the new parameter settings are calculated in several minutes and immediately downloaded to the network. Hence, in this case, the optimization algorithm cannot be used to cope with fast network changes. However, as a remarkable advantage, the use of long-term statistical data leads to more robust methods. In addition, as the temporary limitation is given by the measurement periods, there is plenty of time to apply complex optimization methods, which can further improve network performance. Thus, the term that refers to this kind of tasks is off-line tuning methods. From the operator perspective, the most important advantage of off-line methods is that they do not involve additional costs as it happens with Advanced RRM (i.e. on-line methods), where these algorithms are typically located in the base stations and owned by the manufacturer. 46
3.1. PARAMETER SELF-TUNING In this case, upgrading the network equipment represents an additional cost for the operator. Thus, the off-line approach is a more cost-effective solution and, for this reason, it has been adopted in this work. Finally, to assess a self-tuning algorithm, there are many criteria that can be used, such as the solution quality, the convergence speed and the complexity in terms of computational load and storage requirements. The solution quality achieved by the algorithm (e.g. given by the final values of KPIs) is the most determining factor from the performance point of view, especially in complex situations (e.g. a cellular network) where the algorithm is not able to find the best solution. The computational load and storage requirements are crucial for on-line algorithms and those executed in the UE side, which are not the case in this thesis. Finally, the convergence speed defines how fast the algorithm approaches the final solution. This criterion is of special interest when the algorithm includes the optimizer, as this entity is typically based on an iterative approach. In this work, the proposed algorithms are off-line and conceived for the OAM system, whose limitations in computational load and storage requirements should not be an issue. Thus, the analysis is focused on performance indicators related to the solution quality and the convergence speed. 3.1.2 Fuzzy Logic This section presents the theoretical basis of the computational intelligence methodology known as Fuzzy Logic. This discipline was initiated in 1965 by Lotfi A. Zadeh [111], professor at the University of California, in Berkeley. Fuzzy Logic emerged as an important tool for system control and complex industrial processes, as well as for home and entertainment electronics, diagnostic systems and other expert systems. Currently, multitude of applications based on Fuzzy Logic have been applied in many different areas, among which are control systems, robotics, medicine, pattern recognition, computer vision, information and knowledge management systems, earthquake prediction, scheduling optimization, etc. As an alternative to Classical Logic, Fuzzy Logic introduces a degree of vagueness when things are evaluated [112]. In real life, there is much knowledge that is ambiguous and imprecise, and human reasoning usually works with this kind of information. In this sense, Fuzzy Logic was designed specifically to imitate the behavior of humans. Additional benefits of Fuzzy Logic include its simplicity and its flexibility. In particular, this methodology can handle problems with imprecise and incomplete data, and it can easily model non-linear functions of arbitrary complexity. On the one hand, classical sets arise from the need of humans to classify objects and concepts. Such sets can be defined as a well-defined set of elements or by a membership function µthat can take the values 0 or 1 from a universe of discourse for all the elements that can belong (or not) to the concerned set. Formally, let Xbe the universe of discourse and xthe elements contained in X. In addition, let suppose that Ais a set that contains some elements in the universe of discourse X. Then, the element xbelongs or does not belong to the set Acan be represented by the following function: 47
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES µA(x) = (1if x∈A 0if x /∈A(3.1) where µA(x)is the membership function corresponding to the set A. On the other hand, the need to work with fuzzy sets comes from the existence of concepts with no clear boundaries in their definition. Unlike classical set theory that classifies the elements into crisp sets, fuzzy set theory has an ability to classify elements into continuous sets using the concept of degree of membership. As a result, the membership function can take values in the range between 0 and 1, and the transition between both values is gradual, unlike in classical sets. Formally, let suppose that Bis a fuzzy set that contains elements in the universe of discourse X. Then, such a fuzzy set is characterized by the membership mapping function: µB(x) : X→[0,1].(3.2) For all x∈X,µB(x)indicates the certainty to which element xbelongs to fuzzy set B. Although the membership function for a particular fuzzy set can be of any shape or type, an appropriate membership function is typically determined by experts in the domain over which the set is defined. In this sense, some membership functions are of special interest for designers, e.g. the triangular, trapezoidal and Gaussian types. As for crisp sets in Classical Logic, relations and operators can also be defined for fuzzy sets in Fuzzy Logic. In particular, these relations are the equality, containment, complement, intersection and union of fuzzy sets. Among these relations, the intersection of fuzzy sets plays a key role in the design of rules for fuzzy controllers, as described in the next section. By definition, the intersection of two sets, Aand B, is the set of elements occurring in both sets, while the operators that implement intersection are referred to as t-norms. The result of a t-norm is a set that contain all the elements belonging to both fuzzy sets, but taking into account the degree of membership, which depends on the specific t-norm. The most popular t-norms are the following: •Min-operator. Formally, it is defined as: µA∩B(x) = min{µA(x), µB(x)},∀x∈X. (3.3) •Product operator. It can be expressed as: µA∩B(x) = µA(x)µB(x),∀x∈X. (3.4) As previously stated, Fuzzy Logic was conceived to imitate the behavior of humans. For this reason, the concept of a linguistic variable plays an important role in Fuzzy Logic. A linguistic variable is a variable whose values are words or sentences from natural language, allowing computation with words instead of numbers. Such a linguistic variable can be a word, a linguistic label or an adjective. For example, let consider the height of the people in a country. In this case, the variable height is a linguistic variable. A possible value for the numeric variable 48
3.1. PARAMETER SELF-TUNING height can be tall or small, meaning that a fuzzy set is associated with a linguistic term or value. In addition, certain adverbs can also be combined with adjectives to modify fuzzy values, e.g. very tall would indicate a person who is “taller” than tall. Thus, linguistic variables allow the translation of natural language into logical, or numerical statements, which facilitates handling human reasoning at the computational level. In practice, most sensory inputs are linguistic variables, or nouns in a natural language, for instance, temperature, pressure, displacement, etc. In this work, KPIs and parameters are the inputs of the proposed algorithms, so that typical linguistic values for these inputs are very high,low,medium, etc. An important feature of Fuzzy Logic is that it provides a framework for handling rules (for control or decision making) which have been expressed in an imprecise form. In this context, linguistic variables are embedded in the rules of an FLC, allowing representation of the human control expertise. More specifically, FLCs consist of a number of conditional IF-THEN rules. For the designer who understands the system, these rules are easy to write, and as many rules as necessary can be supplied to describe the system properly (although only a moderate number of rules are usually needed). Next section provides a short overview of FLCs, focusing on the components of such controllers and some types of fuzzy controllers. 3.1.3 Design of a Fuzzy Logic Controller The design of FLCs is one of the most important application areas of Fuzzy Logic [113]. The main benefit of FLCs is that controlling a system (also called plant) can be performed by using sentences rather than equations. This means that a control strategy can be described in terms of linguistic rules, in a more similar way to human language, instead of using e.g. differential equations. Since the first one was conceived in 1975, a huge number of FLCs have been developed so far for consumer products (e.g. washing machines, video cameras and air conditioners) and industrial processes (e.g. robot control, underground trains and hydro-electrical power plants). Experience has shown that FLCs provide results superior to those obtained by conventional control algorithms. In particular, the methodology of the FLC becomes very useful when the processes are too complex for analysis by conventional quantitative techniques or when the available sources of information are interpreted qualitatively, inexactly, or uncertainly [114]. Designing an FLC includes the definition of the fuzzification and defuzzification processes and the derivation of the database and fuzzy control rules. Visually, Fig. 3.2 shows the block diagram of a generic FLC, which comprises four principal components: a fuzzifier, a defuzzifier, a knowledge base and an inference engine. Firstly, the fuzzifier converts input data into suitable linguistic values, which may be viewed as labels of fuzzy sets. Secondly, the knowledge base is a database and a collection of linguistic statements based on expert knowledge, which is usually expressed in the form of IF-THEN rules. Thirdly, the inference engine performs inference to compute a fuzzy output. Finally, the defuzzifier, which has the opposite meaning of the fuzzifier, provides a non-fuzzy control action from an inferred fuzzy control action. The remaining paragraphs of this section describe in more detail each of these blocks. 49
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES Knowledge base Fuzzification process Defuzzification process Inference engine System (orplant) FuzzyLogic Controller Control action Non-fuzzy input Figure 3.2: Block diagram of an FLC Fuzzification process The fuzzification interface involves the following tasks: (a) measuring the values of input variables, (b) performing a scale mapping that translates the range of values of input variables into the corresponding universes of discourse and (c) finding the fuzzy representation of non-fuzzy input values. In practice, this is achieved through application of the membership functions associated with each fuzzy set defined in the input space. More specifically, the fuzzification process is the assignment of membership values (one for each fuzzy value of the linguistic variable) to a numerical input value. For instance, let consider the linguistic variable “temperature of a room”, which can take the fuzzy values low,medium and high. The input of the system is a crisp value of the temperature of the room, while the output is given by the membership value for each label, as shown in Fig. 3.3. Formally, let Xdenote the universe of discourse for the three fuzzy sets. Hence, the fuzzification process receives the element a∈X, and produces the membership degrees µlow(a),µmedium(a)and µhigh(a). Knowledge base The knowledge base of an FLC is divided into two different blocks: a database and a rule base. The former allows to characterize fuzzy rules and fuzzy data manipulation in the FLC, while the latter provides the dynamic behavior of the FLC through a set of linguistic rules derived from the expert knowledge. Firstly, the database is built from concepts which are subjectively defined and based on experience and engineering judgment. The following aspects are related to the construction of the database in an FLC: •Discretization. It is also referred to as quantization. Its function is to convert a continuous universe into a discrete universe, which is composed of a certain number of segments or quantization levels. In this case, a fuzzy set is defined by assigning membership values to 50
3.1. PARAMETER SELF-TUNING 20 15 25 1 low medium high 0.4 0.6 17 (a) a membershipdegree temperatureoftheroom (universeofdiscourse) fuzzification process non-fuzzy input 17ºC low(0.6) medium(0.4) high(0.0) a (a) Figure 3.3: Fuzzification process through an example each generic element of the new discrete universe. In addition, there exists a trade-off when selecting the number of quantization levels. On the one hand, it should be large enough to provide an appropriate granularity, but, on the other hand, it should be small enough to save memory storage. In this sense, the corresponding mapping that transforms measured variables into values in the discretized universe can be linear, non-linear or both. •Normalization. The normalization of a universe involves a discretization of the universe of discourse into a finite number of segments, with each segment mapped into a suitable segment of the normalized universe. The mapping can also be linear, non-linear or both. •Partition of input and output spaces. A fuzzy partition determines how many fuzzy sets need to be defined. This decision, which determines the granularity of the control achievable by the FLC, depends on the characteristics of the system being controlled and the quality required for the control process. •Completeness. The concept of completeness is related to the fact that the FLC generates an appropriate action for every state in the system. Typically, the completeness is in connection with the design experience and engineering knowledge. •Membership functions. These functions allow to assign the grades of membership to the fuzzy sets. Such an assignment is based on the subjective criteria of the decision. For example, if an input variable is affected by noise, the membership functions should be wide enough to reduce the sensitivity to noise. The membership functions are typically expressed in a functional form, among which are the bell-shaped function, triangle-shaped function and trapezoid-shaped function. Secondly, the rule base comprises a collection of fuzzy rules following a syntax of the type IF-THEN to set the control strategy, i.e.: IF (a set of conditions are satisfied) THEN (a set of consequences can be inferred),(3.5) where the first part of the conditional statement is referred to as the antecedent, while the second part is known as the consequent. In particular, the antecedent is a condition in its application 51
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES domain and the consequent is a control action for the system under control. Furthermore, several linguistic variables can be included in the antecedents and the conclusions of the rules. The proper selection of the state variables included in the antecedent and the control variables included in the consequent is essential to the characterization of the operation of the FLC. An example rule would be “if pressure is very high, then open the valve very much”. The main benefit of using this kind of rules is that these rules provide a natural framework for the characterization of human behavior and decisions analysis. In this sense, many experts state that fuzzy control rules provide a convenient way to express their domain knowledge. To formulate these fuzzy rules, for example, an interrogation of experienced experts or operators by using a carefully organized questionnaire can be used. This explains the fact that FLCs are implemented by using fuzzy IF-THEN rules. Inference engine Once the values of the input variables have been converted to fuzzy values through the fuzzification process, the inference engine identifies which rules are triggered and calculates the fuzzy values of the output variables. In other words, this process links the fuzzified inputs to the rule base in order to produce a fuzzified output for each rule. This means that, a degree of membership has to be determined for the output sets which are part of the consequents in the fuzzy rules. Such a degree of membership is calculated from the degrees of membership in the input sets and the relationships between the input sets. These relationships are established by logic operators that combine the sets in the antecedent. The output fuzzy sets in the consequent are then combined to produce one overall membership function for the output of the rule. To explain the inferencing process, let assume that Aand Bare two input fuzzy sets in the universe of discourse X1and Cis a fuzzy set in the universe of discourse X2. Let also consider that the following rule is defined: IF (Ais aand Bis b) THEN (Cis c).(3.6) The values of µA(a)and µB(b)are available for the inference engine, since they have been derived from the fuzzification process. Thus, the inferencing process starts with the calculation of the degree of truth of each rule in the rule base. The degree of truth specifies the triggering strength of a particular rule. It is calculated by combining the antecedent sets using a specific operator, among which are the min-operator and the product operator for the intersection relation as previously stated. In this example, assuming the min-operator, the degree of truth αkfor the rule kis calculated as: αk= min{µA(a), µB(b)}.(3.7) The following step of the inferencing process is to determine a single fuzzy value for each output ci∈Cwhich has been activated. In general, the final fuzzy value corresponding to the output 52
3.1. PARAMETER SELF-TUNING ci, denoted as βi, is computed using the max-operator, i.e.: βi= max ∀k{αki}.(3.8) where αkiis the degree of truth of the rule k, which has activated the output ci. The final result of the inference engine is a set of fuzzified output values. In this sense, the rules that are not activated have a degree of truth equal to zero. In addition, rules can include a weighting factor in the range [0,1] to represent the degree of confidence in that rule. Such factors derived from the expert knowledge are applied when the fuzzy rules are aggregated to produce a non-fuzzy value in the defuzzification process. Defuzzification process This process establishes a relationship between the space of fuzzy control actions defined over the output universe of discourse and the space of crisp (non-fuzzy) control actions. The degree of truth of a rule represents the degree of membership to the sets present in the consequent. Given the degrees of truth from a set of activated fuzzy rules, the defuzzification process converts the output of the fuzzy rules into a scalar (non-fuzzy) value. To calculate such a scalar value, two different approaches can be used. The first approach, based on the well-known Mamdani-type fuzzy rule [115], implements rules in which the consequent is another fuzzy variable [e.g. see (3.6)]. The second one, known as Takagi-Sugeno approach, utilizes rules whose consequent is a polynomial function of the inputs [116]. Mamdani approach. In this approach, there are some methods to find a scalar that represents the action to be taken: •Max-min method. In this case, the rule with the highest degree of truth is selected and then, the activated consequent membership function is determined. Finally, the centroid of the area under that function is calculated and the horizontal coordinate of that centroid is the output of the FLC. •Averaging method. In this approach, the average of the degrees of truth considering all the activated rules is first calculated. Then, each membership function is clipped at the average. Finally, the centroid of the composite area is calculated and its horizontal coordinate is used as the output of the FLC. •Root-sum-square method. The membership functions are scaled such that the peak of each function is equal to the maximum degree of truth that corresponds to that function. Then, the centroid of the composite area under the scaled functions is computed and its horizontal coordinate is the output of the FLC. •Clipped center of gravity method. In this method, the membership functions are clipped at the corresponding degree of truth of the rules. Then, the centroid of the composite area is calculated and the horizontal coordinate is the output of the FLC. 53
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES The calculation of the centroid of the trapezoidal areas depends on the domain (i.e. discrete or continuous) of the membership functions. In the discrete domain, where a finite number of values, denoted as nx, are defined, the output of the defuzzification process is computed as: output =Pnx i=1 xiµC(xi) Pnx i=1 µC(xi).(3.9) where xiis each possible value. In the case of a continuous domain, the output is given by the following expression: output =Rx∈Xxµ(x)dx Rx∈Xµ(x)dx .(3.10) where Xis the universe of discourse. Takagi-Sugeno approach. Formally, a typical rule for this approach follows the generic expression [117]: IF (X1is A1and ... and Xnis An) THEN (Y=p0+p1X1+ ... + pnXn).(3.11) where X1,...,Xnare fuzzy input variables; Aiindicates one of the fuzzy sets defined for the linguistic variable Xi;Yis the output variable, and p0,...,pnare parameters. Thus, the main difference between the Takagi-Sugeno approach and the Mamdani approach is that rather than having a fuzzy consequent, each rule consequent is a mathematical function. Furthermore, this method has been extended to non-linear functions. Given a set of activated rules and their corresponding degree of truth α, the output crisp value is calculated as a weighted average of the rule outputs, i.e.: output =PN i=1 αi·f(X1, ..., Xn) PN i=1 αi ,(3.12) where Nis the number of rules and f(X1, ..., Xn)is some mathematical function of the inputs. The main benefits of the Takagi-Sugeno approach is that a more dynamic control is provided, FLCs are computationally more efficient and best suited to mathematical analysis and they work well with optimization and adaptive techniques. For these reasons, the FLCs proposed in this work are based on this approach. Example To conclude this section, an illustrative example of the operation of an FLC is provided. In particular, the FLC is based on the Takagi-Sugeno approach [116] previously explained. Let suppose that the controller is described by the two rules IF (xis A1and yis B1) THEN (zis f1(x,y)=k1) (3.13) 54
3.1. PARAMETER SELF-TUNING & IF (xis A2and yis B2) THEN (zis f2(x,y)=k2),(3.14) from which the following elements can be identified: •Two input variables, xand y, defined over the universe of discourse Xand Y, respectively. •Two fuzzy sets, A1and A2, defined for the variable x. •Two fuzzy sets, B1and B2, defined for the variable y. •One output variable, z. •Two constant functions, f1and f2, defined for the variable z. The membership functions defined for each fuzzy set of the input variables are shown in Fig. 3.4. Then, the basic operation of the FLC is as follows: •Step 1. The fuzzification process calculates the membership value for each fuzzy set by applying the associated membership function as shown in Fig. 3.5(a). •Step 2. The inference engine computes the degree of truth for each fuzzy rule through the combination of the fuzzified inputs using the min-operator, as shown in Fig. 3.5(b). In particular, the expressions used to calculate the degrees of truth are: α1= min{µA1(x0), µB1(y0)}(3.15) & α2= min{µA2(x0), µB2(y0)}.(3.16) •Step 3. Finally, the defuzzification process calculates the non-fuzzy output as a weighted average of the rule constant outputs. In particular, the equation used to produce the output value is: output =α1·k1+α2·k2 α1+α2 .(3.17) A (x) X 1 A2 x B (y) Y 1B2 y Figure 3.4: Membership functions of the input fuzzy sets in the example 55
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES x1 x2 y(t) ^ y(t) x(t) x(t+1) socialvelocity cognitivevelocity inertia velocity newvelocity Figure 3.7: Determination of the new position of a particle in a particle swarm algorithm [113] describe the rule base and the membership functions in the controller. The objective function used to rate the solutions is defined as a weighted sum of two filtered quality indicators in mobile networks, dropping and blocking rates. Results show that specific strategies for the auto-tuning process can be implemented by correctly choosing the cost function that guides the particle swarm optimization in the FLC. 3.2.5 Reinforcement Learning RL is an area of machine learning based on leading an agent to take actions in an environment in order to maximize a cumulative reward. The two most important features of RL are the trial-and-error search and the fact that actions may affect not only the immediate reward but also the rewards of the successive situations [101]. RL is different from other learning approaches, for example, the supervised learning that is typically used in Neural Networks. In this latter case, learning is carried out by using previously collected examples, also called the training data set. This kind of learning is not suitable for learning from interaction. In addition, in interactive problems, it is usually very difficult to obtain a training data set that is correct and representative of all the situations in which the agent has to act. Thus, in those cases, an agent learning from its own experience remains as the only solution. Beyond the agent and the environment, the following elements can be identified in RL: a policy, a reward function, a value function and, optionally, a model of the environment. Firstly, the policy defines how the agent has to act in a given time. In other words, it is a mapping between perceived states of the environment and actions to be taken from those states. Secondly, the reward function defines the goal in an RL problem. More specifically, it is a mapping between each perceived state and a scalar or reward that indicates the intrinsic desirability of being in that state. However, the objective of the agent is not to obtain the maximum immediate reward, 62
3.2. OPTIMIZATION OF THE SELF-TUNING PROCESS but to maximize the total reward that the agent receives in the long run. For this reason, whereas the reward function indicates what is good in an immediate sense, the value function specifies what is good in the long run. In particular, the value function is a mapping between each perceived state and the total amount of reward that an agent can expect to accumulate over the future, starting from that state. Finally, the model of the environment imitates the behavior of the environment. In this sense, RL can learn by trial-and-error and, at the same time, learn a model of the environment. To illustrate some of the previous concepts, the basic scheme of an RL problem is shown in Fig. 3.8, where a general environment responds at time t+ 1 to the action taken at time t. A key concept in RL is the trade-off between exploration and exploitation. When the agent has to take actions, it should select those actions tried in the past which produced a lot of reward and the only way to discover them is to try actions that have not yet been selected. Thus, there exists a trade-off between exploration and exploitation as the agent must exploit the current knowledge to obtain reward, but it also has to explore other actions which could be best in the long run. In this sense, the objective of the agent is to maximize the received reward in the long run, that is, the sum of the rewards obtained from all situations or states that will be visited in the future: Rt=rt+1 +γrt+2 +γ2rt+3 +... = ∞ X k=0 γkrt+k+1,(3.20) where ris the numerical reward obtained at each time step as a consequence of taking an action and γis the discount rate that determines the importance of future rewards. The Markov property As previously stated, the agent makes its decisions as a function of the state. In this context, there exists an important property of environments and their state signals which is called the Markov property. Prior to its definition, some aspects related to the state signal need to be discussed. The state signal includes all the information which is available to the agent. However, action a reward r environment state sttt st+1 rt+1 agent Figure 3.8: The basic elements in an RL problem 63
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES it should not be expected to inform the agent of everything about the environment, or even everything that would be useful for it in making decisions. In this sense, an appropriate state signal is the one that summarizes past information compactly, but also keeping the relevant part of the information. Roughly speaking, the Markov property is fulfilled when a state signal retains all relevant information. In this situation, the response of the environment at time step t+ 1 depends only on the state and action at time t, in which case the dynamics of the environment can be defined by specifying only: Pr st+1 =s0, rt+1 =rst, at,(3.21) where Pr{·} denotes the probability of its argument, sthe state of the environment, s0any state in the system, rthe received reward and athe action taken by the agent. If an environment has the Markov property, then it is possible to predict the next state and expected next reward given the current state and action. An RL task that satisfies the Markov property is called a Markov decision process (MDP). If the state and action spaces are finite, then it is called a finite MDP, which is defined by a set of states, a set of actions and the dynamic of the environment. The latter is specified by the so-called transition probabilities and the expected value of the next reward. Given any current state sand action a, the transition probability of each possible next state s0is: Pa ss0=Pr st+1 =s0st=s, at=a.(3.22) Similarly, given any current state sand action a, together with any next state s0, the expected value of the next reward is: Ra ss0=Ert+1st=s, at=a, st+1 =s0.(3.23) where E{·} means the expected value of its argument. These two quantities, i.e. the transition probabilities and the expected value of the next reward, define the most important aspects of the dynamics of a finite MDP. Optimal value functions Most RL algorithms search for value functions that estimate how beneficial it is for the agent to be in a given state. As previously stated, the value of a state sis given by the expected cumulative reward that can be received from such a state. A state-value function, named V(s), is defined in RL to determine the benefit of being in a state s. The value obviously depends on the states visited by the agent, which in turn depends on the policy followed. A policy function πis a mapping from states to actions in order to determine the behavior of the agent, where π(s, a)is the probability of taking action afrom state s. In this manner, the value of a state s 64
3.2. OPTIMIZATION OF THE SELF-TUNING PROCESS following the policy πis defined by: Vπ(s) = Eπ{Rt|st=s} =Eπ(∞ X k=0 γkrt+k+1 st=s),(3.24) where Eπ{·} means the expected value under policy π. Similarly, an action-value function, named Q(s, a), is defined in RL to quantify the value of taking action a, when starting from state s. If the agent follows the policy π, then it is formally expressed as: Qπ(s, a) = Eπ{Rt|st=s, at=a} =Eπ(∞ X k=0 γkrt+k+1 st=s, at=a).(3.25) The two previous functions, Vπand Qπ, can be estimated from experience. In addition, an important property of these functions is that they satisfy particular recursive relationships. More specifically, for any policy πand any state s, the following condition holds between the value of sand the value of its possible successor states: Vπ(s) = Eπ{Rt|st=s} =Eπ(∞ X k=0 γkrt+k+1 st=s) =X a π(s, a)X s0 Pa ss0"Ra ss0+γEπ(∞ X k=0 γkrt+k+2 st+1 =s0)# =X a π(s, a)X s0 Pa ss0Ra ss0+γV π(s0),(3.26) which is called the Bellman equation for Vπ. Moreover, the value function Vπis the unique solution to its Bellman equation. Solving an RL problem is equivalent to find a good policy that gives a high reward in the long-term. An optimal policy always has an expected value greater (or equal) than other policies for all states. Likewise, the optimal policies share the same state-value and action-value functions, called V∗and Q∗, respectively. In particular, the optimal state-value function V∗is defined as: V∗(s) = max πVπ(s),(3.27) for all s∈S, where Sis the set of states. Similarly, the optimal action-value function Q∗is defined as: Q∗(s, a) = max πQπ(s, a),(3.28) 65
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES for all s∈Sand a∈A(s), where A(s)is the set of possible actions in state s. Since the function Q∗provides the expected return for taking action ain state sand thereafter following an optimal policy, it can be expressed in terms of V∗as follows: Q∗(s, a) = E{rt+1 +γV ∗(st+1)|st=s, at=a}.(3.29) The Bellman equation for V∗can be rewritten without making reference to any specific policy. In that case, it is called the Bellman optimality equation, which expresses that the value of a state under an optimal policy must be equal to the expected return for the best action taken from that state, i.e.: V∗(s) = max a∈A(s)Qπ∗(s, a) = max aEπ∗{Rt|st=s, at=a} = max aE{rt+1 +γV ∗(st+1)|st=s, at=a} = max aX s0 Pa ss0Ra ss0+γV ∗(s0).(3.30) The Bellman optimality equation for Q∗is: Q∗(s, a) = Ert+1 +γmax a0Q∗(st+1, a0) st=s, at=a =X s0 Pa ss0Ra ss0+γmax a0Q∗(s0, a0).(3.31) For finite MDPs, the Bellman optimality equation has a unique solution independent of the policy. In fact, the Bellman optimality equation is a system of equations, one for each state, so that if there are Nstates, then there are Nequations with Nunknown variables. If the dynamic of the environment is known (i.e. Pa ss0and Ra ss0are available), then in principle this system of equations for V∗can be solved by using any method for solving systems of non-linear equations. Once the system of equations has been solved, it is relatively easy to find an optimal policy. If V∗is available, then those actions that appear best in the next step will be optimal actions. In other words, any policy that is greedy with respect to V∗is an optimal policy. The importance of V∗lies on the fact that if it used to evaluate the short-term consequences of actions, then a greedy policy is in fact optimal in the long-term, since V∗already takes into account the reward consequences of all possible future behaviors. If Q∗is available, the selection of optimal actions is still easier because the agent does not have to search for actions for the next step, but simply to find any action that maximizes Q∗(s, a). Thus, at the expense of representing a function of state-action pairs [i.e. Q∗(s, a)], instead of just of states [i.e. V∗(s)], the optimal action-value function allows to select optimal actions without the need of knowing anything about possible successor states and their corresponding values (i.e. the dynamic of the environment). 66
3.2. OPTIMIZATION OF THE SELF-TUNING PROCESS By solving the Bellman optimality equation, it is possible to provide a way to find an optimal policy and, thus, to solve the RL problem. However, such a solution is rarely directly useful. In practice, there are three assumptions that are rarely satisfied: (a) accurate knowledge of the dynamic of the environment, (b) enough computational resources to complete the computation of the solution, and (c) the Markov property. To solve the problem in an approximate way, many different decision-making methods can be applied, for example, heuristic search methods and dynamic programming. In this context, many RL methods can be clearly seen as an approximate way of solving the Bellman optimality equation, using actual experienced transitions in place of knowledge of the expected transitions. The three most important classes of methods for solving an RL problem are dynamic programming, Monte Carlo methods and temporal-difference methods. Each class of methods has both advantages and disadvantages. In particular, dynamic programming methods, which aim to solve the Bellman equation, are well developed mathematically, but a complete and accurate model of the environment is required. Monte Carlo methods attempt to estimate value functions and discover optimal policies. They are conceptually simple and a model is not required, but they are not appropriate for step-by-step incremental computation. To explain this, note that Monte Carlo methods are based on averaging sample returns, so that this kind of methods is only applicable for episodic tasks. Due to this, experience is divided into episodes and it is only upon the completion of an episode that value estimates and policies are changed. For this reason, it is said that Monte Carlo methods are incremental in an episode-by-episode sense, but not in a step-by-step sense. Finally, temporal-difference methods are fully incremental and a model is not required, but they are more complex to analyze. The three classes of methods also differ in some other aspects, such as the efficiency and speed of convergence, and they can be combined in order to obtain the benefits of each one. Q-Learning algorithm A mechanism for finding an optimal policy consists of following a generalized policy iteration, based on alternating two interacting processes: policy evaluation and policy improvement. The policy evaluation brings the value function closer to that for the current policy, while the policy improvement uses this new value function to enhance the policy in terms of expected value. This concept is illustrated in Fig. 3.9. The result of such an iterative process is that both policy and value function approach to optimality. As previously stated, an RL method appropriate for step-by-step incremental computation is the temporal-difference learning. In this methodology, policy evaluation is carried out from the observed reward, rt+1, and the estimate V(st+1). To calculate the new state-value function, the following expression can be used: V(st)←V(st) + η[rt+1 +γV (st+1)−V(st)],(3.32) where ηis the step-size parameter or learning rate involved in the incremental update. If the 67
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES greedy () ** Figure 3.9: Basic scheme of a generalized policy iteration [101] action-value function is used instead, then it is calculated by: Q(st, at)←Q(st, at) + η[rt+1 +γQ(st+1, at+1)−Q(st, at)].(3.33) Policy improvement is achieved by selecting actions whose current action-value is the greatest from that state, that is, making the policy greedy by: a(s) = arg max kQ(s, k).(3.34) The overall process converges to both the optimal value function and an optimal policy if all state-action pairs are visited an infinite number of times and the policy becomes greedy in the limit [101]. Q-Learning is a popular temporal-difference algorithm in which the learned Q(s, a)directly approximates the optimal Q∗(s, a)independently of the policy followed by the agent [129]. The update of the action-value function corresponds to the equation: Q(st, at)←Q(st, at) + η[rt+1 +γmax aQ(st+1, a)−Q(st, at)].(3.35) In this case, Qapproximates the optimal action-value function, Q∗, without depending on the policy followed. Adaptation of Q-Learning to FLCs The rule base of an FLC can be optimized by applying RL techniques. In particular, a fuzzy version of the Q-Learning algorithm is proposed in [130] to optimize the consequent part of fuzzy 68
3.2. OPTIMIZATION OF THE SELF-TUNING PROCESS rules in an FLC. Such an adaptation of Q-Learning allows to process continuous state and action spaces by a simple discretization of the action-value function. Consequently, the so-called discrete q-values can be stored in a look-up table as only a finite set of state-action values is needed. An additional advantage of this approach is that prior knowledge can be easily introduced in the fuzzy rules speeding up the learning process. The discretization of the action space forces the agent (i.e. the FLC) to choose one action (a consequent) among Jfor rule i. Suppose that there are Nfuzzy rules defined for the FLC. Let a[i, j]be the jth possible action in rule iand q[i, j]its associated q-value stored in the look-up table. Hence, the representation of the continuous Q(s, a)is equivalent to determine the q-values for each rule consequent, then to interpolate for the continuous input vector. More specifically, the fuzzy Q-Learning algorithm is implemented by the following steps: 1. Initialize the q-values in the look-up table. The following assignment is usually used when there is no prior knowledge: q[i, j] = 0,1≤i≤N and 1≤j≤J, (3.36) where q[i, j]is the q-value, Nis the number of rules and Jis the number of actions per rule. 2. Select an action for each activated rule iwith nonzero degree of truth. For instance, actions can be selected using the so-called -greedy policy, i.e.: ai= arg max kq[i, k]with probability 1−, (3.37) or ai=random{ak, k = 1,2, ..., J}with probability , (3.38) where aiis the consequent of rule iand is a parameter that establishes the trade-off between exploration and exploitation in the algorithm (e.g. = 0 means that there is no exploration, that is, the best action is always selected). 3. Calculate the global action inferred by the FLC: a(t) = N X i=1 αi(s(t)) ·ai(t),(3.39) where a(t)is the inferred action at time step t,αi(s(t)) is the degree of truth for rule iand ai(t)is the selected action for that rule. The degree of truth is the distance between the input state s(t)and the rule i, calculated as: αi(s(t)) = L Y j=1 µij(sj(t)),(3.40) 69
CHAPTER 3. ADAPTIVE OPTIMIZATION TECHNIQUES where Lis the number of FLC inputs, µij(sj(t)) is the membership function for the jth FLC input and rule i. 4. Approximate the Q-function from the current q-values and the degree of truth of the rules: Q(s(t), a(t)) = N X i=1 αi(s(t)) ·q[i, ai],(3.41) where Q(s(t), a(t)) is the value of the Q-function for the state s(t)and the action a(t)in iteration t. 5. Leave the system to evolve to the next state, s(t+ 1). 6. Observe the reinforcement signal, r(t+1), and compute the value of the new state denoted by Vt(s(t+ 1)), i.e.: Vt(s(t+ 1)) = N X i=1 αi(s(t+ 1)) ·max kq[i, ak].(3.42) 7. Calculate the error signal: ∆Q=r(t+ 1) + γ·Vt(s(t+ 1)) −Q(s(t), a(t)),(3.43) where r(t+ 1) is the reinforcement signal, γis a discount factor, Vt(s(t+ 1)) is the value of the new state and Q(s(t), a(t)) is the value of the Q-function for the previous state and the action performed from that state. 8. Update q-values by an ordinary gradient descent method: q[i, ai]←q[i, ai] + η·∆Q·αi(s(t)),(3.44) where ηis a learning rate. 9. Repeat the above-described process starting from step 2 for the new current state until the convergence of the algorithm is achieved. Once the Q-Learning algorithm has finished, the fuzzy rules are generated by selecting those consequents that obtained the highest q-value in the look-up table. As a summary, Algorithm 3.2 briefly describes the steps of the Fuzzy Q-Learning algorithm. Finally, as stated in Chapter 2, several works applying both non-fuzzy and fuzzy Q-Learning algorithms in wireless network optimization problems are available in the literature, showing the effectiveness of the combination of FLCs and Q-Learning in this context. 70
3.2. OPTIMIZATION OF THE SELF-TUNING PROCESS Algorithm 3.2. Fuzzy Q-Learning 1. Initialize q-values: q[i, j] = 0,1≤i≤N and 1≤j≤J. 2. Select an action for each activated rule (-greedy policy): ai= arg maxkq[i, k]with probability 1−, ai=random{ak, k = 1,2, ..., J}with probability . 3. Calculate the global action: a(t) = N P i=1 αi(s(t)) ·ai(t). 4. Approximate the Q-function from the current q-values and the degree of truth of the rules: Q(s(t), a(t)) = N P i=1 αi(s(t)) ·q[i, ai]. 5. Leave the system to evolve to the next state, s(t+ 1). 6. Observe the reinforcement signal, r(t+ 1), and compute the value of the new state denoted by Vt(s(t+ 1)): Vt(s(t+ 1)) = N P i=1 αi(s(t+ 1)) ·maxkq[i, ak]. 7. Calculate the error signal: ∆Q=r(t+ 1) + γ·Vt(s(t+ 1)) −Q(s(t), a(t)). 8. Update q-values by an ordinary gradient descent method: q[i, ai]←q[i, ai] + η·∆Q·αi(s(t)). 9. Repeat the above-described process starting from step 2for the new current state until the convergence is achieved. 3.2.6 Justification of the selected technique Amongst the techniques explained in the previous sections, RL has been the selected one in this thesis. The main reasons for discarding the other alternatives as well as the reasons for choosing an RL method are discussed in the following paragraphs. Firstly, although Neural Networks have been successfully applied in many applications, this artificial intelligence technique also has some limitations and disadvantages. On the one hand, neural networks are especially appropriate for prediction, function approximation, classification, pattern recognition, and clustering, which are not the problems tackled in thesis, mainly focused in developing control techniques. On the other hand, an important drawback is that neural networks require a large diversity of training for real-world operation, which can be a severe constraint in complex systems such as wireless networks. In addition, neural networks cannot be trained a second time, in the sense that it is very hard to add new data to an existing network. Finally, they require a lot of computational resources and high processing time for large neural networks. Secondly, although genetic algorithms are a method very easy to understand which practically does not demand a high level of skills in mathematics and they are easily transferred to 71
CHAPTER 4. SIMULATIONS TOOLS −1 0 1 2 3 4 5 6 7 −1 0 1 2 3 4 5 6 x (km) y (km) 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 Figure 4.2: Simulation scenario Figure 4.3: Simulation scenario with wrap-around Fig. 4.3 shows the simulation scenario with the wrap-around technique. Finally, it is necessary to define the set of interfering cells for each cell of the scenario. For each cell, an ordered set of interfering cells are constructed in terms of power received from each interferer by static system-level simulations. Spatial Traffic Distribution Users can be spatially distributed in both an uniform or non-uniform way over the scenario. In the case of uniform spatial distribution, users are located in whatever point of the scenario with the same probability. However, to reproduce a realistic situation, it is recommended to use a non-uniform distribution. 78
4.1. THE LTE MACRO SYSTEM-LEVEL SIMULATOR Figure 4.4: Spatial traffic distribution The typical spatial distribution in urban areas can be described by a log-normal distribution at a cell level. The traffic is created adding to such a distribution a Gaussian random variable [142]. Fig. 4.4 shows the probability of starting a call in any location of the scenario. It is observed that the spatial traffic distribution has a central peak, creating a congested area with higher traffic density. It is noted that the spatial traffic distribution will be slightly affected by the mobility model, explained in the next section. Mobility Model The proposed mobility model does not define any constraint about user directions, as users can freely move over the scenario. More precisely, the mobility model considers random constant paths for the users in the simulation scenario. Users move at constant speed, set to 3, 10 or 50 km/h. This model also includes the effect of the wrap-around technique, which means that when a user reaches the limit of the original scenario, it appears in the correct position of this scenario. Traffic Model The service modeled in the simulator is VoIP. This service is defined as a source generating packets of 40 bytes every 20 ms [143], reaching a bit rate of 16 kbps. As it will be described later, the radio resource allocation in the simulator is performed for time intervals of 10 ms. For this reason, the voice service has been implemented as users that transmit packets of 20 bytes every 10 ms. For this service, it is also necessary to determine when a call is dropped, that is, when the service is interrupted. Such an event occurs when a user does not receive packets during a specific time interval. In particular, user packets are not scheduled when the connection quality 79
CHAPTER 4. SIMULATIONS TOOLS is below a certain threshold or there are not enough resources, so the call may be dropped. 4.1.2 Physical layer Channel model The mobile radio channel can be described as a time-varying linear filter [144]. Therefore, it can be represented in the time domain by its impulse response, h(τ, t), where τstands for delay of each path in h, and the amplitude of each path varies with time t. Also, the channel can be characterized by the time-variant transfer function, H(f, t), which is related with impulse response through the Fourier transform with respect to the delay variable τ. When the behavior of the channel is randomly time variant, the above-mentioned channel functions become stochastic processes. A realistic approach to the statistical characterization of such a channel may be accomplished in terms of correlation of channel functions since it enables channel output autocorrelation to be determined. Channel autocorrelation functions are related through Fourier transform as well. For typical physical channels, time fading statistics can be assumed stationary over short periods of time and channel correlation function is invariant under a translation in time t, thus being categorized as wide-sense stationary (WSS). In addition, frequency-selective behavior is stationary in frequency fbeing the autocorrelation function invariant under frequency translations. This condition is termed uncorrelated scattering (US), and most practical channels satisfy it fairly well. Autocorrelation functions of wide-sense stationary uncorrelated scattering (WSSUS) channels exhibit the property that the time-variant transfer function autocorrelation is stationary both in time tand frequency fvariables, i.e. its value does not depend on the absolute time or frequency considered but only on the time or frequency shift between time or frequency points of observation. As a consequence, a WSSUS channel can be simulated generating the impulse response, h(τ, t), with stationary variation in time tfor each path and no cross-correlation between different values of delay τ(i.e. generating independent stochastic processes for different paths). Stationarity is achieved by applying Doppler filters to the amplitude time tvariation on each path. These filters perform spectrum shaping according to Doppler effect experimented by any radio signal propagating from a transmitter to a moving receiver (or vice versa). Afterwards, the frequency transfer function, H(f, t), can be computed easily by applying the Fourier transform to the impulse response with respect to delay variable. To provide the possibility of simulating non-constant speed mobiles in the future, fading realizations cannot be performed over time as an independent variable. Alternatively, space variables have to be used so that channel varies according to the current position of the mobile at each iteration of simulation. Therefore, a fading channel spatial grid has been generated. This 80
4.1. THE LTE MACRO SYSTEM-LEVEL SIMULATOR grid provides channel responses for every physical position in the simulated scenario, regardless of mobiles speed. Narrow band fading grid is generated to get a Lord Rayleigh universe [145]. In other words, following Clarke’s model [144], a spatial bidimensional complex Gaussian variable is filtered by a bidimensional Doppler filter. The bandwidth of 2-D Doppler filter can be obtained as a function of spatial grid resolution and wavelength size. Once narrowband channel behavior for each spatial position is obtained, extension to wideband is possible performing the same procedure for every path in power delay profiles described in the specification for Extended Typical Urban (ETU), Extended Pedestrian A (EPA) and Extended Vehicular A (EVA) channels in [146]. Thus, different (uncorrelated) Rayleigh universes are generated for each delay in wideband channel scenario. This results in a distance-variant impulse response h(τ, d)(autocorrelation) of the channel instead of a time-variant impulse response h(τ, t)described in [144] as one of the four system functions for complete WSSUS channel characterization. The only difference is the time to distance (tto d) variable change made. A realization of the function is shown in Fig. 4.5. Since the simulator requires channel realizations for different frequency bands (corresponding to OFDM subcarriers), the distance-variant impulse response has to be transformed into a distance-variant transfer function H(f, d)at each position, by applying Fourier transform with respect to delay variable τ. An example of this function can be seen in Fig. 4.6. The only remaining step is to extend the space variable dof the generated function H(f, d)to a bidimensional (x, y)space variable, obtaining H(f, x, y), a tridimensional function that provides frequency response for each spatial position given by coordinates, xand y. Figure 4.5: Generated bidimensional channel impulse response for ETU channel model in [146] 81
CHAPTER 4. SIMULATIONS TOOLS Figure 4.6: Generated distance-variant transfer function for ETU channel model in [146] Radio propagation channel The simulator includes two alternatives for obtaining propagation calculations. As a first option, the calculations are performed at each iteration and whenever necessary (e.g. in the function that evaluates the channel conditions of each link or in the admission control function). Alternatively, propagation calculations are made from a set of pre-computed matrices. In this case, it is not necessary to perform the calculations during the simulation. For the definition of the pre-computed propagation matrix, the scenario is divided into a grid, whose resolution is given by the correlation distance of the slow fading (20 m). To know the values of the propagation loss a user is experiencing, it is only necessary to read the position of the matrix corresponding to the position occupied by the user in the scenario relative to every base station and then interpolate it with other values of the matrix depending on the relative position in the grid. The propagation matrices include the path loss calculations and the slow fading. In both options, the radio propagation model is the COST 231 extension of Okumura-Hata model [147]. This model is applicable for frequencies in the range from 1500 to 2000 MHz. The effective height of the base station or eNB antenna has been set to 30 m, while the effective height of the UE antenna has been set to 1.5 m. With these assumptions and setting the operating frequency to 2 GHz, the expression for the propagation loss as a function of the distance is given by: L= 134.79 + 35.22 log d, (4.1) where drepresents the distance in km between the UE and the eNB which the user is connected to. 82
4.1. THE LTE MACRO SYSTEM-LEVEL SIMULATOR In addition to the propagation loss, the simulator includes a slow fading model based on the fact that the local average of the radio signal envelope can be modeled by a log-normal distribution, i.e. the local average, in dB, is a Gaussian random variable. The standard deviation of the distribution depends on the considered environment. A typical value for the macrocell urban area analyzed is 8 dB [148]. For the choice of the propagation matrices, the value of shadowing is included in these matrices. The other alternative requires some additional calculations. The dynamic nature of the simulator leads to the implementation of a correlation model between the successive samples which represent the slow fading. An ARMA(1,1) model [149] has been selected for the simulator in this work, zt=θzt−1+ (1 −θ)at,(4.2) where ztrepresents the slow fading sample at the current simulation step, zt−1is the slow fading sample at the previous simulation step, atis a Gaussian random variable uncorrelated with zt and θand (1 −θ)are the coefficients of the ARMA(1,1) model. The coefficients of this model are determined from the probability that a user terminal suffers fading caused by the same obstacle at the time interval ∆/v. That probability can be modeled as an exponential distribution: θ=P(τ < ∆/v) = exp(−∆·λ),(4.3) where ∆is the distance moved by the user terminal at a time interval, vis the UE velocity and λis the interruption rate of the line of sight. The interruption rate of the line of sight, λ, is the inverse of the correlation distance. A typical value of the correlation distance for the macrocell urban area simulated is 50 m [150]. Finally, the Gaussian random variable, at, must be defined based on its mean and standard deviation. This variable provides a statistical distribution of zero mean and a standard deviation, σa, that relates to the standard deviation of the slow fading, σz, as follows: σ2 z=sinh(∆ ·λ/2) cosh(∆ ·λ/2) ·σ2 a.(4.4) Once the propagation calculations have been carried out, it is possible to study the link quality experienced by each user in terms of SIR. The next section describes the process to calculate the value of SIR for each user. 83
CHAPTER 4. SIMULATIONS TOOLS 4.1.3 Link layer SIR calculation The SIR is a representative measurement of the link quality that the user is experiencing. To calculate the SIR in the simulator, it is first necessary to calculate the interference experienced by each user. It is assumed that intra-cell interference is negligible in LTE because the scheduler assigns different frequencies and time slots to each user. Thus, only co-channel inter-cell interference due to the interfering cells using the same subcarriers is considered. This requires knowing the signal arriving to each user from all interfering cells. To calculate the interference from each base station to the terminal, the channel response is not taken into account, but only the path loss and slow fading are considered here. The SIR calculation for a given subcarrier k,γk, is computed using the expression proposed in [151], γk=P(k)ׯ G×N N+Np×RD NSD/NST ,(4.5) where P(k)represents the frequency-selective fading power profile value for the kth subcarrier, ¯ G includes the propagation loss, the slow fading, the thermal noise and the experienced interference, Nis the Fast Fourier Transform size used in the OFDM signal generation, Npis the length of the cyclic prefix, RDindicates the percentage of maximum total available transmission power allocated to the data subcarriers, NSD is the number of data subcarriers per TTI and NST is the number of total useful subcarriers per TTI. If it is assumed that the multipath fading magnitudes and phases are constant over the observation interval, the frequency selective fading power profile value for the kth subcarrier can be calculated using the expression: P(k) = paths X p=1 MpApexp (j[θp−2πfkTp]) 2 ,(4.6) where pis the multipath index, Mpand θprepresent the amplitude and the phase values of the multipath fading respectively, Apis the amplitude value corresponding to the long-term average power for the pth path, fkis the relative frequency offset of the kth subcarrier within the spectrum, and Tpis the relative time delay of the pth path. In addition, the fading profile is assumed to be normalized such that E[P(k)] = 1. The value of ¯ Gis calculated from the expression: ¯ G=Pmax gn(UE)×gUE P LUE,n×SHU E,n Pnoise +PN k=1,k6=nPmax ×gk(UE)×gUE P LUE,k×SHUE,k ,(4.7) 84
4.1. THE LTE MACRO SYSTEM-LEVEL SIMULATOR where gn(UE)is the antenna gain of the serving base station in the direction of the user UE, gUE is the antenna gain of the user terminal, Pnoise is the thermal noise power, P LUE,k is the propagation loss between the user and the eNB k,SHUE,k is the loss due to slow fading between the user and the eNB kand Nis the number of interfering eNBs considered (set to 43 in the simulator). A PRB is the minimum amount of resources that can be scheduled for transmission in LTE. As a PRB comprises 12 subcarriers, it is necessary to translate those SIR values previously calculated for each subcarrier into a single scalar value. This can be made using the Exponential Effective SIR Mapping, which is based on computing the effective SIR by the equation: SIReff =−βln 1 Nu Nu X k=1 exp −γk β!,(4.8) where βis a parameter that depends on the MCS used in the PRB [152] assuming that all subcarriers of the PRB have the same modulation and Nuindicates the number of subcarriers used to evaluate the effective SIR. The values of βhave been chosen so that the block error probability for all the subcarriers are similar to those obtained for the effective SIR in a additive white Gaussian noise (AWGN) channel [153]. The value of βfor a particular MCS is shown in Table 4.1. Table 4.1: Values of βdepending on the Modulation and Coding Scheme Modulation Coding βfactor QPSK 1/3 1.49 QPSK 2/5 1.53 QPSK 1/2 1.57 QPSK 3/5 1.61 QPSK 2/3 1.69 QPSK 3/4 1.69 QPSK 4/5 1.65 16QAM 1/3 3.36 16QAM 1/2 4.56 16QAM 2/3 6.42 16QAM 3/4 7.33 16QAM 4/5 7.68 64QAM 1/3 9.21 64QAM 2/5 10.81 64QAM 1/2 13.76 64QAM 3/5 17.52 64QAM 2/3 20.57 64QAM 17/24 22.75 64QAM 3/4 25.16 64QAM 4/5 28.38 85
CHAPTER 4. SIMULATIONS TOOLS Once the effective SIR has been calculated, the BLER showing the connection quality can be derived. There exist curves that establish the relationship between the values of SIR and BLER defined for an AWGN channel for every modulation and coding rate combination. These curves can also be used to calculate the BLER because inter-cell interference is equivalent to AWGN as the value of βhas been selected for this purpose. Then, given the value of BLER and taking into account the MCS used in the transmission, it is possible to calculate the value of throughput, Ti, for each user as follows: Ti= (1 −BLER(SIRi)) ×Di TTI,(4.9) where Diis the data block payload in bits [154], which depends on the MCS selected for the user in that time interval and BLER(SIRi)is the value of BLER obtained from the effective SIR. Link Adaptation Before explaining the Link Adaptation function, the 3GPP standardized parameter known as CQI needs to be described. Such an indicator represents the connection quality in a subband of the spectrum. The resolution of the CQI is 4 bits, although a differential CQI value can be transmitted to reduce the CQI signaling overhead. Thus, there is only a subset of possible MCS corresponding to a CQI value [155]. QPSK, 16QAM and 64QAM modulations may be used in the transmission scheme. In the simulator, the CQI is reported by the user to the base station each iteration (100 ms). Based on CQI values, the link adaptation module selects the most appropriate MCS to transmit the information on the PDSCH depending on the propagation conditions of the environment. To quantify the link quality for each user and for each subband of the spectrum, the CQI index is used to provide this information. If the experienced BLER value is required to be smaller than a specific value given by the service, it is possible to establish a SIR-to-CQI mapping that allows to select the most appropriate MCS from a given value of SIR [140]. The standard 3GPP defines a 5-bit MCS field of the downlink control information to identify a particular MCS. This leads to a greater variety of possible MCSs. For simplicity, the developed LTE simulator includes only the same set of MCS given by the CQI index. From the effective SIR value, the index CQI is calculated and the MCS can be determined for the next time interval. Resource Scheduling The Resource Scheduling can be decomposed into a time-domain and frequency-domain scheduling. On the one hand, it is necessary to determine which user transmits at the following time interval. On the other hand, the frequency-domain scheduler selects those subcarriers within the system bandwidth whose channel response is more suitable for the user transmission. For this purpose, the channel response for each user and for each subcarrier of the system bandwidth has to be estimated. Such a piece of information is given by the channel realizations generated in 86
4.1. THE LTE MACRO SYSTEM-LEVEL SIMULATOR the initialization phase of the simulation, assuming a perfect estimation of the channel response. To select the most appropriate frequency subband for the user, the CQI index is used. The developed simulator includes different strategies for radio resource scheduling. In all of them, the CQI parameter gives the information of the channel quality experienced by each user. Likewise, scheduling is done for each cell at each iteration following the configured strategy [156]. The scheduling algorithms implemented in the simulator are: •Best Channel Scheduler (BC): in this scheduler, both time-domain and frequency-domain scheduling are done for a more efficient use of resources. At each iteration, all users are sorted based on the quality experienced for each PRB, which is obtained from CQI values. Once the users are sorted, the allocation will proceed until there are not available radio resources or no more users to transmit. The resource allocation is made following the expression: ˆ i[n] = arg max i{rik[n]},(4.10) where ˆ iis the selected user iand rik is the estimated achievable throughput for PRB k and user iobtained from the CQI. This scheduling algorithm maximizes the overall system efficiency because the resource allocation is done looking for the combinations PRB-user with better channel conditions. The disadvantage of this algorithm is that harms users with bad channel conditions. Thus, if a user is far from the serving eNB or it has a deep fading for prolonged periods of time, it cannot be scheduled and it can suffer significant delays. •Round Robin to Best Channel Scheduler (RR-BC): this scheduler uses different strategies for time-domain and frequency-domain scheduling. For time-domain scheduling, the Round Robin method is applied. Thus, users are selected cyclically without taking into account the channel conditions experienced by each of them. Then, each PRB is assigned to the user with a higher potential transmission rate for that PRB (transmission rate is estimated based on the user’s CQI value for each PRB). At each iteration and for each base station, the expressions to be evaluated are: ˆ i[n+ 1] = (ˆ i[n] + 1) mod Nu(4.11) and ˆ k[n] = arg max k{rik[n]},(4.12) where ˆ iis the selected user, Nuis the number of users and ˆ krepresents the PRB selected. In this case, the goal is to maximize system efficiency, but trying not to harm users with unfavorable channel conditions. •Large Delay First to Best Channel Scheduler (LDF-BC): this scheduler is similar to the 87
CHAPTER 4. SIMULATIONS TOOLS 45 50 55 60 65 70 75 80 85 0 2 4 6 8 10 12 14 16 18 20 Traffic load level (%) Call Dropping Ratio (%) RR−BC LDF−BC Figure 4.8: CDR as a function of the traffic load for two scheduling schemes 45 50 55 60 65 70 75 80 85 0 1 2 3 4 5 6 7 8 9 10 Traffic load level (%) Call Blocking Ratio (%) RR−BC LDF−BC Figure 4.9: CBR as a function of the traffic load for two scheduling schemes computationally-efficient dynamic system-level simulator for enterprise LTE femtocells has been developed, including a three-dimensional office scenario, specific mobility and traffic and propagation models for indoor environments. As part of the work in this thesis, the indoor mobility model has been designed for this simulator. The simulation scenario is of 3×2.6 km, comprising three tri-sectorized macrocells in the same site. To avoid border effects in the simulation, the simulator incorporates the wrap-around technique [141]. Inside the coverage area of one macrocell, an office building with dimensions 50 m×50 m has been placed. The number of floors inside the building is configurable. The floor plan is the same for all floors. Fig. 4.10 shows the layout of one of the floors. Magenta circles reflect femtocells positions, lines are the walls (different colors represent their thickness), and 94
4.2. THE LTE FEMTO SYSTEM-LEVEL SIMULATOR Figure 4.10: A floor diagram black diamonds are work stations. The outdoor mobility model adopted for the simulator is very simple, since movement only needs to be reflected in a large-scale. Users move with a random direction that does not change and a constant speed of 3 km/h. The indoor mobility model developed for the simulator as part of this thesis is described in the next section. Finally, more details about the simulator for enterprise LTE femtocell scenarios, such as the indoor propagation model, can be found in [159]. 4.2.1 Indoor mobility model Indoor environments are complex as small user movement may have a strong impact on the signal levels received from base stations. The indoor mobility model developed for the deployment scenario is an extension of the mobility model described in [160]. Users switch between stationary and moving state, but they cannot change floors. A set of location points are defined in the scenario layout so that the user can select them as destination points. Each point has a probability of keeping the user stationary in that location. The time each user spends in the stationary state follows a geometric distribution. Points located in office rooms emulate desks, so that they have associated a high probability of keeping the user there. In contrast, points located in corridors have this probability equal to zero, emulating passing locations for the users. Unlike the approach in [160], where the user switches to the moving state leaving a desk and moving to the corridor, the destination point can be either a desk in the office room or the corridor door. The user speed in the moving state is equal to 1 km/h. If the corridor door is reached, the user selects the next destination point located in the corridor using a non-uniform distribution. Such a distribution depends on how often the user visits each room. For instance, coffee rooms, meeting rooms, and toilets should be more likely visited. Location points within these rooms are treated the same as office rooms. When the mobile reaches a destination point within a room, it is changed into the stationary state. 95
CHAPTER 4. SIMULATIONS TOOLS The user trajectory is a straight line between source and destination. Thus, there must be direct line of sight with the destination point. For this reason, some auxiliary destination points have been defined within the corridor to avoid crossing walls, as the corridor is not rectangular, as in [160]. Finally, it is noted that the user trajectories are pre-computed and stored in a file. Such a piece of information is used as an input during the simulation. Pre-computing trajectories, unlike calculating them in real time during simulation, reduces computationally load. An additional benefit is that pre-computed trajectories can be reused for different experiments, obtaining repeatability. An example of user trajectories based on the proposed mobility model is shown in Fig. 4.11. 4.3 The HSPA/LTE macro/pico system-level simulator To assess the proposed algorithms for TS in the context of HetNets, a dynamic system-level simulator for HSPA/LTE macro/pico scenarios has been used. The utilization of the simulator in this thesis has been exclusively from the user-level perspective. The simulated scenario consists of a heterogeneous cellular environment, whose network layer structure can be configured by the user. In this sense, different RATs (HSPA, LTE), cell sizes (macrocells, picocells) and frequencies can be deployed with the aim of simulating a heterogeneous environment. The mathematical framework supporting the simulator is described in [161]. 4.4 Conclusions In this chapter, firstly, a computationally-efficient dynamic system-level simulator for LTE macrocells has been described. This simulator includes the main characteristics of the RAT as well as the RRM algorithms which provide notable improvements in the efficient use of the available Figure 4.11: User trajectories computed by the indoor mobility model 96
4.4. CONCLUSIONS radio resources. For this purpose, the simulator has been implemented so that simulations require a low computational cost. In addition simulations are composed of epochs to evaluate the modification of network parameters performed by optimization algorithms. At physical and link layers, the design of the simulator has been focused on the calculation of several indicators with the purpose of evaluating the connection quality in a mobile communication. Those indicators are required in the execution of RRM functions. Hence, it is essential that these indicators reflect accurately the behavior of a real network. To achieve this goal, an OFDM channel model has been performed to characterize the temporary and frequency variation of the radio transmission environment for each user during the simulation. The main functions of RRM have also been described in this chapter. At link level, the previous calculated indicators are inputs of the Link Adaptation and Dynamic Scheduling functions. At network level, the main functions are admission control and mobility management, whose parameters can be modified to evaluate optimization algorithms. Simulation results have shown network performance in terms of several indicators for different traffic load levels and scheduling schemes. The next part of this chapter has been devoted to a computationally-efficient dynamic system-level simulator for enterprise LTE femtocells, whose scenario has been briefly described. In addition, the indoor mobility model developed for this simulator has been explained in more detail, since it has been part of the work in this thesis. Finally, a dynamic system-level simulator for HSPA/LTE macro/pico scenarios has also been mentioned in this chapter, as it has been used in the context of this thesis to evaluate the proposed algorithm for TS purposes. 97
CHAPTER 4. SIMULATIONS TOOLS 98
Chapter 5 Mobility load balancing In this chapter, different MLB algorithms are proposed. One group of these algorithms is conceived for macrocells scenarios in which the objective is to alleviate persistent congestion problems, such as a population increase in a city area. The second group of proposed MLB algorithms is designed to solve localized congestion problems in enterprise femtocell scenarios, where a thorough deployment is not typically performed. In addition, the developed algorithms are intended for voice service. Prior to describing the proposed algorithms, an introduction to the mobility parameters and procedures mainly related with the MLB algorithms is provided in Section 5.1. Then, Section 5.2 is devoted to MLB algorithms in macrocell scenarios, while Section 5.3 presents MLB algorithms in enterprise femtocell scenarios. Both sections follow a similar structure, where the specific problem, the proposed solution and its performance assessment are addressed. Finally, Section 5.4 summarizes the major findings of this work. 5.1 Introduction Load balancing is a major issue addressed in the field of SON, which can be solved by sharing traffic between adjacent cells. This problem is tackled here as a Self-Optimizing task that can be solved by tuning specific network parameters, like those involved in the HO process. The HO process is responsible for transferring an ongoing call from one cell to another. By adjusting HO parameters settings, the service area of a cell can be modified to send users to neighboring cells. Thus, the size of the congested cell is reduced while adjacent cells increase in size taking users from the congested cell edge. As a result of a better matching between the spatial distribution of traffic demand and network resources, more users could be accepted in the crowded area so that the call blocking probability would be reduced [54]. In LTE, when a user is in connected mode, the network is responsible for deciding when 99
CHAPTER 5. MOBILITY LOAD BALANCING performing an HO to maintain the ongoing connection. The network also determines for how long each mobile terminal has to send signal measurements back to their serving eNBs, as well as the interval time to perform each measurement. One of the most widely used algorithms for the HO-triggering decision is the Power Budget HO, which is equivalent to the 3GPP A3 event [35]. This algorithm triggers the execution of the HO procedure if the following condition is fulfilled for a specific time period determined by the TTT parameter: RSRPj> RSRPi+HOMi→j,(5.1) where RSRPiand RSRPjare the averaged values of RSRP measured for serving cell iand target cell jrespectively, and HOMi→jis the HOM defined between cell iand cell j. The HOM determines the area where the users connected to a cell would perform an HO toward a neighboring cell. In the context of MRO, the parameter HOMi→jand the symmetric HOMj→i are usually set to the same positive value so that certain symmetric region between the two cells is ensured in order to avoid unnecessary HOs. In Fig. 5.1(a), the RSRP of the neighbor and the serving cells are represented. For instance, if a user connected to cell imoves to cell j, it connects to cell jwhen the RSRP from cell jis equal to the RSRP from cell iplus HOMi→j. As in this case the HOMs are assumed to be symmetric, the same value is applied to the opposite situation (i.e. when the user moves from cell jto i,HOMj→iis used). Maintaining the symmetric region, both HOMs could also be jointly tuned (e.g. HOMi→jis increased while HOMj→iis decreased) so that the service area of these cells is modified, for instance, for MLB purposes. In this case, both HOMs should be modified with the same magnitude to preserve the hysteresis region. However, those variations in HOM should have opposite sign to modify the service area of the two cells. In Fig. 5.1(b), it is observed that HOMi→jhas been increased, while HOMj→i has been decreased. As a result, the service area of the cell iis larger. Finally, it is noted that, since HOMs are defined on an adjacency basis, cell service areas cannot be only re-sized but also re-shaped. 5.2 MLB algorithms for voice service in macrocells This section is devoted to MLB algorithms developed in the context of future wireless networks. In particular, the self-optimization of an FLC to solve persistent congestion problems in macrocell scenarios is investigated. Such a congestion problem is due to an uneven spatial traffic distribution, e.g. when the center of a city becomes crowded. The optimization process is carried out by the fuzzy Q-Learning algorithm with the goal to reduce the blocking probability for voice services while the call dropping is controlled according to network operator constraints. Results are evaluated in a non-uniformly distributed traffic scenario, where the FLC is optimized to fulfill the call dropping constraint. The rest of the section is organized as follows. Firstly, the problem is formulated, where the system model is described, including the main system measurements and involved parameters. After this, the structure of the proposed self-tuning scheme as well as the application of the fuzzy 100
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS cell i RSRP HOMji HOMij cell j cell icell j symmetricregion cell i RSRP HOMji HOMij cell j cell icell j symmetricregion loadbalancing (a)MRO (b)MLB Figure 5.1: Adjustment of HOM for HO optimization and load balancing purposes Q-Learning algorithm to the proposed FLC are explained. The last part of the section describes the simulation setup and discusses the results. 5.2.1 Problem formulation During the last years, cellular networks have experienced a large increase in size and complexity. Generally, network planning provides proper dimensioning of radio resources during the design phase of the RAN. However, as traffic demand changes over time, both traffic demand and network resource dimensioning become misaligned, thus leading to an inefficient use of resources. To cope with such a problem in a cost-effective manner, self-optimizing techniques remains as the best solution rather than adding new resources. To model the misalignment between traffic demand and network resources throughout the 101
CHAPTER 5. MOBILITY LOAD BALANCING time, an assumption made in this study is that the spatial traffic distribution during the network planning stage is uniform, leading to a deployment of cells evenly distributed in the scenario. However, as network evolves, the matching between the spatial distribution of traffic demand and network resources becomes poorer. For this reason, a different spatial traffic distribution is assumed during the operational phase, i.e. when the MLB algorithm is applied to the network. In particular, the center of the selected scenario has become more crowded than the periphery (e.g. emulating the center and the periphery of a small city), so that performance assessment has been carried out based on such a non-uniform spatial traffic distribution. It is worth mentioning that the proposed algorithm can be applied to other non-uniform spatial traffic distributions, since those distributions would be variants of that used in this work, in which the size and the number of hotspots (crowded areas) included in the scenario would be different. As results are expected to be equivalent, a simplified traffic distribution allows a better visualization of the results. Load balancing has an impact on the GoS, which includes call accessibility and maintainability. As a result of a better exploitation of the system capacity, GoS is improved. Although HO-based load balancing is an effective method to share traffic in cellular networks, it may cause negative effects on call maintainability. If the HOM is decreased, the target cell would increase the probability to be more preferred than the serving cell (even if the connection quality is worse), so that some users could be handed over to the target cell. Those users, usually located in the cell edge, will experience worse radio conditions in the target cell as a result of applying such a traffic sharing technique. Thus, negative values of HOMs would increase the risk of dropping. The main contribution of this work is the design of an algorithm to control quality performance for real-time traffic during the load balancing process. Typically, studies found in the bibliography addressing the load balancing problem aim to provide more capacity to a fixed number of users in the system so that they experience higher instant throughput o lower delay, but accessibility is usually not tackled. In future networks, accessibility will be an important feature as services such as VoIP calls are expected to be widely used. In addition, the connection quality loss experienced by some users when load balancing is carried out and controllability of such an effect have not been properly addressed in the literature. This work addresses how much the connection quality can be decreased when load balancing is carried out depending on the operator policy. Thus, flexibility and easiness from the network operator perspective is provided by simply adjusting a single parameter. Unlike the design in [9], the FLC proposed in this work balances traffic load in a intra-system LTE scenario where call dropping becomes a key issue. In addition, HOMs of the standard HO algorithm are adjusted in this work, instead of tuning load thresholds, as proposed in [9]. Finally, another assumption is that the voice call is the service considered in this work, since it is expected to be widely used in the future due to the successful and existing applications based on VoIP. In principle, the proposed algorithm can be applied to any real-time service (voice call, video calls, etc.) for which call dropping is defined as a KPI. For other services (e.g. data services), KPIs such as the throughput could be used to estimate the connection quality of the user. 102
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS Network model and system measurements The proposed self-tuning scheme is applicable to any RAT for next generation cellular networks. In particular, an LTE downlink macro-cellular network providing constant bit rate service is considered here. Each cell is controlled by an eNB (i.e. a base station) and all the cells use the same frequency band. Following the 3GPP standard specifications, the eNBs are interconnected by the X2 interface, enabling direct communication between them. Thus, measurements such as traffic load can be easily exchanged between eNBs over the X2 interface and faster HOs can be performed. The main network level functionalities are admission control and HO. The admission control is responsible for checking the availability of free PRBs in the candidate cell before accepting a call. A ‘worst-case’ criterion has been taken to accept calls, i.e. the user is finally accepted if the highest number of PRBs needed to maintain a connection (worst-case PRB requirement) is less or equal than the number of PRBs available in the candidate cell. If the condition is not satisfied by any candidate cell, then the user connection is blocked. On the other hand, the HO allows user mobility across the network. The call dropping model also plays an important role because the optimization algorithm attempts to control the occurrence of this event. A call is dropped when a percentage of data packets are dropped during a specific time interval. Packet dropping may occur not only because there is a poor connection quality, but also because there are not any available resources to be scheduled. The most important outcome derived from sharing traffic between cells is that the call blocking is reduced, especially in those cells highly loaded. To quantify the call blocking, network operators usually use the CBR, which in this work is defined as in (4.26). Additionally, the load balancing may lead to an increment in the call dropping, which is usually quantified by the CDR, defined here as in (4.25). Since the goal is to solve persistent congestion problems, and not temporary traffic fluctuations, the previous measurements are collected during a long period (i.e. above 15 minutes). As a result, fast changes in network loading are filtered. Note that some situations can lead to a more dynamic traffic distribution, such as the half-time in a football match or a concert. In those cases, faster actions in the order of seconds or few minutes would be needed, so that the previous measurements would not be appropriate due to lack of accuracy and, instead, instantaneous measurements should be employed, for instance, the current load of the system, which is closely related to the call blocking that could be expected. 103
CHAPTER 5. MOBILITY LOAD BALANCING Table 5.2: Candidate consequents for the fuzzy rules Candidate Rule CBRi-CBRjHOM action 1 Unbalancedi→jH EL, VL, Z 2 Unbalancedi→jM EL, VL, L 3 Unbalancedi→jL VL, Z, VH 4 Balanced H VL, Z, VH 5 Balanced M Z 6 Balanced L VL, Z, VH 7 Unbalancedj→iL Z, VH, EH 8 Unbalancedj→iM H, VH, EH 9 Unbalancedj→iH VL, Z, VH L: Low, M: Medium, H: High, EL: Extremely Low, VL: Very Low, Z: Zero, VH: Very High, EH: Extremely High 80% of time the best action is selected). In this case, the FLC starts from a non-optimized set of rules and the Q-Learning algorithm progressively determines the best fuzzy rules for the FLC, preserving system performance. 5.2.3 Simulation setup The dynamic system-level simulator for LTE macrocells explained in Chapter 4 has been used to perform simulations. This simulator first runs a module of parameter configuration and initialization, where a warm-up distribution of users is created to provide meaningful network statistics from the first simulation iteration. Then, several optimization loops or epochs are executed to emulate the tuning process. Each epoch comprises 14,000 simulation steps, equivalent to 23 min of actual network time. Each simulation step includes updating user positions, propagation computation, generation of new calls, and radio resource management algorithms. At the end of each epoch, measurements and reliable statistics are obtained to be used in the following optimization loop. Finally, the main statistics and final results are shown. The simulated scenario includes a macro-cellular environment with a regular layout consisting of 19 tri-sectorized sites evenly distributed in the scenario, as shown in Fig. 4.2. The main simulation parameters are summarized in Table 5.3. Only the downlink is considered in the simulation, as it is the most restrictive link. For simplicity, the service provided to users is the voice call as it is the main service affected by the tuning process. A non-uniform spatial traffic distribution is assumed in order to generate the need for load balancing. It is assumed that, as traffic demand changes over time, both the spatial distribution of traffic demand and the network resource dimensioning have become misaligned. In this scenario, the center has become more crowded than the periphery (e.g. emulating the center and the periphery of a small city). Thus, there are several cells with a high traffic density whereas surrounding cells have a low traffic density. The average traffic load is set to 75%. It is expected that HO parameter changes performed by the FLCs manage to relieve congestion in the affected area. 110
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS Table 5.3: Simulation parameters Parameter Configuration Cellular layout Hexagonal grid, 57 cells (3x19 sites), cell radius 0.5 km Transmission direction Downlink Carrier frequency 2.0 GHz System bandwidth 1.4 MHz Frequency reuse 1 Propagation model Okumura-Hata with wrap-around Log-normal slow fading, σsf = 8 dB and correlation distance=50 m Channel model Multipath fading, EPA model Mobility model Random direction, constant 3 km/h Service model real-time constant bit rate service (voice call), poisson traffic arrival, mean call duration 120 s, 16 kbps Base station model Tri-sectorized antenna, SISO, EIRPmax = 43dBm Scheduler Time domain: Round-Robin Frequency domain: Best Channel Power control Equal transmit power per PRB Link Adaptation Fast, CQI based, perfect estimation Handover Time-To-Trigger=100 ms HOM: [−24,24]dB Traffic distribution Unevenly distributed in space Time resolution 100 TTI (100 ms) Epoch time 23 min Discount factor γ0.95 Learning rate η0.1 Three load balancing methods are simulated. The initial situation is common to all methods, meaning an unbalanced traffic distribution. The first approach is carried out by the non-optimized FLC and it is used as a benchmark. In this case, the inference engine is intuitively designed from the knowledge operator explained in Section 5.2.2. The other two methods are based on the optimization carried out by the fuzzy Q-Learning algorithm also described in Section 5.2.2. The first approach, named as UEE, requires two simulations as the FLC is optimized with 100%exploration (= 1) in a first simulation and then used without the optimization module in a second simulation to evaluate performance. The second approach, named as BEE, requires only a simulation as the FLC is optimized while performance is enhanced. In this case, the optimization is carried out with 20%exploration (= 0.2) and the Q-Learning optimizer entity is teamed with the FLC in the evaluation simulation. Within each epoch, network and optimization parameters such as HOMs, action q-values and FLC consequents, remain unchanged. The learning rate ηcontrols the speed of the learning process and it is set to 0.1 in order to provide a good trade-off between accumulated and fresh q-values. The discount factor γis set to 0.95 to take into account future rewards when performing 111
CHAPTER 5. MOBILITY LOAD BALANCING an action. 5.2.4 Performance results To compare the developed FLCs, a utility function named Uis proposed: U= [CBR + (1 −CBR)·CDR]·100,(5.7) which is a metric that aggregates both KPIs to provide an estimation of the user dissatisfaction. Such indicators, CBR and CDR, consider the total number of blocked and dropped calls in the network, respectively. The first experiment is related to the non-optimized FLC, which is applied to the initial unbalanced situation. As a result, the CBR is highly decreased to the disadvantage of an uncontrolled increase in the CDR, which would be unacceptable for operators, suggesting the need for a FLC optimization that finds a better trade-off between those indicators. The values of the CBR and the CDR are 1.4% and 8.0%, respectively, when such indicators remain stationary after load balancing. In this situation, the value of Uis 9.3%. The following experiment is related to the optimization approach named as UEE. To find the optimal fuzzy rules by using 100%exploration, three simulations starting from the unbalanced situation have been carried out with = 1 for different values of CDRth. As a result, three different sets of fuzzy rules have been extracted from the optimization process. Table 5.4 shows the optimal rule consequents determined by selecting the candidate action with the highest qvalue for each rule and they are compared with the rule base for the non-optimized FLC described in Section 5.2.2. When increasing CDRth, rules 3 and 9 are modified to leave the HOM at the same value, instead of returning it to the default value. As a result, lower and higher values of HOM can be achieved, increasing the risk of call dropping. Another consequence of increasing CDRth is that higher changes in HOMs are allowed, as rules 2 and 8 are modified to perform larger increments in HOMs. Table 5.4: Optimized fuzzy rules Non-optimized Action for Action for Action for Rule FLC CDRth = 0.06 CDRth = 0.1CDRth = 0.15 1 EL VL VL VL 2 VL L L VL 3 L VH Z Z 4 Z VL VL VL 5 Z Z Z Z 6 Z VH VH VH 7 EH VH VH VH 8 VH H H VH 9 H VL Z Z 112
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS Fig. 5.6 represents a realization of the q-value evolution of the consequents for optimized rules 1-4 during a simulation when CDRth = 0.06 (rules 6-9 are obviated due to symmetry). As it can be observed, those consequents obtaining higher q-values are the optimal actions that can be identified in Table 5.4. The initial HOM values are set to 3 dB, i.e. the default value, corresponding to the fuzzy set named ‘Medium’, so that rules 2 and 8 are triggered from the first iterations, while others are triggered from iterations 10-15. In addition, the situations that trigger rules 1 and 7 are less likely to happen than others, slowing down the convergence of the algorithm for those rules. Regarding the rule 1, the algorithm converges around iterations 80-90, from which the optimal policy for the FLC can be derived. It can be seen that there is clearly one candidate action better than the others for each rule, meaning that the algorithm leads to a specific set of rules ensuring the convergence of the algorithm. In addition, the convergence time (around 32 hours) is enough to solve persistent congestion problems. The operator firstly has to identify the time frames during a day in which there exists congestion. Generally, the congestion situation is also repeated every day (e.g. from Monday to Saturday) in the same locations, so that the optimization process can be performed over days. Then, the operator performs the optimization process in the target area during the congestion periods. 20 40 60 80 100 0 2 4 Epoch q−value Rule 1 EL VL N 20 40 60 80 100 0 10 20 Epoch q−value Rule 2 EL VL L 20 40 60 80 100 −4 −2 0 2 4 6 Epoch q−value Rule 3 VL N VH 20 40 60 80 100 −5 0 5 10 Epoch q−value Rule 4 VL N VH Figure 5.6: Evolution of candidate action q-values (UEE) 113
CHAPTER 5. MOBILITY LOAD BALANCING To compare evaluation of the three FLCs optimized by the UEE method, Fig. 5.7 and Fig. 5.8 show the temporary evolution of the global CBR and global CDR, respectively, when the optimized controller is used without the optimization module during the load balancing process (100% exploitation). Also performance of the non-optimized FLC from the first experiment is shown as a benchmark. All the simulations start from an initial situation of a global CBR close to 6% and a global CDR of 2%. After a few iterations, the network converges to a stationary state given by these indicators. The non-optimized FLC leads to the lowest CBR, but also to the highest CDR. When the FLC is optimized with a higher CDRth the global CBR decreases further after load balancing, while increasing the global CDR. This is because the FLC leads to higher and lower values of HOMs in some adjacencies, so that the cell service area suffers larger variations. Regarding the CBR in Fig. 5.7, the final values are 3.5, 3.1 and 2.4% when CDRth is 0.06, 0.10 and 0.15, respectively. The optimization process also has an impact on the speed of the adaptation process. When CDRth = 0.15, the FLC achieves stationarity faster, as higher changes in HOMs are performed by the FLC (action of rules 2 is VL instead of L). In particular, stationarity in CBR is achieved around iteration 25 (9.58 hours), while the fall time (time required for the response to rise from 10% to 90% of its final value) is equal to 4.98 hours. In the case of CDRth = 0.06 and CDRth = 0.10, the slope of the CBR transient response is very similar (action of rule 2 is L for both cases), but stationarity is achieved before when CDRth = 0.06 because its final value is higher. In particular, stationarity in CBR is achieved around iteration 35 (13.4 hours) and 38 (14.5 hours), and the fall time is 6.52 and 8.43 hours when CDRth is equal to 0.06 and 0.10, respectively. The same reasoning could be applied to the evolution of CDR, shown in Fig. 5.8. As it can be observed, there is no overshoot or oscillation in the temporary response of these performance indicators. Note that most of the congestion is relieved in 10-12 hours, so that slow temporary changes (e.g. in the order of days) could be managed by the FLC. 0 10 20 30 40 50 60 0.01 0.02 0.03 0.04 0.05 0.06 0.07 Epoch Call Blocking Ratio CDRth=0.06 CDRth=0.10 CDRth=0.15 Non−optimized Figure 5.7: Temporary evolution of the CBR 114
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS 0 10 20 30 40 50 60 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 Epoch Call Dropping Ratio CDRth=0.06 CDRth=0.10 CDRth=0.15 Non−optimized Figure 5.8: Temporary evolution of the CDR To analyze the impact on the HO procedure, the evolution of the HOM and the number of HOs per epoch in a specific adjacency (cells 31 and 32) for two cases of UEE method (CDRth = 0.10 and CDRth = 0.15) are shown in Fig. 5.9 and Fig. 5.10, respectively. This adjacency is located close to the congestion situation, so that the HOM will be adjusted as part of the load balancing process. As expected, adaptation of HOM is faster when CDRth is set to 0.15, because the consequent for rule 2 (rule 8) is VL (VH) instead of L (H). The impact on HOs is that the number of outgoing calls from cell 32 to cell 31 is higher than the number of incoming call to cell 32 from cell 31 after the load balancing because the algorithm attempts to offload cell 32. The optimization process greatly simplifies the selection of fuzzy rules when the network operator imposes a GoS constraint related to the call maintainability, typically in the CDR. By adjusting the CDRth, the network evolves to a new state with different trade-off between global CBR and CDR. Table 5.5 shows the evaluation of the utility function, U, for the three optimized FLCs and the non-optimized FLCs (mean and 95% confidence interval (CI) are shown). The best FLC configuration is the one with the lowest value of U(i.e. the lowest user dissatisfaction), corresponding to the FLC optimized with CDRth = 0.06. It is noted that optimizing the FLC with a CDRth lower than 0.06 leads to negative action q-values, as reinforcement signal rewards with negative values most of the time. In this case, finding the optimal fuzzy rules becomes a complicated task because all rules present similar q-values. Thus, the minimum value for Uis achieved when the FLC is optimized with CDRth = 0.06. As this configuration brings out a good trade-off between performance indicators, additional results are shown below for CDRth = 0.06. Firstly, the CBR per cell is depicted in Fig. 5.11. The row of bars at the back corresponds to the initial situation of load imbalance, while the row of bars in the front corresponds to the end of a simulation after load balancing. It is observed that there are cells with high CBR and others with negligible CBR at the beginning of the simulation. The traffic load is shared between cells after load balancing and CBR is more equalized than the initial situation. For instance, the CBR 115
CHAPTER 5. MOBILITY LOAD BALANCING 10 20 30 40 50 60 70 80 90 100 -4 -3 -2 -1 0 1 2 3 4 Epoch [dB] HOMargin(Cell32-Cell31) CDRth=0.10 CDRth=0.15 Figure 5.9: Evolution of the HOM in an adjacency for different values of CDRth 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 Epoch NumberofHOs CDRth=0.10 Fromcell32tocell31 Fromcell31tocell32 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 Epoch NumberofHOs CDRth=0.15 Fromcell32tocell31 Fromcell31tocell32 Figure 5.10: Evolution of the number of HOs in an adjacency for different values of CDRth in cells 33 and 40 is decreased from approximately 20% to 7%. This means a reduction of 65% in those highly congested cells. The CDR per cell is represented in Fig. 5.12, where both the initial and final situation are also shown. As the service area is greatly decreased in congested cells, the users connected to these cells are very close to the base station, experiencing good quality. Thus, the CDR measured in those congested cells (e.g. cells 33, 39 and 40) is decreased after load balancing. In the case of less loaded cells, the CDR is increased as they receive users from congested cells in worse quality conditions. Such an increment means a greater proportional increase taking into account that these cells have a low number of users and receive many users from congested cells. The next experiment analyzes the BEE optimization approach, where the FLC is used 116
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS Table 5.5: Utility function for different values of CDRth CDRth CBR CBR CDR CDR U (mean) (95% CI) (mean) (95% CI) 0.06 3.48% [3.45, 3.51] 4.17% [4.14, 4.21] 7.5% 0.10 3.06% [3.02, 3.10] 4.89% [4.84, 4.93] 7.8% 0.15 2.43% [2.40, 2.46] 6.01% [5.97, 6.04] 8.2% non-opt 1.44% [1.40, 1.49] 8.04% [7.98, 8.09] 9.3% 10 20 30 40 50 0 0.05 0.1 0.15 0.2 0.25 Cell Call Blocking Ratio Last iteration First iteration Figure 5.11: CBR per cell using the FLC optimized by UEE 10 20 30 40 50 0 0.05 0.1 0.15 0.2 Cell Call Dropping Ratio Last iteration First iteration Figure 5.12: CDR per cell using the FLC optimized by UEE 117
CHAPTER 5. MOBILITY LOAD BALANCING starting from an unbalanced traffic distribution and the fuzzy rules are progressively modified to obtain more reward according to the reinforcement signal. The parameter is set to 0.2 in order to keep a trade-off of 20% exploration and 80% exploitation. In this case, the best action is executed most of the time and load balancing can be performed while call quality is preserved according to the specific value of CDRth. Fig. 5.13 represents an example of the qvalue evolution for each consequent of the rules 1-4 when CDRth = 0.06 and = 0.2. As it can be observed, the Q-Learning algorithm only needs 35-40 epochs to find the best action and use it most of the time, since low degree of exploration has been adopted. Considering that the best fuzzy rules are those receiving the highest q-value, the optimal configuration obtained from this last simulation is similar to that derived from the UEE approach with CDRth = 0.06 and = 1. This conclusion is drawn from comparing the rule base of the two approaches, where all rules of both FLC configurations match except rule 7, whose consequent is VH in the case of UEE approach and EH in the BEE approach. Such a matching suggests that both UEE and BEE optimization achieve similar system performance. The following paragraphs compare performance of the proposed methods. In summary, Fig. 5.14 shows the performance of the proposed methods in terms of the 20 40 60 80 0 2 4 Epoch q−value Rule 1 EL VL N 20 40 60 80 0 20 40 Epoch q−value Rule 2 EL VL L 20 40 60 80 0 10 20 Epoch q−value Rule 3 VL N VH 20 40 60 80 0 10 20 30 Epoch q−value Rule 4 VL N VH Figure 5.13: Evolution of candidate action q-values (BEE) 118
5.2. MLB ALGORITHMS FOR VOICE SERVICE IN MACROCELLS 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 0.1 0.01 0.015 0.02 0.025 0.03 0.035 0.04 0.045 0.05 0.055 Call Dropping Ratio Call Blocking Ratio Non−optimized UEE CDRth=0.06 UEE CDRth=0.10 UEE CDRth=0.15 BEE CDRth=0.06 Figure 5.14: Performance of the proposed methods global CBR and global CDR. All the simulations start from the same network state and evolve to a different state corresponding to the end of the simulation and represented by a specific symbol in Fig. 5.14. Each curve represents the intermediate states in which the HOMs are dynamically modified according to the FLC behavior. It is observed that all methods reduce CBR by increasing CDR. However, the non-optimized approach leads to an excessive CDR, highlighting the need of optimizing the FLC. Using the FLCs optimized by the UEE method, the increase in CDR can be controlled depending on the network operator constraint set by CDRth. Increasing CDRth leads to a higher value of global CDR, but still far from the nonoptimized case. In the case of the BEE approach, non-optimal actions are explored during the load balancing process, so that the trajectory followed by the FLC is slightly worse in terms of CDR. In addition, exploration leads to small oscillations in performance once the load balancing is performed, as shown in Fig. 5.14. The utility function, U, takes the value 7.7% when BEE is performed with CDRth = 0.06, showing that it is not a better solution than applying UEE for the same value of CDRth (U= 7.5% was obtained in that case), but close to it. The major benefit of the BEE approach is that it can be applied in any scenario without a previous (UEE) optimization phase, which usually requires the use of simulation tools. For instance, suppose that the traffic load of the scenario now is lower than in the previous simulations. More precisely, the average traffic load is set to 60% instead of 75%. Likewise, the network operator desires to keep the same quality constraint given by CDRth = 0.06. The UEE optimized FLC with CDRth = 0.06 would not fully exploit the FLC because the scenario now is different (i.e. the FLC would not be optimized for that traffic condition, but for the previous one). In the case of the BEE optimization approach, the FLC would select new optimal actions leading to a lower value of CBR and speeding up the load balancing process while preserving the same constraint in CDR. For instance, optimal consequent EL is selected for rule 1 instead of VL, VL is selected for rule 2 instead of L, and Z is selected for rule 3 instead of VH. Fig. 5.15 shows that the decrease in the global CBR produced by BEE is faster and greater than that 119
APPENDIX A. SUMMARY (SPANISH) carga en este tipo de escenarios, en donde además la planificación no suele seguir un riguroso análisis, por lo que se puede optimizar el uso de los recursos aportando mayores ganancias al sistema. Al igual que MLB, MRO ha sido considerado de forma especial por el 3GPP debido a su relevancia desde el punto de vista de los operadores. Los principios básicos de MRO están definidos en [43]. Las mejoras planteadas para esta función y su aplicación en el contexto de redes heterogéneas y escenarios inter-tecnología se describen en [53]. En la comunidad investigadora, la optimización de traspasos ha ido ganando cierta atención [76][77], centrándose en el desarrollo de algoritmos SON [12][13][14][15][16] y su estudio en escenarios inter-tecnología [78][79]. Sin embargo, estas referencias no incluyen un análisis profundo sobre los principales parámetros de los traspasos (margen de traspaso y Time-to-Trigger) en diferentes situaciones de carga, velocidad de usuario, etc., por lo que en este sentido se requiere mayor investigación. Dado que tanto MLB como MRO son funciones muy importantes en el contexto de SON, la coordinación de ambas funciones también juega un papel esencial y por ello se ha ganado el interés de la comunidad investigadora [85][86][87]. No obstante, debido a que la coordinación depende en cierta manera de las implementaciones de ambas funciones, la investigación llevada a cabo hasta el momento se considera escasa. En este sentido, dado que el objetivo de MLB en esta tesis es hacer frente a situaciones de congestión persistentes y adaptarse a variaciones lentas del tráfico, la coordinación de MLB y MRO actuará a un ritmo más lento que los algoritmos de coordinación propuestos en la literatura, tratándose por tanto de un problema diferente. Posteriormente, se realiza un análisis del estado del arte en el ámbito de redes heterogéneas y direccionamiento de tráfico [17][18][19][20][21]. Sin embargo, la mayoría de las referencias son estudios previos al concepto más amplio de red heterogénea, en donde no solo existen diferentes tecnologías de acceso radio, sino también diferentes tamaños de celdas, frecuencias de uso, elementos de red (p.ej. relays), etc. Además, las tecnologías que se utilizan para evaluar las técnicas de direccionamiento de tráfico son típicamente GSM, UMTS y WLAN, de manera que se requiere un mayor estudio en escenarios con nuevas tecnologías de acceso radio, tales como LTE y LTE-A. A.2.2 Técnicas adaptativas para la optimización de parámetros de red La última parte del Capítulo 2 realiza un análisis de la literatura actual sobre los métodos de optimización que se utilizan para mejorar las prestaciones de las técnicas de auto-ajuste de parámetros. En primer lugar, se estudia el estado del arte en controladores de lógica difusa (Fuzzy Logic Controller, FLC) como principal método heurístico para el problema de autoajuste [91][94][95][96][75]. La principal razón de usar este tipo de técnicas es que son mucho más atractivas desde el punto de vista del operador debido a su bajo coste, ya que no requieren costosas herramientas de simulación para su funcionamiento. En segundo lugar, se ha realizado un análisis de la bibliografía sobre las técnicas de optimización que permiten optimizar el funcionamiento de un FLC, en particular, las reglas difusas que definen su dinámica. En el contexto de las redes móviles, es probable que los expertos tengan dificultades para ajustar estas reglas, o bien sea 222
A.3. AUTO-OPTIMIZACIÓN DE REDES DE COMUNICACIONES MÓVILES necesario adaptar dinámicamente la conducta del FLC a las variaciones del entorno. Por esta razón, el uso de métodos de optimización para FLC es una cuestión muy importante que ha sido abordada en la literatura [97][98][99][9][19][100]. Entre estas técnicas, en esta tesis se ha valorado especialmente el aprendizaje por refuerzo [101], debido a que es un método particularmente apropiado para aprender mediante la interacción con el entorno, sobre todo en aquellas situaciones en las que es complicado obtener ejemplos representativos de todas las situaciones en las que el FLC debe actuar, como ocurre en los escenarios de redes móviles. Dentro del aprendizaje por refuerzo, se destaca el algoritmo de Q-Learning y sus aplicaciones en el ámbito de optimización de las redes móviles [102][103][104][105][106][107][108]. Además, la combinación de FLC y el algoritmo de Q-Learning es un potente mecanismo de optimización en donde las facilidades de modelado del FLC que ofrecen un alto nivel de abstracción se añaden a la capacidad del Q-Learning para adaptar/entrenar el controlador. En el Capítulo 3, se realiza en primer lugar una introducción a la lógica difusa [111], centrándose posteriormente en el auto-ajuste de parámetros mediante FLC [113][116]. En segundo lugar, se realiza un estudio sobre las principales técnicas usadas para la optimización de FLC. Para cada técnica, se proporciona una breve introducción así como una descripción de cómo aplicar la técnica para optimizar FLC, incluyendo sus ventajas e inconvenientes. En particular, las técnicas estudiadas son las redes neuronales, los algoritmos genéticos, el enjambre de partículas y el aprendizaje por refuerzo [117][113]. De entre todas estas técnicas, se justifica la elección del uso de aprendizaje por refuerzo para optimizar FLC en esta tesis. A.3 Auto-optimización de redes de comunicaciones móviles En la segunda parte de la tesis se presentan las principales aportaciones para cada uno de los problemas abordados. El objetivo común es diseñar técnicas para el auto-ajuste de parámetros en redes de comunicaciones móviles. Para evaluar dichas técnicas, se ha desarrollado un simulador LTE dinámico de nivel de sistema. Aunque también se han hecho uso de otros simuladores, una importante parte del diseño de este simulador se ha realizado en el marco de esta tesis, de manera que el Capítulo 4 se dedica a la descripción del mismo. El resto de capítulos están dedicados a cada uno de los problemas que se abordan en la tesis. En particular, el Capítulo 5 trata el problema del balance de carga o MLB, tanto en escenarios de macroceldas como en escenarios de femtoceldas corporativas. El Capítulo 6 está dedicado al problema de la optimización de traspasos o MRO así como a la coordinación entre MLB y MRO. En el Capítulo 7 se estudian mecanismos de direccionamiento de tráfico en redes heterogéneas. En todos estos capítulos se sigue una estructura similar en la que, tras describir el problema, se presentan las técnicas propuestas y se exponen los resultados derivados de su evaluación mediante herramientas de simulación especialmente apropiadas para los algoritmos diseñados. 223
APPENDIX A. SUMMARY (SPANISH) A.3.1 Simulador LTE dinámico de nivel de sistema El Capítulo 4 presenta la herramienta principal de simulación que se ha utilizado en esta tesis para evaluar y validar las técnicas desarrolladas. En el proceso de diseño se ha hecho especial énfasis en desarrollar un modelo abstracto de los niveles más bajos de la arquitectura, es decir, los niveles físico y de enlace. El motivo fundamental es la necesidad de crear un simulador computacionalmente eficiente, de forma que permita simular largos períodos de tiempo en base a los problemas abordados. En particular, en esta tesis, se ha diseñado el modelo abstracto del nivel de enlace, mientras que el diseño de otros niveles del simulador, así como la parte de implementación, corresponden a otros trabajos. El Capítulo 4 comienza describiendo el nivel físico, en donde se hace uso de una serie de realizaciones de canal OFDM generadas previamente con el fin de caracterizar fielmente el canal radio, incluyendo el desvanecimiento multicamino. Tras esto, se describe el nivel de enlace, caracterizado por los cálculos que permiten obtener la relación señal a interferencia (Signal-toInterference Ratio, SIR), a partir de la cual se calcula la tasa de error de bloque (Block Error Rate, BLER) experimentada por el usuario. A continuación, se describen las dos funciones principales de este nivel, es decir, la adaptación de enlace y la planificación. Mientras que la primera permite seleccionar el esquema de modulación y codificación más adecuado para cada transmisión, la segunda asigna los recursos radio disponibles a los usuarios basándose en las condiciones de canal a lo largo de las diferentes sub-bandas de frecuencias. Posteriormente, se definen las funciones más importantes a nivel de red, tales como el control de admisión, el control de congestión y los traspasos. Finalmente, se presentan algunos resultados que permiten validar el simulador. A.3.2 Balance de carga mediante movilidad En el Capítulo 5 se describen las técnicas de balance de carga propuestas en esta tesis. El primer grupo de técnicas desarrolladas está pensado para ser aplicado en escenarios de macroceldas, mientras que el segundo grupo está diseñado para escenarios de femtoceldas corporativas. El capítulo comienza describiendo el algoritmo de movilidad y los parámetros típicamente utilizados en MLB para mitigar las situaciones de cogestión. Tras esta breve introducción, la primera parte del capítulo se dedica al problema de balance de carga en macroceldas, en particular, cuando el problema se debe a congestiones persistentes como por ejemplo las debidas a un crecimiento imprevisto de población en un área determinada. Una vez que se describe el problema, se justifica la necesidad del algoritmo propuesto y se describe tanto el modelo de red empleado como las medidas de sistema utilizadas. A continuación, se presenta el esquema de auto-ajuste (FLC) diseñado, mostrando sus entradas y salidas, así como las reglas que definen su comportamiento. Seguidamente, se describe el proceso de optimización mediante el uso del algoritmo de Q-Learning particularizado al problema en cuestión. En este sentido, se establecen una serie de supuestos prácticos que permiten aumentar la eficacia del controlador diseñado. Por último, se muestran los resultados de la evaluación y validación de las técnicas propuestas mediante el uso del simulador LTE diseñado en el Capítulo 4. 224
A.3. AUTO-OPTIMIZACIÓN DE REDES DE COMUNICACIONES MÓVILES La segunda parte del Capítulo 5 trata el problema de balance de carga en escenarios de femtoceldas corporativas. Tras describir brevemente el problema y justificar la solución, se presentan dos diseños de FLC extraídos de la literatura que serán posteriormente combinados con los nuevos diseños propuestos en esta tesis para mejorar sus prestaciones. Los FLC no solo modifican los márgenes de traspaso, sino también la potencia de transmisión de las femtoceldas. Seguidamente, se explican las métodos propuestos, los cuales están basados en FLC y Q-Learning, así como la combinación con los FLC ya existentes. Finalmente, se presentan los resultados de evaluar las diferentes estrategias en un simulador específico de femtoceldas. A.3.3 Optimización de la movilidad El Capítulo 6 dedica una primera parte a la optimización de traspasos en redes de comunicaciones móviles, mientras que la segunda parte se dedica a la coordinación de esta función con la estudiada en el capítulo anterior, el balance de carga. El Capítulo 6 comienza con una descripción de los principales parámetros presentes en el algoritmo de traspaso y de las medidas que se utilizarán en el posterior diseño del algoritmo SON. Tras esto, se describe el diseño del FLC propuesto para la optimización de traspasos. En particular, se presentan tres configuraciones diferentes del FLC, en función de las reglas y los valores de salida que se definen. La principal diferencia entre estas configuraciones es la magnitud del cambio en el parámetro a modificar. El diseño del controlador se basa en un análisis de sensibilidad de los dos parámetros más importantes del algoritmo de traspaso, es decir, el margen de traspaso y el Time-to-Trigger. Dicho análisis se lleva a cabo para diferentes situaciones de carga y velocidad de usuario en una red LTE. Posteriormente, se evalúan las prestaciones del FLC y se analiza el impacto de modificar dicho parámetro según diferentes niveles de optimización, por ejemplo, a nivel de red, de celda o de adyacencia, así como el impacto del ruido en las medidas que deciden cuando se dispara un traspaso. La segunda parte del Capítulo 6 trata el problema de la coordinación entre las funciones propuestas de MLB y MRO. La solución adoptada se basa en inhibir una de las funciones cuando su actuación (por ejemplo, valores extremos en el parámetro a modificar) puede empeorar las prestaciones de la red debido a la actuación de la otra función. Para ello, se definen una serie de umbrales que permiten detectar aquellas situaciones en las que este hecho ocurre. Finalmente, se evalúan las prestaciones del algoritmo de coordinación en una de estas situaciones que requiere resolver el conflicto entre ambas funciones. A.3.4 Direccionamiento de tráfico en redes heterogéneas En el Capítulo 7 se estudian diferentes técnicas para el direccionamiento de tráfico (TS) en redes heterogéneas. El estudio se divide en dos partes: la primera se centra en el desarrollo de técnicas mediante ajuste estático de los parámetro de red, es decir, a través de una serie de análisis de sensibilidad que permitan encontrar los ajustes óptimos; la segunda parte se dedica al desarrollo de técnicas mediante ajuste dinámico, las cuales modifican los parámetros de forma adaptativa para hacer frente a las variaciones de contexto en la red. 225
APPENDIX A. SUMMARY (SPANISH) En primer lugar, se describe el problema haciendo hincapié en las principales cuestiones que surgen en un escenario de redes heterogéneas cuando el operador desea redirigir el tráfico en la red de acuerdo a una determinada política. Seguidamente, se describe el escenario bajo el cual se evaluarán las técnicas propuestas. Dicho escenario debe ser representativo de las principales características que definen una red heterogénea. Tras esto, se presentan según la terminología 3GPP dos algoritmos que son muy importantes en este tipo de escenarios, la reselección de celda mediante prioridades absolutas y el algoritmo de traspaso inter-tecnología. Así mismo, se especifican las medidas de red que se utilizarán a lo largo del capítulo. Una vez que se han introducido los conceptos previos necesarios, el capítulo se centra en el enfoque estático, comenzando por un breve estudio de las técnicas existentes para realizar TS. Tras justificar el uso de las técnicas en la presente tesis, se investiga cómo ajustar las prioridades absolutas del algoritmo de reselección de celda en función de las características del escenario. A continuación, se describe el algoritmo propuesto que optimiza los parámetros de traspaso inter-tecnología mediante una serie de análisis de sensibilidad. Una vez explicado el enfoque estático, el estudio se centra en el enfoque dinámico, el cual se basa en el diseño de un FLC optimizado por el algoritmo de Q-Learning que permite soportar diferentes políticas del operador y adaptarse a cambios en la red. Dicho algoritmo se basa en el ajuste de los parámetros de traspaso inter-tecnología. Finalmente, se muestran los resultados de evaluar los algoritmos tanto del enfoque estático como del dinámico. A.4 Evaluación La evaluación y comparación de las distintas alternativas propuestas en esta tesis se describe en la última parte de los Capítulos 5, 6 y 7. Para cada una de ellas, se han utilizado diversas metodologías, medidas, figuras de mérito, etc. Las pruebas se han realizado mediante el uso de simuladores dinámicos de red, adaptados a las particularidades del problema a resolver. Respecto al simulador LTE principalmente utilizado para la evaluación de los algoritmos desarrollados, los resultados de las pruebas realizadas para su validación se describen al final del Capítulo 4. A.4.1 Resultados Balance de carga mediante movilidad. Para analizar los algoritmos propuestos de balance de carga en escenarios de macroceldas, se define una figura de mérito que tiene en cuenta tanto el bloqueo experimentado en la red como la caída de llamadas. De esta manera, dicha figura de mérito ofrece una medida sobre la insatisfacción de usuarios en la red. Las simulaciones presentan un tráfico espacial no uniforme con el objetivo de focalizar la carga en el centro del escenario y generar cierta congestión. Los algoritmos propuestos se comparan con un caso de referencia, caracterizado por el uso de un FLC sin optimizar para balance de carga. En los resultados, se presenta la evolución en el tiempo de un parámetro característico del algoritmo de optimización y los principales indicadores relativos al bloqueo y la caída de llamadas, así como su representación a nivel de celda. Los resultados muestran que es posible reducir significativamente el bloqueo en la red mientras se mantiene la 226
A.4. EVALUACIÓN caída de llamadas en las celdas vecinas bajo un cierto nivel establecido por el operador. En el caso del balance de carga en escenarios de femtoceldas corporativas, se define una figura de mérito similar al caso anterior, en la que igualmente se mide el empeoramiento de la calidad de la conexión como efecto indeseado al realizar balance de carga en la red. Además, para comparar los algoritmos propuestos, se hace uso de un modelo de horizonte infinito [164] que permite evaluar respuestas temporales en donde las muestras más lejanas en el tiempo se consideran menos importantes que las más cercanas. El estudio también incluye un análisis de la carga de señalización introducida por la modificación de los parámetros involucrados, así como la evolución temporal de estos parámetros y otros indicadores de interés. Los resultados destacan los beneficios de combinar FLCs, que proporcionan una respuesta rápida, con diseños propuestos en la tesis (basados en FLC y Q-Learning), que proporcionan buenas prestaciones a largo plazo. Optimización de la movilidad. La descripción de los resultados sobre optimización de la movilidad incluye en primer lugar un análisis de sensibilidad de los dos parámetros más característicos del procedimiento de traspaso. El análisis se lleva a cabo para diferentes niveles de carga de la red y diferentes velocidades de usuario. Se asume además que la distribución espacial de tráfico es uniforme. Los resultados muestran que la red es más sensible a variaciones en el margen de traspaso que en el Time-toTrigger. Posteriormente, se presentan los resultados de evaluar el FLC propuesto, cuyo diseño está basado en los resultados del análisis de sensibilidad previo. Para ello, se ha considerado que la época de optimización o intervalo de tiempo que hay entre dos modificaciones del parámetro involucrado debe ser suficientemente largo para obtener medidas fiables de los indicadores de prestaciones para el tamaño de red, el modelo de tráfico y la carga de red seleccionados para las simulaciones. También se ha definido una figura de mérito que incluye un indicador relativo a las llamadas caídas y otro relativo a la carga de señalización debida a los traspasos. Para facilitar la visualización de los resultados, se presentan en una misma gráfica los dos principales indicadores de prestaciones para este problema. Los resultados muestran que la configuración propuesta para el FLC proporciona valores más bajos de caída de llamadas que otras configuraciones evaluadas, a la vez que se mantiene un coste de señalización similar. Este hecho es especialmente apreciable cuando la velocidad de los usuarios del escenario es elevada. Para la evaluación del esquema propuesto de coordinación entre MLB y MRO se analiza la evolución temporal de los indicadores principales y se realizan medidas no solo a nivel de red sino también a nivel de celda. En este caso, la distribución espacial de tráfico es no uniforme con el objetivo de crear congestión en la red. Los resultados indican que la coordinación es necesaria, haciéndose efectiva en determinadas situaciones en las que algún indicador de la red (p.ej. de bloqueo o de señalización de traspasos en la red) provoca que el parámetro a modificar (el margen de traspaso) alcance valores grandes. Direccionamiento de tráfico en redes heterogéneas. Las prestaciones de los algoritmos propuestos para el direccionamiento de tráfico en redes heterogéneas se han evaluado en un simulador HSPA/LTE dinámico de nivel de sistema que incluye celdas de diferentes tamaños, en particular, macroceldas y picoceldas. En los resultados, se 227
APPENDIX A. SUMMARY (SPANISH) definen los indicadores utilizados para la evaluación de dichos algoritmos, como por ejemplo la medida del tráfico cursado en la red, definido según el servicio proporcionado a los usuarios. Otros indicadores están relacionados por ejemplo con la distribución de usuarios a través de las redes presentes en el escenario, así como la utilización de recursos en dichas redes. El estudio también ha tenido en cuenta diferentes niveles de penetración de los terminales LTE. Además, para analizar la capacidad de adaptación del controlador propuesto, se ha incluido un cambio en la distribución espacial de usuarios que afecta al número de usuarios presentes en los hotspots o lugares donde existe una alta concentración de usuarios. Los resultados muestran que las técnicas propuestas son efectivas para modificar la distribución de usuarios, especialmente cuando las áreas de cobertura de las redes están completamente solapadas (p.ej. cuando las estaciones bases comparten los emplazamientos). Además, el controlador propuesto es capaz de modificar la distribución de usuarios de acuerdo a las políticas del operador, así como de adaptarse a cambios contextuales en la red. A.4.2 Conclusiones En el Capítulo 8 se resume el trabajo realizado durante la tesis. En particular, se describen los resultados y aportaciones para cada uno de los problemas abordados y se proponen líneas de continuación. Las principales aportaciones de la tesis son las siguientes: •Propuesta de algoritmos para el balance de carga mediante movilidad en macroceldas. Los diseños propuestos están basados en la optimización mediante Q-Learning de un FLC que trata de resolver problemas de congestión persistente en redes inalámbricas de próxima generación. El objetivo principal es reducir el bloqueo para servicios de voz controlando al mismo tiempo la degradación causada en la calidad de la conexión de aquellos usuarios afectados por el proceso. •Propuesta de algoritmos para el balance de carga mediante movilidad en femtoceldas corporativas. El objetivo principal consiste en mejorar, mediante el uso de técnicas de optimización (en este caso, Q-Learning), las prestaciones de diseños existentes basados en FLC aprovechando la velocidad que poseen estos controladores. La degradación de la calidad de la conexión también es un factor clave en el desarrollo de los esquemas propuestos. •Propuesta de algoritmos para la optimización de la movilidad. Las principales contribuciones en esta área son: –Un análisis de sensibilidad de los principales parámetros de traspaso realizado para diferentes niveles de carga de la red y diferentes velocidades de usuario. –El diseño de un FLC que modifica adaptativamente los márgenes de traspaso para la optimización de la movilidad. Para ello, se han considerado diferentes configuraciones del FLC, así como diferentes niveles de optimización (a nivel de red, de celda y de adyacencia). También se ha analizado el impacto del ruido en las medidas de nivel de señal que se utilizan para lanzar un traspaso. 228
A.5. LISTA DE PUBLICACIONES –El diseño de un esquema de coordinación entre el FLC propuesto para la optimización de la movilidad y el algoritmo (FLC combinado con Q-Learning) propuesto para el balance de carga mediante movilidad. Dicha coordinación trata de coordinar dos funciones cuyos cambios en la red se producen lentamente (en torno a minutos, horas). •Propuesta de algoritmos para el direccionamiento de tráfico en redes heterogéneas. Las principales contribuciones en esta área son: –Un análisis de la asignación de prioridades absolutas a las redes de una red heterogénea en el algoritmo de reselección de celda. Se estudian algunas cuestiones que surgen del ajuste estático de los parámetros de reselección de celda desde el punto de vista del direccionamiento de tráfico. –Un procedimiento para seleccionar los umbrales óptimos de los eventos del 3GPP que controlan la ejecución de los traspasos inter-tecnología. Dicha selección resulta especialmente tediosa en escenarios con diferentes tamaños de celdas y cuando los eventos que controlan la ejecución de traspasos se basan en los niveles de señal recibida por el terminal. –El diseño de un esquema basado en FLC y optimizado mediante Q-Learning que proporciona al operador la flexibilidad de elegir diferentes políticas de direccionamiento de tráfico y permite cierta adaptación a las variaciones del entorno (p.ej. en la distribución de usuarios) de forma automática. Para ello, el algoritmo propuesto modifica los umbrales de los eventos que controlan la ejecución de los traspasos inter-tecnología. A.5 Lista de publicaciones La siguiente lista presenta las publicaciones relacionadas con esta tesis: Revistas Derivadas de la tesis [I] P. Muñoz, R. Barco, D. Laselva y P. Mogensen. Mobility-based Strategies for Traffic Steering in Heterogeneous Networks. IEEE Communications Magazine, vol. 51, no. 5, pp 54-62, May 2013. [II] P. Muñoz, R. Barco y I. de la Bandera, Optimization of Load Balancing using Fuzzy Q-Learning for Next Generation Wireless Networks, Expert Systems With Applications (Elsevier), vol. 40, no. 4, pp 984-994, Marzo 2013. [III] P. Muñoz, D. Laselva, R. Barco y P. Mogensen. Adjustment of mobility parameters for traffic steering in multi-RAT multi-layer wireless networks, EURASIP Journal on Wireless 229
APPENDIX A. SUMMARY (SPANISH) Communication and Networking, 2013:133, May 2013. [IV] P. Muñoz, R. Barco e I. de la Bandera. On the Potential of Handover Parameter Optimization for Self-Organizing Networks. IEEE Transactions on Vehicular Technology, aceptado en 2013. [V] P. Muñoz, R. Barco, J. M. Ruiz, I. de la Bandera y A. Aguilar. Fuzzy Rule-based Reinforcement Learning for Load Balancing Techniques in Enterprise LTE Femtocells. IEEE Transactions on Vehicular Technology, aceptado en 2012. [VI] J. M. Ruiz, S. Luna-Ramírez, M. Toril, F. Ruiz, I. de la Bandera, P. Muñoz, R. Barco, P. Lázaro y V. Buenestado, Design of a Computationally Efficient Dynamic System-Level Simulator for Enterprise LTE Femtocell Scenarios, Journal of Electrical and Computer Engineering, vol. 2012, Diciembre 2012. [VII] P. Muñoz, I. de la Bandera, F. Ruiz, S. Luna-Ramírez, R. Barco, M. Toril, P. Lázaro y J. Rodríguez, Computationally-Efficient Design of a Dynamic System-Level LTE Simulator, International Journal of Electronics and Telecommunications, vol. 57, no. 3, pp 347-358, Septiembre 2011. Relacionadas con la tesis [VIII] R. Barco, P. Lázaro y P. Muñoz. A Unified Framework for Self-Healing in Wireless Networks. IEEE Communications Magazine, Vol.50 (12), pp.134-142. Diciembre 2012. Conferencias y Workshops Derivadas de la tesis [IX] P. Muñoz, I. de la Bandera, R. Barco, M. Toril, S. Luna-Ramírez y J.M. Ruiz, “Sensitivity Analysis and Self-Optimization of LTE Intra-Frequency Handover”, 6th Scientific Meeting, COST action IC1004, Málaga (España), Febrero 2013. [X] P. Muñoz, R. Barco, I. de la Bandera, M. Toril y S. Luna-Ramírez, “Optimization of a Fuzzy Logic Controller for Handover-based Load Balancing”, Int. Workshop on SelfOrganising Networks (IWSON), IEEE Vehicular Technology Conference (VTC) Spring, Budapest, Hungría, Mayo 2011. [XI] P. Muñoz, I. de la Bandera, R. Barco, F. Ruiz, M. Toril y S. Luna-Ramírez, “Estimation of Link-Layer Quality Parameters in a System-Level LTE Simulator”, 5th International Conference on Broadband and Biomedical Communications (IB2COM) 2010. Málaga (España). 230
A.5. LISTA DE PUBLICACIONES Diciembre, 2010. [XII] P. Muñoz, I. de la Bandera, R. Barco, M. Toril y S. Luna-Ramírez, “Optimización del balance de carga en redes LTE mediante el algoritmo de Q-Learning difuso”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (España), Septiembre, 2011. [XIII] I. de la Bandera, P. Muñoz, R. Barco, M. Toril y S. Luna-Ramírez, “Auto-ajuste del margen de handover en redes LTE”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (España), Septiembre, 2011. [XIV] P. Muñoz, I. de la Bandera, R. Barco, F. Ruiz, M. Toril y S. Luna-Ramírez, “Diseño del Nivel de Enlace para un Simulador LTE”, XXV Simposio de la Unión Científica Internacional de Radio (URSI 2010), Bilbao (España), Septiembre, 2010. Relacionadas con la tesis [XV] J.M. Ruiz, S. Luna-Ramírez, M. Toril, F. Ruiz, I. de la Bandera y P. Muñoz, “Analysis of Load Sharing Techniques in Enterprise LTE Femtocells”, Int. Workshop on Femtocells, Wireless Advanced 2011, Londres, Reino Unido, Junio 2011. [XVI] J. Rodríguez, I. de la Bandera, P. Muñoz y R. Barco, “Load Balancing in a Realistic Urban Scenario for LTE Networks”, Int. Workshop on Self-Organising Networks (IWSON), IEEE Vehicular Technology Conference (VTC) Spring, Budapest, Hungría, Mayo 2011. [XVII] J. Rodríguez, I. de la Bandera, P. Muñoz y R. Barco, “Balance de carga en red LTE en un entorno urbano realista”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (España), Septiembre, 2011. [XVIII] G. Jiménez, R. Barco, P. Muñoz y M. Toril, “Optimización del Umbral de Traspaso para Femtoceldas LTE”, XXVI Simposio de la Unión Científica Internacional de Radio (URSI 2011), Leganés (España), Septiembre, 2011. [VI, VII, XI, XIV] se dedican a las herramientas de simulación que facilitan la evaluación de las prestaciones de los algoritmos propuestos. [II, X, XII, XVI, XVII] tratan el problema de balance de carga en escenarios de macroceldas, mientras que en escenarios de femtoceldas este problema se estudia en [V, XV, XVIII]. El problema de la optimización de los traspasos se aborda en [IV, IX, XIII]. [VIII] se dedica a la auto-curación en el campo de redes auto-organizativas. En [I, III], se describen las técnicas propuestas para el direccionamiento de tráfico en el contexto de redes heterogéneas. El autor ha sido el autor principal de todas las contribuciones “derivadas de la tesis” excepto [VI, XIII], siendo un contribuidor clave en las mismas. En [VI], el autor escribió la sección dedicada al modelo de movilidad en interiores y estuvo involucrado en el diseño de la capa de enlace del simulador. En [XIII], el autor participó activamente en las etapas de diseño 231
BIBLIOGRAPHY [73] J. Ruiz-Avilés, S. Luna-Ramírez, M. Toril, and F. Ruiz. Fuzzy logic controllers for traffic sharing in enterprise LTE femtocells. In Proc. of IEEE 75th Vehicular Technology Conference (VTC), 2011. [74] S. M. Tseng and W. Z. Tseng. Drop rate optimization by tuning time to trigger for WCDMA systems. In Proc. of Second International Conference on Ubiquitous and Future Networks (ICUFN), 2010. [75] C. Werner, J. Voigt, S. Khattak, and G. Fettweis. Handover Parameter Optimization in WCDMA using Fuzzy Controlling. In Proc. of IEEE 18th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), 2007. [76] P. Legg, G. Hui, and J. Johansson. A Simulation Study of LTE Intra-Frequency Handover Performance. In Proc. of IEEE 72nd Vehicular Technology Conference (VTC), 2010. [77] Yejee Lee, Bongjhin Shin, Jaechan Lim, and Daehyoung Hong. Effects of time-to-trigger parameter on handover performance in SON-based LTE systems. In Proc. of 16th AsiaPacific Conference on Communications (APCC), 2010. [78] Q. Song, Z. Wen, X. Wang, L. Guo, and R. Yu. Time-Adaptive Vertical Handoff Triggering Methods for Heterogeneous Systems. Computer Science, 5737/2009:302–312, 2009. [79] A. Awada, B. Wegmann, D. Rose, I. Viering, and A. Klein. Towards Self-Organizing Mobility Robustness Optimization in Inter-RAT Scenario. In Proc. of IEEE 73rd Vehicular Technology Conference (VTC), 2011. [80] R. Combes, Z. Altman, and E. Altman. Coordination of autonomic functionalities in communications networks. In CoRR abs/1209.1236, 2012. [81] INFSO-ICT-216284 SOCRATES. Framework for the development of self-organisation methods. Technical Report Deliverable D2.4, Version 1.0.3, September, 2008. [82] L. C. Schmelz, M. Amirijoo, A. Eisenblaetter, R. Litjens, M. Neuland, and J. Turk. A coordination framework for self-organisation in LTE networks. In Proc. of IEEE International Symposium on Integrated Network Management (IM), 2011 IFIP, pages 193–200, 2011. [83] INFSO-ICT-216284 SOCRATES. Final Report on Self-Organisation and its Implications in Wireless Access Networks. Technical Report Deliverable D5.9, Version 1.0, December, 2010. [84] B. Sas, K. Spaey, I. Balan, K. Zetterberg, and R. Litjens. Self-Optimisation of Admission Control and Handover Parameters in LTE. In Proc. of IEEE 73rd Vehicular Technology Conference (VTC), Spring, 2011. [85] A. Lobinger, S. Stefanski, T. Jansen, and I. Balan. Coordinating Handover Parameter Optimization and Load Balancing in LTE Self-Optimizing Networks. In Proc. of IEEE 73rd Vehicular Technology Conference (VTC), Spring, 2011. [86] W. Li, X. Duan, S. Jia, L. Zhang, Y. Liu, and J. Lin. A Dynamic Hysteresis-Adjusting 238
BIBLIOGRAPHY Algorithm in LTE Self-Organization Networks. In Proc. of IEEE 75th Vehicular Technology Conference (VTC), Spring, 2012. [87] Y. Li, M. Li, B. Cao, Y. Wang, and W. Liu. Dynamic optimization of handover parameters adjustment for conflict avoidance in long term evolution. China Communications, 10(1):56– 71, 2013. [88] S. Horrich, S. Ben Jamaa, and P. Godlewski. Adaptive Vertical Mobility Decision in Heterogeneous Networks. In Proc. of Third International Conference on Wireless and Mobile Communications (ICWMC), 2007. [89] D. Turina and A. Furuskar. Traffic Steering and Service Continuity in GSM-WCDMA Seemless Networks. In Proc. of the 8th International Conference on Telecommunications (ConTEL), 2005. [90] I. de la Bandera, S. Luna-Ramírez, R. Barco, F. Ruiz, M. Toril, and M. Fernández-Navarro. Inter-system Cell Reselection Parameter Auto-Tuning in a Joint-RRM Scenario. In Proc. of 5th International Conference on Broadband Communications and Biomedical Applications, IB2COM, 2010. [91] S. Luna-Ramírez, M. Toril, F. Ruiz, and M. Fernández-Navarro. Adjustment of a Fuzzy Logic Controller for IS-HO Parameters in a Heterogeneous Scenario. In Proc. of the 14th IEEE Mediterranean Electrotechnical Conference, MELECON, 2008. [92] T. Chandra, W. Jeanes, and H. Leung. Determination of optimal handover boundaries in a cellular network based on traffic distribution analysis of mobile measurement reports. In Proc. of 47th IEEE Vehicular Technology Conference (VTC), volume 1, pages 305–309, 1997. [93] J. Steuer and K. Jobmann. The use of mobile positioning supported traffic density measurements to assist load balancing methods based on adaptive cell sizing. In Proc. of 13th IEEE Int. Symp. on Personal Indoor and Mobile Radio Communications, volume 3, pages 339–343, 2002. [94] M. Toril and V. Wille. Optimization of Handover Parameters for Traffic Sharing in GERAN. Wireless Personal Communications, 47(3):315 –336, 2008. [95] S.B. ZahirAzami, G. Yekrangian, and M. Spencer. Load Balancing and Call Admission Control in UMTS-RNC, using Fuzzy Logic. In International Conference on Communication Technology Proceedings (ICCT), 2003. [96] J. Rodríguez, I. de la Bandera, P. Muñoz, and R. Barco. Load Balancing in a Realistic Urban Scenario for LTE Networks. In Proc. of IEEE 73rd Vehicular Technology Conference (VTC), 2011. [97] K. C. Foong, C. T. Chee, and L. S. Wei. Adaptive Network Fuzzy Inference System (ANFIS) Handoff Algorithm. In Proc. of the International Conference on Future Computer and Communication (ICFCC), 2009. [98] L. Giupponi and R. Agustí and J. Pérez-Romero and O. Sallent. A Novel Approach for Joint 239
BIBLIOGRAPHY Radio Resource Management Based on Fuzzy Neural Methodology. IEEE Transactions on Vehicular Technology, 57(3):1789–1805, 2008. [99] A. Çalhan and C. Çeken. An optimum vertical handoff decision algorithm based on adaptive fuzzy logic and genetic algorithm. Wireless Personal Communications, pages 1–18, 2010. [100] M. Dirani and Z. Altman. Self-Organizing Networks in Next Generation Radio Access Networks: Application to Fractional Power Control. Computer Networks, 55(2):431–438, 2011. [101] R. S. Sutton and A. G. Barto. Reinforcement Learning: an Introduction. MIT Press, 1998. [102] Y. H. Chen, C. J. Chang, and C. Y. Huang. Fuzzy Q-Learning Admission Control for WCDMA/WLAN Heterogeneous Networks with Multimedia Traffic. IEEE Transactions on Mobile Computing, 8(11):1469 –1479, 2009. [103] J. Nie and S. Haykin. A dynamic channel assignment policy through Q-learning. IEEE Transactions on Neural Networks, 10(6):1443 – 1455, 1999. [104] A. Galindo-Serrano and L. Giupponi. Distributed Q-learning for aggregated interference control in cognitive radio networks. IEEE Transactions on Vehicular Technology, 59(4):1823 – 1834, May 2010. [105] El-Sayed M. El-Alfy and Yu-Dong Yao. Comparing a class of dynamic model-based reinforcement learning schemes for handoff prioritization in mobile communication networks. Expert Systems with Applications, 38(7):8730–8737, July 2011. [106] R. Razavi, S. Klein, and H. Claussen. A Fuzzy Reinforcement Learning Approach for SelfOptimization of Coverage in LTE Networks. Bell Labs Technical Journal, 15(3):153–175, 2010. [107] A. Galindo-Serrano and L. Giupponi. Downlink femto-to-macro interference management based on Fuzzy Q-Learning. In Proc. of International Symposium on Modeling and Optimization in Mobile, Ad Hoc and Wireless Networks (WiOpt), 2011. [108] A. Galindo-Serrano, L. Giupponi, and M. Majoral. On implementation requirements and performances of Q-learning for self-organized femtocells. In Proc. of IEEE Global Telecommunications Conference (GLOBECOM), 2011. [109] M. Haddad, Z. Altman, S.E. Elayoubi, and E. Altman. A Nash-Stackelberg Fuzzy QLearning Decision Approach in Heterogeneous Cognitive Networks. In Proc. of IEEE Global Telecommunications Conference (GLOBECOM), 2010. [110] M. Toril. Self-Tuning Algorithms for the Assignment of Packet Control Units and Handover Parameters in GERAN. PhD thesis, Communications Engineering Department, ETSIT, University of Málaga, 2007. [111] L. A. Zadeh. Fuzzy Sets. Information and Control, 8:338–353, 1965. [112] T. Ross. Fuzzy logic with engineering applications. Wiley, 2010. 240
BIBLIOGRAPHY [113] A. P. Engelbrecht. Computational Intelligence: An Introduction. John Wiley & Sons, 2007. [114] Chuen Chien Lee. Fuzzy logic in control systems: fuzzy logic controller. I. IEEE Transactions on Systems, Man and Cybernetics, 20(2):404–418, 1990. [115] E. H. Mamdani and S. Assilian. An Experiment in Linguistic Synthesis with a Fuzzy Logic Controller. International Journal of Man-Machine Studies, 7:1–13, 1975. [116] T. Takagi and M. Sugeno. Fuzzy Identification of Systems and its Application to Modeling and Control. IEEE Transactions on Systems, Man, and Cybernetics, 15(1):116–132, 1985. [117] R. C. Eberhart and Y. Shi. Computational Intelligence Concepts to Implementations. Elsevier, 2007. [118] K. A. Smith and M. S. Palaniswami. Static and dynamic channel assignment using neural networks. IEEE Journal on Selected Areas in Communications, 15(2):238–249, 1997. [119] D. Gómez-Barquero, D. Calabuig, J. F. Monserrat, N. García, and J. Pérez-Romero. Hopfield Neural Network-Based Approach for Joint Dynamic Resource Allocation in Heterogeneous Wireless Networks. In Proc. of IEEE 64th Vehicular Technology Conference (VTC), Fall, 2006. [120] C.-T. T. Lin and C.-S. G. Lee. Neural-network-based fuzzy logic control and decision system. IEEE Transactions on Computers, 40(12):1320–1336, 1991. [121] F. Herrera, M. Lozano, and J. L. Verdegay. Tuning Fuzzy Logic Controllers by Genetic Algorithms. International Journal of Approximate Reasoning, 12:299–315, 1995. [122] M. A. Lee and H. Takagi. Integrating design stage of fuzzy systems using genetic algorithms. In Proc. of Second IEEE International Conference on Fuzzy Systems, volume 1, pages 612– 617, 1993. [123] C. L. Karr and E. J. Gentry. Fuzzy control of pH using genetic algorithms. IEEE Transactions on Fuzzy Systems, 1(1), 1993. [124] J. Kinzel, F. Klawonn, and R. Kruse. Modifications of genetic algorithms for designing and optimizing fuzzy controllers. In Proc. of the First IEEE Conference on Evolutionary Computation, volume 1, pages 28–33, 1994. [125] S. Ghost, A. Konar, and A. K. Nagar. Dynamic Channel Assignment Problem in Mobile Networks Using Particle Swarm Optimization. In Proc. of Second UKSIM European Symposium on Computer Modeling and Simulation (EMS), pages 64–69, 2008. [126] P. Garcia-Diaz, S. Salcedo-Sanz, J. Plaza-Laina, and A. Portilla-Figueras. A discrete Particle Swarm Optimization Algorithm for Mobile Network Deployment Problems. In Proc. of IEEE 17th International Workshop on Computer Aided Modeling and Design of Communication Links and Networks (CAMAD), pages 61–65, 2012. [127] L. Su, P. Wang, and F. Liu. Particle swarm optimization based resource block allocation algorithm for downlink LTE systems. In Proc. of 18th Asia-Pacific Conference on Communications (APCC), pages 970–974, 2012. 241
BIBLIOGRAPHY [128] H. Dubreil, Z. Altman, V. Diascorn, J.-M. M. Picard, and M. Clerc. Particle swarm optimization of fuzzy logic controller for high quality RRM auto-tuning of UMTS networks. In Proc. of IEEE 61st Vehicular Technology Conference (VTC), Spring, volume 3, pages 1865–1869, 2005. [129] C. Watkins and P. Dayan. Technical Note: Q-Learning. Machine Learning, 8(3):279 –292, 1992. [130] P. Y. Glorennec. Fuzzy Q-learning and dynamical fuzzy Q-learning. In Proc. of the Third IEEE Conference on Fuzzy Systems, volume 1, pages 474–479, 1994. [131] H. Dubreil. Méthodes d’optimisation de contrôleurs de logique floue pour le paramétrage automatique des réseaux mobiles UMTS. PhD thesis, ENST - COMELEC Communication et Electronique, 2005. [132] D.P. Rini, S. M. Shamsuddin, and S. S. Yuhaniz. Particle Swarm Optimization: Technique, System and Challenges. International Journal of Computer Applications, 14(1):19–27, 2011. [133] Q. Bai. Analysis of Particle Swarm Optimization Algorithm. Computer and Information Science, 3(1):180–184, 2010. [134] E. S. Ali and S. M. Abd-Elazim. Statistical Assessment of New Coordinated Design of PSSs and SVC via Hybrid Algorithm. International Journal of Engineering and Advanced Technology (IJEAT), 2(3):647–654, 2013. [135] M. I. Jiménez Vega. Simulador dinámico de red de comunicaciones móviles de segunda generación. Master’s thesis, Communications Engineering Department, ETSIT, University of Málaga, 2004. [136] M. Guerrero Navarro. Simulador dinámico de red de comunicaciones móviles de tercera generación. Master’s thesis, Communications Engineering Department, ETSIT, University of Málaga, 2005. [137] J. Wu, Z. Yin, J. Zhan, and W. Heng. Physical Layer Abstraction Algorithms Research for 802.11n and LTE Downlink. In Proc. of International Symposium on Signals Systems and Electronics (ISSSE), 2010. [138] J. Olmos, A. Serra, S. Ruiz, M. García-Lozano, and D. Gonzalez. Link Level Simulator for LTE Downlink. In Proc. of 7th European Meeting COST-2100 - Pervasive Mobile & Ambient Wireless Communications, TD(09)779, 2009. [139] J. C. Ikuno, M. Wrulich, and M. Rupp. System Level Simulation of LTE Networks. In Proc. of IEEE 71st Vehicular Technology Conference (VTC), Spring, 2010. [140] C. Mehlführer, M. Wrulich, J. Colom Ikuno, D. Bosanska, and M. Rupp. Simulating the Long Term Evolution Physical Layer. In Proc. of 17th European Signal Processing Conference (EUSIPCO 2009), 2009. [141] T. Hytönen. Optimal Wrap-around Network Simulation. Technical Report , Helsinki, 242
BIBLIOGRAPHY University of Technology Institute of Mathematics, 2001. [142] B. Ahn, H. Yoon, and J. W. Cho. A Design of Macro-micro CDMA Cellular Overlays in the Existing Big Urban Areas. IEEE Journal on Selected Areas in Communications, 19(10):2094 – 2104, 2001. [143] Next Generation Mobile Networks (NGMN) Alliance. NGMN Radio Access Performance Evaluation Methodology, Version 1.0, January 2008. www.ngmn.org . [144] J. D. Parsons. The Mobile Radio Propagation Channel. Pentech, 1992. [145] W. C. Jakes. Microwave Mobile Communications. Wiley, 1974. [146] 3GPP. Evolved Universal Terrestrial Radio Access (E-UTRA); User Equipment (UE) radio transmission and reception (Release 11), version 11.4.0 (2013-03). TS 36.101. [147] E. Bonek. Tunnels, corridors, and other special environments. In L. Correira E. Damosso, editor, COST Action 231: Digital mobile radio towards future generation systems, pages 190–207. European Union Publications, Brüssel, 1999. [148] U. Gotzner, A. Gamst, and R. Rathgeber. Spatial traffic distribution in cellular networks. In Proc. of IEEE 48th Vehicular Technology Conference (VTC), Spring, volume 2, pages 1994–1998, 1998. [149] D. Huo. Simulating slow fading by means of one dimensional stochastical process. In Proc. of IEEE 46th Vehicular Technology Conference (VTC), Spring. ’Mobile Technology for the Human Race’, volume 2, pages 620–622, 1996. [150] M. Gudmundson. Correlation model for shadow fading in mobile radio systems. Electronics Letters, 27(23):2145–2146, 1991. [151] 3GPP. Feasibility study for Orthogonal Frequency Division Multiplexing (OFDM) for UTRAN enhancement (Release 6), version 6.0.0 (2004-06). TR 25.892. [152] 3GPP. System Analysis of the Impact of CQI Reporting Period in DL SIMO OFDMA (R1-061506). 3GPP TSG-RAN WG1 45, Shanghai, China, May 2006. [153] E. Tuomaala and H. Wang. Effective SINR Approach of Link to System Mapping in OFDM/Multi-Carrier Mobile Network. In Proc. of 2nd International Conference on Mobile Technology, Applications and Systems, 2005. [154] 3GPP. OFDM-HSDPA System level simulator calibration (R1-040500). 3GPP TSG-RAN WG1 37, Montreal, Canada, May 2004. [155] 3GPP. Evolved Universal Terrestrial Radio Access (E-UTRA); User Equipment (UE) conformance specification Radio transmission and reception Part 1: Conformance Testing; (Release 11), version 11.0.1 (2013-03). TS 36.521. [156] J. T. Entrambasaguas, M. C. Aguayo-Torres, G. Gomez, and J. F. Paris. Multiuser Capacity and Fairness Evaluation of Channel/QoS-Aware Multiplexing Algorithms. IEEE Network, 21(3):24–30, 2007. 243
BIBLIOGRAPHY [157] D. Hong and S. S. Rappaport. Traffic Model and Performance Analysis for Cellular Mobile Radio Telephone Systems with Prioritized and Nonprioritized Handoff Procedures. IEEE Transactions on Vehicular Technology, 35(3):77–92, 1986. [158] M. Kazmi, O. Sjobergh, W. Muller, J. Wierok, and B. Lindoff. Evaluation of InterFrequency Quality Handover Criteria in E-UTRAN. In Proc. of IEEE 69th Vehicular Technology Conference (VTC), Spring, 2009. [159] J. M. Ruiz-Avilés, S. Luna-Ramírez, M. Toril, F. Ruiz, I. de la Bandera, P. Muñoz, R. Barco, P. Lázaro, and V. Buenestado. Design of a Computationally Efficient Dynamic System-Level Simulator for Enterprise LTE Femtocell Scenarios. Journal of Electrical and Computer Engineering, 2012, 2012. [160] ETSI. Universal Mobile Telecommunications System (UMTS), Selection procedures for the choice of radio transmission technologies of the UMTS, version 3.2.0 (1998-04). TR 101 112. [161] I. Viering, M. Döttling, and A. Lobinger. A Mathematical Perspective of Self-Optimizing Wireless Networks. In Proc. of International Conference on Communications (ICC ’09), 2009. [162] WINNER II. Channel Models. Part II. Radio Channel Measurement and Analysis Results. D1.1.2. v1.0, WINNER II IST project, 2007. [163] T. Sorensen, P. Mogensen, and F. Frederiksen. Extension of the ITU channel models for wideband (OFDM) systems. In Proc. of IEEE 62nd Vehicular Technology Conference (VTC), 2005. [164] L. Kaelbling, M. Littman, and A. Moore. Reinforcement learning: a survey. Journal of Artificial Intelligence Research, 4:237 – 285, 1996. [165] 3GPP. Evolved Universal Terrestrial Radio Access (E-UTRA) User Equipment (UE) procedures in idle mode, version 10.3.0 (2011-09). TS 36.304. [166] 3GPP. Technical Specification Group Radio Access Network; Radio Resource Control (RRC); Protocol Specification, version V11.0.0 (2011-12). TS 25.331. [167] M. N. Halgamuge, H. L. Vu, K. Ramamohanarao, and M. Zukerman. A call quality performance measure for handoff algorithms. Int. J. Commun. Syst., 24:363–383, 2011. [168] 3GPP. Evolved Universal Terrestrial Radio Access (E-UTRA); Requirements for support of radio resource management, version 11.4.0 (2013-03). TS 36.133. [169] P. Muñoz, R. Barco, and I. de la Bandera. Optimization of load balancing using fuzzy Q-Learning for next generation wireless networks. Expert Systems with Applications, 40(4):984–994, 2013. [170] P. Muñoz, R. Barco, I. de la Bandera, M. Toril, and S. Luna-Ramírez. Optimization of a Fuzzy Logic Controller for Handover-Based Load Balancing. In Proc. of IEEE 73rd Vehicular Technology Conference (VTC), 2011. 244
BIBLIOGRAPHY [171] P. Muñoz, R. Barco, J. Ruiz-Avilés, I. de la Bandera, and A. Aguilar. Fuzzy Rule-based Reinforcement Learning for Load Balancing Techniques in Enterprise LTE Femtocells. IEEE Transactions on Vehicular Technology, accepted in 2012. [172] P. Muñoz, R. Barco, and I. de la Bandera. On the Potential of Handover Parameter Optimization for Self-Organizing Networks. IEEE Transactions on Vehicular Technology, accepted in 2013. [173] P. Muñoz, D. Laselva, R. Barco, and P. Mogensen. Adjustment of mobility parameters for traffic steering in multi-RAT multi-layer wireless networks. EURASIP Journal on Wireless Communication and Networking, 2013:133, 2013. [174] P. Muñoz, R. Barco, D. Laselva, and P. Mogensen. Mobility-based Strategies for Traffic Steering in Heterogeneous Networks. IEEE Communications Magazine, 51(5):54–62, 2013. 245