An adaptive cross-layer framework for multimedia delivery over heterogeneous networks
Full text
FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO An Adaptive Cross-Layer Framework for Multimedia Delivery over Heterogeneous Networks Daniel Oancea Programa Doutoral em Engenharia Electrotecnica e de Computadores Supervisor: Maria Teresa Andrade (Professor Auxiliar) December 11, 2014
c Daniel Oancea, 2014
An Adaptive Cross-Layer Framework for Multimedia Delivery over Heterogeneous Networks Daniel Oancea Programa Doutoral em Engenharia Electrotecnica e de Computadores Approved by . . . : President: Name of the President (Title) Referee: Name of the Referee (Title) Referee: Name of the Referee (Title) 2014
Abstract Multimedia computing has grown from a specialist field into a pervasive aspect of everyday computing systems. Methods of processing and handling multimedia content have evolved considerably over the last years. Where once specific hardware support was required, now the increased power of computing platforms has enabled a more softwarebased approach. Accordingly, nowadays the particular challenges for multimedia computing come from the mobility of the consumer allied to the diversity of client devices, where wide-ranging levels of connectivity and capabilities require applications to adapt The major contribution of this work is to combine in an innovative way, different technologies to derive a new approach, more efficient and offering more functionality, to adapt multimedia content, thus enabling ubiquitous access to that content. Based on a reflective design, we combine semantic information obtained from different layers of the OSI model, namely from the network and application levels, with an inference engine. The inference engine allows the system to create additional knowledge which will be instrumental in the adaptation decision taking process. The reflective design provides facilities to enable inspection and adaptation of the framework itself. The framework implementation is totally device-agnostic, thus increasing its portability and consequently widening its range of application. The definition of this approach and the consequent research conducted within this context, led to the development of a putative framework called Reflective Cross-Layer Adaptation Frameworg (RCLAF). The RCLAF combines ontologies and semantic technologies with a reflective design architecture to provide support in an innovative way towards multimedia content adaptation. The developed framework validates the feasibility of the principles argued in this thesis. The framework is evaluated in quantitative terms, addressing the range of adaptation possibilities it offers, thus highlighting its flexibility, as well as in qualitative terms. The overall contribution of this thesis is to provide reliable support in multimedia content adaptation by promoting openness and flexibility of software. i
ii
Resumo Os ´ ultimos anos registaram uma evoluc¸˜ ao significativa na ´ area das aplicac¸ ˜ oes multim´ edia, a qual deixou de ser vista como como uma ´ area especializada, passando a constituir uma parte integrante e omni-presente dos sistemas de computac¸˜ ao que utilizamos diariamente. Para tal contribuiu significativamente o progresso que se operou nas abordagens e m´ etodos de processamento de conte´ udo multim´ edia. Onde anteriormente era necess´ ario recorrer a hardware espec´ ıfico para fazer o tratamentos dos dados multim´ edia, hoje em dia o crescente poder computacional permite obter o mesmo resultado usando apenas software. Assim, actualmente, os grandes desafios que se colocam na ´ area das aplicac¸ ˜ oes multim´ edia, est˜ ao relacionados com a mobilidade dos consumidores aliada ` a diversidade dos dispositivos m´ oveis e a consequente variedade de ligac¸ ˜ oes de rede de que podem usufruir. Estes desafios levam a que sejam consideradas formas de adaptar dinˆ amicamente as aplicac¸ ˜ oes e os conte´ udos a que os consumidores acedem, por forma a satisfazer diferentes requisitos. A contribuic¸˜ ao mais importante deste trabalho ´ e a de combinar de forma inovadora diferentes tecnologias para obter uma abordagem mais eficiente e com um maior leque de funcionalidades, para o desafio de adaptar de forma dinˆ amica aplicac¸ ˜ oes e conte´ udos multim´ edia, permitindo assim o consumo ub´ ıquo desses conte´ udos. A abordagem adoptada, ´ e baseada num modelo reflectivo, fazendo uso de metadados semˆ anticos obtidos de diferentes camadas do modelo OSI, em espec´ ıfico das camadas de rede e aplicac¸˜ ao, e de um motor de inferˆ encia. Este permite obter conhecimento adicional a partir dos metadados semˆ anticos recolhidos, auxiliando os mecanismos de decis˜ ao de adaptac¸˜ ao. Por outro lado, o modelo reflectivo permite, para al´ em de adaptar o conte´ udo, com que seja possivel adaptar a pr´ opria plataforma onde correm as aplicac¸ ˜ oes multim´ edia. Esta plataforma ´ e completamente agn´ ostica ` a arquitectura hardware e sistema operativo do dispositivo onde est´ a a correr, aumentando dessa forma a sua portabilidade e o ˆ ambito de aplicac¸˜ ao. A definic¸˜ ao desta abordagem e a investigac¸˜ ao realizada nesse contexto, levaram ao desenvolvimento da plataforma Reflective Cross-Layer Adaptation Framework (RCLAF), combinando ontologias e tecnologias semˆ anticas com uma arquitectura reflectiva, oferecendo capacidades inovadoras de adaptac¸˜ ao de conte´ udos e aplicac¸ ˜ oes multim´ edia. Esta plataforma permite validar todos os conceitos e princ´ ıpios defendidos nesta tese. A sua avaliac¸˜ ao ´ e efectuada de forma quantitativa, relativa ` as possibilidades de adaptac¸˜ ao que oferece e distinguindo desta forma a sua versatilidade, quer de forma qualitativa, relativa ` a forma como essas possibilidades s˜ ao implementadas. Em termos globais, a contribuic¸˜ ao desta tese ´ e a de oferecer suporte fi´ avel e flex´ ıvel para adaptac¸ao de conte´ udos multim´ edia, promovendo a abertura e flexibilidade do software. iii
iv
Acknowledgements I would like to express my thanks to those who helped me in various aspects of conducting research and the writing of this thesis. Firstly, the author would like to express his thanks to FCT Portugal who supported his work through the FCT grant no. SFR/BD/29102/2006. Secondly, thanks are directed to the INESC Porto and Faculty of Engineering of University of Porto who offered the logistic support author need for concluding the work. Thirdly, author’s colleagues within the Telecommunication and Multimedia Unit of INESC Porto must be thanked for the work environment. Above all, the author’s supervisor, Professor Dr. Maria Teresa Andrade, must be thanked for providing a constructive criticism in which the ideas presented within this thesis were formed over the last few years. Special thanks must also go to the author’s family, who has provided inestimable support throughout, to Miguel and Javier for the time spent together, to Vasile Palade for the support and to Evans for interesting talks around software engineering. Finally, the author must thank his girlfriend, Regina, for her unwavering love and support through a difficult and challenging time. Daniel Oancea v
xii LIST OF FIGURES 5.13 The UML activity diagram of the reasoning operation . . . . . . . . . . . 97 5.14 The UML sequence diagram of the reify operation . . . . . . . . . . . . . 98 5.15 The UML activity diagram of the reifyNetInfo operation . . . . . . . . . . 99 5.16 The CLS ontology model . . . . . . . . . . . . . . . . . . . . . . . . . . 102 5.17 The AS ontology model . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 5.18 The AS taxonomy hierarchy . . . . . . . . . . . . . . . . . . . . . . . . 103 5.19 A multi-step Decision Model . . . . . . . . . . . . . . . . . . . . . . . . 104 6.1 The UML package diagram of the RCLAF architecture . . . . . . . . . . 106 6.2 The class diagram of the OWL-manager package .............115 6.3 An overview of the webOntFact package..................118 6.4 The CLS hierarchymodel..........................120 6.5 Class and Property representation of CodecParameter . . . . . . . . . . . 122 6.6 The AS hierarchymodel...........................124 6.7 AnASmodelexample ...........................125 6.8 Process Diagram of Rule Execution . . . . . . . . . . . . . . . . . . . . 129 6.9 An overview of the ServerFact package ..................130 6.10 An overview of the terminal package....................134 6.11 The UML smart API package diagram . . . . . . . . . . . . . . . . . . . 135 6.12 Class diagram for Android API . . . . . . . . . . . . . . . . . . . . . . . 136 7.1 Architecture for an adaptive VoD application . . . . . . . . . . . . . . . 141 7.2 Terminal User Interface . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 7.3 Terminal GUI - getMedia . . . . . . . . . . . . . . . . . . . . . . . . . . 144 7.4 Terminal GUI: Select operation . . . . . . . . . . . . . . . . . . . . . . . 146 7.5 termGUI - playing the content . . . . . . . . . . . . . . . . . . . . . . . 146 7.6 termGUI adaptation process . . . . . . . . . . . . . . . . . . . . . . . . 147 7.7 termGUI the adapted content . . . . . . . . . . . . . . . . . . . . . . . . 148 7.8 Playing the video in Android emulator . . . . . . . . . . . . . . . . . . . 149 7.9 The interaction of the CLMAE component with Terminal and ServerFact components.................................149 7.10 Declaration of the ubuntu terminal into scenario ontology . . . . . . . . . 151 7.11 Declaration of the Reklam’s variation (reklam 1 into scenario ontology) . 152 7.12 The multistep adaptation decision process . . . . . . . . . . . . . . . . . 153 7.13 The parking lot topology for network simulator with Evalvid . . . . . . . 162 7.14 Overall frame loss for the testing scenario described in Section 7.1.2.4 . . 163 7.15 End-to-End delay for reklam5 and reklam3 .................163 7.16 Jitter for reklam5 and reklam3 .......................164 7.17 Various rates captured at receiver for reklam5 and reklam3 variations . . . 165 7.18 Get media response time . . . . . . . . . . . . . . . . . . . . . . . . . . 172 7.19 Introspect network response time . . . . . . . . . . . . . . . . . . . . . . 173 7.20 Introspect media characteristics response time . . . . . . . . . . . . . . . 173 A.1 The ADTE execution time for various media variations. . . . . . . . . . . 188 A.2 The CPU usage of the ADTE reasoning process for several media variations189 A.3 Memory usage of the ADTE when 1,5,10,20,50 and 100 terminals are instantiated into scenario ontology.....................190
LIST OF FIGURES xiii B.1 CPU time and allocated memory for getMedia method when is access it by1and5users...............................192 B.2 CPU time and allocated memory for getMedia method when is access it by10and20users .............................193 B.3 CPU time and allocated memory for getMedia method when is access it by 50 respectively 100 users . . . . . . . . . . . . . . . . . . . . . . . . 194 B.4 CPU time and allocated memory for introspectNet method when is access itby1and5users..............................195 B.5 CPU time and allocated memory for introspectNet method when is access itby10and20users ............................196 B.6 CPU time and allocated memory for introspectNet method when is access itby50and100users............................197 B.7 CPU time and allocated memory for introspectMediaCharact method when is access it by 1 and 5 users . . . . . . . . . . . . . . . . . . . . . . . . . 198 B.8 CPU time and allocated memory for introspectMediaCharact method when is access it by 10 and 20 users . . . . . . . . . . . . . . . . . . . . . . . 199 B.9 CPU time and allocated memory for introspectMediaCharact method when is access it by 50 and 100 users . . . . . . . . . . . . . . . . . . . . . . . 200 B.10 CPU time and allocated memory for reflect method when is access it by 1and5users ................................201 B.11 CPU time and allocated memory for reflect method when is access it by 10and20users ...............................202 B.12 CPU time and allocated memory for reflect method when is access it by 50and100users ..............................203
xiv LIST OF FIGURES
List of Tables 2.1 Adaptation techniques to control multimedia delivery . . . . . . . . . . . 22 2.2 Cross-Layer approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 2.3 An example of initial state and goal state declarations . . . . . . . . . . . 29 2.4 scaling operation for image.yuv picture................... 29 2.5 The adaptation plan for image.yuv spatial scaling . . . . . . . . . . . . . 29 2.6 Support for adaptation of current multimedia applications/frameworks . . 41 3.1 General comparison of the MPEG-7 ontologies . . . . . . . . . . . . . . 52 3.2 General comparasion of multimedia ontology . . . . . . . . . . . . . . . 56 4.1 Adaptive middleware requirements in terms of language requirements . . 74 4.2 Adaptive middleware requirements . . . . . . . . . . . . . . . . . . . . . 75 4.3 Paradigms used over adaptive architectures . . . . . . . . . . . . . . . . 76 4.4 Adaptive middleware categorized by adaptation type . . . . . . . . . . . 77 5.1 The ServFactImpl services ......................... 89 5.2 The TerminalMediaPlayer operations ................... 90 5.3 The OntADTEImpl services......................... 95 6.1 The stateSignal flags ............................111 6.2 The ReasonerManagerInterf methods used to query the ontologies . . . . 117 6.3 The MPEG-21 segments . . . . . . . . . . . . . . . . . . . . . . . . . . 121 6.4 The TBox statements used in AS ontology . . . . . . . . . . . . . . . . . 126 7.1 Characteristics of the media variations used in demonstration example . . 142 7.2 termGUI’s characteristics..........................153 7.3 The network bandwidth values used in example 7.1.2.4 . . . . . . . . . . 153 7.4 Arguments evaluation of the scenario ontology SWRL rule atom (built-In) 155 7.5 Arguments evaluation of the AS ontology SWRL rule 6.2 atom (built-In) lessThan (?cmeasured,?smeasured) ....................159 7.6 Arguments evaluation of the AS ontology SWRL rule 6.6 atom (built-In) lessThan(?update,?val) ...........................159 7.7 CLMAE decision parameters . . . . . . . . . . . . . . . . . . . . . . . . 161 7.8 Measured average MOS in a sliding interval of 25 frames (test1) . . . . . 162 7.9 Average MOS for testing scenario described in Section 7.1.2.4 . . . . . . 163 7.10 Fulfillment of requirements . . . . . . . . . . . . . . . . . . . . . . . . . 167 7.11 The reasoning method invocation response times: same machine . . . . . 174 xv
xvi LIST OF TABLES 7.12 The reasoning method invocation response times: distributed environment ....................................175 7.13 The runInitialSetUp method execution . . . . . . . . . . . . . . . . . . . 175 7.14 The methods executed inside runInitialSetUp method . . . . . . . . . . . 175 7.15 The methods executed inside stateSignal method .............176 7.16 The methods executed inside reifyNetInfo method.............176 7.17 The execution time of the ADTE methods on scenario and AppState ontology ....................................177 7.18 The adaptation operation execution time . . . . . . . . . . . . . . . . . . 177 7.19 Execution time of the methods getMedia and select on Android emulator . 177 B.1 The Terminal HotSpot Methods . . . . . . . . . . . . . . . . . . . . . . 191
xvii
xviii Abbreviations Abbreviations ADSL Asymmetric Digital Subscriber Line AP Access Point ATM Asynchronous Transfer Mode AVO Audiovisual Objects BS Base Station CAC Call Admission Control CSCW Computer Supported Cooperative Work DAML DARPA Agent Markup Language DS Differentiated Service EJB Enterprise Java Beans FDDI Fiber Distributed Data Interface GPRS General Packet Radio Service GSM Global System for Mobile communications HSPDA High Speed Downlink Packet Access gBSD generic Bitstream Syntax Description IETF Internet Engineering Task Force IPv4 Internet Protocol version 4 IPTV Internet Protocol Television ISDN Integrated Services Digital Network MAC Medium Access Control MLPS Multiprotocol Label Switching MOP Metaobject Protocol NM Network Monitor OIL Ontology Inference Layer OWL Ontology Web Language PDA Personal Digital Assistant PPV Pey-Per-View QoS Quality of Service RSVP Resource Reservation Protocol RTP Real-Time Transmission Protocol RTCP Real-Time Control Transmission Protocol SD Standard Definition SHOE Simple HTML Ontology Extensions SNR Signal-to-Noise Ratio SP Service Provider TCP Transport Control Protocol UDP User Data Protocol UEP Unequal Error Protection UMA Universal Multimedia Access UMTS Universal Mobile Telecommunications System VDP Video Data Protocol VoD Video on Demand VSDL Very-high-bit-rate Digital Subscriber Line W3C World Wide Web Consortium WAN Wide Area Networks
Chapter 1 Introduction Nowadays, multimedia consumption is no longer bound to special purpose devices such as analog TV’s and radios. With the “digital age” the number of multimedia-enabled terminals on the market has increased greatly. The new generation of set-top-boxes are capable of rendering high-definition video content and of accessing a wide range of interactive multimedia services, such as PPV, VoD, Time-Shifting, or even Internet content. At the same time, the capabilities of mobile devices continue to increase steadily. The result is the appearance of smart-phones with multimedia capabilities, such as: the iPhone, and Windows or Android-OS based phone terminals. These developments have led to a various number of distribution channels for providers of multimedia content. Now, the content providers (CP) can deliver their content to the final users via standard broadcast networks (e.g., satellite, cable and terrestrial) or via Internet access technologies (e.g., ADSL, VSDL, FTTH, GPRS and UMTS). However, when CPs use the Internet to deliver multimedia content, they are faced with the following problem: the Internet is still a best-effort network, thus there are no Quality of Service (QoS) guarantees. 1.1 Context, Motivation and Objectives With today’s IPTV technology, content that was traditionally distributed through broadcast channels (e.g., radio stations, TV stations, and news agencies) can now be distributed over the Internet too. This trend opens up new business models for CPs, and offers users easier, wider, and flexible access to any kind of information, such as the latest news headlines, via the Internet. However, the Internet is still a best-effort network which means that users do not obtain a guaranteed bit rate or delivery time, as such parameters depend on the instantaneous traffic conditions. This makes the distribution of multimedia content over the Internet problematic: compared with data transfer (e.g., FTP), multimedia content is characterized by the high bit rates it requires, burstiness (imposing random peaks 1
2Introduction of workload on the network), and by its own sensitivity to delay, jitter and data loss. “Onthe-fly” media consumption or live consumption of a TV channel over the Internet may experience problems caused by network bandwidth fluctuations or bottlenecks. The CPs have two ways to deliver the content: through the Service Providers (SP) (e.g., cable companies, Internet Server Providers, or ISP carriers) or directly, using the public Internet infrastructure. In the former scenario, the CP is not faced with the aforementioned problems because the SP’s access networks (e.g., local loops) are dimensioned in such a way as to ensure the QoS to the customers (e.g., packet prioritization techniques as we will see in the comming chapter). In the latter scenario, so called over-the-top technologies (OTT) , the CP cannot guarantee QoS to their customers, because the traffic the CP generates is mixed (shared) with all the other traffic on the Internet. To be able to successfully deploy the new business models associated to the Internet, CPs must be able to offer a reliable and good quality service to their customers. Therefore, to allow CPs to take advantage of the Internet to deliver their contents to the final user, solutions are required to overcome the identified problems. This is precisely the objective of this thesis—to investigate approaches and propose an efficient solution that may enable the distribution of good quality audiovisual content to end users in heterogeneous networked environments, using the public Internet infrastructure. As will be explained in the next section, we believe that such a solution can be based on the ability to dynamically adapt the content to the varying conditions of the network and consumption context. 1.2 Proposed Approach Considerable effort has been and is still being made worldwide by the scientific community to find solutions to the identified challenges. As a consequence of such efforts, several frameworks have been proposed to handle the delivery of multimedia content over the Internet. This is the case of the frameworks koMMA, CAIN, Enthrone or DCAF, whose comprehensive descriptions are given in Section 2.4). These frameworks take decisions and adapt the content that is being delivered, taking into consideration the constraints imposed by the consumption context, notably: the terminal capabilities; the condition of the network; and the preferences of the user. Accordingly, they already provide some intelligence to the decision making process. However, they lack support for the expression of semantically richer descriptions of the context and user preferences, as well as the possibility of reasoning on top of that information and deriving additional knowledge. This could be very important for enabling wiser and flexible decision making, enabling the CPs to offer better quality services and thus contributing to increase user satisfaction. For example, different decisions could be taken for the same terminal and network constraints, depending on the type of content being transmitted or on the situation of the user when consuming the content. Another fragility that these frameworks present comes from the
1.2 Proposed Approach 3 fact that they do not include flexible mechanisms to automatically reconfigure the client application in an adequate manner to start receiving a video stream with different characteristics, when content is adapted on-the-fly. Although in the meantime, some of the problems mentioned have been solved by other solutions, at the time the research was conducted and thesis was written, they were still challenges. We argue that the use of an intelligent approach that is able to use data collected from the network and application level to infer new knowledge, together with reflective techniques, can provide an effective and efficient solution to the indicated limitations. It is our belief that using an ontology-based knowledge system as a main data structure for an adaptation decision engine, where all the content and data for a specific adaptation decision request are stored, will offer wiser and more flexible support capabilities for content adaptation in multimedia applications. The information collected from the network and application level (e.g., network, terminal and user and consumption characteristics) will be captured into ontologies and, using a state of the art logical reasoning service, as for instance rule engines or ontology reasoners, we will be able to infer logical facts (decisions) concerning multimedia content adaptation. The use of reflection techniques provides mechanisms to automatically inspect the framework and consequently re-configure it to start receiving an adapted video bitstream, with different characteristics from the one that was being presented before the adaptation. We take into consideration two different VoD scenarios of consumption of high quality multimedia content via the Internet. In the first scenario, a user arrives home and decides to see the video news via headlines. Using a notebook, which is wired to its home network, the user accesses the VoD news portal of that user’s SP. The user then chooses the preferred headline story and starts to consume the media content at Standard Definition (SD) resolution (640x480). Suppose that while the user is watching the video, the user’s spouse uses another PC connected to the same home network, and starts to download a large data file. This download will increase the data traffic over their home access network, with consequences on the quality of the streaming. To keep seeing the video, the SP must react and take adaptive measures: changing the video bit rate and video resolution. In the second scenario, the user starts watching the video at SD resolution on a notebook wired to its home network. It is a nice evening and at a particular moment, the user decides to move to the backyard to relax. The user takes the PDA, connects it to the WiFi home network, and resumes the play-out of the video. Unfortunately, the PDA does not support such a high resolution and, since the user is still willing to see the video, the stream characteristics must be changed: bit rate and resolution. We integrate the content, media and network information and use an inference engine in a reflective architecture to solve the aforementioned adaptation problems.
10 Introduction
Chapter 2 Related Work on Multimedia Adaptation This chapter presents the state-of-the-art on multimedia content adaptation. It is split into three parts, taking into consideration that multimedia adaptation operations could be implemented at different layers of the ISO/OSI reference model, The first part, highlights the adaptation techniques at the transportlevel. In the second part, the most relevant techniques and approaches encountered at the application-level are presented, whereas in the third part, the cross-layer approaches are exposed. In our opinion, multimedia adaptation can be tackled by integrating a knowledgemodel and an inference engine in a reflective cross-layer architecture. Accordingly, Chapter 3 and Chapter 4 are dedicated to these technologies. These two chapters are not intended to be a survey, but are rather meant to allow identifying the key issues behind adaptation. The main contribution of the present chapter is the identification of the relevant design approaches techniques which serve as basis for designing and implementing the RCLAF framework, emphasizing their strengths and limitations based on an analysis of the published literature. 2.1 Introduction Before presenting existing work on multimedia adaptation techniques it is necessary and important to define the relevant meaning of the term adaptation. The adaptation process is used across many fields and communities and we need to clarify its meaning and why do we need adaptation in multimedia content delivery. 11
12 Related Work on Multimedia Adaptation 2.1.1 Definitions In the Oxford Dictionary 1, the meaning of the noun adaptation is explained as follows: “1) the action or process of adapting or being adapted; 2) a film, television drama, or stage play that has been adapted from a written work and 3) process of change by which an organism or species becomes better suited to its environment.” The concept of adaptation thus refers to behavior change. In computer science the adaptation is seen as the capability of a system to configure itself according with changing conditions. Conditions in this context are the surrounding parts of the system, which we call environment (e.g., network bandwidth, user preferences, terminal characteristics). Thus, we can formulate a more in-depth definition of adaptation concerning software systems: Definition 1. Adaptation is a system process which modifies its behavior to enable or improve its external interactions with individual users based on information acquired about its user(s) and its environment. As the software market continuously evolves, the software systems need to adjust or improve. The implementation of human-centered middleware and framework platforms require instant (real-time) adaptation due to their exposure to dynamics changes. Rasenack [Ras09], classifies the adaptation process in two branches: adaptation environment and adaptation behavior. adaptation envinronment behaviour resources operating system component technology interface functionality Figure 2.1: Rasenack’s adaptation branches [Ras09] The adaptation environment aims at handling adaptation of software components 2to their operating environment. As the Figure 2.1 shows the first sub-branch of the environment consists in resources and is related with adaptation of software components that are running on different hardware machines. A classical example is the adaptation of software components installed on a personal computer to be able to run as well on mobile devices (e.g., PDA, Smartphone). The second sub-branch is the operating system, and refers to the software components that should adapt to different operating systems. It is 1http://oxforddictionaries.com 2The terminology software component refers to encapsulated software functionality such as: one or more logical processes, organization-related processes or tasks
2.1 Introduction 13 very well known that the input-output (I/O) operations of UNIX 3and Windows 4operating systems are different. Therefore, the software components that are developed to run on both should have an adaptive filter (e.g., a factory method) for accessing correctly the I/O of the operating system. The third sub-branch is the component technology in that the components are developed and deployed. Some notable examples are the EJBs components, which are expected to run on suitable EJB servers. If they are used in a different way they have to adapt to the target environment. The adaptation behavior branch deals with the adaptation of software components behavior to support the assemble of applications. The software systems encapsulate several software components. These components can be developed using different standards or standards that are not sufficiently consistent in use. Therefore, integrating these components in an application could raise some problems due to the fact that some of them could have incompatible interfaces. In that situation, it will not be possible to exchange messages between them. In this research work, the adaptation plays an important role in the consumption of multimedia content. Software systems designed for multimedia applications may experience dynamic changes, especially when they operate in mobile environments. For example when users are moving around and the network bandwidth becomes fluctuant, it is desirable to automatically and dynamically trigger an adaptation process which will allow them to successfully continue to consume media content even in situations of bandwidth constraints. This multimedia content adaptation is seen as a logical-process that aims at changing multimedia content characteristics such as: spatial or temporal video frame characteristics, video bit-rate, encoding format, etc. A spatial adaptation could be the adjustment of the video frame resolution, whereas the temporal adaptation deals with changing the refresh rate of video frames. Adaptation of bit rate may be necessary when there is limitation in the available transmission bandwidth. The last referred adaptation is undertaken in cases where a target device does not support a particular multimedia content format. There are four architectural design solutions supporting distributed content adaptation: server-side,client-side,proxy-based and service-oriented approach [CL04]. The adaptation decisions taken by the content adaptation process could be implemented at different layers of the ISO/OSI reference model. In Section 2.2 the techniques and approaches used to implement the adaptation decisions are presented. The architectural design solutions are described in Section 2.3. 3A trademark of Open Group 4A trademark of Microsoft corporation
14 Related Work on Multimedia Adaptation 2.2 Adaptation Techniques 2.2.1 Transport-Layer Adaptation At the transport-layer, to efficiently deliver multimedia over heterogeneous networks, it is important to estimate the status of the underlying networks so that multimedia applications can adapt accordingly. The IETF has proposed many service models and mechanisms to meet the demands of QoS multimedia applications. RTP/RTCP [SCFJ98] seems to be the “de facto” standard for delivering multimedia over the Internet, offering capabilities for monitoring the transmission. RSVP [ZBHJ97] allows to reserve resources presents difficulties in terms of implementation. To overcome the implementation problem, Differentiated Services(DS) [BCS94] are introduced. With DS is possible to create several differentiated service classes by marking packets differently using the IPv4 byte header called Type-of-Service or “DS filed”. Another protocol, Multi protocol Label Switching (MPLS) [RVC01] marks the packets with labels at the ingress of an MPLS-domain. This way, applications may indicate the need for low-delay, high-throughput or low-loss-rate services. One of the most important issues in heterogeneous networks is to detect current available bandwidth and perform efficient congestion control. A proper congestion control scheme should maximize the bandwidth utilization and at the same time should avoid overusing network resources which may cause congestions. Traffic Engineering [AMA+99] with Constraint-Based Routing [CNRS98] tackles these issues by managing how the traffic is flowing through the network and identifying routes that are subject to some constraints, such as bandwidth or delay. Another approach for the transmission of the multimedia data is the use of UDP (transport content) in conjunction with TCP [BTW94] (for control, play, pause, etc), VDP [CTCL95] that permits audio and video presentations to start before the source files are downloaded completely, or delivery over differentiated service network [FSLX06]. The last approach uses an adaptation algorithm that overcomes the shortcoming of besteffort networks and the low bandwidth utility ratio problem of the existing differentiated service network scheduling scheme [CT95]. The CAC algorithm [NS96] is a solution to dynamically adjust the bandwidth of ongoing calls [KCBN03]. 2.2.2 Application-Layer Adaptation There are many topics that may be addressed at the application level to adapt media delivery quality. More specifically, recent progresses on adaptive media codecs, error protection schemes, power saving approaches, resource allocation and OTT’s adaptive streaming solutions are reviewed in this section.
2.2 Adaptation Techniques 15 Media codecs have the ability to dynamically change the coding rate and other coding parameters to adapt the varying network conditions (bandwidth, loss, delay, etc.). Scalable coding techniques were introduced to realize this type of media adaptation. The major technique to achieve this scalability is the layered coding technology, which divides the multimedia information into several layers. For example, the video coding techniques which utilize the discrete cosine transform (DCT), such as AVC/H.264(MPEG-4 part 10 of specification) [ISJ06] and MPEG-4, categorize layered coding techniques into three classes: temporal,spatial and SNR scalability [Li01]. Video codecs, with temporal or spatial scalability encode a video sequence into one base layer and multiple enhancement layers. The base layer is encoded with the lowest temporal or spatial resolution, and the enhancement layers are coded upon the temporal or spatial prediction to lower layers. SNR scalable codecs encode a video sequence into several layers with the same temporal and spatial fidelity. The base layer will provide the basic quality, and enhancement layers are coded to enhance the video data quality when added back to the base layer [Gha99]. However, the layered codec requires the enhancement layers to be fully received before decoding. Otherwise, they will not provide any quality improvement at all. Besides the varying network conditions, there are also packet losses and bit errors in heterogeneous networks. Therefore, a protection scheme is essential for improving endto-end media quality. Automatic Repeat Request (ARQ) and Forward Error Correction (FEC) are two basic error correction mechanisms. FEC is a channel coding technique protecting the source data at the expense of adding redundant data during the transmission. FEC has been commonly suggested for applications with strict delay requirements such as voice communications [BLB03]. In the case of media transmission, where delay requirements are not that stricts, or round-trip delay is small (e.g., video/audio delivery over a single wireless channel), ARQ is applicable and usually plays a role as a complement to FEC. In mobile environment (WLAN), Unequal Error Protection (UEP) and ARQ are applied [LvdS04]. As mentioned previously, layered scalable media codec usually divides media into a base layer and multiple enhancement layers. Since the correct decoding of enhancement layers depends on the errorless receipt of the base layer it is natural to adopt UEP for layered scalable media. Specifically stronger FEC protection can be applied to the base and lower layer data while weak channel coding protection level is applied to the higher layer parts. How to decide the distribution between source codes and channel codes is a problem of bit allocation, which takes into consideration the existing limited resources. An analytic model describing the relation between media quality and source/channel parameters must be developed. Optimal bit allocation can be addressed by numeric methods such as dynamic programming and mathematical functions like penalty functions or Lagrange multiplier. Several bit allocation schemes have been developed, taking in consideration different kinds of scalable media codec channel models into account [ZZZ04, ZWX+04,
16 Related Work on Multimedia Adaptation ZZZ03]. As well, to achieve the optimal end-to-end quality, by adjusting the source and channel coding parameters, the Join-Source-Channel-Coding (JSCC) schemes may be used [ZEP+04]. On wireless environments, the packet losses are caused by the network congestion and wireless transmission errors. Since different losses lead to different perceived QoS at the application level, Yang [YZZZ04] proposed a loss differentiated rate-distortion based bit allocation scheme which minimizes the end-to-end video distortion. In addition to optimize the quality of media streaming for mobile users over wireless environments, we also need to consider the constraints imposed by limited battery power. How to achieve a good user’s perceived QoS while minimizing the power consumption is a challenge. In order to maintain a certain transmission quality, larger transmission rate in wireless channels inherently needs more power and more power also allows adopting more complicated media encoding algorithms with higher complexity and thus can achieve better efficiency. Therefore, an efficient way to obtain optimal media quality is to jointly consider source-channel coding and power consumption issues. Usually, JSCC are targeted at minimizing the power consumed by a single user. For instance, a low-power communication system for image transmission has been developed in [GAS+99]. In [ELP+02] quantization and mode selection have been investigated. The power consumption for source coding and channel coding, has been considered in [ZJZZ02]. In WLAN scenarios, the use of APIs placed in strategic locations, permits to carry out dynamic adaptation of transmitted data in a distributed fashion [LR05]. Also, a feasible solution in wireless networks is the use of mobile agents in the adaptation process [BBC02, BF05, BC02]. Mobile agents act as device proxies over the fixed network, negotiate the proper QoS level and dynamically tailor VoD flows depending on the terminal characteristics and user preferences. QoS adaptation can be performed on demand by the transport agents along the communication path when required by receivers [PLML06]. In mobile systems rather than rely on the system to manage resource transparently, applications must themselves adapt to prevailing network characteristics. This is the goal of the Odyssey platform [NS99]. Noble sustains that structuring an adaptive mobile system as a collaboration between the operating systems and the application can be a powerful solution [Nob00]. Recently, the streaming media industry shift from the classic streaming protocols such as RTSP, MMS, RTMS back to the HTTP-based delivery [Zam09]. The main idea behind this move was to try to adapt the media delivery to Internet and not vice-versa, adapting Internet to the new streaming protocols. The efforts finding an adaptive streaming solution for media delivery on Web were concentrated around the HTTP protocol, since the Internet was build and optimized around it. A new delivery service has been developed, named HTTP-based adaptive streaming [BFL+12] and is based on HTTP progressive download,
2.2 Adaptation Techniques 17 which is used today by popular video sites such as YouTube 5or Vimeo 6. The term progressive comes from the fact that the clients are able to play back the multimedia content while the download is still in the progress. In the HTTP-based adaptive streaming implementation, the multimedia content source (e.g., video or audio) is cut into shorts chunks, usually 2 up to 4 seconds long. For videos, the partitioning is made at Group of Pictures (GoP) boundary. Each resulted segment has no dependencies on past or future chunks. As the media chunks are downloaded the client plays back them sequentially. If there is more than one variation (source encoded at multiple bitrates) of the same media source chunk, the client can estimate the available network bandwidth and decide which chunk (smaller or bigger) to download ahead of time. Microsoft and Apple have implemented different prototypes of the HTTP-based adaptive streaming and we will look more closely on them in the following section along with a briefly introduction on OTT streaming. 2.2.2.1 Over-the-top technologies The OTT streaming is the one of the trends has emerged in the multimedia streaming market over the past several years. In the OTT technology, the multimedia service (e.g., video or audio streaming) is delivered to the customers through the Internet. This is the reason the service is also referred in literature as Internet Video. The term over-the-top stems from the fact that the service is not managed (provided) by the customer’s ISPs. The ISP may be aware of the content of the IP packets which traverse the access network by using filtering mechanisms, but is not the one who generates them. This is the case when the customers access user-generated (amateur) video content such the one hosted on YouTube or the one professionally generated by entities such as TV stations (e.g., BBC 7), news agencies (e.g., Reuters 8, Associated Press 9) or VoD (e.g., movies rental service) over the Internet through which movie studios promote their commercial offer. The OTT model opened new horizons for multimedia business and today we see that video on web sites is more a necessity than a feature. Cisco’s Visual Networking Index (VNI) predicts that over the next few years, 90 percent from the Internet traffic will be video related. Since the OTT model plays an important role in streaming the video objects over the Internet, in the following lines the most important OTT implementations are presented. Microsoft Smooth Streaming has been included in the Internet Information Service (ISS) package, since version 7 10. The smooth streaming uses MPEG-4 Part 14 (ISO/IEC 14496-12) for encoding the streams. The basic unit in MPEG-4 is called “box” and can 5www.youtube.com 6https://vimeo.com/ 7www.bbc.co.uk 8http://www.reuters.com/ 9http://www.ap.org/ 10http://www.iis.net/media
18 Related Work on Multimedia Adaptation encapsulate data and metadata. The MPEG-4 standard allows the MPEG-4 boxes to be organized in fragments and therefore, the smooth streaming is writing the files as a series of metadata/data box pairs. Each GoP is defined as a MPEG-4 fragment and is stored within a MPEG-4 file. and for each bitrate source variation is created a separate file. Therefore, each available bitrate has associated a specific file. To retrieve the MPEG-4 fragments from the server, the smooth player, called Silverlight , needs to download them in a logical sequence. While the HTTP-based adaptive streaming splits into multiples file chunks the multimedia content resource, in the smooth streaming the content (e.g., video or audio) is virtually split into small fragments. The client before asking the server for a particular fragment with a particular bitrate needs to know what fragments are available. It gets this information via a manifest metadata file when it starts the communication session with server. The manifest file encapsulates information about the available streams such as: codecs, video resolutions, available audio tracks, subtitles, etc. Once the client knows what are the resources available, it sends a HTTP request (GET) to the server for the fragments. Each fragment is downloaded separately within a request session that specifies in the header the bitrate and the fragment’s time offset. The smooth streaming server uses an indexed manifest file, which maps the available MPEG-4 files to the bitrates of the fragments to determine in which file to search. Then it reads the appropriate MPEG-4 file and based on its indexes, figures out which fragment box corresponds to the requested time offset. The chunk is then extracted and sent over the wire to client as a standalone file. The adaptive part, which switches between chunks with different bitrates, is implemented on the client site. The Silverlight application looks at fragments download times, rendered frame rates, buffer fullness and decides if there is a need of chunks with higher or lower bitrates from the server. Apple Hypertext Transfer Protocol live streaming in comparison with Smooth Streaming is using another approach when it comes to fragmentation. The Apple’s fragmentation 11 is based on the ISO/IEC 13818-1 MPEG2 Transport Stream file format, which means the segments with different bitrates are stored in different transport stream files not in only one file as in Microsoft’s implementation. The transport stream files are divided in chunks, by default with duration of 10 seconds, by a stream segmentation process. As in the Smooth streaming approach, the client is using a manifest (XML metadata) file which inherits the MP3 playlist standard format to keep record of the media chunks available on the server site. The manifest file is organized as an indexed relational data structure directing the client to different streams. The client monitors continuously the network conditions and if the bandwidth drops, checks the manifest file for the location of additional streams and 11http://tools.ietf.org/html/draftpantos-http-live-streaming
2.2 Adaptation Techniques 19 gets the next chunk of the media data (e.g., audio or video) encoded at a lower bitrate. The Figure 2.2 shows a manifest file. Manifest file 200Kbp 1Mbps #EXTM3U #EXT-X-STREAM-INF:PROGRAM-ID=1, BANDWIDTH=200000 http://localhost/myFile.m3u8 #EXT-X-STREAM-INF:PROGRAM-ID=1, BANDWIDTH=1000000 http://localhost/myFile.m3u8 ts1_1.ts ts1_2.ts ts1_3.ts ts2_1.ts ts2_2.ts #EXTM3U #EXT-X-MEDIA-SQUENCE:0 #EXT-X-TARGETDURATION:10 #EXTINF:10, http://localhost/ts1_1.ts #EXTINF:10, http://localhost/ts1_2.ts #EXTINF:10, http://localhost/ts1_3.ts #END-X-ENDLIST Figure 2.2: Apple HTTP Live streaming metainfo data example As is shown in Figure 2.2 each transport stream file in the manifest starts with EXTM3U tag which distinguishes from regular MP3 playlists. The tag EXT-X-STREAM-INF, which always precedes the EXTM3U tag, is used to indicate the corresponding bitrate of a particular media segment using the BANDWIDTH attribute. There are two different streams defined in the Figure 2.2 for the same media resources: one at 200Kbps and other at 1Mbps. Each stream is composed by a sequence of segments which are downloaded in the client buffer for a smooth playing. Adobe FLV/F4V, similar to Microsoft and Apple implementations, supports among the HTTP, streaming over the RTMP protocol for delivering live content (TV or radio broadcasts) [MSS12]. When the HTTP protocol is used, Adobe Air 12 is required as streaming server and Flash Player as the client player. In case of RTMP, Flash Media Server is used as server component and Flash Media Player as client component. The Adobe’s implementation uses multiple files, which are encoded in VP6 , H.264 , AAC and MP3 standardized formats, for generating a multi-bitrate playback. The files are “sliced” into fragments and used from a manifest file format (XML metadata), called Flash Media Manifest (F4M ). The files for live or VoD streams are placed on a HTTP server which receives fragment requests from the client and returns the appropriate fragment from the file. The Adobe 12http://www.adobe.com/ro/products/air.html
26 Related Work on Multimedia Adaptation to match those constraints. In [Gec97] such a solution is described, where the application keeps track of the user profile and customizes the content according to its profile. Other important initiatives take into account the resource negotiation (e.g., network bandwidth, processor cycles, battery life), through a specific API, as in the case of the Odyssey platform[NS99]. The work conducted in [CM03] uses a function to adapt the bandwidth fluctuations in such a way as to maximize the client multimedia consumption experience for the given environment. The main drawback of this approach is the fact that the tasks are distributed only on the client side. 2.3.3 Proxy-Based Adaptation The Proxy-based adaptation approaches use an intermediate node, placed between the server and the client, to perform content adaptation as it is shown in Figure 2.5. Server Clients 2. content request 5. streaming the content adapted 3. adapt the stream wired wired wireless Proxy 1. content request + terminal and network characteristics 3. streaming content 4. adapt the stream Internet Internet Figure 2.5: Adaptation carried out by a proxy The proxy node practically distillates the data exchanged between the server and the client. It receives the characteristics of the target terminal and asks the server (on behalf of the client) the content that the client wants to consume. The server sends the desired content to the proxy which will deal with content adaptation. In multimedia content adaptation, several proxy-based solutions were proposed [BC02, LR05]. In [BC02] Bellavista proposes a mobile agents based middleware called ubiQoS. The mobile agents used in this middleware are named SOMA (Secure and Open Mobile Agents). The ubiQoS design provides accessibility to VoD services from any connection point. There are two types of SOMA: ubiQoS proxies which provide adaptation management and ubiQoS processors which are dedicated to perform adaptation task. In [LR05], a platform called APPAT (Adaptation Proxy Platform) uses distributed adaptation over many proxies. It was designed for applications where the data is transmitted between several participants. However, this platform is application-specific and suffers in terms of scalability. 2.3.4 Service-Oriented Adaptation This architectural design (Figure 2.6) is based on Web services [WK04].
2.4 Multimedia Adaptation Frameworks and Applications 27 Adaptation Service Clients 2. content request 5. streaming the content adapted 3. adapt the stream wired wired wireless Broker 1. content request + terminal and network characteristics 4. streaming the content adapted 3. adapt the content Adaptation Service Adaptation Service Internet Internet Figure 2.6: Adaptation carried out by the adaptation service The main idea in this approach is the fact that the adaptation decision is carried out by a service distributed over a network. A special component of the architecture, called Broker, is designated to deal with identifying and selecting the most appropriate adaptation service. The MUSIC project [GRWK09] (Self-Adapting Applications for Mobile Users in Ubiquitous Computing) embraced this approach and enhanced it further by designing a framework that extends compositional adaptation by considering dynamically discovered services.The proposed framework is able to configure the application according to varying contextual conditions, taking into account user context, service properties and service level agreements of available services. This solution is quite flexible as Service discovery protocols allow advertising any new adaptation service to the Broker and insertion of new services or replacement of existing ones is easily accomplished as in a componentbased application. If a service disappears when it was under use, an adaptation process is triggered. However, adopting this service oriented solution implies, as was mentioned, the existence of an additional component (the Broker) to deal with the complexity of selecting the most appropriate adaptation service based on a given user request. 2.4 Multimedia Adaptation Frameworks and Applications This section examines some notable research frameworks and applications that address the concerns of multimedia content adaptation. We identify the main issues behind these tools to facilitate the design and the implementation of the putative architecture described in Section 5. 2.4.1 koMMA framework The koMMa framework [JLTH06] has been implemented in the course of an official ISO/IEC MPEG Core Experiment (CE) into so-called Conversions and Permissions amendment[] of MPEG-21 Digital Item Adaptation (DIA). The main idea behind this research project is that the multimedia content adaptation should be undertaken by more than one software tool for various user preferences, terminal capabilities, network characteristics
28 Related Work on Multimedia Adaptation or coding formats. The adaptation decision process is designed to automatically construct an adaptation plan for media resources that fulfill the device constraints. content description multiemdia content tool descriptions adaptation tools pool of adaptation tools pool of multimedia resources Adaptation decision taken engine Adaptation engine Adapted multiemdia content usage envinronment description Multimedia Server request response adaptation plan Client Figure 2.7: The koMMa framework architecture The overall architecture of the koMMa framework, which uses the server-side approach, can be seen in Figure 2.7 [JLTH06]. The framework comprises two major components: adaptation decision engine and adaptation engine. The former component is responsible for finding an adequate sequence of transformation steps (e.g., adaptation plan), that can be applied on a multimedia content. Then the adaptation plan is sent to the adaptation engine which applies the transformation steps on multimedia content. To find an adequate sequence it is being used the state-space planning problem [Bra90], where actions are applied on initial state to reach the goal state. The goal is to apply a set of transformation operations on a given multimedia resource such that the goal state is reached. In this way, the multimedia resource is converted into a format which comply with the user’s device characteristics. The sequence of transformation steps can be seen as sequence of adaptation services that are dealing with multimedia content adaptation. The composition of the services (actions) in such way to enable the automatic execution of the adaptation task is made it through OWL-S [MBH+04] which is being used as a knowledge representation mechanism for capturing semantics of transformations tools and algorithms. Thus, in the domain created, the initial state corresponds to an original multimedia content described under the MPEG-7 specification while the goal state corresponds to an adapted version of that media which fits the user’s needs described according to the MPEG-21 specification. The actions are operations that act directly on media content (e.g., convert the media from one format to another) and they are expressed in terms of input, output, preconditions and effects or simple IOPE . For consideration, imagine the following adaptation scenario: a picture with a given
2.4 Multimedia Adaptation Frameworks and Applications 29 resolution (640x480) must be resized to 320x240. Accordingly with the adaptation plan described above, the initial state and goal state are described in Table 2.3 as follows: initial state jpegImage(//path/to/image.yuv),width(640),height(480) goal state jpegImage(//path/to/image.yuv),horizontal(320),vertical(240) Table 2.3: An example of initial state and goal state declarations These descriptions show the spatial scaling operation on image.yuv picture applying the IOPE approach. The scaling operation is described in Table 2.4, as follows: Input: imageIn,oldWidth,oldHeight,newWidth,newHeight Output: imageOut Preconditions: jpegImage(imageIn),width(oldWidth),height(oldHeight) Effects: jpegImage(imageOut),width(newWidth),height(newHeight), horizontal(newWidth),vertical(newHeight) Table 2.4: scaling operation for image.yuv picture Given this information, the koMMA framework computes an adaptation plan that may look like (Table 2.5): 1.read(//path/to/image.yuv,outImage1) 2.spatialScale(outImage1,640,480,320,240,outImage2) 3.write(outImage2,//path/to/out put/image.yuv) Table 2.5: The adaptation plan for image.yuv spatial scaling 2.4.2 CAIN framework The CAIN framework [LM07] proposes a content adaptation approach which is based on Constraints Satisfaction Problem (CSP) [MS98]. A list of content adaptation tools (CAT) are applied in series, representing the adaptation chain for a given use case. Each CAT, may integrate different adaptation approaches such as: transcoding, scalable content, temporal summarization or include semantic driven adaptation [Giv03]. The overall adaptation process is shown in Figure 2.8. The input, at the CAIN invocation, is composed by the following information: the media resource itself, the description of the media resource in terms of MPEG-7 metadata, a MPEG-21 BitStream Syntax Description (BSD) compliant content description and the usage environment. The later is described by the MPEG-21 UED and comprises: the user characteristics, the device capabilities and the network characteristics. The CAIN framework will parse the metadata and will extract the necessary information for the Decision Module (DM). Based on this information, the DM will decide which CATs will be invoked to adapt the multimedia content.
30 Related Work on Multimedia Adaptation Media CAT capabilities description Adapted Media MPEG-7,MPEG-21 Content description MPEG-7, MPEG-21 Adapted Media description MPEG-21 Usage Envinronment Decision Module Cat1 Cat2 Cat n CAIN Figure 2.8: CAIN Adaptation Process The output of the framework will be the adapted content and the media description in terms of MPEG-7 and MPEG-21 specifications. In the adaptation decision process the DM uses the CAT capabilities description module which stores information about the adaptation capabilities of each CATs. It stores information about which kind of adaptation operations they can perform and a list of parameters that are involved in these operations such as: accepted input and output format, frame-rate, resolution, bit-rate. Figure 2.9 [LM07], represent a hypothetical adaptation process carried out by the DM. As input is given: a content description of a video resource that is available in a specific format and bit-rate; a mandatory usage environment description which describes the constraints to be imposed on the adapted media content and desired usage environment which describes the user preferences for the better video consuming experience. The DM uses the CAT capabilities descriptor to find the CAT that fulfills the constraints (at least the mandatory ones if all cannot be fulfilled). The CSP is being used by the DM in this finding. Based on the content description were defined a set of variables as follows: Foas the initial video format, Boas the initial bit-rate format, Fnas a terminal accepted format and Bnas the network maximum network bitrate:
2.4 Multimedia Adaptation Frameworks and Applications 31 Content: Format: MPEG-2 bitrate: 28000 bits/s Mandatory: Format: MPEG-4 bitrate: <=1200 bits/s Desired output: bitrate: >=700 bits/s Input: Format: MPEG-1, MPEG-2 bitrate: 15000-80000 bits/s Output: Format: JPEG bitrate: unbounded Input: Format: MPEG-2 bitrate: 10000-50000 bits/s Output: Format: MPEG-4 bitrate: 1000-1700 bits/s Input: Format: MPEG-2, MPEG-4 bitrate: unbounded Output: Format: MPEG-4 bitrate: 500-1400 bits/s CAT1 CAT2 CAT3 Adapted Content Content Description Usage Envinronment Figure 2.9: An example of a context where the DM chooses the best CAT for adaptation Fo=MPEG −2Fn=MPEG −4 Bo=28000 Bn<=1200 Bn>=700 (2.1) Based on the CATs that already exists, were defined as well as set of variables FIiand FOirepresenting the input and the output of each CAT. In the same way are defined the BIiand BOivariables which represent the input and the output range accepted by each CATi. Therefore, following the example depicted in Figure 2.9, the following domain is being defined for each variable: FI1=MPEG −1,MPEG −2FO1=JPEG BI1= [15000...80000]BO1=unbounded FI2=MPEG −2FO2=MPEG −4 BI2= [10000...50000]BO2= [1000...1700] FI3=MPEG −2,MPEG −4FO3=MPEG −4 BI3=unbounded BO3= [500...1400] (2.2)
32 Related Work on Multimedia Adaptation The variable domains from formula 2.1 which are constrained by Bnlimits are transformed as follows by the CAIN: Fo=MPEG −2Fn=MPEG −4 Bo=28000 Bn= [min...1200] Bn= [700...max](2.3) Taking into account the CAT capabilities, three rules 2.4 have been defined for the DM choosing process. The rules respect the classical pattern, comprising two parts: antecedent part and consequent part. If the antecedent part is satisfied, the consequent part is evaluated as true as well. Fo∈FI1∧Bo∈BI1∧Bn∩BO1→CAT1 Fo∈FI2∧Bo∈BI2∧Bn∩BO2→CAT2 Fo∈FI3∧Bo∈BI3∧Bn∩BO3→CAT3(2.4) In the antecedent part of the rules, the symbol ∈appears whenever a parameter takes a value and the other one takes a range; the intersection ∩is used when both parameters are sets of values. The DM applies the CPS approach on each antecedent part of the rules to find which CATs better fit. The solution is: CAT1=f alse,CAT2=true,CAT3=true and indicates that only CAT2and CAT3are suitable to be invoked by the adaptation process. The set of possible solutions represents the mandatory constraints and at this point, there is no unique solution to the adaptation problem. Applying the CAT2constraints, the output variables take the following domain: Fn=MPEG −4 Bn= [100...1200](2.5) and applying the CAT3, the following: Fn=MPEG −4 Bn= [500...1400](2.6) In order to choose the solution, the DM takes into consideration the desirable constraints shown in Figure 2.9 and apply them on the remaing set (2.5 and 2.6). The way
2.4 Multimedia Adaptation Frameworks and Applications 33 how the desirable constraints are applied differ quietly from mandatory ones. First, a priority list from the desirable constraints is being made and then, the following algorithm is being used: 1) take the first constraint and try to fulfill it; 2) if after applying this constraint, there is no feasible adaptation that fulfills the requirements, then ignore these constraints, else keep the constraint and reduce the range of the domain of the variables in formula (2.5) and (2.6) accordingly; 3) run the algorithm again with the rest of the constraints. In the context example shown in Figure 2.9, we have only one desirable constraint, which is: Bn= [700...max]. Running the algorithm aforementioned, the output variables of the CAT2reach the following domain: Fn=MPEG −4 Bn= [700...1200](2.7) and the following for the CAT3: Fn=MPEG −4 Bn= [700...1400](2.8) Finally, since more than one CAT reached the desired target, an optimization step is applied to select the CAT which will perform the adaptation. This optimization step takes into account a prioritized list proposed by the media server provider that basically consists of a list of optimization elements that refers to media characteristics such as: bit-rate, resolution. The optimization goes through the following steps: 1) apply the optimization element to each solution; 2) if remains only one solution, then select this solution and break loop; 3) run the algorithm again in the same manner with the rest of the optimization elements. Therefore, the final decision depends on the preferences of the media server provider regarding the aforementioned maximization elements. If the server provider chooses to maximize the bit-rate, the maximization element maxim(Bn)increases the values of the CAT2as: Fn=MPEG −4 Bn= [1200](2.9) and the following for the output of CAT3, :
34 Related Work on Multimedia Adaptation Fn=MPEG −4 Bn= [1400](2.10) The DM will choose the CAT3over the CAT2, and this is the final solution since the CAT adaptation tool will produce an output which has the highest bit-rate (1400). 2.4.3 Enthrone project The ENTHRONE project [tesa] proposes a solution to deliver an audio-video service over the heterogeneous networks, targeting different kind of terminals with different characteristics (e.g., PCs, PDAs). As mentioned in Section 2.2.3, the ENTHRONE adaptation decision process uses information from different layers of the ISO/OSI reference model (cross-layer information). The adaptation process is driven by device characteristics, network and user preferences. Using these characteristics and preferences, expressed as MPEG-21 Part 7 metdata [tesc], an adaptation engine performs an appropriate decision to adapt the multimedia content. As it can be seen in Figure 2.9, the ENTHRONE adaptation decision engine (ADE) comprises two modules: ADE Processor and ADTE Wrapper. The ADE Processor acts as an interface for the ADTE Wrapper. It prepares the necessary data that later on will be used by the ADTE Wrapper such as: AdaptationQoS (AQoS) [21004] and User Constraint Description (UCD) according to the MPEG-21 standard. The ADTE Wrapper integrates the ADTE tool [DMW05] (ADTE-HP) developed by the Hewlett Packard (HP) laboratories. ADTE tool ADTE Wrapper ADE Processor AQoS UEDs UED Validation Schemes UED XSLTs Decision AQoS+UCD Decision ADE Figure 2.10: The Enthrone ADE architecture The ADTE-HP software tool provides utilities for multimedia adaptation based on the MPEG-21 DIA specification that aims to standardize various metadata including those
2.4 Multimedia Adaptation Frameworks and Applications 35 supporting decision-taking and constraint specifications. This tool comprises two software modules as it is depicted in Figure 2.11: •ADTE (Adaptation Decision Taking Engine) which takes Digital Items [tesc] (DI) containing AQoS and UCD (UED) as inputs to yield decisions as outputs. •BAE (Bitstream Adaptation Engine) which takes as inputs the decisions from the ADTE to perform actually bit-stream adaptation. Figure 2.11: The HP’s ADTE model The ADTE Wrapper of the ENTHRONE’s ADE uses only the first module (ADTE engine) for the purpose of taking adaptation decisions. The ADTE Wrapper feeds up the ADTE with AQoS normative metadata (which supports decision-taken) and with UCD normative metadata (which represents explicit adaptation constraints). AQoS provides the means to express the relation between resources or constraints, adaptation operations and media quality. The combined use of these tools allows them to decide which measures must be taken and when to adapt the multimedia content. The UCD constraints may also be implicitly specified by the UED description that covers display capabilities, audio/video capabilities, terminal and network characteristics. The mechanism underlying the decision-taking process is based on a constrained optimization problem [PEGW03] involving algebraic variables that represent any combination of adaptation parameters, media characteristics, usage environment. The MPEG-21 DIA part 7 provides for the decision-taking functionality to be differentiated with respect to sequential logical segments corresponding to partitioning such as a group of pictures (GOP), frame, referred as the adaptation unit. All the adaptation units can be defined as: I[n] = i0[n],i1[n],i2[n],...iM−1[n]where n=0,1,2,... and M represents the set of variables. For each adaptation unit, the optimization problem can by seen as: Maximixe or Minimize On,j(I[n],H[n]),j=0,1,2...Jn−1 Subject to: Ln,k(I[n],H[n]) = true,k=0,1,2...Kn−1 (2.11)
42 Related Work on Multimedia Adaptation whereas Enthrone and DCAF takes the decision for a given variation of the content, selecting values for different encoding parameters (e.g., spatial and temporal resolutions, bitrate), KoMMa and CAIN selects sets of elementary adaptations that together can transform the content in order to satisfy the constraints, running an optimization process to discover the optimal set of adaptations. Still, the frameworks lack the support for the expression of semantically-richer descriptions of the context and user preferences, as well as the possibility to reason on top of that information and derive additional knowledge. Additionally, the paradigm they follow is that of adapting the content only and not the behavior of the application itself. 2.5 Summary This chapter has presented current approaches for managing multimedia delivery quality. These approaches work at different levels: transport, application or a combination of both levels (cross-layer). We have described several service models and mechanisms proposed by IETF and also some protocols (e.g. RTP/RTPC, RSVP, DS, MLPS, VDP) that operate at the network level. We also reviewed some solutions (e.g., media codecs, error protections schemes, power saving approaches, resource allocation) and several frameworks and OTT implementations (smooth streaming, Apple Hypertext Transfer Protocol live streaming and Adobe FLV/F4V) that work at the application level. Finally, we have presented the cross-layer design principles and examined how they were applied in various multimedia adaptation projects. The chapter ends with an analysis of the most representative frameworks and applications used in multimedia content adaptation areas.
Chapter 3 Ontologies in Multimedia Multimedia data is produced, processed and stored digitally. Indexing this data to make it searchable is imperative. This indexing requires the multimedia content to be annotated in order to create metadata which contains a concise and compact description of the features of the content. 3.1 Introduction The ontologies provide fundamental form for knowledge representation about the real world. In computer science, as already mentioned in 3.2, the ontologies define a set of representational primitives with which you can model a particular domain of knowledge. Currently, there is a standard that standardizes tools or ways to describe and classify the multimedia content: MPEG-7 [tesb]. However, the descriptors used by these standards are far from what users want. Consequently, the research trends in multimedia were focused in the last decade to bring as well the ontologies in this area. The high-level description of the content (semantics such as: places, actors, events, objects) that ontology provides, reduces the conceptual gap between the user and the machine. The “conceptual gap” refers to the mismatch between the low-level information that can be extracted from a given multimedia content such as visual (e.g., texture, camera motion, spatial and temporal resolution, video codecs) and audio (e.g., spectrum, audio codec) and the interpretation considered high-level information that each user makes in a given scenario on this data. For example, we can consider the following scenario where a user wants to search in a digital library, for videos that have more than one variation and a bit-rate more than 384Kbps. A such query would not be possible on a digital library which is based only on MPEG-7 constructs. Instead, a less accurate interrogation such as Give me the videos, having more than one variation and all videos with bit-rate more than 384Kbps can be constructed, but will prove false results, since will be returned all videos described by more than one variation and all that have the bit-rate more than 384Kbps. Building a 43
44 Ontologies in Multimedia knowledge-domain for videos with particular bit-rates with MPEG-7 constructs will allow us to express such queries. 3.2 Background in Ontologies The term Ontology has a long story. Its original philosophical sense refers to the subject of existents and its objective is to determine what entities and types of entities that exist in the real word, and thus to study the structure of the world. The study of ontology can be tracked back in antiquity to the work of Plato and Aristotle. The Aristotle’s ontology offers primitive categories or concepts such as substanceand quality, which was presumed to account for All That Is. In contrast, in computer and information science, ontology defines a set of representational primitive used to describe and represent an area or domain of knowledge. Ontologies are used by people, databases and applications that need to share the domain information (a domain is just a specific subject area or area of knowledge such as medicine, tool manufacturing, real estate, automobile repair, financial management, etc). Ontologies include computer-usable definition of basic concepts in a given domain and the relationships that can be established among them. The representational primitives are: classes, representing general things; relationships that can exist among the class members and; the properties (attributes) those classes may have. Ontologies are expressed using a logical-based language. In this way a meaningful and detailed distinction can be made between classes, properties and relationships. To clarify, consider the listing depicted in Figure 3.1 about the author of this dissertation. (a) Information about the author (b) A graph representation of the sentences from (a) Figure 3.1: RDF graph example Figure 3.1b is graphical representation of information from Figure 3.1a: assertions form a non-oriented graph, with subjects and objects of each statement as nodes (classes members), and predicates (relationships) as edges. There are two types of nodes: resources and literals. Literals represent data values such as numbers or strings and cannot be the subjects of the statements, only objects (e.g., Oancea). On the other side, the resources can be either subjects or objects (e.g., Daniel). The direction of the arrow points from the subject of statement to the objects of statement. This data model is used by
3.2 Background in Ontologies 45 the Semantic Web, and it is formalized in the language called the Resource Description Framework (RDF) [KC04]. Although the RDF is recognized as a language for ontologies, the RDF is rather limited: doesn’t have the ability to describe the cardinality constraints (e.g., Daniel having at least one supervisor) a feature that can be found in almost all conceptual modelling languages, or to describe a simple conjunction of classes such as Maria Teresa Andrade is supervising Daniel. Therefore, it was concluded that a more expressive language was needed, which led to the appearance of several proposals for “ontology languages” including Simple HTML Ontology Extensions (SHOE), OIL and DAML+OIL (DARPA Agent Markup Language - Ontology Inference Layer). The W3C recognized that an ontology language standard would be a prerequisite for the development of the Semantic Web and thus decided to set-up a working group to develop such a standard. The result of this activity group was the Web Ontology Language (OWL) standard 1. OWL exploits the work done on OIL and DAML+OIL and also tights the integration of those languages with RDF. Three subsets of OWL have been defined, with decreasing functionality, expressiveness and complexity: OWL-Full, OWL-DL (Description Logic) and OWLLite. OWLFull is the full, unrestricted OWL specification. OWL-DL introduces a number of restrictions on the use of the OWL-Full such as the separation of the classes and individuals with a precise scope: make the OWL-DL decidable. The OWL-Lite is basically an OWL-DL with a subset of its language elements. Implementing the full OWL specifications is not feasible because there is no algorithm capable of providing complete inference over a complex OWL-Full large knowledgebase. Basically, OWL-Full does not provide any guarantees of successfully reaching a conclusion within a bounded period of time. Instead, the OWL-DL contains subsets of the OWL-Full language that provide some more restricted expressive power in exchange for more attractive and feasible computational characteristics. The “decidability” of the OWL-DL comes from the use of description logic. The description logic holds the rules to construct valid knowledge representations that are decidable, which actually produce an answer. It derives from first-order logic. OWL ontology contains definition of classes, individuals and relationships (properties) between them. An individual can be seen as an object and a property as a binary relationship between two objects. A class is a collection of individuals. Starting from this observation, the ontologies that have been developed within the context of the RCLAF framework (the RCLAF’s ontologies), Cross-Layer-Semantic (CLS) and Application State (AS), are implemented using OWL-DL and more details about their implementation can be found in Section 5.6.1, respectively Section 5.6.2. 1www.ontology.org
46 Ontologies in Multimedia 3.3 MPEG-7 Ontologies MPEG-7 is intended to describe audiovisual information regardless of storage, coding, display, transmission, medium, or technology. As stated previously, the MPEG-7 XML Schemas, which define 1182 elements, 417 attributes and 377 complex types [TCL+07], does not provide formal grounding for semantics of its elements. To overcome this shortcoming of formal semantics in MPEG-7, several multimedia ontologies represented in OWL language have been proposed [Hun01, TPC07, GC05, ATS+07]. In the following we describe the main characteristics of these ontologies. 3.3.1 Hunter Hunter was the first initiative to build multimedia ontologies translating the existing standards such as MPEG-7 [Hun01] into RDF . Later on, the resulted ontology was translated into OWL in terms of language and harmonized using the ABC upper ontology [LH01] for applications in digital libraries [Hun02]. The ontology was build through reverse-engineering of the existing XML Schema definitions. All the classes, properties between them and semantic definitions were constructed following this approach. For simplifying this process, only a core subset of the MPEG-7 specifications was used. In the first phase, were identified the top basic entities (e.g., Image, Video, Audio, AudioVisual, Multimedia) and the relationship between them applying top-down approach and then were determined the hierarchies of classes. To show how the translation process took place, consider as example the schema definition for the complex type Person. The Listing 3.1 shows how the concept Person is defined in MPEG-7. <complexType name=”PersonType”> <complexContent> <e x t e n s i o n ba se =”mpeg7:AgentType”> <sequence> <el em en t name=”Name” t y p e =” mpeg7:PersonNameType ” /> <el em en t name=”Affiliation” minOccurs=”0” maxOccurs=” unbounded ”> <complexType> <c h o i c e> <el em en t name=”Organization” t y p e =”mpeg7:OrganizationType”/> <el em en t name=”PersonGroup” t y p e =”mpeg7:PersonGroupType” /> </ c h o i c e> </ complexType> <el em en t name=” Address ” t y p e =”mpeg7:PlaceType”/> </ sequence> </ e x t e n s i o n> </ complexContent> </ complexType> <complexType name=”PersonNameType”>
3.3 MPEG-7 Ontologies 47 <sequence> <c h o i c e minOccurs=”1” maxOccurs=” unbounded ”> <el em en t name=”GivenName” t y p e =” s t r i n g ” /> <el em en t name=” FamilyName ” t y p e =” s t r i n g ” /> </ c h o i c e> </ sequence> </ complexType> Listing 3.1: The definition of the PersonDS in MPEG-7 Schema According with the schema definition, the PersonType is an AgentType and its defined as a sequence of PersonNameType and Affiliation. To be able to say that “the person is an agent” in RDF, the Affiliation was defined as a class and PersonType as a subclass of it. The sequence of elements that defines the PersonType in RDF were translated as a list of data properties. Therefore, each element from this sequence was mapped to a data property as follows: the element Name mapped to data property name and Affiliation to affiliation. The Figure 3.2 shows how the children elements of the PersonType were translated into RDF concepts. Figure 3.2: Hunter’s MPEG-7 ontology definition for Person The Affiliation as its defined in Listing 3.1 can have values which are instantiations either of the OrganizationType or the PersonGroupType. The RDF provides a way to define multiple possible ranges by using the unionOf. The Listing 3.2 shows how the unionOf axiom restriction was used to map the choice definition from the PersonDS. <r d f s : C l a s s r d f : I D =”Affiliation”> <rd fs:co mm ent>E i t h e r an O r g a n i z a t i o n o r a PersonGroup</ rdf s:com me nt> <daml:u nion Of r d f : p a r s e T y p e =” d a m l : c o l l e c t i o n ”> <r d f s : C l a s s r d f : a b o u t =”#Organization”/> <r d f s : C l a s s r d f : a b o u t =”#PersonGroup”/> </ daml:unionOf>
48 Ontologies in Multimedia </ rdfs:Class> <r d f : P r o p e r t y r d f : I D =”affiliation”> <rdfs:label>affiliation</ rdfs:label> <r d f s : d o m ain r d f : r e s o u r c e =” # Pe rso n ” /> <r d f s : r a n g e r d f : r e s o u r c e =”#Affiliation”/> </ rdf:Property> Listing 3.2: The definition of the Affiliation class and affiliation property in Hunter’s MPEG-7 ontology The current version of this ontology is an OWL full language 2. Among the aforementioned top-basic entities, were defined also descriptors for storing information about production, creation and usage. This ontology usually has been used in decomposition of images. Having the foundation on ABC ontology, enables queries for abstract concepts such as subclasses of events or agents, which constitute what MPEG-7 names “narrative world” [tesb], to retrieve media objects or segments of media objects. 3.3.2 DS-MIRF The DS-MIRF framework [TPC07] aims to facilitate development of knowledge-based multimedia applications utilizing the MPEG-7 and MPEG-21. The ontology proposed within DS-MIRF framework, provides a formal semantic in OWL-DL language to MPEG7 description and Classification Schemas . The DS-MIRF’s ontology has been designed manually. The XML complex-type constructs are mapped into ontology as OWL classes which represent group of individuals. This individuals are interconnected sharing the same properties. The XML simple datatypes are defined in separate XML schema and are imported in the DS-MIRF ontology. The XML elements are kept in the rdf:IDs of the corresponding OWL classes. For exemplification, an excerpt from MPEG-7 description schema showing the AgentType complex data type has been taken (see the Listing 3.3). <complexType name=”AgentType” abstract=” t r u e ”> <complexContent> <e x t e n s i o n ba se =”mpeg7:DSType”> <sequence> <el em en t name=” Ic on ” t y p e =”mpeg7:MediaLocatorType” minOccurs=”0” maxOccurs=” unbounded ” /> </ sequence> < / e x t e n s i o n> </ complexContent> </ complexType> Listing 3.3: The AgentType definition in MPEG-7 MDS 2http://metadata.net/mpeg7
3.3 MPEG-7 Ontologies 49 Therefore, this ontology will capture the AgentType XML complex datatype as an OWL class in OWL-DL language as shown in Listing 3.4. The Icon element of the MPEG-7 type MediaLocatorType is represented by the Icon OWL object property and relates class instances from domain AgentType to range MediaLocatorType. <o w l : C l a s s r d f : I D =”AgentType”> <r d f s : s u b C l a s s O f r d f : r e s o u r c e =”#DSType” /> </ o w l : C l a s s> <o w l : O b j e c t P r o p e r t y r d f : I D =” Icon ”> <r d f s : d o m ain r d f : r e s o u r c e =”#AgentType” /> <r d f s : r a n g e r d f : r e s o u r c e =” # MediaLocatorType ” /> </ owl:ObjectProperty> Listing 3.4: The AgentType OWL class in DS-MIRF’s ontology Also, several XML constructs such as sequence element order or the default values for attributes are captured in this ontology. Therefore, this ontology makes possible returning to an original MPEG-7 description from a RDF data. The generalization of this approach, which is closes to Hunter’s one, led to development of a model for capturing the semantics of any XML Schema in OWL-DL language [TC07]. This ontology has been used in OWL domain ontologies such as Soccer and Formula 1 [TPC07] to demonstrate how the knowledge can be integrated in the general-purpose constructs of MPEG-7. 3.3.3 Rhizomik The Rhizomik [GC05] aims that fully capture the semantics of the MPEG-7 standard, and it is thus considered the most complete one with respect to the ones mentioned in this section. The designed ontology covers all the MPEG-7 standard, CS and TV Anytime 3. A tool, called XSD2OWL, in this scope has been designed into the ReDeFer 4project. This tool captures a good part of the semantics of the XML schemas. The XML construct names are maintained into the resulted ontology. The definitions of the XML elements and datatypes have been translated using the approach detailed in [GC05]. The resulted MPEG-7 ontology, when the XSD2OWL tool was applied on MPEG-7 schemas, comprises 2372 classes and 975 properties and is expressed in OWL-Full language. Moving her to OWL-DL language required manually adjustment over 23 properties (rdf:Property), because them have both data and object type range. This anomaly occurred because correspondent XML elements are both defined as containers of complex and simple types. The manually adjustment implied utilization of two distinct OWL properties (e.g., owl.DataProperty and owl:ObjectProperty). 3http://www.tv-anytime.org 4http://rhizomik.net/redefer
50 Ontologies in Multimedia The Rhizomik automatized the translation process of XSD metadata into OWL compliant knowledge. This approach has been used in conjunction with other XML schemas in the Digital Right Management(RDS) domain, such as MPEG-21 [GGD07]. 3.3.4 COMM Compared with the approaches described above in which efforts were channeled in the direction of how to “translate” MPEG-7 standard to RDF/OWL, the COMM (Core Ontology of Multimedia) re-designed completely MPEG-7 according to the intended semantics, by taking into account the DOLCE 5as foundation ontology and two design patterns: one for contextualization called Description and Situation (D&S) and the another one for information objects called Ontology for Information Objects (OIO). The COMM ontology is expressed using the OWL-DL language in OWL terms and provides MPEG-7 standard compliance, semantic and syntactic interoperability, separation of concerns, modularity and extensibility. The COMM aims is to enable and facilitate multimedia annotation. COMM covers the most important parts of the MPEG-7 standard. However, some parts of the MPEG-7 standard have not yet been considered (e.g., navigation and access) and can be formalized analogously by using other multimedia patterns: Decomposition Pattern,Content Annotation Pattern,Media Annotation Pattern and the Semantic Annotation Pattern. 3.3.5 Analysis In the previously section, we have described several MPEG-7 multimedia ontologies that are relevant for this work. Table 3.1 gives a general overview of the state-of-art in these ontologies. They are focus on the semantics of the MPEG-7 standard and do not provide direct support for multimedia content adaptation. The RCLAF framework makes use of ontologies not only to capture information in media domain, but also in application domain in order to be able to provide support for adaptation. The semantics of the contextual information (e.g., terminal user characteristics, network environment) in RCLAF help the framework to understand all these participating entities, how they are interacting with each other and make important decisions related to adaptation. Therefore, the media semantics captured in media domain ontology should be linked with other domains that capture information about characteristics of the terminal, network and the application. The media resource is one of the main actors which participate in RCLAF’s multimedia consumption scenario and should be captured into domain ontology. We have chosen to represent the media resource following the MPEG-7 standard because increased over the Internet in the last years the content annotated in this format. The ability of domain 5http://www.loa.istc.cnr.it/DOLCE.html
3.4 Ontology-based Multimedia Content Adaptation 51 ontology to link with other domains was as well, an important criteria to design the media resource domain of the RCLAF’s knowledge-model described in Section 5.6. Investigating the ontologies presented previously, we have seen that they provide ways to integrate with external domains at certain level. Hunter’s and COMM ontologies use as foundation the ABC respective the DOLCE ontologies that provide basics to relate them with other domains. Hunter’s ontology uses either semantics from MPEG-7 (e.g., depicts) or defines external properties that use an MPEG-7 class (e.g., mpeg7:Multimedia). The COMM ontology uses a specific pattern called Semantic Annotation Pattern that puts under the dolce:Particular or owl:Thing class any external domain. Sub-classing one of the MPEG-7 SemanticBaseType such as: places, events or agents, provides integration capabilities to DS-MIRF ontology. However, the RCLAF’s multimedia domain ontology only needs to capture information that is useful for RCLAF framework to take adaptation decisions, such as: codec, bitrate, resolution (for video contents). This lowlevel data is encapsulated into the media profile descriptor of the MPEG-7 MDS. Although the COMM proposed a new vision to build multimedia ontologies, by the use of an upper ontology (DOLCE) to provide extensibility with respect to the multimedia vocabulary, is used mainly for annotation and do not provides semantics for describing aforementioned descriptor. In fact, Hunter’s and Rhizomik are the only ontologies that provide at certain level formal semantics for this descriptor. In conclusion, there is no ontology, at least from the best of our knowledge able to describe the multimedia content with information we need and in the same time links with domains that describes network and terminal characteristics. Thus, one of the tasks of this work is to take benefits of the appropriate multimedia ontologies to develop a media resource domain ontology, called MPEG7Prot4 that covers this gap. 3.4 Ontology-based Multimedia Content Adaptation The utilization of the ontologies to support adaptation is not a new idea. There are several ontologies-based models used to capture concepts and relationship between entities in domain of multimedia content adaptation. This section reviews the ones that are representative for this work and points out how the RCLAF’s ontology differ from them. 3.4.1 MULTICAO The MULTICAO [BA09] proposes an ontology-driven content adaptation engine. This engine is a sub-system module of the VISNET II platform 6. It uses ontologies to share 6http://www2.inescporto.pt/utm-en/projects/projects/visnet-ii-en/
58 Ontologies in Multimedia
Chapter 4 Reflection in Multimedia In this chapter we present the fundamental techniques of reflection and reflective framework, and examine several works that have applied reflection to the domain of multimedia computing framework. It starts by introducing a list of requirements for middleware adaptation. Then continues examine closely the paradigms for adaptation which offer powerful techniques for performing system-wide dynamic adaptation for framework implementations. The chapter concludes with an analysis about how the requirements for middleware adaptation are met by several reflective implementations. 4.1 Introduction Multimedia delivery chain involves different entities such as: service providers, network providers and user’s profiles. Together, these entities can compose complex scenarios with a variety of different terminals with different characteristics for playing a specific media. These devices may in turn use different network connections (e.g., ADSL, Cable, Wi-Fi etc) with particular characteristics over which users can consume multimedia content. Moreover, as the mobile users moves the network conditions change. Hence such a systems must be able to adapt multimedia content according with user’s environments, thus reacting to changes that may occur (e.g., traffic congestion, switching between terminals) during the multimedia delivery. This problem raises the need for adaptation in multimedia consumption scenarios and new adaptation models which brings flexibility and adaptability must be developed. 4.2 Requirements for middleware adaptation In this section we describe the requirements that a middleware need to fulfil in order to support adaptation in multimedia applications. First we discuss the requirements for a modelling language that is used to implement adaptive middleware. Then, we presented 59
60 Reflection in Multimedia requirements for a middleware as the execution environment for adaptive multimedia applications. The language requirements are: •Expressiveness — The language should have enough expressivity in order to provide support for configuration and reconfiguration of the multimedia middleware. •Modularity — An adaptive architecture should be modularized, meaning that the architecture should be composed by several sub-models rather than be a monolithic model to cover all concepts for adaptation decision measures. •Tool Support — Should provide facilities for the design and analysis of a system. Furthermore, tool(s) should also provide support for code generation. •Ease of Use — the design is made by the software engineers, thus is important to consider if a language is easy to be used. This requirement influences the modelling paradigm, the notation and the tool support. •Separation of Concerns — Fulfilling this requirement helps the programmers to tackle different design and development issues in a clearer manner and therefore, the overall complexity of the system is diminished. The middleware requirements are: •Consistency — The dynamic changes into middleware platform may lead to inconsistent states. To overcome these situations, mechanisms for consistency check should be introduced to preserve the consistency of the running system. For example when a middleware reacts to external stimuli (e.g., fluctuation of the network bandwidth), the adaptive decision measures should be performed atomically and correctly to guarantee that the application will end-up in a full and coherent functional state. •Performance — The middleware platforms should have a good and predictable performance. The adaptation should not compromise the middleware’s efficiency. Furthermore, the middleware platforms require flexibility in order to allow developers to choose between more control or better performance. •Flexibility — The adaptive middleware should be flexible enough in order to allow the adaptation to be added, removed or modified at runtime. This requirement implies that fact that all conceivable adaptation scenarios cloud not be foreseen during the application development stage. •Extensibility — The middleware platforms must be developed in such manner in order to support the inclusion of the new strategies, constraints or technologies. As the
4.3 Paradigms for Adaptation 61 initial platform requirements change, evolution of the design should be supported. In the case of multimedia applications, the designed middleware should support evolution of media types, media interaction types etc. •Open distributed system benefits — The multimedia middleware should solve problems related with heterogeneity and distribution. In consequence, properties, such as openness, portability and interoperability must be preserved. The approach to open the structure of the architecture and implement adaptation should follow separation of concerns by providing a principled way to tackle the problem independently of the operating system or middleware. The design principles used by the software engineers should be applied to ensure that the application is portable (platform and language independent). The interoperability between devices should be achieved by using a standardized set of protocols, technologies and data formats. The above requirements will be used later on, more specifically in Section 4.5.4 to evaluate and compare several architectures which are representative for this work. The next section introduces the main paradigms used to support adaptive applications while the following sections survey how these paradigms are used into adaptive architectures. 4.3 Paradigms for Adaptation A number of programming paradigms [SSRB00] have contributed to create the architectural shape of the adaptive middleware frameworks. In addition to object-oriented paradigm, four paradigms play important roles in supporting adaptive applications: componentbased design,aspect-oriented paradigm,computational reflection and software design patterns [SSRB00]. Each of them are described in the following lines. Component based design (CBS) advocate to reuse your pre-fabricated software components that may be combined with other vendor’s components (third party libraries) to implement software applications. A software component can be seen as a software unit that can be independently produced and deployed. The components specifies clearly what they require and what they provide. The independent development of the software components enables the late composition (known also as late binding) which plays an important key role for adaptive systems. The late composition provides a way for coupling compatible software components at the run-time through their interfaces. The CBS facilitate the creation of adaptive systems customized to specific application domains. It enables creation of software components or reuse of the one that already exists and creation of replaceable units already tested and bug free. This approach reduces the software-development costs and at the same time, elevate interoperability between enterprise systems. Currently, are three major players in the market of enterprise software solutions:
62 Reflection in Multimedia •Sun Microsystems (actual Oracle) proposed Enterprise Java Bean(EJB) [Oraa] which is a middleware component model that enables Java developers to use of-the-self Java components or beans. The EJB component model supports adaptation by automatically supporting services such as transactions or security for distributed applications. •Object Management Group (OMG) proposes CORBA Component Model(CCM) [Cob00] that can be considered as a cross-platform, cross-language super set of EJB. The CCM supports adaptation by enabling injection of adaptive code into component containers (e.g., the component themselves remain intact). •Microsoft Corporation introduced a proprietary solution DCOM [Cor] based on COM+ [Cor] server application infrastructure. This technology has been deprecated in favor of Microsoft .NET framework [Cor]. The Aspect-oriented paradigm advocates that the complex software applications are composed of different intervened cross-cutting concerns [Ste06]. The cross-cutting concerns are properties or areas of interest such as QoS, security, fault tolerance etc. In Object-Oriented Programming the things are abstracted among classes in an inheritance tree. In Aspect-Oriented Programming (AOP) , the cross-cutting concerns are scattered among different classes and this paradigm enables separation of them during the development process. Dur ring the compilation process or at the run-time an aspect can be used to weave different aspects of the program together to add a new behavior to the application. Customized middleware versions based on AOP can be generated for specific application domains. Yang [YCS+02] and David [cDLBs01] proposes a two-step approach to dynamically weave the aspects: during the compilation uses a static AOP weaver and during the run-time uses reflection. However, since the cross-cutting concerns are scattered all over the classes in AOP complicates the development and maintenance of the applications. The computational reflection is a particular case of the open implementation paradigm and it was originated early by the Brian Cantwell Smith [Smi84]. Reflection is the capability of a system to reason about itself and act upon this information. This is known as the CCSR (Causally Connected Self Representation [Smi84]). The self-representation gives the system ability to answer questions about itself and perform actions upon itself. There are several benefits introduced by the casual connection. Firstly, the self-representation provides an accurate representation of the system and secondly, the system is able to perform both self-introspection and self-adaptation. Basically, a reflective system comprises two levels: base-level and meta-level. The base-level deals with business logic of the application while the meta-level deals with the system’s representation. The changes made are made at the meta-level via this selfrepresentation are reflected in the underlying base-level, and vice versa. The process of
4.3 Paradigms for Adaptation 63 making the base-level accessible at the meta-level is known as reification. Operations that involve introspection and to make changes to the meta-level are called Meta-Object Protocol (MOP). Software design patterns provide a way to reuse best designs practiced over the years. The goal of software design patterns is to offer communication insight and experience about common problems that software developers face with during the time and their know “solution”. The paradigms introduced in this section address only part of the adaptation techniques used by numerous recent and ongoing adaptive middleware projects. In multimedia, in order to manage the context and environment changes, the adaptation model must be able to adapt itself based upon reasoning on current environment conditions and its current behavior. We believe that reflection represent a suitable solution to development of such adaptive systems. The reflection provides mechanisms to introspect the structure and behavior of the application frameworks. Reflection has been predominantly applied to language design. Nowadays, a wide variety of reflective languages are available such as: Sun’s core Java Reflection library [Orab] or OpenC++ [Chi95]. In this section we concentrate on the application of reflection to the design of framework-based multimedia adaptation systems. We focus in particular on the general techniques involved, which are common to architecture solutions described in section 4.5. 4.3.1 The need of Reflection A fist approach in direction of flexibility and adaptability was to open up the system implementation, applying the open implementation paradigm [GK96]. This paradigm comes in contradiction with software engineering paradigm in which the implementation details are hidden from the user. The resulted peace of software, that hides implementation details from the user, can be see it as a black box. This approach undoubtedly brings some benefits to users such as better understanding of the software functionalities, easy of use but fails to enhance the level of portability and feasibility due to the fact that is not possible to access and modify the software internals. Hiding implementation details makes the systems impossible to change when adaptation is need it. The open implementation paradigm is centered around two interfaces. The first interface is used for accessing the basic module’s functionality whereas the latter provides a way to access and change the module’s internals. The computational reflection is a particular case of the open implementation paradigm and it was originated early by the Brian Cantwell Smith [Smi84].
64 Reflection in Multimedia 4.3.2 Reflection types There are two main types of reflection [KCBC02]: structural and behavioral. The structural reflection is concerned with the ability of a system to inspect and modify (e.g., modify the structure of an object in such way to add new behavior at the run-time) the system’s internal architecture. This type of reflection focuses on how the system is constructed. As examples of structural reflection we can mention the set of the operations the system supports, the abstract data types within the system. The context or the meta-data also can be see it as a form of structural reflection: provides additional information about the underlying system (e.g., location, current battery level, network and device characteristics). The behavioral reflection is concerned with activity in underlying system. In particular, is dealing with arrival and dispatching when an invocation of an operation take place. Typical mechanisms provided include the use of interceptors that support the reification of the process of invocation. The behavioral reflection allows for reification of that mechanisms that handle the execution of the system. Thus, this type of reflection is concerned how the execution of the system is taken. In addition to this types of reflection, there are two styles of reflection: procedural and declarative [Smi84]. The former represents the system using a program written in the same language as the system while the latter provides a set of statements to represent the system. The declarative style offers a better high-level representation of the system comparatively with the procedural style. However, the declarative reflection needs a mechanism to interpret the statements. Consequently, the casual connection (relationship between the user’s action and the implementation of the system) is more difficult to realize (e.g., such mechanism should guarantee that the representation is the same with the actual state of the system). In procedural reflection, the casual connection is easily achieved since the representation of the system is part of the representation itself. However, the combination of both style in designing the reflective system is possible. 4.4 Reflective Architectures for Multimedia Adaptation As previously mentioned in Section 2.1, the role of adaptation is to modify the behavior of the application after the application is developed in response to changes that may occur in functional requirements or operation conditions. Depending on when the application enables the adaptation, the reflective architectures could be categorized as static reflective architectures and dynamic reflective architectures. The former triggers the adaptation during the compilation or start up while the latter enables the adaptation during the run time. The MetaSockets, described in details in Section 4.5.2.4, loads adaptive code during run-time to adapt to wireless network loss rate changes.
4.5 QoS Enabled Middleware Frameworks 65 Another categorization of the architectures that apply principles of reflection to achieve adaptation, takes into account the application domain. Based on application domain we can divide the architectures in: QoS oriented,critical and lightweight middleware. The QoS oriented middleware is used mainly in real-time or multimedia applications that are required to meet deadline and adhere to QoS contracts such as: video conferencing, Internet telephony, Internet television, Internet VoD etc. The critical provides support for distributed applications that need to be correctly operational. In this category we meet military applications for command and control or applications designed for medical area. The lightweight middleware is designed for those applications that need a small footprint to run on limited resources devices such as set-top-boxes, smart phones or industrial controllers. The critical and lightweight middleware are beyond the scope of this dissertation. In the reminder of this section, we present the major middleware architectures designed for QoS provisioning. They helped to contour the shape of the RCLAF framework described in the following chapters. 4.5 QoS Enabled Middleware Frameworks QoS-oriented middleware supports distributed applications that require quality-of-service. The Sadjadi [SM03] classifies the QoS-oriented middleware into: real-time,stream oriented,reflection-oriented and aspect oriented middleware. Since the aspect-oriented middleware do not use computational reflection to achieve the adaptation, we will focus only on the first three categories. Should be mention here that although all the examples presented in the following sections employ computational reflection, due to the fact that some of them have primary focus on other application domains (e.g., real-time or streamoriented) we do not classify them into the reflection-oriented category. 4.5.1 Real-Time Oriented The real-time middleware needs the meet the deadlines (e.g. operational deadlines from event to system response) that are defined within the real-time applications. The real-time middleware shall guarantee the response within strict time constraints. An example of real-time system where its application within the context can be considered as critical is military distributed control systems where a failure may lead to loss of life. 4.5.1.1 DynamicTAO The dynamicTAO [KRL+00] is a reflective CORBA Object Request Broker (ORB) supporting run-time distributed reconfiguration. It was developed as a part of the K2 project [Cam99] (University of Illinois) which aimed to develop a distributed operating system based on
66 Reflection in Multimedia adaptive architecture. The dynamicTAO was developed as extension of the TAO middleware platform [SC98]. The TAO respects the CORBA standard and encapsulates explicit information about the ORB internal engine. In particular, the TAO uses a configuration file that keeps the strategies that the ORB uses to implement strategies such as: scheduling, concurrency, security and monitoring. At the run-time, the selected strategies are loaded by the ORB engine. This dynamic customization of the ORB is achieved through a collection of entities known as: component configurators (e.g., DomainConfigurator, TAOConfigurators). The role of these configurators is to maintain dependencies between a component and other system components. For example, the DomainConfigurator keeps references to instances of ORB while the TAOConfigurators have the responsibility to attach or detach components implementing the aforementioned strategies. The MOP of dynamicTAO presents several features. First, is capable to transfer components across the distributed system. If there is a case when a particular component is not available on the local system, that component can be fetched from remote repositories. Secondly, the MOP loads or unloads modules that encapsulate different ORB behaviors (strategies). This allows strategy to be added or removed from the middleware. Finally, the ORB configuration can be inspected and modified dynamically to provide dynamic adaptation of the internal ORB engine. The dynamicTAO system does not make use of meta-objects for the purpose of reifying aspects of the ORB. Rather, component configurators represent the meta-level entities, which provide facilities for the inspection and reconfiguration of the ORB. Interestingly, the 2K platform offers a framework for hierarchical resource management. Although the framework provides means for admission control and reservation of resources, little support is offered for the dynamic reconfiguration of resources. 4.5.1.2 Orbix The Orbix [Cor09] designed by the IONA Technologies (now part of Progress Software) is a CORBA ORB compliance middleware. Beyond the implementation of the standard it also provides enterprise-class (software that provides high speed and high reliability) that resides at the core of distributed systems such as: billing systems, multimedia news delivery, airport runway illumination system, telephones systems etc. Orbix/E provides developers with a solution for simplified application modeling and development, and the power to create and deploy robust mobile and embedded computing applications quickly and easily. Orbix/E is a customizable middleware allowing to developers to generate customized versions of it, and it is configurable because of its ability to parse configuration files during the application start-up time, for example, to load optional plug-gable protocols.
4.5 QoS Enabled Middleware Frameworks 67 4.5.2 Stream Oriented The stream-oriented middleware provides to the multimedia application developers a continuous data streaming abstraction. Blair has conducted several research works about the role of computational reflection in middleware. He started within the ADAPT project [BCD+97] which aims to apply reflection to design multimedia application which can be dynamically adapted in response to environment conditions. He continued this work in OpenORB project [BCRP98] and combined the reflection with component-based design in OpenORB v2 [BCA+01]. The MetaSockets [SMK03a] is dealing with adaptation of the multimedia streams. In the following lines these projects are described. 4.5.2.1 ADAPT framework The ADAPT framework [BCD+97] is the result of the work carried out within the ADAPT project and proposes an adaptive framework platform for mobile multimedia applications. The main idea behind the proposed framework is a communication abstraction which captures the concept of information flow. In this scope were been defined two concepts: stream interfaces and explicit bindings. The stream interfaces are defined in terms of flows. The flow it is seen in this work as a point of consumption or production of a continuous media type. Each flow encapsulates information about its name, the multimedia type it handles (e.g., MPEG−4 video codec) and the direction (e.g., in or out). Out is for producer while in is for consumer. The ADAPT framework takes the concept of binding and extends it to a further concept, named open binding. The binding is the act through which is established a logical association between two objects that intend to interact. Unlike CORBA, where the binding is implicit (the logical association it is made by the CORBA infrastructure and is not visible to programmers) in ADAPT framework is explicit. This means that the bindings are created, managed and invoked as any other objects by the programmers. There are two types of bindings defined:operational bindings and stream bindings. The binding model assumes that two objects are connected through another object called binding object. The Figure 4.1 illustrates this model. The interfaces on the binding object are connected to interfaces of the objects. This coupling is called local binding. The idea of explicit binding, is to provides support for QoS in terms of monitoring and control. For example, a binding can specify the QoS for a multimedia streaming flow in terms of network parameters such as jitter, delay, throughput, latency and packet loss. The open binding concept introduced by the ADAPT framework promotes the idea that it is important to have access to the internals of binding objects. Through this approach, the programmers may introspect and modify the behavior of a binding in a principled manner. An open binding comprises a set of objects which are connected together through the local binding concept explained above. The binding objects that compose
74 Reflection in Multimedia presented in previously section are language dependent and may influence their efficiency and portability (e.g., FlexyNet is written in Java language while the ADAPT, OpenORB and OpenORB v2 are implemented in C/C++). The granularity of reflection in the midllewares presented here ranges from systems consist of fewer, larger components (cf. coarse-grained) to the ones that are made up of several small components (cf. fine-grained). For example, the FlexiNet offers an objectbased approach for reflection whereas dynamicTAO only allows ORB-wide reflection. The tool support is providing meta-information and code generators to facilitate the software engineering process. None of the works presented provides a full integrated tool support (e.g., with UML editors) to hep development process. Middleware Language Requirements Expressiveness Modularity Tool Support Easy of Use Separation of concerns DynamicTao +/−−+/−+/− Orbix + + +/−+/− ADAPT + + +/− OpenORB + + +/− OpenORB v2 + + +/− MetaSockets + + +/−+/− FlexiNet −+ +/−+ OpenCORBA +/−+ +/−+/− Table 4.1: Adaptive middleware requirements in terms of language requirements Table 4.2 illustrates how the middlewares are categorized with respect to requirements that they need to fulfill to achieve adaptation. We have observed that the coarse-grained approaches are generally focused to provide high performance to applications, whereas the adaptive systems that are fine-grained oriented are more concerned with the range of adaptation possibilities. Regarding the open distributed system benefits, the FlexiNet and OpenCORBA are the approachs that best meet this requirement. Consequently, extensibility is also supported by these approaches. Another thing that is worth to be mentioned is the fact that some works provides consistency. The dynamicTAO is the only approach which offers support in this concern. There are two approaches which best cope with resource management. These are FlexiNet and dynamicTAO. In the FlexiNet abstractions are defined for resource, resource
4.5 QoS Enabled Middleware Frameworks 75 Middleware Middleware Requirements Consistency Performance Flexibility Extensibility Open distributed system benefits DynamicTao + + +/−−+/− Orbix −+− − +/− ADAPT +/−+/−+ + OpenORB +/−+ + + +/− OpenORB v2 +/−+ + + +/− MetaSockets −+/− − − +/− FlexiNet +/− − +/−+ + OpenCORBA −+/−+ + Table 4.2: Adaptive middleware requirements pools and resource managers and therefore any king of resource may be modeled. In dynamicTAO, the resource management is tackled by a load-balancing approach. The resources are reserved for instances within the ORB, but there is little support for dynamically changing the resources allocated to these ORB instances. The strategy to detect when a task is experiencing a lack of resources for these frameworks imply changing sample rate, introducing filters and do not take into account resource reconfiguration. The Table 4.3 sketches the paradigms employed by each middleware to achieve adaptation. The table shows that computational reflection and component-based design have been relatively more studied in the adaptive middleware research than aspect-oriented programming and software design patterns. This approach provides flexibility for the dynamic replacement of components and promotes software reusability. In Table 4.4 categorizes the adaptive middleware projects using the adaptation type. This table illustrates that researh work conducted in adpative middleware solutions has exploited both static and dynamic adaptations. As it can be seen the trend is toward dynamic adaptation.
76 Reflection in Multimedia Adaptive Middleware Ref CBD AOP SDP QoS Middleware Real Time DynamiCTAO x x Orbix x x Stream Oriented ADAPT x OpenORB x x OpenORB v2 x x MetaSockets x x x x Reflection Oriented FlexiNet x x OpenCORBA x Table 4.3: Paradigms used over adaptive architectures 4.6 Summary In this chapter has been introduced the concept of middleware. Then were introduced the paradigms that used to build adaptive applications. Several reflective architectures considered representative in building the shape of the RCLAF architecture are presented. It has been shown that these architectures offers ad-hoc solution to obtain the adaptation that multimedia application require. Reflection has been presented as a way to achieve adaptation. However, it has been argued that minimum support for multimedia application is provided in terms of stream communication and resource management. The next chapter introduces the overal design of a reflective framework prototype (called RCLAF) proposed in this thesis for tackling the multimedia content adaptation.
4.6 Summary 77 Adaptive Type Static Dynamic Orbix DynamicTAO ADAPT OpenORB OpenORB v2 MetaScockets FelxiNet OpenCORBA Table 4.4: Adaptive middleware categorized by adaptation type
78 Reflection in Multimedia
Chapter 5 Architecture of RCLAF This chapter presents the overall design of a conceptual framework for adapting multimedia content in a distributed environment. The RCLAF is an attempt to implement adaptation support for multimedia delivery over heterogeneous networks in a reflective and inspired manner. This chapter is intended to provide a complete description of the general design of the RCLAF framework, and provide an understanding of the fundamental concepts and approach. It has been developed based on the requirements identified in Section 5.2 and relies on the concept of reflective middleware. The organization of this chapter is as follows. In Section 5.2, it starts with the challenges that have led the author to design and implement the RCLAF middleware, then continues with the design approach for the implementation of the middleware platform, next shows how the ontologies are to be used in adaptation decisions and to maintain the state of the application, and finally ends with a short recapitulation. 5.1 Introduction Recently, frameworks have gained popularity among developers. Generally, a framework is a piece of software which automates tasks and aligns with an elegant architectural solution. It describes the interfaces implemented by the framework components and the interaction between these components (e.g., flow control). The advantage of using frameworks is the fact that when the framework classes are inherent, the design patterns are automatically inherent as well, which leads to the rapid creation and deployment of software solutions. In addition, frameworks contains features that are commonly needed for the development of enterprise applications. Having these features already embedded will save us all a lot of time when we start writing the implementation code. 79
80 Architecture of RCLAF 5.2 Challenges Multimedia applications for tackling the problems identified in Section 1.1 require flexibility from the underlying system. Current frameworks do not provide flexibility, because the developers do not need to know any details about the object from which the service is being required. As a matter of fact they do not get to know the details about the system platform that supports the interaction. Thus, the desired flexibility cannot be ensured. This black-box 1approach is not adequate for developing flexible applications. Hence, it is necessary to design a distributed architecture that will the expose internal structure and allow modifications. Taking these observations into consideration, the proposed framework should provide two basic capabilities: configuration and reconfiguration. The configuration will be used to select specific operations at initialization time, while the reconfiguration changes the behavior of the framework at run-time. For example, considering the scenarios described in Section 1.1, the configuration can be applied to establish the communication according to the channel between the user’s terminals and the streaming video server. Once the communication is established and the user starts to consume the desired multimedia content, the reconfiguration capability will allow us, through a specific component (adapter), to change the video stream characteristics (e.g., reduce the bit rate, resolution, etc.) in order to reduce the video quality and to be able to adapt to, e.g., network congestion. The reconfiguration can also be applied when the user switches between terminals. An adapter will be introduced in order to adapt the stream according to the new terminal characteristics. The adaptation can be undertaken by the operating system or at the application level. However, there are some problems with these methods. Firstly, the adaptations made by the operating system is platform dependent and requires a deep understanding of the internals of the operating system. Moreover, these days the trend is to let as much as possible of the flexibility be in the applications, in order to satisfy a large variety of requirements. Secondly, the adaptations developed at the application-level cannot be reused since they are application specific. The above observations led us to emphasize that adaptation should be carried out by a reflective architecture that ensures at the same time platform independence and isolation from the implementation details. We can go further and assert that traditional frameworks are not suitable for designing flexible applications, and therefore an open implementation of frameworks could be an appropriate way to overcome this problem. The reflective architecture which this dissertation proposes must cope with a list of 1A black box is a device, system, or object which can be viewed solely in terms of its input, output, and transfer characteristics without any knowledge of its internal working mechanism.
5.3 Design Model 81 important requirements in order to support media adaptation for fixed and mobile computing. These requirements are: •Provide dynamic reconfiguration facilities: to tackle the changes that may occur in the application’s behavior and operating context at run-time. A framework supporting dynamic reconfiguration needs to detect changes and either reallocate resources, or tell the application to adapt to the changes. In this way, it can run efficiently under a broad range of conditions. •Provide flexibility: by providing a separation of concerns. The separation of distinct features that overlap in functionality as little as possible is highly desirable for building open systems [GK96]. The modularity and encapsulation of the programming shall be taken into consideration when software engineers develop the adaptive applications. •Provide asynchronous interaction: to tackle the problems regarding latency and disconnected communication that can arise. A client using asynchronous communication issues a request and continues operating and then collects the result at the appropriate time. This type of interaction reduces the consumption of network bandwidth and increases the scalability of the system. •Lightweight framework: this needs to be considered when deploying applications for mobile devices and to avoid designing applications that are to heavy to run on mobile devices with limited resources. •Context awareness: this is an important requirement for building efficient adaptive systems. The context of mobile users is usually determined by their current location, which, in turn, defines the environment where the computation associated with the unit is performed. The context in multimedia adaptation may include information that comes from the application (e.g., resolution, CPU power capabilities, media decoding capabilities) and network layer of the ISO/OSI reference model (e.g., available network bandwidth). In the following sections the design decisions of the RCLAF architecture are given in detail. 5.3 Design Model The requirements identified in the previous section and the analysis performed in Chapter 2 represent the basis for developing the RCLAF framework, which aims at supporting multimedia adaptation in heterogeneous networks. The framework developed by the author uses a distributed architecture for tackling the adaptation problems identified in the
82 Architecture of RCLAF usage scenarios from Section 1.1. For designing this framework, three important highlevel abstraction models were considered: •The programming model is fundamentally concerned with the mechanisms used to construct multimedia applications which perform multimedia content adaptation. How to perform the adaptation is greatly dependent on how the applications are constructed. Therefore the programming model has a major impact on designing the applications and on the tasks that they should carry out. •The knowledge model deals with those aspects referring to what kind of information the system needs in order to make adaptation decisions and how this information should be represented. In order to make adaptation decisions, the framework needs to capture a substantial amount of information that comes from different layers of the ISO/OSI stack model. A suitable knowledge model will facilitate this capture. •The decision model is concerned with the adaptation decision-making process. Finding a model to extract the best solution from a list of possible ones is challenging. A suitable decision model will determine the quality of the multimedia adaptation decision process. In the sections that follow, a detailed description of these three models is provided. 5.4 Pictorial Representation The interoperability framework involves the following types of entities: •Client Terminal devices that can be used by the clients (users) for multimedia content consumption. In this category one has nowadays smart-phone devices (e.g., the iPhone), PDA’s, notebooks, laptops, desktop computers, and TV set-top-boxes. In this thesis, the author took into consideration only the PCs (portable or desktop) and those devices capable of running Android OS. This decision was based on economic reasons: the Terminal component of the RCLAF framework which is completely implemented in the Java language will be more easily tested on a computer or on a PC mobile device emulator such as an Android emulator. •Service Provider entities which provide services to the customers, such as the distribution of analog or digital TV signals, Internet access, and phone service. They have contract agreements for content distribution with Content Providers. •Content Provider entities involved in the media production itself. The most common are Radio and TV stations, news agencies, and film studios.
5.4 Pictorial Representation 83 Client Service Provider Content Provider Figure 5.1: Actors entities relationship The relationships between these entities are illustrated in Figure 5.1. The framework proposed in this thesis comprises three software components: CLMAE, ServerFact, and Terminal, which are described as follows: •CLMAE stands for Cross-Layer Multimedia Adaptation Engine, and is responsible for multimedia adaptation decisions. For making adaptation decisions, this component uses a knowledge domain represented by a set of ontologies united under the acronym CLS (Cross-Layer Semantic). •ServerFact is the component which runs on the server side and facilitates the access to the media content dispensed by the Content Providers; it also manages the video streaming server. •Terminal is the component deployed on the user terminal and is responsible for creating (instantiating) a multimedia player to consume the multimedia content selected by the client. Figure 5.2 depicts the RCLAF’s processing nodes and how these software components are distributed across them. Client Machine Tomcat Application Server Tomcat Application Server Terminal Media Player CLMAE ServerFact VLCVoDServer VLC Application Apache Web Server <xml> MPEG-7 </xml> Streaming Server SOAP over HTTP SOAP over HTTP SOAP over HTTP Telnet HTTP Figure 5.2: The RCLAF’s UML Deployment Diagram
90 Architecture of RCLAF Figure 5.6: The UML sequence diagram for reflect service implementation bean of the ServFact service application. Then, using the prepareVoD message, it configures the streaming server in VoD mode. 5.5.1.2 Terminal The Terminal module unites a number of sub-components which are intended to be installed on the client terminal. These sub-components are: TerminalMediaPlayer, TerminalFactory, TerminalApp, and TerminalCallBack, and together compose a media player capable of interacting via SOAP with the CLAME and with the ServerFact modules of the RCLAF architecture. The TerminalMediaPlayer is the core component of the Terminal module and wraps a media player. This media player provides a set of operations through which it interacts with the ServerFact and the CLMAE modules. Table 5.2 summarize these operations. Operation Description connect Connects to the SP services select Checks if the selected media it is suitable for playing play Play the media stop Stop the media switch Switch terminal with other Table 5.2: The TerminalMediaPlayer operations
5.5 Overview of the Programming Model 91 Through the connect operation, the Terminal connects to the SP and retrieves the media content available. The UML sequence message chart of this operation is depicted in Figure 5.7. Figure 5.7: The UML sequence diagram of the connect request As shown, when the connect operation is invoked, the TerminalMediaPlayer constructs a SOAP message and sends it to the ServFactImpl component (the service bean implementation of the ServFact component). The ServFactImpl creates the object ServFactContentBind, which binds the available media content (XML meta-data) into content objects, and then calls the getMediaItems method for getting the available media items. The select operation is used to check whether a media content, received within a connect operation, can be played by the client terminal. As shown in the flow logic of this operation (Figure 5.8), the TerminalMediaPlayer creates and uses the TerminalFactory object for grabbing information about the terminal’s characteristics. The TerminalFactory follows the Abstract Factory design 3. An instrospector (e.g., the introspect operation) is used to grab information about the terminal’s characteristics and IDs (unique identifier). Then, a SOAP message containing this information is built by the TerminalMediaPlayer and sent to the OntADTEImpl component which represents the service bean implementation for the CLMAE component. The TerminalCallback object notifies the TerminalMediaPlayer when the response from OntADTEImpl is ready, and once the response is received, an adapter is used (e.g., the adapter operation) to implement the decision taken by the adaptation decision engine. The adapter instantiates a video player and prepares (e.g., by setting the player with the right video and audio decoding codecs) the player for consuming the resource. 3The purpose of the Abstract Factory is to provide an interface for creating families of related objects, without specifying concrete classes
92 Architecture of RCLAF The decision to use the Abstract Factory design comes from the fact that the Terminal component could be running on different hardware platforms (e.g., laptop or PDA) and therefore it is necessary to isolate the concrete objects that are generated. Making this isolation, the terminal can change the implementation from one factory (object) to another depending on the device platform. A detailed description of the implementation of this creational factory pattern is provided in Section 6.3.3. To consume a media resource, the TerminalMediaPlayer uses the play command. In the play command scenario depicted in Figure 5.9, when the player starts consuming the media content, the TerminalMediaPlayer sends a SOAP message to the OntADTEImpl component of the CLMAE module to “register” the state of the terminal (e.g., playing) in the AS ontology. The AS ontology is wrapped by the ADTEImpl component, and the OntADTEImpl creates a reference to this component and “registers” the application state through the isrtOP operation. Then an OWL reasoner is called (ResonerManager) to reason against the AS ontology to check whether a “terminal switching event” has arisen. In the affirmative case, the response of the playControl message will contain the position from where the player needs to start playing, otherwise, it will return a confirmation that the application state has been successfully registred. The TerminalMediaPlayer will invoke the playingMonitor operation which constructs a SOAP request notification message (ctrlMsg) and sends it to the OntADTEImpl. The OntADTEImpl will notify the terminal if during the play operation, for example, the network environment characteristics have changed and therefore an adaptation measures must be undertaken. The adaptation decision that comes from the meta-level, more exactly from the CLMAE component, is implemented through this message. When the TerminalMediaPlayer stops the media player, it informs the application that the state has changed and now the terminal is no longer playing. It constructs a SOAP message and sends it to the sendSignal bean implementation of the OntADTEImpl service, as shown in Figure 5.10. The OntADTEImpl receives the message, creates a reference to the ADTEImpl object, and calls the removeOP operation to remove the “playing” state associated with the player from the AS ontology. In case the client wants to switch terminals, the TerminalMediaPlayer provides the switch operation, through which the Terminal notifies the RCLAF framework (Figure 5.11). The TerminalMediaPlayer creates a reference to the SendSignal object, sets the signal flag (e.g., an integer number associated with the switch operation) and constructs a SOAP message with this flag. Then, the message is sent by the TerminalMediaPlayer by invoking the sendSignal bean implementation of the OntADTEImpl service.
5.5 Overview of the Programming Model 93 Figure 5.8: The UML sequence diagram for select operation
94 Architecture of RCLAF Figure 5.9: The UML sequence diagram of the play command Figure 5.10: The UML sequence diagram for stop operation
5.5 Overview of the Programming Model 95 Figure 5.11: The UML sequence diagram for the switch command 5.5.2 The structure of the meta-level The CLMAE component is located at the meta-level. As a consequence of the programming model described above, the basic constructs for the CLMAE component are metaobjects. The interfaces provided by these meta-objects are called meta-object protocols, or simply MOP . Each meta-object is associated with an individual base-level object for limiting the scope of reflection to the reified objects. The meta-object protocol may be used to dynamically access the meta-objects. The main sub-component of the CLMAE component is the OntADTEImpl, which provides several methods (services) through which the base-level of the RCLAF framework reifies information to the CLMAE module. These services are enumerated in Table 5.3: Operation Description reasoning Provides adaptation solution reifyNetInfo Reify network characteristics from base-level stateSignal Signal the state of the application ctrlMsg Message control playControl Play message control Table 5.3: The OntADTEImpl services This component acts as a wrapper around the ontologies (e.g., CLS and AS) used by the RCLAF architecture. This sub-component provides an interface which comprises two types of methods: operational ones and signalling ones. The operational methods offer access to the internal structure, while the signalling method informs the CLMAE
96 Architecture of RCLAF about changes (e.g., decreases in network bandwidth or switches between terminals) that may occur during the multimedia consumption process. This component is modeled as a web service interface and provides three operational methods: reasoning, reifyNetInfo, and playCTRL, and also one signalling method: stateSignal. By invoking the operational methods, we will be able to reify information from the base-level. Figure 5.12: The UML sequence diagram of the reasoning operation The reasoning method is called by the terminal every time it tries to consume a particular media resource (e.g., audio or video). The message sequence flow of the reasoning service bean implementation is shown in Figure 5.12. When the SOAP message constructed by the select method of the Terminal component (see Figure 5.8, operation 1.6) reaches the endpoint implementation, the OntADTEImpl service extracts the information from it (e.g., terminal and network characteristics) and inserts it into the CLS ontology. Then, it creates a reference to the ADTEImpl object, which acts as a wrapper for the CLS ontology. This wrapper provides an interface for inserting information into the CLS ontology as OWL entities. Afterwards, it creates an ReasonerManager object and calls the fireUpReasonerForData method, which invokes the OWL reasoner (Pellet) to reason against the AS ontology to “detect” whether the client terminal was changed. Finally, the OntADTEImpl service calls the runInitialSetUp which will reason against the CLS ontology for adaptation decision measures.
5.5 Overview of the Programming Model 97 To support the switching terminal scenario described in Section 1.1, the RCLAF framework should be aware if the client that starts consuming a multimedia resource is a new client or is a client who only changed the terminal or the access network. Figure 5.13: The UML activity diagram of the reasoning operation The RCLAF framework uses the AS ontology to add this feature. Each terminal or client has an associated id which is stored in the AS ontology. The AS ontology is responsible for storing information about the state of the application—a detailed description of this will be given in Section 5.6.2. To check whether a client has switched terminals, the RCLAF framework compares the user id and associated terminal id with the ones stored in the AS ontology and to check whether the client has switched networks, compares the terminal id with the IP addresses. The activity diagram depicted in Figure 5.13 shows how the RCLAF framework detects whether the client changed the network. The RCLAF framework checks first whether the terminal id about to be inserted already exists in the AS ontology. To check the terminal id, the RCLAF framework uses an OWL reasoner and reasons against the AS ontology. If the terminal id is not found
98 Architecture of RCLAF in the AS ontology, it is inserted and considered as a new client. Otherwise, the framework checks whether the IP address of the terminal has changed. If it has, the RCLAF framework decides that the client has changed access networks. The reifyNetInfo signalling method is used to reify information about network bandwidth from the base-level to the meta-level of the RCLAF framework. The sequence flow is shown in Figure 5.14. When the SOAP message reaches the reifyNetInfo implementation, the OntADTEImpl extracts the network id and its associated bandwidth. Then it searches in the AS ontology for the terminal that is using that network id and invokes the OWL reasoner to see whether the terminal needs a multimedia adaptation. The activity diagram of this process is depicted in Figure 5.15. Figure 5.14: The UML sequence diagram of the reifyNetInfo operation 5.6 The Knowledge Model The knowledge model uses a set of ontologies to describe the entities that are involved in the multimedia distribution and adaptation process. These ontologies are used by the
5.6 The Knowledge Model 99 Figure 5.15: The UML activity diagram of the reifyNetInfo operation
106 Implementation of RCLAF The RCLAF framework integrates a multimedia knowledge-domain represented using the OWL language and an inference engine for reasoning against this knowledge. This cross-layer reflective architecture follows a reflective architecture pattern which defines mechanisms for changing the structure and the behavior of the framework during execution time. The RCLAF framework is split into two parts: the meta-level, which makes the framework self-aware, and the base-level, which includes the application logic. In this way, the RCLAF has a highly structured approach which offers separation of concerns. 6.3 Component Programming As mentioned in Chapter 5, the RCLAF framework is composed of three modules: Terminal,CLMAE and ServerFact. These components comprise a comprehensive set of Java interfaces, abstract classes and classes that are grouped into packages as shown in Figure 6.1. Figure 6.1: The UML package diagram of the RCLAF architecture The webOntFact package contains classes and interfaces of the CLMAE component. The WebServFact package contains the classes of the module ServerFact and the terminal classes of the Terminal module. The manager package contains a set of classes and interfaces that make use of OWL-API version 2.2.0 3package. These are used to manipulate the RCLAF’s OWL ontologies. 6.3.1 CLMAE package This section describes the CLMAE programming model. This component incorporates the CLS and AS ontologies. The CLS ontology is the main component of the CLMAE 3http://sourceforge.net/projects/owlapi/
6.3 Component Programming 107 and represents the “heart” of the entire framework because it stores the pertinent information (e.g., device characteristics and network features) necessary to build the knowledgedomain which is meant to be used to dynamically adapt the multimedia content. In addition to the adaptation process, the AS ontology keeps track of the state of the application. To manipulate these two ontologies at the programming level, we exploit the benefits of the Pellet OWL-API package. This package provides a minimal set of classes and interfaces to facilitate ontology manipulations. However, the package does not provide direct methods for inserting, replacing or deleting entities or properties in ontologies. To overcome this shortcoming, the OWL-manager package has been created. In Section 6.3.1.1, a comprehensive description of this package will be given. The following lines describe the business logic behind the Web Service methods summarized in Table 5.3. The reasoning service is invoked when a terminal wants to play a particular media resource or when an external event (e.g., network traffic congestion) occurs. Based on the information provided by the terminal (e.g., device and access network characteristics), the reasoning bean implementation should be able to take the necessary adaptation measures in order to ensure a smooth video playing experience. The reified information sent by Terminal contains information about the characteristics of the device and the network. This data is encapsulated into a SOAP message and sent over wire. Listing 6.1 shows the format of the SOAP message request used in the reasoning service. <s oa p en v: E nv el o pe xm ln s: so ap en v =” h t t p : / / schemas . xmlsoap . org / soap / e nvelo pe / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ” xmlns:ter=”termCharact”> <soapenv:Hea d e r /> <soapenv:Body> <ont:request> <ter:Terminal> <TermCapability id=”?”> <DataIO T r a n s f e r S p e e d =”?” BusWidth=”?” MaxDevices=”?” MinDevices =”?”/> <Decoding i d =”?”> <Format>?</ Format> <Co decP ara me ter i d =”?”/> </ Decoding> <Display ActiveResolution=”?” B i t P e r P i x e l =”?” ContrastRatio=”?” Gamma=”?” MaxBrightness=”?” A c t i v e D i s p l a y =”?” Co lo rC ap ab le =” ? ” FiledSequentialColor=”?” StereoScopic=”?” sRGB=”?”> <CharacterSetCode>?</ CharacterSetCode> <C ol or Bi tD ep th re d =”?” g re en =”?” b l u e =”?”/> <ColorPrimaries>?</ C o l o r P r i m a r i e s> <Renderin gF or ma t h r e f =”?”/> <S c r e e n S i z e h o r i z o n t a l =”?” vertical=”?”/> <Mode refreshRate=”?”> <R e s o l u t i o n i d =”?” horizontal=”?” vertical=”?” activeResolution=”?”/>
108 Implementation of RCLAF <Si z eCh a r h o r i z o n t a l =”?” vertical=”?”/> </ Mode> </ Display> <Enc oding i d =”?”> <Format>?</ Format> <Co decP ara me ter i d =”?”/> </ Encodin <Power AverageAmpereConsumption=”?” B a t t e r y C a p a c i t y R e m a i n i n g =”?” BatteryTimeRemaining=”?” RunningOnBatteries=”?”/> <S to r ag e I n p u t T r a n s f e r R a t e =”?” O u t p u t T r a n s f e r R a t e =”?” S i z e =”?” W r i t e b l e =”?”/> </ TermCapability> <AccessNetwork id=”?” maxCapacity=”?” minGuaranteed=”?” inSequenceDelivery=”?” e r r o r D e l i v e r y =”?” errorCorrection=”?” maxPacketSize=”?”/> </ ter:Terminal> <T e r m i n a l I d> <IPAddress>?</ I PAd dre ss> <MACAddress>?</ MACAddress> <HostName>?</ HostName> </ T e r m i n a l I d> <mediaItem>?</ mediaItem> </ o n t : r e q u e s t> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.1: The reasoning message format request As can be seen, the request message carries the following data: 1. Information regarding the terminal characteristics, such as terminal and access network characteristics, within the Terminal XML tags; 2. Information within the TerminalId XML tags with respect to the identification of the terminal, such as its IP address, MAC address, and hostname. 3. The name of the media content that the terminal wants to play, specified inside of the mediaItem XML tags. Annex C.3 gives an example of a reasoning SOAP Message request. When the request message reaches the end-point address, the reasoning service transforms the data suitable for transferring into a data object, a process known as unmarshalling, and extracts the information from it. Then, it verifies whether the termId has already been inserted into the AS ontology. If so, it runs the OWL2 reasoner, searching for new facts about the termId. The work flow activity of this process was drawn in Figure 5.13 and the equivalent implementation is shown in Listing 6.2. i f ( s t a t e . ont . o n t
6.3 Component Programming 109 . c o n t a i n s I n d i v i d u a l R e f e r e n c e ( URI . c r e a t e ( bf . t o S t r i n g ( ) ) ) ) { ReasonerManagerInterf reasonerState = new ReasonerManager ( s t a t e . on t ) ; t r y { Map<OWLIndividual , Set<OWLConstant>> e v e n t s = r e a s o n e r S t a t e . fireUpReasonerForData(” d e t e c t e d E v e n t s ” ) ; Set<Entry<OWLIndividual , Set<OWLConstant>>> e v e n t s = e v e n t s . entrySet () ; i f ( e v e n t s . isEmpty ( ) ) { Response r e s u l t = r u n I n i t i a l S e t U p ( ) ; return result ; }else { f o r ( Entry<OWLIndividual , Set<OWLConstant>> e n t r y : e v e n t s ) { i f ( e n t r y . getKey ( ) . e qu a l s ( h o s t I d ) ) { Set<OWLConstant>e n t r y 1 = e n t r y . g et Va lu e ( ) ; f o r ( OWLConstant o wl Co ns ta nt : e n t r y 1 ) { i f ( ow lC on st an t . t o S t r i n g ( ) . e q u a l s I g n o r e C a s e ( ” Changed t h e net work a c c e s s ! ” ) ) {/ / N o ti f y t e r m i n a l } } }else { Response r e s u l t = r u n I n i t i a l S e t U p ( ) ; return result ; } } } }catch ( E x c e p t i o n e ) {throw new F a u l t ( e . getMessage ( ) , ” F a u l t I n f o ! ” ) ; } }else { Response r e s u l t = r u n I n i t i a l S e t U p ( ) ; return result ; } Listing 6.2: A snipped from reasoning Web Service implementation The runInitialSetUp method used here inserts all the information needed by the ADTE to take adaptation decision measures. Part of this information comes within the SOAP request message. The other part is grabbed from the base-level, by invoking the getMediaCharact and getServerNetCharact methods as shown in Listing 6.3. These methods, in turn, call the ServerFact’s introspection interfaces: introspectMediaCharact respectively introspectNet. The runInitialSetUp passes the name of the media content for which it is searching on to the introspectMediaCharact interface and retrieves the characteristics of the media. . . . getMediaCharact(” t e s t ” , mediaItem ) ; I nt r os p N et R e sp i n t r o s p N e t = g e t S e r v e r N e t C h a r a c t ( ) ; ReasonerManagerInterf reasoner = new ReasonerManager ( a d t e . o nt ) ; try { r e s u l t = r e a s o n e r . f i r e U p R e a s o n e r ( ” p l a y ” ) ; }catch ( E x c e p t i o n e ) {throw new F a u l t ( e . getMessage ( ) , ” F a u l t I n f o ! ” ) ; }
110 Implementation of RCLAF R e s u l t S e t r e s u l t 1 = r e a s o n e r . g e t B e s t B i t R a t e ( ) ; Response r e s p o n s e = new Response ( ) ; i f ( r e s u l t 1 . hasNext ( ) ) { Q u e r y S o l u t i o n qr = r e s u l t 1 . n e x t ( ) ; Q u er yS ol ut io n qr2 = r e s u l t 1 . n e x t S o l u t i o n ( ) ; / / Get t h e name of t he v ideo r e s o u r c e . S t r i n g v i de o = qr . g e t R e s o u r c e ( ” v i d e o ” ) . getLocalName () ; / / Get t h e l o c a l URI . S t r i n g l o c a l = qr . g e t L i t e r a l ( ” l o c a t i o n ” ) . g e t S t r i n g ( ) . t ri m ( ) ; / / Get t he a v er ag e b i t R a t e . int avg = q r . g e t L i t e r a l ( ” avg ” ) . g e t I n t ( ) ; / / Get t h e v id e o f o rm a t . S t r i n g videoForm = qr . g e t L i t e r a l ( ”videoFormat”) . g e t S t r i n g ( ) . t r i m ( ) ; / / Get o n l i n e URI . S t r i n g l o c a l 2 = qr2 . g e t L i t e r a l ( ” l o c a t i o n ” ) . g e t S t r i n g ( ) . t r i m ( ) ; i f ( vi d e o != n u l l && l o c a l != n u l l && avg != 0 && l o c a l 2 != n u l l && videoForm != n u l l ){ a d a p t ( v id eo , l o c a l ) ; /∗Sends t h e r t s p URI t o t e r m i n a l ∗/ r e s p on se . setRts pU RI ( l o c a l 2 ) ; /∗Inform t h e t e r m i n a l about what codec t yp e ∗p l a y e r must i n s t a n t i a t e . ∗/ r e s p o n s e . set Code cTyp e ( videoForm ) ; . . . Listing 6.3: A code snippet from runInitialSetUp method implementation Once the information is instantiated into the CLS ontology, the runInitialSetUp method calls the ReasonerManager, which is an OWL 2 reasoner implementation (based on Pellet API), of the ReasonerManagerInterf. The reasoner looks to see if new facts are inferred along the OWL object property play, and returns the response as a collection (Java Set API interface). The RCLAF framework uses the ‘the best bitrate’ as its adaptation strategy when more than one solution is returned. The getBestBitRate method implements this strategy by using the SPARQL language. A brief description about how this method is implemented will be given in Section 6.3.1.1. This method sorts by bitrate in descending order, the media item variations that can be played by the terminal in actual conditions, and returns the first one from the result set, representing the item variation with the best bitrate. The adaptation decision must be implemented at the base-level of the architecture. On the server-side, the ADTE configures the streaming server with the right media, while on the terminal-side, it instantiates the media player with the right codecs. The set-up of the streaming server is achieved through the adapt method which takes as parameters the name of the VoD channel that streaming server should create and the URI of the video resource that will ‘feed’ the server. The adapt method calls, in turn, the ServerFact’s reflect interface to configure the streaming server in VoD mode. On the terminal-side, the
6.3 Component Programming 111 ADTE constructs a SOAP response message as shown in Listing 6.4, containing information about the RTSP address of the VoD channel (inside of the rtspURI XML tags) and the name of the codec (inside of the codecType XML tags) that the terminal media player should instantiate for playing. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ”> <soapenv:Hea d e r /> <soapenv:Body> <ont:response> <rtspURI>?</ rts pURI> <codecType>?</ codecType> </ ont:response> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.4: The reasoning message response format The stateSignal service informs the RCLAF framework in which state the Terminal is. There are four application states defined: three of them are as follows: STOP signals the fact that the terminal has stopped playing the stream; EXIT informs the framework that the terminal has been disconnected from the SP; SWITCH signals the fact that the user wants to switch terminals. For each of these flags, there is associated an integer value (see Table 6.1) which will facilitate, as will presented in the following lines, the business-logic implementation of the OntADTEImpl’s stateSignal interface. The Terminal uses these flags to signal the state to the RCLAF framework by invoking the stateSignal service. The Terminal sends the id of the terminal that uses the flag (inside of the termId XML tags) and the name of the media resource on which ‘he acts’ (inside of the videoId XML tags) as shown in Listing 6.1. The SOAP message is unmarshalled into a data object when it arrives at service bean implementation. Depending on the value of the flag, several execution paths are defined using the switch statement, as the snippet code from Listing 6.6 shows. Flags Description Associated value STOP signals stop playing 2 EXIT signals the exit 3 SWITCH signals switching terminal 4 Table 6.1: The stateSignal flags <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ”> <soapenv:Hea d e r /> <soapenv:Body> <ont:sendSignal> <te r m I d>?</ t e rm I d>
112 Implementation of RCLAF <videoId>?</ videoId> <int>?</ i n t> </ ont:sendSignal> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.5: StateSignal message request format When the flag STOP is sent by the terminal, the stateSignal service calls the ReasonerManager and removes from the AS ontology the OWL data property isConsuming associated to this terminal id. In the case of EXIT, the stateSignal service removes the OWL entity from the AS ontology that represents the id of the terminal which has been disconnected from the SP. If the SWITCH flag is received, the terminal id is saved into the AS ontology as ‘paused’ and tracks the current position in the stream being played. . . . switch ( f l a g . g e t I n t ( ) ) { c a s e 2 : / / STOP s t a t e . removeOP ( termId , ”isConsuming” , v i d e o I d ) ; try { s t a t e . on t . man . save O ntol o g y ( s t a t e . o n t . on t ) ; }catch . . . break ; c a s e 3 : / / EXIT ReasonerManagerInterf reasoner = new ReasonerManager ( s t a t e . on t ) ; Set<Entry<OWLIndividual , Set<OWLIndividual>>> e n t r y = null ; try { e n t r y = r e a s o n e r . f i r e U p R e a s o n e r ( ”isBeingUsedBy”) . entrySet () ; }catch ( E x c e p t i o n e ) { } f o r ( Entry<OWLIndividual , Set<OWLIndividual>> e n t r y 2 : e n t r y ) { f o r ( OWLIndividual owlIndv : e n t r y 2 . ge tV al u e ( ) ) { i f ( owlIndv . getURI ( ) . g et Fr ag me nt ( ) . e q u a l s I g n o r e C a s e ( t er mI d ) ) { s t a t e . removeIndv ( e n t r y 2 . getKey ( ) . getURI ( ) . g et Fra gm en t ( ) ) ; } } } s t a t e . removeIndv ( t e rm I d ) ; synchronized ( look ) { } break ; c a s e 4 : / / SWITCH s t a t e . i s r t D P ( term Id , ” i s P a u s i n g ” ,true) ; . . . try { s t a t e . on t . man . save O ntol o g y ( s t a t e . o n t . on t ) ; }catch (UnknownOWLOntologyException e) { } . . . break ; . . .
6.3 Component Programming 113 Listing 6.6: A snippet code from stateSignal implementation The PlayCtrl service is invoked by the terminal before it starts playing the media resource. This service verifies whether the user that wants to consume the media resource is a new user or a client who switched terminals. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ”> <soapenv:Hea d e r /> <soapenv:Body> <ont:playReq> <te r m I d>?</ t e rm I d> <videoId>?</ videoId> </ o n t : p l a y R e q> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.7: PlayCtrl request message format The terminal id (termId XML tags) and the video that the user wants to play (videoId XML tags) are encapsulated into a SOAP message, as Listing 6.7) shows. They are instantiated into the AS ontology and are ‘linked’ through the OWL object property isPlaying. Then, the service calls the OWL reasoner to see if the new facts are inferred along the OWL object property swithTerminal. If new facts are inferred, a ‘switch terminal event’ is triggered by the RCLAF framework and starts to interrogate the AS ontology, the current position from where it left the stream. A snipped Java code showing how the service detects ‘switch terminal’ events is listed in Listing 6.8. The PlayCtrl service notifies the Temrminal component about the position in the current stream through a SOAP message response (see Listing 6.9). boolean switchEvent = false ; . . . switchEvent = reasoner1 . fireUpReasoner(” s w i t h T e r m i n a l ” ) . entrySet () ; }catch ( E x c e p t i o n e1 ) { } f o r ( Entry<OWLIndividual , Set<OWLIndividual>> entry : switchEvent ) { i f ( e n t r y . getKey ( ) . getURI ( ) . ge tF ra gm en t ( ) . e q u a l s I g n o r e C a s e ( play Req . g e t V i d e o I d ( ) ) ) { f o r ( OWLIndividual e n t r y 2 : e n t r y . g et Va lu e ( ) ) { i f ( e n t r y 2 . getURI ( ) . ge tF ra gm en t ( ) . e q u a l s I g n o r e C a s e ( playReq . getTermId ( ) ) ) { switchEvent = true ; } } } . . . Listing 6.8: A snipped code from PlayCtrl service implementation
114 Implementation of RCLAF If new facts are not inferred, the RCLAF framework treats the terminal as a new one, and associates its playing state into the AS ontology through the isPlaying OWL object property. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ”> <soapenv:Hea d e r /> <soapenv:Body> <ont:playResp> <currentPosition>?</ c u r r e n t P o s i t i o n> </ ont:playResp> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.9: PlayCtrl response message format The reifyNetInfo service follows the work flow described in Figure 5.15. It is a oneway messaging service (fire-and-forget messaging) which basically means that the service client sends the message and then closes the connection. The network probes and sends the measured network bandwidth and the network’s id (see Listing 6.10) and based on this information, the reifyNetInfo service searches the ‘bandwidth drops!’ events. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ont=” h t t p : / / a r e s . i n e s c n . p t : 8 0 8 0 / O nt ol o gy Fa c t / s e r v i c e s / On t ol og y Fa c t . xsd4 ”> <soapenv:Hea d e r /> <soapenv:Body> <ont:probe> <bandwidthMeasured>?</ bandwidthMeasured> <net ID>?</ netID> </ o n t : p r o b e> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.10: ReifyNetInfo request message format 6.3.1.1 OWL-Manager package This package was defined and implemented to facilitate the CLMAE’s interaction with the CLS and AS ontologies. Figure 6.2 shows the class diagram of the OWL-manager package. Based on the Pellet’s OWL-API, the OWL-manager API provides the ADTE abstract class which facilitates direct access to the operations that are usually made on ontologies, such as inserting, replacing, or deleting entities or properties. Below is a snipped code of the insertion of a data property (Listing 6.11). p u b l i c voi d i s rt D P ( S t r i n g indv , S t r i n g prop , boolean d a t a ) { StringBuffer bf = new S t r i n g B u f f e r ( ont . o nt . getURI ( ) . t o S t r i n g ( ) ) ; bf . append ( ”#”) ;
6.3 Component Programming 115 Figure 6.2: The class diagram of the OWL-manager package
122 Implementation of RCLAF Figure 6.5: Class and Property representation of CodecParameter </ owl:DatatypeProperty> Listing 6.13: The CodecParameterFillRate subclases representationin using OWL 2 specifications The classes VertexRate,BitRate,MemoryBandwidth and BuferSize have been represented in the same way. The data properties were defined like functional properties in OWL, because we suppose that in our scenario, an individual of entity bitRate, for instance, can have only one value. Similar to the CodecParameter schema, we may have other transformations for schema models. This schema has been chosen to prove that we are able to “translate” a typical MPEG-21 description schema into an OWL 2 representation model. Once we have defined the UED ontology, we went to the next step in our development, which involves defining an ontology for our scenario. The ontology was constructed using an OWL 2 specification and incorporates the knowledge requirements of the scenario described in Section 1.1. Modeling a multimedia adaptation scenario is complicated. We defined a terminal as being an object. Each terminal object can either be a PDA, Notebook, MobilePhone, DesktopPC or a SetTopBox. The Terminal concept defines a terminal with some particular characteristics and has a relationship with MediaResource concepts through the play property, which states that the actual terminal is able to play that particular resource(s). The instances of Terminal can be described through the hasTermCapability property, which establishes a relationship between a terminal and that terminal’s capabilities. In the same way, the media characteristics are represented through the hasMediaCharact property. The objective in our adaptation scenario is to find the right multimedia content that fits the user’s choice and characteristics. We developed an initial set of rules to achieve this
6.3 Component Programming 123 using the SWRL editor provided by the Protege tool 9. This editor permits direct editing of SWRL code with its inherent constraint to binary predicates. The logic captured by these rules is shown in Equation 6.1, in which the variables are prefaced with question marks. hasBitRate(?f ormat,?bitRate)∧hasFileFormat(?f ormat,?f ileForm)∧ hasFrame(?f ormat,?f rame)∧hasMediaFormat(?pro file,?f ormat)∧ hasMediaPro f ile(?resrc,?pro f ile)∧hasDecding(?termCap.?aCdc)∧ hasDecoding(?termCap,?vCdc)∧hasDisplay(?termCap,?display)∧ hasMode(?display,?mode)∧hasResolution(?mode,?res)∧ hasAccessNetwork(?term,?net)∧hasMediaCharact(?media,?resrc)∧ hasTermCapability(?term,?termCap)∧hasAudioNameFormat(?fileForm,?aForm)∧ hasAverage(?bitRate,?avg)∧hasHeight(?f rame,?height)∧ hasVideoNameFormat(?fileForm,?vForm)∧hasWidth(?f rame,?width)∧ hasAudioCodec(?aCdc,?audio)∧hasHorizontalSizeC(?res,?horiz)∧ hasMinGuaranteed(?net,?bandwidth)∧hasVerticalSizeC(?rest,?vert)∧ hasVideoCodec(?vCdc,?video)∧equal(?audio.?aForm)∧ equal(?video,?vForm)∧greaterThanOrEqual(?bandwidth,?avg)∧ greaterT hanOrEqual(?horiz,?width)∧ greaterT hanOrEqual(?vert,?height)⇒play(?term.?media) (6.1) To be able to play a chosen media on a particular terminal, the following requirements must be met: 1) the media is encoded in a format which is supported by the terminal device, 2) the resolution the terminal is capable of is at least equal to that of the chosen media, and 3) the user access network bandwidth is capable of supporting the media bitrate flow. The atoms equal(?audio,?aForm)and equal(?video,vForm?)are used to check the first requirement. The variables audio and video are references to the terminal device audio and video codec capabilities, while the aForm and vForm variables describes the codec of the media chosen. The second requirement is checked through the following atoms: greaterT hanOrEqual(?horiz,width)and greaterThanOrEqual(?vert,?height). The horiz and vert atoms variables describe the terminal device resolution, while the width and height describe the frame resolution of the media chosen. Finally, the last requirement is checked by the SWRL rule atom greaterThanOrEqual(?bandwidth,?avg), where the 9Protege is a free, open source ontology editor and knowledge-base framework.
124 Implementation of RCLAF bandwidth variable holds the bandwidth link value between the streaming server and the client and the avg holds information about the media average bitrate. When the aforementioned requirements are met, the SWRL rule 6.1 is triggered by the OWL reasoner and will infer along the play object property the fact that the terminal is able to play the chosen media. In some cases, the chosen media may be described by various variations and therefore only those variations that meet the requirements are chosen. In this case, the result of the OWL reasoning process will be a set of variations and how the CLMAE will choose the final solution from this set is an optimization problem which is solved by applying an adaptation policy (e.g., the best solution is the media variation with the best bit-rate) which will extract the solution that will be decided on. The set of the proposed solutions is ordered by bit-rate using the getBestBitRate() method (Figure 6.2). We chose the first variation from the set because having the best bit-rate, it will provide the best video experience to the user. 6.3.1.3 AS ontology The role of the AS ontology is to maintain the state of the base-level application. The RCLAF framework must be aware at all times of the state of the terminal, what media content the terminal is consuming, and what are the network conditions being used by the terminal for media delivery. The state of a terminal is one of the following states: playing, paused, stopped. These states must be reified by the AS ontology. The class hierarchy is shown in Figure 6.6: Figure 6.6: The AS hierarchy model The AS ontology models the principal actors that participate in the multimedia consuming scenarios described in Section 1.1: Video, Streaming Server, Network and Terminal/Client. The top classes (concepts) of the AS ontology have in their turn subclasses. All ABox (cf. assertion boxes) statements are described as follows. Video is characterized by Resolution and videos may be available in various video resolutions. The Network class
6.3 Component Programming 125 models all the characteristics of the network links. In this category are: local loops 10, Internet network paths, and private networks. Two local loops are defined: ServerLocalLoop which represents, at the conceptual level, the network link that connects the server to the Internet, and the ClientLocalLoop, which defines the network link between the ISP and the customer’s premises. Streaming server represents a physical (hardware board) or software application for video streaming purposes. The AS ontology contains also the TBox (cf. terminologies boxes) statements, which use classes and properties to describe data instances (ABoxes). Table 6.4 lists the TBox statements used. The ABox and TBox statements model the contextual requirements for tracking the state of the application. A declaration such as: “A streaming server streams a video resource to a terminal through a specific network path” can be logically separated into four context entities: StreamingServer, Video, Terminal and NetworkPath. These four parts are represented in the AS ontology by four properties respectively: serverConnectedTo, isUsingVideo, isConsuming and isBeingUsedBy. Figure 6.7 illustrates an example describing the state of a stream consumed by a particular terminal in the AS ontology. as:StreamingServer as:Instance as:Video as:ServerLocalLoop as:NetworkPath as:Network owl:Class rdf:individual rdf:property isa as:serverConnectedTo as:isUsingVideo as:isPartOf as:isUsingNetwork as:Terminal as:isBeingUsedBy Legend: Figure 6.7: An AS model example The AS ontology is implemented using the OWL-DL language. A set of SWRL rules have been added on top of OWL to leverage the expressivity and to preserve the decidability of the knowledge model. The RCLAF framework must be able to react to external stimuli, such as: network bandwidth drops or switching terminals. These stimuli are captured in the AS ontology and, together with an inference engine, the RCLAF framework is able to extract the necessary information to decide whether an external stimuli will trigger or not a multimedia adaptation process. A set of eight SWRL rules have been defined to help the AS ontology in this matter. As I mentioned before, for our scenarios two external stimuli are taken into consideration: 10The Local loop or subscriber line is the physical link or circuit that connects the demarcation point of the customer’s premises to the edge of the carrier or telecommunications service provider’s network.
126 Implementation of RCLAF Concept Role (properties) Concept/Value Types unionOf(ServerLocalLoop, isPartOf Network ClientLocalLoop) Network isUsingNetwork NetworkPath NetworkPath isBeingUsedBy Terminal Terminal isConsuming Stream Terminal isUsingNetPath NetworkPath Terminal clientConnectedTo ClientLocalLoop StreamingServer isUsingVideo Video StreamingServer serverConnectedTo ServerLocalLoop StreamingServer switchTerminal Terminal NetworkPath bandwidth rdf:positiveInteger NetworkPath bandwidthEvent rdf:string ClientLocalLoop clientAvgBandwidth rdf:positiveIneteger unionOf (ClientLocalLoop, clientMeasuredBandwith rdf:positiveInteger ClientLocalNet) Terminal detectedEvents rdf:string StreaminServer doAction rdf:string Terminal hasCurrentHardwAddress rdf:string Terminal hasCurrentIPAddress rdf:string Terminal hasHardwareAddr rdf:string Terminal hasIPAddrr rdf:string unionOf (ClientLocalLoop, hasNetId rdf:string ServerLocalLoop) Terminal isPausing rdf:boolean ServerLocalLoop serverAvgBandwidth rdf:positiveInteger ServerLocalLoop serverMeasuredBandwidth rdf:positiveInteger Terminal trackCurrentTime rdf:int NetworkPath useChannel rdf:positiveInteger Table 6.4: The TBox statements used in AS ontology
6.3 Component Programming 127 changing the network bandwidth, on either the server-side or client-side link, and switching terminals. When network bandwidth fluctuation occurs on the clientor server-side, the RCLAF framework must be able to check if the current available network bandwidth is enough for smooth playing. In order to check that, the new value is added to the AS ontology using the clientMeasuredBandwith or serverMeasuredBandwith properties, and the rules 6.2 or 6.3 are fired by the Pellet OWL reasoner depending on where the bandwidth has been measured (i.e., on the client-side or on the server-side, respectively) to determinate the new available network bandwidth between the streaming server and the client terminal. The captured logic is depicted in rules 6.2 or 6.3, in which variables are prefaced with question marks. NetworkPath(?path)∧isUsingNetwork(?path,?client)∧ isUsingNetwork(?path,?server)∧clientMeasuredBandwidth(?client,?cmeasured)∧ serverMeasuredBandwidth(?server,?smeasured)∧lessT han(?cmeasured,?smeasured) ⇒bandwidth(?path,?cmeasured) (6.2) NetworkPath(?path)∧isUsingNetwork(?path,?client)∧ isUsingNetwork(?path,?server)∧clientMeasuredBandwidth(?client,?cmeasured)∧ serverMeasuredBandwidth(?server,?smeasured)∧lessT han(?smeasured,?cmeasured)∧ ⇒bandwidth(?path,?smeasured) (6.3) Once the new bandwidth value is obtained, it is compared with the current one. If the new value is less, then the SWRL rules 6.4 and 6.6 will be “fired” by the Pellet OWL reasoner . A “network bandwidth dropped!” event will be trigged together with the name of the terminal that needs adaptation. bandwidth(?path,?updateVal)∧useChannel(?path,?val)∧ lessThan(?updateVal,?val)⇒bandwidthEvent(?path,“droped!00) (6.4) Otherwise the reasoner will “fire” rule 6.5 that generates a “doAction” event which means “continue streaming!” and no action will be taken.
128 Implementation of RCLAF bandwidth(?net,?val)∧useChannel(?stream,?ch)∧ lessThan(?ch,?val)⇒doAction(?stream,“continue streaming!”) (6.5) isConsuming(?term,?stream)∧isUsingNetPath(?term,?path)∧ bandwidth(?path,?updateVal)∧useChannel(?path,?val)∧ lessThan(?updateVal,?val)⇒detectedEvents(?term,”needs adaptation!”) (6.6) The SWRL rules 6.7 and 6.8 are sequentially fired by the Pellet reasoner when the client switches between terminals. The switch event is triggered if the reasoner detects a different Hardware Address from the one used currently by the terminal. isConsuming(?term,?video)∧isConsuming(?term1,?video)∧ hasHardwareAddr(?term,?mac)∧hasHardwareAddr(?term1,?mac1)∧ isPausing(?term,true)∧notEqual(?mac,?mac1)⇒switchTerminal(?video,?term1) (6.7) hasCurrentIPAddress(?term,?currentIP)∧hasHardwareAddr(?term,?mac)∧ hasIPAddress(?term,?ip)∧notEqual(?ip,?currentIP) ⇒detectedEvents(?term,”Changed the network access!”) (6.8) 6.3.1.4 Inferring execution process To execute the SWRL rules described in the previously sections, in either Pellet 11 or another rule engine, they must first be translated into the corresponding rule-language. In doing so, engine specific characteristics need to be taken into account that lie outside the scope of the SWRL. The engine will evaluate the content of the working memory and fire rules when the antecedents are satisfied. The firing of a rule can result in the “assertion” of 11Pellet is an OWL 2 reasoner. Pellet provides standard and cutting-edge reasoning services for OWL ontologies.
6.3 Component Programming 129 new facts into working memory. These facts must be transferred back as OWL knowledge into the OWL ontology. The overall process diagram can be seen in Figure 6.8. Figure 6.8: Process Diagram of Rule Execution 6.3.2 ServerFact package The ServerFact is located at the base-level of the architecture and controls the streaming server. Like CLMAE, the ServerFact component was developed and deployed as a Java web-service application into an Apache Tomcat container and provides services for the Terminal and CLMAE components. The class diagram overview of the ServerFact package is shown in Figure 6.9. The ServerFact bean implementation is centered around the ServFactImp class, which offers the business logic for the services enumerated in Table 5.1. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ”> <soapenv:Hea d e r /> <soapenv:Body /> </ so a p e n v: E n v el o p e> Listing 6.14: Get media request message format <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ser=” h t t p : / / a r e s . i n e s c n . p t / S e r v F a c t . wsdl . xsd1 ”> <soapenv:Hea d e r /> <soapenv:Body> <ser:mediaItemsWS> <s e r : I t e m i d =”?”> <s y n o p s i s>?</ s y n o p s i s> </ ser:Item> </ ser:mediaItemsWS> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.15: Get media response message format
130 Implementation of RCLAF Figure 6.9: An overview of the ServerFact package To reduce program complexity and to allow a realistic reflective implementation, the introspectors (e.g., introspectMediaCharact,introspectNet) and the adaptor (e.g., reflect) are already created as web service methods and are available to the meta-level. The CLMAE accesses them only through a valid access key. This key is obtained by invoking the createIntrospector(), respectively, createAdaptor() service. The request SOAP messages for invoking these two services are shown in Listing 6.16, respectively, Listing 6.17. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ser=” h t t p : / / a r e s . i n e s c n . p t / S e r v F a c t . wsdl ”> <soapenv:Hea d e r /> <soapenv:Body> <ser:createIntrospector /> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.16: Create Introspector request message format <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ser=” h t t p : / / a r e s . i n e s c n . p t / S e r v F a c t . wsdl ”> <soapenv:Hea d e r /> <soapenv:Body> <ser:createAdaptor /> </ soapenv:Body> </ so a p e n v: E n v el o p e>
6.3 Component Programming 131 Listing 6.17: Create Adaptor request message format Until we will have standardized the creation of the web services on demand, this is a way to ensure the correct setting of introspectors and adapters, which is critical for the operation of the RCLAF framework. The createIntrospector service responds with a SOAP message which has the format depicted in Listing 6.18 and createAdaptor with the message from Listing 6.19. The key is a string data type which is randomly generated and is specified inside the tag out of the SOAP body. <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ser=” h t t p : / / a r e s . i n e s c n . p t / S e r v F a c t . wsdl ”> <soapenv:Hea d e r /> <soapenv:Body> <ser:createIntrospectorResponse> <out>?</ o u t> </ ser:createIntrospectorResponse> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.18: Create Introspector response message format <s oa p en v: E nv el o pe x ml ns :s oa pe nv =” h t t p : / / schemas . xmlsoap . org / soap / en v elo p e / ” xmlns:ser=” h t t p : / / a r e s . i n e s c n . p t / S e r v F a c t . wsdl ”> <soapenv:Hea d e r /> <soapenv:Body> <ser:createAdaptorResp> <out>?</ o u t> </ ser:createAdaptorResp> </ soapenv:Body> </ so a p e n v: E n v el o p e> Listing 6.19: Create Adaptor response message format The introspectMediaCharact() method uses the class ServFactContentBind which, in turn, uses the JAXB API 12 to access the SP available media content described in the XML format using MPEG-7 Schema specifications. The most obvious benefit of this is the fact that the unmarshalling process of the meta-data content builds a tree of content objects which are necessary for the implementation of the introspectMediaCharact() as a web service. The content tree is not a DOM representation of the media content and therefore the content trees produced through the JAXB are more efficient in terms of memory use than DOM-based trees. Listing 6.20 shows the SOAP message format produced by the CLMAE to invoke the introspectMediaCharact service. As can be seen, the SOAP message includes the element 12JAXB allows Java developers to access and process XML data without having to know XML or XML processing.