Agora: A Distributed Language Model Framework With API-Call Support for Integrated Climate Forecasting
Full text
Received 6 April 2025, accepted 24 April 2025, date of publication 8 May 2025, date of current version 19 May 2025. Digital Object Identifier 10.1109/ACCESS.2025.3568028 Agora: A Distributed Language Model Framework With API-Call Support for Integrated Climate Forecasting ALEXANDRA UDRESCU AND DAN-MATEI POPOVICI Computer Science Department, National University of Science and Technology POLITEHNICA Bucharest, 060042 Bucharest, Romania Corresponding author: Dan-Matei Popovici ([email protected]) This work was supported in part by the European Union through the FUTURAL Project-Empowering the FUTure through innovative Smart Solutions for rURAL areas (HORIZON EUROPE) under Project 101083958, and in part by Unitatea Executivă pentru Finant .area Învăt ,ământului Superior, a Cercetării, Dezvoltării Si Inovării (UEFISCDI) through the Project FUTURAL-Soluţii inteligente inovatoare pentru zonele rurale (Orizont Europa Institutii) under Project 020234823. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. ABSTRACT We introduce Agora, a Generative AI-driven system that delivers expert answers and recommendations on climate and agriculture, transforming complex data into clear, natural language explanations. While built for the rural domain, Agora is highly adaptable and can be deployed across various domain applications. It operates as a ‘‘mixture-of-experts’’ language model system, selectively utilizing multiple fine-tuned large language models for inference. By dynamically integrating external data through API calls, Agora ensures real-time, contextually relevant responses. Agora is built for extensibility—it seamlessly integrates new APIs and domains without requiring a full system retrain. Developed entirely with open-source large language models from the LLaMA family, Agora remains open and adaptable, allowing anyone to extend and enhance its capabilities. Optimized for accessibility, Agora runs efficiently on commodity GPUs without compromising performance. By eliminating the need for expensive hardware like NVIDIA’s A100, it makes text generation more affordable and widely accessible. Agora outperforms closedsource models, achieving 78% accuracy on our question-answering benchmark. This result is achieved via dynamic API integration, which pulls in real-time external data, making responses more adaptive, precise, and context-aware. INDEX TERMS API-call orchestration, API-call support, forecast generation, large language model, model fine-tuning, natural language processing. I. INTRODUCTION In many Eastern European countries, particularly Romania, a significant gap exists between the technological advancements and practices prevalent in rural areas. Despite the widespread availability of Internet connectivity Romania ranks within the top 54 of 140 countries for mobile Internet speed [1] and accessibility to expert data, smart services, and cutting-edge information remains low in these regions. A report [2] from the Black Sea Basin Programme revealed that approximately one million people in Romania, along with their families, are disconnected from modern advancements. This is further supported by the statistics The associate editor coordinating the review of this manuscript and approving it for publication was Arianna D’Ulizia . shown in Figure 1. The same report highlights that 97% of Romanian farms are microand subsistence enterprises, typically family owned, covering up to 10 ha. These farms employ at least half the agricultural workforce available in Romania. Although specialized weather forecasts, seasonal cropping information, and insights into the effects of climate change are readily available [3],[4],[5], many potential beneficiaries in rural Romania are not utilizing them. This disconnect arises because many rural users are unfamiliar with the technologies underpinning these services. The process of installing and navigating apps, datasets and API clients, coupled with the growing complexity of technical interfaces such as satellite and radar projections or multimodel forecasts can be daunting. Therefore, converting complex information into a clear, accurate, and specific 84112 2025 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support FIGURE 1. Awareness about smart farming applications [2]. text-based format is akey obstacle to the widespread adoption of modern smart farming practices in the near future. As part of a broader effort to support rural communities [6], our objective was to develop a system capable of accessing expert data from diverse third-party sources, reasoning about such data, and delivering actionable conclusions and recommendations. By leveraging the power of Large Language Models (LLMs), which excel in processing, understanding, and generating human language, we aim to provide rural communities with easy access to expert knowledge and insights without requiring deep technical expertise. LLMs, built on the groundbreaking Transformer architecture [7], have become state-of-the-art tools, attracting unprecedented attention and research effort, and are now widely adopted across diverse applications. They power chatbots capable of human-like conversations and improve programming tools with features such as code generation and explanation. The latest language models, with hundreds of billions of parameters, demonstrate advanced reasoning abilities and can perform logical reasoning to some extent. However, language models are not out-of the box solutions for many applications including ours. LLMs suffer from train-data dependency - they are constrained by the data they have been trained on, thus limiting their reasoning abilities, as illustrated in [8];in-context reasoning limitations - they struggle with fundamental reasoning tasks such as calculating minimum, maximum, or average values across data series, and often fail to decide when to perform factual checks instead of performing text generation; prohibitive training costs - they demand extensive memory and GPU resources, with fine-tuning costs starting at hundreds of dollars and escalating to hundreds of thousands for training from scratch. In this paper, we introduce Agora, an advanced system powered by language models specifically designed to address queries related to weather forecasts, climate, and crop management, while effectively overcoming the challenges mentioned earlier. Agora can invoke third-party APIs to improve its answers. Moreover, it can efficiently orchestrate multiple API calls, allowing it to trigger additional requests when needed to retrieve supplementary data, thereby ensuring more comprehensive and accurate answers to user queries. For instance, to answer a question such as: Can the cultivation of tomatoes thrive in the climatic conditions around the city of Piteşti, Romania? Agora recognized that simply retrieving data on optimal sowing conditions for tomatoes is insufficient. To deliver a precise response, it also initiates a secondary call to obtain historical climate data for Piteşti. In another example: What is the coolest month in the city of Sibiu? Agora retrieves monthly temperature averages for the given location and sends them to an aggregation API to perform the min operation over the given interval. In addition to answering questions accurately, Agora is highly modular. Instead of relying on a single large-scale model, it utilizes smaller domain-specific language models, each tailored to handle specific types of data. These models are coordinated by a larger interviewer model, that integrates and processes their outputs. The domain-specific models that we trained focused on agriculture, weather, and climate data. However, Agora is flexible and can be deployed for a wide range of tasks, with the ability to scale and incorporate additional domains as needed. Most importantly, by distributing expertise across multiple models, Agora is scallable. Using multiple models with a relatively low number of parameters, it supports training and inference that can be efficiently performed on commodity GPUs. Agora addresses a key gap in current research, which is largely focused on training massive, general-purpose models with broad knowledge and billions of parameters. These large-scale systems—often proprietary—are out of reach for academia, small businesses, and other resource-constrained users. They come with usage costs and lock users into closed ecosystems. In contrast, we believe the future lies in smaller, high-expertise models tailored to specific domains. These models can be trained and deployed on modest hardware, are easier to evaluate, and rely on expert data to produce reliable, targeted responses. Agora is built with this vision in mind. In Section II we outline the challenges encountered during the development of Agora and explain how its design effectively addresses each of these issues. We explain Agora’s implementation in Section III. In Section IV we discuss model training and evaluation. In Section Vwe provide an in-depth review of the existing approaches that follow similar directions. In Section VI we discuss limitations and finally, in Section VII we present conclusions and future work. II. DESIGN CHALLENGES WHEN BUILDING AGORA To begin designing Agora, we first explored the existing domain-specific models relevant to our areas of interest. The development of large language models (LLMs) has made significant progress in fields such as agriculture [9] and climate science [4],[10]. Such models were created by opting for a base language model, which may be either open-source VOLUME 13, 2025 84113
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support (e.g., the LLaMA [11] family) or proprietary (e.g., the GPT-4 family). Next, two main approaches are employed. The first relies on (i) prompt engineering, where the model is guided by instructions and/or examples to answer domain-specific questions, as seen in [4]. This may involve augmenting prompts with relevant dataset fragments and explanatory information to improve response accuracy. Alternatively, (ii) the model can be fine-tuned using a dataset of example queries and the corresponding target answers. Our experience has shown that prompt engineering tends to perform well on large models (with more than 100 billion parameters). Fine-tuning can yield good results with much smaller models (e.g., the LLaMA2 7B model [11]). However, both methods are inherently constrained by their reliance on data encountered during training, which limit their reasoning abilities to that information. Retrieval-Augmented Generation (RAG) [12] is a common approach for addressing this limitation. RAG operates in two phases: first, a retriever analyzes the user query and retrieves relevant data from a database or external source. Second, these data are appended to the user query as context, which the LLM then uses to generate its response. We found that typical retrieval components in RAG are often simple, typically employing basic search methods such as BM25 [13], which use keyword matching and do not scale well for complex numerical datasets. For example, considering the query from the previous section: ‘‘What is the coolest month in the city of Sibiu?’’, a BM25based retriever struggled to identify the specific dataset slice necessary to answer this query. At best, the entire dataset can be appended to a user query. However, this approach is impractical because language models have a limited context window (e.g., LLaMA 2’s 4096-token limit is equivalent to approximately 3000 words). To address these issues, we examined methods that support invoking external APIs during text generation, such as [8]. The authors proposed a language model that has been trained to selectively pause text generation, invoke an API and incorporate the response into ongoing text generation. This methodology which includes dataset creation and the resulting model, is termed Toolformer. Toolformer leverages the inherent reasoning abilities of the small-size 6B-parameter GPT-J model [14] to automatically annotate an existing dataset with API calls. Subsequently, the dataset is used to fine-tune the same model to teach it where to perform API calls and how to select the appropriate parameters. Further details are presented in Section III. By experimenting with Toolformer we identified five essential criteria that were essential in creating Agora: 1) API-call support: Our system needs to learn when and how to perform external API-invocations in order to support its answer-generation process; 2) API-call orchestration: In many cases, API responses may influence the construction of subsequent APIcalls. For instance, for many user queries, we first gather data regarding monthly precipitation, then it, and FIGURE 2. Agora architecture. based on the resulting value search for crops whose optimal conditions match those precipitation averages. The system needs to support such dependencies between API-calls. 3) API scalability: New data sources exposed via new APIs should be easy to integrate without requiring an entire retraining of the system 4) System scalability: Training the entire system should be achievable on commodity GPUs, with minimal costs. 5) Open-source reliance: The system should exclusively utilize open-source models for inference, allowing for easy adaptation and implementation in various environments. Of these five features, Toolformer satisfies only 1., 4. and 5. criteria. In addition, we found that criteria 2. and 3. conflict when attempting to create a single language model. As the type of orchestration between API-calls becomes more elaborate, smaller models tend to under-perform, and the data size required for training increases exponentially with the number of APIs. Therefore, in order to build Agora, we experimented with the idea of creating a multi-model system, that separates the tasks of properly addressing APIcalls and responses, from the task of API-call orchestration. III. AGORA ARCHITECTURE The Agora was a central public space in ancient Greece where citizens gathered to discuss public matters and make decisions. Our system is inspired by Agora’s concept. It consists of a collection of models E1, . . . , Enhenceforth called experts together with a special model MI, which we term the interviewer model. Whenever a user query is received, MI, as well as all the experts generate one token at a time simultaneously on the same input prompt. At specific points during this process, when necessary, MI delegates control to an expert model Eito continue the response. This process is illustrated in Figure 2, which shows the Interviewer model coordinating with three expert models. Each expert is trained to invoke API calls to retrieve data from third-party repositories, databases, or other external sources. For complex APIs, we train a dedicated 84114 VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support FIGURE 3. Agora architecture. model to handle those interactions. In contrast, for simpler APIs that can be learned from examples, a single expert is trained to manage multiple call types. Once an expert model Eicompletes its task, control returns to MI, which resumes text generation until another expert is required or the response is complete. To clarify this text generation workflow, we provide an example in Figure 4. The user question is shown in gray. The text generated by MI is illustrated in blue, whereas that generated with the aid of expert E1is shown in orange. E1is the weather forecast and climate model. To properly address this question, MI notices that temperature information is necessary and uses E1 to generate the appropriate text. Figure 3illustrates how the Interviewer model orchestrates API interactions between two expert models. In Step 1, the Interviewer delegates the initial part of the response to the first expert. This expert issues an API call and uses the retrieved data to generate its contribution (Step 2). In Step 4, the Interviewer invokes a second expert to continue the response. This second expert relies on the output of the first to construct a valid API call— highlighting a dependency between the two experts. Such dependencies can span multiple experts and occur in arbitrary sequences. A concrete example of this interaction is shown in Figure5, which involves two experts, E1and E2.E1specializes in climate-related data, while E2focuses on crop-related knowledge. As before, the question is shown in grey. Since climate data is needed first, the Interviewer starts text generation with E1. When crop-specific information becomes relevant, E2is engaged to continue the response. Finally, MI synthesizes the complete answer based on the outputs from both experts. Next we discuss how each type of model (interviewer/experts) has been built. A. EXPERT MODELS Expert models are trained independently and their task is to perform accurate calls for a designated API (or APIs). In this work we consider four APIs: (i) current date - which is necessary when queries use time expressions FIGURE 4. Generating answers with Agora. relative to the current moment in time, such as now,today, tomorrow, (ii) weather & climate - which retrieves weather forecasts as well as queries related to precipitation, wind and temperature records for up to 10 years in the past, (iii) crops - which retrieves agricultural data regarding to optimal planting periods and precipitation requirements, from a collection of data-sources and (iv) aggregation - which computes maximal, minimal and average values over lists of integers. We trained three experts to handle these four APIs. Each expert model was trained via finetuning, starting from the same base LLaMA-3-8B-Instruct model [15]. LLaMA-3-8B-Instruct is already fine-tuned for instructionfollowing [15], making it a strong foundation for task-specific dialogue and assistant behavior. This means better alignment and usability out of the box—especially for structured, guided outputs like ours. Despite its smaller size, it matches or outperforms larger models such as LLaMA 2–13B and GPT3.5, while remaining lightweight enough to run on a single commodity GPU [15]. Just as importantly, LLaMA-3 is opensource, giving us the flexibility to fine-tune and deploy the model privately, ensuring compliance with sensitive data requirements. The API call representation is inspired from [8]. For a given API a, a call is a pair (fa,i1. . . , in) where fadesignates the call name, while i1, . . . , ikdesignate the kparameters of the call. API calls are encoded as arbitrary word sequences identified only with special tokens marking the start and end of a call, as illustrated below: <API_CALLa>fa(i1, . . . , in)</API_CALLa> Thus, during text generation, whenever the expert model predicts <API_CALLa>as the next token to be generated, the following steps occur: •The model continues generation internally until the complete sequence of the call, ending when </API_CALLa>is produced; •The actual call fa(i1, . . . , in) of API ais performed and the response sequence ris retrieved, and appended to the current context; Through fine-tuning, expert models Eilearn how to construct valid API calls, and by carefully designing the training data, they also learn when API-calls should take place. In Section IV we go into more detail into our approach for training experts as well as dataset instrumentation. More details regarding the API syntax can be found in Appendix A. VOLUME 13, 2025 84115
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support FIGURE 5. Agora with two experts. B. LIMITATIONS WHEN BUILDING ONE EXPERT ACROSS MULTIPLE APIS Using careful instrumentation of the training dataset [8], a model can be trained to perform calls of multiple APIs. However, this approach does not scale well when the number nand the complexity of different supported APIs increases. Suppose ta1is the number of fine-tuning examples that a model needs to train to accurately learn one API type (e.g. the Weather API). If a completely new and independent API type (say Crops) is to be added, another tbexamples would suffice. However, if we would like to capture possible dependencies between the relative position of one API call with respect to the other in the text (e.g. many Crop calls are likely to have a Weather call prior to them, and Crop answers may influence how the Weather calls are being performed), then the training data needs to have entries where one call occurs in relation to another (order O(ta×tb) entries), as well as entries where calls occur independently. Hence, the complete training dataset for learning the two APIs with dependencies between them is as follows: O(ta·tb+ta+tb) (1) This illustrates the API-call scalability problem highlighted in Section II. As the number nof APIs increases, the dataset size required to learn them increases exponentially with respect to n. Moreover, the dependencies between API calls are not solely positional. Consider the example in Fig. 5. To assess whether peach trees are suitable for planting in the area of Constanta, we need to have knowledge about precipitation averages in Constanta from the Weather API. More generally, a response rfrom an API call may influence the manner in which another subsequent call is performed. Training data size is not the only limitation - as the number of fine-tuning 1Our experience shows that ta≤10000. More details in Section IV. examples increases, small models such as LLaMA-3-8b are no longer capable of sustaining such intricate correlations, and their association performance degrades. Thus, our solution replaces the ‘‘single model’’ scenario as well as the massive dataset required for fine-tuning, with several smaller expert models, each adapted to a given domain via a straightforward fine-tuning task. To correlate the text generated by such experts, we require another dedicated model. C. THE INTERVIEWER MODEL The interviewer model (MI) was trained specifically to handle the task of expert model moderation. More specifically, MI is responsible for deciding when an expert LLM should be used to generate parts of the answer, as well as to integrate these parts in the overall answer, whenever necessary. We implemented several options to achieve this moderation. The first, termed sequential control, is shown in Fig. 6. Informally, in this setup, ’’the interviewer asks an expert‘‘.MI generates token sequences x1, . . . , xn. Next, it decided to use expert E1to continue the answer. This is achieved by generating a special token CTRLi(CTRL1in Fig. 6). Expert model E1uses x1, . . . , xnas the initial context (i.e. the sequence of previously-generated tokens). It generates tokens yn+2, . . . , yn+m, followed by a STOP token that returns the control to MI. FIGURE 6. Sequential control passing between models. When formulating a question related to plant cultivation MI might accurately decide to allow the Crops expert model 84116 VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support to continue generation. However, our initial experiments showed that MI does not always exhibit context sensitivity. Oftentimes, the expert is more capable of assessing the context and determining when to start ’’talking‘‘ by generating its own CTRLitoken. Therefore, we also included another scenario, parallel generation, illustrated in Figure 7. Here, the expert modelsteps and interrupts the Interviewer. To achieve this, we have all the expert models continuously generate tokens. As before, token ykis generated by an expert based on the history of tokens x1, . . . , xk−1which represent the current context. When an expert generates a token CTRLi, it preempts the Interviewer. All subsequent tokens up until STOP are part of the user’s answer. In Figure 7, all tokens that the user does not see are indicated by dashed lines. This situation often occurs when an expert decides to perform an API call. We observed that in almost all scenarios such a decision is contextually correct and should be prioritized over the text generated by the Interviewer. FIGURE 7. Parallel generation with multiple models. Finally, our experience has also shown that the expertgenerated text needs to be processed before output. For this reason we introduced Parallel generation with moderation (see Figure 8). Here, we apply a text transformation function gto the sequence of tokens yn+2, . . . , yn+mgenerated by an expert. FIGURE 8. Parallel generation with moderation. If gis the identity function (g(s)=s), then the text generated by the expert is unmodified. If g(·)=ϵ(the empty string), then the entire sequence generated by the expert is effectively hidden from the user. However, this sequence will still be part of the context which MI uses to generate text. In Figure 8, note that token xn+m+2is generated based on the sequence of tokens x1,x2, . . . , xn,yn+2, . . . , yn+m, to which the expert E1contributed with yn+2, . . . , yn+m. This form of hiding tokens is particularly helpful in dealing with APIcalls that produce tabular data as a response. For example, the Weather expert is trained to generate such API calls. We want this type of information to be in the context of the Interviewer to draw a conclusion, but not be explicit for the user. We illustrate this situation using the example shown in Fig. 9. After starting a sentence, MI decides to yield the context to the Current date expert. The control tokens were omitted for brevity. The expert performs a call to retrieve the current date. This call takes no parameters. The actual call will be hidden from the user (illustrated with white boxes in Figure 9), but kept in the context window. After the expert has finished the call, control is resumed by MI which switches to the Weather expert. At this point, the current context contains the location as well as date, which will be used by the second expert model to construct its weather-related API-call. Once the call has finished, text generation control returns to MI, which uses the time and forecast information to produce its conclusion. This example highlights several key traits of our approach: •We use API calls not only to generate text-answers, but also to add contextual information flexibly. By using g, we keep this information in the context so that the Interviewer can reason about it and at the same time hide it from the user. This is akin to dynamic generation of query-dependent prompts. •The example in Figure 9also illustrates the dependency between the answer given by the Current date API, and the construction of the subsequent Weather call. In practice we find many such dependencies, sometimes cascading over three or more calls. For instance, we might need to fetch the current date, based on it identify weather-related data, then perform an average over the result and finally fetch crop-related information based on that average. IV. TRAINING AND EVALUATING AGORA Agora was built as a result of an iterative refinement process, in which we explored different designs to achieve our goals. As mentioned in Section II, these were: (i) API-call support, (ii) the ability to support dependencies between calls i.e. APIcall orchestration, (iii) the ability to easily integrate new APIs - API scalability, (iv) the ability to train the entire system on commodity GPUs (system scalability) and (v) open-source reliance. A. FINE-TUNING A SINGLE MODEL FOR API-CALL SUPPORT Our first step was to apply the Toolformer methodology [8], which consists of creating a single model fine-tuned to address our different types of API-calls. Toolformer selectively inserts API calls into a large dataset, enabling the model to learn when and how to generate them. 1) DATASET Our methodology diverged from that in [8] because of the unavailability of such an existing dataset. Agricultural texts and datasets generally lack the sufficient temporal and spatial data required for accurate weather and climate API calls, with relevant examples occurring too infrequently for queries on optimal sowing conditions. VOLUME 13, 2025 84117
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support FIGURE 9. Illustrating the usage of the function g. To build training data for each of our APIs—Current Date,Weather & Climate,Crops, and Aggregation—we used GPT-4o to generate separate datasets of question–answer pairs annotated with API calls. Ensuring diversity in these datasets was essential for the expert models to generalize effectively. We began by writing hand-crafted question templates that varied in complexity, involving between one and three experts, and capturing different types of dependencies across domains. GPT-4o was then used to produce semantic variations of these templates, using a range of prompts to encourage diverse language styles and phrasings. Finally, GPT-4o instantiated each template by filling in specific details—such as locations, crop types, or climate conditions—to create fully concrete questions. Using the selected parameters, we constructed a correct set of API calls, queried the relevant data and asked GPT-4o to generate an answer, together with the necessary API calls, resulting in a complete question–answer pair. The prompts used for each API are provided in Appendix B-A, along with more details on the additional processing performed on the synthetically generated examples. 2) TRAINING We chose a relatively small, state-of-the-art open-source LLM, specifically LLaMA-3-8B-Instruct, because the LLaMA-3 family consistently outperforms other models of similar size. We performed fine-tuning using LoRA [16] and QLoRA [17] adapters, to fit the memory GPU limitations. The exact hyper-parameters that we used are presented in Appendix B-C. To achieve objective (iv), we selected the NVIDIA AD102 GeForce RTX 4090 GPU, a high-performance, cost-effective, and readily available piece of hardware on the market. We utilized three such GPUs each with 24,576 MB of available VRAM, accessed via the CUDA API. 3) EVALUATION We began by identifying the main categories of userrelevant questions, with a focus on complex queries that require integrating information across multiple domains. This analysis included mapping out all possible dependency relationships between API calls. Our findings show that weather and climate data often serve as foundational inputs, with crop-related queries typically depending on both of these, as well as the current date. Aggregation API calls tend to rely on the results of prior API responses, such as those from weather, climate, or crop services. Based on these insights, we designed a comprehensive set of question templates that systematically reflect the full range of possible dependencies. Using an approach similar to that described in Section IV-A1, we generated diverse questions grounded in these templates. We then used GPT-4o—guided by the prompts described in the Appendix—to produce user queries evenly distributed across the identified categories, including examples involving only a single API call. These queries are distinct from those used in training (see Section IV-A1) and were not seen by the model during fine-tuning. We subjected each of the systems under scrutiny to these questions and manually graded the answers on a scale of 1 to 5. Grades 3 - 5 are assigned to answers where all API calls are correct or only part of them, but the overall answer and underlying reasoning are valid. Grades 1 and 2 refer to 84118 VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support answers in which API calls are invalid and dependencies are misidentified. For each query included in our evaluation, we have deterministic knowledge of both the specific API call(s) that need to be invoked and the dependencies between them. This predefined structure eliminates any ambiguity in the evaluation process, allowing us to assess each response with complete certainty and ensuring a rigorous and objective accuracy analysis. The results are shown in Figure 10.Accuracy refers to the percentage of scored grades from three to five out of the total, while Percentage of perfect answers refers to those grades of five out of the total. FIGURE 10. Agora performance compared to other approaches. Our first experiments consist in applying the Toolformer [8] methodology on our training set, and with our language model choice. Although the results were promising (first row in Figure 10), and better results could be obtained by increasing the size of the dataset or that of the model, our observation was that the single model lacked the ability to associate multiple API calls, even though it had knowledge of each available API. Our attempts at instrumenting the dataset to capture API call dependencies revealed the APIcall scalability problem discussed in Section III-B. B. AGORA - INTERVIEWER WITH MULTIPLE EXPERTS To enhance performance and add modularity (i.e. our objective (iii)), we introduced the Agora system, which distributes the API-call generation tasks across expert models, dependency handling and generating conclusions - to the Interviewer model. We used the same strategy as before to generate the datasets for the training of experts. In our first Agora iteration, we used the same 8B base model for experts and the Interviewer. The training procedure for each expert model follows the details in Section IV-A2. Table 1outlines the APIs managed by each expert. For the interviewer model we used a system prompt that contained a short description of each API, a grammar showcasing their syntax and few-shot examples of their use. This system prompt, which can be found in the Appendix, enables the interviewer to understand how different APIs can TABLE 1. The number of examples used to train each expert model. be linked or combined when a user query lacks sufficient information for a single API call. Moreover, using a system prompt ensures the modularity of the Agora system because adding or removing one or more APIs requires only updating this prompt. We observed a significant improvement in performance and Agora successfully chained multiple API calls, thus achieving objective (ii). The results are shown in Figure 10 (second column). Our subsequent objective was to further increase the system performance. We experimented with a 70B interviewer model along with 8B expert models. This led to strong performance results, as shown in Figure 10 (third column), because the larger model’s enhanced reasoning abilities enabled it to better understand the available APIs and combine them. Simultaneously, new expert models can be trained independently, with no intervention required for the existing ones, and with minimal API descriptions that need to be added to the Interviewer’s system prompt, thus achieving objective (iii). However, running inference on the 70B interviewer model with an NVIDIA AD102 is not feasible because of insufficient GPU memory, requiring us to switch to a more capable A100 with 80GB of memory. 1) MEMORY CONSTRAINTS DURING INFERENCE Using LoRA [16] and QLoRA [17] adapters during finetuning, each expert can be formed by combining a base model with a small adapter, which can be easily plugged in or removed as needed (see Figure 11). This strategy significantly enhances memory efficiency and reduced GPU memory requirements by up to three times, as reported in [16] compared with traditional finetuning methods. This approach offers several advantages for Agora’s architecture, which, in principle, is designed for n+1 models, where nrepresents the number of experts together with the Interviewer. Because all n+1 models must run for each generated token, we can achieve this by loading the n+1 base models on a sufficiently large GPU (or multiple GPUs). However, we can achieve better memory utilization by loading one base model on a GPU, together with kadapters, one for each expert. This means that on each GPU, we can run the inference from kdifferent experts sequentially by simply VOLUME 13, 2025 84119
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support swapping out the corresponding LoRA adapters, as illustrated in Figure 11. This method reduces the number of base models that need to be loaded simultaneously, thus saving computational resources at the expense of increased time owing to sequential loading. Additionally, when the interviewer model is significantly larger than the expert models, we can load the interviewer onto a dedicated GPU. The experts which are much smaller than the Interviewer, can be evenly distributed across the remaining GPUs. FIGURE 11. Agora performance compared to other approaches. Figure 11 showcases several scenarios used during our evaluation. In Scenario 1, four smaller GPUs run inference in parallel, an arrangement we used to assess the initial version of Agora. As the evaluation moved to the larger 70B Interviewer model, we transitioned to more powerful 80GB A100 GPUs. Scenarios 2 and 3 demonstrate how base models and adapters are distributed across available memory. In Scenario 3, for example, a single GPU holds one base model and two adapters, enabling inference for two expert models, E2and E3, which generate tokens sequentially. The adapter-switching time is approximately 20 times shorter than the time needed to generate a token. Despite hardware limitations, Scenarios 2 and 3 illustrate that our multi-model system can still operate efficiently, though with reduced performance due to sequential inference. C. AGORA COMPRESSION The primary limitation of the Agora system, as described in Section IV-B, is the substantial resource demands necessary to achieve optimal performance. This is mainly because of the need to load both the LLaMA-3-70B-Instruct model (Interviewer) and the experts’ base model, LLaMA-3-8BInstruct, resulting in excessive memory consumption. The memory overhead surpassed the capabilities of NVIDIA AD102 GPUs alone. To mitigate this, we developed a new, standalone model, which we trained using the responses provided by Agora. We call this model a ‘‘compression model’’, because unlike Agora, it is a single language model, but is able to reproduce, and even enhance the performance of Agora. To achieve this, we proceeded as follows: •We started with a dataset of questions and used Agora to generate answers. We employed an 70B interviewer and 8B expert models. This dataset includes examples demonstrating individual API usage as well as examples of chained API calls. •Agora has great but not perfect accuracy, hence we carefully filtered it to retain only question-answer pairs with highly accurate multi-API call examples. The final dataset contains 17,000 such entries. •We finetuned a single LLaMA-3-8B-Instruct model with this dataset, resulting our compression model. The pipeline for training the Agora compression model, relying on the previous stages we have descripted, is illustrated in Figure 12. FIGURE 12. All steps required to build the Agora compression model. While this model demonstrated the best performance overall, even slightly surpassing Agora with the 70B interviewer (Fig. 10),there were some trade-offs. The compression model: •cannot be fine-tuned independently of Agora, as there is no existing dataset suitable for this task. Instead, the required dataset must be generated directly using Agora, or a similar tool. •sacrifices modularity (i.e. objective (v)), meaning it cannot - by itself accommodate new APIs without resorting to Agora as previously described. However, we found that fine-tuning the newer versions of a compression model is an acceptable compromise for various applications. The fine-tuning time - performed on an NVIDIA A100, is under two hours. 84120 VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support APPENDIX B PROMPTS FOR BUILDING AGORA A. PROMPTS FOR BUILDING TRAINING DATASETS FOR EXPERTS Below, we present the prompts used to generate datasets for each API, along with sample outputs from the generated data for clarity. Crops and Aggregation: The question-answer pairs in the datasets for the crops and aggregation APIs were generated directly using GPT-4, without any additional filtering, by employing the prompts shown in Figures 18, 19 and 20. Each API is linked to a corresponding back-end function, which either returns the current date or queries an internal database for sowing conditions. For handling the current date and weather & climate APIs, we generated the question and the parameters for the calls using GPT-4o. We began by taking a list of Romanian cities along with their latitude and longitude, which were later needed when interacting with the weather and climate API. FIGURE 18. Prompt for generating dataset for Crops expert dataset. FIGURE 19. Prompt for generating dataset for Aggregation expert dataset. We wrote by hand 190 different time expressions that we considered plausible to be used in a query. Our API provides, by design, hourly data for intervals of up to three days, daily data for intervals of up to two months, and monthly data for longer time spans. If the user’s query includes the name of a country, it is passed as a parameter in the call. If no country is specified, the parameter value defaults to ‘‘???’’, in which case the system assumes the query refers to the largest city with the given name. In this situation, the country name is retrieved automatically from a database. We filtered out examples where data retrieval from our sources was unsuccessful and then tasked GPT-4o-mini with adjusting the current date and all related dates in each example. This ensured that the final dataset does not contain examples tied exclusively to the original creation date. In addition to fully annotated examples, we realized the need for more examples focusing solely on the weather & climate API call parameters. To boost the model’s ability to select the correct parameters for this API, we generated additional examples that contained only the question and the call parameters, omitting the final result and text reasoning about the retreived data. FIGURE 20. Prompt for generating dataset for Weather & climate expert dataset. Table 4outlines the number of examples generated for training each expert. VOLUME 13, 2025 84127
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support TABLE 4. The number of examples used to train each expert model. B. PROMPTS FOR CROSS-DOMAIN QUESTIONS To generate questions that require API calls from both the Weather & Climate and Aggregation experts, we used the prompt shown in Figure 21. We used similar prompts for all other corss-domain questions. FIGURE 21. Prompt for generating dataset for Aggregation expert dataset. Table 5shows the number of examples generated for each combination of API calls that we deemed necessary and that involve at least two experts. TABLE 5. The number of examples with combinations of API calls. C. FINE-TUNING HYPER-PARAMETERS We applied the same training settings consistently across all our fine-tuning experiments. A constant learning rate (LR) was suboptimal, so we adopted a cosine LR scheduler with a maximum value of 1e-4, which provided better convergence. Due to restricted memory, we used a batch size of 1 and trained the model for 2 epochs with a maximum input sequence length of 2048 tokens. We applied gradient clipping at 0.3 and set the weight decay to 0.1 to stabilize the training process. For the LoRA hyperparameters, we have set both alpha and rank (r) to 32. ACKNOWLEDGMENT Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. REFERENCES [1] InvestRomania. (2022). Internet Infrastructure. Accessed: Jul. 22, 2024. [Online]. Available: https://investromania.gov.ro/web/internetinfrastructure/ [2] Business Agency Association (BAA), ‘‘Jointly preparing the conditions in the agricultural and connected sectors in the BSB area for the digital transformation (BSB smart farming),’’ Dunarea de Jos, Univ. Galati, Galati, Romania, Tech. Rep., 2021. Accessed: Dec. 12, 2024. [3] S. Rezayi, Z. Liu, Z. Wu, C. Dhakal, B. Ge, C. Zhen, T. Liu, and S. Li, ‘‘AgriBERT: Knowledge-infused agricultural language models for matching food and nutrition,’’ in Proc. 31st Int. Joint Conf. Artif. Intell., Jul. 2022, pp. 5150–5156, doi: 10.24963/ijcai.2022/715. [4] N. Koldunov and T. Jung, ‘‘Local climate services for all, courtesy of large language models,’’ Commun. Earth Environ., vol. 5, no. 1, p. 13, Jan. 2024, doi: 10.1038/s43247-023-01199-1. [5] T. T. Nguyen, J. Brandstetter, A. Kapoor, J. K. Gupta, and A. Grover, ‘‘ClimaX: A foundation model for weather and climate,’’ in Proc. 40th Int. Conf. Mach. Learn., Jan. 2023, pp. 1–14. [6] Futural Project. (2024). Futural Project–agriculture and Climate. Accessed: Sep. 14, 2024. [Online]. Available: https://futural-project.eu/ [7] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, ‘‘Attention is all you need,’’ Adv. Neural Inf. Process. Syst., vol. 30, pp. 5998–6008, Jun. 2017. [8] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomelí, L. Zettlemoyer, N. Cancedda, and T. Scialom, ‘‘Toolformer: Language models can teach themselves to use tools,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 36, Jan. 2024, pp. 1–16. [9] (2024). KissanAI. Accessed: Jul. 22, 2024. [Online]. Available: https:// kissan.ai/ [10] B. Silva, L. Nunes, R. Estevão, V. Aski, and R. Chandra, ‘‘GPT-4 as an agronomist assistant? Answering agriculture exams using large language models,’’ 2023, arXiv:2310.06225. [11] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, ‘‘LLaMA: Open and efficient foundation language models,’’ 2023, arXiv:2302.13971. [12] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, ‘‘Retrievalaugmented generation for knowledge-intensive NLP tasks,’’ in Proc. Adv. Neural Inf. Process. Syst., Jan. 2020, pp. 9459–9474. [13] S. Robertson and H. Zaragoza, ‘‘The probabilistic relevance framework: BM25 and beyond,’’ Found. Trends Inf. Retr., vol. 3, no. 4, pp. 333–389, 2009. [14] B. Wang and A. Komatsuzaki. (2021). Gpt-j-6b: A 6 Billion Parameter Autoregressive Language Model. Accessed: Aug. 12, 2023. [Online]. Available: https://github.com/kingoflolz/mesh-transformer-jax [15] A. Alford. (2024). Meta Releases Llama 3 Open-Source LLM. Accessed: May. 7, 2024. [Online]. Available: https://www.infoq.vcom/news/2024/ 05/meta-llama-3/ [16] J. E. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, and W. Chen, ‘‘LoRA: Low-rank adaptation of large language models,’’ in Proc. Int. Conf. Learn. Represent., Jan. 2021, pp. 1–16. [17] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, ‘‘QLoRA: Efficient finetuning of quantized LLMs,’’ in Proc. Adv. Neural Inf. Process. Syst., Jan. 2023, pp. 10088–10115. [18] (2024). Agri1. Accessed: Sep. 4, 2024. [Online]. Available: https://www. agri1.ai/en/ [19] S. Chen, G. Long, J. Jiang, D. Liu, and C. Zhang, ‘‘Foundation models for weather and climate data understanding: A comprehensive survey,’’ 2023, arXiv:2312.03014. [20] J. Devlin, M. Chang, K. Lee, and K. Toutanova, ‘‘BERT: Pre-training of deep bidirectional transformers for language understanding,’’ in Proc. NaacL-HLT, Minneapolis, MN, USA, Jan. 2019, vol. 1, no. 2, pp. 4171–4186. 84128 VOLUME 13, 2025
A. Udrescu, D.-M. Popovici: Agora: A Distributed Language Model Framework With API-Call Support [21] Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang, ‘‘Retrieval-augmented generation for large language models: A survey,’’ 2023, arXiv:2312.10997. [22] L. Yuan, Y. Chen, X. Wang, Y. Fung, H. Peng, and H. Ji, ‘‘CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets,’’ in Proc. 12th Int. Conf. Learn. Represent., Jan. 2023, pp. 1–16. [23] Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun, ‘‘ToolLLM: Facilitating large language models to master 16000+ real-world Apis,’’ in Proc. 12th Int. Conf. Learn. Represent., Jan. 2024, pp. 1–19. [Online]. Available: https://openreview.net/forum?id= dHng2O0Jjr [24] S. Gao, Z. Shi, M. Zhu, B. Fang, X. Xin, P. Ren, Z. Chen, J. Ma, and Z. Ren, ‘‘Confucius: Iterative tool learning from introspection feedback by easy-todifficult curriculum,’’ in Proc. AAAI Conf. Artif. Intell., Mar. 2024, vol. 38, no. 16, pp. 18030–18038. [25] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez, ‘‘Gorilla: Large language model connected with massive Apis,’’ 2023, arXiv:2305.15334. [26] R. Yang, S. Lin, Y. Li, S. Zhao, Y. Ge, X. Li, and Y. Shan, ‘‘GPT4Tools: Teaching large language model to use tools via self-instruction,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 36, Jan. 2023, pp. 1–17. [27] W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing. (2023). Vicuna: An Open-source Chatbot Impressing Gpt-4 With 90% Chatgpt Quality. Accessed: Sep. 12, 2023. [Online]. Available: https://lmsys.org/ blog/2023-03-30-vicuna/ [28] Q. Tang, Z. Deng, H. Lin, X. Han, Q. Liang, B. Cao, and L. Sun, ‘‘ToolAlpaca: Generalized tool learning for language models with 3000 simulated cases,’’ 2023, arXiv:2306.05301. [29] C. Qian, C. Han, Y. R. Fung, Y. Qin, Z. Liu, and H. Ji, ‘‘CREATOR: Tool creation for disentangling abstract and concrete reasoning of large language models,’’ 2023, arXiv:2305.14318. [30] T. Cai, X. Wang, T. Ma, X. Chen, and D. Zhou, ‘‘Large language models as tool makers,’’ in Proc. 12th Int. Conf. Learn. Represent., Jan. 2023, pp. 1–14. [31] S. Hao, T. Liu, Z. Wang, and Z. Hu, ‘‘ToolkenGPT: Augmenting frozen language models with massive tools via tool embeddings,’’ in Proc. 37th Conf. Neural Inf. Process. Syst., 2023, pp. 1–13. [Online]. Available: https://openreview.net/forum?id=BHXsb69bSx [32] A. Srivastava et al., ‘‘Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,’’ in Proc. Trans. Mach. Learn. Res., Jan. 2022, pp. 1–95. [33] G. Mialon, R. Dessì, M. Lomelí, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Çelikyilmaz, É. Grave, Y. LeCun, and T. Scialom, ‘‘Augmented language models: A survey,’’ Trans. Mach. Learn. Res., Jan. 2023, pp. 1–33. [34] Y. Talebirad and A. Nadiri, ‘‘Multi-agent collaboration: Harnessing the power of intelligent LLM agents,’’ 2023, arXiv:2306.03314. [35] S. Zejiang Shen, H. Lang, B. Wang, Y. Kim, and D. Sontag, ‘‘Learning to decode collaboratively with multiple language models,’’ 2024, arXiv:2403.03870. [36] Z. Chai, G. Wang, J. Su, T. Zhang, X. Huang, X. Wang, J. Xu, J. Yuan, H. Yang, F. Wu, and Y. Yang, ‘‘An expert is worth one token: Synergizing multiple expert LLMs as generalist via expert token routing,’’ 2024, arXiv:2403.16854. [37] X. Xu, M. Li, C. Tao, T. Shen, R. Cheng, J. Li, C. Xu, D. Tao, and T. Zhou, ‘‘A survey on knowledge distillation of large language models,’’ 2024, arXiv:2402.13116. [38] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Welleck, B. P. Majumder, S. Gupta, A. Yazdanbakhsh, and P. Clark, ‘‘Self-refine: Iterative refinement with self-feedback,’’ in Proc. Adv. Neural Inf. Process. Syst., vol. 36, Jan. 2023, pp. 1–24. [39] Y. Dong, R. Mu, G. Jin, Y. Qi, J. Hu, X. Zhao, J. Meng, W. Ruan, and X. Huang, ‘‘Building guardrails for large language models,’’ 2024, arXiv:2402.01822. ALEXANDRA UDRESCU received the engineering degree in computer science from the Faculty of Automatic Control and Computers, National University of Science and Technology POLITEHNICA Bucharest, where she is currently pursuing the master’s degree. She is a Computer Scientist specializing in algorithms, formal methods, and machine learning. DAN-MATEI POPOVICI received the Ph.D. degree from the National University of Science and Technology POLITEHNICA Bucharest, in 2012. He was a Research Fellow with the Clausthal University of Technology and ICUB (University of Bucharest’s Research Institute). He is currently an Associate Professor with the Computer Science Department, National University of Science and Technology POLITEHNICA Bucharest. His research interests include formal verification techniques for computer networks, computing research education, and NLP using language models. VOLUME 13, 2025 84129