scieee AI-readable full text Open interactive document viewer

ML Field Planner with TAPIS: Configuring and Analyzing AI Models for Animal Ecology - BYOP

Vallabhajosyula, Manikya Swathi; Gautam Gururaj Molakalmuru, Molakalmuru; Agulugaha Isuru, Gamage; Karthikeyan, Neelesh; Clifford, Nick; Khuvis, Samuel; Christian, Garcia; Freeman, Nathan; Stubbs, Joe; Plale, Beth; Ramnath, Rajiv

Abstract

Artificial intelligence (AI) is being increasingly applied across various domains, including animal ecology, genome sequencing, digital agriculture, and biomedical science. The availability of high-quality benchmark datasets has enabled researchers to apply advanced AI models—particularly in computer vision—where labeled imagery from drones or camera traps can address domain-specific challenges. In ecology, image-based workflows often utilize state-of-the-art vision models, such as YOLO, for detection and classification. Perspectives in Machine Learning for Wildlife Conservation [1] demonstrates how MegaDetector (built on YOLO) can filter blank images, reducing human review to under 15% and conserving storage. Modern cyberinfrastructure now integrates both powerful central computing resources and capable edge devices, such as camera traps and drones, that can run AI models in inference mode. Yet, selecting and testing models—balancing accuracy, efficiency, and deployment requirements remains an essential but time-consuming task.The core challenge for domain scientists is navigating the rapidly growing choice of models, the massive volume of images generated in varied environments, and the diversity of data resolutions, noise levels, and ecosystem conditions. Model selection must consider not only performance metrics, such as precision and recall, but also efficiency metrics, including latency, memory footprint, power usage, and CPU/GPU utilization. For example, MegaDetector v6 is claimed to be one-sixth the size of v5, five times faster, and with 12% higher recall. Yet, our experiments show that the untrained v5 is twice as accurate as v6. When fine-tuned with a small, targeted, labeled dataset representative of deployment conditions, v6 improved, but v5 still outperformed it by 8% with minimal fine-tuning. This raises the question: how should ecologists decide which model to deploy for a given ecosystem and device?The ML Field Planner [2], accessed through the TAPIS UI portal, addresses this by enabling experiments and analyses to compare models across edge devices, evaluating both performance and efficiency before deployment. This animal ecology–inspired use case demonstrates how the portal enables ecologists to test workflows, explore trade-offs, and make informed decisions about deployment. All experiments can be run via the TAPIS UI, and results can be visualized there as well. A built-in change agent guides ecologists in discovering suitable models, comparing their results in an easy-to-read format, and selecting configurations aligned with their datasets and hardware. Additional analytics can be performed through JupyterLab notebooks, enabling researchers to examine model behavior and results on individual test images.The portal demonstration will showcase interfaces from the ML Field Planner, the CKN [3] Dashboard for camera trap analytics, the integrated chatbot, and an example inference notebook, illustrating how these tools streamline model evaluation and selection for ecological applications.References:-----------[1] Tuia, Devis, et al. "Perspectives in machine learning for wildlife conservation." Nature Communications 13.1 (2022): 792.[2] Stubbs, Joe, et al. "ML Field Planner: Analyzing and Optimizing ML Pipelines For Field Research." Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration. 2025. 1-9.[3] Withana, Sachith, and Beth Plale. "CKN: An edge AI distributed framework." 2023 IEEE 19th International Conference on e-Science (e-Science). IEEE, 2023

Full text

Title: ML Field Planner with TAPIS: Configuring and Analyzing AI Models for Animal Ecology Authors: Swathi Vallabhajosyula, Neelesh Karthikeyan, Agulugaha Isuru Gamage, Gautam Gururaj Molakalmuru, Samuel Khuvis, Christian Garcia, Nathan Freeman, Joe Stubbs, Beth Plale, Rajiv Ramnath Abstract: Artificial intelligence (AI) is being increasingly applied across various domains, including animal ecology, genome sequencing, digital agriculture, and biomedical science. The availability of high-quality benchmark datasets has enabled researchers to apply advanced AI models—particularly in computer vision— where labeled imagery from drones or camera traps can address domain-specific challenges. In ecology, imagebased workflows often utilize state-of-the-art vision models, such as YOLO, for detection and classification. Perspectives in Machine Learning for Wildlife Conservation [1] demonstrates how MegaDetector (built on YOLO) can filter blank images, reducing human review to under 15% and conserving storage. Modern cyberinfrastructure now integrates both powerful central computing resources and capable edge devices, such as camera traps and drones, that can run AI models in inference mode. Yet, selecting and testing models— balancing accuracy, efficiency, and deployment requirements remains an essential but time-consuming task. The core challenge for domain scientists is navigating the rapidly growing choice of models, the massive volume of images generated in varied environments, and the diversity of data resolutions, noise levels, and ecosystem conditions. Model selection must consider not only performance metrics, such as precision and recall, but also efficiency metrics, including latency, memory footprint, power usage, and CPU/GPU utilization. For example, MegaDetector v6 is claimed to be one-sixth the size of v5, five times faster, and with 12% higher recall. Yet, our experiments show that the untrained v5 is twice as accurate as v6. When fine-tuned with a small, targeted, labeled dataset representative of deployment conditions, v6 improved, but v5 still outperformed it by 8% with minimal fine-tuning. This raises the question: how should ecologists decide which model to deploy for a given ecosystem and device? The ML Field Planner [2], accessed through the TAPIS UI portal, addresses this by enabling experiments and analyses to compare models across edge devices, evaluating both performance and efficiency before deployment. This animal ecology–inspired use case demonstrates how the portal enables ecologists to test workflows, explore trade-offs, and make informed decisions about deployment. All experiments can be run via the TAPIS UI, and results can be visualized there as well. A built-in change agent guides ecologists in discovering suitable models, comparing their results in an easy-to-read format, and selecting configurations aligned with their datasets and hardware. Additional analytics can be performed through JupyterLab notebooks, enabling researchers to examine model behavior and results on individual test images. The accompanying figure (next page) shows sample screenshots from the ML Field Planner, the CKN [3] Dashboard for camera trap analytics, the integrated chatbot, and an example inference notebook, illustrating how these tools streamline model evaluation and selection for ecological applications. References: [1] Tuia, Devis, et al. "Perspectives in machine learning for wildlife conservation." Nature Communications 13.1 (2022): 792. [2] Stubbs, Joe, et al. "ML Field Planner: Analyzing and Optimizing ML Pipelines For Field Research." Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration. 2025. 1-9. [3] Withana, Sachith, and Beth Plale. "CKN: An edge AI distributed framework." 2023 IEEE 19th International Conference on e-Science (e-Science). IEEE, 2023