scieee AI-readable full text Open interactive document viewer

Statistics Jumpstart for Agricultural Research

Hewa Kattuge, Sameera Madusanka

Abstract

As an undergraduate, a master's student, and now a PhD candidate, I faced the same challenge many students struggle with: the problem of statistics. For years, statistics felt like a barrier—confusing calculations that created more difficulties simply because I lacked proper guidance. Once I learned the basic paths, everything changed. Deep statistical theory is essential for a PhD, but most undergraduates don’t need that level of detail. They simply need to use statistics correctly to finish their research. That realization inspired this book. Think of statistical analysis like a vehicle: to reach your research conclusion, you don’t need to build the engine—you just need to know which vehicle to take. This book is for the hardworking final-year student with limited statistical background who wants to complete their research without struggling with numbers. My aim is to show you the right “Car” to take so you can move confidently from your research question to your final conclusion.

Full text

Statistics Jumpstart for Agricultural Research A ‘No-Worries’RGuide to Finishing Your Research H. K. S. Madusanka 2025 ISBN 978-624-92891-1-6 Copyright ©2025 H. K. S. Madusanka All rights reserved. ISBN: [978-624-92891-1-6] i Other Books by H. K. S. Madusanka •Farm Machinery & Mechanization: An Engineering Logic (ISBN 978-624-92891-0-9) Acknowledgment Here is the revised acknowledgements section incorporating your additions: I wish to express my deepest gratitude to Prof. (Mrs.) A. G. B. Aruggoda, Professor in Agricultural Biology, Department of Agricultural and Plantation, The Open University of Sri Lanka, for believing in me and encouraging me to put this effort into text form. I am also sincerely thankful to The Open University of Sri Lanka for providing the academic foundation and environment that inspired this work. I also extend my gratitude to Prof. (Mrs.) D.D.M. Jayasundara and Dr. (Ms.) D.M.P.V. Dissanayaka from the Faculty of Science, Kelaniya University, for teaching and guiding me in learning and practicing statistics. Finally, my heartfelt appreciation goes to my mother, my wife, and my three wonderful children for their love, patience, and unwavering support throughout this journey. iii Preface “Statistics is the tool that builds your research conclusion. You don’t need to know how to forge the hammer, but you must know how to strike the nail correctly” As an undergraduate, a master’s student, and now a PhD candidate, I faced the same challenge many students struggle with: the problem of statistics. For years, statistics felt like a barrier—confusing calculations that created more difficulties simply because I lacked proper guidance. Once I learned the basic paths, everything changed. Deep statistical theory is essential for a PhD, but most undergraduates dont need that level of detail. They simply need to use statistics correctly to finish their research. That realization inspired this book. Think of statistical analysis like a vehicle: to reach your research conclusion, you dont need to build the engine—you just need to know which vehicle to take. This book is for the hardworking final-year student with limited statistical background who wants to complete their research without struggling with numbers. My aim is to show you the right “Car” to take so you can move confidently from your research question to your final conclusion. H. K. S. Madusanka 2025 iv Contents 1 Our “No-Worries” Philosophy 2 1.0.1 Our Philosophy: Driving a Car, Not Building an Engine ....... 2 1.0.2 How to Use This Book: Find Your “Car” ................ 3 1.0.3 The “No-Worries” Promise ........................ 4 2 Installing Your “Car” (R & RStudio) 5 2.1 Why Drive This Car? (The Specs) ....................... 5 2.1.1 Step 1: Installing the “Engine” (R) ................... 6 2.1.2 Step 2: Installing the “Dashboard” (RStudio) ............. 7 2.1.3 The 4 Panes of RStudio ......................... 7 2.2 How to “Drive” (Running Code Properly) ................... 9 2.2.1 How to “Turn the Key” (The Run Button) ............... 9 2.2.2 The Magic Symbol: Comments (#)................... 10 2.2.3 Save Your Script ............................. 11 3 Putting “Fuel” in the “Car” (Loading Your Excel Data) 12 3.0.1 The 3 Golden Rules (Cleaning Your Excel File) ............ 12 3.0.2 The Most Important Rule: The Shape (Wide vs. Long) ........ 13 3.0.3 The readxl Package: Your High-Quality “Fuel Hose” ......... 14 3.0.4 The One Line of Code (How to Load Your Data) ........... 15 3.0.5 “My Data Won’t Load!” (Simple Troubleshooting) ........... 18 3.1 The Most Important Key: The Dollar Sign ($) ................. 19 3.1.1 The Logic: The Cabinet and the Folder ................. 19 3.1.2 Why are we telling you this now? .................... 20 4 Installing Your “Upgrades” (The Common Tools) 21 4.0.1 Your “Graphing Tool” (ggplot2) ..................... 21 4.0.2 Your “Data Handling Tool” (dplyr) ................... 22 4.0.3 The Code: How to Install Them All at Once .............. 22 v CONTENTS 5 The Rules of the Road (Normality & p-values) 24 5.0.1 Rule #1: The “p-value” (Your GPS) .................. 24 5.0.2 Rule #2: The Safety Check (Normality) ................ 24 5.0.3 Rule #3: The Hierarchy of Truth .................... 28 5.0.4 Our “Driver’s Manual” is Complete! ................... 28 6 The “Survey Car” (Questionnaires & Counts) 31 6.0.1 “Use This Car If...” ............................ 31 6.0.2 The Rules of the Road: Ethics & Sensitive Data ............ 32 6.0.3 Constructing the Car: Designing the Questionnaire .......... 32 6.0.4 Avoiding the Potholes: Bias & “Soft” Statistics ............ 33 6.1 Who is Coming? (Sampling Methods) ...................... 34 6.1.1 Method 1: Simple Random Sampling (The Lottery) .......... 34 6.1.2 Method 2: Stratified Sampling (The “Fair Share”) ........... 35 6.1.3 Method 3: Convenience Sampling (The “Easy Way”) ......... 35 6.2 The “How Many?” Question (Sample Size) ................... 36 6.2.1 The Scientific Way (When you know the Total Population) . . . . . . 36 6.2.2 The “Rule of 30” ............................. 37 6.2.3 Quality over Quantity .......................... 37 6.3 The Journey: Visualizing Your Survey ..................... 37 6.3.1 The “Photo”: The Stacked Bar Chart .................. 38 6.4 The Destination: The Chi-Square Test ..................... 39 6.4.1 Wait! Why Not ANOVA? ........................ 39 6.4.2 The Code ................................. 40 6.4.3 Reading the Answer ........................... 40 6.5 Test Drive: Real-World Examples ........................ 41 6.5.1 Scenario 1: The “Adoption” Study (Education vs. Tech) ....... 41 6.5.2 Scenario 2: The “Pest Score” Study (Visual Ratings) ......... 43 6.5.3 Scenario 3: The “Consumer Market” Study (Yes/No) ......... 45 6.5.4 Scenario 4: The “Likert Scale” Study (Opinion Mining) ........ 46 6.5.5 Summary Checklist for the Survey Car ................. 47 7 The “Two-Group Car” (A vs. B) 48 7.0.1 Introduction: Are You in the Right Place? ............... 48 7.1 Putting Fuel in the Car: The Data Structure Trap ............... 50 7.2 Driving Road A: The T-Test (Parametric) ................... 51 7.2.1 The Code ................................. 51 7.2.2 Understanding the Syntax: The Tilde (~) ............... 51 vi CONTENTS 7.2.3 Reading the Output ........................... 52 7.2.4 How to Write This in Your Report ................... 52 7.3 Driving Road B: The Wilcoxon Test (Non-Parametric) ............ 54 7.3.1 How it Works (The “Ranking” Engine) ................. 54 7.3.2 The Code ................................. 54 7.3.3 Reading the Output ........................... 54 7.3.4 How to Write This in Your Report ................... 55 7.4 The “Photo” of Your Trip: The Box Plot .................... 56 7.4.1 Understanding the Box Plot ....................... 56 7.4.2 The Code ................................. 56 7.4.3 Pro Tip: Adding the Raw Points .................... 57 7.5 The “Special Road Condition”: Paired Data (The “Before & After”) . . . . . 59 7.5.1 The Difference: Independent vs. Paired ................. 59 7.5.2 The Code: Just add paired = TRUE .................. 60 7.6 Test Drive: Real-World Examples ........................ 61 7.6.1 Scenario 1: The “High-Yield” Corn (Normal Data) .......... 61 7.6.2 Scenario 2: The “Sick Chicken” (Non-Normal Data) .......... 62 7.6.3 Scenario 3: The “Before & After” Soil Test (Paired Data) ....... 63 7.6.4 Scenario 4: The “No Difference” Result (Packaging) .......... 64 7.6.5 Summary Checklist for the Two-Group Car ............... 64 8 The “Multi-Group Car” (A vs. B vs. C -CRD) 65 8.1 Is this the Right Car for You? .......................... 65 8.1.1 Use this car if. . . ............................. 65 8.1.2 STOP! The “Field Layout” Check (Crucial) .............. 65 8.2 The Danger: Why Not Just Use T-Tests? .................... 67 8.3 The Safety Check (Condensed) ......................... 68 8.3.1 The Code ................................. 68 8.3.2 The Decision ............................... 68 8.4 Driving Road A: The One-Way ANOVA .................... 68 8.4.1 The “Bus” Logic ............................. 68 8.4.2 The Code (2 Steps) ............................ 68 8.4.3 Reading the Output ........................... 69 8.5 The “Who Won?” Step: Choosing Your Post-Hoc Test ............ 69 8.5.1 Option 1: Tukey’s HSD (The “Gold Standard”) ............ 69 8.5.2 Option 2: Dunnett’s Test (The “Control Freak”) ............ 70 8.5.3 Option 3: DMRT (Duncan’s Multiple Range Test) ........... 70 8.6 How to Report the Results (The “Letters”) ................... 70 vii CONTENTS 8.6.1 What Do the Letters Mean? ....................... 71 8.6.2 Getting the Letters Automatically .................... 71 8.6.3 Reading the Output ........................... 72 8.7 Driving Road B: The Kruskal–Wallis Test ................... 72 8.7.1 The “Tank” Logic ............................. 72 8.7.2 The Code ................................. 73 8.7.3 Reading the Output ........................... 73 8.8 The “Who Won?” Step: Post-Hoc for Road B ................. 73 8.8.1 The Code ................................. 73 8.8.2 Reading the Output ........................... 74 8.9 The “Who Won?” Step: Post-Hoc for Road B (Dunn’s Test) ......... 75 8.9.1 Required package ............................. 75 8.9.2 The Code ................................. 75 8.9.3 Reading the Output ........................... 75 8.10 Getting the Numbers: Means and Standard Deviations ............ 77 8.10.1 The Code ................................. 77 8.10.2 Reading the Output ........................... 77 8.10.3 Building Your Final Thesis Table .................... 78 8.11 Test Drive: Real-World Examples ........................ 79 8.11.1 Scenario 1: The Variety Trial (Standard Approach) .......... 79 8.11.2 Scenario 2: The New Chemical (Control Focused) ........... 80 8.11.3 Scenario 3: The Strict Professor (DMRT) ................ 81 8.11.4 Scenario 4: The Insect Count (Not Normal) .............. 81 8.11.5 Summary Checklist for the Multi-Group Car .............. 82 9 The “Relationship Car” (Do X and Y Move Together?) 84 9.1 Is this the Right Car for You? .......................... 84 9.1.1 Use this car if. . . ............................. 84 9.2 The Golden Rule: Correlation =Causation ................... 84 9.2.1 The “Ice Cream & Shark” Example ................... 85 9.2.2 The Agriculture Warning ......................... 85 9.3 The Journey: Visualizing the Relationship ................... 85 9.3.1 Reading the Map ............................. 86 9.4 The Safety Check: Pearson vs. Spearman .................... 87 9.4.1 Checking the Road ............................ 87 9.5 The Destination: The Code (cor.test)..................... 87 9.5.1 Road A: Pearson (Normal Data) ..................... 87 9.5.2 Road B: Spearman (Not Normal) .................... 88 viii 1.0.3 The “No-Worries” Promise Here is our simple, personal promise to you. In this book, we will: •NEVER show you a scary mathematical formula. •ALWAYS talk to you in plain, simple English. •ONLY use examples about agriculture (yield, fertilizer, farms, surveys). •FOCUS on getting you the answer and the graph you need to finish your report. You can do this. You’ve got this. Now, let’s get your keys. 4 Chapter 2 Installing Your “Car” (R & RStudio) Welcome to the “garage”! This is the only technical part of the book, and we’ll be done in about 15 minutes. We just need to install two free programs on your computer. Think of it this way: R is the powerful “engine” for your car, and RStudio is the clean, easy-to-use “dashboard” and “steering wheel.” We install the engine first, then the dashboard. 2.1 Why Drive This Car? (The Specs) Before we get our hands dirty, you might be wondering: “Why can’t I just use Excel? Why do I have to learn this?” That is a fair question. Here is the spec sheet for your new vehicle and why it is the standard for agricultural research: •It is Free (Cost): Unlike other statistics software (like SPSS, SAS, or Minitab) which can cost hundreds of dollars for a license, R is 100% free. It is “Open Source,” meaning it belongs to the community, not a corporation. You will never lose access to your work after you graduate. •It is High-Performance: Excel is great for data entry, but it struggles with complex statistics or large datasets. R is built for heavy lifting. It can handle thousands of rows of crop data, complex experimental designs, and genetic data without crashing. •It is Versatile: R is not just for math. You can use it to create publication-quality graphs, design maps for your field trials, and even write reports (like this book!). It is an all-in-one tool. •Availability & Community: R runs on everything—Windows, Mac, and Linux. Because it is the global standard for science, if you get stuck, you can simply Google your error message and find thousands of people ready to help. 5 2.1. WHY DRIVE THIS CAR? (THE SPECS) •User-Friendliness (With a Twist): R by itself can be scary (it looks like a hacker’s screen). That is why we use RStudio. RStudio gives you the buttons, menus, and visual aids that make R friendly and approachable. In short: We are trading a bicycle (Excel) for a Ferrari (R). It takes a little longer to learn how to drive, but you will get to your destination much faster and in better style. 2.1.1 Step 1: Installing the “Engine” (R) R is the core program that does all the heavy lifting (the statistics). It’s completely free, and we get it from its official source: the Comprehensive R Archive Network (CRAN). 1. Go to the CRAN website. Open your web browser and visit: https://cran. r-project.org. See Figure 12.2. 2. Find the download link. Near the top of the page, you’ll see “Download and Install R.” Click on the link for your computer’s operating system (Windows, macOS, or Linux). 3. Download and run the installer. •For Windows: Click the “base” link, then click the big “Download R for Windows” button. •For macOS: Click the “R-x.x.x.pkg” link that matches your version of macOS. 4. Install the program. Once the file is downloaded, run it and click “Next” through all the default options. You dont need to change any settings. Figure 2.1: The official CRAN homepage (https://cran.r-project.org/). Thats it—the “engine” is now installed! Youll probably never open this program directly. We will always use RStudio, which well install next. 6 2.1. WHY DRIVE THIS CAR? (THE SPECS) 2.1.2 Step 2: Installing the “Dashboard” (RStudio) RStudio is the program we will open every time. Its our “dashboard.” Think of it as a clean, friendly interface that sits on top of R and makes it a hundred times easier to use. 1. Go to the RStudio website. Open your web browser and visit: https://posit.co/ download/rstudio-desktop/. (The company is called Posit, but the program itself is still called RStudio.) 2. Find the free version. Youll see a few different editions. You only need the one called “RStudio Desktop.” Its completely free—click the “Download” button for it. See Figure 2.2. 3. Download and run the installer. The website usually detects your computer automatically and shows the correct link for Windows or macOS. Click to download. 4. Install the program. Just like with R, run the downloaded file and click “Next” or “Install” through all the default options. Figure 2.2: The official RStudio download page (Download RStudio Desktop). 2.1.3 The 4 Panes of RStudio When you open RStudio, it may look empty. To prepare it for work, go to: File > New File > R Script. Now you should see 4 Windows (Panes). This is your cockpit. 7 2.1. WHY DRIVE THIS CAR? (THE SPECS) Figure 2.3: RStudio interface: 1. Source (Top-Left) for writing code. 2. Environment (TopRight) for data. 3. Console (Bottom-Left) for running code. 4. Plots (Bottom-Right) for graphs. 8 2.2. HOW TO “DRIVE” (RUNNING CODE PROPERLY) 1. Top-Left (The Script): This is your Notepad. You write and save your code here. It is permanent. 2. Bottom-Left (The Console): This is the Engine. This is where the code actually runs. 3. Top-Right (The Environment): This is your Storage. You will see your data (e.g., my_data) here. 4. Bottom-Right (Files/Plots): This is your Gallery. Your graphs will appear here. 2.2 How to “Drive” (Running Code Properly) Crucial Rule: Never type your code directly into the Console (Bottom-Left). If you close RStudio, it disappears forever. The Correct Way: 1. Write your code in the Script (Top-Left). 2. Run the code to send it to the Console. 2.2.1 How to “Turn the Key” (The Run Button) There are two ways to run the code you wrote in the script. Method 1: The Mouse (Good for beginners) 1. Highlight the code you want to run. 2. Click the green Run button at the top of the script window. 9 2.2. HOW TO “DRIVE” (RUNNING CODE PROPERLY) Figure 2.4: Highlight your code, then click the green Run button. Method 2: The Shortcut (The Pro Way) 1. Place your cursor on the line you want to run. 2. Press Ctrl + Enter (Windows) or Cmd + Enter (Mac). This runs the line and moves to the next one automatically. WARNING: THE “MISSING LINE” TRAP A common error happens when you highlight only part of a multi-line command. If you have a code block connected by a plus sign (+), like this: Your Code ggplot(my_data, aes(x = Yield)) + geom_histogram() You must highlight ALL connected lines before clicking Run. If you only run the first line, R sees the +and waits for more. You will see a +symbol in the Console. The Fix: Click inside the Console, press Esc to cancel, then highlight the full block and run it again. 2.2.2 The Magic Symbol: Comments (#) Sometimes you want to write a note to yourself, like “This calculates the average,” but you don’t want R to try and run it as code. To do this, use the hash symbol: # 10 2.2. HOW TO “DRIVE” (RUNNING CODE PROPERLY) R ignores anything you type after a #. Using Comments # This is a comment. R will ignore it. mean(my_data$Yield) # This code will run Why use this? •Notes: Remind yourself what the code does. •Safety: If a line of code is broken, put a #in front of it to “turn it off” without deleting it. 2.2.3 Save Your Script Before closing RStudio, always save your script (File > Save As...). Name it something like My_Thesis_Analysis.R. When you reopen RStudio, simply open the file and continue working. 11 Chapter 3 Putting “Fuel” in the “Car” (Loading Your Excel Data) This is the most important step in our whole “Driver’s Manual.” A car can’t go anywhere without fuel, and R can’t do anything without your data. The good news is that your data is probably in an Excel file (.xlsx). This chapter will show you, step by step, how to get your hard-earned data from that Excel file into R so we can start our “journey.” If you follow these steps, it will work—every time. 3.0.1 The 3 Golden Rules (Cleaning Your Excel File) This is the part where 99% of all problems happen. R is a very powerful engine, but it gets confused easily. It cannot read files made for humans (with pretty titles and merged cells). It can only read files that are pure data. Before you even try to load your file, open it in Excel and make sure it follows these three Golden Rules. •Golden Rule 1: No Merged Cells. No Titles. R has no idea what a merged cell is. Your file must not contain any titles, subtitles, or merged cells. •Golden Rule 2: Headers in Row 1. Only Row 1. Your column names (such as “Yield,” “Treatment,” or “Plot”) must appear in the very first row. •Golden Rule 3: Just the Data. The rest of the file should be a clean, rectangular table of your data. No notes, no total rows, and no pictures. Our “No-Worries” Advice: Make a copy of your Excel file. Keep your decorative original, but create a version called data_for_R.xlsx and clean it up. 12 Visual comparison of a “Bad” Excel file (left) versus a “Good,” R-ready file (right). BAD (for humans) GOOD (for R) Title: “Fertilizer Trial 2024” Variety, Plot, Treatment, Yield (Merged cell across A,B,C,D) A, 1, Control, 45.2 Date: “October” A, 2, Nitro, 55.1 Plot Treatment Yield (kg) A, 3, Phos, 50.8 1 Control 45.2 B, 4, Control, 44.9 2 Nitro 55.1 B, 5, Nitro, 56.2 3 Phos 50.8 B, 6, Phos, 51.3 Note: “Plot 2 had pests” (...and so on...) Figure 3.1: A “Bad” Excel file with merged cells and titles (left) versus a “Good,” R-ready data file (right). Your file must look like the one on the right. 3.0.2 The Most Important Rule: The Shape (Wide vs. Long) Even if your file is clean, it might still be the wrong shape. The Problem: Wide Format (Human Friendly) In Excel, it feels natural to put Group A in one column and Group B in another: •Column A: Yield_Variety_A •Column B: Yield_Variety_B This is called the wide format. It looks nice for humans, but R hates it. Functions such as t.test() and ggplot() do not work easily with this structure. Table 3.1: Wide Format (Human Friendly) — DO NOT USE Yield_Variety_A Yield_Variety_B 12 15 14 18 11 16 13 17 13 3.1. THE MOST IMPORTANT KEY: THE DOLLAR SIGN ($) You must tell R: “Go to the my_data cabinet, open it, and grab the Yield folder.” In R code, we write it like this: my_data$Yield •my_data = The Cabinet •$= The Key (Open it!) •Yield = The Folder Figure 3.7: The Logic of $: It acts as the key to open your specific dataset and find the column inside. 3.1.2 Why are we telling you this now? Because every “Car” in the Showroom works this way. •In the Survey Car (Chapter 6), you will use my_data$Zone. •In the Two-Group Car (Chapter 7), you will use my_data$Treatment. •In the Relationship Car (Chapter 9), you will use my_data$Rainfall. Whenever you see a $in the code later in this book, just remember: Cabinet $ Folder. 20 Chapter 4 Installing Your “Upgrades” (The Common Tools) Great job! You’ve got your “car” (R), your “dashboard” (RStudio), and you know how to put “fuel” in it (load your Excel data). You are 90% of the way there. The “base model” of R is powerful, but it’s not very pretty or easy to handle. To make our journey smooth and to produce beautiful “photos“ (graphs) for our final report, we need to install a few final “upgrades.” In the R world, these “upgrades” are just called packages. We are going to install the two most helpful packages in all of R. Don’t worry, it’s just one line of code. 4.0.1 Your “Graphing Tool” (ggplot2) The first upgrade is called ggplot2. The name is a bit strange, but its job is simple: it makes beautiful graphs. The built-in graphs in R are... well, they’re ugly. They look like they were made in the 1990s. The graphs you’ll make with ggplot2 are professional, clean, and will make your final report look fantastic. We will use this package for every single “photo” (graph) we make in the “Showroom” (Part 2). Think of it as our high-end “digital camera” for the journey. 21 Figure 4.1: Capturing insights, one graph at a time—with ggplot2 in focus! 4.0.2 Your “Data Handling Tool” (dplyr) The second upgrade is called dplyr. This is our “Swiss Army Knife” for data. Sometimes, you’ll need to do simple jobs before you can make a graph. For example, what if you want to: •Calculate the average (mean) yield for each of your fertilizer types? •Filter your data to only look at “Plot A”? •Sort your results from highest to lowest? dplyr is the tool that makes all these “data handling” jobs simple and logical. It’s the “power-steering” and “cruise control” for our data analysis. 4.0.3 The Code: How to Install Them All at Once Okay, let’s add these two “upgrades” to our “garage.” Just like we did with readxl in the last chapter, we use the install.packages() command. The great news is we can install them all at the same time. We just need to give R a “list” of the upgrades we want. In R, a simple list is made by putting things inside c(...) (which stands for “combine”). 1. Open RStudio. 2. Go to the Console (the window on the left, with the ‘>’ symbol). 3. Type this one line exactly as it is written here, and press Enter: Your Code install.packages(c(“ggplot2”, “dplyr”)) 22 R will connect to the internet (you must be online for this) and install both packages. It might take a minute or two. You will see a lot of text flash by. As long as you don’t see a big red “Error” message, you are all set. Remember: You only have to do this one time, ever. Just like readxl, these “upgrades” are now permanently installed in your “garage.” 23 Chapter 5 The Rules of the Road (Normality & p-values) Before we enter the “Showroom” (Part 2) to pick a specific car, there are two universal rules you must know. These rules apply to almost every car you will drive in the next section. Instead of repeating them in every single chapter, we are writing them here once.Mark this page. You will need to come back to it often! 5.0.1 Rule #1: The “p-value” (Your GPS) In every test we run (T-test, ANOVA, Chi-Square), R will give you a number called the p-value. This number tells you if your result is “Real” (Significant) or just “Luck” (Not Significant). THE UNIVERSAL P-VALUE RULE If p < 0.05 (Small Number): SIGNIFICANT. There is a real difference or relationship. (We usually want this!) If p > 0.05 (Big Number): NOT SIGNIFICANT. The groups are the same. It was just random chance. 5.0.2 Rule #2: The Safety Check (Normality) Note: This rule applies only to Measuring (Height, Weight, Yield). If you are doing Surveys (Counting), you can skip this. In statistics, the “Road” is the shape of your data. We must check the road before we pick a car. •The Ferrari (T-Test/ANOVA) needs a smooth, normal road. 24 •The Land Rover (Wilcoxon/Kruskal) handles rough, non-normal roads. We have three ways to check the road. Table 5.1: Example of “Yield” data. We will check if this column is Normal. Group Yield A 12 A 14 ... ... B 15 B 18 ... ... Step 1: Visual Check A (The Histogram) First, let’s look at the shape. We use a histogram. We want to see whether the data looks like a “bell” (high in the middle, low on the sides). Your Code library(ggplot2) ggplot(my_data, aes(x = Yield)) + geom_histogram(bins = 10, fill = “skyblue”, color = “black”) + theme_minimal() 25 Figure 5.1: Left: Normal (Bell Shape) = Smooth Road. Right: Skewed (Messy) = Rough Road. Step 2: Visual Check B (The QQ Plot) Sometimes histograms are messy. A better visual tool is the QQ plot (Quantile-Quantile plot). Think of this like checking whether a fence is straight. •Normal Data: The dots will follow the straight diagonal line. •Not Normal Data: The dots will curve off at the ends (like a banana). Your Code ggplot(my_data, aes(sample = Yield)) + stat_qq() + stat_qq_line() + theme_minimal() 26 Figure 5.2: Left: Normal (Dots on the line). Right: Not Normal (Dots curve away). Step 3: The Engine Check (Shapiro-Wilk Test) Pictures can be subjective. To be scientific, we ask R for a hard number. The Shapiro-Wilk test gives us a p-value. Your Code # Check normality of the Yield column shapiro.test(my_data$Yield) The Result: R Output Shapiro-Wilk normality test data: my_data$Yield W = 0.97, p-value = 0.234 The Decision (The Fork in the Road) Now look at the p-value from the Shapiro test. This is your GPS. 27 THE NORMALITY DECISION Scenario A: p-value >0.05 (Big Number) Verdict: Your data is NORMAL (Smooth Road). Action: Use the Standard Car (T-Test, ANOVA). Scenario B: p-value <0.05 (Small Number) Verdict: Your data is NOT NORMAL (Rough Road). Action: Use the All-Terrain Car (Wilcoxon, Kruskal-Wallis). 5.0.3 Rule #3: The Hierarchy of Truth Sometimes your eyes (Histogram) and the math (Shapiro) disagree. Who wins? THE HIERARCHY 1. Visuals (Subjective): These are like looking at the sky to guess the weather. Good for a quick look. 2. The Test (Objective): This is like looking at a Barometer. It gives you a precise number. The Rule: Always trust the Shapiro-Wilk p-value. It is the safest way to defend your thesis to an examiner. Figure 5.3: Visuals rely on your eyes (Subjective). The Test relies on math (Objective). Trust the math. 5.0.4 Our “Driver’s Manual” is Complete! That’s it. You’ve done it! You have your “car” (R), your “dashboard” (RStudio), your “fuel” (your data), and your “upgrades” (ggplot2 and dplyr). Your “Driver’s Manual” is complete. 28 You are now 100% ready to start your real “journey.” In the next section, Part 2: The “Showroom,” you don’t have to read everything. You just find the one chapter that matches your research, and we’ll “drive” to your answer. Figure 5.4: Your data science vehicle is ready: R as the car, RStudio as the dashboard, data as fuel, and ggplot2 +dplyr as your upgrades. 29 6.2. THE “HOW MANY?” QUESTION (SAMPLE SIZE) •Our Advice: If you use this, be honest! Write in your report: “Due to time constraints, convenience sampling was used.” 6.2 The “How Many?” Question (Sample Size) “How many surveys do I need?” This is the number one question students ask. 6.2.1 The Scientific Way (When you know the Total Population) Sometimes, you know exactly how many people are in your group. This number is called N. •Example: “There are exactly 500 registered farmers in the local cooperative.” (N= 500) If you know N, you should use a formula to calculate your sample size (n). Examiners love this because it proves you are being scientific. The easiest formula to use is Yamane’s Formula. Yamane’s Formula n=N 1 + N(e2) Where: •n= The sample size you need. •N= The total population size (e.g., 500). •e= The “Margin of Error” you accept. (In agriculture, we usually accept 0.05 or 5%). Example Calculation Let’s say you have 500 farmers (N= 500) and you want a standard error of 0.05 (e= 0.05). n=500 1 + 500(0.052) n=500 1 + 500(0.0025) n=500 1+1.25 n=500 2.25 =222.2 So, you need to survey 223 farmers. 36 6.3. THE JOURNEY: VISUALIZING YOUR SURVEY Figure 6.6: Using Yamane’s Formula to turn a known Population (N) into a Sample Size (n). Note: If this number is too big for your budget (e.g., you can’t afford 223 surveys), you can increase your error (e) to 0.07 or 0.10, but you must explain this in your report! 6.2.2 The “Rule of 30” In statistics, magic starts to happen when n= 30. If you have fewer than 30 people in a group, the math (the “Central Limit Theorem”) struggles to work. •Guideline: Try to get at least 30 respondents for each group you want to compare. •Example: If comparing Village A vs. Village B, try to get 30 from A and 30 from B (Total = 60). 6.2.3 Quality over Quantity Remember this rule: 50 honest surveys are better than 200 bad ones. Don’t fake data to hit a high number. A smaller, well-conducted survey is perfectly acceptable for an undergraduate project. 6.3 The Journey: Visualizing Your Survey Okay, you have your data. You loaded it into R (Chapter 3). Now let’s look at it. We assume your data is in a table called my_data, with columns Zone (Wet/Dry) and Preference (Agree/Disagree). 37 6.3. THE JOURNEY: VISUALIZING YOUR SURVEY 6.3.1 The “Photo”: The Stacked Bar Chart We want to compare the counts. A Stacked Bar Chart is the best way to show “How much of Zone 1 said Agree“ vs ”How much of Group Zone 2 said Agree.” The Code: Your Code library(ggplot2) ggplot(my_data, aes(x = Zone, fill = Preference)) + geom_bar(position = “fill”) + labs(title = “Farmer Preferences by Zone”, y = “Proportion”, x = “Zone”) + theme_minimal() What this code does: •aes(x=Zone, fill=Preference): Puts the Zones on the bottom, and uses color (fill) to show the answers. •position=“fill”: This makes all bars the same height (100%), so you can easily compare percentages. Figure 6.7: The result: A Stacked Bar Chart. It clearly shows if one Zone has more “Agree” responses (Blue) than the other. 38 6.4. THE DESTINATION: THE CHI-SQUARE TEST 6.4 The Destination: The Chi-Square Test The graph looks like there is a difference... but is it real? Or just luck? To find out, we need a statistical test. 6.4.1 Wait! Why Not ANOVA? Before we drive, you might be asking: “Why don’t we use ANOVA? Isn’t that the standard test for research?” This is the most common mistake in survey analysis. You cannot use ANOVA (or a T-Test) here. Here is why: 1. The “Average” Problem: ANOVA compares Averages (Means). •You can find the average height of a corn plant. •But you cannot find the average of “Yes” and “No.” 2. The “Normality” Problem: ANOVA assumes your data follows a “Bell Curve” (Normal Distribution). •Survey data is Categorical. It fits into “buckets.” •You cannot check for “Normality” on categories. Because your data is Counts (Categorical), not Measurements (Continuous), we need the Chi-Square Test. Figure 6.8: The Decision: Use ANOVA for “Weighing/Measuring” (Left). Use Chi-Square for “Sorting/Counting” (Right). 39 6.4. THE DESTINATION: THE CHI-SQUARE TEST 6.4.2 The Code Running the test takes two simple steps. First, we turn our data into a small table (a “Contingency Table”). Then, we test that table. # Step 1: Make a simple table of counts # Syntax: table(Row_Variable, Column_Variable) Your Code my_table <- table(my_data$Zone, data$Preference) # Step 2: Run the test Your Code chisq.test(my_table) 6.4.3 Reading the Answer R will give you an answer in the Console. It will look like this: R Output Pearson’s Chi-squared test data: my_table X-squared = 4.2, df = 1, p-value = 0.0404 The Rule: Look at the p-value. •p < 0.05: The difference is Significant. (The Zones truly have different preferences). •p > 0.05: The difference is Not Significant. (Any difference on the graph is just random chance). In our example, p= 0.04. That is less than 0.05. Conclusion: There is a significant difference in preference between the Wet Zone and Dry Zone. You can write this in your report! 40 6.5. TEST DRIVE: REAL-WORLD EXAMPLES 6.5 Test Drive: Real-World Examples You know how to drive the car. Now let’s take it on a few different roads. Here are three specific scenarios you might face in your research. Find the one that looks like yours! 6.5.1 Scenario 1: The “Adoption” Study (Education vs. Tech) The Situation: You are studying a new precision farming technology. You surveyed 100 farmers. You want to know: “Does the Education Level of the farmer (Primary vs. Secondary) affect their Adoption Level (Low, Medium, or High)?” The Data: •Column 1: Education (Categories: “Primary”, “Secondary”) •Column 2: Adoption (Categories: “Low”, “Med”, “High”) Data Preview (CSV Format) ID Education Adoption 1 Primary Low 2 Secondary High 3 Secondary Med 4 Primary Low 5 Secondary High ... ... ... Figure 6.9: Scenario 1: Comparing two groups (Education) across three outcomes (Adoption Level). 41 6.5. TEST DRIVE: REAL-WORLD EXAMPLES The Code: Your Code # Create the table adopt_table <- table(my_data$Education, my_data$Adoption) # Run the test chisq.test(adopt_table) The Result: R Output Pearson’s Chi-squared test data: adopt_table X-squared = 35.714, df = 2, p-value = 1.757e-08 The Conclusion: The p-value (1.757e−08) is < 0.05. Answer: Yes! Education level significantly affects the adoption of technology. (You can likely see on your graph that Secondary education farmers have higher adoption). 42 6.5. TEST DRIVE: REAL-WORLD EXAMPLES 6.5.2 Scenario 2: The “Pest Score” Study (Visual Ratings) The Situation: You are comparing two Tomato Varieties (Variety A and Variety B). You didn’t count the bugs; you just gave each plant a Visual Score: •0 = No Damage •1 = Little Damage •2 = Heavy Damage Warning: Even though 0, 1, and 2 are numbers, they are actually Categories (Buckets). You cannot find the “Average Score.” You must use Chi-Square! The Data: •Column 1: Variety (A, B) •Column 2: Score (0, 1, 2) Data Preview (CSV Format) ID Variety Score 1 A 0 2 A 1 3 B 2 4 A 0 5 B 1 ... ... ... The Code: Your Code # Create the table pest_table <- table(my_data$Variety, my_data$Score) # Run the test chisq.test(pest_table) The Result: R Output Pearson’s Chi-squared test data: pest_table X-squared = 5.8, df = 2, p-value = 0.055 43 6.5. TEST DRIVE: REAL-WORLD EXAMPLES The Conclusion: The p-value (0.055) is > 0.05. It is very close, but it is bigger. Answer: There is NO significant difference in pest resistance between Variety A and Variety B. One is not better than the other. 44 6.5. TEST DRIVE: REAL-WORLD EXAMPLES 6.5.3 Scenario 3: The “Consumer Market” Study (Yes/No) The Situation: You are selling organic mushrooms. You surveyed customers in the City (Urban) and the Village (Rural). You asked: “Are you willing to pay $5.00 for this pack? (Yes/No).” The Data: •Column 1: Location (City, Village) •Column 2: Will_Buy (Yes, No) Data Preview (CSV Format) ID Location Will_Buy 1 City Yes 2 Village No 3 City Yes 4 City Yes 5 Village No ... ... ... The Code: Your Code # Create the table market_table <- table(my_data$Location, my_data$Will_Buy) # Run the test chisq.test(market_table) The Result: R Output Pearson’s Chi-squared test data: market_table X-squared = 24.1, df = 1, p-value = 9.14e-07 The Conclusion: Wait! What is 9.14e-07? This is scientific notation. It means 0.000000914. This is much, much smaller than 0.05. Answer: There is a Highly Significant difference. (Your graph will likely show that City people are much more willing to pay than Village people). 45 7.2. DRIVING ROAD A: THE T-TEST (PARAMETRIC) Figure 7.3: Visualizing the T-Test: It checks if the “peaks” of the two bell curves are far enough apart to be real. 7.2.3 Reading the Output R will give you a lot of numbers. Ignore most of them. Focus on the p-value. R Output Welch Two Sample t-test data: my_data$Yield by my_data$Variety t = -2.45, df = 17.8, p-value = 0.024 alternative hypothesis: true difference in means is not equal to 0 95 percent confidence interval: ... The Decision Rule: •p < 0.05: Significant Difference. (The Treatment worked!) •p > 0.05: No Significant Difference. (The averages are statistically the same). In our example, p= 0.024. This is less than 0.05. Conclusion: Variety A and Variety B are significantly different. 7.2.4 How to Write This in Your Report You need to sound professional. Here is a template you can copy: 52 7.2. DRIVING ROAD A: THE T-TEST (PARAMETRIC) “An independent samples t-test was conducted to compare the yield of Variety A and Variety B. The results showed a significant difference (t=−2.45, p = 0.024). Variety B had a higher mean yield than Variety A.” 53 7.3. DRIVING ROAD B: THE WILCOXON TEST (NON-PARAMETRIC) 7.3 Driving Road B: The Wilcoxon Test (Non-Parametric) You are here because: Your Shapiro–Wilk test gave a p-value < 0.05. Your data is Not Normal. The road is rough, messy, or has outliers (extreme values). You cannot use the Ferrari (T-Test). You need the Land Rover. The appropriate tool here is the Wilcoxon Rank Sum Test (also called the Mann–Whitney U Test). 7.3.1 How it Works (The “Ranking” Engine) Why is this car so tough? Because it uses a different kind of logic. It does not look at the actual values; it looks at the ranks. Imagine a race: •Data: 10 seconds, 11 seconds, 500 seconds (the “500” is a crazy outlier). •T-Test: Tries to average 10, 11, and 500. The average becomes huge (173). The outlier destroys the analysis. •Wilcoxon: Converts them to places: 1st, 2nd, 3rd. •Result: The difference between “2nd” and “3rd” is just one step. The “500” no longer breaks the math. 7.3.2 The Code The code is almost identical to the T-Test. We simply change the function name. Your Code # Syntax: wilcox.test(Measurement ~ Group, data = my_data) wilcox.test(Yield ~ Variety, data = my_data) 7.3.3 Reading the Output R will give something like this: R Output Wilcoxon rank sum test with continuity correction data: my_data$Yield by data$Variety W = 25, p-value = 0.031 alternative hypothesis: true location shift is not equal to 0 Decision Rule: 54 7.3. DRIVING ROAD B: THE WILCOXON TEST (NON-PARAMETRIC) •p < 0.05: Significant Difference (Groups differ). •p > 0.05: No Significant Difference (Groups are similar). In our example, p= 0.031 <0.05.Conclusion: There is a significant difference between Variety A and Variety B. 7.3.4 How to Write This in Your Report Because this is a non-parametric test, we report medians instead of means. “A Wilcoxon rank sum test was conducted to compare the yield of the two varieties. The results showed a significant difference (W= 25,p= 0.031). The median yield of Variety B was significantly higher than that of Variety A.” 55 7.4. THE “PHOTO” OF YOUR TRIP: THE BOX PLOT 7.4 The “Photo” of Your Trip: The Box Plot You already have your p-value (the destination), but you cannot put a p-value on a poster. You need a picture. When comparing two groups of measurements, the Box Plot is the undisputed champion. Why? Because a bar chart only shows the mean, while a box plot shows the median, the spread, and the outliers. 7.4.1 Understanding the Box Plot A box plot looks like a box with “whiskers” coming out of it. Here is how to read it: •Thick Middle Line: The Median (the middle value). •The Box: Contains the middle 50% of your data (the Interquartile Range). A tall box means your data is spread out. A short box means your data is consistent. •The Dots: These are Outliers — extreme values far from the rest. Figure 7.4: Anatomy of a Box Plot: The heavy line is the median. The dots represent outliers. 7.4.2 The Code We use our faithful camera: ggplot2. 56 7.4. THE “PHOTO” OF YOUR TRIP: THE BOX PLOT Your Code library(ggplot2) ggplot(my_data, aes(x = Variety, y = Yield, fill = Variety)) + geom_boxplot() + labs(title = “Comparison of Variety A vs B”, y = “Yield (kg/ha)”, x = “Variety”) + theme_minimal() Figure 7.5: The result: A clean, professional graph for your report. You can clearly see whether one variety tends to yield more. 7.4.3 Pro Tip: Adding the Raw Points Sometimes examiners like to see the individual observations on top of the box. You can do this by adding one extra layer: geom_jitter(). Advanced Code (Optional) ggplot(my_data, aes(x = Variety, y = Yield, fill = Variety)) + geom_boxplot(alpha = 0.5) + geom_jitter(width = 0.2) + theme_minimal() 57 7.4. THE “PHOTO” OF YOUR TRIP: THE BOX PLOT Figure 7.6: The result: A clean, professional graph for your report. You can clearly see whether one variety tends to yield more with geom_jitter(). This places small dots over the box plot representing every individual plant or animal measured. It looks clear, transparent, and very scientific! 58 7.5. THE “SPECIAL ROAD CONDITION”: PAIRED DATA (THE “BEFORE & AFTER”) 7.5 The “Special Road Condition”: Paired Data (The “Before & After”) Before we close the garage on Chapter 6, we must talk about one special road condition. Are you comparing two totally different groups of objects? Or are you measuring the same object twice? 7.5.1 The Difference: Independent vs. Paired •Independent (The Standard Road): You compare Plot A vs. Plot B, or Cow Group 1 vs. Cow Group 2. The samples are separate. •Paired (The Special Road): You measure the same object at two different times. Agriculture examples of paired data: •Animal science: Weighing the same cow before feeding and after feeding. •Soil science: Measuring pH in a plot, adding lime, then measuring pH in the same plot a month later. •Post-harvest: Measuring fruit firmness at harvest, then measuring the same fruit after one week of storage. Figure 7.7: Independent vs. Paired samples: Left shows two different cows (independent); right shows the same cow measured twice (paired). The math you use is different. 59 7.5. THE “SPECIAL ROAD CONDITION”: PAIRED DATA (THE “BEFORE & AFTER”) 7.5.2 The Code: Just add paired = TRUE If your data are paired, you must tell R. If you run the standard (independent) test, R will treat paired observations as independent and your p-value will be wrong (often too large). For paired data it is often easiest to keep the measurements in two columns in Excel: •Column 1: Weight_Before •Column 2: Weight_After Scenario A: Paired t-test (if normal) Your Code # Syntax: t.test(column1, column2, paired = TRUE) t.test(my_data$Weight_Before, my_data$Weight_After, paired = TRUE) Scenario B: Paired Wilcoxon (if not normal) Your Code wilcox.test(my_data$Weight_Before, my_data$Weight_After, paired = TRUE) 60 7.6. TEST DRIVE: REAL-WORLD EXAMPLES 7.6 Test Drive: Real-World Examples You have the keys. Now let’s drive. Here are four real research scenarios. Find the one that looks like yours! 7.6.1 Scenario 1: The “High-Yield” Corn (Normal Data) The Situation: You are comparing the yield (kg/ha) of two corn varieties: “Hybrid X” and “Local.” You have 15 plots for each. The Check: You ran shapiro.test() and got p= 0.45. The data is Normal. The Goal: Drive Road A (The T-Test). Your Code # Syntax: t.test(Measurement ~ Group, data = my_data) t.test(Yield ~ Variety, data = my_data) The Result: R Output Welch Two Sample t-test data: Yield by Variety t = 3.12, df = 26.4, p-value = 0.0042 mean in group Hybrid X: 5200 mean in group Local: 4100 The Interpretation: Since p= 0.0042 <0.05, there is a significant difference. Hybrid X (Mean = 5200) yielded significantly more than Local (Mean = 4100). 61 8.3. THE SAFETY CHECK (CONDENSED) 8.3 The Safety Check (Condensed) Before we choose between the Bus (ANOVA) and the Tank (Kruskal–Wallis), we must check the road condition: Normality. For a full explanation of Normality and why it matters, see Section 5.0.2. 8.3.1 The Code Run the ShapiroWilk test on your measurement column. Your Code shapiro.test(my_data$Yield) 8.3.2 The Decision Look at the p-value: THE FORK IN THE ROAD If p-value >0.05 (Normal): →Go to Section 8.4 (ANOVA). If p-value <0.05 (Not Normal): →Go to Section 8.7 (Kruskal–Wallis). 8.4 Driving Road A: The One-Way ANOVA You are here because: Your data is Normal (p > 0.05). We use the One-Way ANOVA (Analysis of Variance). 8.4.1 The “Bus” Logic ANOVA asks a simple question: “Is the variation BETWEEN the groups bigger than the variation WITHIN the groups?” If the piles (groups) are truly different, ANOVA will detect it. 8.4.2 The Code (2 Steps) In R, we first build the model using aov(), and then we look at the summary. 68 8.5. THE “WHO WON?” STEP: CHOOSING YOUR POST-HOC TEST Your Code # Step 1: Build the Model (save as ’my_model’) my_model <- aov(Yield ~ Variety, data = my_data) # Step 2: Get the Results summary(my_model) 8.4.3 Reading the Output R will show a table. Look for the column Pr(>F) — this is your p-value. R Output Df Sum Sq Mean Sq F value Pr(>F) Variety 2 154.2 77.10 12.45 0.00012 *** Residuals 27 167.2 6.19 The Decision: •p>0.05 (Not Significant): Stop. All groups are statistically the same. •p<0.05 (Significant): Proceed. At least one group is different. ANOVA does not tell you which one — for that, you need a Post-Hoc Test. 8.5 The “Who Won?” Step: Choosing Your Post-Hoc Test This is one of the most important sections for Agriculture students. ANOVA told you: “There is a difference.” Now you need a Post-Hoc Test (Latin: “after this”) to determine which varieties differ. You have a Menu of Three Tests. Choose the one that fits your research goal. 8.5.1 Option 1: Tukey’s HSD (The “Gold Standard”) Use this when: You want to compare everyone vs. everyone. (e.g., Is A better than B? Is B better than C? Is A better than C?) This is the safest and most respected test. 69 8.6. HOW TO REPORT THE RESULTS (THE “LETTERS”) Your Code # Syntax: TukeyHSD(your_model) TukeyHSD(my_model) 8.5.2 Option 2: Dunnett’s Test (The “Control Freak”) Use this when: You want to compare all treatments only against a single Control. (e.g., Water (Control) vs. three new sprays.) You do not care about comparisons such as Spray 1 vs. Spray 2. Required Package: DescTools You only need to install this package once. After that, simply load it using library(). Your Code install.packages(“DescTools”) # Install once library(DescTools) DunnettTest(x = my_data$Yield, g = my_data$Variety, control = “Water”) Note: The control group name must match exactly. 8.5.3 Option 3: DMRT (Duncan’s Multiple Range Test) Use this when: Your Professor or Journal specifically asks for DMRT. Warning: DMRT is a very liberal test (it easily finds significance). Many modern statisticians prefer Tukey, but many Agriculture Journals still use DMRT. Required Package: agricolae You only need to install this package once. After that, load it with library(). Your Code install.packages(“agricolae”) # Install once library(agricolae) duncan.test(my_model, “Variety”, console = TRUE) 8.6 How to Report the Results (The “Letters”) In Agriculture, we do not usually paste raw p-values into our report. We use Letter Notation (a, b, c). 70 8.6. HOW TO REPORT THE RESULTS (THE “LETTERS”) Figure 8.3: The Menu: Tukey compares everyone (all arrows). Dunnett compares only vs. Control (one direction). DMRT ranks treatments. 8.6.1 What Do the Letters Mean? •Different letters (a vs. b): These groups are significantly different. (A is statistically higher or lower than B.) •Same letters (a vs. a): These groups are not significantly different. (They are statistically the same.) •Combo letters (ab): This group is “in the middle.” It is not different from group “a” and not different from group “b.” 8.6.2 Getting the Letters Automatically The standard TukeyHSD() function gives you p-values, but converting them to letters manually is a headache. Since this is agricultural research, we use the agricolae package, which generates the letters automatically. Your Code library(agricolae) # Syntax: HSD.test(model, “Group”, console = TRUE) HSD.test(my_model, “Variety”, console = TRUE) 71 8.7. DRIVING ROAD B: THE KRUSKAL–WALLIS TEST 8.6.3 Reading the Output R will print a table showing the ranked means and their letters: R Output Variety Yield Groups Fert_C 65.4 a Fert_A 62.1 a Fert_B 45.2 b Control 30.1 c Interpretation: •Fert C and Fert A both have the letter “a.” They are the top performers and are not different from each other. •Fert B has the letter “b.” It is significantly lower than A and C. •Control has the letter “c.” It is significantly lower than all other treatments. How to write this in your Thesis: “Means followed by the same letter are not significantly different according to Tukey’s HSD test.” 8.7 Driving Road B: The Kruskal–Wallis Test You are here because: Your data is NOT normal (p < 0.05). You cannot use the ANOVA Bus — it will break down on this rough road. Instead, we use the Kruskal–Wallis test. 8.7.1 The “Tank” Logic Think of this as “ANOVA on Ranks.” This test cheats a little bit to be safer. Instead of looking at the actual numbers (which might have crazy outliers), it looks at the ranks. Imagine a race: •Data: 10 seconds, 11 seconds, 500 seconds (The “500” is a crazy outlier). •ANOVA: Tries to average 10, 11, and 500. The average is huge (173). The outlier breaks the math. 72 8.8. THE “WHO WON?” STEP: POST-HOC FOR ROAD B •Kruskal-Wallis: Converts them to places: 1st, 2nd, 3rd. •Result: The difference between “2nd” and “3rd” is just 1 step. The “500” doesn’t break the math anymore. 8.7.2 The Code The syntax mirrors ANOVA but uses a different function name. Your Code # Syntax: kruskal.test(Measurement ~ Group, data = my_data) kruskal.test(Yield ~ Variety, data = my_data) 8.7.3 Reading the Output R Output Kruskal-Wallis rank sum test data: Yield by Variety Kruskal-Wallis chi-squared = 14.2, df = 2, p-value = 0.0008 The Decision: •If p > 0.05:Not significant — stop here; all groups are similar. •If p < 0.05:Significant — at least one group differs. Proceed to the post-hoc step to find out who. 8.8 The “Who Won?” Step: Post-Hoc for Road B If the Kruskal–Wallis test is significant (p < 0.05), you know that someone differs, but not who. You cannot use Tukey (Tukey assumes normal data). Instead use pairwise nonparametric comparisons with an adjustment for multiple testing. We recommend the pairwise Wilcoxon test with a Bonferroni (or other) p-value adjustment. 8.8.1 The Code Your Code pairwise.wilcox.test(my_data$Yield, my_data$Variety, p.adjust.method = “bonferroni”) 73 8.8. THE “WHO WON?” STEP: POST-HOC FOR ROAD B 8.8.2 Reading the Output R returns a matrix (upper or lower triangle) of adjusted p-values. R Output Fert_A Fert_B Fert_B 1.000 - Fert_C 0.004 0.002 How to read the grid: •Fert_A vs. Fert_B: p= 1.000 (not significant). •Fert_A vs. Fert_C: p= 0.004 (significant). •Fert_B vs. Fert_C: p= 0.002 (significant). Conclusion: Fertilizer C differs significantly from both A and B; A and B do not differ. 74 8.9. THE “WHO WON?” STEP: POST-HOC FOR ROAD B (DUNN’S TEST) 8.9 The “Who Won?” Step: Post-Hoc for Road B (Dunn’s Test) If your Kruskal–Wallis test was significant (p < 0.05), you know someone won, but you don’t know who. You cannot use Tukey (Tukey assumes normal data). For the “Tank” (Kruskal–Wallis), we use the specific partner tool called Dunn’s Test. 8.9.1 Required package Required package: dunn.test Install the package once; thereafter just load it with library(). 8.9.2 The Code Dunn’s test is not built into base R, so install the small upgrade (package) and run the test. Your Code install.packages(“dunn.test”) # Install once library(dunn.test) # Syntax: dunn.test(Value, Group, method = “bonferroni”) dunn.test(my_data$Yield, my_data$Variety, method = “bonferroni”) Note: We use method = “bonferroni” to adjust p-values and reduce false positives. 8.9.3 Reading the Output R will print a list of pairwise comparisons. Focus on the P.adjusted column. R Output Comparison Z P.unadj P.adjusted 1. Fert_A - Fert_B -0.52 0.6012 1.0000 2. Fert_A - Fert_C 2.85 0.0041 0.0123 * 3. Fert_B - Fert_C 3.10 0.0019 0.0057 * How to read the table: •Fert_A vs. Fert_B: P.adjusted = 1.0000 (not significant). 75 8.9. THE “WHO WON?” STEP: POST-HOC FOR ROAD B (DUNN’S TEST) •Fert_A vs. Fert_C: P.adjusted = 0.0123 (<0.05) — significant. •Fert_B vs. Fert_C: P.adjusted = 0.0057 (<0.05) — significant. Conclusion: Fertilizer C is significantly different from both A and B. A and B are not different. 76 8.10. GETTING THE NUMBERS: MEANS AND STANDARD DEVIATIONS 8.10 Getting the Numbers: Means and Standard Deviations Your ANOVA gave you a p-value. Your Post-Hoc test gave you letters (a, b). But when you build the final table for your thesis, you also need the actual descriptive statistics. Most Agriculture Journals prefer the format: Mean ±Standard Deviation (SD) or Mean ±Standard Error (SE). We use our “Swiss Army Knife” tool: dplyr (see 4.0.2). 8.10.1 The Code We instruct R to: 1. Take the dataset. 2. Group the data by Variety. 3. Calculate the Mean, SD, and SE for each group. Your Code library(dplyr) # Create a summary table my_summary <- my_data %>% group_by(Variety) %>% summarise( Average = mean(Yield), SD = sd(Yield), SE = sd(Yield) / sqrt(n()) ) # View the table print(my_summary) Note: SE is calculated as SD/√n. Use whichever your supervisor or journal prefers. 8.10.2 Reading the Output R produces a clean summary table: 77 Chapter 9 The “Relationship Car” (Do X and Y Move Together?) 9.1 Is this the Right Car for You? 9.1.1 Use this car if. . . You have TWO continuous numbers (measurement vs. measurement). You want to know if they are “holding hands” (related). •“As X goes up, does Y go up?” (Positive relationship) •“As X goes up, does Y go down?” (Negative relationship) Agriculture Examples: •Rainfall vs. Yield: Does more rain give more grain? •Temperature vs. Growth: Does hotter weather reduce growth? •Age vs. Milk: Do older cows give less milk? STOP AND CHECK Do you have a category? If one of your columns is text (like “Variety A” or “Control”), STOP. Go back to Chapter 7(Two Groups) or Chapter 8(Multiple Groups). This chapter is ONLY for Number vs. Number. 9.2 The Golden Rule: Correlation =Causation Before coding anything, you must memorize this rule. It is the most famous rule in all of statistics: 84 9.3. THE JOURNEY: VISUALIZING THE RELATIONSHIP Just because two things move together does not mean one causes the other. 9.2.1 The “Ice Cream & Shark” Example In the summer, ice cream sales go up. In the summer, shark attacks go up. •Correlation: Yes — both lines increase together. •Causation? No — eating ice cream does not cause shark attacks. •The Real Cause: Summer (more people swim). Figure 9.1: Correlation between ice cream sales and shark attacks does not imply causation, as both are caused by the third variable: summer. 9.2.2 The Agriculture Warning Be careful in your thesis! •Wrong: “My correlation proves Nitrogen caused the yield increase.” (Maybe you also watered those plants more.) •Right: “There was a strong positive relationship between Nitrogen and yield.” 9.3 The Journey: Visualizing the Relationship The “photo” for this trip is the scatter plot. It places each observation as a dot on the graph. 85 9.3. THE JOURNEY: VISUALIZING THE RELATIONSHIP We use our camera: ggplot2. Your Code library(ggplot2) ggplot(my_data, aes(x = Rainfall, y = Yield)) + geom_point() + theme_minimal() 9.3.1 Reading the Map Look at the pattern (shape) of the dots: •Uphill slope (/): Positive correlation (as Rain increases, Yield increases). •Downhill slope (\): Negative correlation (as Pests increase, Yield decreases). •Messy cloud (O): No correlation (no relationship). Figure 9.2: Visualizing Correlation: Positive, Negative, and No Relationship in Agricultural Data. 86 9.4. THE SAFETY CHECK: PEARSON VS. SPEARMAN 9.4 The Safety Check: Pearson vs. Spearman Just like in previous chapters, we have a “Ferrari” and a “Land Rover” for the journey. •The Ferrari (Pearson Correlation): The standard test. It is powerful and precise, but assumes your data is normal and has a linear pattern. •The Land Rover (Spearman Correlation): A rank-based test. It handles outliers, non-normal data, and curved (non-linear) relationships very well. 9.4.1 Checking the Road You must check the normality of both your X variable and your Y variable. (See 5.0.1 if you need a refresh on interpreting p-values). Your Code shapiro.test(my_data$Rainfall) shapiro.test(my_data$Yield) THE FORK IN THE ROAD Scenario A: BOTH are Normal (p > 0.05) →Use Pearson. Scenario B: ONE (or both) is NOT Normal (p < 0.05) →Use Spearman. 9.5 The Destination: The Code (cor.test) We use the function cor.test(). It gives both the p-value (significance) and the correlation coefficient (strength). 9.5.1 Road A: Pearson (Normal Data) Your Code cor.test(my_data$Rainfall, my_data$Yield, method = “pearson”) The Result: 87 9.6. READING THE DASHBOARD: THE “R” VALUE R Output Pearson’s product-moment correlation data: my_data$Rainfall and my_data$Yield t = 3.45, df = 28, p-value = 0.002 alternative hypothesis: true correlation is not equal to 0 cor = 0.55 95 percent confidence interval: 0.21 to 0.78 9.5.2 Road B: Spearman (Not Normal) Your Code cor.test(my_data$Pests, my_data$Yield, method = “spearman”) The Result: R Output Spearman’s rank correlation rho data: my_data$Pests and my_data$Yield S = 189, p-value = 0.0009 alternative hypothesis: true rho is not equal to 0 rho = -0.68 95 percent confidence interval: -0.85 to -0.42 9.6 Reading the Dashboard: The “r” Value The test output gives you a p-value, which tells you whether the relationship is statistically significant. It also gives you the correlation coefficient. For Pearson this appears as cor, and for Spearman it appears as ρ(rho). Both represent the same concept, so in this chapter we use the common symbol rfor interpretation. 9.6.1 The Scale: −1to +1 •Exactly +1:Perfect positive match (twin brothers). •Exactly 0: No relationship at all (strangers). •Exactly −1:Perfect negative match (complete opposites). 88 9.7. HOW TO WRITE THIS IN YOUR REPORT 9.6.2 The Agriculture Guide to Strength In real data, values are rarely exactly 1 or 0. Use this guide: Value of r(positive or negative) Strength Description 0.0 – 0.3 Weak / Negligible 0.3 – 0.5 Low Correlation 0.5 – 0.7 Moderate Correlation 0.7 – 0.9 Strong Correlation 0.9 – 1.0 Very Strong Correlation 9.7 How to Write This in Your Report When reporting correlation results, you must include: •the p-value, •the rvalue, and •the interpretation (strength + direction). Template: “A [Pearson/Spearman] correlation test showed a [strong/weak], [positive/negative] relationship between Variable A and Variable B (r= 0.85, p < 0.05).” Example: “A Pearson correlation test showed a strong, positive relationship between Rainfall and Yield (r= 0.78, p = 0.002).” 89 9.8. TEST DRIVE: REAL-WORLD EXAMPLES 9.8 Test Drive: Real-World Examples You know the rules. You know the warning (Correlation =Causation). Now let’s see how this works in real agricultural research. 9.8.1 Scenario 1: The “Good Match” (Positive / Pearson) The Situation: You measured the Nitrogen Rate applied and the final Grain Yield. You want to know whether they are related. The Check: Both variables are normal (p > 0.05). The Plan: Drive Road A (Pearson). Your Code cor.test(my_data$Nitrogen, my_data$Yield, method = “pearson”) R Output Pearson’s product-moment correlation data: my_data$Nitrogen and my_data$Yield t = 8.4, df = 18, p-value = 0.0001 sample estimates: cor = 0.89 The Interpretation: •p < 0.05: Significant relationship. •r = 0.89: Very strong, positive correlation. •Report: “There is a very strong, positive relationship between Nitrogen and Yield.” 90 9.8. TEST DRIVE: REAL-WORLD EXAMPLES 9.8.2 Scenario 2: The “Bad Influence” (Negative / Spearman) The Situation: You counted Thrips (Pests) on chili plants and measured their Marketable Yield. The Problem: Thrips counts are highly skewed — not normal. The Expectation: As pests go up, yield should go down. The Plan: Drive Road B (Spearman). Your Code cor.test(my_data$Thrips, my_data$Yield, method = “spearman”) R Output Spearman’s rank correlation rho data: my_data$Thrips and my_data$Yield S = 4500, p-value = 0.003 sample estimates: rho = -0.72 The Interpretation: •p < 0.05: Significant. •r = -0.72: Strong, negative relationship. •Report: “There is a strong negative relationship between pest count and yield.” 91 9.8. TEST DRIVE: REAL-WORLD EXAMPLES 9.8.3 Scenario 3: The “Strangers” (No Relationship) The Situation: You checked whether the Farmer’s Age is related to their Rice Yield. The Check: Data is normal. The Plan: Drive Road A (Pearson). Your Code cor.test(my_data$Age, my_data$Yield, method = “pearson”) R Output Pearson’s product-moment correlation t = 0.4, df = 28, p-value = 0.65 sample estimates: cor = 0.08 The Interpretation: •p > 0.05: Not significant. •r = 0.08: Extremely weak relationship (almost zero). •Report: “There is no significant relationship between farmer age and rice yield.” 9.8.4 Summary Checklist for the Relationship Car •Did you confirm both variables are numerical? •Did you check Normality using ShapiroWilk? •Did you choose Pearson (Normal) or Spearman (Not Normal)? •Did you check both r(Strength) and p(Significance)? •Did you remember that Correlation is not Causation? If yes, you are done. You have conquered relationship chapter 92 Chapter 10 The “Prediction Car” (Linear Regression) 10.1 Is this the Right Car for You? 10.1.1 “Use this car if...” You have TWO Continuous Numbers, and you want to Predict one based on the other. The Difference from Chapter 8: •Correlation (9): Asks “Are they related?” (Answer: Yes/No). •Regression (Here): Asks “What is the Equation?” (Answer: y=mx +c). Agriculture Example: •“If I add exactly 10kg of Nitrogen, exactly how much Yield will I get?” •“If I apply exactly 2 L of pesticide mixture per plot, exactly how many insects will be controlled?” •“If I irrigate each plant with exactly 5 L of water, exactly how much increase in growth will I see?” •“If I apply exactly 20 tons of compost per hectare, exactly how much improvement in soil organic matter will occur?” 10.2 The Logic: The Line of Best Fit Regression tries to draw a straight line through your data. We need two numbers: 93 10.6. TROUBLESHOOTING: WHAT IF THE ENGINE FAILS? The Diagnosis: Your data is unstable at high values. Example: Small farms always produce 2 tons. Big farms produce anywhere from 10 to 50 tons. The “Variance” is not equal. The Solution: Your p-values are unreliable. You often need a Log Transformation (shrinking the big numbers). Advice: Do not just run the model. Consult a statistician about “Transforming Data.” THE GOLDEN RULE OF SAFETY If your diagnostic plots show these problems, STOP. Do not put the linear equation in your thesis. It is scientifically incorrect. It is better to say: “The assumptions of linear regression were not met, so the model was not used,” than to publish a wrong equation. 100 10.7. READING THE DASHBOARD (THE RESULTS) 10.7 Reading the Dashboard (The Results) If all 4 Pillars are safe, now check the results. Your Code summary(my_model) R Output Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 150.2 12.5 12.0 0.0001 *** Nitrogen 12.5 1.1 11.3 0.0002 *** ... Multiple R-squared: 0.854 10.7.1 Step 1: Is it Real? (The p-value) Look at the column marked Pr(>|t|) in the Nitrogen row. •Value: 0.0002 •Meaning: This is much smaller than 0.05. •Conclusion: The relationship is Significant. Nitrogen definitely affects Yield. 10.7.2 Step 2: What is the Equation? (The Estimate) Look at the column marked Estimate. This tells you the math. A. The Intercept Row (150.2): This is where the line starts. “If you add 0 kg of Nitrogen, your field will still yield 150.2 kg.” B. The Nitrogen Row (12.5): This is the Slope. It is the most important number in the chapter. “For every 1 unit increase in Nitrogen, the Yield goes up by 12.5 units.” Your Final Equation: Y ield = 150.2 + 12.5(Nitrogen) 10.7.3 Step 3: How Strong is it? (R-squared) Look at the bottom line: Multiple R-squared. 101 10.8. THE “PHOTO”: SCATTER PLOT WITH REGRESSION LINE •Value: 0.854 •Meaning: This is a score out of 1.0. •Translation: “Our model explains 85.4% of the variation in yield.” (This is an excellent score. If it was 0.10, your model would be very weak). 10.8 The “Photo”: Scatter Plot with Regression Line We use ggplot2 again, but this time we add a “smooth” line to display the fitted regression equation. Your Code library(ggplot2) ggplot(my_data, aes(x = Nitrogen, y = Yield)) + geom_point() + geom_smooth(method = “lm”, color = “red”) + theme_minimal() Figure 10.7: Scatter plot with fitted regression line showing how Yield changes with increasing Nitrogen levels. 10.9 Test Drive: Real-World Examples You now have the keys to the Prediction Car. Lets drive it on different agricultural roads. 102 10.9. TEST DRIVE: REAL-WORLD EXAMPLES Here are four real scenarios you will face in applied research. 10.9.1 Scenario 1: The “Feed Efficiency” Trial (Positive Slope) The Situation: You fed 20 pigs different amounts of a “High Protein Mix” (05 kg/day). You measured their Daily Weight Gain (g). Goal: Predict weight gain based on feed amount. Figure 10.8: Scenario 1: Does more feed result in more meat? (Positive relationship). Step 1: The Code Your Code pig_model <- lm(Gain ~ Feed, data = my_data) summary(pig_model) Step 2: The Output 103 10.9. TEST DRIVE: REAL-WORLD EXAMPLES R Output Coefficients: (Intercept) 200.5 ... 0.001 ** Feed 50.2 ... 0.00004 *** Multiple R-squared: 0.92 The Interpretation: •Slope (50.2): Each 1 kg of feed increases daily gain by 50.2 g. •Equation: Gain = 200.5 + 50.2×Feed. •Conclusion: The feed is highly effective (R2= 0.92). 104 10.9. TEST DRIVE: REAL-WORLD EXAMPLES 10.9.2 Scenario 2: The “Ripening” Study (Negative Slope) The Situation: You exposed bananas to different concentrations of Ethylene Gas (0, 10, 50, 100 ppm). You measured Days to Ripen. Expectation: More gas fewer days (negative relationship). Figure 10.9: Scenario 2: Does ethylene gas make fruits ripen faster? (Negative relationship). Step 1: The Code Your Code gas_model <- lm(Days ~ Ethylene, data = my_data) summary(gas_model) Step 2: The Output R Output Coefficients: (Intercept) 14.0 ... 0.001 ** Ethylene -0.12 ... 0.002 ** Multiple R-squared: 0.85 105 10.9. TEST DRIVE: REAL-WORLD EXAMPLES The Interpretation: •Slope (-0.12): Each 1 ppm ethylene reduces ripening time by 0.12 days. •Equation: Days = 14.0−0.12 ×Ethylene. •Conclusion: Ethylene significantly accelerates ripening. 106 10.9. TEST DRIVE: REAL-WORLD EXAMPLES 10.9.3 Scenario 3: The “Weed Competition” Study (Economic Loss) The Situation: You recorded Weed Density (weeds/m2) in wheat plots and measured final Grain Yield. Goal: Quantify how much yield is lost per weed. Figure 10.10: Scenario 3: How much grain do weeds steal? (Economic prediction). Step 1: The Code Your Code weed_model <- lm(Yield ~ WeedCount, data = my_data) summary(weed_model) Step 2: The Output 107 10.9. TEST DRIVE: REAL-WORLD EXAMPLES R Output Coefficients: (Intercept) 4000.0 ... 0.001 ** WeedCount -25.5 ... 0.001 *** Multiple R-squared: 0.78 The Interpretation: •Slope (-25.5): Each additional weed per m2reduces yield by 25.5 kg. •Equation: Yield = 4000 −25.5×Weeds. 108 10.9. TEST DRIVE: REAL-WORLD EXAMPLES 10.9.4 Scenario 4: The “Null” Result (Moon Phase Myth) The Situation: Some farmers believe a Full Moon increases germination. You tested Moon Visibility (%) vs. Germination Rate. Figure 10.11: Scenario 4: Testing a myth — does the moon affect germination? (No relationship). Step 1: The Code Your Code moon_model <- lm(Germination ~ MoonPerc, data = my_data) summary(moon_model) Step 2: The Output R Output Coefficients: (Intercept) 85.2 ... 0.001 ** MoonPerc 0.01 ... 0.654 Multiple R-squared: 0.01 109 11.4. THE SAFETY CHECK (CONDENSED) 11.4 The Safety Check (Condensed) Before we choose our car, we must check the road condition (Normality). (If you forgot how to interpret Histograms or Shapiro-Wilk, please go back and read Chapter 5: The Rules of the Road 8.3.) 11.4.1 The Code Run the Shapiro-Wilk test on your Yield column. Your Code shapiro.test(my_data$Yield) 11.4.2 The Decision Look at the p-value: THE FORK IN THE ROAD If p-value > 0.05 (Normal): →Go to Section 10.5 (RCBD ANOVA). If p-value < 0.05 (Not Normal): →Go to Section 10.7 (Friedman Test). 116 11.5. ROAD A: THE RCBD ANOVA (PARAMETRIC) 11.5 Road A: The RCBD ANOVA (Parametric) You are here because: Your data is Normal (p > 0.05). We still use ANOVA, but the formula is different from Chapter 7. This time we must include the Block in the model so R can subtract the hill effect. 11.5.1 The Code (The Plus Sign) The magic symbol here is the plus sign (+). We tell R: “Yield depends on Treatment PLUS Block.” Your Code # Syntax: Yield ~ Treatment + Block my_model <- aov(Yield ~ Treatment + Block, data = my_data) summary(my_model) 11.5.2 Reading the Output (Two P-values) In Chapter 7, we only had one important row. Now, RCBD produces two rows that matter. R Output Df Sum Sq Mean Sq F value Pr(>F) Treatment 2 154.2 77.1 12.45 0.001 ** Block 3 45.5 15.1 2.10 0.120 Residuals 18 ... Row 1: Treatment (Did the fertilizer work?) Focus on the Treatment row (may appear as Variety,Fert, etc.). •p < 0.05: Significant. Your fertilizers differ. You have a winner. •p > 0.05: Not Significant. All fertilizers performed the same. This is the row that answers your research question. Row 2: Block (Did the field matter?) Now check the Block row. •p < 0.05 (Significant): The Hill Mattered. Interpretation: Conditions differed across the field. Verdict: Excellent! You were correct to use blocks. 117 11.5. ROAD A: THE RCBD ANOVA (PARAMETRIC) •p > 0.05 (Not Significant): The Hill Didnt Matter. Interpretation: The field was fairly uniform. Verdict: Thats fine — you were simply extra cautious. 11.5.3 Special Scenario: The “Strong Block, Weak Treatment” Sometimes you will see a result like this: R Output Df ... F value Pr(>F) Treatment 2 ... 0.54 0.650 (Not Significant ) Block 3 ... 8.90 0.002 ** (Significant ) What this means: 1. Your Fertilizers (Treatment) are all the same. None of them is better than the other. 2. BUT your Field (Block) was very uneven. The top of the hill behaved very differently from the bottom. STOP HERE Because the Treatment is not significant, you cannot run a Post-Hoc test (Tukey or DMRT). •Do not try to find a “winner” among the fertilizers — there isnt one. •Do not run Tukey on the Blocks — we are not interested in comparing blocks. Your Report: “The blocking was effective (p < 0.05), indicating significant field variation. However, there was no significant difference between the treatments (p > 0.05).” The Decision: If the Treatment p-value is significant (<0.05), proceed to the next section to find out which treatment won (Post-Hoc). 118 11.6. THE FINISH LINE: POST-HOC & MEANS 11.6 The Finish Line: Post-Hoc & Means You are here because: •Your ANOVA Treatment p-value was significant (<0.05). •You need to know which fertilizer was the best. We need two things for your thesis: the Letters (a, b, c) and the Numbers (Mean ± SD). 11.6.1 Step 1: The Letters (Choosing Your Test) In Chapter 7, we learned about the menu of post-hoc tests. We use the same menu here, but remember: we test the Treatment, not the Block. CRITICAL RULE When running these tests, R will ask “Which group?” You must type “Treatment” (or whatever your fertilizer column is named). Do not type “Block”. We do not compare blocks. Option 1: Tukey’s HSD (The Standard) Use this to compare every treatment with every other treatment. Your Code library(agricolae) # Syntax: HSD.test(model, “Group”, console=TRUE) HSD.test(my_model, “Treatment”, console=TRUE) Option 2: DMRT (Duncan’s Test) Use this if your professor specifically asks for “Duncan’s.” Very common in older agriculture journals. Your Code library(agricolae) duncan.test(my_model, “Treatment”, console=TRUE) Option 3: Dunnett’s Test (Control-Focused) Use this if you want to compare treatments only against the “Control” group. 119 11.6. THE FINISH LINE: POST-HOC & MEANS Your Code library(DescTools) DunnettTest(x = my_data$Yield, g = my_data$Treatment, control = “Control”) 11.6.2 Step 2: The Numbers (Mean ±SD) Next, we calculate the numerical values to put in your final table. We use dplyr and group by Treatment only. Your Code library(dplyr) my_summary <- my_data %>% group_by(Treatment) %>% summarise( Average = mean(Yield), SD = sd(Yield) ) print(my_summary) 11.6.3 Building the Final RCBD Table Combine the Average (from Step 2) and the Letters (from Step 1). Table 11.1: Effect of Fertilizers on Yield (RCBD Analysis) Treatment Yield (Mean ±SD) Fertilizer A 65.4±2.1a Fertilizer B 62.1±3.5ab Fertilizer C 55.2±1.8b Control 40.1±4.0c Note: Means with the same letter are not significantly different. 120 11.7. ROAD B: THE FRIEDMAN TEST (NON-PARAMETRIC) 11.7 Road B: The Friedman Test (Non-Parametric) You are here because: Your data is NOT Normal (p < 0.05). The standard ANOVA assumes a smooth road. Since your road is rough (skewed data), we use the “All-Terrain” vehicle for blocked designs: The Friedman Rank Sum Test. 11.7.1 The Logic: Racing Inside the Block How does this test handle the “Hill Effect” without complicated math? It converts the data into Ranks, but it does so separately inside each block. Think of each block as a small race: •Block I (Dry / Low Yield): –Yields: Fert A (2 tons), Fert B (3 tons), Fert C (1 ton) – Ranks: B = 1st, A = 2nd, C = 3rd •Block II (Wet / High Yield): –Yields: Fert A (20 tons), Fert B (30 tons), Fert C (10 tons) –Note: These numbers are huge compared to Block I! – Ranks: B = 1st, A = 2nd, C = 3rd The Magic: The Friedman test ignores the enormous difference between “3 tons” and “30 tons.” It only sees that Fertilizer B keeps finishing in 1st place. 11.7.2 The Code (The Vertical Bar) The syntax is unique. We use a vertical bar (|) to separate the Treatment from the Block. Translation: “Analyze Yield by Treatment, grouped within Block.” Your Code # Syntax: friedman.test(Yield ~ Treatment | Block, data = my_data) friedman.test(my_data$Yield ~ my_data$Treatment | my_data$Block) 121 11.8. THE “WHO WON?” STEP: POST-HOC FOR ROAD B 11.7.3 Reading the Output R Output Friedman rank sum test data: my_data$Yield, my_data$Treatment, and my_data$Block Friedman chi-squared = 10.8, df = 2, p-value = 0.0045 The Decision: •p < 0.05: Significant — at least one treatment consistently ranks higher. (Proceed to Post-Hoc). •p > 0.05: Not Significant — the rankings appear random. 11.8 The “Who Won?” Step: Post-Hoc for Road B If your Friedman test was significant (p < 0.05), you know that at least one treatment is different. However, Friedman does not tell you which treatment differs. To find the winner, we use the Pairwise Wilcoxon Test. 11.8.1 The Logic: Why not just separate tests? You might be thinking: “Why can’t I just run a Wilcoxon test for A vs. B, and another for B vs. C?” The Answer: The “Lottery Ticket” Problem. Every time you run a statistical test, you accept a small risk (usually 5%) of making a mistake—finding a difference that is not real. •If you run one test, your risk is 5%. •If you run three tests, your risk becomes roughly 15%. •If you run ten tests, the risk becomes very large. It works like buying lottery tickets: the more tests you run, the more likely you are to “win” a fake significant result by pure chance. 11.8.2 The Solution: The Penalty Card (Bonferroni) To solve this, we apply a mathematical penalty called the Bonferroni Adjustment. It makes the significance level stricter, which cancels the “luck factor” and ensures your results are trustworthy. 122 11.8. THE “WHO WON?” STEP: POST-HOC FOR ROAD B 11.8.3 The Code Your Code # Syntax: pairwise.wilcox.test(Yield, Treatment, p.adjust = “bonferroni”) pairwise.wilcox.test(my_data$Yield, my_data$Treatment, p.adjust.method = “bonferroni”) 11.8.4 Reading the Output (The Matrix) R produces a triangular grid of p-values. R Output Pairwise comparisons using Wilcoxon rank sum test Fert_A Fert_B Fert_B 1.000 - Fert_C 0.004 0.002 How to read the grid: •Fert_A vs. Fert_B: p = 1.000 (Not significant — they are equal). •Fert_A vs. Fert_C: p = 0.004 (Significant — C differs from A). •Fert_B vs. Fert_C: p = 0.002 (Significant — C differs from B). Conclusion: Fertilizer C is the clear winner. Fertilizers A and B are statistically the same and lower performing. 123 11.9. TEST DRIVE: REAL-WORLD EXAMPLES 11.9 Test Drive: Real-World Examples You have the keys to the 4Œ4 (RCBD). Now lets drive it across real agricultural terrains. Here are four practical scenarios you will face in the field. 11.9.1 Scenario 1: The Hillside Trial (Standard RCBD) The Situation: You tested 3 maize varieties on a sloping field. The Blocks: You divided the hill into Top, Middle, and Bottom blocks. The Check: Data is Normal (p > 0.05). The Plan: RCBD ANOVA →Tukeys HSD. Figure 11.3: Scenario 1: Blocking by elevation removes the slope effect. Your Code model <- aov(Yield ~ Variety + Block, data = my_data) summary(model) R Output Df Sum Sq Mean Sq F value Pr(>F) Variety 2 150.2 75.1 12.5 0.001 ** Block 2 85.5 42.7 7.1 0.004 ** Interpretation: •Variety (p= 0.001): Significant. Varieties differ. Proceed to Tukey. •Block (p= 0.004): Significant. The hill affected yield. Good blocking choice! 124 11.9. TEST DRIVE: REAL-WORLD EXAMPLES 11.9.2 Scenario 2: The Old vs. Young Cow Trial (Friedman) The Situation: You tested 3 diets (A, B, C) on dairy cows. The Problem: Young cows eat less; old cows eat more. The Blocks: You grouped cows by age (Heifers, 1st Lactation, Mature). The Check: ShapiroWilk p= 0.001 (Not Normal). The Plan: Friedman Test Pairwise Wilcoxon. Figure 11.4: Scenario 2: Blocking by age ensures fair comparison of diets. Your Code friedman.test(Milk ~ Diet | AgeGroup, data = my_data) R Output Friedman rank sum test Friedman chi-squared = 9.8, df = 2, p-value = 0.007 Interpretation: The p-value is 0.007 (<0.05) — the diets differ significantly after accounting for age. 125 12.4. THE ENGINE: THE MAGIC ASTERISK (*) Your Code str(my_data) R Output (Success) $ Variety : Factor w/ 2 levels "V1","V2": 1 1 2 ... $ Fertilizer : Factor w/ 2 levels "N0","N1": 1 2 1 ... $ Block : Factor w/ 3 levels "1","2","3": 1 1 1 ... Now you are ready to start the engine. 12.4 The Engine: The Magic Asterisk (*) How do we tell R to check for an interaction? We use the asterisk (*) symbol. The Logic: •Variety + Fertilizer = Check them separately. •Variety * Fertilizer = Check them separately AND check whether they interact. 12.4.1 The Code Your Code # Scenario A: Uniform Field (CRD) model <- aov(Yield ~ Variety * Fertilizer, data = my_data) # Scenario B: Blocked Field (RCBD) model <- aov(Yield ~ Variety * Fertilizer + Block, data = my_data) summary(model) 12.5 Reading the Dashboard (The Hierarchy Rule) R will give you three P-values (plus one for Block if included). 132 12.6. THE “PHOTO”: THE INTERACTION PLOT R Output Df Sum Sq Mean Sq F value Pr(><F) Variety 2 15.4 7.7 12.5 0.001 ** Fertilizer 1 45.2 45.2 55.1 0.000 *** Var:Fert 2 8.5 4.2 6.1 0.004 ** Residuals 24 ... 12.5.1 The Golden Rule: Look at the Bottom First In a Two-Way ANOVA, you must follow a strict hierarchy. Always look at the Interaction Row first (Var:Fert). THE DECISION TREE Step 1: Check the Interaction (Var:Fert) p-value. Scenario A: It is significant (p < 0.05). Action: Ignore the rows above it. The main effects depend on each other. The numbers here can’t tell you the full story. You must visualize the result to understand it (see Next Section). Scenario B: It is not significant (p > 0.05). Action: The interaction did not occur. You can now safely interpret the Variety and Fertilizer rows separately (just like a normal ANOVA). 12.6 The “Photo”: The Interaction Plot If your interaction was significant, you must see it to understand it. We need a graph that shows how the two factors behave together. 12.6.1 The Code (interaction.plot) R has a built-in tool that is very easy to use for this job. 133 12.6. THE “PHOTO”: THE INTERACTION PLOT Your Code # Syntax: interaction.plot(x.factor, trace.factor, response) interaction.plot( x.factor = my_data$Fertilizer, trace.factor = my_data$Variety, response = my_data$Yield ) Figure 12.2: The Three Faces of Interaction. Left: Parallel lines mean No Interaction (Main effects are independent). Center: Crossing lines mean Strong Interaction (Main effects are misleading). Right: Non-parallel lines show Interaction (The magnitude of the effect changes). 12.6.2 Reading the Photo •Parallel Lines: No interaction. The fertilizer works the same for both varieties. •Crossing Lines (X-shape): Strong interaction. The fertilizer reverses the ranking (good for one, bad for another). •Non-Parallel Lines (< or > shape): Interaction. The fertilizer helps one variety more than the other. 134 12.7. THE FINISH LINE: POST-HOC FOR FACTORIALS 12.7 The Finish Line: Post-Hoc for Factorials This is where 90% of students get stuck. How you run the post-hoc depends entirely on the Decision Tree from Section ??. 12.7.1 Scenario A: Interaction IS Significant (p < 0.05) The situation: The lines crossed. The rule: You cannot test “Variety” alone — you must test the combinations. The goal: Compare VarA:N0 vs VarA:N1 vs VarB:N0 ... You have two common choices for the test (check with your supervisor): Option 1: Tukey’s HSD (The gold standard) Your code library(agricolae) # We list BOTH factors inside c() HSD.test(model, c("Variety", "Fertilizer"), console = TRUE) Example output you will see: R output Study: Yield ~ Variety * Fertilizer HSD Test for Yield ... groups Yield groups Var_A:High 65.2 a Var_B:High 50.1 b Var_B:Low 48.5 b Var_A:Low 20.2 c Interpretation: •Var_A with High Fert is the clear winner (group “a”). •Var_A with Low Fert is the clear loser (group “c”). •Conclusion: Variety A is risky — excellent when fertilized, poor when not. Option 2: DMRT (Duncan’s test) Use this if your supervisor prefers Duncan’s test. The code is almost identical: 135 12.7. THE FINISH LINE: POST-HOC FOR FACTORIALS Your code duncan.test(model, c("Variety", "Fertilizer"), console = TRUE) 12.7.2 Scenario B: Interaction is NOT Significant (p > 0.05) The Situation: The lines were parallel. The factors act independently. The Rule: You can separate them. You do not need to test the combinations. The Goal: Find the best Variety (ignoring fertilizer) AND the best Fertilizer (ignoring variety). You run two separate tests. You may use either Tukey or DMRT. Option 1: Tukey’s HSD (The Standard) Your Code # Test 1: Compare Varieties HSD.test(model, "Variety", console = TRUE) # Test 2: Compare Fertilizers HSD.test(model, "Fertilizer", console = TRUE) Option 2: DMRT (Duncan’s Test) Use this if your professor prefers Duncan’s grouping letters. Your Code # Test 1: Compare Varieties duncan.test(model, "Variety", console = TRUE) # Test 2: Compare Fertilizers duncan.test(model, "Fertilizer", console = TRUE) The Output: You will get two separate tables. •Table 1 (Variety): e.g., Var A is “a” and Var B is “b” (A is better overall). •Table 2 (Fertilizer): e.g., High is “a” and Low is “b” (High performs better overall). 136 12.8. TEST DRIVE: REAL-WORLD EXAMPLES 12.8 Test Drive: Real-World Examples You have the keys to the "Combo Car." Now let’s drive it. Here are four specific scenarios you will face in the field. 12.8.1 Scenario 1: The "Magic Combination" (Significant Interaction) The Situation: You tested 2 Maize Varieties (Hybrid, Local) at 2 Nitrogen levels (High, Low). The Result: The Hybrid loved the Nitrogen, but the Local variety fell over (lodged) with too much Nitrogen. R Output Variety p = 0.001 ** Nitrogen p = 0.001 ** Var:Nitrogen p = 0.003 ** (Significant) The Interpretation: •The lines cross. The interaction is significant (p < 0.05). •Golden Rule: Ignore the main effects. •Conclusion: High Nitrogen is good, but only if you use the Hybrid variety. 137 12.8. TEST DRIVE: REAL-WORLD EXAMPLES 12.8.2 Scenario 2: The "Independent Effects" (No Interaction) The Situation: You tested 3 Spacing distances (S1, S2, S3) and 2 Weeding methods (Hand, Chemical). The Result: Chemical weeding was better than Hand weeding, and S1 was better than S3. They did not interfere with each other. R Output Spacing p = 0.002 ** Weeding p = 0.040 * Space:Weed p = 0.650 (Not Sig) The Interpretation: •The lines are parallel. Interaction is not significant (p > 0.05). •Conclusion: You can make two separate statements: 1. Chemical weeding is better. 2. Spacing S1 is better. The combinations do not need to be checked. 138 12.8. TEST DRIVE: REAL-WORLD EXAMPLES 12.8.3 Scenario 3: The "Slope" Factorial (RCBD) The Situation: You tested 2 Varieties ×2 Fertilizers. The Field: The field had a slope, so you used Blocks. Your Code model <- aov(Yield ~ Var * Fert + Block, data = my_data) R Output Var p = 0.020 * Fert p = 0.030 * Block p = 0.004 ** (Significant) Var:Fert p = 0.002 ** (Significant) The Interpretation: •Block (p= 0.004): The slope matteredgood thing you included it. •Interaction (p= 0.002): The varieties responded differently to the fertilizer. •Action: Run Tukey on the combinations (Var:Fert). 139 12.8. TEST DRIVE: REAL-WORLD EXAMPLES 12.8.4 Scenario 4: The "Null" Result (Nothing Worked) The Situation: You tested 2 Pruning methods on 2 Mango Varieties to see if sweetness (Brix) increased. R Output Pruning p = 0.450 Variety p = 0.320 Prun:Var p = 0.890 The Interpretation: •Interaction (p= 0.890): No interaction. •Main Effects (p > 0.05): No difference in pruning or variety. •Conclusion: Neither pruning method nor variety affects mango sweetness. 12.8.5 Summary Checklist for the Combo Car •Did you convert all columns to factors? •Did you use the asterisk (*) in your code? •Did you check the interaction p-value first? •Did you make an interaction plot? If yes, you are done. Congratulations! 140 PART 3: “Roadside Assistance” (The Glovebox) 141