{"id":436,"date":"2023-03-21T10:53:32","date_gmt":"2023-03-21T10:53:32","guid":{"rendered":"https:\/\/www.editage.com\/blog\/?p=436"},"modified":"2026-07-31T22:03:40","modified_gmt":"2026-07-31T16:33:40","slug":"hypothesis-testing-different-types-for-biomedical-researchers","status":"publish","type":"post","link":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/","title":{"rendered":"Hypothesis Testing &#038; NHST: Definition, Steps, Tips, Examples"},"content":{"rendered":"\r\n<p><strong>Key takeaways:<\/strong><\/p>\r\n<ul class=\"[li_&amp;]:mb-0 [li_&amp;]:mt-1 [li_&amp;]:gap-1 [&amp;:not(:last-child)_ul]:pb-1 [&amp;:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1\" dir=\"ltr\">\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">Hypothesis testing infers something about a population from a sample; NHST is the specific procedure behind most published statistics. In NHST, you assume no effect, compute a test statistic, and judge how improbable the observed data would be if that assumption held. It quantifies uncertainty rather than eliminating it.<\/li>\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">The null hypothesis carries the burden of proof because it&#8217;s the narrow, falsifiable claim. You never confirm the alternative; you only establish that chance alone struggles to account for what you saw, and a non-significant result is not evidence that no effect exists.<\/li>\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">The p value is routinely misread. It&#8217;s the probability of data this extreme <em>assuming<\/em> the null is true, not the probability the null is true, not a measure of effect magnitude, and not a guarantee of replication. The 0.05 cutoff is convention, and biomedical research that has crucial consequences for patients often demands a stricter threshold.<\/li>\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">Two errors trade off against each other: false positives (rate \u03b1) and false negatives (rate \u03b2). Power, usually targeted at 80%, is your chance of catching a real effect. The only way to reduce both error types at once is a larger sample.<\/li>\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">Test choice follows the data structure: z-test or one-sample t-test against a known value, independent t-test for two unrelated groups, paired t-test for repeated measurement of the same subjects, chi-square for categorical associations, ANOVA for three or more group means, with non-parametric alternatives when normality fails.<\/li>\r\n<li class=\"font-claude-response-body whitespace-normal break-words pl-2\">The framework&#8217;s well-known weaknesses \u2014 binary significant\/not-significant thinking, p-hacking, publication bias, and confusing statistical with clinical significance \u2014 are mitigated by pre-registration, adequate power, replication, and always reporting effect sizes and confidence intervals alongside p values.<\/li>\r\n<\/ul>\r\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_85 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Glossary_of_Key_Terms\" >Glossary of Key Terms<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_Is_Hypothesis_Testing\" >What Is Hypothesis Testing?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_Is_Null_Hypothesis_Significance_Testing_NHST\" >What Is Null Hypothesis Significance Testing (NHST)?<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#The_Four_Core_Components_of_NHST\" >The Four Core Components of NHST<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#How_NHST_Differs_from_Simply_%E2%80%9CTesting_a_Hypothesis%E2%80%9D\" >How NHST Differs from Simply &#8220;Testing a Hypothesis&#8221;<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Why_%E2%80%9CNull_Hypothesis%E2%80%9D_and_Not_Just_%E2%80%9CHypothesis%E2%80%9D\" >Why &#8220;Null Hypothesis&#8221; and Not Just &#8220;Hypothesis&#8221;?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#The_Two_Schools_of_Thought_Behind_NHST\" >The Two Schools of Thought Behind NHST<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Core_Concepts_and_Key_Terms\" >Core Concepts and Key Terms<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_are_Null_and_Alternative_Hypotheses\" >What are Null and Alternative Hypotheses?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_is_the_p_value\" >What is the p value?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Significance_Level_%CE%B1\" >Significance Level (\u03b1)<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#One-Tailed_vs_Two-Tailed_Tests\" >One-Tailed vs. Two-Tailed Tests<\/a><ul class='ez-toc-list-level-4' ><li class='ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_1\" >Example 1<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_2\" >Example 2<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Types_of_Statistical_Tests_in_NHST\" >Types of Statistical Tests in NHST<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Step-by-Step_Guide_to_Hypothesis_Testing\" >Step-by-Step Guide to Hypothesis Testing<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Worked_Examples\" >Worked Examples<\/a><ul class='ez-toc-list-level-4' ><li class='ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_1_Does_a_New_Drug_Lower_Blood_Pressure\" >Example 1: Does a New Drug Lower Blood Pressure?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_2_Gender_and_Voting_Preferences\" >Example 2: Gender and Voting Preferences<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-20\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_3_Comparing_Anxiety_Across_Three_Therapy_Modalities\" >Example 3: Comparing Anxiety Across Three Therapy Modalities<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-21\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Example_4_Exact_Binomial_Test_for_Treatment_Efficacy\" >Example 4: Exact Binomial Test for Treatment Efficacy<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-22\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_are_Type_I_and_Type_II_Errors\" >What are Type I and Type II Errors?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-23\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Common_Misconceptions_About_P_Values\" >Common Misconceptions About P Values<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-24\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Limitations_of_Hypothesis_Testing\" >Limitations of Hypothesis Testing<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-25\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Best_practices_to_mitigate_these_limitations\" >Best practices to mitigate these limitations:<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-26\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Frequently_Asked_Questions_FAQs\" >Frequently Asked Questions (FAQs)<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-27\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_is_a_null_hypothesis\" >What is a null hypothesis?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-28\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_is_statistical_power_and_why_does_it_matter\" >What is statistical power, and why does it matter?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-29\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#How_is_a_confidence_interval_related_to_a_hypothesis_test\" >How is a confidence interval related to a hypothesis test?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-30\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_does_it_mean_to_%E2%80%9Cpre-register%E2%80%9D_a_study\" >What does it mean to &#8220;pre-register&#8221; a study?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-31\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#When_should_I_use_a_non-parametric_test_instead_of_a_t-test_or_ANOVA\" >When should I use a non-parametric test instead of a t-test or ANOVA?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-32\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_is_effect_size_and_which_measures_are_commonly_reported\" >What is effect size, and which measures are commonly reported?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-33\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#What_is_the_difference_between_a_one-sample_and_two-sample_test\" >What is the difference between a one-sample and two-sample test?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-34\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Why_has_NHST_been_criticised_and_what_are_the_proposed_alternatives\" >Why has NHST been criticised, and what are the proposed alternatives?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-35\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#Can_hypothesis_testing_be_used_with_observational_data_or_only_with_experiments\" >Can hypothesis testing be used with observational data, or only with experiments?<\/a><\/li><\/ul><\/li><\/ul><\/nav><\/div>\n<h2 class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\"><span class=\"ez-toc-section\" id=\"Glossary_of_Key_Terms\"><\/span><strong>Glossary of Key Terms<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n<div class=\"overflow-x-auto w-full px-2 mb-6 print:overflow-x-visible\" dir=\"ltr\">\r\n<table class=\"min-w-full border-collapse text-sm leading-[1.7] whitespace-normal\">\r\n<thead class=\"text-left\">\r\n<tr>\r\n<th class=\"text-text-100 border-b-0.5 border-[hsl(var(--border-300)\/0.6)] py-2 pr-4 align-top font-bold\" scope=\"col\">Term<\/th>\r\n<th class=\"text-text-100 border-b-0.5 border-[hsl(var(--border-300)\/0.6)] py-2 pr-4 align-top font-bold\" scope=\"col\">Meaning<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Hypothesis testing<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Using sample data to judge whether evidence supports a claim about a population<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">NHST<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Null Hypothesis Significance Testing \u2014 the formal procedure of assuming no effect, then testing how unlikely the data are under that assumption<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Null hypothesis (H\u2080)<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">The default position of no effect, no difference, no relationship<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Alternative hypothesis (H\u2081 \/ H\u2090)<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">The researcher&#8217;s claim that something real is happening<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Test statistic<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">A value (z, t, \u03c7\u00b2, F) expressing how far the sample result sits from what H\u2080 predicts, in standard-error units<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">P value<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Probability of obtaining results this extreme or more so, given that H\u2080 is true<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Significance level (\u03b1)<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">The pre-specified false-positive rate you&#8217;ll tolerate \u2014 commonly 0.05, often stricter in clinical work<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Critical value<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">The test-statistic cutoff separating &#8220;reject&#8221; from &#8220;fail to reject&#8221;<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Degrees of freedom (df)<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">A sample-size-linked count that selects the correct reference distribution<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">One-tailed test<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Used when the hypothesis specifies a direction (increase or decrease only)<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Two-tailed test<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Used when a difference in either direction would be of interest<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Type I error<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">False positive \u2014 rejecting a true null; occurs at rate \u03b1<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Type II error<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">False negative \u2014 failing to reject a false null; occurs at rate \u03b2<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Statistical power (1 \u2212 \u03b2)<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Probability of detecting a real effect; conventionally planned at 0.80<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Effect size<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Magnitude of a difference or association independent of sample size \u2014 Cohen&#8217;s d, Pearson&#8217;s r, \u03b7\u00b2, odds ratios<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Confidence interval<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">A range conveying both direction and magnitude; a 95% CI excluding the null value corresponds to p &lt; 0.05 two-tailed<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Z-test<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Continuous data, large samples, known population SD<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">t-test<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Continuous data with unknown SD \u2014 one-sample, independent-samples, or paired versions<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Chi-square test<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Tests association between categorical variables<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">ANOVA<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Compares means across three or more groups<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Non-parametric tests<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Distribution-free alternatives (Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis) for skewed or small-sample data<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Fisher school<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Treats the p value as a continuous measure of evidence, with no fixed cutoff<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Neyman-Pearson school<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Pre-specifies \u03b1 and makes a binary decision, controlling long-run error rates<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Bayesian approach<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Updates prior beliefs with data, reporting posterior probabilities or Bayes factors<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">P-hacking<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Exploiting analytic flexibility until a result crosses significance, inflating false positives<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Publication bias<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Over-representation of positive findings because null results go unpublished (the file-drawer problem)<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Pre-registration<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Publicly recording hypotheses and analysis plans before data collection, e.g. on OSF or ClinicalTrials.gov<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">HARKing<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Hypothesising After Results are Known \u2014 retrofitting hypotheses to observed data<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Equivalence testing<\/td>\r\n<td class=\"border-b-0.5 border-[hsl(var(--border-300)\/0.3)] py-2 pr-4 align-top\">Demonstrating an effect is small enough to be unimportant, rather than merely non-significant<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"What_Is_Hypothesis_Testing\"><\/span><strong>What Is Hypothesis Testing?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>Hypothesis testing is a statistical method used to determine whether there is enough evidence in sample data to draw conclusions about a population. Instead of collecting data from an entire population, you take a sample and test whether the evidence supports or contradicts an assumption about that population.<\/p>\r\n\r\n\r\n\r\n<p>In everyday terms: you have a hunch, you collect data, and you ask whether the data are consistent with &#8220;nothing is going on&#8221; or whether they strain credulity enough to suggest something real is happening.<\/p>\r\n\r\n\r\n\r\n<p>For example, if a company says its website gets 50 visitors each day on average, hypothesis testing can be used to look at past visitor data and see if this claim is true or if the actual number is different. <a href=\"https:\/\/www.geeksforgeeks.org\/data-science\/understanding-hypothesis-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<p>In the social and biomedical sciences, the stakes are higher. We use NHST to ask things like:<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li>Does a new antidepressant actually reduce symptoms more than a placebo?<\/li>\r\n<li>Do boys and girls differ in academic self-efficacy?<\/li>\r\n<li>Does social isolation increase mortality risk?<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"What_Is_Null_Hypothesis_Significance_Testing_NHST\"><\/span><strong>What Is Null Hypothesis Significance Testing (NHST)?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>Null Hypothesis Significance Testing is the specific, formalised framework that underlies the vast majority of statistical tests published in scientific journals. While &#8220;hypothesis testing&#8221; is a broad term covering many philosophies, NHST refers to one particular procedure: you begin by assuming the null hypothesis is true, collect data, compute a test statistic, and then ask how probable your observed result would be under that assumption.<\/p>\r\n\r\n\r\n\r\n<p>The word <em>null<\/em> is key. It does not mean &#8220;zero&#8221; in a casual sense \u2014 it means the hypothesis of <em>no effect, no difference, no relationship<\/em>. Everything in NHST is organised around building a case against this default position.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"The_Four_Core_Components_of_NHST\"><\/span><strong>The Four Core Components of NHST<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Component<\/strong><\/td>\r\n<td><strong>Role in NHST<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td><strong>Null hypothesis (H\u2080)<\/strong><\/td>\r\n<td>The default claim of &#8220;no effect&#8221; that must be disproven<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Alternative hypothesis (H\u2081)<\/strong><\/td>\r\n<td>The research claim the investigator hopes the data will support<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Test statistic<\/strong><\/td>\r\n<td>A number summarising how far the sample result is from what H\u2080 predicts<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/p-value-statistics-hypothesis-testing-definition-meaning\/\" target=\"_blank\" rel=\"noreferrer noopener\">P value<\/a><\/strong><\/td>\r\n<td>The probability of observing a result this extreme or more extreme, <em>given that H\u2080 is true<\/em><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"How_NHST_Differs_from_Simply_%E2%80%9CTesting_a_Hypothesis%E2%80%9D\"><\/span><strong>How NHST Differs from Simply &#8220;Testing a Hypothesis&#8221;<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>It is worth distinguishing NHST from the broader scientific notion of hypothesis testing. Scientists test hypotheses all the time through prediction, experimentation, and observation. NHST is specifically a <em>probabilistic decision rule<\/em> applied to sample data. Its output is not a verdict on whether a theory is correct; it is a statement about whether the data are unusual enough, under a particular null model, to warrant further scrutiny. A p value below 0.05 is not a discovery but instead it is a signal that the null hypothesis struggles to explain your data.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Why_%E2%80%9CNull_Hypothesis%E2%80%9D_and_Not_Just_%E2%80%9CHypothesis%E2%80%9D\"><\/span><strong>Why &#8220;Null Hypothesis&#8221; and Not Just &#8220;Hypothesis&#8221;?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>The null is set up as a straw man precisely because it is falsifiable in a probabilistic sense. You cannot prove that a drug <em>works<\/em>: there are infinite ways it could work, at varying magnitudes. But you can ask a narrow, testable question: <em>is the observed improvement consistent with pure chance?<\/em> If the answer is &#8220;barely,&#8221; you reject the null. The burden of proof sits firmly with the null, and the researcher accumulates evidence against it.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"The_Two_Schools_of_Thought_Behind_NHST\"><\/span><strong>The Two Schools of Thought Behind NHST<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>There are two classical schools of thought on how best to use the <a href=\"https:\/\/www.editage.com\/blog\/p-value-statistics-hypothesis-testing-definition-meaning\/\">p-value<\/a>: the Fisher school and the Neyman-Pearson school. There is also a <a href=\"https:\/\/www.editage.com\/insights\/10-steps-to-get-started-with-bayesian-statistics-in-biomedical-research\">Bayesian way<\/a> to interpret the p-value, but that presents a whole other set of dilemmas. <a href=\"https:\/\/www.sjsu.edu\/faculty\/gerstman\/StatPrimer\/hyp-test.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>School<\/strong><\/td>\r\n<td><strong>Core Idea<\/strong><\/td>\r\n<td><strong>How Decision Is Made<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td><strong>Fisher<\/strong><\/td>\r\n<td>P value = continuous measure of evidence against H\u2080<\/td>\r\n<td>Smaller p = stronger evidence; no fixed threshold<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Neyman-Pearson<\/strong><\/td>\r\n<td>Pre-specify \u03b1; control long-run error rates<\/td>\r\n<td>Reject or don&#8217;t reject based on \u03b1 threshold<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Bayesian<\/strong><\/td>\r\n<td>Update prior beliefs with new data<\/td>\r\n<td>Posterior probability, Bayes factors<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<p>Modern practice, especially in journals, blends Fisher and Neyman-Pearson, often awkwardly. Understanding which framework you are working in matters enormously for interpretation.<\/p>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Core_Concepts_and_Key_Terms\"><\/span><strong>Core Concepts and Key Terms<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_are_Null_and_Alternative_Hypotheses\"><\/span><strong>What are Null and Alternative Hypotheses?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>The null hypothesis (H\u2080) is a statement of &#8220;no difference,&#8221; &#8220;no association,&#8221; or &#8220;no treatment effect.&#8221; The alternative hypothesis (H\u2090) is a statement of &#8220;difference,&#8221; &#8220;association,&#8221; or &#8220;treatment effect.&#8221; H\u2080 is assumed to be true until proven otherwise. However, H\u2090 is the hypothesis the researcher hopes to bolster. <a href=\"https:\/\/www.sjsu.edu\/faculty\/gerstman\/StatPrimer\/hyp-test.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Term<\/strong><\/td>\r\n<td><strong>Symbol<\/strong><\/td>\r\n<td><strong>Plain-English Meaning<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td>Null hypothesis<\/td>\r\n<td>H\u2080<\/td>\r\n<td>&#8220;Nothing is going on; any observed difference is chance&#8221;<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Alternative hypothesis<\/td>\r\n<td>H\u2081 \/ H\u2090<\/td>\r\n<td>&#8220;Something real is happening&#8221;<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Significance level<\/td>\r\n<td>\u03b1<\/td>\r\n<td>The false-positive rate you are willing to tolerate<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>P value<\/td>\r\n<td>p<\/td>\r\n<td>Probability of observing these data (or more extreme) if H\u2080 were true<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Test statistic<\/td>\r\n<td>Z, t, \u03c7\u00b2, F<\/td>\r\n<td>How many standard errors your sample result sits from H\u2080<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Critical value<\/td>\r\n<td>\u2014<\/td>\r\n<td>The test-statistic threshold that demarcates &#8220;reject&#8221; from &#8220;fail to reject&#8221;<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Degrees of freedom<\/td>\r\n<td>df<\/td>\r\n<td>A count tied to sample size; used to find the correct reference distribution<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_is_the_p_value\"><\/span><strong>What is the p value?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>The P value answers the question: &#8220;If the null hypothesis were true, what is the probability of observing the current data or data that is more extreme?&#8221; Note that the P value is NOT the probability that the hypothesis (or any other hypothesis) is right or wrong. In fact, it assumes the null hypothesis is right! <a href=\"https:\/\/www.sjsu.edu\/faculty\/gerstman\/StatPrimer\/hyp-test.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<p>This distinction is crucial and perpetually misunderstood. More on it in the misconceptions section below.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Significance_Level_%CE%B1\"><\/span><strong>Significance Level (\u03b1)<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>The significance level (\u03b1) represents how sure we want to be before saying the claim is false. Usually, we choose 0.05 (5%). Choosing \u03b1 = 0.05 means accepting a 5% chance of wrongly rejecting a true null hypothesis, i.e., a false alarm. <a href=\"https:\/\/www.geeksforgeeks.org\/data-science\/understanding-hypothesis-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<p>In biomedical contexts where a wrong decision could harm patients, researchers often set \u03b1 = 0.01 or even 0.001.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"One-Tailed_vs_Two-Tailed_Tests\"><\/span><strong>One-Tailed vs. Two-Tailed Tests<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>A one-tailed test is used when we expect a change in only one direction: either up or down, but not both. A two-tailed test is used when we want to see if there is a difference in either direction, higher or lower.<\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Test Type<\/strong><\/td>\r\n<td><strong>When to Use<\/strong><\/td>\r\n<td><strong>Example Hypothesis<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td><strong>Right-tailed<\/strong><\/td>\r\n<td>Expecting an increase<\/td>\r\n<td>H\u2081: \u03bc &gt; 50<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Left-tailed<\/strong><\/td>\r\n<td>Expecting a decrease<\/td>\r\n<td>H\u2081: \u03bc &lt; 50<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Two-tailed<\/strong><\/td>\r\n<td>Any difference, direction unknown<\/td>\r\n<td>H\u2081: \u03bc \u2260 50<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_1\"><\/span><strong>Example 1<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n<p>A sociologist testing whether immigrants score <em>differently<\/em> (not just higher or lower) on a civic knowledge test compared to native-born citizens would use a two-tailed test, since the direction of difference is theoretically uncertain.<\/p>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_2\"><\/span>Example 2<span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n<p>A pharmacologist testing whether a new antihypertensive <em>lowers<\/em> blood pressure (not raises it) would use a one-tailed (left-tailed) test.<\/p>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Types_of_Statistical_Tests_in_NHST\"><\/span><strong>Types of Statistical Tests in NHST<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>Choosing the wrong test is one of the most common errors in applied research. The decision depends on the type of data (continuous vs. categorical), the number of groups, and whether population variance is known.<\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Test<\/strong><\/td>\r\n<td><strong>Data Type<\/strong><\/td>\r\n<td><strong>Groups<\/strong><\/td>\r\n<td><strong>When to Use<\/strong><\/td>\r\n<td><strong>Example<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td><strong>Z-test<\/strong><\/td>\r\n<td>Continuous<\/td>\r\n<td>1 or 2<\/td>\r\n<td>Large sample (n &gt; 30), known population SD<\/td>\r\n<td>Comparing national exam mean to a known standard<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/t-test-definition-assumptions-formula-calculation\/\">One-sample t-test<\/a><\/strong><\/td>\r\n<td>Continuous<\/td>\r\n<td>1<\/td>\r\n<td>Small sample, unknown population SD<\/td>\r\n<td>Testing if a clinic&#8217;s mean wait time differs from 30 min<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/t-test-definition-assumptions-formula-calculation\/\">Independent samples t-test<\/a><\/strong><\/td>\r\n<td>Continuous<\/td>\r\n<td>2<\/td>\r\n<td>Comparing means of two unrelated groups<\/td>\r\n<td>Depression scores in therapy group vs. control<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/t-test-definition-assumptions-formula-calculation\/\">Paired t-test<\/a><\/strong><\/td>\r\n<td>Continuous<\/td>\r\n<td>2 (related)<\/td>\r\n<td>Same subjects measured twice<\/td>\r\n<td>Blood pressure before vs. after drug<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/chi-square-test-types-explained-for-biomedical-researchers\/\">Chi-square test<\/a><\/strong><\/td>\r\n<td>Categorical<\/td>\r\n<td>2+<\/td>\r\n<td>Association between categorical variables<\/td>\r\n<td>Gender vs. vaccine hesitancy (Yes\/No)<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong><a href=\"https:\/\/www.editage.com\/blog\/anova-types-uses-assumptions-a-quick-guide-for-biomedical-researchers\/\">ANOVA<\/a><\/strong><\/td>\r\n<td>Continuous<\/td>\r\n<td>3+<\/td>\r\n<td>Comparing means of \u22653 groups<\/td>\r\n<td>Anxiety scores across 3 therapy modalities<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>One-tailed tests<\/strong><\/td>\r\n<td>Any<\/td>\r\n<td>Any<\/td>\r\n<td>Directional hypothesis is pre-specified<\/td>\r\n<td>New drug expected to reduce tumour size<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Step-by-Step_Guide_to_Hypothesis_Testing\"><\/span><strong>Step-by-Step Guide to Hypothesis Testing<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>Every NHST follows the same logical structure. Below is the canonical seven-step procedure:<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>Step 1: State the hypotheses.<\/strong> Define H\u2080 and H\u2081 in precise, testable terms before looking at the data.<\/li>\r\n<li><strong>Step 2: Choose the significance level (\u03b1).<\/strong> Pre-specify \u03b1, usually 0.05. Changing it after seeing results invalidates the test.<\/li>\r\n<li><strong>Step 3: Select the appropriate statistical test.<\/strong> Match the test to your data structure (see table above).<\/li>\r\n<li><strong>Step 4: Collect and organize the data.<\/strong> Gather a representative sample. Poor data quality produces misleading p values regardless of the test.<\/li>\r\n<li><strong>Step 5: Compute the test statistic.<\/strong> Calculate how far your sample result lies from what H\u2080 predicts, in units of standard error.<\/li>\r\n<li><strong>Step 6: Determine the p value and make a decision.<\/strong> If p-value \u2264 \u03b1 \u2192 reject H\u2080. If p-value &gt; \u03b1 \u2192 insufficient evidence to reject H\u2080, which is not proof that H\u2080 is true. <a href=\"https:\/\/www.geeksforgeeks.org\/data-science\/understanding-hypothesis-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/li>\r\n<li><strong>Step 7: Interpret results in plain language.<\/strong> Report the effect size, direction of difference, and p value. State the conclusion in the context of the original research question.<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Worked_Examples\"><\/span><strong>Worked Examples<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_1_Does_a_New_Drug_Lower_Blood_Pressure\"><\/span><strong>Example 1: Does a New Drug Lower Blood Pressure?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n\r\n\r\n\r\n<p>A pharmaceutical team recruits 10 hypertensive patients and measures systolic blood pressure before and after a 4-week course of a new antihypertensive.<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>H\u2080:<\/strong> The drug has no effect on blood pressure (mean difference = 0)<\/li>\r\n<li><strong>H\u2081:<\/strong> The drug reduces blood pressure (mean difference &lt; 0)<\/li>\r\n<li><strong>Test:<\/strong> Paired t-test (same patients measured twice)<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<p>Using a paired t-test, with before-treatment values averaging around 122 mmHg and after-treatment values around 117 mmHg, the t-statistic is approximately -9. With degrees of freedom = 9, the p-value is approximately 0.0000085: far below the significance threshold of 0.05. The researchers reject the null hypothesis. There is statistically significant evidence that the average blood pressure before and after treatment differs. <a href=\"https:\/\/www.geeksforgeeks.org\/data-science\/understanding-hypothesis-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_2_Gender_and_Voting_Preferences\"><\/span><strong>Example 2: Gender and Voting Preferences<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n\r\n\r\n\r\n<p>A political scientist wants to know whether gender and voting preference (Candidate A vs. Candidate B) are related in a random sample of 400 voters.<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>H\u2080:<\/strong> Gender and voting preference are independent<\/li>\r\n<li><strong>H\u2081:<\/strong> Gender and voting preference are associated<\/li>\r\n<li><strong>Test:<\/strong> Chi-square test of independence<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td>\u00a0<\/td>\r\n<td><strong>Votes for A<\/strong><\/td>\r\n<td><strong>Votes for B<\/strong><\/td>\r\n<td><strong>Total<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td>Men<\/td>\r\n<td>95<\/td>\r\n<td>105<\/td>\r\n<td>200<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>Women<\/td>\r\n<td>130<\/td>\r\n<td>70<\/td>\r\n<td>200<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Total<\/strong><\/td>\r\n<td><strong>225<\/strong><\/td>\r\n<td><strong>175<\/strong><\/td>\r\n<td><strong>400<\/strong><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<p>If the chi-square statistic yields p = 0.003 &lt; 0.05, H\u2080 is rejected. The researcher concludes there is a statistically significant association between gender and candidate preference in this sample.<\/p>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_3_Comparing_Anxiety_Across_Three_Therapy_Modalities\"><\/span><strong>Example 3: Comparing Anxiety Across Three Therapy Modalities<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n\r\n\r\n\r\n<p>A clinical psychologist recruits 90 patients with generalized anxiety disorder and randomly assigns them to cognitive-behavioural therapy (CBT), mindfulness-based therapy (MBT), or a waitlist control. Post-treatment anxiety scores are compared.<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>H\u2080:<\/strong> Mean anxiety scores are equal across all three groups (\u03bc\u2081 = \u03bc\u2082 = \u03bc\u2083)<\/li>\r\n<li><strong>H\u2081:<\/strong> At least one group mean differs<\/li>\r\n<li><strong>Test:<\/strong> One-way ANOVA (three groups, continuous outcome)<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<p>If F(2, 87) = 8.4, p = 0.0004 &lt; 0.05, H\u2080 is rejected. Post-hoc tests (e.g., Tukey&#8217;s HSD) then identify <em>which<\/em> pairs of groups differ significantly.<\/p>\r\n\r\n\r\n\r\n<h4><span class=\"ez-toc-section\" id=\"Example_4_Exact_Binomial_Test_for_Treatment_Efficacy\"><\/span><strong>Example 4: Exact Binomial Test for Treatment Efficacy<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h4>\r\n\r\n\r\n\r\n<p>Suppose a treatment has an expected success rate of 0.25. A researcher claims she has a new treatment with improved efficacy and tests it in 3 patients. If all 3 patients respond, P = Pr(X = 3) = 0.0156. This would be rare if the true success rate were only 25%, so the evidence against H\u2080 is deemed significant. If only 2 of 3 respond, P = Pr(X \u2265 2) = 0.1406 + 0.0156 = 0.1562. This observation is not unusual under H\u2080, so the evidence is deemed non-significant. <a href=\"https:\/\/www.sjsu.edu\/faculty\/gerstman\/StatPrimer\/hyp-test.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"What_are_Type_I_and_Type_II_Errors\"><\/span><strong>What are Type I and Type II Errors?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>Every NHST decision carries two possible error types. Understanding them is essential for designing studies and interpreting results responsibly.<\/p>\r\n\r\n\r\n\r\n<p>A <a href=\"https:\/\/www.editage.com\/blog\/type-i-type-ii-errors-hypothesis-testing\/\">Type I error<\/a> occurs when we reject the null hypothesis although that hypothesis was true. A <a href=\"https:\/\/www.editage.com\/blog\/type-i-type-ii-errors-hypothesis-testing\/\">Type II error<\/a> occurs when we fail to reject the null hypothesis even though it is false. <a href=\"https:\/\/www.geeksforgeeks.org\/data-science\/understanding-hypothesis-testing\/\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Decision<\/strong><\/td>\r\n<td><strong>H\u2080 is Actually True<\/strong><\/td>\r\n<td><strong>H\u2080 is Actually False<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td><strong>Reject H\u2080<\/strong><\/td>\r\n<td>\u274c Type I Error (False Positive): rate = \u03b1<\/td>\r\n<td>\u2705 Correct (True Positive)<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Fail to Reject H\u2080<\/strong><\/td>\r\n<td>\u2705 Correct (True Negative)<\/td>\r\n<td>\u274c Type II Error (False Negative): rate = \u03b2<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>Type I error (\u03b1):<\/strong> Concluding a new antidepressant works when it actually doesn&#8217;t: leading to unnecessary prescription and costs.<\/li>\r\n<li><strong>Type II error (\u03b2):<\/strong> Concluding a drug doesn&#8217;t work when it actually does: a missed therapeutic opportunity.<\/li>\r\n<li><strong>Statistical Power (1 \u2212 \u03b2):<\/strong> The probability of correctly detecting a real effect. Power is typically set at 0.80 in study planning, meaning researchers accept a 20% chance of missing a real effect.<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<p><strong>The trade-off:<\/strong> Lowering \u03b1 to reduce false positives increases \u03b2 (more false negatives), and vice versa. The only way to reduce both simultaneously is to increase sample size.<\/p>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Common_Misconceptions_About_P_Values\"><\/span><strong>Common Misconceptions About P Values<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>The interpretation of P values is a minefield. The man who introduced it as a formal research tool, the statistician and geneticist R.A. Fisher, could not explain exactly its inferential meaning. He proposed a rather informal system that could be used, but he never could describe straightforwardly what it meant from an inferential standpoint. <a href=\"https:\/\/www.sjsu.edu\/faculty\/gerstman\/StatPrimer\/hyp-test.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">\u00a0<\/a><\/p>\r\n\r\n\r\n\r\n<p>Here are the most dangerous misconceptions, with corrections:<\/p>\r\n\r\n\r\n\r\n<figure class=\"wp-block-table\">\r\n<table>\r\n<thead>\r\n<tr>\r\n<td><strong>Misconception<\/strong><\/td>\r\n<td><strong>What People Think<\/strong><\/td>\r\n<td><strong>What Is Actually True<\/strong><\/td>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td>&#8220;p &lt; 0.05 means the result is important&#8221;<\/td>\r\n<td>A small p means a big, important effect<\/td>\r\n<td>P values say nothing about <a href=\"https:\/\/www.editage.com\/blog\/effect-size\/\">effect size<\/a> or practical importance<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>&#8220;p = 0.04 proves the alternative hypothesis&#8221;<\/td>\r\n<td>H\u2080 is false; H\u2081 is true<\/td>\r\n<td>We only conclude the data are unlikely under H\u2080; we don&#8217;t confirm H\u2081<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>&#8220;p = 0.06 means no effect exists&#8221;<\/td>\r\n<td>Failing to reject H\u2080 proves it is true<\/td>\r\n<td>Absence of evidence \u2260 evidence of absence<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>&#8220;p is the probability H\u2080 is true&#8221;<\/td>\r\n<td>p = P(H\u2080 is true | data)<\/td>\r\n<td>p = P(data this extreme | H\u2080 is true): very different<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>&#8220;p &lt; 0.05 is always the right threshold&#8221;<\/td>\r\n<td>0.05 is a universal law of nature<\/td>\r\n<td>\u03b1 is a convention; the right threshold depends on the stakes<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>&#8220;Replication is guaranteed by a small p&#8221;<\/td>\r\n<td>The finding will reappear in future studies<\/td>\r\n<td>A single p value makes no guarantee about reproducibility<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Limitations_of_Hypothesis_Testing\"><\/span><strong>Limitations of Hypothesis Testing<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n\r\n\r\n\r\n<p>NHST has attracted intense criticism over the past few decades, especially in light of the replication crisis in psychology and biomedicine. Key limitations include:<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li><strong>Binary thinking:<\/strong> Forcing a rich continuum of evidence into &#8220;significant&#8221; or &#8220;not significant&#8221; loses information and encourages all-or-nothing interpretation.<\/li>\r\n<li><strong><a href=\"https:\/\/www.editage.com\/insights\/have-you-fallen-prey-to-data-dredging\">P-hacking<\/a> and researcher degrees of freedom:<\/strong> Flexible data collection, analysis choices, and selective reporting inflate the false-positive rate far above the nominal \u03b1.<\/li>\r\n<li><strong>The file-drawer problem and <a href=\"https:\/\/www.editage.com\/insights\/publication-and-reporting-biases-and-how-they-impact-publication-of-research\">publication bias<\/a>:<\/strong> Studies that fail to reject H\u2080 are less likely to be published, biasing the published literature toward positive findings.<\/li>\r\n<li><strong>Conflation of statistical and practical significance:<\/strong> A study of 100,000 patients may find that a drug lowers blood pressure by 0.5 mmHg with p &lt; 0.0001: statistically overwhelming, clinically irrelevant.<\/li>\r\n<li><strong>Data quality dependence:<\/strong> The accuracy of the results depends on the quality of the data. Poor-quality or inaccurate data can lead to incorrect conclusions.<\/li>\r\n<li><strong>Context limitations:<\/strong> Hypothesis testing doesn&#8217;t always consider the bigger picture, which can oversimplify results and lead to incomplete insights.<\/li>\r\n<li><strong>Assumption violations:<\/strong> Most standard tests assume <a href=\"https:\/\/www.editage.com\/blog\/normality-test-methods-of-assessing-normality\/\">normally distributed data<\/a>, independent observations, and equal variances. Violations can distort p values.<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Best_practices_to_mitigate_these_limitations\"><\/span><strong>Best practices to mitigate these limitations:<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<ul>\r\n<li>Pre-register hypotheses and analysis plans (e.g., on OSF or ClinicalTrials.gov)<\/li>\r\n<li>Report effect sizes and <a href=\"https:\/\/www.editage.com\/blog\/what-is-confidence-intervals-and-why-is-it-important\/\">confidence intervals<\/a> alongside p values<\/li>\r\n<li>Use <a href=\"https:\/\/www.editage.com\/blog\/sample-size-and-statistical-power-definition-formulas-calculations-worked-examples\/\">sufficiently powered studies<\/a> (plan for \u226580% power)<\/li>\r\n<li>Replicate findings before drawing firm conclusions<\/li>\r\n<li>Consider Bayesian approaches or equivalence testing where appropriate<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<h2><span class=\"ez-toc-section\" id=\"Frequently_Asked_Questions_FAQs\"><\/span><strong>Frequently Asked Questions (FAQs)<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\r\n<h3 class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\"><span class=\"ez-toc-section\" id=\"What_is_a_null_hypothesis\"><\/span><strong>What is a null hypothesis?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n<p class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">The null hypothesis (written H\u2080) is the default assumption in a statistical test: that there&#8217;s no effect, no difference, or no relationship in the population you&#8217;re studying. It&#8217;s the claim you try to find evidence <em>against<\/em>.<\/p>\r\n<p class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">For example, if you&#8217;re testing whether a new drug lowers blood pressure, the null hypothesis says it doesn&#8217;t; any difference you observe between the treatment and control groups is just random variation. The alternative hypothesis (H\u2081) says the drug does have an effect.<\/p>\r\n<p class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">You collect data and calculate how likely your results would be if H\u2080 were true. That probability is the p-value. If it&#8217;s small enough (commonly below 0.05), you reject the null hypothesis.<\/p>\r\n<p class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">Two things worth remembering:<\/p>\r\n<ol>\r\n<li class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">You never &#8220;prove&#8221; the null true; you only fail to reject it.<\/li>\r\n<li class=\"font-claude-response-body break-words whitespace-normal\" dir=\"ltr\">Rejecting H\u2080 tells you an effect likely exists, not that it&#8217;s large or important.<\/li>\r\n<\/ol>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_is_statistical_power_and_why_does_it_matter\"><\/span><strong>What is statistical power, and why does it matter?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>Statistical power (1 \u2212 \u03b2) is the probability that a test will correctly detect a true effect when one exists. A study with 50% power has only a coin-flip chance of finding a real effect. Low power wastes resources and produces unreliable findings. Power depends on <a href=\"https:\/\/www.editage.com\/blog\/sample-size-and-statistical-power-definition-formulas-calculations-worked-examples\/\">sample size<\/a>, effect size, and \u03b1. Most disciplines target at least 80% power during <a href=\"https:\/\/www.editage.com\/blog\/types-of-study-designs-in-biomedical-research\/\">study design<\/a>, requiring formal power calculations before data collection.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"How_is_a_confidence_interval_related_to_a_hypothesis_test\"><\/span><strong>How is a confidence interval related to a hypothesis test?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>A 95% <a href=\"https:\/\/www.editage.com\/blog\/what-is-confidence-intervals-and-why-is-it-important\/\">confidence interval<\/a> (CI) and a two-tailed test at \u03b1 = 0.05 convey equivalent information: if the CI excludes the null value (e.g., zero for a mean difference), the corresponding p value will be below 0.05. CIs are often preferred because they communicate both the direction and the magnitude of the effect, not just whether it passed a threshold. Reporting both the p value and the CI is considered best practice.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_does_it_mean_to_%E2%80%9Cpre-register%E2%80%9D_a_study\"><\/span><strong>What does it mean to &#8220;pre-register&#8221; a study?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>Pre-registration means publicly documenting your hypotheses, data collection plan, and analysis strategy before collecting data, typically through platforms like ClinicalTrials.gov (biomedical) or the Open Science Framework (social sciences). This prevents researchers from unconsciously adjusting their hypotheses or analysis methods after seeing results (HARKing: Hypothesising After Results are Known), which inflates the false-positive rate and undermines reproducibility.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"When_should_I_use_a_non-parametric_test_instead_of_a_t-test_or_ANOVA\"><\/span><strong>When should I use a non-parametric test instead of a t-test or ANOVA?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>Parametric tests like <a href=\"https:\/\/www.editage.com\/insights\/what-biomedical-researchers-need-to-know-about-t-tests\">t-tests<\/a> and <a href=\"https:\/\/www.editage.com\/blog\/anova-types-uses-assumptions-a-quick-guide-for-biomedical-researchers\/\">ANOVA<\/a> assume the data are approximately normally distributed. When sample sizes are small and data are strongly skewed, heavily bounded (e.g., Likert scales with small n), or contain extreme outliers, non-parametric alternatives are more appropriate. Common examples include the Mann-Whitney U test (instead of independent t-test), Wilcoxon signed-rank test (instead of paired t-test), and Kruskal-Wallis test (instead of one-way ANOVA). <a href=\"https:\/\/www.editage.com\/insights\/an-introduction-to-non-parametric-tests-for-biomedical-researchers\">Non-parametric tests<\/a> sacrifice some statistical power in exchange for fewer distributional assumptions.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_is_effect_size_and_which_measures_are_commonly_reported\"><\/span><strong>What is effect size, and which measures are commonly reported?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>Effect size quantifies the <em>magnitude<\/em> of a difference or association, independent of sample size. Common measures include Cohen&#8217;s d (standardised mean difference; d = 0.2 small, 0.5 medium, 0.8 large), Pearson&#8217;s r (correlation), \u03b7\u00b2 (eta-squared, for ANOVA), and odds ratios or relative risks (for categorical outcomes in clinical research). Reporting effect sizes alongside p values allows readers to judge whether a statistically significant finding is also practically or clinically meaningful.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"What_is_the_difference_between_a_one-sample_and_two-sample_test\"><\/span><strong>What is the difference between a one-sample and two-sample test?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>A one-sample test compares a single group&#8217;s mean (or proportion) to a known or hypothesised population value. Example: testing whether the mean birth weight in a hospital differs from the national standard of 3.2 kg. A two-sample test compares the means (or proportions) of two independent groups.<\/p>\r\n\r\n\r\n\r\n<p>Example: testing whether mean depression scores differ between patients receiving CBT and those receiving pharmacotherapy. When the two sets of measurements come from the same individuals at different times (e.g., pre- and post-intervention), a paired test is used instead.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Why_has_NHST_been_criticised_and_what_are_the_proposed_alternatives\"><\/span><strong>Why has NHST been criticised, and what are the proposed alternatives?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>Critics argue that the rigid p &lt; 0.05 threshold encourages dichotomous thinking, incentivises p-hacking, and obscures effect sizes. The American Statistical Association issued statements in 2016 and 2019 urging researchers to move beyond &#8220;statistically significant.&#8221; Proposed alternatives and complements include:<\/p>\r\n\r\n\r\n\r\n<ul>\r\n<li>Bayesian inference (expressing results as updated probability distributions),<\/li>\r\n<li>estimation-based approaches (reporting effect sizes and CIs without binary cutoffs),<\/li>\r\n<li>equivalence testing (demonstrating that an effect is small enough to be unimportant),<\/li>\r\n<li>and false discovery rate (FDR) control in large-scale genomic or neuroimaging studies.<\/li>\r\n<\/ul>\r\n\r\n\r\n\r\n<p>Many journals now require effect sizes and CIs in addition to p values.<\/p>\r\n\r\n\r\n\r\n<h3><span class=\"ez-toc-section\" id=\"Can_hypothesis_testing_be_used_with_observational_data_or_only_with_experiments\"><\/span><strong>Can hypothesis testing be used with observational data, or only with experiments?<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\r\n\r\n\r\n\r\n<p>NHST applies to both <a href=\"https:\/\/www.editage.com\/blog\/types-of-experimental-research-designs\/\">experimental<\/a> and <a href=\"https:\/\/www.editage.com\/blog\/observational-study\/\">observational<\/a> data, but the conclusions that can be drawn differ. Randomised controlled trials (RCTs) allow causal inference: if the test rejects H\u2080, the intervention is likely the cause. In observational studies (e.g., <a href=\"https:\/\/www.editage.com\/blog\/questionnaire-survey-research\/\">survey data<\/a>, <a href=\"https:\/\/www.editage.com\/blog\/cohort-study\/\">cohort studies<\/a>), NHST can detect associations but cannot establish causation because of potential <a href=\"https:\/\/www.editage.com\/blog\/confounding-variables-identification-definition-types-examples\/\">confounding<\/a>. A statistically significant association between coffee consumption and reduced Parkinson&#8217;s disease risk, for instance, does not by itself prove that coffee is protective because unmeasured lifestyle confounders may explain the association.<\/p>\r\n","protected":false},"excerpt":{"rendered":"So, what is Hypothesis testing? In plain English, Hypothesis Testing is a way to figure out if an idea or assumption about a group of people is actually true, based on the data that's available. Hypothesis testing is widely used in biomedical research to confirm whether there's a connection between different variables, like if having a certain disease affects levels of a specific biomarker in the body. ","protected":false},"author":2,"featured_media":438,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_ayudawp_aiss_exclude":false,"_ayudawp_aiss_summary":"While \"hypothesis testing\" is a broad term covering many philosophies, NHST refers to one particular procedure: you begin by assuming the null hypothesis is true, collect data, compute a test statistic, and then ask how probable your observed result would be under that assumption. The null hypothesis (H\u2080) is a statement of \"no difference,\" \"no association,\" or \"no treatment effect.\" The alternative hypothesis (H\u2090) is a statement of \"difference,\" \"association,\" or \"treatment effect.\" H\u2080 is assumed to be true until proven otherwise. The P value answers the question: \"If the null hypothesis were true, what is the probability of observing the current data or data that is more extreme?\" Note that the P value is NOT the probability that the hypothesis (or any other hypothesis) is right or wrong.","_ayudawp_aiss_summary_provider":"manual","_ayudawp_aiss_summary_hash":"5a4ee3264f22a571129157ccc4e5b07eae76ade1"},"categories":[14],"tags":[23,24],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v20.6 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Hypothesis Testing &amp; NHST: Definition, Steps, Tips, Examples<\/title>\n<meta name=\"description\" content=\"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Hypothesis Testing &amp; NHST: Definition, Steps, Tips, Examples\" \/>\n<meta property=\"og:description\" content=\"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\" \/>\n<meta property=\"og:site_name\" content=\"Educational Articles For Researchers, Students And Authors - Editage Blog\" \/>\n<meta property=\"article:published_time\" content=\"2023-03-21T10:53:32+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-31T16:33:40+00:00\" \/>\n<meta property=\"og:image\" content=\"http:\/\/www.editage.com\/blog\/wp-content\/uploads\/2023\/03\/Importance-Of-Binomial-Nomenclature-1.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1920\" \/>\n\t<meta property=\"og:image:height\" content=\"1080\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Editor Editor\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Editor Editor\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"18 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\"},\"author\":{\"name\":\"Editor Editor\",\"@id\":\"https:\/\/www.editage.com\/blog\/#\/schema\/person\/194519c669bbbc38e9ed47cc02c5a44f\"},\"headline\":\"Hypothesis Testing &#038; NHST: Definition, Steps, Tips, Examples\",\"datePublished\":\"2023-03-21T10:53:32+00:00\",\"dateModified\":\"2026-07-31T16:33:40+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\"},\"wordCount\":3859,\"publisher\":{\"@id\":\"https:\/\/www.editage.com\/blog\/#organization\"},\"keywords\":[\"Statistical Analysis Services\",\"Statistical Review Services\"],\"articleSection\":[\"Research Tips\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\",\"url\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\",\"name\":\"Hypothesis Testing & NHST: Definition, Steps, Tips, Examples\",\"isPartOf\":{\"@id\":\"https:\/\/www.editage.com\/blog\/#website\"},\"datePublished\":\"2023-03-21T10:53:32+00:00\",\"dateModified\":\"2026-07-31T16:33:40+00:00\",\"description\":\"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.editage.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Hypothesis Testing &#038; NHST: Definition, Steps, Tips, Examples\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.editage.com\/blog\/#website\",\"url\":\"https:\/\/www.editage.com\/blog\/\",\"name\":\"Educational Articles For Researchers, Students And Authors - Editage Blog\",\"description\":\"Get insightful educational articles from the world of academia for researchers, students and authors. Visit Editage Blog for helpful content and tips on getting published and writing articles that are up to international journal publication standards. Click here to find out more!\",\"publisher\":{\"@id\":\"https:\/\/www.editage.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.editage.com\/blog\/?s={search_term_string}\"},\"query-input\":\"required name=search_term_string\"}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/www.editage.com\/blog\/#organization\",\"name\":\"Educational Articles For Researchers, Students And Authors - Editage Blog\",\"url\":\"https:\/\/www.editage.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.editage.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/www.editage.com\/blog\/wp-content\/uploads\/2022\/08\/editage-logo.png\",\"contentUrl\":\"https:\/\/www.editage.com\/blog\/wp-content\/uploads\/2022\/08\/editage-logo.png\",\"width\":394,\"height\":82,\"caption\":\"Educational Articles For Researchers, Students And Authors - Editage Blog\"},\"image\":{\"@id\":\"https:\/\/www.editage.com\/blog\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.editage.com\/blog\/#\/schema\/person\/194519c669bbbc38e9ed47cc02c5a44f\",\"name\":\"Editor Editor\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.editage.com\/blog\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/33094b932a69316d705f8302c2f84d82?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/33094b932a69316d705f8302c2f84d82?s=96&d=mm&r=g\",\"caption\":\"Editor Editor\"},\"url\":\"https:\/\/www.editage.com\/blog\/author\/admin-2\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Hypothesis Testing & NHST: Definition, Steps, Tips, Examples","description":"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/","og_locale":"en_US","og_type":"article","og_title":"Hypothesis Testing & NHST: Definition, Steps, Tips, Examples","og_description":"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.","og_url":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/","og_site_name":"Educational Articles For Researchers, Students And Authors - Editage Blog","article_published_time":"2023-03-21T10:53:32+00:00","article_modified_time":"2026-07-31T16:33:40+00:00","og_image":[{"width":1920,"height":1080,"url":"http:\/\/www.editage.com\/blog\/wp-content\/uploads\/2023\/03\/Importance-Of-Binomial-Nomenclature-1.jpg","type":"image\/jpeg"}],"author":"Editor Editor","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Editor Editor","Est. reading time":"18 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#article","isPartOf":{"@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/"},"author":{"name":"Editor Editor","@id":"https:\/\/www.editage.com\/blog\/#\/schema\/person\/194519c669bbbc38e9ed47cc02c5a44f"},"headline":"Hypothesis Testing &#038; NHST: Definition, Steps, Tips, Examples","datePublished":"2023-03-21T10:53:32+00:00","dateModified":"2026-07-31T16:33:40+00:00","mainEntityOfPage":{"@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/"},"wordCount":3859,"publisher":{"@id":"https:\/\/www.editage.com\/blog\/#organization"},"keywords":["Statistical Analysis Services","Statistical Review Services"],"articleSection":["Research Tips"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/","url":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/","name":"Hypothesis Testing & NHST: Definition, Steps, Tips, Examples","isPartOf":{"@id":"https:\/\/www.editage.com\/blog\/#website"},"datePublished":"2023-03-21T10:53:32+00:00","dateModified":"2026-07-31T16:33:40+00:00","description":"Learn the basics of hypothesis testing, what is the null hypothesis, what is the alternative hypothesis, p value, confidence intervals, effect size.","breadcrumb":{"@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/www.editage.com\/blog\/hypothesis-testing-different-types-for-biomedical-researchers\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.editage.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Hypothesis Testing &#038; NHST: Definition, Steps, Tips, Examples"}]},{"@type":"WebSite","@id":"https:\/\/www.editage.com\/blog\/#website","url":"https:\/\/www.editage.com\/blog\/","name":"Educational Articles For Researchers, Students And Authors - Editage Blog","description":"Get insightful educational articles from the world of academia for researchers, students and authors. Visit Editage Blog for helpful content and tips on getting published and writing articles that are up to international journal publication standards. Click here to find out more!","publisher":{"@id":"https:\/\/www.editage.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.editage.com\/blog\/?s={search_term_string}"},"query-input":"required name=search_term_string"}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.editage.com\/blog\/#organization","name":"Educational Articles For Researchers, Students And Authors - Editage Blog","url":"https:\/\/www.editage.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.editage.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/www.editage.com\/blog\/wp-content\/uploads\/2022\/08\/editage-logo.png","contentUrl":"https:\/\/www.editage.com\/blog\/wp-content\/uploads\/2022\/08\/editage-logo.png","width":394,"height":82,"caption":"Educational Articles For Researchers, Students And Authors - Editage Blog"},"image":{"@id":"https:\/\/www.editage.com\/blog\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/www.editage.com\/blog\/#\/schema\/person\/194519c669bbbc38e9ed47cc02c5a44f","name":"Editor Editor","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.editage.com\/blog\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/33094b932a69316d705f8302c2f84d82?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/33094b932a69316d705f8302c2f84d82?s=96&d=mm&r=g","caption":"Editor Editor"},"url":"https:\/\/www.editage.com\/blog\/author\/admin-2\/"}]}},"jetpack_featured_media_url":"https:\/\/www.editage.com\/blog\/wp-content\/uploads\/2023\/03\/Importance-Of-Binomial-Nomenclature-1.jpg","_links":{"self":[{"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/posts\/436"}],"collection":[{"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/comments?post=436"}],"version-history":[{"count":11,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/posts\/436\/revisions"}],"predecessor-version":[{"id":1810,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/posts\/436\/revisions\/1810"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/media\/438"}],"wp:attachment":[{"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/media?parent=436"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/categories?post=436"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.editage.com\/blog\/wp-json\/wp\/v2\/tags?post=436"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}