Biostatistics: Basic Statistical Concepts, Sampling Methods, and Variables

Basic Concepts of Statistics

  • Definition of Statistics (الإحصاء):

    • Statistics is the scientific method that specializes in collecting data and facts about a specific phenomenon or a specific group of phenomena.

    • It encompasses organizing and classifying this data and facts in a manner that facilitates their analysis and interpretation.

    • It provides the framework for drawing conclusions and making sound decisions in light of the analyzed results.

  • Main Branches of Statistics:

    • Descriptive Statistics (الإحصاء الوصفي):

    • Concerned with summarizing and describing data using a set of statistical tools.

    • Goal: Organize and summarize data in a way that helps to understand it quickly and effectively.

    • Inferential Statistics (الإحصاء الاستدلالي):

    • Used to interpret data and draw general conclusions from a sample of data.

    • Used specifically for hypothesis testing and predicting future trends.

    • Key elements include Hypothesis testing and Data analysis.

  • Scope of Application of Statistics (مجال تطبيق علم الإحصاء):

    • Used in every field that involves scientific research.

    • Applications encompass:

    • Agricultural

    • Industrial

    • Psychological

    • Social

    • Sports and Youth

    • Medical

    • Economic

    • Administrative

    • Engineering

Stages of the Statistical Method

  • Stages of the Statistical Method (مراحل الطريقة الإحصائية):

    • Stage 1: Defining a problem or research hypothesis.

    • Stage 2: Collecting data and information about the phenomenon or phenomena related to the research.

    • Stage 3: Classifying, tabulating, and displaying data (جمع وتصنيف وتبويب البيانات).

    • Stage 4: Calculating statistical indicators, such as estimating the parameters of the research community (population).

    • Stage 5: Analyzing the study data and reaching the results.

    • Stage 6: Interpreting the results and making a decision regarding the research hypotheses.

Statistical Population and Statistical Term

  • Statistical Population / Community (المجتمع الإحصائي):

    • Comprises all components of the phenomenon that is the subject of research or study.

    • These components are supposed to share a certain characteristic or set of characteristics.

    • Components may be a living organism or anything else.

  • Classification of Populations:

    • Finite Population (المجتمع المحدد):

    • A population that consists of a limited and known number of individuals or elements.

    • All of its elements can be counted accurately.

    • Infinite Population (المجتمع غير المحدد):

    • A society/population that consists of an unlimited or unknown number of individuals or elements.

    • Completely counting all elements is impossible due to the sheer size or unlimited nature of the population.

  • Statistical Term / Unit / Item (المفردة الإحصائية):

    • The smallest unit in the statistical community.

Data Collection Sources and Methods

  • Sources of Research Data (مصادر البيانات):

    • Historical Sources (المصادر التاريخية):

    • Data and information stored and collected by various state agencies and institutions as a result of surveys conducted by these agencies or bodies for their own purposes.

    • Examples include:

      • Population census data

      • Agricultural and industrial production statistics

      • Foreign and domestic trade statistics

      • Statistics on students graduating from Iraqi universities

    • Field Sources (المصادر الميدانية):

    • Used when it is not possible to obtain data from historical sources.

    • Involves obtaining data directly from its original sources in the field.

  • Methods of Collecting Data from the Field (أساليب جمع البيانات من الميدان):

    • Comprehensive Inventory Method (أسلوب الحصر الشامل):

    • Data is collected on all elements of the statistical community.

    • The statistical community must be specific and finite (every element in it can be observed).

    • Examples:

      • General population census

      • Inventory of establishments and industrial units in Iraq

    • Sample Method (أسلوب العينات):

    • Means collecting data and information about a specific group of statistical community items.

    • This chosen group of items is called a sample.

    • Chosen in a way that ensures it accurately represents the statistical community.

Random Sampling Methods

  • Random Sample (العينة العشوائية):

    • A group of individuals selected from the statistical community in which the researcher has no interference in the selection process.

  • Types of Random Samples:

    • 1. Simple Random Sample (العينة العشوائية البسيطة):

    • Chosen randomly to ensure that every single item in the statistical community has the same probability of appearance in the sample.

    • Condition: The statistical community must be homogeneous with respect to the characteristic(s) related to the research.

    • Mathematical Procedure:

      • Assume a homogeneous, finite statistical population with NN items.

      • When choosing a simple random sample of size nn, each item has an equal probability of appearance equal to 1N\frac{1}{N}.

      • The total number of unique simple random samples that can be selected from this population is calculated using the combinations formula:         CnN=N!n!(N−n)!C_n^N = \frac{N!}{n!(N-n)!}

    • Example:

      • A homogeneous statistical population consists of N=4N = 4 items: AA, BB, CC, DD.

      • Goal: Select a simple random sample of n=3n = 3 items.

      • Number of simple random samples:         C34=4!3!(4−3)!=246=4C_3^4 = \frac{4!}{3!(4-3)!} = \frac{24}{6} = 4

      • Factorial expansion: 4!=4×3×2×1=244! = 4 \times 3 \times 2 \times 1 = 24

      • Probability of drawing any individual item: 14=1N\frac{1}{4} = \frac{1}{N}

      • The 44 possible simple random samples are:

      • A,B,CA, B, C

      • A,B,DA, B, D

      • A,C,DA, C, D

      • B,C,DB, C, D

    • 2. Stratified Random Sample (العينة الطبقية العشوائية):

    • Considered the best and most accurate sampling type for representing a heterogeneous statistical population.

    • In many cases, population items are non-homogeneous regarding the studied trait (e.g., studying family income where families are high-income, medium-income, or low-income).

    • Drawing a simple random sample from a heterogeneous population is invalid because it leads to biased estimates toward one class.

    • Procedure:

      • Partition the heterogeneous population NN into LL strata based on common characteristics among items within each stratum.

      • Stratum sizes satisfy: N1+N2+⋯+NL=NN_1 + N_2 + \dots + N_L = N.

      • Draw a simple random sample from each stratum proportional to its size in the population (Proportional Distribution Method / طريقة التوزيع المتناسب).

    • Proportional Distribution Formulas:

      • Weight of stratum hh in the population:         Wh=NhNW_h = \frac{N_h}{N}         for h=1,2,3,…,Lh = 1, 2, 3, \dots, L, where W1+W2+⋯+WL=1W_1 + W_2 + \dots + W_L = 1

      • Size of simple random sample drawn from stratum hh:         nh=n×Wh=n×NhNn_h = n \times W_h = n \times \frac{N_h}{N}

      • Total sample size nn:         n=n1+n2+⋯+nLn = n_1 + n_2 + \dots + n_L

      • Note that ratio of contribution is preserved:         nhn=NhN\frac{n_h}{n} = \frac{N_h}{N}

    • Example:

      • Population of N=2200N = 2200 families.

      • High-income families (N1N_1): 700700

      • Medium-income families (N2N_2): 900900

      • Low-income families (N3N_3): 600600

      • Required: Draw a stratified random sample of n=110n = 110 families using proportional distribution.

      • Calculation of weights:         W1=7002200=722W_1 = \frac{700}{2200} = \frac{7}{22}         W2=9002200=922W_2 = \frac{900}{2200} = \frac{9}{22}         W3=6002200=622W_3 = \frac{600}{2200} = \frac{6}{22}

      • Calculation of sample sizes:         n1=722×110=35n_1 = \frac{7}{22} \times 110 = 35         n2=922×110=45n_2 = \frac{9}{22} \times 110 = 45         n3=622×110=30n_3 = \frac{6}{22} \times 110 = 30

      • Proportionality confirmation:         W1=722=35110W_1 = \frac{7}{22} = \frac{35}{110}         W2=922=45110W_2 = \frac{9}{22} = \frac{45}{110}         W3=622=30110W_3 = \frac{6}{22} = \frac{30}{110}

    • Homework Exercise:

      • A community consists of N=1000N = 1000 people distributed by income into three classes: 200200 low-income, 550550 middle-income, and 250250 high-income people. Draw a random sample of n=100n = 100 people representing all classes of society.

    • 3. Systematic Random Sample (العينة العشوائية المنتظمة):

    • Population of NN items is divided into nn groups, each group containing kk items, where:       k=Nnk = \frac{N}{n}

    • Items must be arranged according to a specific system (e.g., ascending, descending, or address order along a street).

    • Selection Method:

      1. Select one item randomly from the first group of kk items (let its sequence number be aa).

      2. Select subsequent items by adding interval kk sequentially:

      • First item: aa

      • Second item: a+ka + k

      • Third item: a+2ka + 2k

      • Continuing up to the last item.

      1. All items after the first are separated by equal intervals kk

    • Number of Possible Samples:

      • The number of systematic random samples that can be selected from the population equals the number of items in each group, kk.

    • Example:

      • N=24N = 24 students arranged in descending order of grades.

      • Desired sample size n=6n = 6 students.

      • Step 1: Calculate interval kk:         k=Nn=246=4k = \frac{N}{n} = \frac{24}{6} = 4

      • Students divided into 66 groups of 44 students each:

      • Group 1: Items 1,2,3,41, 2, 3, 4

      • Group 2: Items 5,6,7,85, 6, 7, 8

      • Group 3: Items 9,10,11,129, 10, 11, 12

      • Group 4: Items 13,14,15,1613, 14, 15, 16

      • Group 5: Items 17,18,19,2017, 18, 19, 20

      • Group 6: Items 21,22,23,2421, 22, 23, 24

      • Step 2: Select a student randomly from Group 1 (e.g., sequence number 33).

      • Step 3: Add k=4k = 4 sequentially to obtain remaining sequences:

      • Student 1: Sequence 33

      • Student 2: 3+4=73 + 4 = 7

      • Student 3: 7+4=117 + 4 = 11

      • Student 4: 11+4=1511 + 4 = 15

      • Student 5: 15+4=1915 + 4 = 19

      • Student 6: 19+4=2319 + 4 = 23

      • Final systematic random sample sequence: (3,7,11,15,19,23)(3, 7, 11, 15, 19, 23).

    • Homework Exercises:

      • H.W. 1: How do we select a regular random sample of n=50n = 50 students from a population of N=500N = 500 students?

      • H.W. 2: How do we choose a regular random sample from a population of N=3000N = 3000 teachers, where the desired sample size is n=500n = 500 teachers?

    • 4. Multistage Sampling (المعاينة متعددة المراحل):

    • The population is divided into hierarchical sampling units across multiple stages:

      • Stage 1: Divide population into primary units; select a simple random sample of primary units.

      • Stage 2: Divide selected primary units into smaller secondary units; select a simple random sample of secondary units.

      • Stage 3: Divide selected secondary units into smaller units; select a simple random sample from them.

      • Continue subdivision and random selection until reaching the final units from which data is collected.

    • Example: Estimating average sugar consumption of an Iraqi family (statistical unit = family):

      • Stage 1: Divide Iraq into governorates (primary units); randomly select a sample of governorates.

      • Stage 2: Divide selected governorates into districts (secondary units); randomly select a sample of districts.

      • Stage 3: Divide selected districts into sub-districts (nawahis); randomly select a sample of sub-districts.

      • Stage 4: Divide selected sub-districts into residential neighborhoods (mahallat); randomly select a sample of neighborhoods.

      • Stage 5: Divide selected neighborhoods into alleys (azqah); randomly select a sample of alleys.

      • Stage 6: Divide selected alleys into residential houses; randomly select houses to reach the families for data collection.

Non-Random Sampling Methods

  • Non-Random Samples (العينات غير العشوائية):

    • A group of items selected from the statistical population in a way where the researcher interferes in the selection process based on research considerations.

  • Types of Non-Random Samples:

    • 1. Intentional / Purposive Sample (العينة العمدية):

    • Selected deliberately based on prior belief that items are best suited to represent the study population.

    • Example: Studying ways to develop football sports — deliberately selecting a sample of football specialists due to their experience in sports management.

    • 2. Quota Sample (العينة الحصصية):

    • Statistical population is divided into several strata based on criteria relevant to the study.

    • An intentional (non-random) sample is chosen from each stratum proportional to the stratum's size in the population.

    • The total sum of these intentional sample sizes forms the quota sample size.

Variables and Their Classification

  • Definition of Variable (المتغيرات):

    • Any characteristics, number, or quantity capable of being measured or counted.

    • Also referred to as a data element (e.g., age, gender, expenses, country of birth).

    • Called a variable because its value varies across units in the population or over time.

    • Represented by symbols such as XX, YY, ZZ, etc.

  • Classification of Variables:

    • 1. Quantitative Variables (المتغيرات الكمية):

    • Variables that take numerical values and can be measured.

    • Discrete Variables (المتغيرات المنفصلة):

      • Take specific numerical values, often integers (صحيحة).

      • Examples: Number of students in a classroom, number of cars in a parking lot.

    • Continuous Variables (المتغيرات المتصلة):

      • Can take any numerical value within a given range, often including decimals.

      • Examples: Weight of people, height of buildings.

    • 2. Qualitative / Descriptive Variables (المتغيرات النوعية / الوصفية):

    • Variables describing categories or attributes that cannot be measured numerically.

    • Examples: Gender (male/female), eye color (blue, brown, green), social status (rich, middle class, poor).

  • Classification Identification Exercises:

    • Intelligence score (درجة الذكاء): Descriptive / qualitative variable (متغير وصفي\text{متغير وصفي}

    • Type of college (نوع الكلية): Descriptive / qualitative variable (متغير وصفي\text{متغير وصفي}

    • Phone number (رقم الهاتف): Discrete quantitative variable (متغير كمي منفصل\text{متغير كمي منفصل}

    • Family monthly income (الدخل الشهري للأسرة): Continuous quantitative variable (متغير كمي متصل\text{متغير كمي متصل}

    • Blood type (فصيلة الدم): Descriptive / qualitative variable (وصفي\text{وصفي}

    • Number of children in the family (عدد أطفال الأسر): Discrete quantitative variable (كمي منفصل\text{كمي منفصل}

    • Skin color (لون البشرة): Descriptive / qualitative variable (وصفي\text{وصفي}

    • Gender (الجنس): Descriptive / qualitative variable (وصفي\text{وصفي}

    • Cotton yield (محصول القطن): Continuous quantitative variable (كمي متصل\text{كمي متصل}

    • Student grade (درجة الطالب): Discrete quantitative variable (كمي منفصل\text{كمي منفصل}

    • Number of cars in the parking lot (عدد السيارات في موقف السيارات): Discrete quantitative variable (كمي منفصل\text{كمي منفصل}

    • Number of students in class (عدد الطلاب في الفصل): Discrete quantitative variable (كمي منفصل\text{كمي منفصل}