Comprehensive Study Notes on Statistical Enquiry and Data Collection Techniques

Statistical Enquiry and Classification of Data

A statistical enquiry is formally defined as the process of investigating a specific topic where the data is collected exclusively in quantitative form. This investigative process allows for the objective analysis of information based on numerical evidence rather than qualitative observations. Depending on the objective of the study, statistical enquiries are categorized into two primary types: general purpose and specific purpose.

General purpose enquiries involve data that is collected to serve multiple functions or to be used across various fields of study. Examples of this include the national population census, national income data, and school records. These datasets provide a broad foundation of information that can be accessed by various researchers for different analytical needs. Conversely, specific purpose enquiries are characterized by data collected with a very narrow and defined objective in mind. Examples include conducting a survey to determine a student's favorite subject or performing research into consumer preferences for a particular product. In these cases, the data is tailored specifically to answer a singular research question.

Primary and Secondary Data Sources

Data sources are fundamentally divided into primary and secondary categories based on the origin of the information. Primary data refers to data that is collected first-hand by the investigator for the specific purpose of the current study. It is often referred to as raw data because it has not undergone any previous statistical treatment. The merits of primary data include its high level of accuracy and its originality; however, the demerits are significant as the collection process is typically very expensive and highly time-consuming.

Secondary data consists of information that has already been collected by other individuals or organizations for a purpose other than the current investigation. While primary data is "first-hand," secondary data is "second-hand" because it is already available in existing records. The primary advantages of secondary data are that it is easy to collect, less costly, and saves a significant amount of time. Nevertheless, there are notable demerits, such as the possibility that the data may not be accurate or is less reliable for the current researcher's specific needs compared to data they would have collected themselves.

A comparison between primary and secondary data highlights these distinctions: primary data is first-hand and more reliable but expensive and time-consuming, requiring personal investigation. Secondary data is already collected and documented in books or reports, making it cheap and quick to access, though it is generally considered less reliable than primary sources.

Methods of Primary Data Collection

There are several distinct methods for collecting primary data, each with its own methodology and set of trade-offs. The first is Direct Personal Investigation, a method where the investigator personally contacts the respondents and gathers information directly from them. The merits of this approach include the high degree of accuracy and reliability of the data obtained, as well as the flexibility to adjust questions based on the respondent's reactions. It allows for personal questioning which can uncover deeper insights. However, the demerits are that it is exceptionally time-consuming and expensive compared to other methods.

Telephone Interviews represent a second method of data collection. In this scenario, the investigator contacts respondents over the telephone to ask questions and obtain information. This method is praised for being quick and less expensive than personal visits, but it suffers from a lower response rate, as respondents may be less willing to engage in a conversation over the phone than in person.

Mailing Questionnaires involve sending a predetermined list of questions to respondents via post or email, which the respondents fill out and return. This method is characterized by low costs and the ability to achieve wide geographical coverage. The effectiveness of this method depends heavily on the quality of the questionnaire. However, it frequently suffers from a low response rate and the possibility of misinterpretation of the questions by the respondents, as there is no investigator present to clarify the meaning of the queries.

Design of Questionnaires and Pilot Surveys

The creation of an effective questionnaire is a critical component of data collection. A good questionnaire should feature short and simple questions that are clear to the respondent. The flow of the document should move from general questions to more specific ones, and the designer must be careful to avoid asking personal or sensitive questions that might alienate the respondent. These features ensure that the data collected is both relevant and high-quality.

Before launching a full-scale study, researchers often conduct a Pilot Survey. A Pilot Survey is a preliminary, small-scale trial survey conducted before the main survey to test the effectiveness of the data collection process. Its primary purpose is to test the questionnaire for any flaws or ambiguities. By identifying and correcting these issues early, a pilot survey improves the overall accuracy of the final study and helps eliminate wasted time and costs that would have occurred if errors were discovered during the main survey.

Sources of Secondary Data and Precautions

Secondary data can be obtained from two main types of sources: published and unpublished. Published sources include data that has already been collected and officially released by organizations, governments, or institutions for public use. Typical examples include government reports, magazines, and newspapers. Unpublished sources refer to data that has been collected but has not been officially published for the public. This information is kept in internal records and can sometimes be obtained from offices when requested. Examples of unpublished sources include government office records, school records, and internal company records.

When utilizing secondary data, researchers must exercise specific precautions to ensure the integrity of their study. They must check the reliability of the source of the data and verify the original method of collection used. Furthermore, they must check the data for accuracy and evaluate its overall suitability for the current purpose. Failing to take these precautions can lead to flawed conclusions based on erroneous or irrelevant information.

Techniques of Data Collection: Census vs. Sampling

There are two main techniques for collecting data regarding a population: the Census Method and the Sample Method. In the Census Method, information is collected from every single unit or individual within the entire population under study. The primary advantage of this method is that it provides highly detailed and comprehensive information. However, its disadvantages are that it is extremely time-consuming and requires a massive amount of labor and resources.

The Sample Method is a technique where only a part of the population, known as a sample, is selected and studied. The results and observations derived from this sample are then used to represent the characteristics of the whole population. The major advantages of the sample method are that it saves both time and cost. The disadvantages include the fact that the results may be less accurate than a full census and the possibility of bias, leading to wrong conclusions if the sample does not correctly represent the population.

Random and Non-Random Sampling Methods

Sampling methods are broadly categorized into random and non-random techniques. Random sampling is a method where every unit of the population has an equal chance of being selected for the sample. Its main advantage is the lack of bias, though it may not be suitable for extremely large populations. Within random sampling, several specific variations exist:

Stratified Sampling involves dividing the population into different groups, called strata, based on a common characteristic. Samples are then selected from each group. An example of this is dividing a student body based on subjects, such as Science, Commerce, and Humanities, and then sampling from each.

Cluster Sampling is a method where the population is divided into groups called clusters. Instead of selecting individuals, a few clusters are randomly selected, and every individual within those selected clusters is studied. For example, if one wants to study students in a city, individual schools could be treated as clusters.

Systematic Sampling involves arranging the population in a specific order and selecting units at a regular interval. If there are N=100N = 100 students and a sample of n=10n = 10 is needed, the sampling interval, denoted as kk, is calculated as: k=10010=10k = \frac{100}{10} = 10 A starting student is selected randomly between 11 and 1010. For instance, if the 3rd3^{rd} student is picked, then every 10th10^{th} student thereafter is selected for the study: 3rd3^{rd}, 13th13^{th}, 23rd23^{rd}, up to the 93rd93^{rd} student.

Non-Random Sampling occurs when units do not have an equal chance of selection. This category includes Purposive Sampling (also known as Judgmental Sampling), where the researcher specifically selects samples that they believe suit the purpose of the study. This is generally only useful when small units are involved. Convenience Sampling involves selecting samples because they are easily available or convenient to the researcher. Quota Sampling involves dividing the population into groups and selecting a fixed number, or quota, from each group. While this ensures representation of different groups, the selection within those groups is non-random. An example would be specifically selecting 4040 males and 6060 females for a survey. Non-random methods generally carry a high risk of bias and a lack of representativeness.