CHAPTER 2 PSYCH 101
Conducting Research in Psychology
THE NATURE OF SCIENCE
Science is about testing intuitive assumptions regarding how the world works, observing the world, and being open-minded to unexpected findings. Some of science’s most important discoveries happened only because the scientists were open to surprising and unexpected results. Fundamentally, science entails collecting observations, or data, from the real world and evaluating whether the data support our ideas or not. The Stanford Prison Experiment fulfilled these criteria, and we will refer to this example several times in our discussion of research methods, measures, and ethics. Common Sense and Logic Science involves more than common sense, logic, and pure observation. Although reason and sharp powers of observation can lead to knowledge, they have limitations. Consider common sense, the intuitive ability to understand the world. Often common sense is quite useful: Don’t go too close to that cliff. Don’t rouse that sleeping bear. Don’t eat food that smells rotten. Sometimes, though, common sense leads us astray. In psychology, our intuitive ideas about people’s behavior are often contradictory or flat-out wrong. For example, most of us intuitively believe that who we are is influenced by our parents, family, friends, and society. It is equally obvious, especially to parents, that children come into the world as unique people, with their own temperaments, and people who grow up in similar environments do not have identical personalities. To what extent are we the products of our environment, and how much do we owe to heredity? Common sense cannot answer that question, but science can.
Rationalism is the view that using logic and reason is the way to understand how the world works. Logic is also a powerful tool in the scientist’s arsenal, but it can tell us only how the world should work, not how the world actually works. Sometimes the world is not logical. A classic example of the shortcoming of logic is seen in the work of the ancient Greek philosopher Aristotle. He argued that heavier objects should fall to the ground at a faster rate than lighter objects. Sounds reasonable, right? Unfortunately, it’s wrong. For 2,000 years, however, the argument was accepted simply because the great philosopher Aristotle wrote it and it made intuitive sense. It took the genius of Galileo to say, “Wait a minute. Is that really true? Let me do some tests to see whether it is true.” He did and discovered that Aristotle was wrong (Crump, 2001); the weight of an object does not affect its rate of speed when falling. Science combines logic with research and experimentation.
The Limits of Observation
Recall from The Origins of Psychology (in chapter “Introduction to Psychology”) that empiricism is the view that our observations and experience, not pure rea son and logic, are another path to knowledge. Science is empirical in that it is based on observations and experience. Science relies on observation, but even observation can lead us astray. Our knowledge of the world comes through our five senses, but they can be fairly easily fooled, as any good magician or artist can demonstrate—and as we explore in some detail in the chapter “Sensing and Perceiving Our World”. Even when we are not being intentionally fooled, the way in which our brains organize and interpret sensory experiences may vary from person to person. Another problem with observation is that people tend to generalize from their observations and assume that what they witness in one situation applies to all similar situations. Imagine you are visiting another country for the first time. Let’s say the first person you have any extended interaction with is rude, and a second, briefer interaction goes along the same lines. Granted, you have lots of language difficulties; nevertheless, you might conclude that all people from that country are rude. After all, that has been your experience. Those, however, were only two interactions, and after a couple of days you might meet other people who are quite nice. The point is that one or two cases are not a solid basis for a generalization. Scientists must collect numerous observations and conduct several studies on a topic before generalizing their conclusions.
What Is Science?
Is physics a science? Few would argue that it is not. What about biology? Psychology? Astrology? How does one decide? Now that we have looked at some of the components of science and explored their limitations, let’s consider the larger question: What is science? People often think only of the physical sciences as “science,” but science comes in at least three distinctflavors (Feist, 2006b):
• physical science,
• biological science, and
• social science.
As we mentioned in the chapter “Introduction to Psychology”, psychology is a social science.
The physical sciences study the world of things—the inanimate world of stars, light, waves, atoms, the Earth, compounds, and molecules. These sciences include physics, astronomy, chemistry, and geology. The biological sciences study plants and animals in the broadest sense. These sciences include biology, zoology, genetics, and botany. Finally, the social sciences study humans, both.
Dr. Andrew Wakefield published a scientific paper claiming that autism spectrum disorder was often caused by vaccines for measles, mumps, and rubella. There were many problems with the paper from the outset, not the least of which was its small, unrepresentative sample size (12 children). Many scientists and medical panels could not confirm the results and were highly skeptical of Dr. Wakefield’s findings. Unfortunately, the paper created quite a bit of publicity, and many parents ignored standard vaccination schedules, leading to numerous deaths from preventable diseases. After a 7-year investigation, the original paper was deemed fraudulent by the British Medical Journal and retracted (withdrawn). The investigation concluded that Dr. Wakefield had altered the results of his study to make vaccines appear to be the cause of autism spectrum disorder. as individuals and as groups. These sciences include anthropology, sociology, economics, and psychology.
Science is as much a way of thinking or a set of attitudes as it is a set of procedures. Scientific thinking involves the reasoning skills required to generate, test, and revise theories (Koslowski, 1996; Kuhn, Amsel, & O’Loughlin, 1988; Zimmerman, 2007). What we believe or theorize about the world and what the world is actually like, in the form of evidence, are two different things. Scientific thinking keeps these two things separate. In other words, scientists remember that belief is not the same as reality. There are three attitudes central to scientific thinking. The first is to question authority—including scientific authority. Be skeptical (see Figure 2). Don’t just take the word of an expert; test ideas yourself. The expert might be right, or not. That advice extends to textbooks—including this one. Wonder. Question. Ask for the evidence. Be a critical thinker. Also question your own ideas. Make your own observations—be empirical. Our natural inclination is to really like our own ideas, especially if they occur to us in a flash of insight. As one bumper sticker extols, “Don’t believe everything you think.” Believing something does not make it true. The second attitude of science is open skepticism (Sagan, 1987). Doubt and skepticism are hallmarks of critical and scientific reasoning. The French philosopher Voltaire put scientific skepticism most bluntly: “Doubt is uncomfortable, certainty is ridiculous”; however, skepticism for skepticism’s sake is also not scientific, but stubborn. Scientists are ultimately open to accepting whatever the evidence reveals, however bizarre it may be and however much they may not like it or want it to be the case. For example, could placing an electrical stimulator deep in the brain, as if it were a switch, turn off depression? That sounds like a far-fetched treatment, worthy of skepticism, but it does work for some people (Mayberg et al., 2005). Confirming Voltaire’s assertion that doubt is uncomfortable, brain imaging evidence suggests that doubt and skepticism are associated with areas of the brain involved in the sensation of taste and disgust (and belief with reward and pleasure), so doubt is a less pleasant state than belief (Harris, Sheth, & Cohen, 2008; Harris et al., 2009; Shermer, 2011). The third scientific attitude is intellectual honesty. When the central tenet of knowing is not what people think and believe, but rather how nature behaves, then we must accept the data and follow them wherever they take us. If a researcher falsifies results or interprets them in a biased way, then other scientists will not arrive at the same results if they repeat the study. Every so often we hear of a scientist who faked data in order to gain fame or funding. For the most part, however, the fact that scientists must submit their work to the scrutiny of other scientists helps ensure the honest and accurate presentation of results.
All sciences—whether physics, chemistry, biology, or psychology—share the general properties of open inquiry that we have discussed. Let’s now turn to the specific methods scientists use to acquire new and accurate knowledge of the world.
The Scientific Method
Science depends on the use of sound methods to produce trustworthy results that can be confirmed independently by other researchers. The scientific method by which scientists conduct research consists of five processes: Observe, Predict, Test, Interpret, and Communicate (O-P-T-I-C; see the Research Process for this chapter, In the observation and prediction stages of a study, researchers develop expectations about an observed phenomenon. They express their expectations as a theory, defined as a set of related assumptions from which testable predictions can be made. Theories organize and explain what we have observed and guide what we will observe (Popper, 1965). To put it simply, theories are not facts—they explain facts. Our observations of the world are always either unconsciously or consciously theory-driven, if you understand that theory in this broader sense means little more than “having an expectation.” In science, however, a theory is more than a guess. Scientific theories must be tied to real evidence, they must organize observations, and they must generate expectations that can be tested systematically. A hypothesis is a specific, informed, and testable prediction of what kind of outcome should occur under a particular condition. For example, consider the real life study that suggests that caffeine increases sex drive in female rats (Guarraci & Benson, 2005). The hypothesis may have been phrased this way: “Female rats that consume caffeine will have more couplings with male rats than female rats that do not consume caffeine.” This hypothesis predicts that a particular form of behavior (coupling with male rats) will occur in a specific group (female rats) under particular conditions (the influence of caffeine). The more specific a hypothesis is, the more easily each component can be changed to determine what effect it has on the outcome.
To test their hypotheses (the third stage of the scientific method), scientists select one of a number of established research methods, along with the appropriate measurement techniques. Selecting the methods involves choosing a design for the study, the tools that will create the conditions of the study, and the tools for measuring responses (such as how often each female rat allows a male to mounther). One basic principle of all scientific research is that measures and tools need to be both reliable and valid. Reliability means the test or measure gives us a consistent result over time or between different raters. Validity means when a scientist claims to measure a particular concept, such as sex drive for example, she really is measuring that concept and not something else. We will examine each of these elements in the next section “Research Designs in Psychology”. In the fourth step of the scientific method, scientists use mathematical techniques to interpret the results and determine whether they are significant (not just a matter of chance) and whether they closely fit the prediction. Do psychologists’ ideas of how people behave hold up, or must they be revised? Let’s say that the caffeine-consuming female rats coupled more frequently with males than did nonconsuming females. Might this enhanced sexual interest hold for all rats or just those few we studied? Statistics, a branch of mathematics we will discuss shortly, helps answer that question. The fifth stage of the scientific method is to communicate the results. Generally, scientists publish their findings in a peer-reviewed professional journal. Following a standardized format, the researchers report their hypotheses, describe their research design and the conditions of the study, summarize the results, and share their conclusions. In their reports, researchers also consider the broader implications of their results. What might the effects of caffeine on sexuality in female rats mean for our understanding of caffeine, arousal, and sex in female humans? Publication also serves an important role in making research findings part of the public domain. Such exposure not only indicates that colleagues who reviewed the study found it to be credible but also allows other researchers to repeat and/or build on the research. It is important to point out, however, that not all published scientific papers are of equal quality (that is, use equally reliable and valid measures and techniques).
Research Process\
Observe. Researchers working with rats observe their behavior.
Predict. Propose a hypothesis based in theory: Caffeine will make female rats seek more couplings with males.
Test. Collect data: How often do females on caffeine allow males to mate and how often do females not on caffeine allow males to mate?
Interpret. Analyze resulting data to confirm or disconfirm the theory-prediction.
Communicate. Publish findings of “Caffeine and mating behavior in female rats” via peer review process.
As when reading any kind of information, one should always ask “how did they come to that conclusion?” and “on what evidence did they draw that conclusion?”
Replication is the repetition of a study to confirm the results. The advancement of science hinges on replication. No matter how interesting and exciting results are, if they cannot be duplicated, the original findings may have been accidental. Whether a result holds or not, new predictions can be generated from the data, leading in turn to new studies. For example, recently, a team of more than 50 social psychologists from around the world replicated 13 classic findings in social psychology and found that although 10 of the 13 findings were replicated, the strength of the findings decreased (Klein et al., 2014). Three did not replicate and the field now knows that they were by chance and cannot be trusted. The other 10 findings are real (if a bit smaller in size) and can be built upon. This finding confirms the cumulative nature of scientific progress.
What Science Is Not: Pseudoscience
Do you believe that the planets and stars determine our destiny, that aliens have visited Earth, or that the human mind is capable of moving or altering physical objects? Astrology, unidentified flying objects (UFOs), and extrasensory perception (ESP) are certainly fascinating topics to ponder. As thinking beings, we try to understand things that science may not explain to our satisfaction. Many of us are willing to believe things that science and skeptics easily dismiss. For example, in 2015 a Chapman University poll of 1,500 representative American adults found the following (Ledbetter, 2015):
• Fifty percent endorsed at least one paranormal belief (e.g., ghosts, ESP, astrology, etc.).
• Forty-one percent believed that spirits inhabit haunted places.
• Twenty-seven percent believe the dead can communicate with the living.
• Twenty-one percent believe aliens have visited Earth.
Similarly, a 2015 survey of of more than 1,000 American adults reported that 72% believed in angels, 42% in demonic possession, and 62% believed the Earth was made 6,000 years ago (creationism), only 45% believed in the theory of evolution, and 56% believed in UFOs (What People, 2009). People often claim there is “scientific evidence” for certain unusual phenomena, but that does not mean the evidence is truly scientific. There is also false science, or pseudoscience. Pseudoscience refers to practices that appear to be and claim to be science but, in fact, do not use the scientific method to come to their conclusions. What makes something pseudoscientific comes more from the way it is studied than from the content area. According to Derry (1999) pseudoscience practitioners
1. make no real advances in knowledge,
2. disregard well-known and established facts that contradict their claims,
3. do not challenge or question their own assumptions,
4. tend to offer vague or incomplete explanations of how they came to their conclusions, and
5. tend to use unsound logic in making their arguments
THE SCIENTIFIC METHOD. The scientific method consists of an ongoing cycle of observation, prediction, testing, interpretation, and communication (OPTIC). Research begins with observation, but it doesn’t end with communication. Publishing the results of a study allows other researchers to repeat the procedure and confirm the results. Philosophy, art, music, and religion, for instance, are not pseudosciences because they do not claim to be science. Pseudoscientific claims have been made for alchemy, creation science, intelligent design, attempts to create perpetual motion machines, astrology, alien abduction, psychokinesis, and some forms of mental telepathy. Perhaps the most pervasive pseudoscience is astrology, which uses the positions of the sun, moon, and planets to explain an individual’s personality traits and to predict the future. There simply is no credible scientific evidence that the positions of the moon, planets, and stars and one’s time and place of birth have any influence on personality or life course (Hartmann, Reuter, & Nyborga, 2006; Shermer,1997; Zarka, 2011), yet about one in four American adults believe in astrology. Overall, telekinesis, astrology, alien abduction explanations of UFOs, and creation science, to name a few, meet the criteria for pseudoscience. For instance, astrology or ESP as fields of study are very much the same in their knowledge and ideas as they were 50 years ago, neither doubts and questions their own assumptions, methods, or results, and the conclusions are often vague and non-testable. In all fairness, there have been some peer-reviewed reliable observations of UFOs and some scientifically sound evidence for precognition (anticipating the future) and telepathy (Bem, 2011; Bem & Horonton, 1994; Bem, Palmer, & Broughton, 2001; Rosenthal,1986). However, other attempts to replicate these findings have not been successful (Galak, LeBeouf, Nelson, & Simmons, 2012; Rouder & Morey, 2011). Had they been replicated, we would be forced to accept them, at least tentatively. Remember, open skepticism is the hallmark of science. If there is scientifically sound evidence for something—even if it is difficult to explain—and it has been replicated, then we haveto accept it. The key is to know how to distinguish sound from unsound evidence.
RESEARCH DESIGNS IN PSYCHOLOGY
Science involves testing ideas about how the world works, but how do we design studies that test our ideas? This question confronts anyone wanting to answer a psychological question scientifically.
Principles of Research Design. Like other sciences, psychology makes use of several types of research designs—plans for how to conduct a study. The design chosen for a given study depends on the question being asked. Some questions can best be answered by randomly placing people in different groups in a laboratory to see whether a treatment causes a change in behavior. Other questions have to be studied by questionnaires or surveys. Still other questions can best be answered simply by making initial observations and seeing what people do in the real world. Sometimes researchers analyze the results of many studies on the same topic to look for trends. In this section, we examine variations in research designs, along with their advantages and disadvantages. We begin by defining a few key terms common to all research designs in psychology. A general goal of psychological research is to measure change in behavior, thought, or brain activity. A variable is anything that changes, or varies, within or between individuals. People differ from one another on age, gender, weight, intelligence, level of anxiety, and extraversion, to name a few psychological variables.
Psychologists do research by predicting how and when variables influence each other. For instance, a psychologist who is interested in whether girls develop verbal skills at a different rate than boys focuses on two variables: gender and vocabulary. All researchers must pay careful attention to how they obtain participants for a study. The first step is for the researchers to decide the makeup of the entire group, or population, in which they are interested. In psychology, populations can be composed of, for example, animals, adolescents, boys or girls of any age, college students, or students at a particular school. How many are older than 50 or younger than 20? How many are European American, African American, Asian American, Pacific Islander, or Native American? How many have high school educations, and how many have college educations? Can you think of a problem that would occur if a researcher tried to collect data directly on an entire population? Because most populations are too large to survey or interview directly, researchers draw on small subsets of each population, called samples. A sample of a population of college students, for instance, might consist of students enrolled in one or more universities in a particular geographic area. Research is almost always conducted on samples, not populations. If researchers want to draw valid conclusions or make accurate predictions about a population, it is important that their samples accurately represent the population in terms of age, gender, ethnicity, or any other variables of interest. When a poll is wrong in predicting who will win an election, it is often because the polled sample did not accurately represent the population.
Descriptive Studies
Ideas for studies often start with specific and personal experiences or events one person being painfully shy; someone rushing onto train tracks to rescue a person who had fallen in front of an ongoing train; or, having personal experience with trauma. These experiences can be and often are the driving force behind a person’s desire to study them more systematically. The point is that single events and single cases often lead to new ideas and new lines of research. When a researcher is interested in a question or topic that is relatively new to the field, often the wisest approach may be to use a descriptive design. In general, in descriptive designs the researcher makes no prediction and does not try to control any variables. She simply defines a problem of interest and describes as carefully as possible the variable of interest. The basic question in a descriptive design is, What is variable X? For example, What is love? What is genius? What is apathy? The psychologist makes careful observations, often in the real world outside the research lab. Descriptive studies usually occur during the exploratory phase of research, in which the researcher is looking for meaningful patterns that might lead to predictions later on; they generally do not involve testing hypotheses. Survey research is an exception since it can involve testing predictions. The researcher then notes possible relationships or patterns that maybe used in other designs as the basis for testable predictions (see Figure 5). Four of the most common kinds of descriptive methods in psychology are case studies, naturalistic observations, qualitative research/interviews, and surveys.
Case Study Psychotherapists have been making use of insights gained from individual cases for more than 100 years. A case study involves the observation of one person, often over a long period of time. Much wisdom and knowledge of human behavior can come from careful observation of one individual over time. Because case studies are based on one-on-one relationships, often lasting years, they offer deep insights that surveys and questionnaires often miss. Sometimes studying the lives of extraordinary individuals, such as van Gogh, Lincoln, Marie Curie, Einstein, or even Hitler, can tell us much about creativity, greatness, genius, or evil. An area of psychology called psychobiography combines psychology with history to understand human behavior through the study of individual lives in historical context (Elms, 1993; Runyan, 1982; Schultz, 2005). Like other descriptive research, case studies and psychobiographies do not test hypotheses but can be a rich source for them. One has to be careful with case studies, however, because not all cases are generalizable to other people. That is why case studies are often a starting point for the development of testable hypotheses.
Characteristics of descriptive studies
What type of questions might be researched? Single variable, such as How do people flirt?
What is the most suitable method of answering the question? Case study, observation, survey, or interviews/qualitative research
What is the best use for this kind of study? To find patterns that might lead to predictions for more complete research project To start describing observed flirtation behavior
What is the main limitation of this kind of study? Hypotheses are not tested Cannot look at cause and effect
Naturalistic Observation A second kind of descriptive method is naturalistic observation, in which the researcher observes and records behavior in the real world. The researcher tries to be as unobtrusive as possible so as not to influence the behavior of interest. Naturalistic observation is more often the design of choice in comparative psychology by researchers who study nonhuman behavior (especially primates) to determine what is and is not unique about our species. Developmental psychologists occasionally also conduct naturalistic observations. For example, the developmental psychologist Edward Tronick of Harvard University has made detailed naturalistic observations of the infants of the Efe people in Zaire. He has tracked these children from 5 months through 3 years to understand how the Efe culture’s communal pattern of child rearing influences social development in children (Tronick, Morelli, & Ivey, 1992). Although the traditional Western view is that having a primary caregiver is best for the social and emotional well-being of a child, Tronick’s research suggests that the use of multiple, communal caregivers can also foster children’s social and emotional well-being. The advantage of naturalistic observation is that it gives researchers a look at real behavior in the real world rather than in a controlled setting—such as in a laboratory, where people might not behave naturally. Few psychologists use naturalistic observation, however, because conditions cannot be controlled and cause-and-effect relationships between variables cannot be demonstrated.
Qualitative Research Letting people say what they want in responding to questions is the essence of an interview. Interviews occur between two people, one asking the questions and the other answering, usually in open-ended answers. Sometimes interview questions are predetermined or structured and sometimes they are spontaneous or unstructured. Interviews are an example of qualitative research, which involves data gathered from open-ended and unstructured answers rather than quantitative or numeric answers. The advantage to qualitative research, namely the open-ended and flexible answers, is also its disadvantage. How does one interpret, summarize, and make sense of a person’s interview answers? Just as importantly, how do we compare one person’s answers to another’s and get a general sense of what the trends are? These are difficulties of qualitative research.
Survey Research Surveys do not have these difficulties because more often than not, they restrict the possible answers to some kind of numeric rating scale, such as 1 for “completely disagree,” 3 for “neither disagree nor agree,” and 5 for “completely agree.” Research that collects information using any kind of numeric and quantifiable scale and often has limited response options is referred to as quantitative research. These are structured and quantitative answers and can be summarized and calculated for trends and averages. Yet, answers are restricted to a few categories and sometimes the response options are too limited and do not capture the person’s true ideas or attitudes. A common example of a limited response option is when a respondent has to answer on a 1 to 5 scale from “completely disagree” to “completely agree.” Sometimes survey research is descriptive and exploratory and other times it may propose and test hypotheses. There is another concern with survey research. Think about your own response when you are contacted via phone or email about participating in a scientific survey. Many of us don’t want to participate and ignore the request. So how does a researcher know that people who participate are not different from people who don’t participate? Maybe those who participate are older or younger, have more education or less education. In other words, we need to know that the information we collect comes from people who represent the group we are interested in, which is known as a representative sample (see Figure 6). Sampling is the procedure researchers use to obtain participants from a population.
SAMPLING. For practical reasons, research is typically conducted with small samples of the population of interest. If a psychologist wanted to study a population of 2,200 people (each face in the figure represents 100 people), he or she would aim for a sample that represented the makeup of the whole group. Thus, if 27% of the population were blue, the researcher would want 27% of the sample population to be blue, as shown in the pie chart on the left. Contrary to what many students think, representative does not mean that all groups have the same numbers.
Americans were shocked by Alfred Kinsey’s initial reports on male and female sexual behavior. Kinsey was the first researcher to survey people about their sexual behavior. For better or worse, his publications changed attitudes about sex.
The well-known Kinsey surveys of male and female sexual behavior provide good examples of the strengths and weaknesses of survey research (Kinsey, Pomeroy, & Martin, 1948; Kinsey et al., 1953). Make no mistake—just publishing such research caused an uproar in both the scientific community and the general public at the time. Kinsey reported, for instance, that up to 50% of the interviewed men but only about half as many (26%) of the women had had extramarital affairs. Another widely cited finding was that approximately 10% of the population could be considered homosexual. The impact of Kinsey’s research has been profound. By itself it began the science of studying human sexuality and permanently changed people’s views. For example, Kinsey was the first to consider sexual orientation on a continuum from 0 (completely heterosexual) to 6 (completely homosexual) rather than as an either-or state with only two options. This approach remains a lasting contribution of his studies. By today’s standards, however, Kinsey’s techniques for interviewing and collecting data were rather primitive. He didn’t use representative sampling and oversampled people in Indiana (his home state) and in prisons—both of which led to biased results. In addition, he interviewed people face-to-face about the most personal and private details of their sex lives, also making it more likely they provided less than honest and biased answers.
Correlational Studies. Once an area of study has developed far enough that predictions can be made, but for various reasons people cannot be randomly assigned to groups or variables cannot be manipulated, a researcher might choose to test hypotheses by means of a correlational study. Correlational designs measure two or more variables and their relationship to one another. In this design, the basic question is, Is X related to Y? For instance, “Is sugar consumption related to increased activity levels in children?” If so, how strong is the relationship, and is increased sugar consumption associated (correlated) with increased activity levels, as we would predict, or does activity decrease as sugar consumption increases? Or is there no clear relationship? Correlational studies are useful when the experimenter cannot manipulate or control the variables. For example, it would be unethical to raise one group of children one way and another group another way in order to study parenting behavior. We could use a good questionnaire to find out whether parents’ scores related to their parenting behavior are consistently associated with particular behavioral outcomes in children. In fact, many questions in developmental psychology, personality psychology, and even clinical psychology are examined with correlational techniques. The major limitation of the correlational approach is that it does not establish whether one variable actually causes the other. Parental neglect in childhood might be associated with antisocial behavior in adolescence, but that does not necessarily mean that neglect causes anti social behavior. Some other variable (e.g., high levels of testosterone, poverty, antisocial friends) might be the cause of the behavior. We must always be mindful that correlation is necessary for causation but is not sufficient by itself to establish causation.
CHARACTERISTICS OF CORRELATIONAL STUDIES. These studies measure two or more variables and their relationship to one another.
What type of questions might be researched? Is one variable related to another variable and how strong is the relationship? Is X related to Y? For example: Do certain styles of flirting get better results? How does this differ for men and women?
What is the most suitable method of answering the question? Questionnaire
When is this study design most appropriate? Most useful when the researcher is unable to manipulate the variables to examine questions
What is the main limitation of this kind of study? Cannot look at cause and effect
Psychologists often use a statistic called the correlation coefficient to draw conclusions from their correlational studies. Correlation coefficients tell us whether two variables relate to each other and the direction of the relationship. Correlations range between −1.00 and +1.00, with coefficients near 0.00 indicating that there is no relationship between the two variables. A 0.00 correlation means that knowing about one variable tells us nothing about the other. As a correlation approaches +1.00 or −1.00, the strength of the relationship increases. Correlation coefficients can be positive or negative. If the relationship is positive, then as a group’s score on variable X increases, its score on variable Y also increases. Height and weight are positively correlated—taller people generally weigh more than shorter people. For negative correlations, as one variable increases, the other decreases. Alcohol consumption and motor skills are negatively correlated—the more alcohol people consume, the less physically coordinated they become. To further demonstrate correlation, let’s consider the positive correlation between students’ scores on midterm and final exams. By calculating a correlation, we know whether students who do well on the midterm are likely to do well on the final. Based on a sample of 76 students in one of our classes, we found a correlation of +0.57 between midterm and final exam grades. This means that, generally, students who did well on the midterm did well on the final. Likewise, those who did poorly on the midterm tended to do poorly on the final. The correlation, however, was not extremely high, so there was some inconsistency. Some people performed differently on the two exams. When we plot these scores, we see more clearly how individuals did on each exam (see Figure 8). Each dot represents one student’s scores on both exams. For example, one student scored an 86 on the midterm but only a 66 on the final. When interpreting correlations, it is important to remember that a correlation does not mean there is a causal relationship between the two variables. Correlation is necessary but not sufficient for causation. When one variable causes another, it must be correlated with it, but just because variable X is correlated with variable Y, it does not mean that X causes Y. The supposed cause may be an effect, or a third variable may be the cause. What if hairiness and aggression in men were positively correlated? Would that imply that being hairy makes a man more aggressive? No. In fact, both hairiness and aggressiveness are related to a third variable, the male sex hormone testosterone.
Experimental Studies. Often people use the word experiment to refer to any research study, but in science an experiment is something quite specific. All psychological studies measure behavior, but a true experiment has two unique characteristics:
1. Experimental manipulation of a predicted cause, the independent variable
2. Random assignment of participants to control and experimental groups or conditions, meaning that each participant has an equal chance of being placed in each group
The independent variable in an experiment is an attribute the experimenter manipulates under controlled conditions. The independent variable is the condition the researcher predicts will cause a particular outcome. The dependent variable is the outcome, or response to the experimental manipulation. You can think of the independent variable as the “cause” and the dependent variable as the “effect,” although reality is not always so simple. If there is a causal connection between the two, then the responses depend on the treatment, hence the name dependent variable. Earlier we mentioned the hypothesis that sugar consumption makes kids overly active. In this example, sugar levels consumed would be the independent variable and behavioral activity level the dependent variable. Recall the study of the effect of caffeine on sex drive in rats. Is caffeine the independent or dependent variable? What about sex drive? Figure 10 features other examples of independent and dependent variables.
INDEPENDENT AND DEPENDENT VARIABLES. Remember: The response, or dependent variable (DV), depends on the treatment. It is the treatment, or independent variable (IV), that the researcher manipulates
Random assignment is a method used to assign participants to different research conditions to guarantee that each person has the same chance of being in one group as another. Random assignment is achieved with either a random numbers table or some other unbiased technique. Random assignment is critical, because it ensures that on average the groups will be similar with respect to all possible variables, such as gender, intelligence, motivation, and memory, when the experiment begins. If the groups are the same on these qualities at the beginning of the study, then any differences between the groups at the end are likely to be the result of the independent variable. Experimenters randomly assign participants to either an experimental group or the control group. An experimental group consists of participants who receive the treatment or whatever is thought to change behavior. In the sugar consumption and activity study, for example, the experimental group would receive a designated amount of sugar.
The control group consists of participants who are treated in exactly the same manner as the experimental group but with one crucial difference: They do not receive the independent variable, or treatment. Instead, they often receive no special treatment or, in some cases, they get a placebo, a substance or treatment that appears identical to the actual treatment but lacks the active substance. In a study on sugar consumption and activity level, an appropriate placebo could be an artificial sweetener. The experimental group would receive the treatment (sugar), and the control group would be treated exactly the same way but would not receive the actual treatment. Instead, the control group could receive a food flavored with an artificial sweetener. Experimental and control groups must be equivalent at the outset of an experimental study so as to minimize the possibility that other characteristics could explain any difference found after the administration of the treatment. If two groups of children are similar at the start and if one group differs from the other on activity level after receiving different amounts of sugar, then we can conclude that the treatment caused the observed effect. That is, different levels of sugar consumption caused the differences in activity level. In our hypothetical study on sugar and activity, for instance, we would want to include equal numbers of boys and girls in the experimental and control groups and match them with respect to age, ethnicity, and other characteristics, so that we could attribute differences in activity level following treatment to differences in sugar consumption only. Suppose we didn’t do a good job of randomly assigning participants to our two conditions and the experimental group ended up with 90% boys but the control group had 90% girls. If, after administering the sugar to the experimental group and the placebo (sugar substitute) to the control group, we found a difference in activity, then we would have two possible explanations for the difference: gender and sugar. Either being male or female caused the difference or consuming large amounts of sugar did. In this case, gender would be a confounding variable—an additional variable that the researcher failed to control for in the experimental design (too many males), and one that could be responsible for a change in the dependent variable (activity level). Because most of the people in the experimental group were male and consumed sugar, we do not know whether being male or consuming sugar was responsible for the difference in active behavior. These two variables are confounded and cannot be teased apart. The power of the experimental design is that it allows us to say that the independent variable (treatment) caused changes in the dependent variable, as long as everything other than the independent variable was held constant. Random assignment guarantees group equivalence on a number of variables and prevents ambiguity over whether effects might be due to other differences between the groups.
CHARACTERISTICS OF EXPERIMENTAL STUDIES.
Only in true experimental designs, in which researchers manipulate the independent variable and measure its effects on the dependent variable, can researchers determine cause and effect. By itself, correlation is not causation.
What type of questions might be researched? Does the independent variable cause the dependent variable? Does X cause Y? Do smiles with raised eyebrows versus those without lead to more offers of dates?
What is the most suitable method of answering the question? Random assignments of participants, controlled experimental conditions in a lab setting
What is the best use for this kind of study? Most useful for the researcher to infer cause
What is the main limitation of this kind of study? Results cannot always be applied to the real world
In addition to random assignment to control and experimental groups, a true experiment requires experimental control of the independent variable. Thus, researchers must make sure that all environmental conditions (such as noise level and room size) are equivalent for the two groups. Again, the goal is to make sure that nothing affects the dependent variable besides the independent variable. In our experiment on sugar consumption and activity level, we first must randomly assign participants to either the experimental group (in which participants receive some amount of sugar) or the control group (in which participants receive some sugar substitute). The outcome of interest is activity level, so each group might be videotaped for a short period 30 minutes after eating the sugar or sugar substitute. What if the room where the experimental group was given the sugar was several degrees warmer than the room where the control group received the sugar substitute, and our results showed that the participants in the warmer room were more active? Could we feel confident that sugar led to in creased activity level? No, because the heat in that room may have caused the increase in activity level. In this case, room temperature would be the confounding variable. There is a design that has one quality of an experiment (the manipulation of an independent variable), but not the other (random assignment), a design known as a quasi-experimental design (Shadish, Cook, & Campbell, 2002). It is called “quasi” because it is partly or almost an experimental design. It is often used with human participants, because we cannot randomly assign them to ethically untenable conditions, such as depression or childhood abuse or even personality traits such as extroversion. In these situations, the researchers use preexisting groups but then manipulate a variable to see how the groups respond.
Any knowledge that participants and experimenters have about the experimental conditions to which participants have been assigned can also affect the outcome of an experiment. In single-blind studies, participants do not know the experimental condition to which they have been assigned. This is a necessary precaution in all studies to avoid the possibility that participants will behave in a biased way. For example, if participants know they have been assigned to a group that receives a new training technique on memory, then they might try harder to perform well. This would confound the results. Another possible problem can come from the experimenter knowing who is in which group and unintentionally treating the two groups somewhat differently. This could lead to the predicted outcome simply because the experimenter has biased the results. In double-blind studies, neither the participants nor th researchers (at least the ones administering the treatment) know who has been assigned to which condition. Ideally, then, neither the participants nor those collecting the data should know which group is the experimental group and which is the control group. The advantage of double-blind studies is that they prevent two potential problems with experimental designs: experimenter expectancy effects and demand characteristics. Experimenter expectancy effects occur when the behavior of the participants is influenced by the experimenter’s knowledge of who is in which condition (Rosenthal, 1976, 1994). Demand characteristics are subtle cues given by experimenters to the participants as to how they should behave in the role of participant. These cues may be provided unconsciously, or even consciously. The latter apparently happened, for example, in the Stanford Prison Study when before the study began, Zimbardo suggested to the assigned “guards” that although they cannot abuse or torture, they can create fear, boredom, frustration, lack of privacy, and control (Zimbardo, 2007). That is precisely what the guards did.
Longitudinal Studies. People change over time—our brains, thoughts, feelings, personality, motivations. In fact, every aspect of us changes. The only way to study change over time is with longitudinal studies. Longitudinal designs make observations of the same people over time, ranging from months to decades. These kinds of studies are not only useful for studying change over time, but also can be used to study how specific causes affect specific outcomes. For example, we cannot randomly assign young children to live in abusive or neglectful home environments. But some children do and some do not. So if we follow these groups over time, we can examine the outcomes of such living conditions on thought and behavior. Longitudinal studies often combine observational, correlational, and quasi-experimental techniques.
Twin-Adoption Studies. One of the most important questions in psychology is how much of our thought, behavior, personality, and mental health stems from built-in biological forces (nurture) and how much is from learning and the environmental (nature). The nature-nurture topic runs throughout the book and our answer is always it is “both/and” not “either/or.” How do researchers conduct research to answer these questions? They primarily use two different methods: twin-adoption studies and gene-by-environment studies.
Twin-Adoption Studies The best way to untangle the effects of genetics and environment is to study twins who are adopted or not and compare them to other siblings who are adopted or not, which is what twin-adoption studies do. We will be discussing twin-adoption research throughout the rest of the book, so it is important that you understand the logic of how these studies tease apart the extent to which nature and nurture is involved in creating differences or similarities between people The easiest way to understand how twin-adoption research teases nature and nurture effects apart is to realize that there are three forms of similarity: genetic (nature), environmental (nurture), and trait. Genetic similarity varies based on degree of relationship, for example in cases of twins, siblings, and parents and their children. Genetic similarity ranges from 100% with identical twins (split from a single egg) to 0% with unrelated people, such as adopted siblings. In be tween 100% and 0% we have fraternal twins (two eggs fertilized by two sperm), siblings, and parents and their children. These all share 50% of their genes. Environments vary in many ways, the most obvious being raised in the same house (together) or not (apart). Researchers who make use of twin-adoption methods then look for how similar or different these different groups are on a given trait in order to calculate how much of the variation in that trait is due to genetic influence, or its heritability. Researchers cannot for obvious reasons assign people to different genetic or environmental conditions to see what effect each one has on a trait. So they turn to what exists in nature to tease the two apart. The logic of all of this is laid out in Figure 12. Let’s just take two of these hypothetical cases. First, if identical twins (100% genetically alike) are raised apart (different environments) and yet they still are very similar on certain traits such as their height, weight, intelligence, or personality, then we would know that genes play a large role in the variation of those traits. Likewise, if adopted siblings (0% genetically alike) who are raised together (same environment) share strong similarities on these traits, then we have to conclude that environment plays a large role in the variation of those traits.
Gene-by-Environment Studies The second technique in the study of heritability, gene-by-environment interaction research, allows researchers to assess how genetic differences interact with the environment to produce certain behavior in some people but not in others (Moffitt, Caspi, & Rutter, 2005; Thapar, Langley, & Asherson, 2007). Instead of using twins, family members, and adoptees to vary genetic similarity, gene-by-environment studies directly measure genetic variation in parts of the genome itself and examine how such variation interacts with different kinds of environments to produce different behaviors or traits.
Meta-Analysis. As powerful as results may be from an individual study, the real power of scientific results comes from the cumulative overall findings from all studies on a given topic. If a topic or question has been sufficiently studied, researchers may choose to stand back and analyze all the results of the numerous studies on a given topic. For example, a researcher interested in the effects of media violence on children’s aggressive behavior might want to know what all of the research not just one or two studies—suggests. Meta-analysis is a quantitative method for combining the results of all the published and even unpublished results on one question and drawing a conclusion based on the entire set of studies on the topic. To do a meta-analysis, the researcher converts the findings of each study into a standardized statistic known as effect size. Effect size is a measure of the strength of the relationship between two variables. The average effect size across all studies reflects what the literature overall says on a topic or question. For example, a meta-analysis of 28 different studies from the United States, Europe, and Asia compared technology-based learning (e.g., tablets, laptops, and smart phones) to traditional non-technology-based instruction and found that students using tablets learned better than those without. The average effect size for these 28 studies was 0.23, meaning that students using tablets scored 0.23 of a standard deviation (the average variation around the mean) higher on learning outcomes compared to students with no tablets (Tamin et al., 2015). (See “Making Sense of Data with Statistics” later in this chapter.) In short, meta-analysis tells us whether all of the research on a topic has or has not led to consistent findings and what the effect size is. It is more reliable than the results of any single study.
Big Data. More than 1 billion of the world’s 7 billion people use some form of social media (Facebook, Twitter, Instagram, and others). With the increase in storage and speed of the Internet and the proliferation of mobile apps, a whole new way of collecting data on human behavior has arisen—so-called “big data.” Big Data consists of the extremely vast amounts of information from websites and apps that is collected and analyzed by unusually large and sophisticated computer programs. Big Data come mostly from social media, smartphones, and wearable devices, but sometimes from the scientific literature itself (Augur, 2016).
In psychology, compared to the traditional survey, questionnaire, and even experimental techniques, Big Data afford a much more extensive and reliable means of measuring interests, social relationships, personality, emotion, political attitudes, exercise behaviors, brain activity and structure, and language just to name a few (Schwartz & Ungar, 2015). For example, Tan and colleagues (2016) used Big Data to find and narrow down all brain studies on hippocampus size (a brain structure most notably associated with learning and memory) and gender. By doing so, they narrowed the field down to 76 studies with more than 6,000 participants. Further, using meta-analysis, once they controlled for overall brain size (because men are larger and have larger overall brains), the researchers found there was no size difference in the hippocampi of men and women, further eroding the assumption that men and women have different brains. This study took advantage of both Big Data and meta-analysis.
CHALLENGING ASSUMPTIONS IN THE OBJECTIVITY OF EXPERIMENTAL RESEARCH
You don’t have to be a scientist to understand that it would be wrong and unethical for an experimenter to tell participants how to behave and what to do. Even for the participants to know what group they are in or what the hypotheses of the study are is bad science and biases behavior. Can what the experimenter knows change the behavior of the participants? In a classic case of scientific serendipity, Robert Rosenthal’s PhD thesis challenged the assumption that experimenters who randomly assign animals or people to conditions and manipulate an independent variable are being quite objective—that is, these procedures assure objective results. He discovered the assumption of objectivity was wrong when he set out to conduct a study on perceived success and intelligence. Rosenthal hypothesized that people who believed they were successful would be more likely to see success in others. To test this idea, he conducted an experiment in which he told one group of participants they had done well on an intelligence test and another group they had done poorly on an intelligence test. Rosenthal randomly assigned participants to be in one of these conditions (there was also a neutral control condition in which participants received no feedback on the intelligence test). Then he asked all groups to look at photographs of people doing various tasks and rate how successful they thought the people in the photos were. He reasoned that people who are told they did well on an intelligence test should see more success in photographs of people doing various tasks than people who are told they did not do well on the test.
As a good scientist, Rosenthal compared the average test scores of the participants assigned to different conditions before giving them any feedback on their performance—that is, before the experimental treatment. The reason is simple: If the treatment causes a difference in behavior for the different groups, the researcher needs to make sure the groups started out behaving the same way before treatment. To Rosenthal’s dismay, the groups did differ before receiving treatment. They were also different in exactly the way that favored his hypothesis! Given random assignment, the only difference in the groups at the outset was Rosenthal’s knowledge of who was in which group. Somehow, by knowing who was in which group, he created behaviors that favored his hypothesis. He was forced to conclude that, even when trying to be “scientific” and “objective,” researchers bias results unintentionally in their favor by subtle voice changes or gestures. Instead of having a wonderful “aha moment” of scientific discovery, Rosenthal had more of an “oh no” moment: “What I recall was a panic experience when I realized I’d ruined the results of my doctoral dissertation by unintentionally influencing my research participants to respond in a biased manner because of my expectations” (Rosenthal, personal communication, April 18, 2010).
Rosenthal decided to systematically study what he came to call experimenter expectancy effects. Through several experiments, he confirmed that experimenter expectancies can ruin even the best-designed studies. Also, he discovered that two other surprising factors can change the outcome of the study as well. First, if the study involves direct interaction between an experimenter and participants, the experimenter’s age, ethnicity, personality, and gender can influence the participants’ behavior (Rosenthal, 1976). Second, Rosenthal stumbled upon a more general phenomenon known as self-fulfilling prophecy. A self-fulfilling prophecy occurs when our belief or expectation that something is going to happen unknowingly makes it happen. For example, if research assistants know the study’s hypothesis, they may unconsciously affect the behavior of the participants and make the predicted outcome more likely to happen. Or if a teacher does not expect a student to do very well and then ignores that student, the student may well not do well, fulfilling the expectation of the teacher. Ten years after Rosenthal’s first publication on experimenter expectancy effect, more than 300 other studies confirmed his results (Rosenthal & Rubin, 1978). Such expectancies affect animal participants as well as humans (Jussim & Harber, 2005; Rosenthal & Fode, 1963). Rosenthal’s demonstration of experimenter expectancy effects and self-fulfilling prophecies also led to the development of double-blind procedures in science. Think about it: If what experimenters know about a study can affect the results, then they’d better be as blind to experimental conditions as the participants are. All of this came to be because Rosenthal “messed up” his dissertation and unintentionally challenged the assumptions of the best way to conduct scientific experiments.
COMMONLY USED MEASURES OF PSYCHOLOGICAL RESEARCH
In addition to different study methods, when psychologists conduct research, they rely on a vast array of tools to measure variables relevant to their research questions. The tools and techniques they use to assess thought and behavior are called measures. Measures in psychological science tend to fall into three categories: self-report, behavioral, and physiological. To study complex behaviors, researchers may employ multiple measures (see Figure 13).
Self-Report Measures
Self-reports are people’s written or oral accounts of their thoughts, feelings, or actions. Two kinds of self-report measures are commonly used in psychology:
• Interviews
• Questionnaires
In an interview, a researcher asks a set of questions, and the respondent usually answers in any way he or she feels is appropriate. The answers are often open ended and not constrained by the researcher. (See the section “Descriptive Studies”, for additional discussion on interviews.) In a questionnaire, responses are limited to the choices given in the questionnaire. In the Stanford Prison Experiment, for example, the researchers used questionnaires to keep track of the psychological states of the prisoners and guards. They had participants complete mood questionnaires many times during the study, so that the researchers could track any emotional changes the participants experienced. The participants also completed forms that assessed personality characteristics, such as trustworthiness and orderliness, that might be related to how they acted in a prison environment (Haney et al., 1973). Self-report questionnaires are easy to use, especially in the context of collecting data from a large number of people at once. They are also relatively inexpensive. If designed carefully, questionnaires can provide important information on key psychological variables. A major problem with self-reports, however, is that people are not always the best sources of information about themselves. Why? Sometimes, as a reflection of the tendency toward social desirability, called social desirability bias, people present themselves more favorably than they really are, not wanting to reveal what they are really thinking or feeling to others for fear of looking bad. Presented with questions about social prejudice, for example, respondents might try to avoid giving answers that suggest they are prejudiced against a particular group. Another problem with self-reports is that we have to assume that people are accurate witnesses to their own experiences. Of course, there is no way to know exactly what a person is thinking without asking that person, but people do not always have clear insight into how they might behave (Nisbett & Wilson, 1977).
COMMONLY USED MEASURES IN PSYCHOLOGY.
Behavioral Measures
Behavioral measures involve the systematic observation of people’s actions either in their normal environment (that is, naturalistic observation) or in a laboratory setting. A psychologist interested in aggression might bring people into a laboratory, place them in a situation that elicits aggressive behavior, and video tape the responses. Afterward, trained coders observe the videos and, using a prescribed method, code the level of aggressive behavior exhibited by each person. Training is essential for the coders, so that they can evaluate the video and apply the codes in a reliable, consistent manner. Behavioral measures are less susceptible to social desirability bias than are self-report measures. They also provide more objective measurements, because they come from a trained outside observer, rather than from the participants themselves. This is a concern for researchers on topics for which people are not likely to provide accurate information in self-report instruments. In the study of emotion, for example, measuring facial expressions from video reveals things about how people are feeling that they might not reveal on questionnaires (Rosenberg & Ekman, 2000). One drawback of behavioral measures is that people may modify their behavior if they know they are being observed and/or measured. The major drawback of behavioral measurement, however, is that it can be time-intensive; it takes time to train coders to use the coding schemes, to collect behavioral data, and to prepare the coded data for analysis. As a case in point, one of the most widely used methods for coding facial expressions of emotion requires intensive training, on the order of 100 hours, for people to be able to use it correctly (Ekman, Friesen, & Hager, 2002)! Moreover, researchers can collect data on only a few participants at once, and therefore behavioral measures are often impractical for large-scale studies. \
Physiological Measures
Physiological measures provide data on bodily responses. For years, researchers relied on physiological information to index possible changes in psychological states—for example, to determine the magnitude of a stress reaction. Research on stress and anxiety often measures electrical changes in involuntary bodily responses, such as heart rate, sweating, and respiration, as well as hormonal changes in the blood that are sensitive to changes in psychological states. Some researchers measure brain activity while people perform certain tasks to determine the speed and general location of cognitive processes in the brain. We will look at specific brain imaging technologies in the chapter “The Biology of Behavior”. Here we note simply that they have enhanced our understanding of the brain’s structure and function tremendously. However, these technologies, and even more simple ones, such as the measurement of heart rate, often require specialized training in the use of equipment, collection of measurements, and interpretation of data. Further, some of the equipment is expensive to buy and maintain. Outside the health care delivery system, only major research universities with medical schools tend to have them. In addition, researchers need years of training and experience in order to use these machines and interpret the data they generate.
MAKING SENSE OF DATA WITH STATISTICS
Once researchers collect data, they must make sense of them. Raw data are difficult to interpret. They are, after all, just a bunch of numbers. It helps to have some way to organize the information and give it meaning. To make sense of information, scientists use statistics, mathematical procedures for collecting, analyzing, interpreting, and presenting numeric data. For example, researchers use physiological measures. Measures of bodily responses, such as blood pressure or heart rate, used to determine changes in psychological state.
statistics. The collection, analysis, interpretation, and presentation of numerical data. statistics to describe and simplify data and to understand how variables relate to one another. There are two classes of statistics: descriptive and inferential.
Descriptive Statistics
The first step in understanding research results involves calculating descriptive statistics, which simply tell researchers the range, average, and variability of the scores. For instance, one useful way to describe data is by calculating the center, or average, of the scores. There are three ways to calculate an average the mean, median, and mode. The mean is the arithmetic average of a series of numbers. It is calculated by adding all the numbers together and dividing by the number of scores in the series. An example of a mean is your GPA, which averages the numeric grade points for all of the courses you have taken. The median is the middle score, which separates the lower half of scores from the upper half. The mode is the most frequently occurring score. Sometimes scores vary widely among participants, but the mean, median, and mode do not reveal anything about how spread out—or how varied—scores are. For example, one person’s 3.0 GPA could come from getting B’s in all his courses, while another person’s 3.0 could result from getting A’s in half her classes and C’s in the other half. The second student has much more variable grades than the first.The most common way to represent variability in data is to calculate the standard deviation, a statistical measure of how much scores in a sample vary around the mean. A higher standard deviation indicates more variability (more spread); a lower one indicates less variability (less spread). In the example, the student with all B’s would have a lower standard deviation than the student with A’s and C’s. Another useful way of describing data is by plotting, or graphing, their frequency. Frequency is the number of times a particular score occurs in a set of data. A graph of frequency scores is known as a distribution. To graph a distribution, we place the scores on the horizontal axis, or X-axis, and their frequencies on the vertical axis, or Y-axis. When we do this for many psychological variables, such as intelligence or personality, we end up with a very symmetrical shape to our distribution, which is commonly referred to as either a normal distribution or a “bell curve”—because it looks like a bell. Let’s look at a concrete example of a normal distribution with the well known intelligence quotient (IQ). If we gave 1,000 children an IQ test and plotted all 1,000 scores, we would end up with something very close to a symmetrical, bell-shaped distribution. Very few children would score 70 or below, and very few children would score 130 or above. These are infrequent or rare scores. The majority of children would be right around the average, or mean, of 100. In fact, two-thirds (68%, to be exact) would be within 1 standard deviation (15 points) of the mean. These are frequent or common scores. Moreover, about 95% would be within 2 standard deviations, or between 70 and 130. How do we know this? We know it because we know the exact shape of a normal distribution in the general population of people being studied. Knowing the shape of the distribution allows us to make inferences from our specific sample to the general population. For example, because a normal distribution has a precise shape, we know exactly what percentage of scores is within 1 standard deviation of the mean (68%) and how many are within 2 standard deviations of the mean (95%). This is why we know that a mean IQ score of 70 or lower or 130 or higher occurs only 5 times in 100—both are very unlikely to occur by chance. This quality of allowing conclusions or inferences to be drawn about populations is the starting point for the second class of statistics, inferential statistics.
Inferential Statistics
We do not draw any conclusions from descriptive results. We just describe the scores with them. Inferential statistics, however, allow us to test hypotheses and draw a conclusion (that is, make an inference) as to how likely a sample score is to occur in a population. They also allow us to determine how likely it is that two or more samples came from the same population. In other words, inferential statistics use probability and the normal distribution to rule out chance as an explanation for why group scores are different. What is an acceptable level of chance before we say that a score is not likely to occur by chance? Five in 100 (5%) is the most frequent choice made by psychological researchers and is referred to as the probability level. So if we obtain two means and our statistical analysis tells us there is only a 5% or less chance that these means come from the same population, we conclude that the numbers are not just different but statistically different and therefore not likely to be due to chance. Researchers use many kinds of statistical analyses to rule out chance, but the most basic ones involve the comparison of two or more means. To compare just two means, we use a statistic known as the t-test. The basic logic of t-tests is to determine whether the means for your two groups are so different that they are not likely to come from the same population. If our two groups are part of an experiment and one is the experimental group and the other the control group, then we are determining whether our treatment caused a significant effect, seen in different means. In short, t-tests allow us to test our hypotheses and rule out chance as an explanation. Let’s look at an example, by returning to a question we considered earlier: Does sugar cause hyperactive behavior in children? We will make the commonsense prediction that sugar does cause hyperactive behavior. We randomly assign 100 children to consume sugar (experimental group); another 100 children do not consume sugar (control group). We then wait 30 minutes—to let the sugar effect kick in—and observe their behavior for an additional 30 minutes. We video record each child’s behavior and code it on number of “high activity acts.” If sugar causes activity levels to increase, then the sugar groups number of high activity acts should be higher than those of the no-sugar group. Our data show that the experimental (sugar) group exhibited an average of 9.23 high activity behaviors in the 30 minutes after eating the sugar; the control (no-sugar) group exhibited an average of 7.61 such behaviors. On the face of it, our hypothesis seems to be supported. After all, 9.23 is higher than 7.61. However, we need to conduct a statistical test to determine whether the difference in the number of hyperactive behaviors between our groups of kids who ate sugar versus those who did not really represents a true difference between these two different populations of kids in the real world.
RESEARCH ETHICS
Due to current ethical guidelines, some of the most important and classic studies in psychology could not be performed today. One of them is the Stanford Prison Experiment, which you read about at the beginning of this chapter. This experiment subjected participants to conditions that so altered their behavior the researchers had to intervene and end the study early. In 1971, there were few ethical limitations on psychological research. Since then, and partly as a consequence of studies like the Stanford Prison Experiment, professional organizations and universities have put in place strict ethical guidelines to protect research participants from physical and psychological harm. Ethics are the rules governing the conduct of a person or group in general or in a specific situation; stated more simply, ethics are standards of right and wrong. What are the ethical boundaries of the treatment of humans and animals in psychological research? In psychology today, nearly every study conducted with humans and animals must pass through a rigorous review of its methods by a panel of experts. If the proposed study does not meet the standards, it cannot be approved. Ethics involve rules against scientific misconduct and rules for treatment of human participants and animals.
Scientific Misconduct
Ethical violations and scientific misconduct in science range from honest errors or mistakes to fraud. Errors can be the result of carelessness, bias, or mistakes, but tend not to be intentional. Scientific misconduct, however, is intentional and therefore the most serious ethical violation. According to both the National Science Foundation and the American Psychological Association, scientific fraud or misconduct comes in three forms: plagiarism, falsification, and fabrication (Code of Federal Regulations-689, 2002; Research Misconduct, n.d.). Plagiarism is when someone presents the words or ideas of other people as their own. Falsification is changing, altering, or deleting data. The most serious and blatant form of scientific misconduct and fraud is when a researcher commits scientific fabrication, that is, presenting or publishing scientific results that are made up. Fortunately, fraud and misconduct are relatively rare in science estimates are about 1 in 5,000 to 1 in 23,000 papers need to be taken back or retracted for serious errors or misconduct (Steen, 2010; Steen et al., 2013). One of the worst recent cases of scientific fraud recently involved a social psychologist by the name of Diederik Stapel, who was an up and-coming psychologist, publishing in the top journals. But one of his studies did not provide the results he predicted. In his words: “I said, you know what, I am going to create the dataset” (Bhattacharrjee, 2013). So he sat down at his kitchen table and literally fabricated an entire dataset. It worked. The paper was published in a top psychological journal. Over the next dozen years or so, a total of 55 scientific articles and 10 PhD theses were based on fabricated data. When it was all over, Stapel was fired, lost his career, and admitted “I have failed as a scientist” (Bhattacharrjee, 2013).
Ethical Treatment of Human Participant
Fraud or errors are one thing but ethical issues also arise in the treatment of participants and animals. For example, the classic series of studies by Stanley Milgram in the early 1960s. Milgram’s landmark research on obedience is discussed in more detail in the chapter “Social Behavior”, but we mention it here for its pivotal role in the development of ethical guidelines for human psychological research. Milgram, like many other social psychologists of the mid-20th century, was both fascinated and horrified by the atrocities of the Holocaust and wondered to what extent psychological factors influenced people’s willingness to carry out the orders of the Nazi regime. Milgram predicted that most people are not inherently evil and argued that there might be powerful aspects of social situations that make people obey orders from authority figures. He designed an experiment to test systematically the question of whether decent people could be made to inflict harm on others. Briefly, Milgram’s studies of obedience involved a simulation in which participants were misled about the true nature of the experiment. Thinking that they were part of an experiment on learning, they administered what they thought were electrical shocks to punish the “learner,” who was in another room, for making errors. In spite of protests from the learner when increasingly intense shocks occurred, the experimenter pressured the “teachers” to continue administering shocks. Some people withdrew from the study, but most of the participants continued to shock the learner. After the study, Milgram fully explained to his participants that the learner had never been shocked or in pain at all (Milgram, 1974). Milgram’s study provided important data on how easily decent people could be persuaded by the sheer force of a situation to do cruel things. What is more Milgram conducted many replications and variations of his findings, which helped build knowledge about human social behavior. Was it worth the distress it exerted on the participants? One could ask the same of the Stanford Prison Experiment. Zimbardo appears to have coached the guards by suggesting behaviors they could do and these are the one’s they ended up doing. Although the prison experiment led to some reform in U.S. prisons, it is hard to know whether the deception of the participants and the emotional breakdowns some of them experienced was worth it. What do you think? The Milgram study is one of the most widely discussed studies in the history of psychology. A number of psychologists protested it on ethical grounds (Baumrind, 1964). The uproar led to the creation of explicit guidelines for the ethical treatment of human subjects. Today all psychological and medical researchers must adhere to the following guidelines:
1. Informed consent: Tell participants in general terms what the study is about, what they will do and how long it will take, what the known risks and benefits are, and whom to contact with questions. They must also be told that they have the right to withdraw at any time without penalty. This information is provided in written form and the participant signs it, signifying consent. If a participant is under the age of 18, informed consent must be granted by a legal guardian. Informed consent can be omitted only in situations such as completely anonymous surveys.
2. Respect for persons: Safeguard the dignity and autonomy of the individual and take extra precautions when dealing with study participants, such as children, who are less likely to understand that their participation is voluntary.
3. Beneficence: Inform participants of costs and benefits of participation; minimize costs for participants and maximize benefits. For example, many have argued that the Milgram study was worth the distress (cost) it may have caused participants, for the benefit of the knowledge we have gained about how readily decent people can be led astray by powerful social situations. In fact, many of the participants said that they were grateful for this opportunity to gain knowledge about themselves that they would have not predicted (Milgram, 1974)
4. Privacy and confidentiality: Protect the privacy of the participant, generally by keeping all responses confidential. Confidentiality ensures that participants’ identities are never directly connected with the data they provide in a study.
5. Justice: Benefits and costs must be distributed equally among participants.
In Milgram’s study, participants were led to believe they were taking part in a learning study when, in fact, they were participating in a study on obedience to authority. Is this kind of deception ever justified? The answer (according to the American Psychological Association, APA) is that deception is to be avoided whenever possible, but it is permissible if these conditions are met: It can be fully justified by its significant potential scientific, educational, or applied value; it is part of the research design; there is no alternative to deception; and full debriefing occurs afterward. Debriefing is the process of informing participants of the exact purposes of the study—including the hypotheses—revealing any and all deceptive practices and explaining why they were necessary to conduct the study and ultimately what the results of the study were. Debriefing is required to minimize any negative effects (e.g., distress) experienced as a result of the deception. Deception comes in different shades and degrees. In the Stanford Prison Experiment, all participants were fully informed about the fact that they would be assigned the roles of a prisoner or a guard. In that sense there was no deception. They were not informed of the details and the extent to which being in this study would be like being in a real prison world. They were not told upfront that, if they were assigned to the “prisoner” role, they would be strip-searched. When they were taken from their homes, the “prisoners” were not told this was part of the study. Not informing participants of the research hypotheses may be deceptive but necessary to prevent biased and invalid responses. Not telling participants that they might experience physical pain or psychological distress is a much more severe form of deception and is not ethically permissible. Today, to ensure adherence to ethical guidelines, institutional review boards (IRBs) evaluate proposed research before it is conducted to make sure research involving humans does not cause undue harm or distress. Should Milgram’s study have been permitted? Were his procedures ethical by today’s standards? To this day, there are people who make strong cases both for and against the Milgram study on ethical grounds, as we have discussed. It is harder to justify what Zimbardo did in the prison experiment.
Ethical Treatment of Animals
Human participants are generally protected by the ethical guidelines itemized in the previous section. What about animals? They cannot consent, so how do we ethically treat animals in research? The use of nonhuman species in psychological research is even more controversial than is research with humans. There is a long history in psychology of conducting research on animals. Typically, such studies concern topics that are harder to explore in humans. We cannot, for instance, isolate human children from their parents to see what effect an impoverished environment has on brain development. Researchers have done so with animals. The subfields of biological psychology and learning most often use animals for research. For instance, to determine what exactly a particular brain structure does, one needs to compare individuals who have healthy structures to those who do not. With humans this might be done by studying the behavior of individuals with accidental brain injury or disease and comparing it to the behavior of normal humans. Injury and disease, however, never strike two people in precisely the same way, so it is not possible to reach definite conclusions about the way the brain works by just looking at accidents and illness. Surgically removing the brain structure is another way to determine function, but this approach is obviously unethical with humans. In contrast, nonhuman animals, usually laboratory rats, offer the possibility of more highly controlled studies of selective brain damage. For example, damage could be inflicted on part of a brain structure in one group of rats while another group is left alone. Then the rats’ behaviors and abilities could be observed to see whether there were any differences between the groups. Animals cannot consent to research, and if they could, they would not likely agree to any of this. Indeed, it is an ongoing debate as to how much animal research should be permissible at all. Because animal research has led to many treatments for disease (e.g., cancer, heart disease), as well as advances in understanding basic neuroscientific processes (such as the effects of environment on brain cell growth), it is widely considered to be acceptable. Animal research is acceptable, that is, as long as the general conditions and treatment of the animals is humane. If informed consent is the key to ethical treatment of human research participants, then humane treatment is the key to the ethical use of animal subjects. The standards for humane treatment of research animals involve complex legal issues. State and federal laws generally require housing the animals in clean, sanitary, and adequately sized structures. In addition, separate IRBs evaluate proposals for animal research. They require researchers to ensure the animals’ comfort, health, and humane treatment, which also means keeping discomfort, infection, illness, and pain to an absolute minimum at all times. If a study requires euthanizing the animal, it must be done as painlessly as possible. Despite the existence of legal and ethical safeguards and the importance for medical research in humans, some animal rights groups argue that any and all animal research should be discontinued, unless it directly benefits the animals. These groups contend that computer modeling can give us much of the knowledge sought in animal studies and eliminates the need for research with animals. In addition, current brain imaging techniques, which allow researchers to view images of the living human brain, reduce the need to sacrifice animals to examine their brain structures. As is true of all ethical issues, complex and legitimate opposing needs must be balanced in research. The need to know, understand, and treat illness must be

balanced against the needs, well-being, and rights of participants and animals. Consequently, the debate and discussion about ethical treatment of humans and animals must be ongoing and evolving.