language and reading 1

Page 1: Introduction

  • Course: PSGY1002 - Language and Reading Lecture 1: Word Recognition

  • Instructor: Dr. Ruth Filik

  • Location: Room B53, The University of Nottingham

  • Contact: ruth.filik@nottingham.ac.uk

Page 2: Humor

  • Joke: Two goldfish are sitting in a tank. One says: "You man the guns… I’ll drive!"

Page 3: Lecture Schedule

  • Lecture Topics:

    • Lecture 1: Word Recognition

    • Lecture 2: Sentence Processing

    • Lecture 3: Discourse Processing

  • Background Reading: Eysenck & Keane (2020). Cognitive Psychology: A Student’s Handbook (Eighth Edition).

    • Lecture 1: Chapter 9

    • Lecture 2: Chapter 10

    • Lecture 3: Chapter 10

Page 4: Word Recognition

  • General Issue:

    • Language learners associate visual patterns with meanings.

    • Development of a mental lexicon (memory for words) occurs during this process.

    • So what is the general issue in word recognition? What is it that people actually have to do in order to recognise a word?

      Basically, the reader sees a perceptual pattern that is meaningless in itself - think of seeing a word in a language you don’t know - like this extract from a Georgian book of prayers.

      The pattern somehow conveys meaning because:

      - In learning a language, a person learns to associate visual patterns with meanings.

      - They do this by creating a store of knowledge about the words of the language - the mental lexicon.

      The question then is - How does a particular occurrence of a perceptual pattern get associated with the right meaning?

      So for example, when you read the word ‘pattern’, how do you know what it means?

Page 5: Learning Outcomes

  • Understand common methods for studying visual word recognition.

  • Identify factors influencing word recognition.

  • Learn models of word recognition and their experimental relations.

Page 6: Methods Overview

  • Discuss methods used to study word recognition.

  • There are a number of tasks that researchers use to study word recognition.

    One is eye-tracking, where you simply monitor people’s eye movements as they are reading, and measure how long they spend looking at a particular word.

    Another very commonly used task is the lexical decision task. In this task, participants are presented with strings of letters, some of which form words, and some of which don’t, and are simply asked to decide whether the letter string that they have been presented with forms a word or not. The lexical decision task is often used in conjunction with priming.

    Some people use a naming task. As the title suggests, you simply present someone with a word, and then measure how long it takes them to say it out loud. The longer it takes, the harder it was to recognise the word.

    I’ll now go through each of these methods in a bit more detail.

Page 7: Factors Affecting Word Recognition

  • Investigative Methods:

    • Eye-tracking: Duration spent on words. One method which is commonly used is eye-tracking. In an eye-tracking experiment, the participant would typically read text from a computer screen, while a camera monitors which word they are looking at. This slide gives an example of what your eye does when you are reading a piece of text. So you will typically make one or two fixations on each content word, and then move on to the next one.

      Analysing this data can then give us information like - how long a person spent looking at each word, whether they skipped a particular word, and whether they went back to read it again.

      If they spent a long time looking at a particular word it suggests that word recognition was particularly difficult in this case. On the other hand, if they spent a short time looking at it, or even skipped it altogether (which is common with function words like ‘and’ and ‘the’) this suggests that the word was very easy to recognise.

    • Lexical Decision Task: Time taken to affirm a letter string is a word.Another commonly used task is the lexical decision task. I’ll just give you an example of what that might be like for a participant. During the experiment, they would be presented with a series of letter strings on a computer screen, and be asked to indicate, as quickly and accurately as possibly, whether what they were presented with was a word, or a non-word. They would do this by pressing a button labelled ‘yes’ or ‘no’. So for example…

      The experimenter would simply record how long it took participants to respond in each case. The longer it takes, the more difficult the word was to recognise. Non-words are included so that the correct answer isn’t always ‘yes’.

      If you would like to have a go at this task after the lecture, I have put a link on Moodle which you can click on to try the lexical decision task for yourself.

      Lexical decision is often used in conjunction with priming

    • Naming Task: Time until the participant starts saying a word.

Pages 8-9: Methods Continue

  • Eye-tracking and Lexical Decision Task details.

Page 10: Lexical Decision Task and Priming

  • Context: Typically used with priming (e.g., doctor, pudor, table, tadjid).

  • Priming: Priming is where the participant is 'primed' with a certain stimulus before the actual lexical decision task has to be performed. In this way, it has been shown that participants are faster to respond to words when they are also shown a semantically related prime. For example if participants are simply presented with the word ‘doctor’, and asked to decide whether it is a word or not, it might take them 450 ms to press a button to indicate that it is, indeed, a word. However, if they are also presented with a related word, such as ‘nurse’, the time it will take them to identify ‘doctor’ will be less. In addition, participants are faster to recognise ‘doctor’ in the context of a related word such as ‘nurse’, than in the context of an unrelated word, such as ‘butter’. This is known as a semantic priming effect.

Page 11-12: Naming Task and Sample Questions

  • Setup involves timed responses; sample MCQ relates to eye-tracking.

  • In each trial, participants would be presented with a word on a computer screen and told that they should pronounce it as quickly as possible without stuttering or mispronouncing it. At the beginning of each trial, a warning tone may sound for say 250 msec. Then a word is displayed on the monitor 250 msec following tone offset. When a response is detected by the computer, the word will be erased from the screen. After a response is recorded, there is say a 2-second interval before the tone signalling the next trial.

    To analyse the data, the experimenter would score the trial as either a correct pronunciation of the word, or as an “error” if the word was mispronounced or if an irrelevant noise (such as a cough) triggered the voice-key in the computer. Whether or not the response was correct, and how long it took the participant to say the word out loud, are taken as an indication of how difficult the word was to recognise.

Page 13: Factors Affecting Word Recognition

  • Word frequency, predictability, and neighbourhood effects are key factors.

  • The three factors that we will cover are word frequency, predictability, and neighbourhood effects. In terms of word frequency, studies have shown that commonly used words are recognised more easily than infrequent words. Predictable words are also recognised more easily than those in neutral or misleading contexts. Finally, neighbourhood effects show that word identification can be speeded when similar words exist in the language. I’ll now go through and explain each of these effects in a bit more detail.

Pages 14-19: Word Frequency and Predictability

  • High frequency words recognized easier; predictable context improves recognition rate.

  • Numerous studies have demonstrated that readers find it easier to recognise high frequency words (that is, words which they encounter very often), than low frequency words (so, words that they don’t encounter very often). For example, Schilling et al. used a naming task, a lexical decision task, and an eye-tracking task to investigate the word frequency effect. I’ve put this paper on Moodle, as you might find it useful, particularly if you are considering writing your coursework essay for this module on word recognition.

    The naming task and lexical decision task were used to examine the speed of recognition of words presented in isolation. For example, participants were presented with high frequency words such as “teacher”, or low frequency words such as “armadillo”. In the naming task, they simply had to say the word out loud as soon as it appeared on the screen, and in the lexical decision task, they had to press a button to indicate that the letter string that they were presented with did indeed form a word.

    The eye-tracking task examined how quickly the words were recognised when presented in context. For example, participants would read sentences containing high and low frequency words such as “Amy told the teacher that her dog ate her homework assignment”, and “It is not unusual to see an armadillo cross a road in Texas”. The experimenter would then calculate how long the reader spent looking at the target words such as “teacher” and “armadillo”.

  • As you can see from this table, low frequency words took longer to recognise than high frequency words in all three tasks. That is, people took longer to say the word out loud when it was low frequency than when it was high frequency, as evidenced in a longer naming latency. The lexical decision latency, that is, the time taken to press the button to indicate that the stimulus was in fact a word, was also longer for low frequency words. Finally, the amount of time spent looking at the word in the eye-tracking task, here labelled as ‘first-fixation latency’ was also longer for low frequency words. So, there is quite strong evidence that people find words more difficult to recognise when they encounter them less often.

  • Likewise, several studies have shown that the context in which a word appears can influence how easy it is to recognise, depending on how predictable the context makes it.

    This influence is nicely demonstrated by a classic study that was carried out by Tulving and Gold. In this study, participants read an incomplete sentence, such as “The skiers were buried alive by the sudden …”. And they then had to recognise a single word. The word might be something like ‘avalanche’, which would be predictable following this context, or ‘inflation’ which would be unrelated.

  • The question they addressed in their study was whether increasing the amount of semantically related information that is available before the target word is presented will affect the minimum amount of time needed to identify the word.

    The amount of context that they used was varied.

    Following presentation of the context, a target word was presented at varying exposure durations, starting too brief for recognition, and increasing until the word is recognised.

    They measured the stimulus exposure necessary for recognition with relevant context and with misleading context.

  • In addition, increasing the amount of misleading context, increased the time it took for participants to recognise the target word.

    So, presenting a word after a relevant context that makes it more predictable makes word recognition easier and presenting it after a misleading context that makes it less predictable makes word recognition more difficult.

    We can see from this that predictability has a large effect on word recognition processes

Pages 20-21: Neighbourhood Effects

  • Orthographic and phonological neighbourhoods influence word recognition speed.

  • As well as the frequency with which we encounter a word, and how predictable it is from the context, word recognition can also be influenced by the number of other words in existence that are similar to the word that we are trying to recognise. This is known as the neighbourhood effect. Neighbourhood effects can be caused when there are other words that either look similar or sound similar to the word we are trying to recognise. The first type of neighbourhood effect that I am going to talk about is orthographic neighbourhood effects. The word ‘orthography’ simply refers to information about the spellings of words, or put another way, how the letters are arranged in a word. A word’s orthographic neighbourhood can be defined as the number of words that can be formed by changing one letter of a word. So, for example, two orthographic neighbours of the word ‘tank’ are ‘task’ in which the ‘n’ is changed to an ‘s’, and ‘rank’ in which the ‘t’ is changed to an ‘r’. The term ‘orthography’ simply refers to the way that a word is spelled.

    Some studies have shown that low frequency words are recognised more quickly if there are lots of other words in the language that are spelled similarly, that is, if they have lots of orthographic neighbours. So, a word’s orthographic neighbourhood can affect how it is recognised, for low frequency words at least.

  • Another kind of neighbourhood effect relates to the phonological neighbourhood of a word. So, where orthography refers to how a word is spelled, that is, what it looks like, phonology refers to what a word sounds like. The phonological neighbourhood of a word is the number of words that can be formed by changing one phoneme, which is a unit of sound. So, for example, ‘gait’ is changed to ‘bait’ by changing the ‘g’ sound to ‘b’and ‘gait’ is changed to ‘get’ by changing the ‘ai’ sound to ‘e’.

    Some studies have showed that words with many phonological neighbours are more easily recognised. For example, Yates found this using a variety of tasks such as lexical decision and naming.

Pages 22-23: MCQ on Neighbourhood Effects and Theories

  • Sample MCQ focused on orthographic neighbours.

  • The first model that I will describe is the Logogen model, which was developed in the 60’s and 70’s. Although it was developed some time ago, it does a good job of explaining some of the effects I have discussed so far, as you will see.

    The model is based on the assumption that perceivers have a vast number of specialized recognition units, that each are able to recognize one specific word. The recognition units or ‘word detectors’ are called “logogens”, and these contain information about the sounds of the word, its syntactic and semantic characteristics, and information about word type. All logogens have a scale indicating the activation level of the logogen. The incoming signal is presented to all logogens, and all logogens which match the incoming information are raised in activation. With each following matching sound, the activation increases until a certain critical activation value is reached: the fire threshold. As soon as the activation level of the logogen exceeds the threshold, the logogen fires: at this moment, the word is recognized and all information about the word becomes available. Once a logogen fires, all activation levels of competing logogens immediately decrease to their rest level.

    The logogen is activated in either of two ways:

    – By sensory input.

    – By contextual information. You can see on the model diagram that the logogen system is linked to the cognitive system, which can provide it with information about context. This contextual information can ‘pre-activate’ relevant logogens.

Pages 24-28: Theories of Word Recognition

  • Logogen Model: Words have activation thresholds based on frequency and context.

  • The first model that I will describe is the Logogen model, which was developed in the 60’s and 70’s. Although it was developed some time ago, it does a good job of explaining some of the effects I have discussed so far, as you will see.

    The model is based on the assumption that perceivers have a vast number of specialized recognition units, that each are able to recognize one specific word. The recognition units or ‘word detectors’ are called “logogens”, and these contain information about the sounds of the word, its syntactic and semantic characteristics, and information about word type. All logogens have a scale indicating the activation level of the logogen. The incoming signal is presented to all logogens, and all logogens which match the incoming information are raised in activation. With each following matching sound, the activation increases until a certain critical activation value is reached: the fire threshold. As soon as the activation level of the logogen exceeds the threshold, the logogen fires: at this moment, the word is recognized and all information about the word becomes available. Once a logogen fires, all activation levels of competing logogens immediately decrease to their rest level.

    The logogen is activated in either of two ways:

    – By sensory input.

    – By contextual information. You can see on the model diagram that the logogen system is linked to the cognitive system, which can provide it with information about context. This contextual information can ‘pre-activate’ relevant logogens.

  • Different logogens have different thresholds, depending on factors such as word frequency.

    High frequency words have lower thresholds for firing, and therefore require less stimulus information before the word detector is activated.

    So, a high frequency word such as ‘cat’ will have a much lower threshold than a low frequency word such as ‘cot’ and will therefore be recognised more quickly.

    This is how this model explains the word frequency effects that we covered in the previous part of the lecture.

Page 29-39: Word Superiority Effect

  • Demonstrates improved letter identification within words compared to isolation.

  • Findings from this task typically suggest about a 10% improvement in performance when participants have seen the whole word compared to a single letter. This suggests that it is easier to identify a letter in the context of a word than in isolation. It is also easier to recognise a letter in the context of a word, than if the same letter has been presented as part of a non-word.

    So, the word superiority effect is defined by the finding that performance is better when the letter string forms a word than when it does not. The word superiority effect suggests that information about the word presented can facilitate identification of the letters in that word. The next model of word recognition that I am going to talk about was designed to explain how information from the whole word can improve performance in terms of recognising an individual letter.

Page 40-41: Interactive Activation Model

  • Considers both excitatory and inhibitory connections among letter and word detectors.

  • This model is called the Interactive Activation Model and was developed in the 1980’s by McClelland and Rumelhart.

    The Interactive Activation Model consists of three levels of detectors. The first level contains feature detectors, which can detect letter features, such as vertical or horizontal lines.

    This level then feeds into the letter detector level, in which letters that contain the activated features would also become activated. Finally, the output of the letter detectors would then activate word level detectors which contained these letters. For example, if the initial slot contained the letter ‘w’, this would go on to activate words beginning with ‘w’ such as ‘word’ and ‘work’.

  • All these levels are interconnected and can influence each other. These connections are excitatory for consistent connections, and inhibitory for inconsistent connections. So, if a W is activated in the first letter position, this will feed up the network and activate words beginning with W. If the first letter is a W then the word cannot be ‘fork’, so this word will be inhibited. Inhibition can also occur within levels, in that as the word ‘work’ becomes more activated, it will inhibit ‘word’ and ‘fork’. Activation from the word level will then feed back down to the letter level, where consistent letters like ‘w’ ‘o’ ‘r’ ‘k’ will receive more activation, and inconsistent letters will be inhibited.

    It is this feedback from the word level to the letter level that can explain the word superiority effect. That is, if the participant recognises the word ‘work’ and they are then presented with a ‘k’ or a ‘d’, the ‘k’ will have received activation from with word ‘work’, but the ‘d’ will have been inhibited, making it easier for people to say it was a ‘k’ rather than a ‘d’.

    Although it has been very influential, one disadvantage of the interactive activation model is that it is what is known as a ‘slot coding system’. That is, each individual letter has it’s own slot which it has to fit into in order for the word to be recognised. So, if any of the letters are slightly out of place, for example, if someone was to spell a word incorrectly, you wouldn’t be able to recognise it.

Page 42-43: Transposed Letter Priming

  • Examples show varying recognition times based on letter positioning.

  • there is also some evidence from experimental studies that cannot easily be explained by the Interactive Activation Model.

    This evidence comes from experiments that have showed ‘transposed letter priming’.

    According to the Interactive Activation Model, words with two letters that are simply switched around should be just as hard to recognise as words that have two incorrect letters in them, as in both cases, the wrong letters feed into the wrong slots in the system.

    As a result, ‘j-u-g-d-e’ is no more similar to ‘judge’ and ‘j-u-p-t-e’, as all three options have identical letters in only three of the five letter positions.

    Therefore, if we were to prime a word like ‘judge’ with a word where two of the letters are switched around, this should provide the same level of priming as when two different letters are used. However, Perea and Lupkerdemonstrated that in fact, a word like ‘judge’ is easier to recognise when it is primed by a transposed letter prime than when a substitution prime is used.

    The Interactive Activation Model would have great difficulty explaining the presence of these transposed letter priming effects.


Pages 44-46: Dual-Route Model

  • Highlights routes for recognizing high- and low-frequency words; implications for dyslexia.

  • The final model that I am going to talk about was designed to explain how we read aloud, but it is quite interesting in other respects as well, as it can also be used to explain some of the difficulties experienced by people with dyslexia. This is the Dual Route Model which was proposed by Coltheart et al. (2001).

    According to this model there are two routes by which a word can be read out.

    One route is a “direct route”, which directly links the visually presented word to its entry in the lexicon. This route is typically employed for high frequency or familiar words, and allows the reader to directly access the word, and then pronounce it, since information on how it is pronounced is contained it its entry in the mental lexicon.

  • The second route is a more indirect route which is mainly used for reading low frequency words or non-words.

    By this route, the words are literally ‘spelled out’ using grapheme to phoneme conversion rules, before they can be pronounced. So, for example, if you read the word ‘man’, and you were unfamiliar with this word, it would be spelled out ‘m’ ‘a’ ‘n’ in order for you to know how to pronounce it.

Pages 47-49: Dyslexia Types

  • Phonological and Surface Dyslexia characterized by specific reading challenges.

  • §In the skilled reading system, these two routes develop independently as children learn to read.


    §Problems with each route leads to the formation of different patterns of reading disorder:

    §Developmental surface dyslexia.

    §Developmental phonological dyslexia.

  • The characteristic most often associated with phonological dyslexia is a difficulty with reading non-words or nonsense words.

    That is, non-words (i.e., ‘giph’, ‘polmex’) are significantly harder to read aloud than words.

    The only way to read a novel letter string is to implement some process of decoding.

    Phonological dyslexia assumes a selective deficit in developing the phonological route.

    Applying grapheme-to-phoneme conversion rules has not been mastered or is impaired.

Page 50: Sample MCQ on The Interactive Activation Model

  • Covers the model's applications and significance.

Page 51: Summary of Key Points

  • Methods: Eye-tracking, Naming, Lexical Decision.

  • Factors: Word frequency, Predictability, Neighbourhood.

  • Models: Logogen, Interactive Activation, Dual-route Cascaded.