Detailed Study Notes on Learner Corpora Chapter

Chapter Overview

  • Title: From design to collection of learner corpora

  • Authors: Gaëtanelle Gilquin, Sylviane Granger, Fanny Meunier

  • Publication Date: January 2015

  • DOI: 10.1017/CBO9781139649414.002

  • Citations: 12

  • Reads: 315


1. Introduction

  • Historical Context

    • Development of Second Language Acquisition (SLA) field noted by Gass et al. (1998) in the 1960s or 1970s.

    • Authentic data representing learners' interlanguage used in studies.

    • Limitations of prior SLA studies:

    • Small number of subjects.

    • Limited size of data collected.

    • Illustrative examples of previous studies by Ellis (2008) showing in-depth analysis but questionable generalizability.

  • Introduction of Learner Corpora

    • Expansion of corpus linguistics to interlanguage phenomena.

    • Definition of learner corpus:

    • “A collection of machine-readable authentic texts (including transcripts of spoken data) which is sampled to be representative of a particular language or language variety” (McEnery et al. 2006: 5).

    • Unique aspects:

    • Represents language produced by L2 learners.

    • Seeks to be representative of language variety unlike earlier studies.

    • Nesselhauf’s (2004) definition highlights:

    • “Systematic computerized collections of texts produced by language learners”.

    • Importance of external criteria such as learner level(s) and mother tongue(s).

    • Granger’s (2008) definition emphasizes the elements of naturalness and design criteria in learner corpora.


2. Core Issues

2.1. Learner Corpus Typology

  • Dimensions for Correlation

    • Medium: Written vs spoken texts.

    • Written learner corpora predominant since late 1980s.

    • Spoken corpora increasingly available but labor-intensive to collect.

    • Genre: Varying genres represented but often narrow.

    • Preference for argumentative essays in written corpora.

    • New developments in Language for Specific Purposes (LSP).

    • Target Language: Initially focused on English, but increasingly includes French, German, and Spanish learner corpora.

    • Mother Tongue: Mono-L1 and multi-L1 corpora.

    • Examples: Taiwanese Learner Corpus, International Corpus of Learner Finnish.

    • Chronology: Synchronic (cross-sectional) vs. diachronic (longitudinal) data.

    • Longitudinal databases such as LONGDALE tracking development across time.

    • Quasi-longitudinal data gathered from different proficiency learner groups.

    • Scale: Global vs. local learner corpora.

    • Global collected across diverse populations versus local focused on specific classroom groups.


2.2. Design: Environment, Task, and Learner Variables

  • Importance of Design Criteria

    • Essential due to interlanguage heterogeneity.

    • Considerations include learner's educational environment (foreign vs. second language), tasks performed, and learner characteristics (age, gender, proficiency level).

    • Task Variables:

    • Time constraints, availability of support material, exam conditions affect output.

    • Learner Variables:

    • Variability in language exposure, motivation, and prior knowledge significantly influences data collection.


2.3. Collection of Learner Corpora

  • Recruitment of Participants

    • Typically from students with whom researchers are in contact.

    • Volunteer bias may affect representativeness of data.

    • Written Corpora Collection:

    • Methods include scanning or typing textual data.

    • Necessity to maintain accuracy in capturing learners' written output.

    • Spoken Corpora Collection:

    • High-quality recordings ideally in controlled environments.

    • Challenges include transcription accuracy and overlapping speech.


3. Representative Studies

3.1. CEDEL2 Corpus

  • Compilation of L2 Spanish compositions from English-speaking learners.

  • Highlights design aspects such as content selection, representativeness, and careful consideration of metadata for consistency in data.

3.2. LINDSEI Corpus

  • Polish component challenges in spoken learner data collection and transcription.

  • Illustrates recruitment challenges and variability in learners’ proficiency.

3.3. LeaP Corpus

  • A multilingual corpus examining learner German and English, focusing on prosody.

  • Details multi-level transcription and metadata collection processes, contributing valuable insights into pronunciation in language acquisition.


4. Critical Assessment and Future Directions

  • Identifying Limitations:

    • Recognition of the imbalance in types of available learner corpora.

    • Need for diverse representations including more beginner and young learners.

  • Calls for Improvement:

    • Standardization of metrics and metadata.

    • Focus on objective proficiency measures and learner’s exposure.

  • Future Directions:

    • Possibility for multidimensional learner corpora reflecting diverse learner experiences.

    • Technological advancements facilitating naturalistic language data collection through CMC, gaming, and mobile platforms.


Key Readings

  • Granger, S. (1998): Introduction of the learner corpus concept.

  • Atkins et al. (1992): Guidelines on corpus design criteria.

  • Exploration of various learner corpora, their functionalities, and future implications in SLA and language teaching contexts.


References

  • A comprehensive list of literature covering corpus linguistics, SLA, learner corpora, and data collection methodology, critical for any further research or understanding of the field.