Detailed Study Notes on Learner Corpora Chapter
Chapter Overview
Title: From design to collection of learner corpora
Authors: Gaëtanelle Gilquin, Sylviane Granger, Fanny Meunier
Publication Date: January 2015
DOI: 10.1017/CBO9781139649414.002
Citations: 12
Reads: 315
1. Introduction
Historical Context
Development of Second Language Acquisition (SLA) field noted by Gass et al. (1998) in the 1960s or 1970s.
Authentic data representing learners' interlanguage used in studies.
Limitations of prior SLA studies:
Small number of subjects.
Limited size of data collected.
Illustrative examples of previous studies by Ellis (2008) showing in-depth analysis but questionable generalizability.
Introduction of Learner Corpora
Expansion of corpus linguistics to interlanguage phenomena.
Definition of learner corpus:
“A collection of machine-readable authentic texts (including transcripts of spoken data) which is sampled to be representative of a particular language or language variety” (McEnery et al. 2006: 5).
Unique aspects:
Represents language produced by L2 learners.
Seeks to be representative of language variety unlike earlier studies.
Nesselhauf’s (2004) definition highlights:
“Systematic computerized collections of texts produced by language learners”.
Importance of external criteria such as learner level(s) and mother tongue(s).
Granger’s (2008) definition emphasizes the elements of naturalness and design criteria in learner corpora.
2. Core Issues
2.1. Learner Corpus Typology
Dimensions for Correlation
Medium: Written vs spoken texts.
Written learner corpora predominant since late 1980s.
Spoken corpora increasingly available but labor-intensive to collect.
Genre: Varying genres represented but often narrow.
Preference for argumentative essays in written corpora.
New developments in Language for Specific Purposes (LSP).
Target Language: Initially focused on English, but increasingly includes French, German, and Spanish learner corpora.
Mother Tongue: Mono-L1 and multi-L1 corpora.
Examples: Taiwanese Learner Corpus, International Corpus of Learner Finnish.
Chronology: Synchronic (cross-sectional) vs. diachronic (longitudinal) data.
Longitudinal databases such as LONGDALE tracking development across time.
Quasi-longitudinal data gathered from different proficiency learner groups.
Scale: Global vs. local learner corpora.
Global collected across diverse populations versus local focused on specific classroom groups.
2.2. Design: Environment, Task, and Learner Variables
Importance of Design Criteria
Essential due to interlanguage heterogeneity.
Considerations include learner's educational environment (foreign vs. second language), tasks performed, and learner characteristics (age, gender, proficiency level).
Task Variables:
Time constraints, availability of support material, exam conditions affect output.
Learner Variables:
Variability in language exposure, motivation, and prior knowledge significantly influences data collection.
2.3. Collection of Learner Corpora
Recruitment of Participants
Typically from students with whom researchers are in contact.
Volunteer bias may affect representativeness of data.
Written Corpora Collection:
Methods include scanning or typing textual data.
Necessity to maintain accuracy in capturing learners' written output.
Spoken Corpora Collection:
High-quality recordings ideally in controlled environments.
Challenges include transcription accuracy and overlapping speech.
3. Representative Studies
3.1. CEDEL2 Corpus
Compilation of L2 Spanish compositions from English-speaking learners.
Highlights design aspects such as content selection, representativeness, and careful consideration of metadata for consistency in data.
3.2. LINDSEI Corpus
Polish component challenges in spoken learner data collection and transcription.
Illustrates recruitment challenges and variability in learners’ proficiency.
3.3. LeaP Corpus
A multilingual corpus examining learner German and English, focusing on prosody.
Details multi-level transcription and metadata collection processes, contributing valuable insights into pronunciation in language acquisition.
4. Critical Assessment and Future Directions
Identifying Limitations:
Recognition of the imbalance in types of available learner corpora.
Need for diverse representations including more beginner and young learners.
Calls for Improvement:
Standardization of metrics and metadata.
Focus on objective proficiency measures and learner’s exposure.
Future Directions:
Possibility for multidimensional learner corpora reflecting diverse learner experiences.
Technological advancements facilitating naturalistic language data collection through CMC, gaming, and mobile platforms.
Key Readings
Granger, S. (1998): Introduction of the learner corpus concept.
Atkins et al. (1992): Guidelines on corpus design criteria.
Exploration of various learner corpora, their functionalities, and future implications in SLA and language teaching contexts.
References
A comprehensive list of literature covering corpus linguistics, SLA, learner corpora, and data collection methodology, critical for any further research or understanding of the field.