Lexicography
The main goals of lexicography include the systematic collection, analysis, and presentation of words and their meanings in a comprehensive dictionary format.
Some words have double meanings or their meanings are tricky, which can lead to confusion for users. Therefore, lexicographers must provide clear definitions and usage examples to ensure that readers understand the context in which each meaning applies.
Google translator is good at translating word-for-word basis, while DeepL and ChatGPT are good for translating whole sentences.
Lexeme
Definition: An abstract unit that represents a concept in a language's vocabulary system.
Word-form
Definition: A specific form of a lexeme (e.g., singular/plural, tense).
Lexeme vs Word-form
Example 1:
go, went, gone, goes
Lexemes: 1
Word-forms: 4
Example 2:
dog, dog’s, dogs, dogs’
Lexemes: 1
Word-forms: 4
Lemma
Definition: The form used to represent a lexeme (typically the dictionary form).
Examples:
go, went, gone, goes: One lexeme, lemma: go.
dog, dogs, dog’s, dogs’: One lexeme, lemma: dog.
Classification of Words: Structure
Monomorphemic: A word with only one morpheme (e.g., book, shop).
Polymorphemic: A word made up of more than one morpheme (e.g., bookshop, he likes skiing).
Classification of Words: Meaning
Monosemous vs. Polysemous:
Monosemous: A word with one meaning (e.g., dog).
Polysemous: A word with multiple meanings (e.g., bank—a financial institution or the side of a river).
Polysemous vs. Homonymous:
Polysemy: One word with multiple related meanings.
Homonymy: One word with completely different meanings (e.g., bat—flying animal vs. sports equipment).
Variants:
Homographs: Words that are spelled the same but have different meanings (e.g., lead as a metal and lead as to guide).
Homophones: Words that sound the same but have different meanings (e.g., pair vs. pear).
Basic Sense Relations Between Words
Synonymy: Words with similar meanings (e.g., happy and joyful).
Antonymy: Words with opposite meanings (e.g., hot and cold).
Hyponymy: A relationship where one word is a more specific instance of a broader category.
Hypernym: The broader category (e.g., animal).
Hyponym: A specific instance (e.g., dog).
Metaphor: Using a word to mean something beyond its literal meaning (e.g., time is money).
Metonymy: Using a word associated with something to represent it (e.g., the White House issued a statement).
Prototypes: The most typical example of a concept (e.g., robin as a typical bird).
This format simplifies and organizes the information for clarity. Let me know if you need further adjustments!
Types of corpus studies:
Using ready-made corpora (usually available on the Internet, either freely or by subscription fees)
Building one’s own corpus (particularly a specialist one)
Corpus-based study: a theory (hypothesis), corpus data is used to test, validate or refute it
Corpus-driven study: no theory beforehand; the corpus itself should be the source of our hypotheses about the language
Types of corpora:
Number of languages:
Monolingual corpora: Focuses on a single language, providing insights into its usage and structure.
Bilingual corpora: Contains texts in two languages, useful for translation studies and comparative analysis.
Comparable corpus: A comparable corpus includes texts from multiple languages that are similar in content and genre, allowing for cross-linguistic studies and comparative linguistic research.
Parallel corpus: A type of bilingual corpus where texts in both languages are aligned at the sentence or paragraph level, facilitating direct comparison and aiding in translation.
Mode of communication:
Written
Spoken (spontaneous, semi-spontaneous, written-to-be-spoken)
Mixed
Updates?
Monitor (open-ended, dynamic) corpus
Snapshot (finite-sized, static) corpus
Annotation
Annotated (marked-up) corpus: An annotated (marked-up) corpus is a collection of texts with extra information added. This can include tags that identify grammar, meaning, or language features, helping researchers study language more clearly.
Unannotated (raw, plain) corpus
Choice of texts
General corpus (subtype: national corpus) (A general corpus is a comprehensive collection of texts from various sources within a specific language, often representing the linguistic features of a nation or region. This type of corpus is used to analyze language use, frequency of terms, and syntactic structures across different contexts.)
Specialized (special-purpose) corpus, often built on one’s own (do-it-yourself corpus)
Time perspective
Synchronic corpus (A synchronic corpus captures language at a specific point in time, enabling researchers to study the current state of language use without considering historical changes.)
Diachronic corpus (A diachronic corpus, in contrast, allows researchers to examine language changes over time by analyzing data from different historical periods and comparing the evolution of language use.)
Key features of linguistic corpora
Machine-readability - the ability of a corpus to be processed by computers, facilitating automated analysis and retrieval of linguistic data.
Sampling (text samples)
Balance (NOT a random collection of texts)
Representativeness - the degree to which a sample reflects the characteristics of the larger population, ensuring that the linguistic data is comprehensive and relevant.
Points to consider before corpus use
Corpus structure
Time frame
Corpus size
Corpus structure (COCA)
Written part (~85%), including extracts from webpages, blogs, academic journals, newspapers, magazines, literature (fiction), TV/movies subtitles
Spoken part (~15%): unscripted conversations from TV/radio programs
Corpus structure (BNC)
Written part (~90%), including extracts from regional and national newspapers, specialist periodicals and journals, academic books and popular fiction, published and unpublished letters and memoranda, school and university essays
Spoken part (~10%), including orthographic transcriptions of unscripted informal conversations and spoken language collected in different contexts, ranging from formal business or government meetings to radio shows and phone-ins
Corpus study: basic terminology
Query - A query is what you search for in a corpus (a collection of texts). For example, you might search for a word, phrase, or pattern like "happiness" or "isn't it."
A single word: go
An idiomatic expression: kettle of fish
A foreign expression: en route
An abbreviation: NASA
A place or company name: World Health Organization
Any other item: women tend to
Number of occurrences - This is how many times your query (word or phrase) appears in the corpus. For example, if "happiness" appears 50 times in the text collection, that’s the number of occurrences.
Concordances - Concordances show where and how your query appears in sentences. It’s like a snippet of text showing the word or phrase in context.
Corpus order/ Random order (find sample) - Corpus order means the results are shown in the same order as they appear in the text collection.
KWIC (Key Word in Context) - This is a tool that shows your query (keyword) in the middle, with words before and after it for context.
Example for the word "happiness":"She found happiness in the little things."
"The search for happiness never ends."
Distribution (types of text, temporal aspects) (CHART) - Distribution looks at where and how often your query appears across different types of texts (e.g., novels, news, tweets) or over time (e.g., years or centuries).
Often shown in charts or graphs to make patterns easy to see.
Additional pieces of information about a text (header) - The header contains metadata about the text, such as:
Title
Author
Year of publication
Genre (e.g., fiction, academic)
Word count
Extended context - Extended context means looking at a larger chunk of text around your query to understand its meaning better, such as reading the whole paragraph or page where it appears.
Two types of a corpus-based study
Qualitative study:
an example
The word red
Senses related to: the colour (more than one sense), money, etc.
Also in idiomatic meanings: red herring, red carpet, red eyes, red tape, etc.
Quantitative study:
an example
The suffix -less: meaningless, legless, childless, husbandless, wifeless, dogless, catless
red + NOUN • tell vs say: which is more frequent in English?
Is the frequency of the construction global warming / climate change on the increase or decrease?
The forms with *ism (combination of both quantitative and qualitative study)
Qualitative Corpus Study
Focuses on meaning and context.
Examines how language is used in specific situations or texts.
Looks at smaller samples in detail to understand patterns, themes, or reasons behind language use.
Example: Analyzing how people express gratitude in emails.
Quantitative Corpus Study
Focuses on numbers and patterns.
Counts how often words, phrases, or structures appear.
Uses large amounts of data to find statistical trends.
Example: Measuring how often the word "thanks" appears in different types of texts (e.g., emails, speeches, novels).
In short:
Qualitative = In-depth, focused on meaning.
Quantitative = Broad, focused on numbers.
Subcorpora
A subcorpora is a smaller, specific section of a larger corpus. It’s like dividing a big collection of texts into smaller groups based on certain criteria.
Two types of frequency
Raw frequency: the number of occurrences of a given form
Normalized (relative) frequency: the number of occurrences in a normalized sample (usually 1,000,000 words)
In corpora, there are “wild card”, so
? > for a single character
* > for zero or more characters
Lexicographers are people who compile, edit, and publish dictionaries, utilizing various corpus studies to enhance the accuracy and comprehensiveness of their entries.
Corpus annotation
Extralinguistic (textual) > information about the text (not the language)
Linguistic > information (for the computer rather than the user) about the language itself
Linguistic annotation
1. Lemmatization (Lemma Annotation)
Reduction of words to lemmas, e.g. went_GO
Done automatically (possible errors!)
In BNC/COCA > capitals, e.g. GO, LEARN, DOG, APPENDIX
Aquarium: how many plurals? Which is more common?
Aquarium (BNC): the plural in academic language vs. the language of newspapers. Any interpretation here?
2. POS Annotation (Part-of-Speech Annotation)
Labeling words with their grammatical category (e.g., noun, verb, adjective).
Example: She [PRONOUN] runs [VERB] fast [ADVERB].POS tagging:
must as a noun > must_nn
cause as a noun > cause_nn, a verb > , cause_vv
total as an adjective > total_j
MUST_nn • MUSIC_nn
3. Parsing
Analyzing the grammatical structure of sentences, showing relationships between words (like subject, object, etc.).
Example: A tree diagram showing how "The cat sat on the mat" is structured.
4. Semantic Tagging
Adding tags to indicate the meaning or semantic roles of words.
Example: dog → [ANIMAL], bank → [FINANCIAL INSTITUTION].
5. Discourse Tagging
Annotating larger structures, like how sentences connect or the roles of different parts in a conversation.
Example: Tagging a sentence as an introduction or conclusion.
6. Speech Prosody Tagging
Marking features of speech like intonation, stress, rhythm, and pauses.
Example: Indicating rising tone for a question: "You’re going? [RISING TONE]."
7. Problem-Oriented Tagging
Tagging specific aspects of language for a particular research purpose.
Example: Annotating errors in learner language to study grammar mistakes.
Types and tokens
Type > a unique word form in the corpus (the number of types = the number of different forms)
Token > a single occurrence of a given form (the number of tokens = the number of all the forms taken into consideration)
Examples: APPENDIX, SIT, LEARN, DOG
Frequency vs. Significance in Collocations
1. Frequency
Refers to how often a word or phrase appears.
Example:
"book":
a book, the book
These are frequent but not necessarily significant in meaning corpora-speaking.
2. Significance / Relevance
Refers to how meaningful or characteristic a word’s collocations are.
Example:
"book":
prayer book, cookery book
These are statistically significant collocations, even if they are not frequent.
Equivalence
Equivalence refers to the relationship between a source-language (SL) expression and a target-language (TL) expression in terms of:
Meaning: Ensuring the same idea or concept is conveyed.
Usage: Maintaining appropriate context and cultural relevance.
Problems in Bilingual Lexicography
Overemphasis on Similarities
Similarities between languages often dominate the description, while differences are neglected.
Focus on Source Language (SL)
The content of the source language receives more attention, while the expressional aspects of the target language (TL) are often overlooked.
Culture-Specific Terms
Difficulty in accurately translating terms that are deeply rooted in one culture and lack direct equivalents in another.
Language-Specific Phenomena
Challenges arise with:
Connotation (implied meanings).
Metaphor and metonymy (figurative language).
Pragmatics (contextual use and meaning).
Equivalence types
Semantic Equivalence
Focuses on meaning: ensuring the source-language term and target-language term convey the same concept.
Example:
right (adj.): prawidłowy, właściwy, odpowiedni, słuszny, prawy, etc.
Pragmatic Equivalence
Focuses on usage: ensuring the expression is appropriate in context or function.
Example:
right (excl., used to show understanding or agreement): dobrze, no dobra, fakt, faktycznie, OK, etc.
We have 3 degrees of Equivalence, full, partial and zero.
Definitions
Types of Meaning
Denotative (Conceptual) Meaning: The literal or primary meaning of a word.
Connotative (Associative) Meaning: The additional meanings or associations a word may have.
Emotive Meaning: The emotional response a word evokes.
Collocational Meaning: The meaning that arises from the common combinations of words (collocations).
Methods of Explaining Meaning
Paraphrase: Rewriting the meaning in different words.
True Definition: A precise definition of the concept.
Functional Explanation: Describes the function or use of the word.
True Definitions
Definiendum: The concept or headword being defined.
Definiens: The definition itself.
Types of True Definitions
Intension: The content of a concept (distinctive features).
Example: Intensional Definition
Genus Proximum (hypernym): The general category.
Differentia Specifica: The distinguishing features.
Extension: The range or scope of the concept (specific examples or instances).
Example: Extensional Definition
Fruit: apple, pear, plum, strawberry, etc.
Examples of Extensional Definitions
Furniture: Chairs, tables, beds, cupboards, etc.
Benelux: Belgium, the Netherlands, and Luxembourg.
Synoptic Gospels: The Gospels of Matthew, Mark, and Luke.
Intensional + Extensional Combination
Furniture: The items like chairs, tables, beds, etc., that you use in a room or house for living.
Problems with Intensional Definitions
Errors in Definitions:
Example: Telenowela is incorrectly defined as a short film with a simple plot and few characters, when it refers to a type of soap opera.
Example: Akwitański incorrectly refers to a French city, when it should refer to the Aquitaine region.