Lexicography

The main goals of lexicography include the systematic collection, analysis, and presentation of words and their meanings in a comprehensive dictionary format.

Some words have double meanings or their meanings are tricky, which can lead to confusion for users. Therefore, lexicographers must provide clear definitions and usage examples to ensure that readers understand the context in which each meaning applies.

Google translator is good at translating word-for-word basis, while DeepL and ChatGPT are good for translating whole sentences.

Lexeme

  • Definition: An abstract unit that represents a concept in a language's vocabulary system.

Word-form

  • Definition: A specific form of a lexeme (e.g., singular/plural, tense).

Lexeme vs Word-form

  • Example 1:

    • go, went, gone, goes

      • Lexemes: 1

      • Word-forms: 4

  • Example 2:

    • dog, dog’s, dogs, dogs’

      • Lexemes: 1

      • Word-forms: 4

Lemma

  • Definition: The form used to represent a lexeme (typically the dictionary form).

    • Examples:

      • go, went, gone, goes: One lexeme, lemma: go.

      • dog, dogs, dog’s, dogs’: One lexeme, lemma: dog.


Classification of Words: Structure

  1. Monomorphemic: A word with only one morpheme (e.g., book, shop).

  2. Polymorphemic: A word made up of more than one morpheme (e.g., bookshop, he likes skiing).


Classification of Words: Meaning

  • Monosemous vs. Polysemous:

    • Monosemous: A word with one meaning (e.g., dog).

    • Polysemous: A word with multiple meanings (e.g., bank—a financial institution or the side of a river).

  • Polysemous vs. Homonymous:

    • Polysemy: One word with multiple related meanings.

    • Homonymy: One word with completely different meanings (e.g., bat—flying animal vs. sports equipment).

  • Variants:

    • Homographs: Words that are spelled the same but have different meanings (e.g., lead as a metal and lead as to guide).

    • Homophones: Words that sound the same but have different meanings (e.g., pair vs. pear).


Basic Sense Relations Between Words

  1. Synonymy: Words with similar meanings (e.g., happy and joyful).

  2. Antonymy: Words with opposite meanings (e.g., hot and cold).

  3. Hyponymy: A relationship where one word is a more specific instance of a broader category.

    • Hypernym: The broader category (e.g., animal).

    • Hyponym: A specific instance (e.g., dog).

  4. Metaphor: Using a word to mean something beyond its literal meaning (e.g., time is money).

  5. Metonymy: Using a word associated with something to represent it (e.g., the White House issued a statement).

  6. Prototypes: The most typical example of a concept (e.g., robin as a typical bird).

This format simplifies and organizes the information for clarity. Let me know if you need further adjustments!

Types of corpus studies:

  • Using ready-made corpora (usually available on the Internet, either freely or by subscription fees)

  • Building one’s own corpus (particularly a specialist one)

  • Corpus-based study: a theory (hypothesis), corpus data is used to test, validate or refute it

  • Corpus-driven study: no theory beforehand; the corpus itself should be the source of our hypotheses about the language

Types of corpora:

  • Number of languages:

    • Monolingual corpora: Focuses on a single language, providing insights into its usage and structure.

    • Bilingual corpora: Contains texts in two languages, useful for translation studies and comparative analysis.

      • Comparable corpus: A comparable corpus includes texts from multiple languages that are similar in content and genre, allowing for cross-linguistic studies and comparative linguistic research.

      • Parallel corpus: A type of bilingual corpus where texts in both languages are aligned at the sentence or paragraph level, facilitating direct comparison and aiding in translation.

  • Mode of communication:

    • Written

    • Spoken (spontaneous, semi-spontaneous, written-to-be-spoken)

    • Mixed

  • Updates?

    • Monitor (open-ended, dynamic) corpus

    • Snapshot (finite-sized, static) corpus

  • Annotation

    • Annotated (marked-up) corpus: An annotated (marked-up) corpus is a collection of texts with extra information added. This can include tags that identify grammar, meaning, or language features, helping researchers study language more clearly.

    • Unannotated (raw, plain) corpus

  • Choice of texts

    • General corpus (subtype: national corpus) (A general corpus is a comprehensive collection of texts from various sources within a specific language, often representing the linguistic features of a nation or region. This type of corpus is used to analyze language use, frequency of terms, and syntactic structures across different contexts.)

    • Specialized (special-purpose) corpus, often built on one’s own (do-it-yourself corpus)

  • Time perspective

    • Synchronic corpus (A synchronic corpus captures language at a specific point in time, enabling researchers to study the current state of language use without considering historical changes.)

    • Diachronic corpus (A diachronic corpus, in contrast, allows researchers to examine language changes over time by analyzing data from different historical periods and comparing the evolution of language use.)

Key features of linguistic corpora

  • Machine-readability - the ability of a corpus to be processed by computers, facilitating automated analysis and retrieval of linguistic data.

  • Sampling (text samples)

  • Balance (NOT a random collection of texts)

  • Representativeness - the degree to which a sample reflects the characteristics of the larger population, ensuring that the linguistic data is comprehensive and relevant.

Points to consider before corpus use

  • Corpus structure

  • Time frame

  • Corpus size

Corpus structure (COCA)

  • Written part (~85%), including extracts from webpages, blogs, academic journals, newspapers, magazines, literature (fiction), TV/movies subtitles

  • Spoken part (~15%): unscripted conversations from TV/radio programs

Corpus structure (BNC)

  • Written part (~90%), including extracts from regional and national newspapers, specialist periodicals and journals, academic books and popular fiction, published and unpublished letters and memoranda, school and university essays

  • Spoken part (~10%), including orthographic transcriptions of unscripted informal conversations and spoken language collected in different contexts, ranging from formal business or government meetings to radio shows and phone-ins

Corpus study: basic terminology

  • Query - A query is what you search for in a corpus (a collection of texts). For example, you might search for a word, phrase, or pattern like "happiness" or "isn't it."

    • A single word: go

    • An idiomatic expression: kettle of fish

    • A foreign expression: en route

    • An abbreviation: NASA

    • A place or company name: World Health Organization

    • Any other item: women tend to

  • Number of occurrences - This is how many times your query (word or phrase) appears in the corpus. For example, if "happiness" appears 50 times in the text collection, that’s the number of occurrences.

  • Concordances - Concordances show where and how your query appears in sentences. It’s like a snippet of text showing the word or phrase in context.

  • Corpus order/ Random order (find sample) - Corpus order means the results are shown in the same order as they appear in the text collection.

  • KWIC (Key Word in Context) - This is a tool that shows your query (keyword) in the middle, with words before and after it for context.
    Example for the word "happiness":

    • "She found happiness in the little things."

    • "The search for happiness never ends."

  • Distribution (types of text, temporal aspects) (CHART) - Distribution looks at where and how often your query appears across different types of texts (e.g., novels, news, tweets) or over time (e.g., years or centuries).

    • Often shown in charts or graphs to make patterns easy to see.

  • Additional pieces of information about a text (header) - The header contains metadata about the text, such as:

    • Title

    • Author

    • Year of publication

    • Genre (e.g., fiction, academic)

    • Word count

  • Extended context - Extended context means looking at a larger chunk of text around your query to understand its meaning better, such as reading the whole paragraph or page where it appears.

Two types of a corpus-based study

  • Qualitative study:

    • an example

      • The word red

      • Senses related to: the colour (more than one sense), money, etc.

    • Also in idiomatic meanings: red herring, red carpet, red eyes, red tape, etc.

  • Quantitative study:

    • an example

      • The suffix -less: meaningless, legless, childless, husbandless, wifeless, dogless, catless

      • red + NOUN • tell vs say: which is more frequent in English?

      • Is the frequency of the construction global warming / climate change on the increase or decrease?

      • The forms with *ism (combination of both quantitative and qualitative study)

Qualitative Corpus Study

  • Focuses on meaning and context.

  • Examines how language is used in specific situations or texts.

  • Looks at smaller samples in detail to understand patterns, themes, or reasons behind language use.

  • Example: Analyzing how people express gratitude in emails.

Quantitative Corpus Study

  • Focuses on numbers and patterns.

  • Counts how often words, phrases, or structures appear.

  • Uses large amounts of data to find statistical trends.

  • Example: Measuring how often the word "thanks" appears in different types of texts (e.g., emails, speeches, novels).

In short:

  • Qualitative = In-depth, focused on meaning.

  • Quantitative = Broad, focused on numbers.


Subcorpora

A subcorpora is a smaller, specific section of a larger corpus. It’s like dividing a big collection of texts into smaller groups based on certain criteria.

Two types of frequency

  • Raw frequency: the number of occurrences of a given form

  • Normalized (relative) frequency: the number of occurrences in a normalized sample (usually 1,000,000 words)

In corpora, there are “wild card”, so

  • ? > for a single character

  • * > for zero or more characters

Lexicographers are people who compile, edit, and publish dictionaries, utilizing various corpus studies to enhance the accuracy and comprehensiveness of their entries.

Corpus annotation

  • Extralinguistic (textual) > information about the text (not the language)

  • Linguistic > information (for the computer rather than the user) about the language itself

Linguistic annotation

1. Lemmatization (Lemma Annotation)

  • Reduction of words to lemmas, e.g. went_GO

  • Done automatically (possible errors!)

  • In BNC/COCA > capitals, e.g. GO, LEARN, DOG, APPENDIX

  • Aquarium: how many plurals? Which is more common?

  • Aquarium (BNC): the plural in academic language vs. the language of newspapers. Any interpretation here?

2. POS Annotation (Part-of-Speech Annotation)

  • Labeling words with their grammatical category (e.g., noun, verb, adjective).
    Example: She [PRONOUN] runs [VERB] fast [ADVERB].

  • POS tagging:

    • must as a noun > must_nn

    • cause as a noun > cause_nn, a verb > , cause_vv

    • total as an adjective > total_j

    • MUST_nn • MUSIC_nn

3. Parsing

  • Analyzing the grammatical structure of sentences, showing relationships between words (like subject, object, etc.).
    Example: A tree diagram showing how "The cat sat on the mat" is structured.

4. Semantic Tagging

  • Adding tags to indicate the meaning or semantic roles of words.
    Example: dog → [ANIMAL], bank → [FINANCIAL INSTITUTION].

5. Discourse Tagging

  • Annotating larger structures, like how sentences connect or the roles of different parts in a conversation.
    Example: Tagging a sentence as an introduction or conclusion.

6. Speech Prosody Tagging

  • Marking features of speech like intonation, stress, rhythm, and pauses.
    Example: Indicating rising tone for a question: "You’re going? [RISING TONE]."

7. Problem-Oriented Tagging

  • Tagging specific aspects of language for a particular research purpose.
    Example: Annotating errors in learner language to study grammar mistakes.

Types and tokens

  • Type > a unique word form in the corpus (the number of types = the number of different forms)

  • Token > a single occurrence of a given form (the number of tokens = the number of all the forms taken into consideration)

    • Examples: APPENDIX, SIT, LEARN, DOG

Frequency vs. Significance in Collocations

1. Frequency

  • Refers to how often a word or phrase appears.

  • Example:

    • "book":

      • a book, the book

      • These are frequent but not necessarily significant in meaning corpora-speaking.

2. Significance / Relevance

  • Refers to how meaningful or characteristic a word’s collocations are.

  • Example:

    • "book":

      • prayer book, cookery book

      • These are statistically significant collocations, even if they are not frequent.

Equivalence

Equivalence refers to the relationship between a source-language (SL) expression and a target-language (TL) expression in terms of:

  • Meaning: Ensuring the same idea or concept is conveyed.

  • Usage: Maintaining appropriate context and cultural relevance.

Problems in Bilingual Lexicography

  1. Overemphasis on Similarities

    • Similarities between languages often dominate the description, while differences are neglected.

  2. Focus on Source Language (SL)

    • The content of the source language receives more attention, while the expressional aspects of the target language (TL) are often overlooked.

  3. Culture-Specific Terms

    • Difficulty in accurately translating terms that are deeply rooted in one culture and lack direct equivalents in another.

  4. Language-Specific Phenomena

    • Challenges arise with:

      • Connotation (implied meanings).

      • Metaphor and metonymy (figurative language).

      • Pragmatics (contextual use and meaning).

Equivalence types

  • Semantic Equivalence

    • Focuses on meaning: ensuring the source-language term and target-language term convey the same concept.

    • Example:

      • right (adj.): prawidłowy, właściwy, odpowiedni, słuszny, prawy, etc.

  • Pragmatic Equivalence

    • Focuses on usage: ensuring the expression is appropriate in context or function.

    • Example:

      • right (excl., used to show understanding or agreement): dobrze, no dobra, fakt, faktycznie, OK, etc.

We have 3 degrees of Equivalence, full, partial and zero.

Definitions

Types of Meaning

  • Denotative (Conceptual) Meaning: The literal or primary meaning of a word.

  • Connotative (Associative) Meaning: The additional meanings or associations a word may have.

  • Emotive Meaning: The emotional response a word evokes.

  • Collocational Meaning: The meaning that arises from the common combinations of words (collocations).


Methods of Explaining Meaning

  • Paraphrase: Rewriting the meaning in different words.

  • True Definition: A precise definition of the concept.

  • Functional Explanation: Describes the function or use of the word.


True Definitions

  • Definiendum: The concept or headword being defined.

  • Definiens: The definition itself.


Types of True Definitions

  • Intension: The content of a concept (distinctive features).

    • Example: Intensional Definition

      • Genus Proximum (hypernym): The general category.

      • Differentia Specifica: The distinguishing features.

  • Extension: The range or scope of the concept (specific examples or instances).

    • Example: Extensional Definition

      • Fruit: apple, pear, plum, strawberry, etc.

Examples of Extensional Definitions

  • Furniture: Chairs, tables, beds, cupboards, etc.

  • Benelux: Belgium, the Netherlands, and Luxembourg.

  • Synoptic Gospels: The Gospels of Matthew, Mark, and Luke.

Intensional + Extensional Combination

  • Furniture: The items like chairs, tables, beds, etc., that you use in a room or house for living.


Problems with Intensional Definitions

  • Errors in Definitions:

    • Example: Telenowela is incorrectly defined as a short film with a simple plot and few characters, when it refers to a type of soap opera.

    • Example: Akwitański incorrectly refers to a French city, when it should refer to the Aquitaine region.