Text Mining & NLP Lecture Series Vocabulary

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/102

flashcard set

Earn XP

Description and Tags

Key vocabulary terms spanning text preprocessing, feature engineering, modelling techniques, evaluation metrics, modern transformer architectures, and generative AI concepts covered throughout the lecture series.

Last updated 1:15 AM on 7/7/25
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

103 Terms

1
New cards

Text Mining

AI-driven process of converting unstructured text into structured, analysable data.

2
New cards

Natural Language Processing (NLP)

Field of computer science and AI focused on interactions between computers and human language.

3
New cards

NLP Pipeline

Sequence of steps—Corpus creation, Train/Validation/Test split, Feature Engineering, Task execution—used to build text models.

4
New cards

Corpus

A collection of text documents organised as a dataset for NLP tasks.

5
New cards

Corpora

Plural of corpus: multiple text collections.

6
New cards

Train-Validation-Test Split

Standard method to divide a corpus (e.g., 80 / 10 / 10 %) for model building and evaluation.

7
New cards

K-Fold Cross-Validation

Technique that iteratively partitions data into k subsets to improve reliability of evaluation.

8
New cards

Feature Engineering (Text)

Transforming raw text into numerical representations usable by ML algorithms.

9
New cards

Bag-of-Words (BoW)

Representation where each vocabulary word is a feature; documents become sparse count vectors.

10
New cards

Curse of Dimensionality

Problems arising from very large feature spaces, causing sparse data and overfitting.

11
New cards

Tokenization

Process of splitting text into tokens (words, phrases, symbols).

12
New cards

Compound Token

Multiple-word expression treated as one token, e.g., “United Kingdom.”

13
New cards

Stop Words

Common words (the, and, of, …) often removed because they add little semantic value.

14
New cards

Lowercasing

Normalization step converting all text to lower-case to shrink vocabulary size.

15
New cards

Regular Expression (Regex)

Pattern-matching syntax used to normalize or anonymise text (dates, names, numbers).

16
New cards

Stemming

Heuristic removal of word endings to obtain a crude root (e.g., changing → chang).

17
New cards

Lemmatization

Mapping inflected words to their dictionary root (e.g., changes → change).

18
New cards

Part-of-Speech (POS) Filtering

Keeping only selected grammatical categories (nouns, verbs, adjectives, adverbs) for analysis.

19
New cards

Preprocessing Pipeline

Typical order: normalization → lowercasing → tokenization → stop-word removal → stemming/lemmatization → POS filtering.

20
New cards

Variability (in NL)

Same meaning can be expressed by different sentences (paraphrases).

21
New cards

Ambiguity (in NL)

Single sentence may have multiple interpretations; resolved using context.

22
New cards

Generalization Challenge

Model trained in one domain faces Out-of-Domain or Out-of-Vocabulary inputs when deployed elsewhere.

23
New cards

Term Frequency (TF)

Count of a term’s occurrences within a document.

24
New cards

Inverse Document Frequency (IDF)

log(|D|/n_t); down-weights terms that appear in many documents.

25
New cards

TF-IDF

Weighting scheme TF×IDF elevating informative, rare terms and reducing common ones.

26
New cards

N-Gram

Contiguous sequence of n items from text (unigram, bigram, trigram, …).

27
New cards

Data Sparseness

Increase in zero entries when using large N-grams or large vocabularies.

28
New cards

Keyword Extraction

Task of identifying salient terms in text.

29
New cards

Text Similarity

Measuring how alike two pieces of text are using distance metrics.

30
New cards

Document Classification

Assigning predefined labels to documents.

31
New cards

Sentiment Analysis

Determining the emotional polarity (positive/negative/neutral) of text.

32
New cards

Topic Modelling

Unsupervised discovery of hidden thematic structures in a corpus.

33
New cards

Information Retrieval (IR)

Process of obtaining relevant documents in response to a query.

34
New cards

K-Nearest Neighbour (KNN)

Baseline classifier: label is taken from closest training documents in feature space.

35
New cards

Euclidean Distance

Straight-line metric between two vectors; suited for dense data.

36
New cards

Manhattan Distance

Sum of absolute coordinate differences between vectors.

37
New cards

Cosine Similarity

Normalized dot product measuring angle between vectors; robust for sparse text data.

38
New cards

Jaccard Coefficient

|Intersection| / |Union| similarity of two sets.

39
New cards

Dice Similarity

2×|Intersection| / (|A|+|B|); less sensitive to length differences than Jaccard.

40
New cards

Levenshtein (Edit) Distance

Minimum insertions, deletions, substitutions to transform one string into another.

41
New cards

Soundex

Phonetic hashing algorithm indexing names by pronunciation.

42
New cards

Word Representation

Vector encoding that captures semantic information of a word.

43
New cards

Term-Term Matrix

|V|×|V| co-occurrence counts of words in a corpus.

44
New cards

Point-wise Mutual Information (PMI)

log [p(w,c)/(p(w)p(c))]; measures word-context association strength.

45
New cards

Positive PMI (PPMI)

PMI values with negatives replaced by zero.

46
New cards

Skip-Gram (Word2Vec)

Neural model predicting context words from a target word to learn dense embeddings.

47
New cards

Embedding

Dense, low-dimensional vector representing token semantics.

48
New cards

Semantic Vector Arithmetic

Property where relationships emerge in embedding space (king – man + woman ≈ queen).

49
New cards

Recurrent Neural Network (RNN)

Neural architecture with cyclic connections allowing sequence processing.

50
New cards

Vanishing Gradient

Problem in RNN training where gradients shrink across long sequences, stalling learning.

51
New cards

Bidirectional RNN

RNN processing sequences in both forward and backward directions to exploit full context.

52
New cards

Long Short-Term Memory (LSTM)

RNN variant with gates (forget, input, output) and cell state to capture long dependencies.

53
New cards

Sequence-to-Sequence (Seq2Seq)

Encoder-decoder architecture mapping input sequences to output sequences (e.g., translation).

54
New cards

Attention Mechanism

Method enabling decoder to focus on relevant encoder states via weighted sums.

55
New cards

Self-Attention

Attention applied within the same sequence; token attends to all other tokens.

56
New cards

Transformer

Model composed of self-attention and feed-forward layers; allows parallel processing, no recurrence.

57
New cards

Encoder (Transformer)

Stack of self-attention blocks producing contextual embeddings.

58
New cards

Decoder (Transformer)

Stack that uses masked self-attention and encoder-decoder attention to generate output tokens.

59
New cards

Positional Embedding

Vector added to token embedding to encode order information for Transformers.

60
New cards

Multi-Head Attention

Parallel attention heads capturing different relation subspaces.

61
New cards

Masked Attention

Technique in decoder preventing access to future tokens during training/generation.

62
New cards

Autoregressive Generation

Producing text token-by-token, conditioning each prediction on previously generated tokens.

63
New cards

Large Language Model (LLM)

Very large Transformer (billions of parameters) capable of diverse language tasks.

64
New cards

Temperature (Sampling)

Parameter controlling randomness; higher values yield more diverse outputs.

65
New cards

top_p (Nucleus Sampling)

Sampling from smallest probability mass ≥ p to balance creativity and coherence.

66
New cards

top_k Sampling

Restricts candidate tokens to top k most probable before sampling.

67
New cards

Prompt Engineering

Crafting input instructions to steer LLM behaviour and outputs.

68
New cards

Zero-Shot Prompting

Asking an LLM to perform a task without providing examples.

69
New cards

Few-Shot Prompting

Including a few labelled examples in the prompt to guide the model.

70
New cards

Conversation Buffer Memory

Technique of appending full dialogue history to prompt to maintain context.

71
New cards

Conversation Summary Memory

Using an LLM to condense dialogue history into a shorter summary fed back to the model.

72
New cards

Retrieval-Augmented Generation (RAG)

Framework combining vector search with LLM prompting to generate grounded answers.

73
New cards

Vector Database

Specialised store of embeddings enabling efficient similarity search.

74
New cards

Semantic Search

Information retrieval based on embedding similarity rather than keyword overlap.

75
New cards

Chunking (RAG)

Splitting long documents into smaller parts before embedding and storage.

76
New cards

Quantization (GGUF)

Compression technique reducing parameter precision to run LLMs faster with less memory.

77
New cards

Temperature-Controlled Creativity

Concept where low temperature → deterministic, high temperature → creative outputs.

78
New cards

Agentic AI (LLM Agent)

System where an LLM selects and executes actions/tools autonomously to achieve goals.

79
New cards

ReAct Pattern

Agent loop: Thought → Action → Observation enabling reasoning and tool usage.

80
New cards

Evaluation Metric – BLEU

N-gram overlap score for machine translation output vs reference.

81
New cards

Evaluation Metric – ROUGE

Recall-oriented n-gram overlap metric for summarisation quality.

82
New cards

Evaluation Metric – BERTScore

Semantic similarity metric using BERT embeddings between candidate and reference texts.

83
New cards

Evaluation Metric – METEOR

Machine translation metric incorporating stemming, synonymy and alignment.

84
New cards

Quality Estimation (QE)

Predicting translation quality without reference by modelling sentence-level or word-level scores.

85
New cards

COMET

Neural evaluation metric for machine translation correlated with human judgement.

86
New cards

xCOMET

Multilingual extension providing state-of-the-art QE and evaluation performance.

87
New cards

Tower LLM

Unbabel’s multilingual LLM optimised for translation-related tasks.

88
New cards

Instruction Tuning

Fine-tuning an LLM on (instruction, response) pairs to follow natural language commands.

89
New cards

Masked Language Modelling (MLM)

Pretraining objective where model predicts intentionally masked tokens (used in BERT).

90
New cards

Fine-Tuning

Adapting a pretrained model to a specific downstream task with labelled data.

91
New cards

GGUF

GPT-Generated Unified Format; quantised LLM file enabling efficient local inference.

92
New cards

Hybrid Search

Combining keyword and dense (semantic) search scores for retrieval.

93
New cards

Nucleus

Set of top tokens whose cumulative probability equals top_p threshold in nucleus sampling.

94
New cards

Word Sense Disambiguation

Determining which sense of a polysemous word is used in context.

95
New cards

Sentence Encoder

Model that maps entire sentences to fixed-length embeddings for tasks like similarity and classification.

96
New cards

Embeddings Averaging

Simple method to create sentence vector by averaging constituent word embeddings.

97
New cards

Dense Vector

Representation with most elements non-zero; opposite of sparse vector.

98
New cards

Sparse Vector

Vector with many zero elements typical of Bag-of-Words features.

99
New cards

Cross-Entropy Loss

Common loss function for classification measuring divergence between predicted probabilities and true labels.

100
New cards

Backpropagation

Algorithm for computing gradients of loss with respect to model weights, enabling training.