1/102
Key vocabulary terms spanning text preprocessing, feature engineering, modelling techniques, evaluation metrics, modern transformer architectures, and generative AI concepts covered throughout the lecture series.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Text Mining
AI-driven process of converting unstructured text into structured, analysable data.
Natural Language Processing (NLP)
Field of computer science and AI focused on interactions between computers and human language.
NLP Pipeline
Sequence of steps—Corpus creation, Train/Validation/Test split, Feature Engineering, Task execution—used to build text models.
Corpus
A collection of text documents organised as a dataset for NLP tasks.
Corpora
Plural of corpus: multiple text collections.
Train-Validation-Test Split
Standard method to divide a corpus (e.g., 80 / 10 / 10 %) for model building and evaluation.
K-Fold Cross-Validation
Technique that iteratively partitions data into k subsets to improve reliability of evaluation.
Feature Engineering (Text)
Transforming raw text into numerical representations usable by ML algorithms.
Bag-of-Words (BoW)
Representation where each vocabulary word is a feature; documents become sparse count vectors.
Curse of Dimensionality
Problems arising from very large feature spaces, causing sparse data and overfitting.
Tokenization
Process of splitting text into tokens (words, phrases, symbols).
Compound Token
Multiple-word expression treated as one token, e.g., “United Kingdom.”
Stop Words
Common words (the, and, of, …) often removed because they add little semantic value.
Lowercasing
Normalization step converting all text to lower-case to shrink vocabulary size.
Regular Expression (Regex)
Pattern-matching syntax used to normalize or anonymise text (dates, names, numbers).
Stemming
Heuristic removal of word endings to obtain a crude root (e.g., changing → chang).
Lemmatization
Mapping inflected words to their dictionary root (e.g., changes → change).
Part-of-Speech (POS) Filtering
Keeping only selected grammatical categories (nouns, verbs, adjectives, adverbs) for analysis.
Preprocessing Pipeline
Typical order: normalization → lowercasing → tokenization → stop-word removal → stemming/lemmatization → POS filtering.
Variability (in NL)
Same meaning can be expressed by different sentences (paraphrases).
Ambiguity (in NL)
Single sentence may have multiple interpretations; resolved using context.
Generalization Challenge
Model trained in one domain faces Out-of-Domain or Out-of-Vocabulary inputs when deployed elsewhere.
Term Frequency (TF)
Count of a term’s occurrences within a document.
Inverse Document Frequency (IDF)
log(|D|/n_t); down-weights terms that appear in many documents.
TF-IDF
Weighting scheme TF×IDF elevating informative, rare terms and reducing common ones.
N-Gram
Contiguous sequence of n items from text (unigram, bigram, trigram, …).
Data Sparseness
Increase in zero entries when using large N-grams or large vocabularies.
Keyword Extraction
Task of identifying salient terms in text.
Text Similarity
Measuring how alike two pieces of text are using distance metrics.
Document Classification
Assigning predefined labels to documents.
Sentiment Analysis
Determining the emotional polarity (positive/negative/neutral) of text.
Topic Modelling
Unsupervised discovery of hidden thematic structures in a corpus.
Information Retrieval (IR)
Process of obtaining relevant documents in response to a query.
K-Nearest Neighbour (KNN)
Baseline classifier: label is taken from closest training documents in feature space.
Euclidean Distance
Straight-line metric between two vectors; suited for dense data.
Manhattan Distance
Sum of absolute coordinate differences between vectors.
Cosine Similarity
Normalized dot product measuring angle between vectors; robust for sparse text data.
Jaccard Coefficient
|Intersection| / |Union| similarity of two sets.
Dice Similarity
2×|Intersection| / (|A|+|B|); less sensitive to length differences than Jaccard.
Levenshtein (Edit) Distance
Minimum insertions, deletions, substitutions to transform one string into another.
Soundex
Phonetic hashing algorithm indexing names by pronunciation.
Word Representation
Vector encoding that captures semantic information of a word.
Term-Term Matrix
|V|×|V| co-occurrence counts of words in a corpus.
Point-wise Mutual Information (PMI)
log [p(w,c)/(p(w)p(c))]; measures word-context association strength.
Positive PMI (PPMI)
PMI values with negatives replaced by zero.
Skip-Gram (Word2Vec)
Neural model predicting context words from a target word to learn dense embeddings.
Embedding
Dense, low-dimensional vector representing token semantics.
Semantic Vector Arithmetic
Property where relationships emerge in embedding space (king – man + woman ≈ queen).
Recurrent Neural Network (RNN)
Neural architecture with cyclic connections allowing sequence processing.
Vanishing Gradient
Problem in RNN training where gradients shrink across long sequences, stalling learning.
Bidirectional RNN
RNN processing sequences in both forward and backward directions to exploit full context.
Long Short-Term Memory (LSTM)
RNN variant with gates (forget, input, output) and cell state to capture long dependencies.
Sequence-to-Sequence (Seq2Seq)
Encoder-decoder architecture mapping input sequences to output sequences (e.g., translation).
Attention Mechanism
Method enabling decoder to focus on relevant encoder states via weighted sums.
Self-Attention
Attention applied within the same sequence; token attends to all other tokens.
Transformer
Model composed of self-attention and feed-forward layers; allows parallel processing, no recurrence.
Encoder (Transformer)
Stack of self-attention blocks producing contextual embeddings.
Decoder (Transformer)
Stack that uses masked self-attention and encoder-decoder attention to generate output tokens.
Positional Embedding
Vector added to token embedding to encode order information for Transformers.
Multi-Head Attention
Parallel attention heads capturing different relation subspaces.
Masked Attention
Technique in decoder preventing access to future tokens during training/generation.
Autoregressive Generation
Producing text token-by-token, conditioning each prediction on previously generated tokens.
Large Language Model (LLM)
Very large Transformer (billions of parameters) capable of diverse language tasks.
Temperature (Sampling)
Parameter controlling randomness; higher values yield more diverse outputs.
top_p (Nucleus Sampling)
Sampling from smallest probability mass ≥ p to balance creativity and coherence.
top_k Sampling
Restricts candidate tokens to top k most probable before sampling.
Prompt Engineering
Crafting input instructions to steer LLM behaviour and outputs.
Zero-Shot Prompting
Asking an LLM to perform a task without providing examples.
Few-Shot Prompting
Including a few labelled examples in the prompt to guide the model.
Conversation Buffer Memory
Technique of appending full dialogue history to prompt to maintain context.
Conversation Summary Memory
Using an LLM to condense dialogue history into a shorter summary fed back to the model.
Retrieval-Augmented Generation (RAG)
Framework combining vector search with LLM prompting to generate grounded answers.
Vector Database
Specialised store of embeddings enabling efficient similarity search.
Semantic Search
Information retrieval based on embedding similarity rather than keyword overlap.
Chunking (RAG)
Splitting long documents into smaller parts before embedding and storage.
Quantization (GGUF)
Compression technique reducing parameter precision to run LLMs faster with less memory.
Temperature-Controlled Creativity
Concept where low temperature → deterministic, high temperature → creative outputs.
Agentic AI (LLM Agent)
System where an LLM selects and executes actions/tools autonomously to achieve goals.
ReAct Pattern
Agent loop: Thought → Action → Observation enabling reasoning and tool usage.
Evaluation Metric – BLEU
N-gram overlap score for machine translation output vs reference.
Evaluation Metric – ROUGE
Recall-oriented n-gram overlap metric for summarisation quality.
Evaluation Metric – BERTScore
Semantic similarity metric using BERT embeddings between candidate and reference texts.
Evaluation Metric – METEOR
Machine translation metric incorporating stemming, synonymy and alignment.
Quality Estimation (QE)
Predicting translation quality without reference by modelling sentence-level or word-level scores.
COMET
Neural evaluation metric for machine translation correlated with human judgement.
xCOMET
Multilingual extension providing state-of-the-art QE and evaluation performance.
Tower LLM
Unbabel’s multilingual LLM optimised for translation-related tasks.
Instruction Tuning
Fine-tuning an LLM on (instruction, response) pairs to follow natural language commands.
Masked Language Modelling (MLM)
Pretraining objective where model predicts intentionally masked tokens (used in BERT).
Fine-Tuning
Adapting a pretrained model to a specific downstream task with labelled data.
GGUF
GPT-Generated Unified Format; quantised LLM file enabling efficient local inference.
Hybrid Search
Combining keyword and dense (semantic) search scores for retrieval.
Nucleus
Set of top tokens whose cumulative probability equals top_p threshold in nucleus sampling.
Word Sense Disambiguation
Determining which sense of a polysemous word is used in context.
Sentence Encoder
Model that maps entire sentences to fixed-length embeddings for tasks like similarity and classification.
Embeddings Averaging
Simple method to create sentence vector by averaging constituent word embeddings.
Dense Vector
Representation with most elements non-zero; opposite of sparse vector.
Sparse Vector
Vector with many zero elements typical of Bag-of-Words features.
Cross-Entropy Loss
Common loss function for classification measuring divergence between predicted probabilities and true labels.
Backpropagation
Algorithm for computing gradients of loss with respect to model weights, enabling training.