1/21
A collection of essential terms, algorithms, and evaluation metrics used in the detection of phishing emails via Natural Language Processing as identified in a systematic review of 100 research articles (2006-2022).
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Phishing
A social engineering threat that exploits the ignorance of uninformed internet users to obtain sensitive information from them in a deceiving manner.
Anti-Phishing Working Group (APWG)
A non-profit foundation that records phishing activity; it reported a rise from 44,008 attacks in Q1 2020 to a record high of 260,642 monthly attacks in July 2021.
Support Vector Machines (SVMs)
A heavily utilised supervised learning algorithm for detecting phishing emails that plots data items as points in an n-dimensional space to extract the most appropriate hyper-plane.
TF-IDF
Term Frequency-Inverse Document Frequency; an NLP technique that reveals the significance of a keyword to a document within a textual corpus.
Nazario phishing corpus
The most commonly used dataset for benchmarking phishing email detection methods, appearing in 42 of the analyzed studies.
Gini Index
A purity index used in decision trees to measure the probability of a randomly chosen feature being incorrectly classified.
Entropy
An index proportional to information gain used in decision trees to measure uncertainty.
Random Forest (RF)
An ensemble classifier that makes predictions using a variety of decision trees constructed using a random selection of attributes.
Recurrent Neural Network (RNN)
A deep learning model used for sequential data modelling that learns hidden sequential associations in variable-length input sequences.
Long Short-Term Memory (LSTM)
A polymorphism of RNN developed to overcome gradient exploding and vanishing issues by using gates to influence the state and output.
Principal Component Analysis (PCA)
A technique that extracts mapping from original dimensional space to a smaller dimensional space to ensure minimum information loss.
Latent Semantic Analysis (LSA)
A mathematical procedure for Natural Language Processing designed to embed topics within input documents extracted from the highest feature values.
Chi-square (χ2)
A feature selection procedure that assesses individual features by measuring the linear dependency between an input feature and a target class.
Precision
A classification performance metric calculated as Precision=TP+FPTP.
Recall
A classification performance metric calculated as Recall=TP+FNTP.
F1-measure
A classification performance metric calculated as F1-measure=2×Precision+RecallPrecision×Recall.
Accuracy
A classification performance metric calculated as Accuracy=TP+FP+TN+FNTP+TN.
Bio-inspired computing (BIC)
Optimization algorithms based on natural behaviors (e.g., Grey Wolf or Chicken Swarm) characterized by self-correction and adjustment to changing environments.
Adam optimiser
The most frequently used optimization technique in the reviewed literature, appearing in more than 26% of studies.
Crimeware
A kind of malware defined as software that accomplishes illegal activities intended to generate monetary gains for an assailant.
Sequential minimal optimisation (SMO)
An optimization technique used in 21% of reviewed papers that helps identify the significance of words in textual datasets.
Enron dataset
A public corpus used for email classification research identified in 23 of the reviewed phishing detection studies.