Systematic Literature Review on Phishing Email Detection Vocabulary

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/21

flashcard set

Earn XP

Description and Tags

A collection of essential terms, algorithms, and evaluation metrics used in the detection of phishing emails via Natural Language Processing as identified in a systematic review of 100 research articles (2006-2022).

Last updated 10:42 PM on 6/9/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

22 Terms

1
New cards

Phishing

A social engineering threat that exploits the ignorance of uninformed internet users to obtain sensitive information from them in a deceiving manner.

2
New cards

Anti-Phishing Working Group (APWG)

A non-profit foundation that records phishing activity; it reported a rise from 44,00844,008 attacks in Q1 2020 to a record high of 260,642260,642 monthly attacks in July 2021.

3
New cards

Support Vector Machines (SVMs)

A heavily utilised supervised learning algorithm for detecting phishing emails that plots data items as points in an nn-dimensional space to extract the most appropriate hyper-plane.

4
New cards

TF-IDF

Term Frequency-Inverse Document Frequency; an NLP technique that reveals the significance of a keyword to a document within a textual corpus.

5
New cards

Nazario phishing corpus

The most commonly used dataset for benchmarking phishing email detection methods, appearing in 4242 of the analyzed studies.

6
New cards

Gini Index

A purity index used in decision trees to measure the probability of a randomly chosen feature being incorrectly classified.

7
New cards

Entropy

An index proportional to information gain used in decision trees to measure uncertainty.

8
New cards

Random Forest (RF)

An ensemble classifier that makes predictions using a variety of decision trees constructed using a random selection of attributes.

9
New cards

Recurrent Neural Network (RNN)

A deep learning model used for sequential data modelling that learns hidden sequential associations in variable-length input sequences.

10
New cards

Long Short-Term Memory (LSTM)

A polymorphism of RNN developed to overcome gradient exploding and vanishing issues by using gates to influence the state and output.

11
New cards

Principal Component Analysis (PCA)

A technique that extracts mapping from original dimensional space to a smaller dimensional space to ensure minimum information loss.

12
New cards

Latent Semantic Analysis (LSA)

A mathematical procedure for Natural Language Processing designed to embed topics within input documents extracted from the highest feature values.

13
New cards

Chi-square (χ2\chi^2)

A feature selection procedure that assesses individual features by measuring the linear dependency between an input feature and a target class.

14
New cards

Precision

A classification performance metric calculated as Precision=TPTP+FPPrecision = \frac{TP}{TP + FP}.

15
New cards

Recall

A classification performance metric calculated as Recall=TPTP+FNRecall = \frac{TP}{TP + FN}.

16
New cards

F1-measure

A classification performance metric calculated as F1-measure=2×Precision×RecallPrecision+RecallF1\text{-measure} = 2 \times \frac{Precision \times Recall}{Precision + Recall}.

17
New cards

Accuracy

A classification performance metric calculated as Accuracy=TP+TNTP+FP+TN+FNAccuracy = \frac{TP + TN}{TP + FP + TN + FN}.

18
New cards

Bio-inspired computing (BIC)

Optimization algorithms based on natural behaviors (e.g., Grey Wolf or Chicken Swarm) characterized by self-correction and adjustment to changing environments.

19
New cards

Adam optimiser

The most frequently used optimization technique in the reviewed literature, appearing in more than 2626% of studies.

20
New cards

Crimeware

A kind of malware defined as software that accomplishes illegal activities intended to generate monetary gains for an assailant.

21
New cards

Sequential minimal optimisation (SMO)

An optimization technique used in 2121% of reviewed papers that helps identify the significance of words in textual datasets.

22
New cards

Enron dataset

A public corpus used for email classification research identified in 2323 of the reviewed phishing detection studies.