Comprehensive Study Notes on Forensic Audio Analysis and Speech Acoustics
Fundamentals of Physical Sound and Psychoacoustics
- Nature and Propagation of Sound
- Sound is generated by mechanical vibrations that propagate as pressure waves through a transmission medium such as air.
- At the sound source, a sound wave radiates in all directions.
- Sound energy is transferred to the transmission medium in the form of alternating compressions and rarefactions of air particles.
- The air medium itself does not travel; instead, individual air particles transfer kinetic motion from one particle to another, analogous to waves moving across the surface of the sea.

- Key Physical Parameters of Sound
- Frequency (): Periodic vibration measured in cycles per second. The unit of measurement is the Hertz ().
- Natural sounds span a frequency spectrum between and .
- Accepted human hearing range spans from to .
- Infrasound consists of all frequencies below .
- Ultrasound consists of all frequencies above .
- Amplitude: Represents the sound pressure level (SPL) and directly correlates with perceived volume.

Wavelength (): Physical distance between consecutive identical points in a wave cycle.
Phase: Temporal relationship between two or more wave signals.
Harmonic Content: Presence and relative amplitudes of partial frequencies accompanying the fundamental frequency.
Envelope: Temporal profile describing the evolution of amplitude over time.
- Sensory Parameters of Perception
Pitch (Altezza): Ordering of sounds from low to high. Pitch depends directly on frequency; higher frequency corresponds to higher perceived pitch.
Timbre (Timbro): Characteristic sound quality or "color" allowing differentiation between sound sources sharing the same pitch and volume. Timbre is determined by:
- Attack transient characteristics.
- Spectral envelope structure.
Dynamics (Dinamica): Perception of volume and softness.
Duration (Durata): Temporal length of the sound event.
- Harmonic Content and Auditory Processing
Pure tones (sine waves devoid of harmonics) rarely exist in natural environments.
When an object vibrates, it produces multiple frequencies simultaneously. Harmonics occur at integer multiples of the fundamental frequency ().

Human auditory processing identifies the fundamental frequency () even in complex acoustic signals.
When specific harmonic components are suppressed or missing, the brain reconstructs missing fundamental information provided remaining components are harmonically correlated.
- Sound Envelope (ADSR)
Every sound event possesses a distinct temporal outline divided into four main phases:
- Attack (Attacco): Time taken for sound amplitude to rise from silence to peak maximum.
- Decay (Decadimento iniziale): Time required to drop from the initial peak to the sustain level.
- Sustain (Sustain / Dinamica interna): Steady amplitude level maintained over time.
- Release (Rilascio / Decadimento finale): Time required for amplitude to decay from sustain level back to complete silence once energy stops.

- Phase Interference and Superposition
- In-Phase Signals: Two waves of identical frequency are in phase when positive compression semicycles and negative rarefaction semicycles align exactly in time and space, producing maximum constructive interference (increased amplitude).

- Phase Opposition: Two identical frequency waves are in phase opposition ( shift) when the positive semicycle of one aligns with the negative semicycle of the other, causing complete destructive cancellation.

- Phase Shift: Partial temporal misalignment between waves results in partial additions and partial cancellations, shifting the resulting phase between component boundaries.

- Spectral Envelope and Fourier Analysis
- A spectral envelope represents a curve plotted on a frequency-amplitude plane describing component frequencies within a specific temporal window.
- Spectral characteristics of recordings change continuously over time.
- Fast Fourier Transform (FFT): Algorithm used to decompose time-domain signals into constituent frequency components calculated over small temporal windows (typically to ).
- Fourier Principle: Any arbitrary signal can be expressed mathematically as a weighted linear combination of harmonic functions (sine and cosine) operating at varied frequencies and periods.

- Auditory Perception and Equal Loudness Contours
- Sound intensity level in decibels () is calculated relative to reference pressure :
- Reference human hearing threshold () is approximately (), whereas the threshold of pain is (), providing a dynamic range of .
- ISO 226:2003 Standard: International specification mapping pure-tone sound pressure levels () across frequencies perceived as equally loud by human listeners.
- Human loudness perception is non-linear across frequencies due to physical resonances of the outer ear canal that boost effective pressure in middle-frequency regions.
- Phon / Fon Unit: Measurement unit of perceived loudness. By definition, equals at a frequency of .

Digital Audio Processing and Signal Fundamentals
Digital Audio Parameters
- Sampling Rate: Number of discrete signal samples captured per second, expressed in Hertz ().
- Quantization Depth (Bit Depth): Number of binary bits allocated per sample, determining dynamic resolution and noise floor.
- Channels: Spatial configuration, such as single channel (Mono) or dual channel (Stereo).
Signal-to-Noise Ratio (SNR)
- SNR quantifies the power ratio between desired signal components and background noise.
- When measured in decibels, SNR is defined as:
Ideal SNR calculation requires separate, clean measurement of signal () alone and noise () alone.
In real-world environments, target signal exists concurrently with noise. Comparing instead of distorts measurements, systematically overestimating SNR when actual SNR values are low ().
- Digital Audio Life Cycle
Signal progression spans five distinct stages: Sound Source Microphone Transmission Channel Analog-to-Digital (A/D) Conversion Digital Storage.
A/D Parameters: Quality depends directly on selected sampling frequency (e.g., , ) and bit depth (e.g., , ).
Lossless Encoding: Complete mathematical preservation of original signal data without information discard (e.g., FLAC, ALAC).
Lossy Encoding: Perceptually irrelevant acoustic data is permanently discarded to achieve reduced data file size (e.g., MP3, AAC).
Forensic Audio Management and Legal Framework
Evidence Handling Procedures
- Forensic audio procedures require strict documentation and evidence handling protocols beyond standard audio production studio workflows.
- Ingestion Protocols: Documentation of case circumstances, device provenance, proprietary recording formats, non-standard media types, and exact operational hardware required for proper playback.
- Physical State Documentation: Cataloging hardware damage, physical markings, serial numbers, and storage format specifications.
- Evidence Labeling: Direct marking of physical evidence containers using permanent markers specifying receipt date and examiner initials.
Laboratory Technical Standards
- Technical requirement for regular examiner hearing acuity evaluations.
- Verification and test protocols for software processing tools to guarantee predictable operational bounds.
- Facility Specifications: Acoustically isolated silent control room, minimal Electromagnetic Interference (EMI) ground protection, controlled climate (temperature, humidity, ventilation).
- Hardware Requirements: Systems featuring low-noise A/D and D/A converters, flat frequency response reference monitors, and calibrated professional headphones.
Digital Evidence Chain of Custody
- Digital data is highly susceptible to modification, deletion, or corruption during device boot, network connectivity, opening, saving, or transmission.
- Lifecycle Steps: Identification & Collection Acquisition Preservation Processing Presentation.

Verification and Hashing Standards
- Bitwise verification guarantees absolute digital duplication.
- Cryptographic Hash Properties: One-way function, collision resistance (impossible for two different files to generate identical hash values), extreme sensitivity (minimal data modification causes vast hash value divergence).
- Reference Standards: D.P.C.M. 8 February 1999 (Italian technical rules referencing ISO/IEC 10118-3:1998):
- Dedicated Hash-Function 1 (RIPEMD-160).
- Dedicated Hash-Function 3 (SHA-1).
- Storage Controls: Evidence storage directories must be flagged READ-ONLY, protected via passwords, and mirrored on isolated secondary backup computers.
Italian Legal Framework (Legge n. 48/2008 - Budapest Convention)
- Art. 244 C.P.P. (Inspections): Authorizes judicial authorities to record current conditions, reconstruct pre-existing states, and execute technical operations on computer systems using non-altering preservation measures.
- Art. 247 C.P.P. c. 1-bis (Searches): Mandates technical measures during searches of computer/telematic systems to ensure original data preservation.
- Art. 352 C.P.P. c. 1-bis: Applies preservation requirements to searches executed during in flagrante arrests.
- Art. 359 C.P.P. (Repeatable Technical Operations): Authorizes public prosecutors to appoint technical consultants for non-modifying analysis.
- Art. 360 C.P.P. (Unrepeatable Technical Operations): Mandates immediate notification to defense counsel, suspect, and victim prior to non-repeatable testing where evidence state undergoes permanent modification.
Audio Authenticity vs. Integrity and Analysis Methodologies
Distinction Between Authenticity and Integrity
- Integrity: Verification that digital file information remains completely unchanged and uncorrupted from the moment of collection to final destination.
- Authenticity: Verification that the recording is a genuine, accurate acoustic representation of events captured by a specific acquisition device at a precise physical location.
- Signal processing or lossy compression breaks file integrity without necessarily destroying underlying event authenticity.
Impact of Digital Operations on Authenticity and Integrity
| Process / Operation | Authenticity Maintained | Integrity Maintained |
|---|---|---|
| MP3 Compression | OK | NO |
| Soundcloud / YouTube Upload & Download | OK | NO |
| Dropbox Upload & Download | OK | OK |
| Signal Processing | NO | NO |
| File Open / Save | OK | NO (?) |
- Container-Based Authenticity Analysis
- Filesystem Metadata: OS table entries tracking file ownership, creation date, access date, and modification date.
- Copying creates a new creation timestamp while retaining the original modification timestamp.
- Moving alters directory location paths without updating modification timestamps.
- Internal File Metadata: Structured information embedded inside media containers.
- Container Formats: ISO Base Media File Format, MP4 (.mp4, .m4a, .m4p, .m4b, .m4r), 3GP (.3gp).
- Audio File Formats/Codecs: WAV (PCM/Advanced PCM), WMA, MP3 (MPEG-2 Audio Layer III), AAC.
- Metadata Inspection Software:
- ExifTool (by Phil Harvey): Extracts embedded file headers, firmware parameters, low-pass filter boundaries, and encoder specifications.

* **MediaInfo:** Performs deep structural stream analysis of audio tracks, frame rates, bitrates, and encoding parameters.

- Signal-Based Authenticity Analysis
- Critical Listening: Identification of background noise continuity, clicks, pops, or abnormal transient artifacts in silent environments.
- Spectrogram Visual Analysis: Reveals codec band-pass cutoffs, digital aliasing, environmental background jumps, and splice edit discontinuities.

- Electric Network Frequency (ENF) Analysis
- Electrical power grids exhibit continuous minor fluctuations around nominal frequencies ( or ) due to load imbalances.
- Mains hum couples magnetically or electrostatically into recording devices.
- Geographic Distribution:
- : Europe, Asia (excluding Saudi Arabia), Africa (excluding Liberia), Australia, South America (Argentina, Bolivia, Chile, Uruguay, Paraguay).
- : North and Central America, South America (Ecuador, Venezuela, Peru, Colombia, Brazil).
- Dual nominal standard: Japan utilizes both and .
- Matching recorded ENF curves against power grid database logs provides precise temporal proof of recording date and continuous integrity.
Forensic Audio Enhancement
- Objectives and Processing Limits
- Primary objective is enhancing speech intelligibility or subtle background acoustic events without distorting vocal nuances or introducing artificial artifacts that invalidate evidence in court.

- Time-Domain Processing
- Gain Control & Peak Normalization: Boosts overall signal amplitude to maximum level before digital clipping.
- Automatic Gain Control (AGC): Dynamically attenuates loud passages, amplifies quiet passages, and mutes noise-only regions.
- Dynamic Compressor: Automatically reduces dynamic range when input levels exceed a designated threshold.

* **Compressor Parameters:** Threshold, Attack time, Release time (overly fast settings cause audible volume modulation known as "breathing"), Ratio (), Knee (curve transition softness), Make-Up Gain (boosts signal back to target listening volume).
* **Limiter:** Extreme compression setting with ratios exceeding and instant attack times.
Noise Gate: Mutes signal when input drops below a target threshold, allowing audio through only when speech is active. Multiband noise gates process frequency ranges independently.
- Frequency-Domain Processing
Equalizer (EQ) Filter Types:
- Shelving (LowShelf, HighShelf).
- Bell.
- Band-Pass (LowPass, HighPass).
- Notch.
Equalizer Parameters: Center Frequency, Gain ( cut/boost), Factor (bandwidth quality factor; lower corresponds to wider frequency bandwidth).
Speech Isolation Range: Target speech frequencies typically concentrate between and .
Spectral Subtraction: Subtracts stationary background noise spectra (estimated during conversational speech pauses) from noisy signal frames.
Declipping: Reconstructs clipped peak waveforms caused by A/D converter overload via polynomial interpolation. Introduces synthetic mathematical samples not present in original recordings.
Forensic Transcriptions and Phonetic Fundamentals
Transcription Constraints
- Transcriptions cannot record non-verbal communication, emotional tone, pitch inflections, feelings, or unknown dialectal codes.
- Quality is degraded by background noise, speech distortion, overlap, and regional accents.
- Pre-Transcription Signal Analysis: Examiner checks peak sample levels, min/max RMS levels, total RMS levels, and clipped sample counts.
Phonetic Categorization
- Consonants: Produced by partial or total airflow obstruction in the vocal tract. Defined by Place of Articulation (obstruction location) and Manner of Articulation (obstruction structure).
- Vowels: Produced by vocal tract shape modification without radical airflow obstruction (tongue placement, lip rounding). Defined phonetically by Height, Backness, Rounding, Nasalization, Length, and Dynamicity (Monophthongs, Diphthongs, Triphthongs).
Forensic Speaker Identification and Voice Comparison
Voice as a Biometric Trait
- Properties: Uniqueness (limited by anatomical similarities), Universality, Generality, Stability (evolves over lifespan), Robustness (sensitive to channel, microphone, and compression degradations).
Likelihood Ratio (LR) Evaluative Framework
- Evaluates evidence () under two competing hypotheses: Prosecution Hypothesis (: samples originate from the same speaker) versus Defense Hypothesis (: samples originate from different speakers).
- Verbal Scale and LR Values
| Likelihood Ratio () | Log Likelihood Ratio () | Supported Proposition | Verbal Scale Equivalent |
|---|---|---|---|
| supported over () | Molto forte (Very strong) | ||
| supported over () | Supporto forte / Moderatamente forte | ||
| supported over () | Moderatamente forte | ||
| supported over () | Moderato | ||
| Neutral support ( / ) | Nessun supporto / Lieve / Limitato | ||
| supported over ( / ) | Limitato | ||
| supported over () | Moderato | ||
| supported over () | Moderatamente forte | ||
| supported over () | Molto forte |
Italian Software Applications
- IDEM (Identificazione Vocale): Extracts fundamental frequency () and vowel formants (/i, e, a, o/). Does not distinguish between close-mid and open-mid vowels ([e-ɛ], [o-ɔ]). Operates on a parametric Gaussian model.
- SMART: Extracts acoustic features while accounting for dialectal variation, accented vocalic systems, and diaphasic variations. Operates on a non-parametric model.
Methodological Comparison: Auditory-Acoustic vs. Automatic Speaker Recognition
- Auditory Phonetics + Acoustic Analysis (Au-AC): Combines expert listening (pitch, quality, prosody, dialect, disfluencies evaluated on 10-point scales) with computer-extracted formant measurements ().

- Inter-Software Measurement Variance Example:
| Formant | Software A Measurement () | Software B Measurement () | Delta () |
|---|---|---|---|
| F1 | 385 | 397 | 12 |
| F2 | 1703 | 1716 | 13 |
| F3 | 2220 | 2265 | 45 |
| F4 | 3404 | 3431 | 27 |
| F5 | 4350 | 4446 | 96 |
- System Comparison Table:
| Feature | Auditory-Acoustic Approach (Au-AC) | Automatic Speaker Recognition (ASR) |
|---|---|---|
| Processing Speed | Proportional to material length | Potentially very fast |
| Result Output | Expert evaluation report | Numerical Likelihood Ratio |
| Repeatability | Dependent on analyst skill & conditions | Fully repeatable |
| Error Rate Evaluation | Difficult to evaluate quantitatively | Known statistical error rate |
Source-Filter Theory and Socio-Phonics
- Praat Signal Analysis
- Open-source software utilized for acoustic analysis, phonetic testing, and visual spectral generation.

- Source-Filter Acoustic Model
- Source: Glottal or supraglottal airflow generator (vocal cord vibration driving fundamental frequency ).
- Filter: Modifications produced by vocal tract shape (oral/nasal cavities).
- Resonator: Cavity amplifying specific frequencies into spectral peaks called formants.
- Nodes: Constrictions in the cavity leading to signal frequency increases.
- Antinodes: Cavity openings leading to frequency decreases.
- Vowel Resonance Tube Models:
- Vowel [i]: Tongue constriction at the palatal position ().

* Vowel [a]: Tongue constriction at the pharyngeal position ().

* Vowel [u]: Tongue constriction at the velar position ().

Anatomical Formant Scaled Differences
- Female supralaryngeal vocal tracts are approximately shorter than male tracts, yielding formant frequencies that are on average higher.
- correlates directly with vowel height / degree of jaw opening.
- correlates directly with vowel backness / degree of tongue frontness.
Socio-Phonetics and Cognitive Biases
- Voices encode social markers (age, gender, social class, origin, education).
- Cognitive Biases & Stereotypes: "Hearing what is expected" based on prior identity knowledge.
- Linguistic Priming: Audio ordering affects speaker recognition. Lower recording quality amplifies perceptual biases.