Quantifying the Voice: Clinical Perceptual Scaling and Visi-Pitch Acoustic Analysis
Foundations of Voice Quantification: The Clinician vs. The Machine
The Clinician's Ear: Auditory-Perceptual Evaluation
Nature: This approach is qualitative and listener-dependent.
Purpose: It aims to capture the holistic human experience of voice quality by identifying specific characteristics through the clinician's expertise.
Core Characteristics Tracked:
Roughness.
Breathiness.
Strain.
Primary Tools: Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the GRBAS scale.
The Diagnostic Machine: Instrumental Acoustic Analysis
Nature: This approach is quantitative and driven by biofeedback.
Purpose: Acoustic data is utilized to validate and quantify clinical perception by measuring the exact physical properties of sound waves.
Core Properties Measured:
Frequency.
Amplitude.
Perturbation.
Primary Tools: Visi-Pitch (specifically modules like Real-Time Pitch and Multidimensional Voice Program [MDVP]).
Standardized Perceptual Scales: GRBAS and CAPE-V
The GRBAS Scale
Protocol: Evaluation is conducted using samples of spontaneous speech.
Scale Type: Equal-appearing interval scale.
Scoring Range: to (, ).
Metrics Measured (The GRBAS Acronym):
G (Grade): The overall severity of the voice abnormality.
R (Roughness): Perception of irregular vocal fold vibration.
B (Breathiness): Perception of air leakage through the glottis.
A (Asthenia): Perceived weakness or lack of power in the voice.
S (Strain): Perception of excessive vocal effort or hyperfunction.
The CAPE-V Scale
Protocol: Evaluation is conducted using sustained phonation (vowels) and specific sentences.
Scale Type: Visual Analog Scale (VAS).
Scoring Range: to (, ).
Metrics Measured:
Overall Severity.
Roughness.
Breathiness.
Strain.
Pitch.
Loudness.
Key Insight on Asthenia: Asthenia is explicitly excluded from the CAPE-V. It was determined to be highly influenced by loudness and difficult for clinicians to parse apart perceptually from other factors.
CAPE-V Cutoff Boundaries
Clinical categories are established by translating visual analog lines into distinct segments on a ruler.
Overall Severity / Grade Mapping:
Normal (Consistent with GRBAS 0): .
Mild (Consistent with GRBAS 1): .
Moderate (Consistent with GRBAS 2): .
Severe (Consistent with GRBAS 3): .
Metric-Specific Boundaries:
Roughness: Mild (), Moderate (), Severe ().
Breathiness: Mild (), Moderate (), Severe ().
Strain: Defined primarily in the Normal range ().
Visi-Pitch: Standardized Software Overview
Definition: An online, standardized software program (often part of the Sona-Speech system) designed to assess disordered voice and speech through the use of visual biofeedback.
Primary Clinical Applications:
Routine therapy tasks.
Evaluating progress (e.g., tracking changes in pitch range over time).
Visual cueing for patients.
Determining voice typing.
Specialized use cases: Motor speech disorders, fluency, auditory rehabilitation, and accent modification.
Critical Diagnostic Limitation: Visi-Pitch is intended for assessment and performance tracking. It is not a standalone diagnostic tool. While results provide objective biofeedback regarding current performance, they do not determine the underlying medical etiology or disorder.
The Visi-Pitch Ecosystem: Two Analytical Lenses
Module 1: Real-Time Pitch (The Macro View)
Focus: Evaluation of prosody, vocal range, and connected speech.
Feedback: Provides real-time visual data on pitch, energy, and time parameters.
Key Clinical Tasks Captured:
Habitual Pitch ()
talking out of habitual pitch can lead to nodules
Maximum Phonation Time (MPT).
phonatory function, which is crucial for evaluating vocal health and efficiency in patients. A shortened MPT may indicate vocal strain or pathology.
Monotone Evaluation.
Pitch Range.
Loudness.
Module 2: Multi-Dimensional Voice Program (MDVP) (The Micro View)
Focus: Analysis of vocal fold perturbation and noise values.
Input: Derived from a single sustained vocalization (typically the vowel /a/).
Key Clinical Metrics Captured:
Jitter.
Shimmer.
Relative Average Perturbation (RAP).
Noise-to-Harmonic Ratio (NHR).
Voice Turbulence Index (VTI).
Real-Time Pitch: The Five Macro Tasks and Protocol
Detailed Task Breakdown:
Habitual Pitch (): Represents the average pitch during connected speech. It serves as the natural baseline for the patient.
Maximum Phonation Time (MPT): Measured as the longest sustained /a/ produced on a single breath. This is an indicator of respiratory-phonatory efficiency.
Monotone Evaluation: Assessment of prosody while the patient reads standardized passages (e.g., "The Grandfather Passage" or "The Rainbow Passage"). Flat prosody may indicate neurological or affective issues.
Pitch Range: Evaluated via vocal glides from the patient's lowest to highest possible pitch. A restricted range can indicate pathology or the effects of aging.
Loudness (): Measures vocal intensity in . This assesses if the patient is using appropriate and consistent energy levels during phonation.
Operational UI Process Flow:
Launch: Open Sona-Speech and select the 'Real-Time Pitch' module.
Execute: Follow on-screen instructions for specific tasks (e.g., reading or gliding). Navigate to the 'Energy Tab' specifically for loudness measurements.
Capture: Click 'Stop' and then 'Compute Result Statistics'.
Document: Capture a clear photo of the results screen. Clear the page using the 'X' icon (top right) to prepare for the next trial.
Acoustic Biomarkers: Real-Time Pitch Norms
Habitual Pitch (Average ):
Female Baseline: (Average: ).
Male Baseline: (Average: ).
Maximum Phonation Time (MPT):
General Adult Norm: .
Female Target: .
Male Target: .
Pitch Range:
Female Range: (Musical equivalent: Semitones ).
Male Range: (Musical equivalent: Semitones ).
Clinical Expectation: Approximately octaves ().
Loudness:
Conversational Average: .
Typical Range: .
MDVP Perturbation Matrix and Analysis Protocol
The Golden Rule of MDVP: Analysis requires a Green Signal Only. The patient must phonate a steady, comfortable /a/. If the recording signal enters the "red zone," the extraction of perturbation data will be invalid.
Analysis Execution Steps:
Navigate to Sona-Speech and select "MDVP".
Click 'File' then 'Record' (Shortcut: ).
Patient produces a steady /a/ while the clinician ensures a green signal is maintained.
Click 'Stop'.
Navigate to 'Protocol' and select "Complete MDVP Analysis". Take a photo.
Navigate to 'Protocol' and select "Show MDVP Parameters in Active Window" (Shortcut: ). Take a photo.
Go to 'Window' and select "Purge Active Window" to clear the current data.
MDVP Metric Definitions and Thresholds:
Jitter (%): Measures pitch instability between cycles. Clinically sounds like pitch breaks or roughness. Threshold: < 1.0\text{\textbackslash,\text{\textbackslash,}\%} (Ideally ).
pitch breaks
vocal folds are not openeing and closing in a steady manner
Should be moving in a sequence, steady manner
Shimmer (%): Measures amplitude instability/variability between cycles. Perceived clinically as hoarseness or roughness. Threshold: < 3.5\text{\textbackslash,\text{\textbackslash,}\%} (General cutoff < 5\text{\textbackslash,\text{\textbackslash,}\%}).
amplitude (sometimes high, sometimes low)
Relative Average Perturbation (RAP): Represents smoothed pitch instability. Clinically indicates irregular vocal fold vibration. Threshold: < 0.5\text{\textbackslash,\text{\textbackslash,}\%} (Average observed: ).
Noise-to-Harmonic Ratio (NHR): Compares noise components against periodic sound. Perceived as a breathy, rough quality. Threshold: < 0.19.
Voice Turbulence Index (VTI): Measures high-frequency turbulence. Indicates weak or breathy phonation typical of loose vocal fold adduction. Threshold: < 0.02 (Average observed: ).
Synthesis: Correlating Perceptual and Acoustic Findings
Acoustic Validation of Clinical Perception:
Subjective Finding: Severe Roughness on the CAPE-V.
Objective Correlate: Elevated Shimmer and Jitter on MDVP (reflecting irregular vibration and amplitude variations).
Subjective Finding: Severe Breathiness on the CAPE-V.
Objective Correlate: Elevated NHR and VTI on MDVP (reflecting high noise/turbulence from incomplete glottal adduction).
Subjective Finding: Monotone conversational speech.
Objective Correlate: Restricted Pitch Range and flat variance on Real-Time Pitch modules.
Final Summary Principle: The Visi-Pitch system does not replace the expertise of the clinician's ear; it serves to provide a quantifiable, objective dimension to the clinician's perceptual findings.