Measuring speech in schizophrenia spectrum disorders
A practical guide to speech features, elicitation tasks and study designÂ
Speech is an unusual clinical signal. It is produced naturally, requires no instrumentation beyond a microphone, can be collected in a clinic or a participant’s home, and carries information at three levels at once: the acoustic properties of the voice, the timing of its delivery, and the linguistic structure and meaning of what is said. A few minutes of recorded speech yields hundreds of quantifiable variables across those three levels, none of which requires the participant to do anything unfamiliar.
For research in schizophrenia spectrum disorders (SSDs), that is particularly relevant for the assessment of negative symptoms. Alogia and blunted affect are already inferred from how much a patient says, at what rate, and with what prosodic variation; the rating scales in routine use formalize those judgements rather than replacing the observations underneath them. But those scales require extensive rater training, are ordinarily administered in clinic, are susceptible to rater and expectancy bias, and may not register change falling between severity anchors. Negative symptoms predict functional outcome and quality of life, and no current treatment produces sustained improvement in them — so a measure that is objective, repeatable and collectable remotely has obvious value alongside clinician ratings.
What gets measuredÂ
Automated pipelines extract features at three levels, each with implications for study design.Â
Acoustic features describe the voice signal: mean fundamental frequency and its variance, characterizing pitch and prosodic range; intensity, or perceived loudness; measures of voice quality like jitter and shimmer, which quantify cycle-to-cycle perturbation in frequency and amplitude, and harmonic-to-noise ratio.Â
Timing features describe the rhythm of delivery: speech rate in words per minute, articulation rate in syllables per second, mean pause duration, counts of silent and filled pauses, and the proportion of a recording containing speech rather than silence.Â
Linguistic features describe what was said: lexical diversity and word frequency, the distribution of parts of speech, syntactic complexity from parse-tree depth, semantic coherence between successive utterances calculated from word embeddings, sentiment scored against published affective norms, and speech-graph properties representing how ideas connect across a narrative.Â
A pipeline of this kind returns several hundred variables. Which of them belong in a protocol?
How speech is elicitedÂ
Speech is collected through short structured tasks rather than recorded conversation, which keeps measurement comparable across participants, sites and visits.Â
Several types of tasks have commonly been used in research studies. Picture description asks the participant to describe a scene presented to them, with a follow-up prompt for additional detail. It takes approximately two minutes and returns measures of fluency, lexical richness, syntactic structure and semantic coherence, together with task-performance scoring against the content units present in the image. Open-ended journaling uses a narrative prompt, takes approximately one minute, adds prosodic and sentiment measures to the same linguistic set, and allows for psychologically relevant prompts that can elicit more affective content. Paragraph reading and recall has the participant read a standardized passage aloud and recall it after a delay, roughly a minute each; because the words are supplied, reading isolates motor speech and fluency from language production, while recall adds a measure of episodic memory.Â
Recordings are transcribed by automatic speech recognition (ASR), and acoustic, timing and linguistic features are generated from the audio and transcript together. Â
Lessons from the literatureÂ
The evidence base is now substantial enough to draw practical guidance from. A systematic review published in June pooled diagnostic classification accuracy across 20 schizophrenia studies (n = 2,033) at 84.0% (95% CI 81.1–86.6), with moderate heterogeneity (I² = 54.7%).1 The question is no longer whether speech carries signal, but how to design a study that measures it well. A few lessons:Â
- Choose the right features for the research question. Some features separate patients from controls, some vary with severity within a diagnosis. Decide which question you are asking at the start.Â
- Match the number of features to your sample size. For smaller study populations, reporting the best of several hundred can be a weaker position than naming a handful in advance and reporting them whatever they show. Simpler models tended to outperform deep learning at smaller population scales, and more data is not automatically better data: single-modality models often outperformed multimodal ones. Larger participant samples can support more features, and it may be advantageous to include an unfiltered large set of features for a data-driven analysis or analysis using data-reduction approaches.Â
- Consider establishing reliability before validity. Reliability constrains the magnitude of association a study can detect, but says nothing about whether an association exists. Both need separate evidence, ideally in that order.Â
- The use case of interest will shape the required properties of speech measures — match the standard to your purpose. Acoustic features are consistently the most stable across repeated administration, timing features less so, linguistic features the most variable — a pattern now shown in independent samples using different pipelines.2,3 Diagnostic separation, trait-like characterization or stratification, and within-person longitudinal monitoring set different psychometric bars.Â
- Consider aggregating rather than testing individually. Composite indices and latent factors built across correlated semantic features can be more stable than their constituent features.3,4Â
- Use more than one stimulus per visit where you can. Averaging across trials gives a better estimate, and eliciting more speech tends to improve reliability. For longitudinal designs, if stimuli are not well equated, presenting one repeating and one novel stimulus can balance practice effects against stimulus effects. Other approaches include cycling through multiple novel stimuli at each visit. Â
Two studies, two answersÂ
On the topic of feature selection, two papers published this summer that pose different questions illustrate the point directly.Â
Spilka and colleagues asked which voice features track negative symptom severity within schizophrenia spectrum disorders.2 Sixty-two inpatients completed journaling and picture description tasks at hospital admission and again at discharge. Thirty-eight features were specified in advance based on prior research, and criteria applied in sequence: adequate test-retest reliability, significant correlation with Scale for the Assessment of Negative Symptoms (SANS) total score, and replication of both at the second visit.Â
Four picture-description features (mean pause duration, speech proportion, speech rate and unfilled pauses) met every criterion and three replicated from the journaling task. None correlated with positive or overall symptom severity, supporting their specificity, and none was associated with antipsychotic dose or parkinsonian symptoms. Speech rate was the strongest candidate, also showing convergent validity with an independent negative symptom measure across both tasks and visits. No acoustic feature correlated with negative symptom severity.Â
Tang and colleagues asked whether a brief speech sample can identify individuals needing further psychiatric assessment.5 Their ACES study analysed 266 participants from inpatient and outpatient services. The design is transdiagnostic rather than SSD-specific, so the results describe psychiatric caseness broadly in a heterogenous clinical population. Separating healthy volunteers from participants with any psychiatric disorder reached an F1 of 0.865, with 88% sensitivity and 68% specificity; speech from picture description alone reached F1 of 0.830. A more nuanced second split, between serious mental illness (defined as bipolar disorder or any psychotic disorder, including SSD and major depressive disorder with psychosis) and other psychiatric disorders, reached 0.626 — falling short, the authors say, of the accuracy required for clinical implementation but potentially useful for triage purposes. The most influential features for the first split were acoustic: instability (shimmer) and reduced pitch variance. Split 2 was influenced by articulation rate and filled pauses.Â
The same Winterlight extraction pipeline, overlapping investigators and comparable tasks produced largely non-overlapping feature sets. Acoustic measures allowed for the evaluation of psychiatric case-control differences but weren’t as helpful in resolving the severity of mental illness; timing measures were sensitive to the negative symptoms of schizophrenia. This divergence is not anomalous — it demonstrates the richness of speech as a clinical signal, with the ability to convey information about various aspects of psychopathology in a variety of contexts of use.Â
What this makes practicalÂ
The four markers identified by Spilka and colleagues need little infrastructure. Speech proportion and mean pause duration are computed from the audio waveform alone and need no transcript; speech rate and unfilled pauses depend on whether a word was spoken rather than which word it was. None requires analysis of language content, which lowers transcription burden and reduces the privacy exposure of storing patient language.Â
The authors identify three potential applications, pending further validation: screening or enrichment in trials targeting negative symptoms, remote monitoring between scheduled visits and a quantitative endpoint potentially more sensitive to small changes than a scale with coarse severity anchors. Longitudinal work in the same programme has shown a speech and language component tracking change in thought disorder and negative symptoms across repeated timepoints.6Â
The future directions that emerge from these studies are mostly suited to academic groups: independent cross-site replication, cohorts with healthy comparison groups and normative data, designs long enough to test sensitivity to change rather than stability over days and fairness auditing reported as an outcome.Â
Speech analysis for academic researchÂ
The studies above used Winterlight speech analysis, which will soon be available to academic research groups. Picture description, journaling, and paragraph reading and recall are administered through the Connect platform, recorded, transcribed by ASR, and returned as named acoustic, timing and linguistic features with a feature dictionary defining each one. The features are interpretable rather than opaque model outputs, which is what allows a study to examine the face validity of its predictors and prespecify a small set rather than reporting the best of several hundred. Multiple stimuli are available for each task, supporting the longitudinal with repeat testing study designs described above.Â
If you are considering adding speech assessment to a psychosis cohort, or working out which features to name in a protocol, we are glad to talk it through — just get in touch.
ReferencesÂ
- Parsapoor M, et al. Supervised machine learning classifiers for schizophrenia and bipolar disorder using speech and language: a systematic review, meta-analysis, and novel quality assessment framework. Artif Intell Rev. 2026. doi:10.1007/s10462-026-11587-6
- Spilka MJ, et al. Identification of robust speech-based markers of negative symptom severity in schizophrenia spectrum disorders. Schizophrenia. 2026 (in press). doi:10.1038/s41537-026-00788-1
- Çokal D, et al. What is the retest reliability of computationally extractable speech and language markers? Comput Speech Lang. 2026;100:101981.Â
- Palominos C, et al. A single composite index of semantic behavior tracks symptoms of psychosis over time. Schizophr Res. 2025;279:116–127.
- Tang SX, et al. ACES: ascertaining diagnosis classification with elicited speech in individuals with heterogeneous, comorbid psychiatric disorders. Psychiatry Res. 2026;365:117389. Tang SX, et al. Automated speech and language markers of longitudinal changes in psychosis symptoms. NPP—Digit Psychiatry Neurosci. 2025;3:13.Â

