Skip to content

Can a five-minute voice recording detect anhedonia as accurately as a $500 brain scan?

In our preregistered benchmark it matched the scan: 0.63 AUC against 0.58. Neither cleared the bar convincingly.

Chengdong (Peter) ZhouUSC ’27NIH UGSP ScholarLos Angeles
Anhedonia classification accuracy, voice versus fMRIArea under the ROC curve for six classifiers, with 95 percent bootstrap confidence intervals, plotted against a chance line at 0.50. The three voice models, trained on acoustic features from clinical interview audio, score 0.61 to 0.65 and sit above chance. The three fMRI models, trained on nucleus accumbens activation, score 0.37 to 0.45 and sit at or below chance. The voice models therefore outperform the brain-imaging models on this benchmark.0.30.40.50.60.70.8CHANCEVOICERandom forest0.63Gradient boosted0.62Logistic regression0.56fMRILogistic regression0.58Random forest0.52Gradient boosted0.52

Voice cleared chance, barely (random forest, AUC 0.63, permutation p = .049). Ventral striatal BOLD did not (logistic regression, AUC 0.58, p = .057). The two are statistically non-inferior to each other, which is a statement about how modest both are.

Fig. — AUC-ROC with 95% bootstrap CI, stratified 5-fold CV. Primary preregistered analysis. Voice: eGeMAPSv02, diarized participant speech · DAIC-WOZ, n=142. fMRI: Nucleus accumbens BOLD, BART · ds000030, n=234. The voice result is borderline: p = .049 uncorrected, and it does not survive Bonferroni correction for three classifiers (α/3 = .017). Independent cohorts and different anhedonia instruments (PHQ-8 items 1–2 vs Chapman), so this is a benchmark across datasets, not a within-subject comparison. Both streams overfit heavily (train AUC 1.0, gap > 0.35) on 447 features and 32 positive cases. Stream B is also pipeline-dependent: fMRIPrep 23.x gives 0.45, below chance. Preregistered at osf.io/4d6ey.

Three labs, one question, three angles on it.

My father lost his vision to retinal detachment. He cannot fill out a visual analog scale, which is how most of psychiatry still measures how a person feels. I build instruments that do not require the patient to be a reliable narrator of their own symptoms.

ModelRead Lab · USC

A Rescorla-Wagner simulation of how reward learning collapses when the dopaminergic learning rate falls and effort stops being worth spending.

Model detail
SignalItti Lab · USC Viterbi

Whether that collapse is audible. An open-source pipeline that pulls acoustic biomarkers out of clinical interviews without the audio ever leaving the room.

Signal detail
CircuitNIMH · Experimental Therapeutics

What the circuit is doing while it happens. MEG source localization, gamma-band power, and signal complexity in the insula as a marker of suicidal ideation.

Circuit detail

What exists so far

  1. Jan 2026 – present

    ClinicalWhisper

    Local-first, air-gapped clinical interview transcription and speech-biomarker pipeline. Whisper for transcription, pyannote for diarisation, OpenSMILE for eGeMAPS features.

    ClinicalWhisper on GitHub
  2. 2026

    Cross-Modal Benchmarking of Acoustic Prosody and Ventral Striatal BOLD for Depression-Related Anhedonia Classification

    The two-stream benchmark above, written up: acoustic prosody against ventral striatal BOLD, preregistered and reported with its null.

    Read the preprint

Writing

  1. May 2026
    So here is the question nobody seems to be asking: what is the patient doing during those 72 hours?
    The Plasticity Paradox
  2. May 2026
    I tried an experiment last month. I opened Instagram and attempted to look at the first post on the screen without scrolling.
    The Economics of Attention
  3. March 2026
    The same density that makes the SGV feel like home also traps people inside it.
    The 626

All four essays