Skip to content

Can a five-minute voice recording detect anhedonia as accurately as a $500 brain scan?

Chengdong (Peter) ZhouUSC ’27NIH UGSP ScholarLos Angeles
Summary

In our preregistered benchmark, voice reached 0.63 AUC and the scan 0.58. Voice cleared chance by a narrow margin and the scan did not.1

AUC-ROC with 95 percent bootstrap confidence interval for each classifier, against a chance level of 0.50. The best voice model reached 0.63 (permutation p = .049, not significant after Bonferroni correction) and the best fMRI model 0.58 (p = .057). Both are modest and neither dominates the other.
StreamClassifierAUC95% CI
VOICERandom forest0.630.51 to 0.74
VOICEGradient boosted0.620.50 to 0.73
VOICELogistic regression0.560.45 to 0.67
fMRILogistic regression0.580.51 to 0.66
fMRIRandom forest0.520.45 to 0.60
fMRIGradient boosted0.520.45 to 0.60
Fig. 1 RF random forest · GB gradient boosted · LR logistic regression. AUC-ROC with 95% bootstrap CI, stratified 5-fold CV. Primary preregistered analysis, osf.io/bsvrj.
Datasets, methods and remaining caveat

Voice: eGeMAPSv02, diarized participant speech · DAIC-WOZ, n=142. fMRI: Nucleus accumbens BOLD, BART · ds000030, n=234 of 272 screened. Both streams overfit heavily (train AUC 1.0, gap > 0.35) on 447 features and 32 positive cases. Stream B is also pipeline-dependent: fMRIPrep 23.x gives 0.45, below chance.

Three labs, one question, three angles on it.

Psychiatry still measures how people feel mostly by asking them, often on forms a blind patient cannot fill out. I work on measures that lean less on the patient narrating their own symptoms, starting with the voice.

ModelRead Lab · USC

A Rescorla-Wagner simulation of how anhedonia breaks reward learning: choices that stop following learned value cost far more than slower learning.

Read the model work
SignalItti Lab · USC Viterbi

Whether that collapse is audible. An open-source pipeline that pulls acoustic biomarkers out of clinical interviews without the audio ever leaving the room.

Read the signal work
CircuitNIMH · Experimental Therapeutics

What the circuit is doing while it happens. MEG source localization and Lempel-Ziv signal complexity in the insula, explored as a correlate of suicidal thoughts.

Read the circuit work

Selected work

  1. Zhou, C., Wu, M., Xiang, Y., & Itti, L. (2026). Cross-Modal Benchmarking of Acoustic Prosody and Ventral Striatal BOLD for Depression-Related Anhedonia Classification: A Pre-Registered Study with the ClinicalWhisper Pipeline. bioRxiv preprint, not peer reviewed.DOIBibTeXOSF
  2. Zhou, C. (2026). ClinicalWhisper: a local-first pipeline for clinical interview audio. Open-source software, MIT licence. Zenodo 10.5281/zenodo.20559786.GitHubDOI
  3. Zhou, C. (2026). Computational Psychiatry: An Undergraduate Syllabus. 20 papers, 5 modules, CC BY 4.0. Zenodo 10.5281/zenodo.20559875.GitHubDOI

Writing

  1. May 2026
    So here is the question nobody seems to be asking: what is the patient doing during those 72 hours?
    The Plasticity Paradox
  2. May 2026
    I tried an experiment last month. I opened Instagram and attempted to look at the first post on the screen without scrolling.
    The Economics of Attention
  3. March 2026
    The same density that makes the SGV feel like home also traps people inside it.
    The 626

All four essays

  1. 1Permutation p = .049 for voice (random forest) and .057 for fMRI (logistic regression), both uncorrected. The 95% intervals overlap: both results are modest, and neither dominates the other. ↩