Cephalonauts One

A deep fMRI dataset for decoding naturalistic speech in the human brain.

90

hours of fMRI

3

subjects

30

hours per subject

520 ms

resolution

Brain recordings and speech

Explore a real 30-second recording from the dataset, with measured brain activity synchronized to the original podcast audio.

AUDIO
TRANSCRIPT
TRANSLATION
REAL fMRI DATA · DRAG TO ROTATE
00:0000:30

Dataset overview

Three native French speakers listened to podcasts across 72 scanning sessions. Stimuli were not repeated within participants.

3T whole-brain fMRI358 runsFrench podcasts
Train · 22Val · 1Test · 1

One example fold; each block is a session. The test session rotates across folds.

Audio segment retrieval

Retrieve a 10-second podcast segment from an fMRI frame, among approximately 2,000 candidates from a held-out session. The release includes session-level splits, evaluation metrics, and a neural-network baseline.

Decoding performance improves with additional training data, with no clear saturation over the range studied.

Brain recording→Decoder→Audio candidates
≋ Candidate 01
↗ Matching segment
Target
≋ Candidate 03

Task schematic

Scaling with more data

Retrieval improves as training recordings grow, across all three subjects.

sub-0sub-1sub-2
Top-10 retrieval accuracy (%) ↑0102030051015202530Training data (hours)sub-0 · 1.24 hours · 3.9 ± 1.0% · 1 sessionssub-0 · 3.74 hours · 9.5 ± 1.1% · 3 sessionssub-0 · 7.49 hours · 13.8 ± 1.3% · 6 sessionssub-0 · 12.47 hours · 17.6 ± 1.6% · 10 sessionssub-0 · 18.72 hours · 20.1 ± 2.3% · 15 sessionssub-0 · 27.46 hours · 23.2 ± 2.6% · 22 sessionssub-1 · 1.24 hours · 6.4 ± 1.3% · 1 sessionssub-1 · 3.74 hours · 12.8 ± 2.7% · 3 sessionssub-1 · 7.47 hours · 18.1 ± 2.8% · 6 sessionssub-1 · 12.48 hours · 21.5 ± 3.9% · 10 sessionssub-1 · 18.74 hours · 24.3 ± 3.5% · 15 sessionssub-1 · 27.76 hours · 27.3 ± 4.8% · 22 sessionssub-2 · 1.23 hours · 3.8 ± 1.3% · 1 sessionssub-2 · 3.73 hours · 8.7 ± 1.6% · 3 sessionssub-2 · 7.47 hours · 12.5 ± 2.0% · 6 sessionssub-2 · 12.46 hours · 16.9 ± 2.8% · 10 sessionssub-2 · 18.68 hours · 19.7 ± 4.2% · 15 sessionssub-2 · 26.64 hours · 22.3 ± 4.0% · 22 sessions
Mean ±1 SD across 5 folds · fsaverage5Download data ↗
View values
Top-10 accuracy · mean ±1 SD
SubjectSessionsHoursTop-10 accuracy
sub-011.243.9 ± 1.0%
sub-033.749.5 ± 1.1%
sub-067.4913.8 ± 1.3%
sub-01012.4717.6 ± 1.6%
sub-01518.7220.1 ± 2.3%
sub-02227.4623.2 ± 2.6%
sub-111.246.4 ± 1.3%
sub-133.7412.8 ± 2.7%
sub-167.4718.1 ± 2.8%
sub-11012.4821.5 ± 3.9%
sub-11518.7424.3 ± 3.5%
sub-12227.7627.3 ± 4.8%
sub-211.233.8 ± 1.3%
sub-233.738.7 ± 1.6%
sub-267.4712.5 ± 2.0%
sub-21012.4616.9 ± 2.8%
sub-21518.6819.7 ± 4.2%
sub-22226.6422.3 ± 4.0%

Access the dataset

Preprocessed fMRI in T1w, MNI and fsaverage5 spaces, with aligned audio, transcripts and embeddings.

T1wMNIfsaverage5
Access on Hugging Face ↗

Authentication and dataset access approval required.

Python · sub-1
from datasets import load_dataset

ds = load_dataset(
    "karavela/cephalonauts_one",
    name="sub-1_fsaverage5",
    split="train",
    streaming=True,
)