Cephalonauts One
A deep fMRI dataset for decoding naturalistic speech in the human brain.
hours of fMRI
subjects
hours per subject
resolution
Brain recordings and speech
Explore a real 30-second recording from the dataset, with measured brain activity synchronized to the original podcast audio.
Dataset overview
Three native French speakers listened to podcasts across 72 scanning sessions. Stimuli were not repeated within participants.
One example fold; each block is a session. The test session rotates across folds.
Audio segment retrieval
Retrieve a 10-second podcast segment from an fMRI frame, among approximately 2,000 candidates from a held-out session. The release includes session-level splits, evaluation metrics, and a neural-network baseline.
Decoding performance improves with additional training data, with no clear saturation over the range studied.
Task schematic
Scaling with more data
Retrieval improves as training recordings grow, across all three subjects.
View values
| Subject | Sessions | Hours | Top-10 accuracy |
|---|---|---|---|
| sub-0 | 1 | 1.24 | 3.9 ± 1.0% |
| sub-0 | 3 | 3.74 | 9.5 ± 1.1% |
| sub-0 | 6 | 7.49 | 13.8 ± 1.3% |
| sub-0 | 10 | 12.47 | 17.6 ± 1.6% |
| sub-0 | 15 | 18.72 | 20.1 ± 2.3% |
| sub-0 | 22 | 27.46 | 23.2 ± 2.6% |
| sub-1 | 1 | 1.24 | 6.4 ± 1.3% |
| sub-1 | 3 | 3.74 | 12.8 ± 2.7% |
| sub-1 | 6 | 7.47 | 18.1 ± 2.8% |
| sub-1 | 10 | 12.48 | 21.5 ± 3.9% |
| sub-1 | 15 | 18.74 | 24.3 ± 3.5% |
| sub-1 | 22 | 27.76 | 27.3 ± 4.8% |
| sub-2 | 1 | 1.23 | 3.8 ± 1.3% |
| sub-2 | 3 | 3.73 | 8.7 ± 1.6% |
| sub-2 | 6 | 7.47 | 12.5 ± 2.0% |
| sub-2 | 10 | 12.46 | 16.9 ± 2.8% |
| sub-2 | 15 | 18.68 | 19.7 ± 4.2% |
| sub-2 | 22 | 26.64 | 22.3 ± 4.0% |
Access the dataset
Preprocessed fMRI in T1w, MNI and fsaverage5 spaces, with aligned audio, transcripts and embeddings.
Authentication and dataset access approval required.
from datasets import load_dataset
ds = load_dataset(
"karavela/cephalonauts_one",
name="sub-1_fsaverage5",
split="train",
streaming=True,
)