
Automatic Transcription of Oral Sources
This workflow guides you through the automatic transcription of oral source recordings in SSH research contexts. It covers the full pipeline: audio preparation, model selection and configuration, transcription, speaker diarisation, quality control, manual post-correction, and the export of timestamped, structured transcripts suitable for archival deposit and open publication. The central tool is Whisper, a multilingual, open-source automatic speech recognition (ASR) model developed by OpenAI and released under the MIT licence. The workflow also integrates complementary tools for audio preprocessing (Audacity, FFmpeg), enhanced transcription (WhisperX, faster-whisper), annotation (ELAN), and text encoding (TEI-XML). It applies equally to newly recorded interviews and to the retrospective transcription of archival holdings.
Workflow steps(7)
1 Verify legal authorisation for transcription
2 Assess the recording for transcription readiness
3 Prepare Audio
4 Select and configure the ASR tool
5 Assess and correct the transcript
6 Encode the transcript in TEI-XML (optional. Recommended for archival deposit)
7 Complete the transcription documentation record
The SSH Open Marketplace is maintained and will be further developed by three European Research Infrastructures - DARIAH, CLARIN and CESSDA - and their national partners. It was developed as part of the "Social Sciences and Humanities Open Cloud" SSHOC project, European Union's Horizon 2020 project call H2020-INFRAEOSC-04-2018, grant agreement #823782.


