Skip to main content
Home

Automatic Transcription of Oral Sources

This workflow guides you through the automatic transcription of oral source recordings in SSH research contexts. It covers the full pipeline: audio preparation, model selection and configuration, transcription, speaker diarisation, quality control, manual post-correction, and the export of timestamped, structured transcripts suitable for archival deposit and open publication. The central tool is Whisper, a multilingual, open-source automatic speech recognition (ASR) model developed by OpenAI and released under the MIT licence. The workflow also integrates complementary tools for audio preprocessing (Audacity, FFmpeg), enhanced transcription (WhisperX, faster-whisper), annotation (ELAN), and text encoding (TEI-XML). It applies equally to newly recorded interviews and to the retrospective transcription of archival holdings.

Workflow steps(7)

  1. 1 Verify legal authorisation for transcription

  2. 2 Assess the recording for transcription readiness

  3. 3 Prepare Audio

  4. 4 Select and configure the ASR tool

  5. 5 Assess and correct the transcript

  6. 6 Encode the transcript in TEI-XML (optional. Recommended for archival deposit)

  7. 7 Complete the transcription documentation record

European Union flag

The SSH Open Marketplace is maintained and will be further developed by three European Research Infrastructures - DARIAH, CLARIN and CESSDA - and their national partners. It was developed as part of the "Social Sciences and Humanities Open Cloud" SSHOC project, European Union's Horizon 2020 project call H2020-INFRAEOSC-04-2018, grant agreement #823782.

CESSDACLARINDARIAH-EU