Skip to main content

Speech-to-text transcription

Brazil - Chamber of Deputies

Use case ID: 022

Author: Chamber of Deputies of Brazil

Date: 14 June 2024

Objective:

Transcribe recorded audio or video files into text, which is indispensable for supporting various parliamentary functions such as maintaining official records of parliamentary speeches, and transcribing public hearings, meetings, interviews and other proceedings.

Actors:

  • Stenographers
  • Audio operators or clerk staff
  • Communication unit staff
  • AI system proficient in speech-to-text transcription

Prerequisites:

  • AI model trained for speech-to-text transcription, optionally integrated with diarization and speaker recognition services
  • Rules followed by stenographers
  • Integration with digital services commonly used by users, such as stenographers, audio operators, clerk staff and communication unit staff

Scenario:

  1. When a new parliamentary meeting is recorded in audio or video, the AI system transcribes its contents into text.
  2. The audio operators or clerk staff use the appropriate digital service to ensure accurate recording of speaker identification, time codes and other metadata for the captured speeches.
  3. The AI system combines the transcribed text with the recorded metadata (e.g. speaker identification) to generate a preliminary diarized transcript.
  4. Stenographers review and correct portions of the preliminary diarized transcript to produce the official record of the parliamentary meeting for publication.

Alternate flows:

  • When a new interview is recorded in audio or video, the AI system transcribes its contents into text. Communication unit staff can then revise and correct the resulting transcript before publication or distribution.

Expected results:

  • The process of transcribing speech to text is more efficient.
  • Workload for stenographers and other potential users is reduced.

Potential challenges:

  • Ensuring the AI model accurately transcribes speeches involves recognizing acronyms and idioms, and adhering to official recording rules
  • Detecting potential deterioration in the performance of the AI system over time

Data requirements:

  • Historically recorded speech texts can be used to enhance the performance of AI models
  • Periodic verification of AI ​​model performance

Integrations with other systems:

  • Digital services used by the users
  • Diarization service
  • Speaker recognition service
  • Analytics and reporting tools

Success metrics:

  • Word error rate (WER)
  • Recognition time
  • Speaker diarization accuracy
  • Volume of transcripts assisted by the AI system

 

The Use cases for AI in parliaments collection is published by the IPU’s Centre for Innovation in Parliament as part of the Parliamentary Data Science Hub’s project to create guidelines for AI governance in parliaments.

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International licence. It may be freely shared and reused with acknowledgement of the author and the IPU. 

A use case describes how a system should work. It is used to plan, develop and measure implementation. A use case is not the same as a case study, which is a descriptive text of an actual project’s implementation. Please note that this use case is provided “as is” and neither the IPU nor the author accepts any responsibility for its use.

For more information about the IPU’s work on artificial intelligence, please visit www.ipu.org/AI or contact [email protected]