Speech, vision & media
WhisperX
Adds forced word alignment and speaker diarization around Whisper transcription.
Overview
Research summary
WhisperX combines speech recognition with alignment and speaker diarization components to create timestamped transcripts. It uses faster-whisper for transcription, phoneme-based alignment models to improve word timing, and pyannote tooling when speaker labels are needed. Developers can use the command-line workflow or Python functions to process audio and assemble aligned outputs.
This is useful for searchable recordings, subtitles and meeting-analysis pipelines where approximate segment timestamps are insufficient. Alignment support depends on language-specific models, and speaker diarization requires its own model access and configuration. The pipeline code uses BSD-2-Clause, with external models governed by their respective terms.
Repository summary
- Stars
- 24,335
- Open issues
- Unavailable
- Last push
- 2026-09-26
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
Implementation
Recorded implementation details and interfaces for WhisperX.
Implementation details
Python pipeline using faster-whisper, alignment models and pyannote.audio
- Languages
- Python
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
- Stars
- 24,335
- Open issues
- Unavailable
- Last push
- 2026-09-26
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
- Stars source
- Recorded fallback
- Last push source
- Recorded fallback
Documentation
Recorded references and research provenance for this entry.
Recorded sources 3
- https://github.com/m-bain/whisperX Project page · Linked repository · Research reference
- https://github.com/m-bain/whisperX/blob/main/README.md Research reference
- https://github.com/m-bain/whisperX/blob/main/LICENSE Research reference
Research metadata
- Research date
- 2026-10-02
- Catalogue snapshot
- 2026-10-06