WhisperX
Description
ASR pipeline producing word-level timestamps via wav2vec2 forced alignment plus speaker diarization, useful for accurate subtitle cue timing.
Resource details
- Category path
- Media Tools›Audio & Subtitles›Subtitles & Captions
- Provider
- Not yet classified
- Format
- Not yet classified
- Skill level
- Not yet classified
WhisperX is cataloged in Media Tools, Audio & Subtitles, and Subtitles & Captions. Provider: Not yet classified. Format: Not yet classified. Skill level: Not yet classified.
URL
Related resources
- SubalignerAutomatic subtitle synchronization tool using deep neural networks and forced alignment, with translation support.
- Normalize-AudioBatch EBU R128 loudness normalization tool for MKV files via ffmpeg, preserves video stream and caches loudness analysis.
- mkv2mp4Cross-platform Python GUI for batch MKV→MP4 conversion, with fast remux when streams are compatible and auto-transcode fallback plus SRT→mo…
- A Gamut-Mapping Framework for Color-Accurate Reproduction of HDR ImagesAcademic paper (Sikudova et al.) proposing a gamut-mapping framework for color-accurate HDR image reproduction.
- SceneSegmentation-SCRLOfficial PyTorch implementation of CVPR 2022 Scene Consistency Representation Learning for movie-level scene segmentation (MovieNet/SceneSe…
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.
- 2BiVQA: Double Bi-LSTM No-Reference Video Quality AssessmentNo-reference VQA model for UGC video using double Bi-LSTM, companion code to arXiv:2208.14774.