Wav2Lip
An AI model that achieves accurate lip-syncing in videos. Given an input video of a person and a target speech audio, Wav2Lip generates a video where the person's lip movements match the audio perfectly.
Link
Related resources
- spectreGo audio fingerprinting library built specifically for syncing movie subtitles to audio tracks.
- VQA² (Visual Question Answering for Video Quality Assessment)ACM MM 2025 paper and code for VQA²-Assistant, a unified video/image quality scoring and interpretation model.
- VideoTaggingToolAngular.js web app for managing users, videos, and tagging jobs with an extensible metadata schema for video annotation.
- A Gamut-Mapping Framework for Color-Accurate Reproduction of HDR ImagesAcademic paper (Sikudova et al.) proposing a gamut-mapping framework for color-accurate HDR image reproduction.
- audioprintPython MFCC-based perceptual-hash audio fingerprinting library.
- 2BiVQA: Double Bi-LSTM No-Reference Video Quality AssessmentNo-reference VQA model for UGC video using double Bi-LSTM, companion code to arXiv:2208.14774.
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.