VideoReTalking
A research project (SIGGRAPH Asia 2022) providing code for audio-driven lip synchronization in talking head videos. It enables realistic re-sync of lip movements to new audio on existing footage.
Link
Related resources
- SparseSync: Audio-Visual Synchronisation with Trainable SelectorsBMVC 2022 Spotlight paper and code for sparse-in-space-and-time audio-visual sync detection.
- Eyevinn auto-subtitlesWhisper-based automatic subtitle generation tool that chunks large audio and outputs VTT/SRT/JSON while preserving sync.
- alass — Automatic Language-Agnostic Subtitle SynchronizationRust CLI tool for automatic subtitle sync handling offset drift, ad-break splits, and framerate differences, 88-98% good-sync rate.
- KVQ: NTIRE 2024 Short-form UGC Video Quality Assessment ChallengeCVPR NTIRE 2024 challenge dataset and code for short-form UGC video quality assessment, 600 videos/3600 processed clips.
- AI-Youtube-Shorts-GeneratorLLM-driven highlight detection with Whisper transcription and automatic vertical cropping to turn long videos into shorts, MIT-licensed wit…
- 2BiVQA: Double Bi-LSTM No-Reference Video Quality AssessmentNo-reference VQA model for UGC video using double Bi-LSTM, companion code to arXiv:2208.14774.
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.