TorchAudio
TorchAudio provides a set of tools and building blocks for audio and speech processing, designed to accelerate the development and deployment of machine learning applications in these domains. It offers GPU-compatible, differentiable, and production-ready components, making it va
Link
Related resources
- 5G QoE Prediction: ML-based Quality-Shift Prediction for Video StreamingMachine-learning quality-shift prediction for video streaming over 5G networks, includes 1-sec granularity CLM/YouTube QoE dataset, compani…
- Deep Reinforced Bitrate Ladders (DeepLadder)NOSSDAV 2021 paper on deep-reinforcement-learning-based bitrate ladder construction for adaptive streaming.
- ComfyUI-VideoColorGradingComfyUI node pack for video color grading, integrating into node-based AI compositing workflows.
- VQEG Software ToolsVideo Quality Experts Group's curated index of academic/industry video quality assessment software tools and research collaboration hub.
- audioprintPython MFCC-based perceptual-hash audio fingerprinting library.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.
- 2BiVQA: Double Bi-LSTM No-Reference Video Quality AssessmentNo-reference VQA model for UGC video using double Bi-LSTM, companion code to arXiv:2208.14774.