TorchCodec Documentation — FFmpeg-Backed Video Decoding to Tensors
Description
Meta's library docs for decoding video/audio directly into PyTorch tensors with CPU and CUDA/NVDEC paths, index- and PTS-based frame access, plus encoding APIs.
Resource details
- Category path
- Media Tools›AI & Machine Learning Tools
- Classification
- Not yet classified
TorchCodec Documentation — FFmpeg-Backed Video Decoding to Tensors is cataloged in Media Tools and AI & Machine Learning Tools.
URL
Related resources
- Deep Reinforced Bitrate Ladders (DeepLadder)NOSSDAV 2021 paper on deep-reinforcement-learning-based bitrate ladder construction for adaptive streaming.
- DQ-Ladder: Deep Q-Network Time/Quality-Aware Bitrate LadderDeep reinforcement learning (DQN) approach to bitrate ladder construction achieving 10.3%+ BD-rate gain and 22% decode-time reduction.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- Efficient Bitrate Ladder Construction for Content-Optimized Adaptive StreamingAcademic paper (Katsenou et al.) using ML to predict Pareto-optimal bitrate ladders from spatio-temporal content features.
- AI-Youtube-Shorts-GeneratorLLM-driven highlight detection with Whisper transcription and automatic vertical cropping to turn long videos into shorts, MIT-licensed wit…
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- aws-batch-with-FFmpegReference architecture running FFmpeg containers (ARM64/x86-64/NVIDIA/Xilinx) on AWS Batch with Spot compute environments and SDK/REST job…
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.
- 2BiVQA: Double Bi-LSTM No-Reference Video Quality AssessmentNo-reference VQA model for UGC video using double Bi-LSTM, companion code to arXiv:2208.14774.