VideoCaptioner
VideoCaptioner is a video subtitle processing assistant based on large language models (LLM). It supports speech recognition, subtitle segmentation, optimization, and translation, providing a comprehensive solution for subtitle generation and integration. The tool offers features
Link
Related resources
- Eyevinn auto-subtitlesWhisper-based automatic subtitle generation tool that chunks large audio and outputs VTT/SRT/JSON while preserving sync.
- subtitle-translatorBatch subtitle translation tool supporting .srt/.ass/.vtt/.lrc formats with 17+ LLM providers, bundled format conversion, actively maintain…
- alass — Automatic Language-Agnostic Subtitle SynchronizationRust CLI tool for automatic subtitle sync handling offset drift, ad-break splits, and framerate differences, 88-98% good-sync rate.
- SubalignerAutomatic subtitle synchronization tool using deep neural networks and forced alignment, with translation support.
- whisper-subtitlesAccessibility-focused local subtitle generation using Whisper+OpenVINO, targeting hearing-impaired users across 99 languages.
- Bitrate Ladder Construction via Transfer Learning and Spatio-Temporal FeaturesML-based bitrate ladder optimization achieving 94.1% complexity reduction at only 1.71% BD-rate cost using transfer learning.
- VMAF-torchPyTorch reimplementation of VMAF metric, GPU/gradient-friendly for use in learned codec optimization loops.
- CompressedVQA: Full/No-Reference VQA for Compressed UGC VideoICME 2021 grand challenge-winning full-reference and no-reference video quality models for compressed user-generated content.
- aws-batch-with-FFmpegReference architecture running FFmpeg containers (ARM64/x86-64/NVIDIA/Xilinx) on AWS Batch with Spot compute environments and SDK/REST job…
- Evaluating Video Quality Metrics for Neural and Traditional Codecs (4K/UHD-1)Academic comparison of VMAF, AVQBits, and FasterVQA no-reference quality metrics on 4K content.