ffmpeg-hardware-acceleration

v2026.09.24

Complete GPU-accelerated encoding/decoding system for FFmpeg 7.1 LTS and 8.0.1 (latest stable, released 2025-11-20). PROACTIVELY activate for: (1) NVIDIA NVENC/NVDEC encoding, (2) Intel Quick Sync Video (QSV), (3) AMD AMF encoding, (4) Apple VideoToolbox, (5) Linux VAAPI setup, (6) Vulkan Video 8.0 (FFv1, AV1, VP9, ProRes RAW), (7) VVC/H.266 hardware decoding (VAAPI/QSV), (8) GPU pipeline optimization with pad_cuda, (9) Docker GPU containers, (10) Performance benchmarking. Provides: Platform-specific commands, preset comparisons, quality tuning, full GPU pipeline examples, Vulkan compute codecs, VVC decoding, troubleshooting guides. Ensures: Maximum encoding speed with optimal quality using GPU acceleration.

GitHub
Install command
npx skhub add josiahsiegel/ffmpeg-hardware-acceleration
Markdown
SKILL.md

When to Use This Skill

Activate when GPU acceleration is needed:

  • Encoding speed is critical (10-30x faster than CPU)
  • Processing large batches of videos
  • Real-time encoding for streaming
  • Server-side transcoding at scale
  • Docker containers with GPU passthrough

GPU encoding trades some quality for massive speed. Use -cq, -qp, or -global_quality for quality control.

Quick Reference

PlatformEncoderDecoderDetect Command
NVIDIAh264_nvenc, hevc_nvenc, av1_nvench264_cuvid, hevc_cuvidffmpeg -encoders | grep nvenc
Intel QSVh264_qsv, hevc_qsv, av1_qsvh264_qsv, hevc_qsvffmpeg -encoders | grep qsv
AMD AMFh264_amf, hevc_amf, av1_amfN/A (use software)ffmpeg -encoders | grep amf
Appleh264_videotoolbox, hevc_videotoolboxh264_videotoolboxmacOS only
VAAPIh264_vaapi, hevc_vaapi, av1_vaapiwith -hwaccel vaapiLinux only
Vulkanh264_vulkan, hevc_vulkan, av1_vulkan, ffv1_vulkanVP9, ProRes RAW (8.0+)ffmpeg -encoders | grep vulkan

Current Latest: FFmpeg 8.0.1 (released 2025-11-20). Check with ffmpeg -version.

Hardware Acceleration Overview

Hardware acceleration uses dedicated GPU/SoC components for video processing:

  • NVENC/NVDEC (NVIDIA): dedicated encode/decode engines
  • QSV (Intel): Quick Sync Video on Intel CPUs with integrated graphics
  • AMF (AMD): Advanced Media Framework for AMD GPUs
  • VideoToolbox (Apple): macOS/iOS hardware acceleration
  • VAAPI (Linux): Video Acceleration API (Intel, AMD on Linux)
  • Vulkan Video (Cross-platform, FFmpeg 7.1+/8.0): compute-shader-based codecs

Performance Comparison (2025 Benchmarks)

MethodSpeedQualityPowerUse Case
libx264 (CPU)1xBestHighQuality-critical
libx265 (CPU)0.3xBestVery HighArchival
h264_nvenc10-20xGoodLowReal-time, streaming
hevc_nvenc8-15xGoodLow4K streaming
h264_qsv8-15xGoodVery LowLaptop, efficiency
h264_amf8-15xGoodLowAMD systems

Core Workflow

  1. Detect what you have: ffmpeg -hwaccels, ffmpeg -encoders | grep <api>
  2. Pick a backend: prefer the vendor-native one (NVENC on NVIDIA, QSV on Intel, AMF on AMD, VideoToolbox on Apple). Use Vulkan for cross-platform portability.
  3. Build a full-GPU pipeline where possible:
    • Add -hwaccel <api> and -hwaccel_output_format <api> before -i
    • Use GPU-side filters (scale_cuda, scale_vulkan, vpp_qsv, ...)
    • Use the matching hardware encoder
  4. Tune quality: -preset (NVENC p1-p7), -cq/-qp/-global_quality, lookahead, spatial/temporal AQ
  5. Verify: benchmark with ffmpeg -benchmark and monitor GPU via nvidia-smi dmon, intel_gpu_top, etc.

Minimal Examples per Backend

# NVIDIA NVENC
ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i input.mp4 \
  -c:v h264_nvenc -preset p4 -b:v 5M output.mp4

# Intel QSV
ffmpeg -hwaccel qsv -hwaccel_output_format qsv -i input.mp4 \
  -c:v h264_qsv -preset medium -b:v 5M output.mp4

# AMD AMF
ffmpeg -i input.mp4 -c:v h264_amf -quality balanced -b:v 5M output.mp4

# Apple VideoToolbox
ffmpeg -i input.mp4 -c:v h264_videotoolbox -b:v 5M output.mp4

# Linux VAAPI
ffmpeg -hwaccel vaapi -hwaccel_device /dev/dri/renderD128 \
  -hwaccel_output_format vaapi -i input.mp4 \
  -c:v h264_vaapi -b:v 5M output.mp4

# Vulkan (cross-platform)
ffmpeg -init_hw_device vulkan -i input.mp4 \
  -c:v h264_vulkan -b:v 5M output.mp4

Full GPU Pipeline Pattern (Critical)

ffmpeg -y -vsync 0 \
  -hwaccel cuda -hwaccel_output_format cuda \
  -i input.mp4 \
  -vf scale_cuda=1280:720 \
  -c:v h264_nvenc -preset p4 -b:v 5M \
  -c:a copy \
  output.mp4

Omitting -hwaccel_output_format can cut throughput by up to 50% because decoded frames silently round-trip through CPU memory. See references/gpu-memory-and-troubleshooting.md for memory flow diagrams and best practices.

Cross-Vendor Filter Quick Map

OperationNVIDIAIntelAMD/LinuxCross-platform
Scalescale_cuda, scale_nppvpp_qsv, scale_qsvscale_vaapiscale_vulkan, scale_opencl, libplacebo
Overlayoverlay_cuda--overlay_vulkan, overlay_opencl
Deinterlacebwdif_cudavpp_qsvdeinterlace_vaapibwdif_vulkan
Denoisebilateral_cuda--nlmeans_vulkan, nlmeans_opencl
Chromakeychromakey_cuda--colorkey_opencl
Tonemap (HDR->SDR)--tonemap_vaapilibplacebo, tonemap_opencl
Pad/letterboxpad_cuda (8.0+)--pad_opencl

Use-Case Quick Recipes

Live streaming (low latency)

ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i input \
  -c:v h264_nvenc -preset p3 -tune ll -zerolatency 1 -b:v 6M \
  -f flv rtmp://server/live/stream

VOD (quality target)

ffmpeg -i input.mp4 \
  -c:v hevc_nvenc -preset p6 -tune hq \
  -rc vbr -cq 22 -b:v 0 \
  -rc-lookahead 32 -spatial-aq 1 \
  output.mp4

Batch / parallel encoding

ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i input1.mp4 -c:v h264_nvenc output1.mp4 &
ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i input2.mp4 -c:v h264_nvenc output2.mp4 &
wait

Docker (GPU passthrough)

docker run --gpus all --rm -v $(pwd):/data \
  jrottenberg/ffmpeg:nvidia \
  -hwaccel cuda -hwaccel_output_format cuda \
  -i /data/input.mp4 -c:v h264_nvenc /data/output.mp4

Best Practices

  1. Use full GPU pipelines when possible to avoid CPU/GPU memory transfers
  2. Match decode and encode hardware for best performance
  3. Pick presets deliberately - faster is not always better for quality
  4. Enable lookahead and AQ for quality-critical encodes
  5. Test on target hardware - quality varies by GPU generation
  6. Monitor GPU memory for high-resolution content
  7. Consider power efficiency for laptops and servers
  8. Update drivers regularly for performance and feature improvements

Reference Map

For deep dives on each backend, see:

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

plugins/ffmpeg-core/skills/ffmpeg-hardware-acceleration

Default branch

main

Latest commit

5a1b112

Tree SHA

376c8e0