
Pick a preset, any preset. That one click—or more likely, that one line in an FFmpeg command—sets off a chain reaction. It balances visual quality against file size, and that balance shapes your storage bills, delivery bandwidth, and whether someone squints at a blocky mess or leans back and enjoys the show. I tune encoding pipelines for a living. I’ve dropped cloud bills by 30% just by changing a preset, while keeping perceptual quality at 98%. That’s not marketing fluff. That’s a Tuesday. This piece walks through the mechanics, the trade-offs, and the selection logic engineers and content teams actually need.
What a Preset Actually Controls
Think of a preset as a bundle of encoder knobs that trade compression smarts for speed. The names—ultrafast, veryfast, fast, medium, slow, veryslow—come straight from x264 and x265, and they map to how hard the CPU gets to think. Ultrafast skips most motion-estimation tricks and leans on simple quantization. Veryslow flips on exhaustive motion search, multiple reference frames, adaptive quantization, and a few other heavy hitters.
Each preset tugs on three threads: compression ratio, encoding time, and CPU/GPU load. A faster preset runs fewer analysis passes, so the encoder makes dumber guesses about where bits should go. A slower preset burns cycles studying motion vectors, scene cuts, and texture complexity. The payoff? A smaller file at the same quality, or better quality at the same bitrate.
One common mix-up: the preset does not set the bitrate directly. You still choose a target bitrate or a constant rate factor (CRF). The preset just changes how cleverly the encoder hits that target. With a slow preset and CRF 23, you’ll get a smaller file than ultrafast at CRF 23. Visual quality stays close because the CRF scale aims for constant perceptual quality. The file size shrinks because the encoder allocates bits more intelligently, not because it magically added detail.
CRF vs. Preset: Two Independent Levers

CRF sets the quality target. Lower numbers mean higher quality, bigger files. The preset controls how much computational sweat goes into reaching that quality efficiently. They interact, but they solve different problems. You might pick CRF 18 for a master archive and pair it with a slow preset to keep storage lean. For a live stream, you’d likely run CRF 23 with ultrafast because encoding delay matters more than a few megabytes.
This split matters hard for cost models. Storage costs scale with file size, which depends on both CRF and preset. Compute costs scale with encoding time, which depends almost entirely on the preset. In cloud transcoding pipelines, you pay for instance-hours. A veryslow encode on an AWS c5.4xlarge can run 10x longer than veryfast. That difference adds up fast when you’re pushing thousands of hours of content a month.
Quality Metrics That Matter
Visual quality is not a single tidy number. Three metrics show up in production: PSNR, SSIM, and VMAF. PSNR (Peak Signal-to-Noise Ratio) is quick and simple but doesn’t track what humans actually see. SSIM (Structural Similarity Index) tries to model how our eyes interpret structure. VMAF (Video Multimethod Assessment Fusion), from Netflix, blends several quality models and lines up better with subjective scores.
When I benchmark presets, I stare at VMAF scores at fixed bitrates. For a 1080p video at 5 Mbps, ultrafast might land at VMAF 85. Medium jumps to 92. Veryslow creeps to 94. The leap from ultrafast to medium is noticeable. From medium to veryslow, the gains shrink—often 1–2 VMAF points. Whether that matters depends on your audience. For user-generated clips, 85 might feel fine. For a premium SVOD service, every point above 90 matters.
Bitrate savings flip the story. To hit VMAF 93, ultrafast might demand 8 Mbps. Medium needs 5 Mbps. Veryslow needs 4.5 Mbps. That 3.5 Mbps gap per stream nearly halves CDN costs. Over a million monthly views, the savings get real. That’s why bigger platforms throw compute at slow presets for VOD libraries.
Perceptual Quality and Content Type
Content type tilts the quality-cost balance. Animation and screen recordings compress easily. Fast presets often work fine because flat regions and gentle motion don’t benefit much from exhaustive motion estimation. Sports and action films, with high motion and complex textures, gain more from slow presets. Grainy film footage punishes encoders. That grain looks like random noise, and fast presets treat it as detail worth keeping, which bloats the file. A slower preset with adaptive quantization can tease apart grain and signal, spending bits where they count.
For talking-head videos, facial detail rules. Slow presets hold skin texture and stop blocking around eyes and lips. Fast presets sometimes smear those areas at lower bitrates. If your platform hosts mostly webinars or tutorials, test with medium or slow. The storage savings from slower presets can offset the higher encoding compute cost once a video racks up views.
Cost Modeling: Compute vs. Distribution

The money question: balance one-time encoding compute against recurring storage and delivery costs. Let’s run numbers. Say you have 1,000 hours of source, encoded to three renditions (1080p, 720p, 480p). On a c5.4xlarge instance ($0.68/hour), veryfast encodes at 2x real-time. That’s 500 instance-hours, $340. Veryslow chugs at 0.2x real-time, chewing up 5,000 instance-hours and $3,400.
Now distribution. If veryslow files are 40% smaller, storage drops from 10 TB to 6 TB. At $0.023/GB/month, you save $92/month. CDN delivery at $0.01/GB for 50 TB/month saves $200/month. The $3,060 encoding premium breaks even in about 10 months. After that, you pocket $292/month. For long-lived content, slow presets win on math.
Live encoding flips the table. Latency rules. You can’t run veryslow for a live sports stream—the encode delay would stack seconds per frame. Ultrafast or veryfast are standard. The bitrate penalty stings less because the content disappears fast. Some live encoders lean on hardware acceleration (NVENC, QSV), trading a bit of compression efficiency for real-time speed. Those hardware presets sit near veryfast but with a lighter CPU footprint.
Segmenting Your Library by Value
Smart teams slice content by value. High-value catalog titles earn slow presets because they’ll generate views for years. Breaking news clips get fast presets—speed to publish outweighs storage perfection. This tiered approach squeezes the most from your budget. Implement it by tagging content at ingest: “evergreen” vs. “ephemeral.” Your pipeline picks the preset automatically based on the tag.
ABR (Adaptive Bitrate) ladder design also rubs shoulders with presets. Modern ladders use fewer renditions paired with optimized presets. Instead of 12 fixed bitrate steps, you might encode 5 renditions with per-title optimization and a slow preset. Fewer renditions slash storage and encode time; the slow preset keeps quality steady across the wider bitrate range.
Codec-Specific Preset Behavior
Not all codecs treat presets the same. x264’s presets are mature and well-charted. x265 (HEVC) presets shift the complexity range upward. x265 medium roughly matches x264 slow in CPU hunger and compression gain. H.264 encoding is snappy enough that slow presets are practical on modern hardware. H.265 encoding is 4–10x heavier, so the cost of slower presets bites harder.
AV1 encoders (libaom, SVT-AV1) have their own preset scales. SVT-AV1 presets range from 0 (slowest, best quality) to 13 (fastest). Preset 8 is a reasonable balance for production. AV1’s compression efficiency beats H.264 by about 30%, but encoding time spikes. For VOD, the storage/CDN savings can justify slow AV1 presets. For live, real-time AV1 encoding is still finding its feet, mostly locked to faster presets.
Hardware encoders in GPUs and ASICs offer preset-like tuning, too. NVENC has presets from P1 (fastest) to P7 (slowest). The quality gap between P1 and P7 is narrower than between x264 ultrafast and veryslow. Hardware encoders swap flexibility for raw speed. They’re great for live streaming and real-time transcoding, less so for archival encoding where every byte counts.
Practical Testing Methodology
Don’t just swallow published benchmarks. Your content is its own beast. Run your own tests. Grab 10 representative clips from your library—different genres, motion levels, resolutions. Encode each with 3–4 presets at your target CRF. Measure VMAF, bitrate, and encode time. Plot the results. Hunt for the knee in the curve where slower presets give you less and less.
For a corporate training platform I advised, the sweet spot landed on x264 medium. Slow squeezed only 2% more bitrate reduction but took 3x longer. Their content was mostly slides plus a talking head. The extra compute cost never paid back. For a film archive project, veryslow made sense—those files would sit for decades and stream at high volume.
Automate the grunt work with FFmpeg. A simple script loops over presets and logs VMAF. Use the libvmaf filter. The pattern: set CRF, set preset, encode, run VMAF against the source. Dump results into a CSV. That data becomes your preset policy’s backbone.
FAQ
Does a slower preset always give better quality?
Not exactly. At the same bitrate, slower presets usually produce higher quality. But if you fix quality with CRF, slower presets just shrink files, not necessarily make them prettier. The CRF scale normalizes quality. The win is storage and bandwidth savings, not a visible quality leap.
How do I choose between x264 and x265 presets for 4K content?
For 4K, x265’s compression edge is big—often 40–50% bitrate savings over x264. But x265 slow can run 10x slower than x264 medium. If your library is huge and streaming 4K is a major cost, x265 veryslow on a CPU cluster may pay off. For small libraries or quick turnaround, x264 medium is pragmatic. Test with your actual 4K footage to see if the encode time fits your pipeline.
Can I change presets mid-stream or on existing files?
Nope. Presets bake in at encoding time. You can’t tweak an encoded file’s preset without re-encoding from the original source. For adaptive streaming, you can mix presets across renditions—slow for 1080p, fast for lower resolutions—but each rendition is a separate encode.
What preset should I use for user-generated content platforms?
User-generated content (UGC) lands once and gets watched many times. Start with x264 medium at CRF 22. It balances quality and encode time. If your volume hits millions of uploads a day, consider veryfast to keep the encoding queue from choking. The storage savings from slower presets might get outweighed by the need to process uploads fast. Watch your queue depth and adjust.