Skip to content
transcodelyengineeringpostmortemper-title encodingVMAF

The Day Per-Title Shipped Twice: A libvmaf Timestamp Postmortem

9 min read Dimitar Todorov

Per-title encoding went live at 01:05 UTC, was withdrawn at 03:55 after the first real-worker run measured the search producing a larger file, was fixed in the worker, proven on real machines, and re-enabled at 09:25. The root cause was two timebases disagreeing every third frame. Here is the whole day, the arithmetic, and the three reasons CI could not have caught it.

Per-title encoding is live, and the announcement mentioned that it took two attempts on the same morning. This is the account of that morning. It is here because the interesting part is not that a bug shipped; it is what the bug measured, why a passing integration test was passing, and what kept a broken search from ever reaching an invoice.

All times are UTC, on 15 September 2026.

Timeline

Time Event
01:05 API 5.20.0 live. content_aware.mode = per_title accepted at CreateJob for the first time, on worker 1.28.0.
01:16 to 02:18 First real-worker validation run. Both real-encode per-title cases report an achieved VMAF of 77.226568. The same number, for two different sources with two different targets. The CRF is pinned at the search floor after two iterations, and one output is at 174.1% of the bitrate of its un-searched twin. The cases keep their known-gap tags; per-title is not promoted.
03:55 API 5.20.1 live. per_title withheld again with parameter_unsupported. The reporting plumbing, the pricing rule and the migration stay.
05:54 A complete regression run against live 5.20.1 confirms the withdrawal behaves as designed and introduces no collateral failures.
06:01 Root cause found and reproduced against the worker’s own pinned ffmpeg build, with the fix applied and then reverted to confirm.
07:30 Worker 1.28.1 released and pinned.
07:30 to 07:59 Validated on real workers. Six of six per-title cases pass with distinct, real numbers.
09:25 API 5.21.0 live. per_title accepted again.

Per-title was reachable by customers for roughly two hours and fifty minutes. No customer job is on record as having used it in that window. The failure was found by the validation harness, not by a customer.

What the number was measuring

The per-title search encodes short samples of the source at several CRFs and asks libvmaf to score each one against the untouched original. libvmaf pairs its two inputs through ffmpeg’s framesync, which matches frames by timestamp. The two files the search handed it never carried the same timestamps.

The reference cut was written to Matroska, whose timebase is one millisecond. A 30 fps cut is stamped 33, 67, 100, 133 ms. The sample encode was written to MP4 on an exact frame-rate timebase and stamped 33.333, 66.667, 100.0 ms. The two grids agree only every third frame. Two frames in three were scored against their neighbour rather than against themselves, and libvmaf emitted one comparison more than either file has frames: 451 comparisons for 450 frames, a period-three beat running through the per-frame scores.

The score that came back was a measurement of the misalignment, not of the encode. On the validation suite’s 150-second 1080p fixture, a CRF 18 sample read 77.23 where its true score is 98.53. Every probe then looked equally bad, so the search walked to its CRF floor after two iterations and delivered a file at 174.1% of the un-searched twin’s bitrate. A bigger file, at a 1.5x multiplier, that was in fact worse.

The tell was the identical score. Two sources, two targets, one number to six decimal places. A search that is merely mistuned produces different wrong answers for different inputs. A search that is measuring something other than the encode produces the same one.

Why the existing integration test passed

The severity is phase-dependent. A sample window whose first frame does not land on zero shifts which side of the beat the comparison falls on. That is how the executor’s per-title integration test passed at VMAF 93.04 on a 30-second fixture while the 150-second fixture in the validation suite collapsed. The test was not wrong about what it measured. It was lucky about where its window started.

The fix

Worker 1.28.1 makes two changes.

Both branches of the filtergraph are put on a frame-index timeline (settb=AVTB,setpts=N) before libvmaf sees them. That is not an approximation here: the distorted file is the reference re-encoded, frame for frame, with the reference already resampled to the variant’s frame rate for exactly this reason. Index pairing is the correct pairing.

A frame-count guard refuses to score a sample whose packet count differs from its reference. Index pairing mis-scores silently after a dropped frame, so the mismatch raises an error that routes to the existing static-ladder fallback. That fallback is not billed at the per-title multiplier.

Measured on the same 150-second fixture in the worker’s own ffmpeg build:

CRF chosen Achieved VMAF Probes Bitrate vs the flat-CRF twin
Worker 1.28.0 18 77.227 2 174.1%
Worker 1.28.1 27 93.054 3 63.1%

The new tests do not depend on the phase luck the old one did. The central one compares a pixel-identical Matroska and MP4 pair against the reference scored against itself, so it needs no magic constant: 95.10 against 97.76 before the fix, 97.7638 against 97.7638 after. A second requires strictly falling scores across CRF 18, 28 and 38, which is precisely what two targets reporting one identical 77.226568 failed to produce. A third drives the real search and refuses the floor when the floor already clears the target. A fourth is a cheap argument-level guard that runs without libvmaf at all.

The real-worker run on 1.28.1 then measured 93.043 at CRF 27 and 63.5% of the flat bitrate on the same case, and 90.426 at CRF 29 on a ladder asking for a non-default target of 90.

Why nobody was charged

An output is billed at the per-title multiplier only when an analysis actually chose its CRF. Without one, it bills at 1.0x. That guard was written before the feature could run, for the boring case of a worker too old to perform the search, and it did exactly what it exists for on the morning the search was wrong.

So the exposure was the file, never the invoice: a customer who asked for per-title in that window would have received a larger file at the same request, at the ordinary price. Nobody did, but the guard is why the withdrawal cost nothing and why re-accepting the mode was safe.

Why CI did not catch it

The tests that would have caught this exist. They are a manual gate, and three separate things keep them out of CI. All three have to change before that stops being true.

  1. The worker’s CI test step runs go test -race ./... without -tags=integration, so the libvmaf integration file is not even compiled there.
  2. The Docker integration target names six packages (ffmpeg, ffprobe, packager, chunker, executor and thumbnailer) and does not list the analyzer package, which is where the scorer and its integration tests live. The job’s own path filter omits the same directory, so a change confined to the analyzer does not even trigger the job.
  3. The Docker test stage installs Debian’s ffmpeg, which has no libvmaf. The tests’ guard therefore skips rather than fails. A green run there is not evidence that anything was measured.

The unit tests that do run assert on argument strings. A plausible-but-wrong filtergraph paired with a matching assertion passes them cleanly, which is exactly the shape of this defect.

What changed in process

A known gap is promoted only after a real-worker run. The conformance catalogue’s offline exhaustiveness gate proves that every capability value has a covering case. It never dials a worker. Per-title was correct by that gate and wrong in production. The rule now is that a case loses its known-gap tag only when it has passed against real workers on the exact image that will be pinned.

The API withholds a feature rather than ship a worse output at a multiplier. 5.20.1 exists for that reason alone. The withdrawal was one commit, placed in the mode-only validation rule so all three entry paths (the interceptor, the shared output gate, and the preset path) agreed at once, keeping the plumbing, the pricing rule, the migration and the tagged cases in place. That is why re-enabling it was also one commit.

The manual gate is now documented as manual. The integration test file states the three reasons it cannot run in CI and carries a container recipe that reproduces it in about a minute.

No assertion was weakened at any point in the round trip. The per-title cases kept stating what the feature has to do to be worth selling. That is what let them fail 1.28.0 and certify 1.28.1. A test that had been loosened to pass the first run would have certified the broken search.

What it cost

Three validation sessions ran against real workers and real object storage that day: the full post-release run on 5.20.0, the complete regression on 5.20.1, and the targeted per-title run on 1.28.1.

Run Jobs VM-hours Cost
Post-release, 5.20.0 on worker 1.28.0 127 1.759 €1.87
Regression, 5.20.1 123 1.430 €1.01
Per-title re-qualification, worker 1.28.1 10 0.552 €0.35
Total 260 3.741 €3.23

Three euros and twenty-three cents to find, fix and prove a defect that a customer would otherwise have found by noticing their bill went up and their files got bigger. That is the whole argument for running the validation harness against real machines rather than mocks, and for publishing what it finds.

Open follow-ups

A CI image carrying libvmaf would let the scorer’s integration tests run automatically. Until then the gate is a person running the documented recipe whenever the per-title search or the scorer changes. Adding the analyzer package to the Docker integration target is necessary but not sufficient, because Debian’s ffmpeg would still skip. Both are on the list, and this post will be updated when they land.

One scope note, so this post is not read as more than it is: it covers per_title only. Auto-ABR ladder generation is still not available; content_aware.mode = auto_abr remains rejected at job creation with parameter_unsupported, for the pricing reason the announcement explains, and none of the fixes above change that.

Try it on your own footage

Upload a clip, pick a ladder, and read the manifest that comes back. Test encodes up to 30 seconds are free.

Topics

transcodelyengineeringpostmortemper-title encodingVMAF

Share this article