Lip-sync error is one of the few production defects that viewers notice immediately and forgive slowly. In a webcast pipeline, it can be introduced at any stage: camera capture, contribution encoding, transport, origin packaging, CDN delivery, or the player’s decode and render path. The only way to fix it is to measure it. This article describes a practical method — audio cross-correlation against a reference track — and the packet-level evidence you need to interpret the result.
What the standards actually say
ITU-R BT.1359, Relative timing of sound and vision for broadcasting, is the primary international recommendation for acceptable audio-video timing error. It was approved in November 1998 and remains in force. The recommendation defines the relative timing of sound and vision for broadcasting and is the reference point most broadcast engineers cite when setting tolerances.
ITU-R BT.500, Methodologies for the subjective assessment of the quality of television images, is the companion document for subjective testing. The current version is BT.500-15, approved in May 2023. It is not a lip-sync measurement method, but it defines the viewing conditions and grading scales you need if you want to correlate objective measurements with human perception.
These are the two ITU-R recommendations that matter for this problem. BT.1359 gives you the target; BT.500 gives you the method for validating that your target is perceptually correct.
Why cross-correlation works
Cross-correlation measures the similarity between two signals as a function of the lag applied to one of them. If you have a reference audio track and a captured audio track from the same source, the lag at which the cross-correlation peaks is the delay between them. If you also have a video signal with a known visual event — a clap, a flash, a timecode burn-in — you can compare the audio delay to the video delay and compute the relative offset.
The method is not new. It is the basis of most automated lip-sync measurement tools. What matters in production is the reference signal and the measurement window.
Choosing a reference track
The reference track must be:
- Time-aligned with the video at the point of capture. If the reference is generated separately from the camera, you need a common clock or a clap event to establish the offset.
- Broadband and non-repetitive. Speech is ideal because it has a wide spectrum and no periodic structure that would create ambiguous correlation peaks. Music with a strong beat is worse because it produces multiple peaks at beat intervals.
- Present in the final output. If the reference track is replaced or mixed with other audio, the correlation will degrade.
A spoken-word count with a visual clap at the start is the simplest reliable reference. For automated measurement, a continuous speech track with a known timecode and a visual timecode burn-in is better because it lets you measure drift over time, not just a single offset.
The measurement procedure
At a high level:
- Capture the reference audio and the reference video at the source.
- Capture the output audio and output video at the point you want to measure — typically the player or a capture device at the edge.
- Align the two audio tracks using cross-correlation to find the audio delay.
- Align the two video tracks using a visual event or timecode to find the video delay.
- Subtract the video delay from the audio delay. The result is the lip-sync error.
The cross-correlation itself is straightforward. In practice, you need to window the signals, normalize them, and handle the fact that the output audio may be compressed, resampled, or mixed with other content.
What can go wrong
Cross-correlation is robust to linear distortion but not to nonlinear processing. The following conditions will degrade or invalidate the measurement:
- Lossy compression. AAC, Opus, and MP3 introduce phase and amplitude changes that reduce correlation peak sharpness. At typical streaming bitrates the peak is still detectable, but the confidence interval widens.
- Packet loss and concealment. If audio packets are lost and concealed, the output audio no longer matches the reference. The correlation peak may shift or disappear.
- Adaptive bitrate switching. If the player switches between renditions with different audio encoding parameters, the delay can change mid-stream. A single correlation measurement will not capture this.
- Sample rate conversion. Resampling changes the time base. If the reference and output are at different sample rates, you must resample one to match the other before correlation.
- Audio mixing. If the reference track is mixed with music or effects, the correlation is against a composite signal. The peak may still be present but weaker.
For each of these, the decision is the same: measure at a point in the pipeline where the audio is still uncompressed and unmixed, or accept a wider tolerance and validate with a subjective test.
Packet-level evidence
Cross-correlation gives you a number. Packet captures tell you why. The relevant metrics are:
- RTP timestamp deltas. If the audio and video RTP streams are synchronized, their timestamps should advance at the same rate relative to a common clock. A drift in the delta between audio and video timestamps is a direct indicator of lip-sync error.
- Arrival time jitter. If audio and video packets arrive with different jitter profiles, the playout buffer may introduce different delays. This is common when audio and video take different paths through the network.
- Encoder and decoder timestamps. PTS and DTS values in the transport stream or container tell you what the encoder intended. If the player ignores them or applies its own offset, the error is in the player, not the network.
- Packet loss and retransmission. SRT and RIST both have retransmission mechanisms. If audio and video have different loss profiles, the retransmission delays will differ.
In a production webcast, the most common cause of lip-sync error is not the network. It is the encoder or the player. Contribution encoders sometimes apply different buffering to audio and video. Players sometimes apply audio delay to compensate for video decode latency, and that compensation can be wrong.
Transport protocols and their effect on sync
SRT, RIST, WebRTC, and QUIC all have different mechanisms for handling timing and loss. None of them guarantee lip-sync by themselves. They guarantee delivery; sync is the responsibility of the encoder and decoder.
- SRT uses a timestamp-based retransmission mechanism. It can deliver audio and video with low jitter, but if the sender does not preserve the relative timing of the two streams, the receiver cannot reconstruct it.
- RIST has a similar approach with different retransmission profiles. The same caveat applies.
- WebRTC uses RTP and RTCP. It has a built-in mechanism for lip-sync: the RTP timestamp and the RTCP sender reports allow the receiver to align audio and video. If the sender does not populate these correctly, the receiver cannot align them.
- QUIC is a transport protocol, not a media protocol. It provides streams and reliability, but it does not define how audio and video should be synchronized. That is left to the application.
The practical implication: if you are measuring lip-sync error, you need to know which protocol is in use and whether it preserves the timing information you need. In most cases, the timing information is present but not used correctly.
A decision table for measurement points
| Measurement point | What you can measure | What you cannot measure |
|---|---|---|
| Camera output | Baseline sync between audio and video at capture | Anything downstream |
| Encoder input | Sync after capture and before encoding | Encoder-induced delay |
| Encoder output | Encoder-induced delay and drift | Network and player effects |
| Origin output | Packaging and origin effects | CDN and player effects |
| Player output | End-to-end lip-sync error | Which stage caused it |
Measure at the player for the final number. Measure at the encoder output and origin output to isolate the cause.
Limitations and when to use subjective testing
Cross-correlation is an objective measurement. It tells you the delay between two signals. It does not tell you whether a viewer will notice. ITU-R BT.500 defines the subjective methods for assessing quality, and while it is focused on video, the same principles apply to audio-video sync: controlled viewing conditions, a panel of viewers, and a grading scale.
If your objective measurement says the error is within tolerance but viewers complain, the problem may be that your tolerance is wrong for the content. Fast speech, close-up shots, and percussive sounds make lip-sync errors more noticeable. Wide shots and ambient audio make them less noticeable.
The practical approach is to use cross-correlation for continuous monitoring and BT.500-style subjective testing for validation. If the objective measurement is stable and within BT.1359 limits, and subjective testing confirms it, you have a defensible position.
FAQ
What is the acceptable lip-sync error for broadcast?
ITU-R BT.1359 defines the relative timing of sound and vision for broadcasting. The exact limits are in the recommendation, which is available from the ITU. For production webcasts, the same limits are a reasonable starting point, but you should validate against your own content and audience.
Can I measure lip-sync error without a reference track?
Yes, but it is harder. You can use visual events — a clap, a flash, a mouth opening — and measure the audio delay relative to those events. The accuracy depends on the precision of the event detection. A reference track is more reliable because it gives you a continuous signal to correlate against.
How often should I measure lip-sync error?
For a live production, continuous monitoring is ideal. If that is not possible, measure at the start of each session and after any change to the encoder, transport, or player configuration. Drift can occur over long sessions, so a periodic check is useful.
What tools can I use?
Any audio analysis tool that supports cross-correlation can be used. For packet-level analysis, Wireshark with RTP and SRT dissectors is standard. For automated measurement, you can write a script using a library like FFmpeg for decoding and NumPy or similar for correlation.
Does adaptive bitrate switching affect lip-sync?
It can. If the player switches between renditions with different audio encoding parameters or different buffering, the audio delay can change. This is one of the cases where a single correlation measurement is not enough; you need to measure continuously or at least at each switch.
References
- ITU-R BT.1359, Relative timing of sound and vision for broadcasting. https://www.itu.int/rec/R-REC-BT.1359/en
- ITU-R BT.500, Methodologies for the subjective assessment of the quality of television images. https://www.itu.int/rec/R-REC-BT.500/en