NetEQ’s Accelerate: How Audio Time-Stretching Drifts Lip-Sync Over a Long WebRTC Session

Lip-sync drift in a long WebRTC session is rarely a single catastrophic event. It is usually the slow accumulation of small timeline corrections that the audio pipeline makes to keep its jitter buffer healthy. NetEQ’s Accelerate operation is one of those corrections. Understanding what it does to the audio playout timeline — and what it does not do — is the difference between chasing a phantom codec bug and fixing the actual clock relationship between audio and video.

What NetEQ’s Accelerate actually changes

NetEQ is the adaptive jitter buffer and decoder inside WebRTC’s audio path. Its source tree in the WebRTC repository includes accelerate.cc, preemptive_expand.cc, time_stretch.cc, expand.cc, merge.cc, and decision_logic.cc — the components that decide when to compress or stretch audio to maintain buffer depth. The Accelerate operation shortens the audio signal in the time domain. It does not drop packets and it does not change RTP timestamps on the wire. It changes how many output samples are produced for a given span of decoded input.

The practical consequence: after an accelerate event, the audio renderer has consumed more RTP timestamp duration than the wall-clock time it spent playing. The audio stream’s playout position relative to its own RTP timeline has moved forward. If the video renderer is still pacing against its own RTP timeline and the shared RTCP-derived reference, the two media clocks are now offset by the amount of time removed.

This is not a bug in NetEQ. It is the intended mechanism for recovering buffer depth after jitter or packet loss. The question is whether anything in the WebRTC stack feeds that offset back into video playout.

Why audio and video do not share a playout clock

RFC 3550 specifies that audio and video are transmitted as separate RTP sessions, with separate SSRCs, separate sequence number spaces, and separate timestamp clocks. Synchronized playback is achieved using timing information carried in RTCP packets for both sessions — specifically the NTP timestamp and RTP timestamp pair in RTCP Sender Reports. The RFC is explicit that there is no direct coupling at the RTP level between the audio and video sessions.

That architecture means the receiver must reconstruct a common timeline from two independent RTP timestamp spaces. In WebRTC, this is done by mapping each stream’s RTP timestamps to a common capture clock using the RTCP SR pairs, then scheduling playout against that common clock. The mapping is established at the start of the session and updated as new Sender Reports arrive.

NetEQ’s accelerate operation does not alter the RTP timestamps in the packets it processes. It alters the relationship between decoded sample count and output sample count. From the perspective of the RTP timestamp mapping, the audio stream is still where it always was. From the perspective of the audio device, the stream has moved forward. That gap is the drift.

The feedback question

Does WebRTC adjust video playout when NetEQ accelerates audio? The short answer from the specifications is: not through any mechanism defined in RFC 3550 or RFC 8825. The WebRTC overview (RFC 8825) describes the protocol suite but does not define a cross-media playout correction loop driven by audio jitter buffer operations. The W3C WebRTC Stats specification exposes metrics that can reveal the behavior — jitterBufferDelay, jitterBufferEmittedCount, concealedSamples, concealmentEvents, and totalSamplesReceived on RTCInboundRtpStreamStats — but exposing a metric is not the same as closing a control loop.

In practice, implementations may apply their own A/V sync logic. The specifications do not mandate one, and the behavior is implementation-specific. That is the honest answer, and it is why the drift is worth measuring rather than assuming.

What the stats can tell you

The W3C WebRTC Stats document defines cumulative counters that are designed to be sampled twice and differenced. The relevant ones for this problem:

  • jitterBufferDelay — total time samples have spent in the jitter buffer, in seconds. Divide by jitterBufferEmittedCount to get average buffer delay.
  • jitterBufferEmittedCount — total number of samples emitted from the jitter buffer.
  • concealedSamples — samples generated by concealment (packet loss concealment or expand operations).
  • concealmentEvents — number of concealment events.
  • totalSamplesReceived — total samples received.

These counters do not directly expose accelerate events. The NetEQ source tree includes statistics_calculator.cc and neteq_network_stats_unittest.cc, which track internal operations, but the W3C stats surface does not currently define an acceleratedSamples counter. That means you cannot read accelerate volume directly from getStats(). You can infer it: if jitterBufferDelay stays low while concealedSamples and concealmentEvents climb, the buffer is being managed aggressively. If the audio playout position drifts relative to video, accelerate is a candidate contributor.

A practical measurement approach: sample getStats() every 10 seconds, compute the delta in jitterBufferDelay divided by the delta in jitterBufferEmittedCount, and log it alongside the observed A/V offset. If the average buffer delay is stable but the offset grows, the cause is more likely clock skew between the audio and video capture devices than NetEQ operations.

Clock skew versus NetEQ acceleration

RFC 7273 addresses this directly. It notes that RTP implementations typically assume NTP timestamps are taken using unsynchronised clocks and must compensate for absolute time differences and rate differences. Without a shared reference clock, RTP can time-align flows from the same source at a given receiver using relative timing, but tight synchronization between different receivers or between different senders is not possible.

In a WebRTC session, the audio and video capture devices on the sender may have independent crystal oscillators. A 100 ppm difference between two clocks produces 0.1 ms of drift per second, or 360 ms per hour. That is enough to become visible. NetEQ’s accelerate operation adds to this only when it fires. The two mechanisms are distinct:

Mechanism Effect on audio playout timeline Rate Detectable via
Clock skew (sender audio vs. video capture) Continuous, monotonic offset Proportional to ppm difference RTCP SR NTP/RTP pairs over time
NetEQ Accelerate Stepwise forward jump in audio playout position Event-driven, tied to jitter/loss Indirect via jitterBufferDelay and concealment counters
NetEQ Expand/PLC Stepwise backward or hold in audio playout position Event-driven concealedSamples, concealmentEvents
Video frame drop/repeat Stepwise change in video playout position Event-driven framesDropped, framesDecoded

The decision table matters because the mitigation differs. Clock skew requires a clock discipline mechanism or periodic resynchronization. NetEQ acceleration requires either accepting the offset or implementing a cross-media correction that the specifications do not define.

What you can actually do

There is no standard WebRTC API to disable NetEQ’s accelerate operation. The decision logic in decision_logic.cc and the buffer level filter in buffer_level_filter.cc determine when it fires. What you can control:

  1. Reduce the conditions that trigger accelerate. Accelerate fires when the jitter buffer is deeper than target. A more stable network path — lower jitter, fewer loss bursts — means fewer accelerate events. If you control the contribution encoder, enabling forward error correction or retransmission can reduce the loss that drives buffer growth.
  2. Measure the offset directly. Do not rely on getStats() alone. Capture the audio and video render times at the receiver and compute the offset. If you have access to the sender, compare the RTCP SR NTP timestamps for audio and video to detect clock skew at the source.
  3. Resynchronize periodically if your application allows it. Some conferencing applications insert a brief silence or a video freeze to re-anchor A/V sync. This is application-level, not protocol-level, and it is audible or visible. It is a trade-off, not a free fix.
  4. Check whether your implementation already corrects. Some WebRTC implementations apply a slow audio resampler correction to track the video clock. If yours does, the drift may be bounded. If it does not, the drift is unbounded over a long session.

What the specifications do not say

It is worth being precise about the limits of the sourced material. RFC 3550 defines the RTP and RTCP mechanisms for synchronization. RFC 8825 describes the WebRTC protocol suite. The W3C WebRTC Stats document defines the metrics. None of these documents specify a maximum accelerate ratio, a maximum cumulative offset, or a required cross-media correction loop. The NetEQ source tree shows the components — accelerate.cc, time_stretch.cc, preemptive_expand.cc — but the source code is the implementation, not a specification of behavior.

That means any claim about a specific drift threshold or a specific session duration after which lip-sync becomes perceptible is not supported by the primary sources. The honest position: measure your own pipeline. The mechanisms are known. The thresholds are not universal.

FAQ

Does NetEQ’s Accelerate change RTP timestamps?

No. Accelerate operates on decoded audio samples in the time domain. The RTP timestamps in the packets are unchanged. The effect is on the relationship between decoded sample count and output sample count, which shifts the audio playout position relative to the RTP timeline.

Is there a getStats() metric for accelerated samples?

Not in the current W3C WebRTC Stats specification. The available counters include jitterBufferDelay, jitterBufferEmittedCount, concealedSamples, and concealmentEvents. Accelerate events must be inferred from buffer behavior rather than read directly.

Can I disable Accelerate?

There is no standard API to disable it. The decision logic is internal to NetEQ. You can reduce the conditions that trigger it by improving network stability and reducing packet loss, but you cannot turn it off through a WebRTC API.

Is lip-sync drift more likely from clock skew or from NetEQ?

Clock skew is continuous and proportional to the ppm difference between capture clocks. NetEQ acceleration is event-driven and tied to jitter buffer management. Over a long session with a stable network, clock skew is the more likely dominant contributor. Over a session with bursty loss and deep jitter buffers, NetEQ operations contribute more.

How do I measure A/V offset in a WebRTC session?

Sample getStats() at known intervals, compute the delta in jitterBufferDelay divided by the delta in jitterBufferEmittedCount, and log it alongside the observed A/V offset. If you have access to RTCP Sender Reports, compare the NTP timestamp and RTP timestamp pairs for audio and video to detect clock skew at the source.

Sources