What Is a Jitter Buffer and How Does It Work?
Every voice call you make over the internet relies on a mechanism you've never seen: the jitter buffer. It's the reason your call sounds smooth even when packets don't arrive at perfectly regular intervals β and it's also the reason calls sometimes develop an awkward delay, or suddenly go robotic when the network gets worse. Here's exactly how it works.
A jitter buffer is a small memory queue that holds incoming audio packets briefly before playing them out. It converts an irregular packet stream into smooth audio playback. The tradeoff: a larger buffer absorbs more jitter but adds more call delay. Modern apps use adaptive jitter buffers that grow and shrink automatically based on network conditions.
The Problem It Solves
VoIP audio is transmitted as a stream of small UDP packets, typically sent every 20ms. In a perfect network, each packet arrives exactly 20ms after the last. In the real world, packets experience variable delays β some arrive in 18ms, others in 35ms, occasionally one arrives in 200ms or not at all. This variation is jitter.
If you played out packets the instant they arrived with no buffering, the listener would hear jitter directly as irregular timing between sound segments. Words would stretch and compress unpredictably. The speech would be unintelligible.
The jitter buffer solves this by deliberately introducing a short delay. Packets arrive, sit in the buffer briefly, and are played out at a steady predetermined rate β regardless of how unevenly they arrived. The buffer absorbs the variation so the listener hears smooth audio.
How It Works: Step by Step
The process has four stages. Receive: packets arrive at the buffer from the network. Reorder: if packets arrive out of sequence (common with jitter), the buffer sorts them by sequence number. Hold: packets wait until their scheduled playback time. Play out: packets leave the buffer at a steady 20ms interval, regardless of how they arrived.
When a packet arrives too late to be played in sequence β past its scheduled playback window β the decoder uses Packet Loss Concealment (PLC) to synthesize a replacement audio frame. Modern codecs like Opus do this well for gaps under 100ms; the listener typically hears a brief glitch rather than silence.
Three Types of Jitter Buffer
β Can't adapt to changing conditions
β Set too small: artifacts on bad networks
β Set too large: permanent unnecessary delay
+ Low delay on stable connections
β Call latency increases during network stress
β Algorithm quality varies by implementation
+ Corrects drift without resetting buffer
β More complex to implement
β Limited range for large jitter spikes
The Fundamental Tradeoff: Latency vs Resilience
Every jitter buffer is a dial between two competing goals: low call latency and resilience to jitter. You cannot fully optimize both simultaneously.
Natural conversation
Artifacts on any jitter
Stable connections only
Handles moderate jitter
Occasional brief pauses
Typical home use
Absorbs high jitter
No quality artifacts
Poor networks only
The ITU-T G.114 recommendation states that one-way delay should stay under 150ms for acceptable voice quality β 300ms is the absolute maximum before conversations become significantly awkward. A 400ms jitter buffer alone would exhaust this budget before even accounting for network latency. This is why adaptive buffers grow only as far as necessary, and why fixing your underlying jitter is so much more effective than relying on a large buffer.
If a call starts feeling unnaturally delayed β like there's a growing echo of your own voice, or the other person pauses awkwardly β the adaptive buffer is likely expanding in response to increasing jitter. Audio quality is preserved, but at the cost of latency. The fix is to address the jitter: switch to Ethernet, pause background traffic, or investigate the ISP.
Packet Loss Concealment: What Happens When the Buffer Fails
When a packet arrives past its scheduled playback slot, the jitter buffer treats it as lost and the audio codec applies Packet Loss Concealment. The quality of PLC varies significantly by codec:
| Codec | PLC Quality | Gap Tolerance | Used By |
|---|---|---|---|
| Opus | Excellent | ~120ms undetectable | Zoom, Teams, Discord, WebRTC, WhatsApp |
| AAC-ELD | Excellent | ~80ms acceptable | Apple FaceTime, iMessage audio |
| G.722 | Good | ~60ms acceptable | HD voice, enterprise VoIP |
| G.711 / G.729 | Basic | ~20ms before audible | PSTN, traditional business VoIP, SIP |
This is why modern consumer apps feel more resilient than traditional business VoIP systems on the same network β Opus's PLC is dramatically better than G.711's. On a connection with 2% packet loss, Opus calls are often still intelligible while G.711 calls are already breaking up.
Jitter Buffer Settings in Common Apps
Fully Automatic β No User Settings
Both Zoom and Teams manage their jitter buffers entirely automatically. There are no user-accessible settings. The buffer adapts to observed jitter and can grow to 400ms under poor conditions, prioritizing audio quality over latency. If calls feel delayed, fix the underlying network jitter β the buffer will shrink automatically.
Fully Configurable via rtp.conf
Set jbenable=yes, jbimpl=adaptive, jbmaxsize=200 (ms), jbresyncthreshold=1000. For fixed mode: jbimpl=fixed with jbprefetchcount=60. On stable LAN environments, use adaptive with max 80β100ms for natural conversation feel. Over internet links, 150β200ms max handles typical broadband variability.
Developer-Configurable via API
WebRTC exposes a playoutDelayHint property on RTCRtpReceiver for developer-controlled buffer hints. Example: receiver.playoutDelayHint = 0.04 hints for a 40ms buffer in seconds. Chromium-based browsers honor this reasonably well. Useful for latency-critical applications like real-time collaboration tools where the default adaptive algorithm is too conservative.
How to Get the Best Results
For consumer apps (Zoom, Teams, Discord), the only effective lever is reducing underlying network jitter. The adaptive buffer will then shrink automatically to a low-latency setting. For configurable systems:
- Measure first. Run a jitter test on jitter.is to know your typical and peak values. Set buffer max to 2β3Γ peak jitter with headroom.
- Use adaptive mode. Fixed buffers only make sense on tightly controlled LAN environments with known stable jitter.
- Don't set max too conservatively. A 50ms max that hits the ceiling during occasional 80ms spikes is worse than a 120ms max that rarely gets used.
- On LAN: 30β60ms adaptive β low enough for natural conversation, high enough for any LAN-level variation.
- On home broadband: 80β150ms adaptive handles typical Wi-Fi and cable internet patterns.
The goal isn't a large, forgiving jitter buffer β it's a small one that barely has to do anything because your network jitter is low. A 20ms adaptive buffer on a well-tuned Ethernet connection produces calls indistinguishable from being in the same room. Reduce jitter at the network level; let the buffer shrink as a natural consequence.
Measure your jitter β lower jitter means a smaller buffer and better call quality.
// TEST MY JITTERSummary
A jitter buffer is a receive-side memory queue that converts an irregular packet stream into smooth, steady audio playback. It works by holding incoming packets until their scheduled playback time, reordering out-of-sequence packets, and using Packet Loss Concealment to replace packets that arrive too late.
The fundamental tradeoff is latency versus resilience: a larger buffer absorbs more jitter but adds more delay. Modern adaptive buffers navigate this automatically, growing when conditions deteriorate and shrinking when they improve.
For consumer apps there are no buffer settings to configure β fix the network instead. For configurable systems like Asterisk or WebRTC, use adaptive mode with a max buffer 2β3Γ your peak observed jitter. In all cases, the most effective intervention is reducing jitter at the network level through Ethernet, QoS, and router quality β which allows the adaptive buffer to stay small and call latency to stay low.