· 5 min read
How to Double a Vocal From One Take
Manesh Jayawardhana
CIO & Co-founder
Copy a vocal, delay the copy by 25 milliseconds, pan them apart, and the voice sounds bigger. Delay it by 60 and it sounds like a mistake.
The line between the two is a property of hearing rather than of the mix, and knowing where it sits makes doubling predictable.
The fusion threshold
Below roughly 30 to 40 milliseconds, the ear fuses two nearly identical sounds into one. You perceive a single richer sound, localised toward whichever arrives first — the precedence effect, sometimes called the Haas effect.
Above that window the two separate perceptually and you hear an echo.
The exact threshold varies with the material. Percussive sounds separate earlier because their sharp transients are easy to distinguish; sustained sounds fuse for longer.
For vocals, 15 to 30 milliseconds is the useful range. That is why doubling settings cluster there, and why a delay that sounds like slapback rather than thickness has almost always crossed the threshold.
| Offset | Effect |
|---|---|
| Under 10 ms | Comb filtering, hollow |
| 15-30 ms | Thickening, fused |
| 40-60 ms | Ambiguous, often unpleasant |
| Over 80 ms | Distinct slapback echo |
Timing alone is not enough
A fixed delay produces a mechanical result, because a real second performance does not differ from the first by a constant.
A human singing the same line twice varies continuously — slightly ahead here, behind there, marginally sharper on one vowel, different breath placement, subtly different consonant timing. That continuous variation is what the ear recognises as two performances.
Adding a small pitch offset, around 5 to 20 cents, helps considerably. Adding slow modulation to both the pitch and the timing helps more, because it introduces the drift that a fixed offset lacks.
Even so, artificial doubling has a ceiling. On an exposed lead vocal the difference from a real double is audible to most listeners, and the fix is to record a second take rather than to keep adjusting parameters.
Where it works well
Artificial doubling is genuinely good for:
Backing vocals, where the part is less exposed and width matters more than realism.
Thickening a chorus relative to verses, where the change in texture is the point.
Salvaging a session where re-recording is not possible.
Wide effects on a part that is meant to sound processed.
It is weakest on an exposed lead in a sparse arrangement, which is exactly where people reach for it.
Check it in mono
Doubling that relies on channel differences partly cancels when summed to mono.
A double panned hard left and right with a timing offset produces comb filtering in mono — a hollow, phasey quality on the vocal, in the frequency range where the offset corresponds to a half-wavelength.
Since a large share of listening happens on single-speaker devices, that matters. Check mono before committing, and reduce the width or the offset if the vocal thins out.
Filter the double, do not just pan it
What separates a double that adds width from one that adds mud.
The doubled part occupies exactly the same frequency range as the original, at similar level. Two copies of the same spectrum sum into something thicker in the low mids, which is where a vocal is already most crowded.
High-passing the double more aggressively than the lead removes the low mid buildup and keeps the width, since width perception lives higher up. Rolling off some top on the double as well pushes it behind the lead so the original stays the focus.
Level matters too. A double at the same level as the lead competes with it; one sitting several decibels below adds dimension without ambiguity about which is the main vocal.
Common mistakes to avoid
- Timing offset past the fusion threshold, producing echo instead of thickness.
- Fixed offsets with no pitch variation, which sound mechanical.
- Using it on an exposed lead where a second take was possible.
- Not checking mono, where wide doubling partly cancels.
- Doubling an already-doubled part, which compounds the artificiality.
How to do it with Vocal Doubling Effect Tool
The Vocal Doubling Effect Tool applies both offsets in the browser.
- Load the vocal — it is processed locally.
- Start around 20 to 25 ms with a small pitch offset.
- Compare against the original and against mono.
- If it sounds mechanical on an exposed part, record a second take instead.
Other audio tools are in the tools directory.
Frequently asked questions
Why not just copy the track and delay it?
That is the basic version and it sounds mechanical, because a real double varies continuously in pitch and timing rather than by a fixed amount. Small pitch variation is what makes it convincing.
Is an artificial double as good as a real one?
No, and the difference is audible on an exposed vocal. A real second take differs in phrasing, breath and vowel shape throughout. Artificial doubling is good where re-recording is not an option.
Will it survive mono?
Check it. Wide doubling relies on channel differences that partly cancel when summed, producing a hollow quality. A double that disappears in mono will thin out on a phone speaker.
Final thought
Stay under 30 milliseconds and add a little pitch variation. Past the fusion threshold you are no longer doubling a vocal, you are adding an echo to it.