Skip to content
Euphona

Guide

How pitch correction actually works

Pitch correction tracks a vocal's pitch, decides what note it should be, and resynthesises it. Retune speed, not strength, is what makes it audible.

Three steps, and the second one is the hard one

Every pitch corrector, from the famous ones to the free ones, does the same three things: it works out what pitch the voice is singing at each moment, it decides what pitch it should be, and it resynthesises the audio at the new pitch.

The first step is a solved engineering problem. The third is a solved signal-processing problem. The second one — deciding what the note should be — is a musical judgement that software makes badly by default, and almost every bad-sounding tuned vocal is a consequence of that step, not of the other two.

Step one: finding the pitch

A pitch tracker estimates the fundamental frequency of the voice, typically several hundred times a second. Classical approaches used autocorrelation — sliding the waveform against itself to find the period at which it best matches. Modern trackers use neural networks trained on annotated audio and are markedly better on the difficult cases: breathy singing, growl, vibrato, the moment a note begins.

The output is a continuous pitch curve rather than a list of notes. That distinction matters: a good tool then segments the curve into notes and lets you see and edit both. If you cannot see the pitch curve, you are trusting the tool’s segmentation blindly, and segmentation is where it will get your melisma wrong.

Trackers fail predictably on octave errors — reporting a note an octave off — and on material that is not a single clear voice. Feed one a doubled vocal or a vocal with heavy bleed and it has two fundamentals to choose between and no way to know which you meant. Where the only source is a finished mix, separating the vocal out first is what makes tracking possible at all.

Step two: deciding the target

Now the tool has to decide what each note should be. The naive answer, and the default in most software, is the nearest note in a chromatic scale.

That default is wrong most of the time, and here is why. If the song is in A minor and the singer is aiming for a C, but lands a little flat, the nearest chromatic note might be a B. The corrector will confidently pull the note to B, producing something in tune with nothing and wrong in a way that is much more noticeable than the original flatness.

Which is why you should always tell it the key. Better still:

  • Set the actual scale, so notes outside the key are not available as targets.
  • Exclude notes the melody does not use. If the singer never sings an F♯, removing it from the target set removes an entire category of error.
  • Use a chord track if the tool supports one, so the target follows the harmony rather than a fixed scale. The right note in bar 3 is not the right note in bar 11.
  • Protect the deliberate ones. A blue note, a bent note, a note approached from below on purpose — these are the performance. A tuner locked to a scale will “fix” every one.

Step three: moving the pitch

Once the target is known, the audio has to be re-rendered at the new pitch without changing its duration. Two families of technique do this.

Time-domain methods — PSOLA and its relatives — cut the waveform into short grains aligned to the pitch periods and reassemble them with different spacing. Closer spacing raises the pitch, wider spacing lowers it, and repeating or dropping grains keeps the duration the same. It is computationally cheap and sounds excellent for small shifts on monophonic material, which is exactly the pitch-correction case.

Vocoder and model-based methods decompose the signal into a spectral envelope and an excitation, alter the excitation’s frequency, and resynthesise. More expensive, more flexible, and better at large shifts and at manipulating formants independently.

For the small corrections that make up most tuning work — a few tens of cents — the choice matters far less than what you asked it to do in step two.

Retune speed is the control that matters

This is the single most useful thing to understand about these tools, and it is routinely confused with correction strength.

Strength is how far a note moves toward its target. Retune speed is how quickly it gets there once the tool has decided to move it.

A human voice does not arrive at a pitch instantly. It slides in, overshoots slightly, settles, and adds vibrato. Those transitions carry an enormous amount of the expression in a performance — the portamento between two notes is often the most human thing in a phrase.

Set retune speed slow and the corrector waits, letting the slide happen naturally and only nudging the sustained part of the note. The result is transparent: the singer sounds in tune and nothing else has changed.

Set retune speed to zero and every transition is quantised. The voice jumps between exact pitches with no slide at all, and because human voices never do that, the ear immediately identifies it as processing. That is the famous effect — and it is a deliberate one, used on purpose, not a side effect of tuning too hard.

So: strong correction with slow retune is transparent. Weak correction with instant retune is obviously processed. If a tuned vocal sounds robotic and you did not want it to, the retune speed is what to change.

Formants, and why corrected vocals sound like chipmunks

A voice has two independent things going on. The vocal folds produce a fundamental frequency — the pitch. The throat, mouth and nose then filter that sound, emphasising certain frequency regions called formants. Formants are what make an “ee” sound different from an “ah”, and their positions also tell the listener roughly how large the speaker is.

If you shift the whole signal up in frequency — as a naive speed-change would — the formants move with the pitch, and the listener hears not a higher note but a smaller person. That is the chipmunk effect.

Good pitch correction preserves the formants: it moves the fundamental and leaves the spectral envelope where it was, so the singer sounds like the same person singing a different note. For the small shifts of ordinary tuning this barely matters. For shifts of several semitones it matters enormously, and a tool without formant preservation is unusable for that.

Formant shift is also available as its own control, deliberately — which is how you make a voice sound larger or smaller on purpose, independently of its pitch.

Where the artefacts come from

  • Bad tracking. If the tracker misread the pitch, the correction is applied relative to a wrong reading, and the result is confidently wrong. Breathy and very quiet passages are the usual culprits.
  • Large shifts. Every resynthesis technique degrades with distance. A semitone is nothing; five semitones on a sustained note starts to sound synthetic. If a take is consistently a long way off, re-sing it.
  • Consonants. Sibilants and plosives have no pitch to speak of, and a corrector that tries to tune them produces a characteristic smearing. Good tools detect unvoiced regions and leave them alone. Harsh sibilance is a separate job again, and it belongs with the restoration stage, after pitch rather than before it.
  • Bleed. A vocal track with spill from the room, or a vocal extracted from a mix, gives the tracker two things at once. Tuning the loudest of them drags the other along.
  • Fighting vibrato. Vibrato is deliberate pitch modulation; a fast retune reads it as error and flattens it. Some tools separate vibrato from pitch drift so you can correct one and keep the other, which is the right design.

How to use it well

  • Start with the best take you can get. Correction fixes pitch. It does not fix phrasing, timing, tone or conviction, which are what actually make a vocal good — and most of what goes wrong at the microphone cannot be corrected afterwards at all.
  • Always set the key and scale. The chromatic default is the source of most bad results.
  • Slow the retune unless you specifically want the effect.
  • Tune the notes that are wrong, not every note. Most tools will happily process a whole take; most takes have four or five notes that need it.
  • Listen in context. A vocal soloed will reveal every artefact. A vocal in the mix reveals which ones anyone will hear — and if it is still not sitting right once the tuning is clean, that is usually an EQ question.
  • Keep the untouched take. A week later you will want to compare, and you will sometimes prefer the original.

A last note on names: the technique is pitch correction or vocal tuning, and a tool that does it is a vocal tuner. It is not Auto-Tune — written “autotune” in most searches — which is one specific product, from Antares, and using its name for the whole category is both inaccurate and their trademark. What Auto-Tune actually is covers that product specifically: where it came from, and why the famous robotic sound is a speed setting rather than the sound of correcting something hard.

Keep reading