
Localization refers to our ability to discern the origin of a sound source in space relative to ourselves, including its horizontal and vertical angles, its estimated distance and any perceived movement. We use a number of monaural (one ear) and binaural (two ears) audible cues for these purposes. One such binaural cue is the interaural time difference, or ITD. This refers to the difference in time it takes a sound to reach one ear compared to the other. Sounds located directly in front of or behind us will reach both ears simultaneously. If the angle of the source is moved until the difference is greater than 10 microseconds (10 millionths of a second or 10 µs), a difference in location can be perceived. As a source moves more directly to one side of your head or the other, our ability to discriminate its location using the ITD method diminishes somewhat.
A second binaural mechanism, called the interaural intensity difference, or IID, uses the difference in amplitude caused by the head physically shadowing or masking sounds coming from one side or the other—the interaural intensity difference (IID) is sometimes referred to as the interaural level difference, or ILD. This filtering is called the head-related transfer function (HRTF)—you may see environmental recordists with head-shaped binaural microphones, such as the Neumann KU100 pictured below, trying to recreate this function. Because lower frequencies with longer wavelengths diffract more easily around objects, this mechanism is more effective for higher frequencies. The shape of the pinna (outer ear flap) also filters frequencies depending on their angle of incidence. The pinna is also responsible for our ability to place sounds in the vertical plane using this filtering mechanism. Try folding your ear flap over and see how well you can still place sounds. Sound waves reflecting off the shoulder also provide some location cues. All of these mechanisms are ineffective below approximately 200 Hz, as witnessed by the often out-of-the-way placement of subwoofers in surround sound setups. Taken together, these two cues are known as the known as the duplex theory of localization: the ITD dominates for lower frequencies (up to roughly 1500 Hz), where the head is too small relative to the long wavelengths to cast much of an acoustic shadow, while the IID takes over at higher frequencies, where the head shadows effectively but wavelengths are too short for timing differences to remain unambiguous.

As mentioned earlier, however, the ear canal itself resonates and amplifies frequencies in roughly the 2–5 kHz region, depending on the angle of incidence of the sound. So this mechanism can provide a monaural cue, as slightly turning one's head to increase or decrease the intensity of this resonance is computed by the brain. It may also change the phase of the many reflected signals entering the ear canal off the body and pinnae, and alter the constructive and destructive interference taking place as a result. In fact, a great deal of our ability to localize sound is a psychoacoustic learning mechanism, sometimes referred to as the cone of confusion for sound stimuli that cannot be immediately placed, as they may lie equidistant between the two ears. Prior experience and visual cues aid in resolving some of the ambiguity when it exists.
A psychoacoustic phenomenon to keep in mind when placing loudspeakers is the precedence effect (also known as the Haas Effect), in which a listener receiving the same signal from multiple speakers will place it at the closest speaker, and not in between, unless the time difference between the signals' arrival is less than approximately 35–50 ms (lower for transient clicks, higher for sustained music/speech). Above that threshold, the arrival of the second signal is perceived as an echo of the first. This is why you should try to sit in a central location at a multi-channel electronic music concert! In stadiums, churches, and other large areas with public address systems, signals are often delayed to loudspeakers placed farther away from the origin of the sound, so that listeners sense the sound to be coming from the location they expect it to.
Beyond direction, we also judge how far away a sound source is, and here the cues are somewhat different. The most obvious is simple loudness: all else being equal, a more distant source is quieter. But loudness alone is unreliable—a soft sound nearby and a loud sound far away can reach our ears at the same level—so the brain leans on additional cues. The most powerful is the direct-to-reverberant ratio: as a source moves farther away in a room, the direct sound reaching us straight from the source weakens while the reflected, reverberant energy bouncing off the walls stays relatively constant, so the proportion of reverberation grows with distance. A nearly ‘dry’ sound reads as close; a sound dominated by reflections reads as far. Distant sounds also lose their high frequencies, since air progressively absorbs higher frequencies over long distances, and they lose some of their crispness and transient detail. Interestingly, listeners tend to underestimate the distance of faraway sources, compressing the far field. For composers, these cues are a practical toolkit: adjusting the balance of dry signal to artificial reverb, rolling off the highs, and softening transients will push a sound convincingly into the distance far more effectively than simply turning down its volume.
In judging the apparent size of an acoustic space, the aural cues depend on many factors, including the time elapsed from hearing the source sound to hearing the earliest reflections, the onset of reverberation, the intensity and duration of reverberation, diffusion of high frequencies, and the resonant frequencies of the reverberation. With multi-channel sound and control over artificial reverb, many interesting and novel spatial effects can be created. Some modern multi-speaker, multi-plane sound recording and reproductive systems, such as high-order ambisonics, are based on our ever-expanding knowledge of localization.