Following a conversation in a crowded room is something human hearing does almost effortlessly — a feat auditory researchers call the cocktail-party problem. Reproducing it inside a headset is far harder, because the device is capturing a jumble of overlapping sources through its own microphones and must decide, in real time, how to present them so a wearer can still lock onto one voice. A patent granted to Meta Platforms Technologies, LLC on July 7, 2026, US12677106B1, titled "Statistical provisioning of perceptual audio cues to enhance speech," is directed to a computational approach to exactly that problem.

At the center of the disclosed method is a per-source analysis. For audio arriving from a plurality of sources, the system calculates a pitch similarity for the audio data and an interaural level difference based at least in part on an attenuation level and/or a dynamic range of interaural time differences. Interaural level and time differences — ILD and ITD — are the small discrepancies in loudness and arrival time between the two ears that the brain uses to place a sound in space. By computing them per source rather than for a single mixed signal, the method treats each talker or noise source as a separately locatable object in the artificial-reality environment. The claim ties the interaural level difference to an attenuation level or a dynamic range of interaural time differences, meaning the spatial cue is derived alongside the decision about how much to quiet a source — the two are computed together rather than as independent stages.

Pitch similarity does the work of separating voices from everything else. The specification describes band-passing the signal to roughly 200 Hz to 2500 Hz — the range where the pitch and formant structure of speech lives — and using the resulting similarity measure to decide how strongly to intervene on each source. Because the measure is computed per source and updated continuously, a source that becomes more speech-like over time draws a correspondingly different treatment than the steadier background around it. The named inventors, Antje Ihlefeld and William Owen Brimijoin, II, frame the intervention around perceptual audio cues that the system adjusts in real time to dynamically shape those cues.

Psychoacoustic cues, applied selectively

Where the method departs from a conventional noise gate is in what it does after scoring the sources. Rather than simply muting background audio, it re-synthesizes the scene using cues drawn from how the auditory system already parses speech. Claim 1 enumerates them directly:

the perceptual audio cues comprise one or more of whispered backgrounds, time-dilated vowels, and enhanced sound onsets— Statistical provisioning of perceptual audio cues to enhance speech, US12677106B1

Each cue maps to a known property of intelligibility. "Whispered backgrounds" pushes competing sources toward an unvoiced, lower-salience texture so they recede without vanishing. "Time-dilated vowels" — the disclosure describes stretching vowels on the order of ten percent — lengthens the segments that carry the most linguistic information, buying the listener time to resolve them. "Enhanced sound onsets" steepens the leading edges of speech events, the transients the auditory system relies on to segment a stream into words. The system determines background audio from the audio data, attenuates it, and then generates spatial audio based at least in part on the audio data and the background audio before causing output through one or more speakers. Notably, the abstract frames this as a pipeline that runs for each source of the plurality of sources — pitch similarity, interaural level difference, background determination, and attenuation are each carried out source by source, so a scene with several talkers is handled as several parallel per-source decisions rather than one global filter applied to a summed signal.

Two design choices are worth drawing out for an informed reader. First, the cues are provisioned statistically and in real time, keyed to the per-source pitch-similarity measure — so the amount of whispering, dilation, or onset sharpening scales with how speech-like and how foregrounded a given source is, rather than being applied uniformly. Second, because the pipeline preserves interaural cues while manipulating the content, the enhanced target voice keeps its spatial position: the wearer still perceives it as coming from the right direction in the scene, even as the surrounding sources are pushed back. The audio is captured via microphones of a user's head-mounted display, and the sources may correspond to different entities in the artificial-reality environment.

Where it sits in the Reality Labs stack

The grant reads as one node in a broader cluster of Reality Labs records issued in the same window. On audio, a companion grant, US12669974B1, is directed to creating custom audio mixes for artificial-reality environments, and US12676139B1 covers mixed-reality text narration that reads on-screen text aloud and re-syncs when the text changes — both adjacent to the problem of managing what a wearer hears. On the hardware side, US12674991B2 describes AR-glasses temple-arm components with a speaker seated between front and rear battery cells, the kind of near-ear transducer placement a spatial-audio method ultimately renders through.

The same cohort extends into optics and input. US12663619B2 is directed to a wide-field-of-view optical lens assembly with a piezo-actuated tunable lens exceeding 100 degrees, and US12663864B2 covers EMG-based wrist-wearable gesture control. Taken together, the records describe the sensing, optics, and rendering layers of a head-worn platform. The speech-enhancement grant addresses the audio layer specifically — treating intelligibility not as a fixed filter but as a set of perceptual cues provisioned per source, in real time, so a single voice can be made to stand out from the crowd the headset actually hears.