Repeat Recorder Repeat Recorder
← All articles

Ghislaine Maxwell voice analysis: is the 2026 voice the same person?

A claim has circulated that the Ghislaine Maxwell who was arrested, tried and imprisoned is not the real Ghislaine Maxwell but a body double. Unlike most claims of that shape, this one can be tested, because a recorded voice carries physical markers that a performer cannot choose to change. I took three publicly available recordings, from 1991 at age 29, 2014 at age 52 and 2026 at age 64, and ran them through an AI speaker verification model and a set of acoustic measurements. Every measurement that reflects anatomy rather than performance points the same way.

The quick answer: all three recordings are very probably the same woman. The vocal tract fingerprint, the one measurement a person cannot consciously alter, matches at 0.98 to 0.99 across every pair. Speaker matching runs 0.68 to 0.79, well above the 0.40 to 0.55 a skilled impersonator typically reaches. Voice quality changes steadily in one direction over the 35 years, which is what ageing sounds like, and no recording shows any sign of AI synthesis.

The three recordings, and where each one came from

All three clips were pulled from publicly available YouTube uploads. The 2014 recording matters most, because it sits between the other two in time and is long enough to measure well.

Still frame from the 1991 Ghislaine Maxwell recording

1991, age 29

Source · 37 seconds, 32.4 of them speech

Recorded decades before voice cloning existed, which makes it a clean baseline.

Still frame from the 2014 Ghislaine Maxwell recording

2014, age 52

Source · 60 seconds, 49.5 of them speech

Recorded before the arrest and before deepfakes were possible. The longest sample, so the most reliable of the three.

Still frame from the 2026 Ghislaine Maxwell recording

2026, age 64

Source · 23 seconds, 18.5 of them speech

The recording the body double claim is actually about, and the shortest of the three.

How the analysis was done

Two separate methods were used, because they fail in different ways. One gives a single same-or-different score; the other measures the physical characteristics of the voice so you can see where any difference sits.

Every pair scores above the range a skilled impersonator reaches

Comparing the three full recordings pairwise gives a same-speaker score for each pair. Higher means more alike.

0.786
1991 vs 2014
high
0.685
1991 vs 2026
moderate
0.699
2014 vs 2026
moderate

Repeating the comparison on overlapping windows rather than whole clips holds the same ranking.

PairComparisonsMeanMedianLowestHighestSpread
1991 vs 20142,5830.56830.57400.35200.76520.0632
1991 vs 20269430.50870.50730.32240.65680.0515
2014 vs 20261,4490.53730.53790.36870.67580.0515
Same-speaker scores from overlapping windows. The spread column is the standard deviation, so a bigger number means the score wandered more from window to window.

The vocal tract fingerprint is the part a person cannot fake

A voice resonates through the throat, mouth and nose, and that pattern of resonance can be measured. The shape is anatomy, so unlike pitch or pace it is not something a speaker can decide to change.

0.981
1991 vs 2014
0.980
1991 vs 2026
0.991
2014 vs 2026

Pitch rises then settles, which is what ageing sounds like

Fundamental frequency is how fast the vocal folds vibrate, and it is the single easiest thing on this page for a speaker to change on purpose.

Pitch measurement199120142026
Average132.1 Hz204.3 Hz170.3 Hz
Middle value110.7 Hz198.3 Hz167.8 Hz
How much it varied59.9 Hz33.0 Hz26.5 Hz
Total range439.8 Hz221.7 Hz206.8 Hz

Voice quality moves steadily in one direction across the 35 years

Two more families of measurement describe the texture of a voice: how its energy spreads across frequencies, and how the vocal folds behave.

Voice quality measurement199120142026
Zero crossing rate0.07460.15060.1958
Variation in that rate0.05560.10690.1403
Average loudness0.02740.03390.0329
Variation in loudness0.02330.02900.0302
Voice texture measurement199120142026
Brightness centre1501.3 Hz1981.6 Hz1974.2 Hz
Spread of energy1772.0 Hz1841.7 Hz1596.9 Hz
Upper cut-off3351.0 Hz3878.3 Hz3565.9 Hz
How noise-like0.01230.04230.0601

Speaking speed barely moves across 35 years

Rhythm is measured here as how many syllable onsets occur per second, which is a rough proxy for talking speed.

Rhythm measurement199120142026
Syllables per second4.955.455.44

No AI voice cloning is involved in any of the three

Synthetic speech leaves marks. Two of them are checked here: whether pitch wobbles the way a real larynx wobbles, and whether loudness has natural micro-variation.

Artefact check199120142026Reading
Pitch wobble6.053 Hz6.733 Hz4.945 HzAll natural, 2 to 10 Hz expected
Loudness micro-variation0.008040.010790.01051All natural
Repeat Recorder's record and replay loop: record on the red screen, replay hands-free on the blue screen

Keep your own voice baseline with Repeat Recorder

The reason this analysis can say anything at all is that a 1991 recording existed to compare against. Almost nobody has that for their own voice. Repeat Recorder, my own app, is built for recording yourself and hearing it straight back, and the takes you keep in it become the baseline you do not currently have.

Could a voice actor or an AI have produced these recordings?

Both explanations have to be taken seriously, and they fail for different reasons. The table below separates what a performer can imitate from what they cannot.

Voice characteristicHow fakeableWhat this analysis found
PitchEasyVaries across the three, so it settles nothing
Speaking speed and rhythmEasyConsistent, within 0.5 syllables per second
Accent and intonationModerateBritish accented in all three recordings
Same-speaker scoreHard0.68 to 0.79, against 0.40 to 0.55 for impersonators
Voice quality and harmonicsVery hardNatural patterns, ageing steadily in one direction
Vocal tract resonanceNearly impossible0.980 to 0.991 across every pair
Formant ratiosNearly impossibleSet by anatomy, not by performance
Pitch wobble and shimmerInvoluntaryNatural in all three recordings
Sort by the middle column to separate what a performer can imitate from what they cannot.

What this analysis cannot tell you

These results came from one model and three compressed clips off the internet. That limits what they can settle.

Conclusion: the Maxwell body double claim does not hold up