A claim has circulated that the Ghislaine Maxwell who was arrested, tried and imprisoned is not the real Ghislaine Maxwell but a body double. Unlike most claims of that shape, this one can be tested, because a recorded voice carries physical markers that a performer cannot choose to change. I took three publicly available recordings, from 1991 at age 29, 2014 at age 52 and 2026 at age 64, and ran them through an AI speaker verification model and a set of acoustic measurements. Every measurement that reflects anatomy rather than performance points the same way.
The quick answer: all three recordings are very probably the same woman. The vocal tract fingerprint, the one measurement a person cannot consciously alter, matches at 0.98 to 0.99 across every pair. Speaker matching runs 0.68 to 0.79, well above the 0.40 to 0.55 a skilled impersonator typically reaches. Voice quality changes steadily in one direction over the 35 years, which is what ageing sounds like, and no recording shows any sign of AI synthesis.
The three recordings, and where each one came from
All three clips were pulled from publicly available YouTube uploads. The 2014 recording matters most, because it sits between the other two in time and is long enough to measure well.

1991, age 29
Recorded decades before voice cloning existed, which makes it a clean baseline.

2014, age 52
Recorded before the arrest and before deepfakes were possible. The longest sample, so the most reliable of the three.

2026, age 64
The recording the body double claim is actually about, and the shortest of the three.
How the analysis was done
Two separate methods were used, because they fail in different ways. One gives a single same-or-different score; the other measures the physical characteristics of the voice so you can see where any difference sits.
- All three files were stripped of silence before anything was measured. Each MP3 was loaded at 16 kHz in mono, then WebRTC voice activity detection removed silence and non-speech, which is why the speech durations above are shorter than the clip lengths.
- A neural speaker verification model turned each recording into a 256-number fingerprint, then compared them. Resemblyzer with its pre-trained GE2E model produces one vector per voice, and the closeness of two vectors gives a same-speaker score between 0 and 1.
- Every clip was also chopped into overlapping windows and cross-compared, so one lucky or unlucky second cannot decide the result. That produced between 943 and 2,583 separate comparisons per pair instead of one.
- Six families of acoustic measurements were extracted with librosa. Pitch, vocal tract resonance, voice texture, rhythm, voice quality, and a check for the artefacts that synthetic speech leaves behind.
Every pair scores above the range a skilled impersonator reaches
Comparing the three full recordings pairwise gives a same-speaker score for each pair. Higher means more alike.
high
moderate
moderate
Repeating the comparison on overlapping windows rather than whole clips holds the same ranking.
| Pair | Comparisons | Mean | Median | Lowest | Highest | Spread |
|---|---|---|---|---|---|---|
| 1991 vs 2014 | 2,583 | 0.5683 | 0.5740 | 0.3520 | 0.7652 | 0.0632 |
| 1991 vs 2026 | 943 | 0.5087 | 0.5073 | 0.3224 | 0.6568 | 0.0515 |
| 2014 vs 2026 | 1,449 | 0.5373 | 0.5379 | 0.3687 | 0.6758 | 0.0515 |
- The whole range of 0.68 to 0.79 sits well above the 0.40 to 0.55 a skilled impersonator typically scores on this kind of model. That gap is the reason a voice actor is a poor explanation for these three recordings, before any of the anatomy measurements are considered.
- The 1991 and 2014 pair scores highest at 0.786, comfortably inside the same-speaker range. Both predate any usable voice cloning, so whatever they agree on is a genuine baseline for the person rather than an artefact of technology.
- The 2014 recording matches strongly with both the 1991 and 2026 clips, which chains all three together. A body double swapped in later would have to match a 1991 recording it never heard, through a 2014 recording that also matches.
The vocal tract fingerprint is the part a person cannot fake
A voice resonates through the throat, mouth and nose, and that pattern of resonance can be measured. The shape is anatomy, so unlike pitch or pace it is not something a speaker can decide to change.
- All three pairs exceed 0.97, and the 2014 and 2026 recordings reach 0.991. This is the single strongest result on the page, because it measures the physical shape of a throat rather than anything a performer controls.
- The oldest and newest recordings, 35 years apart, still match at 0.980. Ageing changes how a voice sounds but not the cavity it resonates in, which is exactly the pattern here.
- The same measurement on the three Bin Laden recordings fell to 0.938 across only three years. That contrast is useful: it shows what this method looks like when it is not satisfied, and these Maxwell figures are not that.
Pitch rises then settles, which is what ageing sounds like
Fundamental frequency is how fast the vocal folds vibrate, and it is the single easiest thing on this page for a speaker to change on purpose.
| Pitch measurement | 1991 | 2014 | 2026 |
|---|---|---|---|
| Average | 132.1 Hz | 204.3 Hz | 170.3 Hz |
| Middle value | 110.7 Hz | 198.3 Hz | 167.8 Hz |
| How much it varied | 59.9 Hz | 33.0 Hz | 26.5 Hz |
| Total range | 439.8 Hz | 221.7 Hz | 206.8 Hz |
- Pitch peaks in the 2014 recording at 204.3 Hz and comes back down to 170.3 Hz by 2026, rather than climbing in a straight line. A voice that rises and then settles across three decades fits one person in different situations at different ages, and it is a poor fit for a substitution that would have no reason to land in the middle.
- Pitch also steadies with each recording, varying by 59.9 Hz in 1991 against 26.5 Hz in 2026. Wide swings at 29 and a narrower band at 64 is the ordinary direction of travel for a speaking voice.
- Pitch alone proves nothing, because it is the one measurement here that a speaker can consciously push up or down. It matters only because it is consistent with the vocal tract and voice quality results rather than fighting them.
Voice quality moves steadily in one direction across the 35 years
Two more families of measurement describe the texture of a voice: how its energy spreads across frequencies, and how the vocal folds behave.
| Voice quality measurement | 1991 | 2014 | 2026 |
|---|---|---|---|
| Zero crossing rate | 0.0746 | 0.1506 | 0.1958 |
| Variation in that rate | 0.0556 | 0.1069 | 0.1403 |
| Average loudness | 0.0274 | 0.0339 | 0.0329 |
| Variation in loudness | 0.0233 | 0.0290 | 0.0302 |
| Voice texture measurement | 1991 | 2014 | 2026 |
|---|---|---|---|
| Brightness centre | 1501.3 Hz | 1981.6 Hz | 1974.2 Hz |
| Spread of energy | 1772.0 Hz | 1841.7 Hz | 1596.9 Hz |
| Upper cut-off | 3351.0 Hz | 3878.3 Hz | 3565.9 Hz |
| How noise-like | 0.0123 | 0.0423 | 0.0601 |
- The zero crossing rate climbs in one direction across all three recordings, from 0.075 to 0.151 to 0.196. That measurement tracks breathiness and the loss of harmonic richness, and rising steadily with age is exactly the expected path for one voice.
- Loudness stays in a narrow band the whole way, between 0.027 and 0.034. Consistent vocal effort across 35 years and three completely different settings is a point in favour of one speaker with one habitual delivery.
- The 2014 and 2026 recordings have almost the same brightness, at 1982 Hz and 1974 Hz, while 1991 sits lower at 1501 Hz. That gap is best explained by 1991 recording technology rather than by a different voice, which is also why brightness carries less weight here than the vocal tract numbers.
Speaking speed barely moves across 35 years
Rhythm is measured here as how many syllable onsets occur per second, which is a rough proxy for talking speed.
| Rhythm measurement | 1991 | 2014 | 2026 |
|---|---|---|---|
| Syllables per second | 4.95 | 5.45 | 5.44 |
- The 2014 and 2026 recordings are within one hundredth of each other, at 5.45 and 5.44 syllables per second, and 1991 is only slightly slower at 4.95. Talking speed that stable over 35 years fits one person with one set of speech habits.
- Speaking speed is also one of the easier things for an impersonator to copy, especially with reference tapes to practise against. So this points towards one speaker without being able to settle the question by itself.
No AI voice cloning is involved in any of the three
Synthetic speech leaves marks. Two of them are checked here: whether pitch wobbles the way a real larynx wobbles, and whether loudness has natural micro-variation.
| Artefact check | 1991 | 2014 | 2026 | Reading |
|---|---|---|---|---|
| Pitch wobble | 6.053 Hz | 6.733 Hz | 4.945 Hz | All natural, 2 to 10 Hz expected |
| Loudness micro-variation | 0.00804 | 0.01079 | 0.01051 | All natural |
- All three clips wobble in pitch within the 2 to 10 Hz range that a living larynx produces, and cloned speech usually comes in under 2 Hz. Synthetic voices tend to be too smooth, and none of these three is.
- Loudness varies naturally from moment to moment in all three recordings, where synthesis tends to flatten it out. Both of the artefacts this check looks for are absent, including in the 2026 clip that the body double claim is actually about.

Keep your own voice baseline with Repeat Recorder
The reason this analysis can say anything at all is that a 1991 recording existed to compare against. Almost nobody has that for their own voice. Repeat Recorder, my own app, is built for recording yourself and hearing it straight back, and the takes you keep in it become the baseline you do not currently have.
- Takes are saved into folders you name, so a recording you make today is still findable years later. The whole chain of identity in this article rests on old recordings having survived, and a dated take of yourself reading the same paragraph is the same idea at a personal scale.
- Your own speaking speed is worth checking, because Maxwell's moved by only 0.01 syllables per second between 2014 and 2026. Record the same sentence on two different days, replay both, and you will hear whether your pace is as steady as that or drifts.
- A take plays back about a second after you stop, and it can loop once, a set number of times, or until you stop it. Hearing yourself immediately is what makes small differences audible at all, and hands-free mode keeps the loop running while your hands are busy.
Could a voice actor or an AI have produced these recordings?
Both explanations have to be taken seriously, and they fail for different reasons. The table below separates what a performer can imitate from what they cannot.
| Voice characteristic | How fakeable | What this analysis found |
|---|---|---|
| Pitch | Easy | Varies across the three, so it settles nothing |
| Speaking speed and rhythm | Easy | Consistent, within 0.5 syllables per second |
| Accent and intonation | Moderate | British accented in all three recordings |
| Same-speaker score | Hard | 0.68 to 0.79, against 0.40 to 0.55 for impersonators |
| Voice quality and harmonics | Very hard | Natural patterns, ageing steadily in one direction |
| Vocal tract resonance | Nearly impossible | 0.980 to 0.991 across every pair |
| Formant ratios | Nearly impossible | Set by anatomy, not by performance |
| Pitch wobble and shimmer | Involuntary | Natural in all three recordings |
- A voice actor is a poor explanation, because everything they could imitate is inconclusive here and everything they could not is a match. Pitch and rhythm are imitable and prove nothing either way, while vocal tract resonance at 0.980 to 0.991 is not something practice can deliver.
- An AI clone is a poor explanation for a different reason, which is that the natural imperfections are all still present. Tools like ElevenLabs and VALL-E can fool a verification model, but they tend to smooth out the pitch wobble and loudness variation that all three of these recordings still have.
- Two of the three recordings predate usable voice cloning entirely, so any synthesis theory has to explain only the 2026 clip. That clip has to match a 1991 baseline on anatomy it cannot access, which is a much harder problem than producing a convincing voice.
What this analysis cannot tell you
These results came from one model and three compressed clips off the internet. That limits what they can settle.
- Only one speaker verification model was used, so this is not forensic evidence. A forensic conclusion needs several models, uncompressed source audio and a human expert, and it would still be expressed as a likelihood rather than a fact.
- All three files are compressed MP3s taken from YouTube, so some of the difference measured is compression rather than voice. That affects the brightness and noisiness figures most, and the vocal tract measurements least.
- The three recordings were made in completely different circumstances 35 years apart, and 1991 recording technology alone moves several of these numbers. The brightness gap between 1991 and the two later clips is the clearest example.
- The 2026 clip holds only 18.5 seconds of speech, the least of the three, so the pairs involving it are the weakest on this page. They are also the pairs the body double claim depends on, which is worth stating plainly.
- The tools were Resemblyzer with its GE2E model, librosa, WebRTC voice activity detection, and Python with NumPy and SciPy. Anyone wanting to check the working can reproduce it with the three linked source clips.
Conclusion: the Maxwell body double claim does not hold up
- All three recordings are the same woman. There is no body double in this audio. Every measurement that reflects her actual anatomy agrees, across 35 years and three completely different settings.
- The vocal tract match of 0.980 to 0.991 is the number that settles it, and no impersonator can reach that. A skilled voice actor scores 0.40 to 0.55, because you cannot change the shape of your own throat.
- The 2026 recording is a real human voice, not an AI clone. Its pitch wobble and loudness variation are both natural, and synthesis flattens exactly those two things.