Repeat Recorder Repeat Recorder
← All articles

Osama Bin Laden voice analysis: do the 1997, 2001 and 2004 tapes match?

It has been claimed for years that some of Osama Bin Laden's video and audio messages used a paid actor or a body double rather than Bin Laden himself. That is a question a computer can help with, because a human voice carries physical markers that are hard to imitate. I took three publicly available recordings, from 1997, 2001 and 2004, and ran them through an AI speaker verification model and a set of acoustic measurements. The answer that came back is not the clean yes or no the claim invites.

The quick answer: the 1997 and 2001 recordings are very probably the same man. The 2004 recording is the odd one out. Its pitch is 33% higher than 1997, its voice texture and voice quality measurements sit far away from both earlier tapes, and it scores lowest on speaker matching. Some of that is explained by it being a broadcast tape of unknown origin, and some of it is not. No AI voice cloning is involved in any of the three, because the technology did not exist yet.

The three recordings, and where each one came from

All three clips were pulled from publicly available YouTube uploads. Their differences in length matter later, because a short clip gives a model less to work with.

Still frame from the 1997 Osama Bin Laden recording

1997

Source · 7.0 seconds, 6.6 of them speech

Recorded decades before voice cloning existed, which makes it a clean baseline.

Still frame from the 2001 Osama Bin Laden recording

2001

Source · 95 seconds, 62.1 of them speech

The longest sample by far, so it gives the model the most to measure and is the most reliable of the three.

Still frame from the 2004 Osama Bin Laden recording

2004

Source · 26 seconds, 20.9 of them speech

A tape delivered anonymously to Al Jazeera's Pakistan offices and broadcast on 29 October 2004. Bin Laden stands alone behind a lectern, and the recording equipment and conditions are unknown.

How the analysis was done

Two separate methods were used, because they fail in different ways. One gives a single same-or-different score; the other measures the physical characteristics of the voice so you can see where any difference sits.

The 1997 and 2001 tapes match each other better than either matches 2004

Comparing the three full recordings pairwise gives a same-speaker score for each pair. Higher means more alike.

0.725
1997 vs 2001
moderate to high
0.609
1997 vs 2004
low
0.638
2001 vs 2004
low to moderate

Repeating the comparison on overlapping windows rather than whole clips reproduces the same ranking, and it also shows the 2004 pairs bouncing around more.

PairComparisonsMeanMedianLowestHighestSpread
1997 vs 20011200.65200.65140.58370.72580.0253
1997 vs 2004360.55090.55290.44460.66130.0461
2001 vs 20044800.56200.56410.43160.65780.0403
Same-speaker scores from overlapping windows. The spread column is the standard deviation, so a bigger number means the score wandered more from window to window.

The vocal tract fingerprint is the part a person cannot fake

A voice resonates through the throat, mouth and nose, and that pattern of resonance can be measured. That shape is anatomy, so unlike pitch or pace it is not something a speaker can decide to change.

0.979
1997 vs 2001
0.968
1997 vs 2004
0.938
2001 vs 2004

Pitch rises across all three, which is the wrong direction for an ageing man

Fundamental frequency is how fast the vocal folds vibrate, and it is the single easiest thing on this page for a speaker to change on purpose.

Pitch measurement199720012004
Average116.7 Hz143.6 Hz155.6 Hz
Middle value119.6 Hz142.2 Hz151.6 Hz
How much it varied32.8 Hz14.9 Hz17.7 Hz
Total range110.4 Hz139.0 Hz140.0 Hz

The 2004 tape sounds physically different, not just differently recorded

Two more families of measurement describe the texture of a voice: how its energy spreads across frequencies, and how the vocal folds behave.

Voice texture measurement199720012004
Brightness centre1512.5 Hz1029.3 Hz2649.1 Hz
Spread of energy1604.4 Hz1077.5 Hz1886.1 Hz
Upper cut-off3078.9 Hz1975.6 Hz4758.2 Hz
How noise-like0.03110.00180.0693
Voice quality measurement199720012004
Zero crossing rate0.11350.08610.2502
Variation in that rate0.07670.05150.1148
Average loudness0.01400.02280.0367
Variation in loudness0.00630.00760.0242

Speaking speed is nearly identical across all three

Rhythm is measured here as how many syllable onsets occur per second, which is a rough proxy for talking speed.

Rhythm measurement199720012004
Syllables per second7.127.297.26

No AI voice cloning is involved in any of the three

Synthetic speech leaves marks. Two of them are checked here: whether pitch wobbles the way a real larynx wobbles, and whether loudness has natural micro-variation.

Artefact check199720012004Reading
Pitch wobble5.444 Hz3.942 Hz5.230 HzAll natural, 2 to 10 Hz expected
Loudness micro-variation0.002790.003350.010532004 raised, 3 to 5 times the others
Repeat Recorder's record and replay loop: record on the red screen, replay hands-free on the blue screen

Hear pitch and pace in your own voice with Repeat Recorder

Almost everything measured above is something you can hear once you listen to the same few seconds enough times, and the hard part is the listening rather than the hearing. Repeat Recorder, my own app, exists for that one job: it records you, plays the take straight back about a second later, and starts recording again without you touching the phone.

Could one man's voice really change this much in seven years?

This is where the analysis has to be honest about what it cannot separate. Several ordinary things move these numbers, and one of them moves them a lot.

Possible causeHow much it can moveDoes it explain the 2004 tape?
Ageing about seven yearsMinorNo, because pitch normally falls with age and this rose 33%
Illness or declining healthModeratePossibly, though the widely reported kidney disease would usually lower pitch and add breathiness
Stress or emotional stateModeratePartly, since stress can lift pitch 10 to 20%, but 33% is beyond that
Microphone and roomSignificantPartly, for brightness and noisiness, but not for vocal tract resonance
Formal address rather than conversationMinorPartly for pitch, but it would not treble the zero crossing rate
Broadcast compression and re-encodingSignificantYes, the Al Jazeera chain could move brightness and loudness a long way
Vocal tract anatomyCannot changeUnresolved, since 0.938 is lower than expected over three years but still high
Sort by any column to group the causes that can be ruled out against the ones that cannot.

What this analysis cannot tell you

These results came from one model and three compressed clips off the internet. That limits what they can settle.

Conclusion: two of the three Bin Laden tapes are genuine