top of page
Search

Why Similarity Is Never Proof

The name is almost identical.

The age is approximately right.

The father's name matches.

The person came from the same town.

Surely this must be the person we are looking for.

Perhaps.

But similarity is not proof.

It is a lead.

And confusing the two is one of the easiest ways to make a serious mistake in genealogical research.


Similarity tells us where to look


Genealogical research often begins with very little information.

We may have a name.

An approximate year of birth.

A place.

Perhaps the name of a parent or spouse.

We search available records and find someone who appears to fit.

That is exactly what we hope will happen.

The similarity between the known information and the newly discovered record gives us a candidate worth investigating.

But that is all it gives us.

A candidate.

The next question is not:

What else can I find about this person?

It is:

What evidence demonstrates that this is the person I am looking for?

That distinction may seem small.

In forensic genealogy, it is fundamental.


Names are less unique than we think


People tend to experience names as personal.

Our name identifies us.

So when we find the same name in a historical record, the match can feel significant.

But names exist within populations.

Some are extremely common.

Children are named after grandparents.

Cousins may share names.

Families living in the same community may use the same limited pool of traditional names generation after generation.

The problem becomes even more complicated when names move between languages.

Mosze may become Moshe, Moses, Maurice or Morris.

Chaim may appear as Haim, Hayim, Heim or Henry.

A surname written in Cyrillic may acquire several Polish spellings, another German spelling and several possible English transliterations.

Two records that look different may therefore describe the same person.

But the reverse is equally important.

Two records that look remarkably similar may describe entirely different people.

The name is evidence.

It is rarely enough evidence on its own.


Approximate dates can create false confidence


Dates of birth are another common source of apparent similarity.

Suppose we are looking for someone born around 1900.

We find a person with the correct name born in 1899.

That seems promising.

And it may be.

But "born around 1900" could potentially include several years.

If the name is common, there may be multiple people who fit the same approximate age range.

Historical records also contain incorrect dates.

People changed their ages.

Officials made mistakes.

Later informants remembered incorrectly.

Dates were converted between calendars.

A person may appear with several different years of birth across a lifetime of records.

This flexibility is important when evaluating evidence.

But it also creates danger.

If we are willing to accept a five-year difference in one direction and another five-year difference in the other, our candidate pool becomes much larger than it first appears.

Approximation should help us search.

It should not quietly become proof.


Geography can mislead us too


The person came from the right town.

Another excellent sign.

But what does "from" mean?

Born there?

Lived there?

Registered there?

Came from a nearby village?

Named the nearest city when speaking to a foreign official?

Belonged to an administrative district that later changed?

Historical geography is rarely as neat as a modern map.

This is especially true in areas affected by border changes, war and migration.

A single locality may appear under several names and several jurisdictions.

At the same time, a large city may appear in records as the place of origin for people born in dozens of surrounding communities.

A geographical match can strengthen an identification.

But, like a name or date, it has to be understood in context.


Several similarities are stronger—but still not necessarily proof


This is where things become interesting.

One similarity may be coincidence.

What about three?

The name matches.

The approximate year of birth matches.

The father's name matches.

Surely now we have our person.

Possibly.

The probability may certainly have increased.

But before accepting the identification, we need to ask how distinctive those characteristics actually are.

Was the name common?

Was the father's name common?

How precise is the date?

How large was the community?

Could there have been another family with the same combination of names?

Are the matching details independent, or were they copied from one record into another?

Similarity becomes more meaningful as independent identifying characteristics converge.

But convergence still needs to be tested against the rest of the person's life.


A person is more than a collection of matching fields


One of the limitations of database research is that it encourages us to think in fields.

Name.

Date.

Place.

Father.

Mother.

Spouse.

We search.

The database returns a result.

Five fields match.

Excellent.

But human identity is not a spreadsheet.

A person has a chronology.

If we believe that two sets of records belong to the same individual, their lives should be capable of existing within one reasonably coherent timeline.

Where was the person in 1925?

Where were they in 1935?

Who did they marry?

When did they migrate?

Where were their children born?

Which names did they use?

Who were their parents?

A candidate who matches four database fields but simultaneously appears to be living an incompatible life elsewhere may not be our person.

At that point, the similarities do not disappear.

They simply cease to be sufficient.


Similarity can become dangerous when we want it to be true


There is a psychological element to identification.

Sometimes we want a candidate to fit.

Perhaps we have been searching for hours.

Perhaps the client is waiting.

Perhaps the family has been looking for this person for decades.

Perhaps finding this individual would solve the case.

Then a promising record appears.

The relief is immediate.

Finally.

From that moment, there is a subtle risk.

We stop testing the candidate and start defending them.

A different spelling is explained.

A different date is explained.

A different place is explained.

Another inconsistency appears.

We explain that too.

Each explanation may be perfectly reasonable.

But eventually we need to ask:

Are we following the evidence, or are we protecting the match?

Forensic genealogy requires us to remain willing to lose a good candidate.

Even a very good one.


The strongest test is often contradiction


Once a promising candidate has been identified, I want to know what could prove the identification wrong.

This may sound counterintuitive.

Surely we should search for more matching information?

We should.

But we should also search for contradiction.

If I believe two records describe the same person, what else should be true?

Should the parents match?

Should the marriages align?

Should the migration timeline make sense?

Should the person disappear from one jurisdiction before appearing permanently in another?

Should their siblings, children or addresses provide independent points of connection?

And what evidence would make the identification impossible—or at least highly improbable?

A strong identification should survive this process.

If it does, the similarities become part of a much larger evidential structure.

If it does not, we have learned something equally important.


Proof comes from convergence


A reliable identification is rarely based on one spectacular match.

It is usually built gradually.

A name corresponds.

Then a date.

Then a parent.

Then a sibling appears in an independent record.

A migration document connects the two countries.

A marriage record confirms the parents.

An address appears in two unrelated sources.

A previously unexplained name variation is documented.

Slowly, independent pieces of evidence begin pointing towards the same conclusion.

This is convergence.

And it is fundamentally different from similarity.

Similarity says:

These two people look alike on paper.

Convergence says:

Independent evidence repeatedly connects these documentary identities in ways that are difficult to explain by coincidence.

That is what we are looking for.


And sometimes the evidence diverges


The opposite can happen.

The name matches.

The age is close.

The geography is plausible.

But as the research progresses, the documentary trails begin moving apart.

Different parents.

Different spouses.

Different addresses.

Different chronologies.

Separate families.

At that point, the correct response is not to keep explaining the differences until the original match can somehow be preserved.

It is to recognise that the evidence is diverging.

The similarity was genuine.

The identification was not.

And that distinction matters enormously.


Similarity opens the door


Similarity is essential to genealogical research.

Without it, we would not know where to begin.

Names, dates, places and relationships allow us to identify candidates and construct hypotheses.

But they are starting points.

Not conclusions.

A similar name tells us:

Look here.

A similar date tells us:

Investigate further.

A matching parent tells us:

This candidate deserves serious attention.

Only after the evidence has been examined collectively can we ask whether those similarities amount to a reliable identification.

Because two people can share a name.

They can share an age.

They can come from the same town.

They can even have parents with similar names.

Similarity makes an identification possible.

Evidence makes it sustainable.

And in forensic genealogy, those are not the same thing.



 
 
 

Comments


bottom of page