Showing posts with label DNA. Show all posts
Showing posts with label DNA. Show all posts

Wednesday, April 15, 2015

Getting More From Your Autosomal DNA: Genetic Family Trees

   For years, genealogists have been able to use Y-DNA to validate paternal pedigrees and sort surnames into family groups.  This has been a great advantage for the world of genealogy, but it has been restricted to men and paternal lines.  Autosomal DNA is more inclusive.  Both women and men can take this test and it illuminates the entire family tree as opposed to just the male line.  For those of us that have taken an autosomal test, there are a number of tools that help find cousin matches.  When we find multiple cousins matching the same chunk of DNA, we reach out to our new cousins and attempt to find a common ancestor in our trees.  Many times this is unsuccessful due to incomplete trees.  This is what is called a bottom-up approach.

   What if we used a top-down approach?  What if we started with your 10th great-grandmother?  You’d say autosomal DNA can’t go back that far.  That’s 12 generations ago and the DNA would be diluted to less than 1% of the original amount.  If autosomal DNA behaved mathematically, you’d be correct.  Autosomal DNA behaves more like Legos.  When we inherit DNA from our parents, it’s true that we get 50% from mom and 50% from dad.  That’s where the fairness ends.  When we look at what we inherit from our grandparents (through our parents), it is never 50/50.


   Instead, what we get from our grandparents is a random split.  In the case of the illustration above, this grandchild received a 54/46 split.  This is not uncommon.  See this Slate article.

   Our chromosomes behave like building blocks.  There is a tendency for genes located closely on a chromosome to be inherited together in a block.  This is called gene linkage.  There is no set size for these blocks; size is completely based on the genes that tend to stay together.  Segments around the 2 cM (centiMorgan) size have been found consistently (American Journal of Human Genetics).  The DNA we get from our grandparents come to us in large contiguous sections of hundreds of these blocks.  From generation to generation, the large sections are inherited randomly and unfairly, but the building blocks have a tendency to stay intact and not recombine.  With each generation, there is 50% chance of inheriting or not inheriting a specific block. 

   It’s possible that these 2 cM building blocks are about 25 generations old.  So, when we start with our 10th great-grandparents, they have lots of these blocks that they inherited from their parents and gave to their children.  What we can expect is that their descendants will have an assortment of these blocks from them and other ancestors.  When we examine the autosomal DNA for two dozen of their descendants, we find a set of genetic blocks in common.  No one descendant will have all the available genetic blocks an ancestor has left in the gene pool.  We may find five descendants sharing a block on chromosome one and seven descendants sharing a block on chromosome 12.  With DNA samples from two dozen descendants, about 15 ancestral genetic blocks can be identified.  All of the ancestral genetic blocks taken together uniquely identify your 10th great-grandparents as a couple.  Only their descendants would have this genetic block combination.  (Except in the situation where one set of siblings marries another set of sibling from a different family.)

   When we take the process a step further and analyze the next generation, we start to build a genetic family tree.


The table above shows the genetic blocks identified for Stephen Hopkins and each of his children that had descendants.  For simplicity, only one individual is listed for each column.  Remember that each column of genetic blocks actually represents a married couple: Constance Hopkins and Nicholas Snow, Deborah Hopkins and Andrew Ring, etc.  Each genetic block has a chromosome number and start and end locations.  Blocks in green represent inherited blocks from Stephen to his children.  As we build a genetic family tree, it now becomes possible to take a DNA sample from a living individual and match with Stephen Hopkins.  Once a match with Stephen is found, matches to his children can be checked to see which child the sample descends from.  Generations can be added to the genetic tree until known descendant DNA data has been exhausted.  In the Hopkins family, I was able to extend Constance’s line by a generation to Mary Snow and then then to Mary’s daughter, Mary Paine, before the data ran out.


   Similar to Y-DNA, these sets of genetic blocks (autosomal haplotype) can be used to identify genealogical relationships and sometimes the lack of relationships.  John Hopkins of Connecticut has often been connected as a son of Stephen Hopkins.  When we generate the autosomal haplotype for John and compare it to Stephen, we can see that there is no relation across the board.

   The red blocks indicate John’s DNA segments that have no corresponding segments with Stephen.  The yellow blocks indicate a similar chromosome location, but no genetic match.  Y-DNA gives us the ability to use DNA to see how brothers are potentially connected.  Now autosomal DNA gives us the ability to see how brothers and sisters are potentially connected.


   The autosomal haplotyping process is not a silver bullet that will solve all of our genealogy problems.  It will add to our toolkit as we validate family trees, work through brick-walls and attempt to solve genealogy mysteries.

Reference:

Maglio, MR (2015) Autosomal Haplotypes and the Genetic Reconstruction of Family Trees (Link)

© 2015 Michael Maglio and OriginsDNA. All Rights Reserved.

Thursday, March 19, 2015

Triangulated Small Segments are Identical by Descent

   Autosomal DNA segment matching is a complex issue.  Through testing and observation, it is obvious that some segment matches are false positives.  Computer algorithms will detect any matching allele with no knowledge that the allele is of paternal or maternal origin.


   If we said that the left columns are from the father’s sides and the right from the mother’s, we would see that none of the columns match.  Obviously, we can’t just draw a line down the middle and say one side is the mother’s DNA.  To determine which DNA came from mm and which came from dad, the autosomal results would need to be phased.  To phase the results of an autosomal sample it must be compared to at least one parent result.  By difference, the child result can be split into its paternal and maternal contributions. 


   If it were possible to phase every sample to be matched, false positives by computer algorithm would be eliminated.  Unfortunately, phasing every sample is not always possible.  A person’s parents may be deceased or even unknown.

   Another method of reducing or eliminating false positives is to triangulate each matching segment.  If a segment from autosomal sample A matches the corresponding segment from sample B and sample B matches sample C and sample C matches the original sample A, then the segment is considered triangulated and identical by descent.  How confident are we that the triangulated matches aren’t just a circular series of false positives? 

   Let’s look at segment on chromosome 3 that starts at rs6796502 and is 2.5 cM and 946 SNPs.  For this exercise, any chromosome segment could be used. 

Table 1.  Allele frequencies of 20 loci on chromosome 3.
   On that segment, there are 20 published locations with allele frequencies (NCBI).  Table 1 shows the how often a certain allele combination (AA, AC, AG etc.) appears for a European population.  Based on allele frequency, the most common combination of alleles in this section of chromosome 3 for a population of European descent is listed in Table 2.  I have artificially selected the most common combination to simulate a large portion of the population with European descent.  About 1 in 3,400 or about or about 300,000 people should have this combination. 

Table 2.  Predicted allele combination.
   Imagine for a moment that you roll six dice.  The first die comes up with a one and the second is a two and so on.  The probability of rolling a one on the first die is 1/6 (one side up on a six-sided die).  The probability of rolling a one and then a two is 1/6 times 1/6 or 1/36.  It will happen once every 36 rolls.  The combination illustrated on six dice would happen once in every 46,656 rolls.  Now imagine that is your DNA and we are looking for a match.  The other person would need one through six in the same order.  To calculate that probability we multiply 46,656 by 46,656 and get 2,176,782,336.  DNA matching actual has a better probability of matching.


   Table 3 lists the most common alleles again along with potential alleles that would generate a half match and the corresponding summed frequency.  The probability of the set of 20 potential combinations existing is equal to the product of the frequencies - 0.759.  This probability has to be extrapolated from 20 loci to 946, giving us 2.45x10-6 or 1 in 400,000.  There is a 1 in 400,000 chance of a completely random match on this section of chromosome 3 for the alleles with the highest frequency.  It is well within reason to expect false positives for this one-to-one match.

Table 3.  Probability of a half match within a European population.
   In the event of a three-way match (triangulation), we multiply by 2.45x10-6 again, giving us a probability of 1 in 167 billion.  Now we are outside of what is statistically reasonable.

   The most common set of European alleles doesn't produce the highest probability of a random match.  When the alleles are not the same (AC, AG, CT etc.), there is a higher chance of an autosomal half match.  Table 4 shows an actual set of alleles and the corresponding set of alleles to generate a half match.

Table 4.  Probability of a half match within a European population using actual sample.
   This actual sample takes us from a false positive probability of 1 in 400,000 to 1 in 5,900 (0.000169).  A probability of 1 in 5,900 indicates that we should be seeing completely random matches that have no genetic relationship on a regular basis.  Considering a population of about 1.6 million autosomal tests taken, each of us would have 270 false positive matches on a segment similar to the one shown.     

   Triangulated matches exist for this segment of chromosome 3.  For the probability of this triangulated segment, we multiply by 0.000169 again, giving us 2.87x10-8 or about 1 in 35 million.  Considering the number of results available for matching (about 1.6 million), it is not realistic that we are matching randomly.  In fact, most triangulated matches involve more than three test results.  If four test results are triangulated, the probability goes to 1 in 205 billion.  These probabilities indicate that triangulated results cannot be random and are matching due to common genetic descent.

   I have intentionally used two examples that have a higher probability of having false positive matches.  As soon as we look at matches that don’t have the higher frequency European alleles, the probability of a false positive diminishes. 

Table 5.  Probability of a half match within a European population with a Mediterranean sub-component.
   Table 5 shows a typical set of alleles.  There are two alleles at rs7630053 and rs4558783 that are not typical European and may indicate a Mediterranean ethnicity.  The probability of a one to one match on this segment being a false positive calculates to be 1 in 7 quadrillion. 

   Currently, we cannot examine the allele frequency for every SNP in every match we attempt.  When looking for autosomal matches consider phasing or triangulation.  Phasing the data is very valuable, yet the resources are not always available.  I’ve shown that triangulation eliminates false positives and those matches are statistically identical by descent.  Triangulated small segment matching is very valuable in our research.



References:

Maglio, MR (2015) Autosomal DNA and the Triangulation of Small Segments:  A Statistical Approach (Link)

© 2015 Michael Maglio and OriginsDNA.  All Rights Reserved. 

Thursday, March 5, 2015

Breaking Through the Autosomal DNA Generation Barrier: Connecting to Distant Ancestors

   There has been much debate over the use of small autosomal DNA segments.  It is important to understand where they come from and how they can be used for genetic genealogy.  Small segments are considered noise and false matches.  There are too many small matches to make sense out of, but they are not necessarily false matches.  These segments have been in the population for longer than we thought.  When I match someone at 2 cM it is very likely that they are a 12th cousin, not a 5th cousin.  There is no reason for us to look for small segment matches until we understand where these segments originated.

   When we talk about autosomal DNA, we often over simplify the process of genetic inheritance.  The simple answer is that we inherit half of our DNA from dad and half from mom.  The common message is that with every generation the DNA contribution from an ancestor is randomized and reduced until it is insignificant.  Genetic inheritance is actually much more complex than that.  Complex in a great way.  There is a tremendous amount of ancestral information that we are just beginning to tap into.

   We inherit DNA from our parents and their ancestors in large sections.  Take a look at the graphic below.  Each example is the comparison of a grandchild to a set of paternal grandparents.  You can see in the first example that the grandchild inherited over two-thirds of their grandfather’s first chromosome intact (blue bars).  The remaining section of the first chromosome is from their grandmother.  In the third example, the grandchild has inherited the entire chromosome 14 from their grandmother.  It is physically possible that this grandchild could someday give one of their children the grandmother’s complete chromosome 14.  


In an effort not to over simplify, this is just half the story.  That grandchild has an equal contribution from their maternal grandparents. 

   In the examples above, we can visualize what happens when DNA recombines.  The first example shows where one section of the grandfather’s DNA swapped places with the grandmother’s DNA before it was inherited by the grandchild.  This is called crossover.  In the examples, a) is a single crossover, b) is a double crossover and c) has no crossover.  On average, each of our chromosomes experienced 2 or 3 crossovers before we inherited them.

   Where DNA crossover takes place on a chromosome is not random.  There are approximate locations where the chromosome is more likely to split.  These locations are cleavage sites. 


These locations exist because there are groups of genes along a chromosome that have a tendency to stay together.  These groups are part of gene linkage.  These linked genes only allow for chromosome splits at either end of their linked section.  In my research, the minimum size for one of these gene-linked sections is about 2.5 cM.  These small segments then travel in larger groups.


   In the graphic above, the blue bar represents about a 60 cM match.  The intersection between the black and orange ovals is about 2.5 cM and represents a minimum segment.  In this crossover recombination, the large segment actually split to the right of the minimum segment.  In a future crossover, the chromosome could split on the left side of the minimum segment, giving a large segment bound by the orange oval.

   Why are these minimum segments important?  My research shows that these segments stay in the gene pool for dozens of generations.  Over time, naturally occurring SNP mutations take place.  These minimum inherited segments (MIS) can be differentiated into family groups.

   In my research, I started with 28 well known US colonial surnames and 393 autosomal kits.  For each surname, the associated kits were triangulated.  If three or more kits match on the same segment, you can deduce that it came from a common ancestor.  Each of the surnames investigated had 6 to 13 distinct triangulated segments.  Taken together, these triangulated ancestral segments represent an autosomal haplotype that can be used to identify a descendant’s genetic connection to an ancestor.  Across all of the surnames, these distinct segments appear at recurring locations on each chromosome.  I have listed 21 of these ancestral loci in my paper.

   Not all ancestral segments are the same type.  The segments can be categorized into three groups.  The first category is Common to All.  The surnames in this study are predominantly European.  One segment has been identified on chromosome 2 that triangulates across all surnames.  This segment correlates to a Western Atlantic ethnicity and I call it the Western Atlantic Autosomal Haplotype (WAAH).  The Western Atlantic Autosomal Haplotype should not be confused with ancestry informative markers (AIMs).  The WAAH is composed of about 800 SNPs and there are only about 100 AIMs SNPs in that same stretch of chromosome 2.

   The next category is Shared.  Some segments can be attributed to two or more surnames.  There was considerable intermarriage between US colonial families.  That period was a bottleneck genealogically and genetically.  As two major families married, their combined DNA segments entered the gene pool and were reinforced as their descendants intermarried. 

   The third category is Unique.  These shared segments cannot be attributed to intermarriage of families.  Yet the resulting familial autosomal haplotypes are not composed of a single surname.  In the case of Benjamin Franklin, the genetic proximity to his wife, Deborah Read and his mother, Abiah Folger, may make it impossible to distinguish between Folger, Franklin and Read DNA.  Therefore, the haplotype represents the combined inheritance.  

   Here is one of my case studies.   Augustine Bearse was born in England in 1618 and died in Barnstable, MA before 1697.  The Bearse family was chosen due to my familiarity with the genealogy and the debate surrounding Augustine’s wife.  His wife Mary was supposedly the granddaughter of the Chief of the Cape Cod Native American tribes.  The goal was twofold;  to identify the autosomal haplotype for the Bearse family and determine whether any of the ancestral segments had Native American ethnicity.

   The Bearse study was composed of 48 autosomal samples.  These samples were collected based on claimed genealogical connections.  The triangulated samples generated 8 ancestral loci and indicated an additional 5 loci that had the potential to triangulate with more samples.  The resulting Bearse autosomal haplotype is found below.

Bearse Autosomal Haplotype

   The Bearse haplotype contains the Western Atlantic Autosomal Haplotype (chromosome 2) which is common to all haplotypes in the study.  The other 12 loci are more valuable for genealogical validation.  One of the Bearse descendants triangulates on six of the ancestral segments.  It is highly unlikely that a descendant would match on all of the segments.  Although ancestral segments survive over the generations, the randomness of their distribution makes it difficult for any one person to have received them all.  Yet, triangulating on just one segment unique to Bearse is enough to indicate and validate a relationship.  Lack of a match could mean that an ancestral segment was not inherited or that a non-familial event (adoption, infidelity, etc.) has occurred and the individual’s family tree is incorrect.

   In order to investigate the origins of Augustine’s wife Mary, each ancestry segment from the haplotype was evaluated for ethnicity.  Only the segment on chromosome six at location 55850885 had any Native American ethnicity.  This ancestral segment had not fully triangulated, yet a few of the samples match exactly on Native American SNPs.  With additional samples, the segment could triangulate.  Once validated, the segment might be shared across multiple surnames or unique to Bearse, indicating Native American genes in the Bearse descendants.

   While the amount of autosomal DNA received by each successive generation is only half from each parent, that does not mean that given enough generations a distant ancestor’s genetic contribution will become negligible.  Through genetic linkage, portions of DNA are inherited intact.  Naturally occurring cleavage sites allow for ancestral segments averaging 2.5 cM to be passed from generation to generation as a minimum inherited segment (MIS). 

   Ancestral segment analysis is invaluable for the identification of distant ancestors.  All of the triangulated ancestral locations combine to become a Familial Autosomal Haplotype (FAH) that can be used to validate family history.

   Since finishing my initial research, I have gone on to identify over 50 ancestral loci and over 700 autosomal haplotypes for US colonial ancestors.  Stay tuned for further advances in autosomal research.

References:

Maglio, MR (2015) Minimum Inherited DNA Segment Size and the Introduction of Familial Autosomal Haplotypes (Link)

Website:

© 2015 Michael Maglio and OriginsConnector.  All Rights Reserved.


Monday, April 14, 2014

The DNA of Thomas Jefferson: [Insert Shocking Title Here]

   I’ve been creative with the titles of my articles in the past. It is the first thing people see and it better be eye catching. It’s been said that I'm ‘intentionally provocative’. I enjoy writing about topics that make people think. I draw the line at faulty logic. Many times, I'll draw conclusions from circumstantial evidence, but I always strive for a logical argument.

   Thomas Jefferson was part of the rare y-DNA haplogroup T (formerly K2) and much has been written about his genetics. Not everything written has been logical in its assumptions. Sometimes the story lines misinterpret the underlying science.



“Was Thomas Jefferson the first Jewish President?”

“Thomas Jefferson was Phoenician.”

“If Jefferson was Phoenician, then Charlemagne was also.”

“Thomas Jefferson could have recent origins in the Middle East.”

“Thomas Jefferson’s DNA traced back to Egypt.”

   In these situations the writers took a single data point and ran with it out of context.

   When we look at Jefferson’s DNA and compare it to available records, we only get a handful of matches that don’t tell a complete story. Let’s look at the first headline – was Jefferson Jewish? He didn’t practice Judaism and he wasn’t raised Jewish. He does have one genetic cousin who is a Moroccan Jew, but you have to go back about 2,000 years to find a common ancestor. While haplogroup T does have origins in the Middle-East, I wouldn’t say that it is definitely a Jewish haplogroup. If Jefferson were J1b2, there would be a stronger case tying him to the kohanim Jewish paternal lines. Jefferson also has a Belgian genetic cousin. Perhaps the headline should have been – Was Thomas Jefferson the first Belgian President? Not that exciting. Probably wouldn’t have sold very well.

   Thomas Jefferson was a Phoenician! There are many articles attributing this statement to Spencer Wells as part of his In Search of Adam program in 2005. I can’t find one quote that actual has Wells saying this. In fact in 2008 Wells argued that the Phoenicians were haplogroup J2. Jefferson’s haplogroup T is found in the same places and at the same times as the Phoenician Mediterranean colonies. This may indicate that Jefferson’s ancestors travelled with the Phoenicians as a peer or as a slave. I don’t think that ethnicity by association works.

   If Jefferson was a Phoenician, then so was Charlemagne. This is just plain and simple poor logic and a misunderstanding of genetics. As I mentioned, it doesn’t appear that haplogroup T is Phoenician. While Jefferson may be a descendant of Charlemagne, he is not a direct male descendant. You really need to be a direct male descendant to prove that an ancestor has the same y-DNA. One sample would never be enough to prove Charlemagne’s DNA. Multiple descendant samples and very strong genealogies are required to come close to determining an ancestor’s DNA. You never know where a non-paternal event may pop up.

   Could Jefferson have recent origins in the Middle East? This writer never actually defines recent. We are left to wonder if the Jeffersons lied on their Naturalization applications. Based on ‘time to most recent common ancestor’ calculations, I’d put Jefferson’s ancestors in the Middle East about 3,000 years ago. I guess that’s fairly recent compared to the age of the universe.

   Jefferson’s DNA traced to Egypt! One record match does not make an origin. That one Egyptian genetic cousin actually clusters better with other Moroccan records. This could indicate a back migration from Morocco to Egypt for that one person. A rule of thumb when determining origins is to find clusters of records. Jefferson does have a cluster on either side of the Strait of Gibraltar. This provides a strong argument that Jefferson’s ancestors came through that region and perhaps loitered there for a while. That doesn’t make it his origin.

   How should we define Jefferson’s origins? It is important to define origins with context. Where in Britain did the Jeffersons come from? One biographer puts Jefferson’s family origins in Wales.  Jefferson’s closest British genetic cousin comes from Yorkshire and the Jefferson surname has the highest distribution in Yorkshire. We are still talking about a single point of reference, so I won’t fall into the same trap and pronounce Thomas Jefferson a Yorkie. I can say that it appears that the Jeffersons were British and that the family had been in Britain for at least a 1,000 years. There’s just not enough data to be more certain.

   Jefferson’s tribal DNA does leave a sparse trail of breadcrumbs across Europe in the 1,500 to 2,000 years ago range. There are genetic matches that become increasingly more distant in Belgium, France and Spain. We could connect-the-dots and we probably wouldn’t be far off the migration path. For Jefferson’s European origins, we might say his ancestors were Iberian. If we go back another 500 years, the picture changes to a culture that traveled the Mediterranean. The genetic breadcrumbs are in Morocco, Sicily, Cyprus, Egypt and Turkey. We could talk about Jefferson’s haplogroup T origins. A cluster of data suggests a southern Arabian Peninsula origin about 8,000 years ago.



   If we continue backward in time, Jefferson’s ancestors came from East Africa just like everyone else on the planet. Which ‘origin’ you choose for Jefferson is completely up to your specific agenda. I use genetic genealogy to get a better understanding of the world that we live in and the things we have in common as one species. We may learn through DNA that our ancestors sacked Rome or pillaged the coast of England. That’s history, that’s fascinating, but that’s not who we are today. Unless you personally choose to embrace that history. Jefferson’s distant ancestors may have been Jewish, Phoenician or Egyptian, but that’s not who he was.

   We shouldn’t persecute for the sins of our ancestors or sit on the laurels of their accomplishments. We need to keep moving forward in a positive direction.

#gDNA

Tuesday, November 20, 2012

Stephen Hopkins: Saxon DNA?

   As we approach Thanksgiving, it’s a great time to write about our Mayflower ancestors.  So far, I have found two on my wife's side, Stephen Hopkins and Stephen Hopkins.  Ok, that’s really just one, but I have two lines that trace back to him.   This isn’t unusual, estimates put the count of Stephen Hopkins’ descendants at about 2 million Americans.

   What can Stephen Hopkins’ DNA tell us about his origins and his ancestors?  First, I should say that no one has a sample of Stephen’s DNA.  What we know about Stephen comes from tests completed by his male-line descendants with corroborating genealogical paper trails.  The Hopkins families are members of y-DNA haplogroup R1b, the largest genetic population in Europe.  R1b is often associated with the Celtic and Gallic tribes.  Hopkins’ DNA may be able to shed additional light on his birthplace, extend his genealogy further by tapping into an older family line or tell us about his deep ancestral origins.

   One of the first things I like to do is compare the haplotype, (the numeric markers from a y-DNA test) against a public database like ySearch.org.  The goal is to find other parallel lines of Hopkins with ancestry that predates Stephen.  This would allow us to work forward in time, connecting to Stephen and his father John, breaking through the current brick wall.  Unfortunately, no such records exist.

   What we do get from ySearch is list of genetic cousins and their ancestral locations.  Plotting these locations generates a distribution from Kent to Cornwall across southern England.  The highest concentration of cousins is in the historic Anglo-Saxon kingdom of Wessex.  The current research on Stephen Hopkins has him baptized in Hampshire, the heart of Wessex.

   What kind of R1b was Hopkins?  Was he a Celt, a Gaul, an Anglo-Saxon or something completely different?  One way to get close to the answer is to look at his genetic cousins again.  Since R1b is such a large group, it is important to focus on both the haplotype and SNP that defines his R1b subgroup.  The SNP that best defines Stephen is S493, which on the 2012 haplogroup tree is R1b1a2a1a1a2.  With the explosion of new SNPs identification and the rapidly expanding and changing subgroup nomenclature, researchers are advocating the use of the SNP rather than subgroup as a naming convention.  Let’s call Stephen Hopkins R-S493.

   When I take all these genetic cousins and run them through TribeMapper®, a pattern forms.  Ancestors start to pile up on either side of the English Channel and an approximate date of migration emerges.  Here’s where we pull out our history books.  If the date were about 2,500 years ago, I would say this was a Celtic migration.  If the date were 2,000 years ago, I might say these were Gaels fleeing the Romans.   The calculations come out to be about 1,500 years ago, putting this migration in line with the Anglo-Saxon invasion of Britain.

   Why stop there?  What flavor of Anglo-Saxon are we talking about?  Angle, Saxon, Jute?  The great thing about tribe mapping is that we can continuously turn back the clock and get a new picture.  If we find a Danish connection, then we might say Jutes or an association to the Angeln region of Germany, we could say Angles.  We have to be careful as those names and locations were just a snapshot in time when ancient historians catalogued Germanic tribes.  Those tribes, like all tribes, were just passing through.

   Stephen Hopkins’ DNA points to a genetic cluster in modern day Lithuania and Latvia.   This data most closely correlates to the Saxons and their origins on the Baltic coast.  Continuing this process gives us the following migration map.


   The R-S493 data takes us through Finland, Sweden and back to the mainland Europe to the Iberian Peninsula.  This puts the origin of R-S493 in Iberia about 4,000 years ago ± 500 years.

   We can’t be certain that Stephen Hopkins has Saxon DNA.  We can’t even say that all Saxons were haplogroup R1b.  It’s unlikely that they were a single homogenous ethnic group, but the core of the tribe would have had strong familial and genetic ties.  Were Hopkins’ ancestors at the core of this tribe or part of the fringe, picked up along the way?  A broader study of DNA associated with the same places and times would be required to answer that question.

   If we look at the surname Hopkins, its origins are from Hobbes-kin and even further back to the Germanic name Hrodberht.  Stephen Hopkins and his closest genetic cousins are found in the historic Kingdom of Wessex (West Saxons).  Time-wise, there is a correlation to the Anglo-Saxon invasion of Britain.  We can even make a connection to the proto-Saxons along the Baltic coast.  I’m going out on a limb and calling Hopkins a Saxon.

   That Saxon bloodline remained adventurous and served Stephen well as he voyaged to Bermuda, Jamestown and Plymouth colony.

   It’s never obvious where DNA will lead.  Each tribe mapping is an adventure in itself.


© Michael R. Maglio and OriginsDNA

Thursday, November 15, 2012

Myles Standish: Mayflower DNA

My genealogy has one Mayflower passenger, Stephen Hopkins. Seven other passengers are cousins in one manner or another, Doty, Howland, More, Mullins, Standish, Warren and Winslow. Twenty-four of the Mayflower families have living descendants. I have collected the y-DNA records for fifteen of them. (For more on this process watch this short video) It was no surprise to find eight R1b Celts and four I1 Scandinavians among them. But, the three I2a Balkans intrigued me.

The one name that stood out as I2a was Myles Standish. Every first grader knows that name. My first thought was that Myles was descended from a member of the Roman Legions. Perhaps he was a Scythian or Sarmatian. I needed to identify the Standish family tribe and when they arrived in England. If I was lucky, I’d be able to bracket the immigration of his ancestor to the 1st or 2nd century, the height of the Roman conquest.

As the DNA records started to compile using TribeMapper® analysis, an initial pattern developed showing historic habitation on either side of Hadrian’s Wall. This was the beginning of a great migration story and potentially the end to the dispute of Myles Standish’s origins. Researchers have placed Standish’s birthplace as either Lancashire or the Isle of Man. Based on the data, Lancashire emerges as the most likely location. There was no genetic indication that the Isle of Man was a possibility.

If Standish’s ancestors had been conscripted into the Roman Legion, then I would expect their migration pattern to appear scattered like a diaspora. Fathers and brothers and their descendants would be spread across the Roman empire. There would be no focus for the data points representing the period 2,000 years ago.

The actual data points told a different story. They remained focused. At the end of the last ice age, about 10,000 years ago, Myles Standish’s ancestors were living in the Balkans. As the ice receded, they journeyed up the Danube River, a major migration highway, until they reached the upper Rhine. The upper Rhine was a Neolithic way station for many tribes coming up the Danube or out of Iberia. The area served as a stopover before continuing over the Alps or down the Rhine. The Standish tribe chose to follow the Rhine down to the North Sea.


Between 2,000 and 3,000 years ago, Standish’s ancestors crossed into England and made their way up the Thames to its source. My theory is that they were pushed ever westward by successive waves of immigrants. They found Wales to be well populated already and ventured north to where we find the most recent genetic evidence, in Lancashire.

My initial theory that Standish’s ancestor was brought to England as part of the Roman Legion, to reinforce the troops at Hadrian’s Wall, was wrong. It is always good to have a theory to work toward, but don’t let preconceived ideas get in the way of new evidence. Now that we know that Standish’s origins are pre-Roman we can consider that his family is one of the native tribes of Britain. The most likely Lancashire tribe would be the Setantii, which is a sub-tribe of the Brigantes.

Each one of our ancestors has a unique migration story to tell. Their travels overlap with events that we have read about in history books.

Where did you come from?

© Michael R. Maglio and OriginsDNA

What’s in My gDNA Toolbox

If I were talking about my regular genealogy toolbox, I would be listing links to all the great websites with digital records (e.g. FamilySearch). I would also talk about great repositories like NARA, BPL or the Mass Archives. Or, I would mention tips and techniques like Nearest Neighbor and the Hidden Treasures in old photos.



Now that we are adding DNA as a tool for genealogy, we have to pack a new toolbox.

The Databases – record sources to compare your DNA against

· Ysearch.org – Y-DNA database
· Mitosearch.org – mtDNA database
· FTDNA.com – DNA Project database
· WorldFamilies.net – DNA Project database
· SMGF.org – DNA Project database

The Testing Companies – many different testing companies that are not all equal – do your homework

· FTDNA.com – DNA testing (my favorite)
· 23andMe.com – DNA testing
· SMGF.org – DNA testing
· Ancestry.com – DNA testing
· GeneTree.com - DNA testing

Sources of gDNA Knowledge – There are many areas of genetic genealogy that are open for interpretation. Read everything and come to your own conclusions.

· ISoGG.org – Advocates for the use of genetics as a tool for genealogical research
· Wikipedia - Haplogroup details
· nationalgeographic.com/genographic

Analysis Tools – DNA results love to be compared and analyzed

· hprg.com/hapest5/index.html – Whit Athey’s Haplogroup predictor
· mymcgee.com/tools/ - Dean McGee’s Y-DNA comparison tools
· www.math.mun.ca/~dapike/FF23utils/ - David Pike’s autosomal comparison tools
· http://gedmatch.com/ - Autosomal comparison tools
· PHYLIP – phylogenetic tree creation

DNA Data Management – you need to organize and manage your DNA records

· Legacy Family Tree – supports DNA records (the one I use)
· Family Tree Maker, RootsMagic, Ancestral Quest and The Master Genealogist – supports DNA
· Excel – spreadsheet tools

Misc
· Google Maps – User defined maps – you never know when you might want to build your own custom map

This is hardly an exhaustive list. I use most of these tools on a weekly basis. I’m always looking for new tools (or creating ones that don’t exist).

What's in your toolbox? Let me know what tools you are using.

#gDNA

Armenia, DNA and Ethnicity

Self-identity, what culture or ethnicity do you identify with? Your current culture? Your immigrant ancestor’s culture? Perhaps you identify with a culture buried deep in your DNA.

See this article as a great primer on the differences between - Ethnicity, Nationality, Race, Heritage, and Culture.

Culturally, my wife is an American. I could even say that she is a New Englander. She grew up in an Armenian family, but she doesn’t know the language. What she does identify with is the food and family. Her immigrant grandfather, Reuben, was born in Turkey. Turkish was his nationality, but culturally he associated deeply with the Armenian heritage that was strong in Adana.

Nations redraw their lines, form and dissolve over the course of decades. If you had lived in central Europe over the past few hundred years, one day you might be French and the next day German, only to be French again in a week.

How long does it take us to lose our ethnicity? If I took my family to Armenia and we stayed there for three or four generations, would they think of themselves as Armenian American Armenians. I doubt it. Each generation would absorb the culture around them to a greater degree. Given enough time, some descendants might think that it was just family mythology that they ever lived in the US. We’ve always been here.

We are all immigrants or descendants of immigrants. That goes for the entire planet.

If I look at another side of my wife’s family, they’ve been in America for over 350 years. Their immigrant ancestor, Edward Clark, was ethnically English. In turn, Edward’s immigrant ancestor was Norman and the immigrant ancestor before that was Danish. I can keep going back, Iberia, Asia and Africa. Which culture should they identify with? Nationality is fleeting and uncertain. Ethnicity is in your genes, embrace all the cultures of your ancestors.

About 50,000 years ago, there were no humans in Armenia, or for that matter, Asia Minor. Over the intervening years, folks trickled in from every direction. Let’s look at the current distribution of Armenian y-DNA.


Haplogroup
Percentage
Culture
J1c & J2a
32%
Arabic / Semitic
R1b
25%
Iberian/Gaul
G2a
14%
Caucasus
E1b
8%
Alexandrian
I2a & R1a
8%
Balkan
T1
6%
Mid-eastern
L2a
4%
Dravidian
Q1
1%
Hun


This is a snapshot of modern Armenia. Without analyzing individual haplotypes from this dataset, it is difficult to determine which group arrived first. More than likely each group had multiple waves of immigration across history. I’ve created the map below for you to get a feel for the origin and flow of the major haplogroups.


It’s not unusual for groups J1c, J2a and G2a to have high percentages. Those groups also have their origins in that region. The large portion of R1b can be attributed to the crusaders passing through for hundreds of years. Many of the taverns in this region have signs that say, ‘Alexander the Great slept here’. His empire would have contributed the E1b DNA as they conquered eastward and the Dravidian DNA flowed back toward Greece with the spoils. The Roman and Byzantine influence brought the Balkan DNA. The Huns also stopped by on their way to conquer Eastern Europe.

My wife can count Armenian as part of her heritage, with roots on the Mediterranean coast of Turkey. Someday I will find her living Armenian cousins in order to get DNA tests. Those results will allow me to identify her deeper ancestral ethnicity.

On another line, she is descended from four generations of Sea Captains from Maine with Scottish origins. Should my wife self-identify with all the cultures of her ancestors? Probably not. Should she learn about and understand all those cultures? Definitely. We can pick and choose the best parts of our ancestral heritage and create our own unique ethnic identity. She has a love for the ocean that didn’t come from any early family experience. It’s in her DNA.

Your Mother's Mother

Deep Into DNA*

Your mother’s ancestry can be extremely challenging. Most of us live in a patrilineal society. The wife and the children take the surname of the husband. History shows us that recording the maiden name of the wife was often an afterthought. We have all tried traditional genealogy for our mother’s line. Some of us can go back a couple of generations and have hit brick walls. Some of us have researched a dozen generations. Mitochondrial DNA testing can aid a genealogist in discovering those lost surnames and validating your research.

Two months ago, I wrote about y-DNA and its use in tracing your paternal line. Mitochondrial DNA testing looks at your maternal line. There are many similarities and just as many differences between the two tests.

...continued at The In-Depth Genealogist with a free membership.

*The Deep Into DNA article series is published each month in The In-Depth Genealogist Newsletter and will demystify genetic genealogy and make sense out of DNA testing terminology. Each month we will talk about the types of tests available from major labs and show relevant examples on how to use DNA in your genealogy research.

#gDNA

Your Father's Father

Deep Into DNA*
 

Raise your hand if you’d like to know more about your surname and your father’s ancestry. I’m raising mine!

Most of us live in a patrilineal society. The wife and the children take the surname of the husband. We can’t help but to associate with our father’s family, his clan. The males in the family will inherit the Y chromosome virtually unchanged, though genetically, we can only attribute a small fraction of our overall DNA to that patrilineal line. Psychologically though, 50% of our ethnic identity comes from dad.




We have all tried traditional genealogy for our father’s line. Some of us go back a couple of generations and some of us have researched a dozen generations. A few of us are adopted and know nothing about our paternity. Y-DNA testing has benefits that aid a genealogist in all these situations...

...continued at The In-Depth Genealogist with a free membership.

*The Deep Into DNA article series is published each month in The In-Depth Genealogist Newsletter and will demystify genetic genealogy and make sense out of DNA testing terminology. Each month we will talk about the types of tests available from major labs and show relevant examples on how to use DNA in your genealogy research.

#gDNA

Friday, September 21, 2012

Vandals DNA: Leaving Genetic Graffiti Across Europe

   I have been fascinated by barbarian history since my sixth grade Social Studies teacher, Mr. Rose, handed me the textbook on the subject.  I still read everything I can on the topic.  The biggest draw is the mystery of where each tribe came from, appearing out of the shadow of mythology and for many, disappearing into obscurity.  I’m finding that DNA can help answer the questions of ‘Where are they now?’ and ‘Where did they come from?’

   The word barbarian is first seen in the Greek language as barbaros.  One possible origin of the word is that it was coined from the sound of the language used by these nomadic tribes – ‘bar bar bar’.  The Vandals (Vandali) appear in the history books as early as 166 AD and are described as an East Germanic tribe.  The theories on their origins include having Scandinavian roots in the parish of Vendel, Sweden or Germanic roots with a connection to the word meaning to wander (wandeln).

   Even if the Vandals’ name wasn’t derived from the word wandeln, wandering was what they did the most.  Perhaps chased is a better description.  Around 300 AD, the Goths fought with the Vandals and pushed them west along the Danube River.  By 400 AD, the Huns pushed the Vandals further west to the Rhine River.  At the Rhine, the Vandals fought with the Franks, won and moved into Aquitaine (western France), pillaging and plundering the whole way.

Historical Vandali migration

   I’m reminded of an old cartoon where the barbarian leader is addressing his troops – ‘This time, remember: pillage, then burn.’  The term vandalism is directly attributed to the Vandals’ ruthless pillaging and destruction of culturally significant objects.

   In 409 AD, the Vandals crossed into the Iberian Peninsula, only to be chased by the Visigoths and the Roman army into North Africa in 429 AD.  Over the next decades, the Vandals conquered North Africa and made Carthage their capital.  From there they invaded Sicily, Sardinia, Corsica and in 455 AD they sacked Rome. Fortunes fade, the Vandals were defeated by the Byzantine Empire in 534 AD and disappeared into history.  The name Vandals may have disappeared but the people didn’t.  They were either assimilated into the local cultures or dispersed as slaves, the spoils of war.

   To figure out ‘Who were the Vandals?’, first I had to figure out ‘Where are they now?’  Based on their historic origins, the Vandals would probably fall into a small group of y-DNA haplogroups.  The ethnic descriptions that I’m using are overly simplistic, just enough to give you a feel for the possible cultures present.


Haplogroup
Culture
G2a
Caucasian
I1
Scandinavian
I2a
Danubian
N
Finn
R1a
Balkan
R1b
Celtiberian


   We can’t assume that the Vandals were genetically homogenous.  At various times they were associated and allied with the Alani and Suevi tribes.  Any DNA trail found, could just as easily belong to a group along for the ride.  I started researching all these haplogroups along the Vandals’ 400-year migration route to find a DNA footprint. The key datasets included records from Sicily, Sardinia, Tunisia, Spain, France and Germany.  These locations create a triangle of migratory patterns, clockwise, counterclockwise and dispersion.  If these sets have evidence of Vandals DNA, there should be a counterclockwise flow around Europe and a west to east flow across the Mediterranean.  Immediately I was able to remove haplogroup N from the running, there were no records.


   I have to admit that going into this project I thought that the Scandinavian I1 haplogroup would be my most likely suspect.  I had built a picture in my mind that all the Germanic tribes had come out of Scandinavia.  I researched this group first, looking for genetic flow across time. Usually I work with an individual record and trace backward through time.  For this project, I’m analyzing large groups in multiple datasets.  At a high level I’m searching for DNA that has migrated from Germany, through France and Spain to Tunisia and then to Sicily and Sardinia.

   Using TMRCA (time to most recent common ancestor) and my own TribeMapper techniques, I was able to identify that the Scandinavian I1 haplogroup formed a parallel dispersion pattern.  The y-DNA genes flowed from Germany down into the Italian peninsula and into the Iberian Peninsula at roughly the same time.  Haplogroups G2a and R1a also fall into this category of dispersion.  These DNA records don’t fit the pattern of the Vandals’ migration.

I1, G2a & R1a migration pattern

   The R1b Celtic-Iberian haplogroup proved more difficult to decipher.  This is a major group in Europe, representing over 60% of the Western European population.  For over 10,000 years, there has been a strong flow out of the Iberian Peninsula toward the British Isles, Germany and Scandinavia.  Detecting a counterflow against that tide is problematic.  The apparent direction of migration across the datasets I researched shows a clockwise pattern out of Spain, into France and Germany and then down the Italian peninsula.  R1b doesn’t look like our Vandals, but I’m going to reserve judgment until better analysis tools are developed.

R1b migration pattern

   I’ve left haplogroup I2a, the Danubians, for last because they have the best correlation to the Vandals.  Their genetic migration does show a counterclockwise flow from Germany, through France and Spain and into Sicily and Sardinia.  This DNA can be found in the historic Vandali regions of Aquitaine, Galicia, Lusitania and Andalusia.  Haplogroup I2a is a dominant Germanic group associated with the Danube River, giving them the nickname Danubian.  This matches the Vandals earliest historical references.

I2a migration pattern

   Myles Standish of Mayflower fame was also haplogroup I2a.  Standish is not closely related to the DNA that I am chasing.  His tribe and the Vandals parted ways over 5,000 years ago.

   Have I found the Vandals, Alans or Suevi?  So far, I have been looking at the past 2,000 years.  If I expand the datasets and research back further in time, additional patterns appear.   The DNA takes me to Georgia, Armenia and Iran.  The timing and the location of these records put us in the historic Alani homeland on the Asian steppe.  The Huns were also responsible for driving the Alans west into Europe around 300 AD.  I’ve been looking for Vandals and I’ve found the Alans instead.


   The I2a1 tribal haplotype that I have identified has remained relatively unchanged for thousands of years and has allowed me to follow a migration in and out of Asia and across Europe.  I cannot say that I have found the DNA of all the Alans or that the Alans were only haplogroup I2a.  The correlations that I have made are based on records currently available and it is impossible to say what additional future DNA records may reveal.

   The mystery of the Vandals remains a mystery.  I now have an unexpected peek at a piece of the Alani origins and migrations.  When clients want to know more about their DNA, I can check them against this data.  I’d love to be able to get to a point where I can tell folks - ‘Hey, you’re a Visigoth!’  One of the biggest parts of DNA testing beyond finding family, is connecting ourselves to history, knowing that your ancestors played a role.

© Origin Hunters & OriginsDNA