Showing posts with label haplotype. Show all posts
Showing posts with label haplotype. Show all posts

Thursday, March 5, 2015

Breaking Through the Autosomal DNA Generation Barrier: Connecting to Distant Ancestors

   There has been much debate over the use of small autosomal DNA segments.  It is important to understand where they come from and how they can be used for genetic genealogy.  Small segments are considered noise and false matches.  There are too many small matches to make sense out of, but they are not necessarily false matches.  These segments have been in the population for longer than we thought.  When I match someone at 2 cM it is very likely that they are a 12th cousin, not a 5th cousin.  There is no reason for us to look for small segment matches until we understand where these segments originated.

   When we talk about autosomal DNA, we often over simplify the process of genetic inheritance.  The simple answer is that we inherit half of our DNA from dad and half from mom.  The common message is that with every generation the DNA contribution from an ancestor is randomized and reduced until it is insignificant.  Genetic inheritance is actually much more complex than that.  Complex in a great way.  There is a tremendous amount of ancestral information that we are just beginning to tap into.

   We inherit DNA from our parents and their ancestors in large sections.  Take a look at the graphic below.  Each example is the comparison of a grandchild to a set of paternal grandparents.  You can see in the first example that the grandchild inherited over two-thirds of their grandfather’s first chromosome intact (blue bars).  The remaining section of the first chromosome is from their grandmother.  In the third example, the grandchild has inherited the entire chromosome 14 from their grandmother.  It is physically possible that this grandchild could someday give one of their children the grandmother’s complete chromosome 14.  


In an effort not to over simplify, this is just half the story.  That grandchild has an equal contribution from their maternal grandparents. 

   In the examples above, we can visualize what happens when DNA recombines.  The first example shows where one section of the grandfather’s DNA swapped places with the grandmother’s DNA before it was inherited by the grandchild.  This is called crossover.  In the examples, a) is a single crossover, b) is a double crossover and c) has no crossover.  On average, each of our chromosomes experienced 2 or 3 crossovers before we inherited them.

   Where DNA crossover takes place on a chromosome is not random.  There are approximate locations where the chromosome is more likely to split.  These locations are cleavage sites. 


These locations exist because there are groups of genes along a chromosome that have a tendency to stay together.  These groups are part of gene linkage.  These linked genes only allow for chromosome splits at either end of their linked section.  In my research, the minimum size for one of these gene-linked sections is about 2.5 cM.  These small segments then travel in larger groups.


   In the graphic above, the blue bar represents about a 60 cM match.  The intersection between the black and orange ovals is about 2.5 cM and represents a minimum segment.  In this crossover recombination, the large segment actually split to the right of the minimum segment.  In a future crossover, the chromosome could split on the left side of the minimum segment, giving a large segment bound by the orange oval.

   Why are these minimum segments important?  My research shows that these segments stay in the gene pool for dozens of generations.  Over time, naturally occurring SNP mutations take place.  These minimum inherited segments (MIS) can be differentiated into family groups.

   In my research, I started with 28 well known US colonial surnames and 393 autosomal kits.  For each surname, the associated kits were triangulated.  If three or more kits match on the same segment, you can deduce that it came from a common ancestor.  Each of the surnames investigated had 6 to 13 distinct triangulated segments.  Taken together, these triangulated ancestral segments represent an autosomal haplotype that can be used to identify a descendant’s genetic connection to an ancestor.  Across all of the surnames, these distinct segments appear at recurring locations on each chromosome.  I have listed 21 of these ancestral loci in my paper.

   Not all ancestral segments are the same type.  The segments can be categorized into three groups.  The first category is Common to All.  The surnames in this study are predominantly European.  One segment has been identified on chromosome 2 that triangulates across all surnames.  This segment correlates to a Western Atlantic ethnicity and I call it the Western Atlantic Autosomal Haplotype (WAAH).  The Western Atlantic Autosomal Haplotype should not be confused with ancestry informative markers (AIMs).  The WAAH is composed of about 800 SNPs and there are only about 100 AIMs SNPs in that same stretch of chromosome 2.

   The next category is Shared.  Some segments can be attributed to two or more surnames.  There was considerable intermarriage between US colonial families.  That period was a bottleneck genealogically and genetically.  As two major families married, their combined DNA segments entered the gene pool and were reinforced as their descendants intermarried. 

   The third category is Unique.  These shared segments cannot be attributed to intermarriage of families.  Yet the resulting familial autosomal haplotypes are not composed of a single surname.  In the case of Benjamin Franklin, the genetic proximity to his wife, Deborah Read and his mother, Abiah Folger, may make it impossible to distinguish between Folger, Franklin and Read DNA.  Therefore, the haplotype represents the combined inheritance.  

   Here is one of my case studies.   Augustine Bearse was born in England in 1618 and died in Barnstable, MA before 1697.  The Bearse family was chosen due to my familiarity with the genealogy and the debate surrounding Augustine’s wife.  His wife Mary was supposedly the granddaughter of the Chief of the Cape Cod Native American tribes.  The goal was twofold;  to identify the autosomal haplotype for the Bearse family and determine whether any of the ancestral segments had Native American ethnicity.

   The Bearse study was composed of 48 autosomal samples.  These samples were collected based on claimed genealogical connections.  The triangulated samples generated 8 ancestral loci and indicated an additional 5 loci that had the potential to triangulate with more samples.  The resulting Bearse autosomal haplotype is found below.

Bearse Autosomal Haplotype

   The Bearse haplotype contains the Western Atlantic Autosomal Haplotype (chromosome 2) which is common to all haplotypes in the study.  The other 12 loci are more valuable for genealogical validation.  One of the Bearse descendants triangulates on six of the ancestral segments.  It is highly unlikely that a descendant would match on all of the segments.  Although ancestral segments survive over the generations, the randomness of their distribution makes it difficult for any one person to have received them all.  Yet, triangulating on just one segment unique to Bearse is enough to indicate and validate a relationship.  Lack of a match could mean that an ancestral segment was not inherited or that a non-familial event (adoption, infidelity, etc.) has occurred and the individual’s family tree is incorrect.

   In order to investigate the origins of Augustine’s wife Mary, each ancestry segment from the haplotype was evaluated for ethnicity.  Only the segment on chromosome six at location 55850885 had any Native American ethnicity.  This ancestral segment had not fully triangulated, yet a few of the samples match exactly on Native American SNPs.  With additional samples, the segment could triangulate.  Once validated, the segment might be shared across multiple surnames or unique to Bearse, indicating Native American genes in the Bearse descendants.

   While the amount of autosomal DNA received by each successive generation is only half from each parent, that does not mean that given enough generations a distant ancestor’s genetic contribution will become negligible.  Through genetic linkage, portions of DNA are inherited intact.  Naturally occurring cleavage sites allow for ancestral segments averaging 2.5 cM to be passed from generation to generation as a minimum inherited segment (MIS). 

   Ancestral segment analysis is invaluable for the identification of distant ancestors.  All of the triangulated ancestral locations combine to become a Familial Autosomal Haplotype (FAH) that can be used to validate family history.

   Since finishing my initial research, I have gone on to identify over 50 ancestral loci and over 700 autosomal haplotypes for US colonial ancestors.  Stay tuned for further advances in autosomal research.

References:

Maglio, MR (2015) Minimum Inherited DNA Segment Size and the Introduction of Familial Autosomal Haplotypes (Link)

Website:

© 2015 Michael Maglio and OriginsConnector.  All Rights Reserved.


Wednesday, January 28, 2015

Ghosts of DNA Past: Irish Kings

   In 2006, Laoise T. Moore and the folks at Trinity College in Dublin published a paper famous for identifying the modal haplotype of Irish High King Niall of the Nine Hostages.  In their work, they used seventeen Y-DNA STR markers.  While time to most recent common ancestor (TMRCA) calculations have accuracy issues, having only 17 markers gives a common ancestor over 2,000 years ago.   What the Trinity folks really accomplished was the identification of Niall’s paternal ancestor from over 400 years earlier.  The media in 2006 had a field day in their interpretation that most of Ireland is descended from Niall.  “Niall may be the most prolific male in Irish history.”  Also at 17 markers, there is a very high probability of convergence.  Through normal mutations, haplotypes can change over time to appear similar or identical to other haplotypes.  The lower the number of markers, the higher the chance of convergence.  At that time only high level SNPs were tested to determine haplogroup.  Without terminal SNPs it would have been impossible to recognize convergence, if it existed in the samples.

   In my research on the Kings of Ireland, I have used 67 markers to reduce the chance of convergence and to calculate the age of common ancestors on the descendant side of the target rather than the ancestor side.  I will demonstrate traditional median-joining networks and novel “tribal” markers for the identification of four historic Kings of Ireland.  Did Trinity get Niall’s haplotype correct with the limited data they had at the time?

Ghost:  a manifestation of a dead person

Modal haplotype:  a derived haplotype based on the DNA tests of a group of people

   A modal haplotype is a ghost of a person.  When we look at multiple DNA test results and calculate the mode, by definition we are just taking the values that appear most often.  There is no way to determine if the modal haplotype is the actual haplotype of the historic individual we are researching (short of historic samples).  While the modal is not perfect, it will be close enough at 67 markers for us to determine the genetic “ghost”.

   The septs of Ireland provide us an opportunity to develop genetic genealogy techniques and processes.  Irish surnames are typically patronymic.  The surnames generally take the form of Mac Cárthaigh (McCarthy), meaning son of Cárthaigh or Ui Néill (O’Neill), meaning grandson / descendant of Néill.  Irish septs serve as a collective of related families with shared ancestry and patronymic surnames.  Multiple septs then belong to larger dynasties such as the Eóganachta and the Dál gCais.

   If septs are patrilineal, then Y-DNA haplotypes should be consistent across sept surnames.  Research on the Uí Néill haplotype started with a geographical selection and then a subsequent reduction by sept surnames (Moore et al 2006).  For each target sept, affiliated surnames were identified.  In the case of Uí Néill, the following surnames and associated Y-DNA STR records were accessed from Family Tree DNA projects: O’Neill, Gallagher, Doherty and O’Donnell.  The selection includes 600 records and 5 common European haplogroups.

   Median-joining networks have been in use for over a decade for the visualization of genetic relationships.  The use of them at 67 STR markers has been rare, but it should be the norm.  This first image has the central cluster of a median joining network based on 25 STR markers from the Uí Néill group.  It is just a single cluster with no differentiation.



Figure 1 - Using only 25 STR markers, the Uí Néill network collapses to a single cluster.

When we look at the same group using 67 markers, we get four distinct clusters, each with their own SNP.  The cluster at the far right is predominantly R-L159 and the cluster at the lower right has R-P311/R-L151 nodes.  The cluster at the left contains all of the Uí Néill dynastic surnames, has the majority of nodes and is SNP R-M222, which is consistent with earlier studies.


Figure 2 - View of the Uí Néill network torso showing four distinct clusters.  Three groups on the right are O’Neill only.

As a double check to make sure that I wasn’t seeing some other phenomena, I analyzed three random Irish surnames; Duffy, Kelly and McCormick.  The random sample produced over ten unique clusters with no surname overlap.  This comparison shows that septs are patrilineal and that Y-DNA haplotypes are consistent across sept surnames. 

Figure 3 - Median-joining network of yDNA sampled from three random Irish surnames; Duffy, Kelly and McCormick.  

Re-evaluating the Uí Néill data also shows that Trinity was correct in their identification of a 17-marker Uí Néill haplotype.  New data and new techniques allow us to produce a 67-marker haplotype.


Figure 4 - Sixty-seven STR Uí Néill Modal Haplotype (Niall of the Nine Hostages).

   A different technique that I’d like to illustrate involves the fact that not all STR markers are created equal.  This method takes advantage of “slow” mutating STR markers.  Each marker has its own mutation rate.  By selecting the 15 “slowest” markers with an average mutation rate of 0.00024, a virtual tribal haplotype is created that would be stable within the last 2,000 years (90% probability of 80 generations).  This is an order of magnitude lower than the average rate of 0.0029 used as a constant in typical TMRCA calculations.  The “tribal” markers isolated are DYS426, DYS388, DYS392, DYS455, DYS454, DYS578, DYS590, DYS641, DYS472, DYS594, DYS436, DYS490, DYS450 and DYS640.

   To manipulate the “tribal” haplotype of 15 microsatellites faster the resulting values are concatenated into a string – ex. 12121411119168108101212811.  The “tribal” haplotypes are summarized per surname and plotted to illustrate majority and affinity.


Figure 5 - Uí Néill dynastic haplotypes converted into 15 marker “tribal” haplotypes and summarized.

   The Uí Néill dataset resolved into 37 unique “tribal” haplotypes.  Figure 5 shows that haplotype 12121411119168108101212811 is the most dominant across the Uí Néill surnames.  As with the median-joining network analysis, this “tribal” haplotype is consistent with SNP R-M222. 

   I repeated these two techniques for the Uí Briúin sept using the following surnames and associated Y-DNA records: O’Brien, Hogan, Kennedy and McMahon.  The selection includes 615 records.  The Mac Cárthaigh dataset has the following surnames: McCarthy, Callaghan, Donovan and Sullivan.  The selection includes 319 records.  The Ua Conchobhair data has the following surnames: O’Connor, McManus, Reilly and Rourke.  The selection includes 352 records.

For more details, see my paper at Academia.edu.



Figure 6 - Sixty-seven STR Uí Briúin Modal Haplotype (Brian Boru).


Figure 7 - Sixty-seven STR Mac Cárthaigh Modal Haplotype (McCarthy Eoganachta Kings).



Figure 8 - Sixty-seven STR Ua Conchobhair Modal Haplotype (Last High King Roderick O'Connor).


   Here are a couple of interesting insights from my research.  Niall Noígíallach was High King of Ireland around 378 CE and founder of the Uí Néill dynasty.  Historically, his half-brother Brión, was one of the founders on the Connachta dynasty and an ancestor of the last High King of Ireland, Ruaidrí Ua Conchobair.  If their genealogies are correct, the evidence is in their descendant’s DNA.  The data shows that Uí Néill and Ua Conchobair share the same SNP, R-M222.  The Uí Néill and Ua Conchobair modals are a 6-step match at 67 markers.  There is a 99% probability of a relationship not further than 1,260 years ago.  The results make a strong case for the validity of this historic genealogy.

   Brian Boru, High King of Ireland in 1002 CE, belonged to the Dál gCais dynasty and Tadhg Mac Cárthaigh, the first King of Desmond, belonged to the Eóganachta dynasty.  Ancient genealogies have the Eóganachta and Dál gCais dynasties descended from Ailill Aulom, the son-in-law of legendary king Conn of the Hundred Battles.  The Mac Cárthaighs and Uí Briúins do not share the same SNP (R-L226 vs. R-CTS4466), but by descent they would share a common R-DF13 ancestor.  The Mac Cárthaigh and Uí Briúin modals are an 11-step match at 67 markers.  There is a 99% probability of a relationship not further than 1,920 years ago.  This puts a Mac Cárthaigh-Uí Briúin common ancestor as a contemporary of the legendary Conn.

   New and improved genetic genealogy techniques are invaluable for the identification of historic individuals and the reconstruction of distant family trees at the macro level.

Reference:


Maglio, MR (2015) Identifying Y-Chromosome Dynastic Haplotypes: The High Kings of Ireland Revisited (Link)

Monday, April 30, 2012

My Cousin Otzi: A Story Written in DNA




   There has been a lot in the news lately about Cousin Otzi.   They talk about the fact that he had brown eyes, was lactose intolerant, was suffering from Lyme disease and that he was murdered.  What they don’t talk about was that he liked long walks along the glacier, a nice goat steak every once in a while and that he would give the pelt off his back for a friend.

   As soon as the world learned that they were going to test Otzi’s DNA the conjecture began.  Most folk assumed that Otzi would be part of haplogroup I (one of the earliest groups in Europe) or R1b (the largest genetic group in Western Europe).

   Europe is dominated by haplogroups I1, I2, R1a and R1b.  The rest of the landscape has a scattering of E, G, J and N.

   Otzi’s Y-DNA haplogroup was leaked late last year and confirmed two days ago as G2a2b (formerly G2a4).  My haplogroup is G2a3b.  This means that Otzi and I share a common G2a ancestor.

   G2a2b, G2a3b and G2a are subgroups of G.  Every time a new mutation within a haplogroup is identified a subgroup gets created or expanded.  Here is an example of a long R1b subgroup - R1b1a2a1a1b.

   While Otzi’s haplotype hasn’t been published yet, I did review a number of G2a2b records with the same L91+ mutation.  I ran an MRCA (most recent common ancestor) between my data and this group of Otzi-like folk and a conservative estimate makes our connection about 7,200 years ago.  I can picture our ancestor, and at least two of his sons, sitting around a fire somewhere along the Danube River.

   I look forward to getting to know Cousin Otzi better.