Your DNA, Genomic Research, and Privacy – PediaCast 611

Show Notes

Description

Dr Mykyta Artomov visits the studio as we consider DNA, genomic research, and privacy. Our genetic information helps researchers uncover diseases and develop new treatments. But DNA is also deeply personal. How can researchers learn from genetic data while respecting and protecting patient privacy? Tune in to find out!

Topics

Genomic Medicine
Medical Research
Patient Privacy

Guest

Dr Mykyta Artomov
Principal Investigator
Institute for Genomic Medicine
Nationwide Children’s Hospital

Links

Steve and Cindy Rasmussen Intitute for Genomic Research at Nationwide Children’s

Artomov Lab – Institute for Genomic Medicine at Nationwide Children’s Hospital

Public Platform with 39,472 Exome Control Samples Enables Association Studies without Genotype Sharing (Technical Report) (Nature Genetics)

SCORE: SVD-Based Control Repository

PGS Browser: a Public Platform for Personalized Polygenic Score Analysis and Interpretation (Article) (Nature Communications)

PGS Browser

Genomics 101: An Introduction to Next-Generation Sequencing – PediaCast CME 075

 

Episode Transcript

[Dr Mike Patrick]
This episode of PediaCast is brought to you by the Institute for Genomic Medicine at Nationwide Children’s Hospital.

[MUSIC]

[Dr Mike Patrick]
Hello, everyone, and welcome to another episode of PediaCast. We are a pediatric podcast for moms and dads.

This is Dr. Mike coming to you from the campus of Nationwide Children’s Hospital. We are in Columbus, Ohio.

We’re calling this one Your DNA, Genomic Research and Privacy. I want to welcome all of you to the program. We are so happy to have you with us.

You know, our genetic information can help researchers uncover the causes of disease, develop new treatments, and understand an individual’s risk for having or passing on a genetic disorder. But these discoveries often require enormous amounts of data and DNA is deeply personal. So how can researchers gather the information they need while protecting the people who are sharing their DNA?

Today, we will explore the promise of genetic research, the privacy challenges it creates, and new ways scientists can learn from data without unnecessarily moving it around and exposing it to lots of different people. Of course, in our usual PediaCast fashion, we have a terrific guest joining us in the studio to discuss the topic. Dr. Mykyta Artomov is a principal investigator with the Institute for Genomic Medicine here at Nationwide Children’s Hospital. Before we get to him, I do want to remind you the information presented in every episode of our podcast is for general educational purposes only. We do not diagnose medical conditions or formulate treatment plans for specific individuals. If you’re concerned about your child’s health, be sure to call your health care provider.

Also, your use of this audio program is subject to the PediaCast terms of use agreement, which you can find at PediaCast.org. So, let’s take a quick break. We’ll get Dr. Mykyta Artomov settled into the studio, and then we will be back to talk about your DNA, genomic research, and privacy. It’s coming up right after this.

[MUSIC]

[Dr Mike Patrick]
He studies the code, measures the risk, and builds the tools that turn genetic information into meaningful discovery. Our guest today is Dr. Mykyta Artomov. He is a principal investigator with the Institute for Genomic Medicine at Nationwide Children’s Hospital and an assistant professor of pediatrics at The Ohio State University College of Medicine. Dr. Artomov’s research combines DNA sequencing, clinical information, and statistical modeling to investigate disease susceptibility and to develop safer ways of analyzing and sharing genetic data. That is our topic today, your DNA, genomic research, and privacy.

Before we dive in, let’s offer a warm PediaCast welcome to our guest, Dr. Mykyta Artomov. Thank you so much for stopping by the studio today.

[Dr Mykyta Artomov]
Thank you for the invitation, and I’m really excited to have this discussion today.

[Dr Mike Patrick]
Yeah, I am excited about it as well because this is something that impacts more and more patients and families, and in the future, it’s going to impact even more. We’re just going to keep hearing more and more about this. Let’s just start with the basics, though.

What exactly is genetic data, and why is it so valuable to medical researchers?

[Dr Mykyta Artomov]
That’s a great question. Genetic data refers to the information that we get from the molecules that are called DNA. This is something which encodes the information that is being passed on through the generations.

If you think about the traits such as height, we inherit the height from our parents. Intuitively, taller parents are likely to have tall kids as well. DNA is the molecule where this information is encoded via the sequence of letters, or so-called nucleotides.

Understanding which nucleotides, which specific letters in this DNA sequence are responsible for, in my example, presentation of the height, which I can measure in clinic or at home, is extremely valuable information because this is something which is staying intact throughout the lifetime. Mostly, the DNA that we’re born with does not change throughout the lifetime. Therefore, I can estimate and predict the heritable component of the trait of interest, whether that’s height or a disease, for example, starting from day zero of life.

And therefore, the value of knowing how to link the genetic data to the actual presentation of traits with the further advancement in technology could be used to help you mitigate those inherited risks which each of us possesses and help you gain and get more out of those beneficial traits, such as susceptibility to low cholesterol, for example, to get more of it and extend the healthy lifetime and longevity.

[Dr Mike Patrick]
Yeah, yeah. So, the DNA is a molecule that’s in each of our cells and it has genes on it. You talked about letters and those represent different amino acids or proteins that make up that DNA molecule.

But then that’s a code for things like your height and for what color your eyes are going to be and all of those things. So, once you figure out what part of the genetic code matches up to that trait, then you can start to figure out even without necessarily knowing if someone is in general short or they’re tall, you might be able to figure out it’s likely they’re going to be short or they’re going to be tall based on the code that you’re seeing within the DNA. Is that true?

[Dr Mykyta Artomov]
That is true. And the difficulty of this problem, why the simplification is intuitively probably transparent, but the difficulty comes from the fact that there are three billion letters in that code and that’s a lot of letters. And therefore, understanding what each of those does is still an unsolved problem.

However, we are on the path to understand, well, as you mentioned, what genes are doing. Right. Each of us possesses more than 20,000 genes.

Each of those doesn’t code the protein and protein carries specific function. Some of them release the cholesterol into the bloodstream. Some of them regulate our blood pressure and so on.

So, the more I know about the functionality of the units, whether those are letters or genes or specific domains in this DNA sequence and the code, the better I understand the biological origins of the traits such as height or diseases, which is even more interesting from the clinical standpoint. I understand more about the biology and origins of that disease. How this specifically translates, why this is so important for the practical usage is the following.

So, think about the cardiovascular disease example, right? We understand quite a lot about the etiology of hypertension and myocardial infarctions as a cholesterol that has been built up in the bloodstream, forming the plaque, plaque being destabilized, there is inflammation involved, and then mechanically plaque results in the arterial myocardial infarction and strokes. So therefore, if we understand the origin of that disease, we come up with stadiums, for example.

So, we lower the cholesterol in the bloodstream, therefore decreasing the likelihood of adverse outcomes. We understand much less about things like psychiatric disorders and behavioral health, for example. Therefore, there is, you know, for some of the traits there, there’s very limited available treatment and drug design that is largely challenged by the fact that it’s hard to understand the cell type context, it’s hard to understand what is malfunctioning, it’s hard to understand what is the biological origin of the disease.

So, understanding this origin using the DNA and genetic data is something which makes the discovery of the respective drugs easier and be more reliable.

[Dr Mike Patrick]
And once you discover that a particular drug or medication is best for a particular gene sequence that corresponds to a disease, so let’s say we have a disease and it can show up in lots of sort of subtle different ways depending on what makeup the genes have, because they’re encoding for a little bit of a different protein, but that protein might only be mildly abnormal and might only create then mild symptoms versus it’s a protein that’s not working at all, and so you’re going to see more severe symptoms.

And what medicine you use to treat might be different depending on the genetics of that particular person. And so, then once you figure out what genes correspond to a disease, then you can say, well, if you have this set of genes with this disease, then this medicine might be a little bit better for you. So, it can kind of individualize treatment based on your genetic code.

Is that, again, this is very simplifying but trying to help non-scientists understand.

[Dr Mykyta Artomov]
That actually is also, this is actually more advanced topic than you can think right away, but we’re kind of diving here into something which is called pharmacogenomics, where there is a personalized response to the specific drug types, and some people will respond to one type of drug but not the other. So therefore, yes, the response to a specific type of therapy could be considered a trait, a genetic trait, just like all of the previous examples that we discussed, like height or any other disease. Similarly, the profiling of individuals based on their gene expression, for example, or cell type abundance in their blood also could be a trait that could be studied genetically linked to the specific outcomes of interest from the clinical standpoint.

So therefore, yes, this is true from both standpoint of understanding the origins of the disease, what kind of drug would I like to develop, and what are the specific subgroups of patients that would benefit the most from this specific therapy.

[Dr Mike Patrick]
And in doing this, you obviously then need lots of background information. So, you really need to look at the DNA of lots and lots of people and then see what diseases and how those diseases, you know, whether they’re mild or severe of a particular disease, what it looks like, and then what the DNA looks like in order to start making those kind of determinations. Some of the studies that you guys do are case control and family-based studies.

Can you talk a little bit about those and how those help us identify which genes or changes in those genes are associated with particular diseases and maybe severity of those diseases? How do those case control and family-based studies help you out?

[Dr Mykyta Artomov]
So here, we can continue the, you know, the example of cardiovascular traits that we have already started. And I’ll give you some more specific examples. So, around 2003, a rare family has been found where the abnormal cholesterol, very low cholesterol levels were observed.

Great, that’s apparently no other adverse effects. So hey, great, this is significantly naturally decreasing the risks of the cardiovascular adverse incidents. So, upon investigating the genetic makeup of the DNA, so basically what is so special about the inheritance pattern of the trait within the specific family, and then studying the family helps you to understand this is the same mutation that is responsible for the trait which is shared by everyone within the family.

This is the most likely scenario. All of the individuals are related, so therefore there is just, you know, the same mutation going around. So, what happened is a gene that is called TCSK9 had the mutation in this family, which is completely inactivating, as in simply this gene is not being expressed, probably not produced, no adverse effect, very low cholesterol level.

The problem is generalization of this finding. I found one family where this works. Great, can I make in or infer the, you know, something about how it will work in the population?

On an example of one, probably no. However, this is an extremely valuable lead where in one specific case I find a highly significant mutation which takes care of a very significant health risk. Great, so a few years later, consider another family that has been found where this gene was upregulated due to the mutation.

So, a higher expression of this gene resulting in abnormally high levels of cholesterol. Great, so now on the two families I have found two mutations, two different mutations in the same gene which are basically regulating the cholesterol level within the bloodstream, therefore either significantly decreasing or significantly increasing the risks of the cardiovascular trait. Biologically, this is a very good indication that this gene is very important for this trait.

Still doesn’t answer the question about how does this generalize if I were to talk to many patients in my practice, for example. This is where we pivot from family-based studies where we can search for this highly impactful mutations that are shared by all the members of the family and trace how this goes together with the trait of interest into the cohort-based studies, where we will take a large number of individuals who are in possession of the trait and large number of individuals who are lacking this trait, and those will be the controls for our study, and we’ll compare the genetic makeup between the two groups. So, if I study in that way the cardiovascular incidence of myocardial infarction and stroke, one of the things that comes up is that PCSK9, the gene that I mentioned, is actually frequently mutated among the carriers of the disease. And this is one of the first so-called genome-wide association studies findings in the late 2000s with respect to the genetic origins in the population scale studies for cardiovascular traits.

And this finding is generalizable because this looks at the more common mutations in this gene, telling me that, hey, the mutations in this gene are slightly increasing or decreasing the levels of cholesterol, so I can look at the patterns of those mutations and therefore interpret them in the context of what I’m expecting from the measurable trait. So therefore, as the, you know, solid and practical outcome for people where statins are not as effective as the cardiologists would like to, we use PCSK9 inhibitors, the drugs that are solely based on these findings that I have described, and these are the second- line therapy for people where statins are not effective, therefore significantly mitigating the risk of cardiovascular adverse events.

[Dr Mike Patrick]
Yeah, yeah. It’s really like being a detective and using a lot of data from different sources of both, you know, the lab tests of an individual and what their genetics look like. And once you’re looking at families and see where changes are all in the genetic code are leading to the same disease, and since those are passed on, we see that more in families.

And then once you know what protein that particular thing codes for, then you can start thinking, well, maybe if I shut down that protein or we provide more of that protein or, you know, again, you just start thinking through options and ways that you could overcome what the problem is and then start, that’s how new classes of medicines get made. Is this something I would imagine that there are limitless possibilities of diseases and things to study that will go, you know, long beyond one individual researcher’s lifetime, right? I mean, there’s so much potential here for advancing human medicine based on our genes.

[Dr Mykyta Artomov]
You’re absolutely correct. And, you know, the story that I mentioned was unraveling in early 2000s. And as of right now, such studies have been conducted for thousands of traits across multiple populations and across different, you know, age groups and so on.

So, the data as the sequencing and data generation technology becomes cheaper and more abundant, it is actually becoming feasible to generate more and more data. Therefore, whenever, you know, why do we need so much data? Well, the mutations are not all equal.

Some of them are impacting the trait significantly with the large effect sizes. Some of them are only slightly changing what we’re trying to measure or slightly changing the disease risk, for example. So therefore, the smaller the effect you’re trying to find, the more data you need to see those subtle differences between the healthy and unhealthy individuals.

So therefore, the potential for this is enormous. And this is why there’s a lot of, you know, general investment both from industry and the academia into both generation of the data and the analysis of this data downstream. Because genetics provides this sort of genetic evidence, so-called.

So, something that supports the hypothesis about how the biology of the disease is playing around. And since this evidence is not changing from the, you know, because of the environmental factors, typically our inherited DNA stays the same. That means that once we find that evidence, we’re actually in a better position to produce the medicine or stratify individuals by patterns of risk.

[Dr Mike Patrick]
Yeah. When we’re talking about an individual’s DNA and you’re looking at the DNA to see if there’s particular mutations, often you get other data in addition to the data that you were looking for. So, you usually don’t isolate just one gene.

You know, you may have, you sequence someone’s DNA and you find out there’s lots of data there. And you may be only for that particular patient, only interested in data at a particular location on a particular gene. But we still have all this data that perhaps other researchers could use as they’re trying to find, you know, different mutations that then, you know, result in different diseases.

And so, I guess the more genetic data that we have, the easier it becomes, especially in the case of rare diseases, where you would need, you know, thousands of people in order to find a particular example of the mutation. So, you really, I guess what I’m saying is you just need lots and lots of genetic data out there. But then that becomes sort of a privacy issue because, you know, we all know, you know, as we think about our DNA and, you know, it’s out there on, you know, like ancestry, for example, it seems like there’s, you know, people get a little cringey when you think about other people having access to your DNA data.

And yet that access is so important in order to discover, you know, the causes of diseases and new treatments for diseases and something that’s really going to, you know, benefit humanity. So how do you sort of walk that edge of wanting to respect privacy, but also wanting to advance science? How do you approach that dilemma?

[Dr Mykyta Artomov]
This is a very important, if not the critical question of the genetic research nowadays. And it is also a very sensitive question, because as you mentioned, the privacy and the, you know, ownership of the data by the patient or by the person contributing this data is a cornerstone of anything we can do downstream. And this is the way how it should be handled.

So, therefore, there are multiple protocols for how this data could be utilized and shared. There are multiple boards that are overseeing the research activities within the hospital and on a more, you know, administrative side from the governmental standpoint. But generally, the genetic data is considered to be private.

So, this is a confidential data that is not intended for open share. If we’re talking about a single patient, that is, you know, pretty much always true. However, the National Institutes of Health and supportive organizations are highly encouraging sharing the cumulative data, such as, you know, if I have more than 10 people in my cohort, it’s actually highly encouraged to share the frequency with which I do see the mutations.

I won’t know which person has that mutation, but that will give me the information about how frequent that mutation is in population. Is it one of a kind, and therefore likely doing something important? Or this is something which is present across, you know, multiple individuals, and therefore reflecting just natural variation.

And, you know, as I said, eye color or height, there is no specific, you know, bad or good outcome for it. It’s just natural variation of human beings. So, therefore, there’s a very sensitive balance that needs to be obeyed.

From one standpoint, sharing of the data openly highly benefits the pace of the research at the cost of, you know, exposing private data and so on. So, therefore, there are two possible pathways that are being utilized. A, there are secure pipelines and secure pathways which are highly legally regulated of what can share, you know, what can be shared with whom, when, under which terms, and how exactly this data is going to be protected.

Given the amount of data that is non-genetic that we’re sharing with our banks, social media platforms, and so on, having this coverage is, I’d say, is, you know, taking care of 99% of all possible concerns. If you’re not concerned about, you know, sharing your data with your bank account and, you know, your clinical data with the primary care office and so on, then you probably shouldn’t be more concerned about sharing the genetic data with your research. However, there are other ways of doing that where the secure methods are being developed to take advantage of the existing data without actually exposing the private or individual level data to the research, neither externally nor even between the parties that are being shared the database.

So, you can think of the, you know, different types of approach in this problem. It’s also the cryptography where the data is encrypted and there is data sharing, but nobody can read it really, or the alternatives where the analysis can be performed in a way which does not require sharing of the data. And this is currently the most secure standard for doing the genetic analysis where data formally never leaves the owner’s computer.

Therefore, there is no concern because the data is not being transferred at all.

[Dr Mike Patrick]
How then does that information get shared if it’s just staying in a particular lab? Right.

[Dr Mykyta Artomov]
So, this is the, there are different types of analysis that could be performed. Obviously, there are certain limitations because, you know, you can do only as much without sharing the individual level data. But typically, one of the major directions here is the shared control platforms.

So, we started our discussion from describing the family-based and the cohort-based studies. So, imagine that, you know, a lot of the studies are currently, you know, investigating the cohorts of the specific disease, and we also do so at the Institute for Genomic Medicine. You assemble the clinical, very highly profiled and very well curated data on a specific disease.

You’re interested in what is so different between the genetic makeup of the patients that you have recruited and the general population which does not show, you know, the evidence of this disease. Naturally, in clinical trials, you would go collect the control group. Means more sequencing, means more recruitment, means more money and resources spent on this.

In the other center, studying different disorder, people will enroll their patients and also enroll a healthy group, spend the money on sequencing, spend the money on re-enrollment and so on. It would be so nice if we can use shared controls between the two studies, right? We then can do the sequencing only once and then share the information between the two centers.

So, it’s either the direct sharing of the data which requires all of the legal coverage, time consuming and effort consuming, because we will need to process the same data set on our side computationally twice. So, computational effort also gets factored in here. Or the alternative is building the shared control repository where the data that is permitted for this type of activity is being stored in the aggregated form, and you can send an inquiry.

It’s like, I need controls with this specific properties or and so on. And then get the aggregated information back from the selected controls. So, to give you an example, in case control studies, it is extremely important to have a well-balanced case and control groups.

The reason is because we’re looking basically for a needle in a haystack when we’re searching for specific mutations responsible for the disease. There are many traits that could be immediately different between the cases and controls. Height, age may be different, the origins could be different, right?

So, there’s a huge genetic diversity of people in the world. So, we ideally would like for cases and controls to have a relatively similar genetic background. Otherwise, that’s the first thing we’re going to discover, that there’s a systematic bias between the genetic origins and genetic ancestries.

So, therefore, the methods that are a cornerstone of the shared genetic control platforms is the ability to select a matched group of cases and matched group of controls rather than just get the aggregated frequency of the mutations. This is where a lot of the computational work is being done, and this is where all of the coding, the algebra, and all of the rocket science of genetics starts to appear.

[Dr Mike Patrick]
PW. Yeah, absolutely. So, let me make sure that I have this straight in my brain.

So, if you’re doing a study, and you have your study group, and you have your control group, and you’re looking at the difference in the genomes of to see what’s different on the genetic level of people who have the disease and people who don’t have the disease. And so, if an individual researcher wants to select a control group, there’s a cost involved in terms of sequencing their DNA and finding out, you know, what all their genes look like. But if everyone in, you know, that’s in various labs around the country and around the world, if all of them, if all that DNA could go into a repository, then you could just use your control groups selected from that DNA that’s already there without having to sequence new people.

But then we worry about patient privacy, but does that repository, does that store like the individual person’s like name and their age and, you know, personal identifiers? Or is that DNA that’s available for others to use de-identified?

[Dr Mykyta Artomov]
So, typically for this type of studies, all of the DNA is de-identified, meaning that you can only identify the presence of a person in the database if you already have their DNA in hand, simply by direct comparison, right? So, but that requires you to already have this data in front of you. So, therefore, for the case groups that are being studied, typically the clinical geneticists and people who are assembling this cohort will have the information about the, you know, patient identifiers and identifiable data and so on.

And there is a very specific and very strict protocol on how this data should be handled, stored, destroyed, and how this should be, you know, basically the patient is in charge of how this data could be used. And this is in full charge of how to restrict the usage of this data, even it has been committed to some of the studies already. So, therefore, in case of the local data that has been, you know, assembled on the case cohort, the identifiable data is being treated as a part of the medical records.

So, therefore, it’s highly protected as the personal health information and is not disclosed at any time. All of the other data types are de-identified and only could be used to identify the person if you already have their complete data in front of you. So, therefore, this is considered to be secure from the standpoint of, you know, inability to identify the person behind the DNA sequence.

[Dr Mike Patrick]
Yeah. So, in the original lab, of course, you’re going to have patient identifiers because you took their blood or whatever sample you’re getting the DNA from. But that is handled as private just as it would be in any research project and in any, you know, with the medical record in any clinical department, you’re going to keep that private.

If it goes into one of these repositories of DNA information that then other centers could possibly use as controls, that information is de-identified. And the only way that you could identify someone is if you had their, you know, a large sequence that you could compare to see if it’s an exact match. So, then that leads when people are using the repository for control groups, that information, do they use someone’s entire sequence or just the sequences that they’re interested in, which would then make it even more difficult to identify a particular person because you’re only looking at one small area of their DNA.

[Dr Mykyta Artomov]
So, to be more specific, the concept of the shared control databases has been in the field for a long time, for a decade, if not more. The practical implementations of that are actually very, very few. So, to date, there is a direct way of getting the external genetic data from the National Institutes of Health repositories, and those are de-identified data, and that you can get as the single individual resolution data, so individual level sequences.

These were contributed there under specific consent, so people contributing their data to that were fully informed and made a conscious decision of doing so. One of the challenges of using that data, then you get the full sequence, yes, whatever was sequenced, and the challenge of doing so is you need to process all of this data from the very raw sequences of letters into the analyzable format, which is an expensive endeavor if you’re talking about a large chunk of data, and B, this requires significant expertise. Therefore, a simple clinical question, A, is this mutation more prevalent in controls, becomes, give me a month to process your data, and then I’ll tell you.

So, instead, one of the solutions that was offered by our lab is actually doing all this pre-processing work for you. So, we assembled all of the samples that were consented for broad research use into a single centralized control repository with more than several tens of thousands of people in it. We pre-processed it in a way that this doesn’t require any, you know, research computational expertise from the user standpoint, and then basically now you can ask the questions both about the prevalence of a specific individual mutation or the prevalence of the, you know, mutations in a specific large or entire chunk of the DNA.

So, yes, the answer is you can ask this question in multiple ways, and how easy and how fast you will be able to get the answer will depend on the database that you’re using. Either you’re going for the individual level data and spending all the time processing it or using this kind of unique solutions for the centralized control repository that we have built and getting this response in basically a matter of minutes.

[Dr Mike Patrick]
Yeah. So, I would imagine if you if you have a particular question about a particular area of DNA, then that repository’s computer is going to be able to give you a summary report of what it finds and not necessarily someone’s exact sequence.

[Dr Mykyta Artomov]
That is intentionally built in such a way that the individual level information cannot be extracted out of the control repository. This is done through several kind of lines of protection, starting from the fact that the control data sets of one individual, two individuals, up to 100 individuals are simply not returned. So, if you’re getting any data, this is from at least 100 individuals.

And B, we’re, you know, I’m not going to get into the details of that, but we’re employing specific solutions that are ensuring that there are no two control data sets that we’re releasing that are different just by one sample. So, therefore, you’re, you know, you’re always getting much controls, the summary information, and at any point there is no way to identify a specific mutation belonging to a specific individual within that control repository. Yeah.

[Dr Mike Patrick]
And I think our main goal here as we as we discuss this is that there should be hopefully some comfort in the part of patients and families in terms of sharing their DNA. That it’s not like a researcher at a different facility is going to have access without your knowledge to your sequence exactly and know who it’s coming from. That you’ve really put a lot of thought into de-identifying, only letting researchers have the information that’s pertinent to their work.

And in a sort of a grouped way, again, makes it even more difficult to identify a specific person. So, you know, when folks ask for your DNA for medical research, hopefully they’re explaining all of this to you, but it’s complicated and we’re just trying to spread awareness that there is privacy built in that is trustworthy. Correct?

[Dr Mykyta Artomov]
The clinical genetic research is always a partnership between the patients and families and the providers, clinical geneticists, the genetic counselors, and a large number of team members on the research and hospital facilities that are making all of this genetic research and genetic clinical investigations possible. Therefore, there’s always, just like in any other field dealing with the sensitive data, such as economy or, you know, bank or anything, there’s always a matter of trust between the parties. The data security is one of the cornerstones of clinical research, just like this applies equally to genetic data and clinical data.

Therefore, the building of the cohorts that are being studied is a highly sensitive and highly challenging task for clinicians and for the genetic counselors because, A, this involves a very explicit step of explaining all of the details of how this data is collected, how it’s handled, with whom it will be shared, how it will be used, and empowering and giving all of the instruments to the participant to be in charge and being able to restrict all types of data usage that they’re not comfortable with or even be able to withdraw their consent for, you know, participating in a study. Therefore, this kind of gives all the instruments and the steering wheel, if you want, of this vehicle to the patient. And the decision-making process about the data sharing and so on is extremely limited by the documents that are, you know, describing this to the patient.

Therefore, I feel like, as a field, we’re doing a great work in spreading the awareness and also explaining how we’re using this and how exactly the benefits of using this data are going to impact our society, even though you can see a lot of examples where the studies that are designed, especially in are explicitly designed in such a way that they may not be able to help the patients contributing their data to the study, simply because the solution of the problem may come much later. But you will still see a lot of enthusiasm and a lot of appreciations for the studies and a lot of willingness to participate, because as an outcome of the study, the overall healthcare service and overall health of our society is getting a significant, positive impact. And this is what patients are willing to contribute to.

[Dr Mike Patrick]
Yeah, yeah, absolutely. I also want to point out that a lot of genetic conditions and traits and the way that diseases express themselves, depending on an individual’s DNA, often it’s not just one gene, right? It could even be one area of your code impacts area one.

So, then you have to understand, well, this gene is going to be expressed in a certain way, depending on what this other gene is doing. And so, then it becomes even more complicated in terms of what genes that you’re looking at. And so, when you do hear direct-to-consumer marketing of genetic tests, oftentimes it’s just answering a question of, is there one particular change that’s involved?

But it might not be taking into account all the other changes and all the other data points and the family history and what are your exact symptoms and all the things that the genetic researchers and the clinicians are thinking about, all those different data points, as opposed to these tests that you can send your blood in and get a result back that’s just marketed to consumers. Those may not be as trustworthy because there’s all this data to consider.

[Dr Mykyta Artomov]
Is that correct? That is correct. And it’s never easy to interpret a genetic mutation, whether it is a single one or there’s a hundred in the same individual that you’re trying to interpret.

These are all challenging tasks. However, if you know that the trait is caused by a single gene, it significantly improves the interpretability of the testing results and therefore clinical decision-making. So, screenings for cystic fibrosis, for example, is one of the routine screenings for prenatal tests and so on, simply because we know the gene, we can interpret the mutations, and therefore we can predict the impact.

So how about things that are more challenging, like, you know, things where the cardiovascular system or autoimmune diseases are coming in? As you mentioned, a lot of them are caused not by a single highly impactful mutation, but by a combination or a pattern, if you will, of the mutations that are playing out so unbeneficially in some individuals, or the other way, playing out so beneficially, protecting other individuals from these diseases. So, this is where a single mutation is not enough.

We need to look at not even a dozen, but rather sometimes thousands of mutations to understand that pattern, as, you know, we started with highly so-called polygenic phenotypes, like cardiovascular traits, there’s cholesterol regulation, inflammation, there is the blood pressure regulation that is involved in that. All of those are not regulated by the same genes. Therefore, there is a high variability in terms of the mechanism and genes involved.

Therefore, the concept that is assessing the cumulative impact of this pattern of mutation is something which is currently being kind of slowly rolled into the clinical genetic practice, which is called so-called polygenic rescores. So, this is something which accumulates the impact of many mutations and is used to assess the individual’s susceptibility to this polygenic, highly distributed traits across the DNA. An interpretation of that, as you may imagine, becomes even more data-driven, data-requiring, and challenging.

Therefore, this is where we will see the improvements in the next three, four years, and this is where the area of active research, especially if you add another level of complexity that you mentioned, the clinical pattern that is visible to a physician. Is it important for me to do the genetic test given the clinical state in which I am as a patient? Or is this condition going to benefit from the genetic testing more than the other?

These are the questions that are going to help, answers to which are going to help drive the genetic testing and the clinical genetic interpretation forward.

[Dr Mike Patrick]
Yeah, yeah. It’s already such amazing work that has really impacted so many people’s lives. And as I think about even common diseases, as you’ve been mentioning, cholesterol and blood pressure and things that have been, you know, sometimes difficult to treat in the past, we’re just getting so many more treatments available so that hopefully one of them will work for a given person.

And a lot of that research is being done because of what we’re talking about today in terms of genomic research and figuring out what’s happening at the code level to figure out how to treat a particular disease. So, it’s really, really important work. And I will say, too, that where we have come so far is in a large respect due to all the folks who have contributed their DNA in the past.

And so, as we think about folks who have, you know, trusted genomic research to take care of their data safely, allow other researchers at other places to use that data, has really helped us already to where we are right now. And there’s more to come. And so, I think we still need lots of folks who are willing to provide that data for science and for humanity.

But it’s also our responsibility to make sure that we are taking care of that in a way that respects the individual and, you know, maintains privacy for them with, you know, with their DNA and with what’s being shared in repositories and such. So, it’s complicated, but it is still an important thing for all folks to at least be aware of and to know that the folks doing the research are thinking about their privacy, right?

[Dr Mykyta Artomov]
The privacy is always the cornerstone of the research. And to say, to give you a specific example of how this partnership between the participant and the researcher is playing out, is there is a data set which is called 1000 Genomes. This is the individual level genetic data across entire genome that is made fully public.

Anyone can go and download it. Currently, there are more than 2000 genomes in there. But this data set has been so massively used and has advanced the understanding of both genetic origins of the populations, the specific biological properties of our DNA.

It answered so many questions just by the fact that this is the data that literally anyone, anyone can go and work on. Obviously, this cannot be scaled to, entire healthcare systems or even a single hospital where there’s more than just genetic data being analyzed. However, this shows the power of sharing the data and establishing the robust and controllable boundaries to that that are ensuring the safety in patients’ eyes is something which is a critical task for the researchers and the field of genetics as a whole.

[Dr Mike Patrick]
Yeah. Well, we really appreciate you stopping by and chatting with us about genomic research and patient privacy as it relates to that. We are going to have lots of links in the show notes.

So, if this has sparked curiosity in your mind as a listener, please head over to PediaCastcast.org and look for the show notes for this particular episode. We will have links to the Steve and Cindy Rasmussen Institute for Genomic Research at Nationwide Children’s Hospital and the Artomov Lab, which you run, Dr. Mykyta. And we’ll have a link to that as well.

And then also some of these public repositories that we’ve been talking about. If folks want to go and look and see what it looks like, we’ll have links to some of those as well. And then also, we have talked about genomic research and sort of in broader scope of how it fits into the healthcare system in previous episodes of PediaCast.

And I’ll put a couple of those past episodes in for you as well, if you’d like to explore more on this really interesting topic. And I mean, this is the area that is just changing human healthcare and medicine by leaps and bounds. And there’s much more to come in all of this.

And so, it’s really exciting. And we’re so thankful for you stopping by and sharing your expertise with us. So, once again, Dr. Mykyta Artomov, Principal Investigator with the Institute for Genomic Medicine at Nationwide Children’s Hospital. Thanks so much for chatting with us today.

[Dr Mykyta Artomov]
Thank you so much for the invitation.

[MUSIC]

[Dr Mike Patrick]
We are back with just enough time to say thanks once again to all of you for taking time out of your day and making PediaCast a part of it. We really do appreciate your support. Also, thanks again to our guest this week, Dr. Mykyta Artomov, Principal Investigator with the Institute for Genomic Medicine at Nationwide Children’s Hospital. Don’t forget, you can find us wherever podcasts are found. We are in the Apple Podcast app, Spotify, iHeartRadio, Amazon Music, Audible, YouTube, and most other podcast apps for iOS and Android. Our landing site is PediaCast.org.

You will find our entire archive of past programs there, along with show notes for each of the episodes, our terms of use agreement, and a handy contact page if you would like to suggest a future topic for the program. Reviews are also helpful wherever you get your podcasts. We always appreciate when you share your thoughts about the show.

And we love connecting with you on social media. We are on Facebook, Instagram, Threads, LinkedIn X, and Blue Sky. Simply search for PediaCast.

Don’t forget about our sibling podcast. And if you made it all the way through this episode and you found this fascinating, do check out PediaCast CME. It is similar to this program.

We turn the science up even a few more notches, and we offer free continuing medical education credit for those who listen. So, if you are a physician, a nurse practitioner, a physician assistant, a nurse, a pharmacist, psychologist, social worker, even dentists, we are jointly accredited by all of those professional organizations on PediaCast CME. So, it’s likely we offer the credits you need to fulfill your state’s continuing medical education requirements.

Shows and details are available at the landing site for that program, PediaCastcme.org. You can also listen wherever podcasts are found. Simply search for PediaCast CME.

Thanks again for stopping by. And until next time, this is Dr. Mike saying, stay safe, stay healthy, and stay involved with your kids. So long, everybody.

[MUSIC]

Leave a Reply

Your email address will not be published. Required fields are marked *