Hold a diamond under a light and watch it reflect and refract the rays. Rotate it, and you’ll notice new flares and flashes you wouldn’t have seen before. It’s all coming from within the same diamond. We’re just looking at it from different angles.
It turns out this same sort of approach could help us find new cancer solutions.
Enter Assistant Professors David Guinovart, PhD, and Eric Rahrmann, PhD, and Postdoctoral Associate Mohammad Qaraad, PhD, at The Hormel Institute, University of Minnesota. Together, this research group is building tech tools to key in on overlooked and potentially life-saving clues about how breast cancer operates and spreads so we can better detect, understand, and treat it.
Sleuthing for new cancer clues
Decades of meticulous breast cancer research have already dramatically changed outcomes and survival rates for the better. Even so, by the American Cancer Society’s count, more than 320,000 women will be diagnosed with breast cancer this year, and more than 42,000 are expected to lose their lives to the disease.
There are still more ways to change this. It starts with understanding that breast cancer isn’t actually one singular disease. It comes in different forms or subtypes. Each one’s growth is fueled by different factors, and some spread more aggressively than others. Understanding the nuances of each subtype is essential for cancer care teams to plan the right course of treatment for a patient’s specific breast cancer subtype.
For all we’ve learned about breast cancer, however, there is a vast amount we do not yet understand. We stand at the tip of an iceberg. We walk down a dimly lit hallway, peering through floorboards for clues that might have fallen through the cracks.
Countless cancer studies over the years have resulted in the accumulation of an almost unfathomable amount of data. The Cancer Genome Atlas (TCGA), for example, started in 2006, now contains a treasure trove of 2.5 petabytes (that’s 2.5 million gigabytes) of cancer research data.
Inside this dataset alone, how many new connections might be going missed?
“In the long run, the goal overall is to cure cancer. Unfortunately, cancer is a branch that extends into different worlds, and you end up with a huge tree,” said Dr. Guinovart.
Identifying new biological traits tied to specific breast cancer subtypes could unravel more mysteries from that expansive network of branches: Are there more subtypes or new biomarkers belonging to the breast cancer mosaic that we’re not yet aware of? What new processes could treatments target to stop this cancer in its tracks? Where might this subtype metastasize next?
Finding answers to questions like these first requires looking at the patterns playing out. With Dr. Guinovart’s background in applied mathematics and computer science, and Dr. Rahrmann’s expertise in cancer biology and metastasis, their research team built an algorithm to do just that.
Last year, Dr. Guinovart and Dr. Rahrmann received a two-year, $100,000 Data Science Initiative (DSI) Seed Grant from the University of Minnesota. Their funded project, MOOBI: Multi-Omics Optimization-Based Integration for Enhanced Cancer Research Datasets, aimed to develop a biological analysis method that makes new applications to existing cancer data, like The Cancer Genome Atlas and other publicly available datasets.
The project continues progressing, and the group recently authored a paper appearing in the scientific journal Cancers, which shares more information about the algorithm and its classification approach. The paper was also selected as an Editor’s Choice article by journal editors.
The algorithm discussed in the paper, which they’ve coined “Adaptive Hill Climbing Artificial Lemming Algorithm”, or “AHALA”, moves through datasets like a bloodhound tracking a scent along a woodland trail, sniffing out patterns and clues that could reveal new insights on breast cancer biomarkers.
“There are different types of approaches that you can run for the same dataset to extract meaningful results,” Dr. Guinovart said. “It’s also a good practice to re-analyze existing data with new ideas, new technology, new approaches to see if we can learn something new.”
The team developed scoring systems to shortlist and rank what biological markers to look at to determine different breast cancer subclasses.
Already, the process of looking has come with plenty of promising surprises — with some findings that might have applications beyond breast cancer.
“Not only have we identified what we expected to identify, based on what hundreds of labs have seen, but we have also identified additional targets that have fallen through the cracks,” Dr. Rahrmann said. “Some genes came up that don’t have a known function at all.”
Among their other new leads are potentially new therapeutic biomarkers associated with certain subtypes of breast cancer, as well as possible cancer weaknesses — the kind of information needed for more precise diagnosis and treatment.
“From the biology side of this project, I was surprised by what had been identified that hadn’t been identified before,” Dr. Rahrmann said. “It filtered at a fairly high level of influence. Not just in breast cancer, but in mammary gland development and other aspects of developmental biology that are not yet well-understood. For whatever reason, using these approaches has now filtered to the top of them having importance in determining what the type of breast cancer was. All of them are together here like a cassette that hasn't been appreciated before.”
Verify, verify, verify
If you want promising results from machine learning, just like teaching a bloodhound to track a scent, intentional training on high-quality data from the onset is a must. After all, an algorithm is only as good as the information you give it, the scents it learns to pick up on.
“This new technology is only as smart as we are as a human group. Everything they have been trained on is what they’ve been given by us. So putting in reliable, high-quality information will give us better results in the end,” Dr. Guinovart said. “If you change what happens at an early point, it has ramifications all the way down the line. Taking better care up front can have huge impacts on discovering things that were underappreciated before.”
“We have to think about this up front instead of running to the end,” Dr. Rahrmann added. “In terms of what we’re putting into these systems to train them, we need to be asking ourselves: What are we putting in? What are we giving it to start learning on? We have to do due diligence to make sure we have the best practice possible when we’re putting something in to get the best quality possible.”
That’s also one of the benefits of using public datasets for a project like this, as opposed to generating new scientific experiments alongside their new mathematical model. Not only is it less costly, but using data that’s already been verified after careful scientific review leaves less variables open to chance, and it allows closer scrutiny to evaluate the possible strengths and weaknesses of the algorithm itself.
“If you put new analysis on new data, you’re less likely to convince someone of something else, versus tried-and-true data that everyone’s looked at front- and backwards, and then applying something new. Then, it’s way easier to turn heads and get people to start thinking differently about things,” Dr. Rahrmann explained.
This latest phase of the project is still one of its early steps, one of many along an ongoing journey. Now, it’s entering a stage where Dr. Rahrmann has been able to test and verify some of the identified biomarkers with in-lab experiments.
“Seeing outcomes from our follow-up studies that are validating this is really cool,” Dr. Rahrmann said.
Beyond breast cancer
The mathematical model researchers have developed from this project has the potential to be used for broader applications, like improving testing for other types of cancers or diseases.
“For a computer, the application you’re giving it doesn’t really matter. The interpretation is your own. So you can use this methodology beyond breast cancer subclassification,” said Dr. Guinovart.
It also shows promise for further investigating metastasis, the primary cause of cancer deaths.
“When cancer is diagnosed, it’s often in later stages, when it’s already metastasized. Now that we’ve established some good, solid footing with this project, we can start applying this to the next stages of things like tackling metastasis questions. But first, we needed this. We had to be sure this approach was the right first step to take when looking at the data. With breast cancer-brain metastasis, for example, when breast cancer spreads to the brain, we’ve already started some preliminary analysis with some very interesting leadings so far,” Dr. Rahrmann said.
With enhancing understanding of metastasis as a key next step, the group expects to expand its collaborations. One such source could be MnMet, or the Minnesota Metastasis Research Network, a cancer metastasis research cohort established by Dr. Rahrmann that includes researchers at The Hormel Institute, the University of Minnesota, and Mayo Clinic.
“In the end, we’re still putting together more experts, because we need the view from different angles, different experiences,” said Dr. Guinovart. “With MnMet and MOOBI and all these projects that are ongoing, the goal is the same: adding more people to help solve the problem of cancer.”