Quantum Chemistry of Molecular Evolution

Ph.D. candidate Lázaro A. M. Castanedo and his adviser Chérif F. Matta are trying to answer one of the fundamental questions of molecular evolution: Why did nature pick the particular DNA structure humans have now as a carrier of our genetic information instead of other equally plausible choices? 

“In the absence of constraints, nature’s chemical reactions always choose the path that is the least costly in terms of Gibbs (free) energy,” explains Matta, who is a professor at the Department of Chemistry and Physics at Mount Saint Vincent University. Gibbs (free) energy balances the tendency of nature to increase randomness as time progresses and its tendency to seek the lowest possible energy. “Could nature have made less costly choices than those that led to today’s DNA? If so, why didn’t it?” 

Describing the project, Castanedo, who is pursuing his PhD at Saint Mary’s University in collaboration with Mount Saint Vincent, says he compares the energies of different combinations of building blocks of nucleic acids (that is to say DNA and RNA). In addition, chemistry textbooks provide ample descriptions of the structural and chemical differences between DNA and RNA, but Castanedo and Matta want to know why that is, and if the observed forms are somehow energetically advantageous. 

As an extension of this work, they are also investigating whether some of the nucleotides nature hasn’t chosen have applications in drug development. 

“One of the applications of discovering new nucleotides is that they can be used to produce similar molecular structures for drug discovery,” says Castanedo, who received the 2020 Abe Leventhal Research Bursary from the Alzheimer Society of Nova Scotia for his work in developing molecules to prevent and detect dementias. “They could be used to develop new treatments for Alzheimer’s, cancer, and other diseases.” 

Castanedo says he could do this research in a lab by synthesizing and comparing the components, which would take decades worth of human labour, or, he could approach it theoretically using computers, taking a fraction of the time. 

“We predict what happens in the ‘real world’ using the computational infrastructure of the Digital Research Alliance of Canada and ACENET, which has thousands of cores and a panoply of state-of-the-art computational quantum chemistry software,” he says. “So, you can actually create a model of this molecule on the computer and obtain, among many other properties, its Gibbs energy.”

Castanedo is one of Saint Mary’s heavier users — with 235,000 CPU hours in 2022 alone (about 27 years of CPU time) — because he’s been analyzing a total of 2,530 molecules for this project. 

“For each of these molecules, a CPU needs to run continuously anywhere from, say, 20 to 200 CPU hours. These calculations are computationally intense, as we use high levels of quantum chemical theory, such as density functional theory (DFT), to obtain accurate predictions,” he says.

“They generate a lot of data,” Matta says, “so you need to know what you are looking for — a needle in the haystack so to speak.” 

While preliminary data has already been made publicly available, much of this work is yet to be published. 

Cracking Cannabis Codes

David Joly studies the interaction between plants and micro-organisms.

“We are looking at what genes make plants more resistant to disease and what genes make them more susceptible,” explains Joly, a biology professor at Université de Moncton. “On the pathogen side, we’re trying to determine what genes make a pathogen aggressive with a particular plant and what makes the pathogen detectable by that plant. Plants have an immune system and are able to recognize certain molecules from pathogens, triggering a defence response, a little like we humans do.”

Joly works mostly on cannabis and says some plants are more resistant than others — again, being comparable to humans. Some humans, for example, seem to get the flu every winter, and some simply never seem to get sick.

“If we focus on plants that are more resistant and compare them to plants that are susceptible, can we see differences in the genes?” he says, using tomatoes as an example. “We could take those more resistant plants and use them in a breeding program, crossing them with plants we know produce really juicy tomatoes.”

Joly says he’s starting from the beginning in many ways with cannabis because Canadian researchers have only recently been allowed to study it.

“I have to stick to what’s been authorized to work with,” he says. “So we have to gather as many different plants as possible and test them in our growth cabinets, and then we can sequence their DNA and ultimately use ACENET resources to look at their differences.”

He says he and his team need to screen millions of “letters” in the cannabis genome, and on a small computer, that would take weeks.

“So that’s where we use resources that are available from ACENET,” he says. “What I like about ACENET is the training they offer to take students from zero knowledge of bioinformatics to slowly making them more comfortable with bioinformatics coding. You need to be able to program and code and ACENET teaches them that. I can help them, but we often take advantage of the training from ACENET.”

He says ACENET has been especially instrumental with his undergraduate students, who are almost always new to bioinformatics.

“Even at the graduate level, I have students who arrive here and don’t know much about bioinformatics, so ACENET’s training is still useful there,” he says. “Then we access the different tools and software so we can analyze the data. Every time we encounter problems, the ACENET people are always very useful in helping us find the problem. Sometimes it’s just a semicolon in the coding and they’re patient enough to help us find it.”

Joly says his work would be “very difficult” to do without the services of ACENET.

“You can wash the dishes manually, or you can use the dishwasher, but if you use the dishwasher, you can wash way more dishes in a given amount of time and get other things done while that’s happening,” he says. 

Record Use for UPEI!

Physicist James Polson wanted to tackle a computational project inspired by a collection of experiments done a few years ago by a research group at McGill University. He knew the project would be computationally demanding, but he had no idea that by the end of it, Matthew Kozma, his research student, would have set a Canadian record for compute-cycle use among undergraduate students. When you consider that most high-performance computing work is done by graduate students, postdocs and faculty, it’s noteworthy that an undergraduate — who only worked in the summer months — from a very small Atlantic Canadian university used an amount on par with Canada’s most compute-intensive researchers.

The use was immediately evident to ACENET. “Suddenly a single research group was using more compute resources than entire provinces,” says Greg Lukeman, ACENET CEO.

Polson’s project started with data from the Walter Reisner research group in the Physics Department at McGill University, which does experiments that examine the physical properties of DNA.

“It’s for the purpose of advancing a certain type of nanotechnology designed to manipulate and analyze DNA molecules,” explains Polson, a professor of physics at UPEI. “So they’ve devised this experiment to look at how DNA molecules behave when you squeeze them into very small and heterogeneous spaces.”

Polson had wanted to do a computer simulation project to complement their experiments. He knew what he wanted to do, but the computational methodology was new to him.

“It was on my to-do list for a number of years, but I was always reluctant to give it to an undergraduate student so I figured it would be something I’d have to do myself,” Polson says. That is, until Matt Kozma arrived at UPEI. “Getting the methodology implemented so that you can do these calculations is a really difficult thing. But Matt is very computer savvy, so I thought, ‘if there’s any undergraduate student who can tackle this project, it’s Matthew.’”

Kozma wrote the computer code from scratch, which “is not a trivial thing to do,” according to Polson. “And he was also able to implement the methodology.”

They set out the methodology and code over the course of the summer of 2021, and made sure the program they’d devised was well tested. Then, in the summer of 2022, Kozma really got going. Using all five of Canada’s national supercomputers, his simulations totalled over 30 million CPU hours, setting the all-time Canadian record for usage by an undergraduate student, and more than doubling the previous record. In fact, during his five-month run, he was the third highest user (undergraduate or otherwise) in Canada. Kozma worked elsewhere in the summer of 2023, but when he returned to his studies this year, he and Polson finished up the last bits of work on the study and their paper will be published later this year.

Kozma presented the research at the Atlantic Universities Physics and Astronomy Conference two summers in a row, as well as at the Canadian Undergraduate Physics Conference.

“It’s always great to have the opportunity to present the research, in addition to doing it,” Kozma says.

Ultimately, their research has applications “way downstream” in advancing nanotechnology for manipulating and analyzing DNA molecules.

“Our calculations should help contribute to the improvement in the design of nano-devices for biomolecule analysis,” Polson says. “And yes, far far downstream, this will perhaps lead to important benefits in the health sciences.”

In the meantime, the experimentalists at McGill have expressed considerable interest in the calculations carried out by Kozma and Polson.

Polson says he hopes this story will illuminate the fact that undergraduate researchers can contribute a lot when the right supervision is in place, and when they have access to the right resources.

Understanding Polymers

James Polson spends much of his time trying to understand the mechanisms that make life possible. Polson is an associate professor of physics with the University of Prince Edward Island. He studies polymers – a group of long chainlike molecules composed of many repeated chemical subunits. Proteins and DNA are natural biopolymers that are essential to biological structure and function. Polson uses computer simulations and analytical theoretical methods to study the physical properties of polymers in confined and crowded environments. “We’re carrying out numerically intensive calculations to study model systems under conditions that are relevant to recent experiments using DNA,” he says. Polson’s research falls into three categories. The first is called polymer translocation – the study of what happens when polymer molecules are driven through a narrow hole in a barrier; a process similar to threading a needle. “There’s a lot of interest in this,” he says. “Understanding the basic physics of this process will provide insight useful for developing new translocation-based technologies to sequence DNA.” The second category of research involves studying the effects of squeezing polymers into narrow channels. In such environments, the molecules sometimes sample folded states, and Polson uses simulations to quantify the tendency for polymers to exist in these states. The results will be useful for the development of the technique of genomic mapping, which involves confining DNA in nanochannels. Polson is also studying the propensity of confined polymers to self-segregate – a process that is relevant to chromosome separation in bacteria. “Before a bacterium or a cell divides, it first replicates its chromosomes,” he says. “But there is a lot of uncertainty into the mechanisms involved in pulling those chromosomes apart. One possibility is that entropy provides the main driving force, and our simulations will provide insight into the importance of this effect.” As part of his research, Polson uses the ACENET computer network to create virtual models of polymer molecules to simulate their behaviour under various conditions. While the models are highly simplistic and designed to capture only the most basic features of real molecular systems, the simulations are nevertheless very time-consuming. “What I need from ACENET is computer power. For each calculation I need 100 to 200 processors for one to two days, and dozens of such calculations are needed for any given project.” Polson says the ongoing support he receives from the ACENET staff is crucial to his work. “Every summer I have students working with me. There are a lot of computer skills that they have to learn, but ACENET is really good at giving tutorials for newbies.”

Studying the Duplication of Genes

The duplication of DNA through the process of mitosis is essential to life. But that process doesn’t always go as planned. Sometimes the copies are flawed and the genes contained within them are mutated – the process that drives evolution. Sometimes extra copies of genes are produced. Those extra copies, or gene duplications, work like other gene mutations: sometimes the effect is good; other times it is bad or has no effect at all. Denise Clark is a geneticist and a professor of biology at the University of New Brunswick in Fredericton.She studies those gene duplications as part of her research and she’s using the ACENET computer network to do it. “When two copies of a gene are both functioning within a genome they can take on different roles,” says Clark. “We’re interested in how the the duplicated genes function and also in the mechanism that causes the duplications to arise.” Fruit flies are the primary organism Clark studies. The tiny insects are the gold standard for this kind of genetic research because they reproduce quickly and also because they have already been studied extensively by geneticists for many years. “There has been a determination of the entire DNA sequence, or “genome”, for hundreds individual fruit flies by several fruit fly research groups to look at genetic variation in populations,” she says. “That next-generation DNA sequencing data is available for anyone to look at.” “A couple of years ago I started using ACENET to look for duplicates. If those duplicates are localized – if they’re not present around the world – we can assume they are new.” ACENET provides a crucial piece of the puzzle. The file for one fruit fly genome is gigabytes in size, says Clark. The sheer size of the data that needs to be analyzed would quickly overtax a standard computer system. “I’ve looked at 500 genomes using ACENET. That’s something I couldn’t do on my laptop.” The job is made easier because ACENET has installed a number of genome tools that Clark can use off the shelf. It allows her to look at hundreds of genomes in parallel or string together the tools she wants to use in a pipeline.”I’m not building any new tools, except for the specific pipeline that lists what I need. I start with a genome file that’s gigabytes in size and at the end I’m left with a file that’s hundreds of kilobytes,” she says. Clark says that now that technology to sequence genes has gotten much cheaper and faster, there has been an explosion of new genetic knowledge and ideas recently. “A lot of geneticists’ ideas about genomes and genetic variation have changed since the development of next-generation sequencing technology.”