Blazing new paths in scientific software

A student needed a better way to make sense of thousands of DNA sequences. A professor saw an opportunity to solve a problem that had frustrated biologists for years. Together, they created an open-source software package that gives researchers around the world a powerful new way to visualize how species in a specific geographic area are genetically connected.
Assistant Professor of Biology Andrew Davinack and Rylie Seaberg ’26 recently published “pygenoscape: a Python package for spatial interpolation and visualization of genetic distance landscapes” in Bioinformatics Advances. Their software, called pygenoscape, transforms complex genetic data into interactive three-dimensional landscapes, allowing researchers to see how genetic variation changes across geographic space.
The software grew out of Seaberg’s honors thesis. As she assembled and analyzed thousands of genetic sequences from human and veterinary parasites, she quickly discovered that existing software could not easily handle datasets of that size.
“A similar program called Alleles in Space came out in 2005, but it only works for Windows, has some bugs and can’t handle as much data as our software,” Davinack said. “You write one line of code and it produces your visualization.”
The software helps answer fundamental questions in biology. Are populations becoming genetically isolated simply because they live far apart? Or are rivers, mountains, dams or other barriers preventing animals from interbreeding? Pygenoscape displays those answers as peaks and valleys across a three-dimensional landscape, giving scientists an intuitive way to identify how populations of animals differ genetically across a given landscape.
“If you want to study how genetically related gray wolf populations are in Yellowstone National Park, for example, you can extract DNA from 200 animals and compare their genetic distances,” Davinack explained. “This software takes those data points and overlays them in a three-dimensional space. A completely flat landscape means the wolves are genetically similar, but an uneven landscape interrupted by sharp peaks indicate regions where there is much greater genetic variation.”
Unlike earlier programs, pygenoscape is designed to work with both traditional DNA sequence data and modern genome-scale datasets containing thousands of genetic markers. It is also freely available, making the tool accessible to researchers worldwide. The software arrives at a time when many biological researchers are increasingly turning to Python for large-scale data analysis.
“In the biological sciences, there is a strong preference for software written in R,” Davinack said. “Today, more researchers are seeing the advantages of using a general purpose programming language like Python, and we’re hoping this package helps advance its use in the field.”
The project also demonstrates how undergraduate research can contribute to advances in science.

After completing a computational chemistry internship at Northeastern University, Seaberg approached Davinack about pursuing an honors thesis. Her project, titled “A Comparative Phylogeography of Three Nematode Parasites with Contrasting Life Cycles,” compared genetic differences among parasites that affect humans and animals. As her dataset continued to grow, Davinack realized they needed a more powerful visualization tool.
Once pygenoscape was developed, Seaberg became its first major tester.
“She pulled thousands of sequences from databases,” Davinack said. “We knew if it worked with her dataset, it would easily work with any dataset.”
Davinack credited Seaberg with playing an essential role in refining the software.
“Rylie tested the stripped-down version of pygenoscape using large genetic datasets that she compiled and processed for her honors thesis,” he said. “She also assisted in debugging and validation, and contributed to evaluating how effectively the program recovered known biological patterns.”
The software was ultimately tested on datasets ranging from invasive bumblebees to marine invertebrates before the pair published their findings.
For Seaberg, the experience confirmed where her own interests lie.
“Doing research and writing a thesis helped me realize that I enjoyed looking at data more than specific genes and genomics,” she said.
She also credits the mentorship she received from Davinack, Professor of the Practice of Biology Stacey Nguyen and Professor of Mathematics Michael Kahn.
“They all have a similar mentorship style,” Seaberg said. “They allowed me to do independent work, to learn by making my own mistakes, and they supported me when I needed assistance. Because of the relationships I built with professors, I was afforded multiple opportunities at Wheaton—including serving as a teaching assistant.”
This fall, Seaberg will begin pursuing a master’s degree in data science at Boston University.
Beyond academic research, Davinack sees important practical applications for pygenoscape, particularly in conservation biology.
“In the aquatic sciences, many dams are being removed due to weakening functionality,” he said. “These visualizations can tell you whether those particular structures are causing genetic variation differences among aquatic species. They can give conservation managers the tools to determine whether a particular conservation strategy should move forward.”
The project also highlights the strengths of Wheaton’s interdisciplinary bioinformatics program, which combines biology, mathematics, computer science, biological chemistry and environmental science. Rather than following a single prescribed path, students design programs that reflect their interests and often tackle original research questions.
“If they choose to pursue a research project, they get to craft it themselves, which provides more room to be intellectually curious,” Davinack said. “Helping design bioinformatics software packages is one of the unique experiences bioinformatics students can obtain if they choose to do honors-level research in my lab.”
That combination of computational skills and scientific knowledge is increasingly valuable as biology becomes more data-intensive.
“Once you’re a bioinformatician, you can work at nonprofits, hospitals, sequencing facilities, biotech startups or organizations with very specific missions, such as the National Cancer Institute,” Davinack said. “There’s a need for more people who know how to work with a lot of data, especially with the rise of AI.”