Skip to main content

Data Science professor awarded grant to support partnership with Internet Archive and preserve digital history

Assistant Professor of data science Alexander Nwala leads a project to preserve web resources referenced in U.S. television news. (Photo illustration by Stephen Salpukas)

We've all heard that what goes on the internet lasts forever, but what does that really mean? In reality, much of the internet quietly disappears over time, vanishing into a digital abyss of broken links, deleted pages and forgotten stories. For future generations, piecing together our digital history may prove nearly impossible.

To help fight the loss of digital media, Dr. Alexander Nwala is leading a collaborative research project, CiteCast, to preserve web resources referenced in U.S. television news. Backed by a new $750,000 grant from the Institute of Museum and Library Services, Nwala, in collaboration with Old Dominion University and the Internet Archive, will develop AI models that automatically identify, archive and catalog these referenced web pages in a publicly accessible database, ensuring they remain available to researchers, journalists and future generations.

The project reflects a broader mission of preserving media that might otherwise be lost.

Nwala, an assistant professor of data science, is working with Dr. Yanfu Zhang, assistant professor of computer science, to develop an AI model capable of detecting U.S. local and national TV news site web resources. From there, researchers at Old Dominion will use that model to analyze the past decade of U.S. local and national television news broadcasts, and the Internet Archive will use the information to store in their WayBack machine.

"This is inherently a preservation project where we want to preserve those primary resources because they are valuable," Nwala explains.

Beyond the technical tools themselves, the project is driven by a curiosity about how information moves between media forms.

"We have so many different interesting research questions we want to understand, which is just tracking how much information flows between these two massive modes of mass communication," said Nwala.

As the internet has grown, the line between the web and other forms of media has blurred, and the constant exposure to multiple media sources at once has rapidly changed how information reaches the public.

"News still plays a huge role in the decisions we make." expressed Nwala. "What we are doing is just showing that there's a connection between two different sources of information, one being television news, one being the web, so other folks just knowing the volume of traffic flowing between the news and the web is valuable."

Television news broadcasts are often picky with what information they include in their segments, so anything pulled from the internet tends to carry more significance.

"For TV news editors to take out valuable TV news time to cite something on the web tells you that you should pay attention to it, that there's value there," explained Nwala.

Instead of letting that information float into the abyss, CiteCast is working to preserve this information in the Wayback Machine.

In the early stages of the project, Nwala is working with Zhang and the William & Mary computer science department to build the AI tool that will serve as the backbone for CiteCast, training the model to identify web resources.

For Nwala, this work goes far beyond a simple citation.

"There's value in preserving [TV web citations] because there have been so many instances where you always have to go back in time to learn about something, and if no one is preserving it, unfortunately you're going to lose a lot," emphasized Nwala. "We ought to preserve it for future scientists, for future researchers, for the future public that may not even know it's valuable now."