A pandemic of papers
When COVID-19 emerged in early 2020, it triggered not only a global public health crisis but also an unprecedented surge in scientific output. Within weeks, preprint servers and journals were flooded with new coronavirus-related studies, rapidly expanding an already substantial body of research on coronaviruses. For researchers, clinicians, and policymakers trying to make sense of this rapidly growing literature, this was both a blessing and a burden: a wealth of knowledge was becoming available, but it was arriving far too quickly to read, assess, and prioritize effectively.
This was the challenge that the BIP! Services team set out to address with BIP4COVID19. Working intensively under the extraordinary circumstances of the pandemic, the team released the first version of the service, together with version 0.1 of its underlying dataset, in March 2020, at a time when the COVID-19 outbreak was reaching its peak worldwide.
Using impact indicators to assist COVID19-related literature search
BIP4COVID19 was an open dataset accompanied by a web application, hosted at bip.covid19.athenarc.gr, designed to help people explore COVID-19-related scientific literature more effectively. Ιts core idea was straightforward: add impact indicators to existing COVID-19 literature collections, giving users additional signals for identifying articles that were attracting attention from the scientific community and beyond.
The project was later described in the paper "BIP4COVID19: Releasing impact measures for articles relevant to COVID-19". It built on two well-known open literature resources available at the time:
-
CORD-19, the COVID-19 Open Research Dataset released by the Semantic Scholar team
-
LitCovid, a curated collection of COVID-19 literature maintained by the National Library of Medicine
By cleaning, deduplicating, and integrating records from these sources, together with information from PMC and social networks, the team assembled an open dataset providing impact indicators for COVID-19-related literature that eventually grew to contain over 600,000 unique articles.
BIP4COVID19 included the traditional citation-based impact indicators already available through other BIP! Services, like BIP! Finder and BIP! DB: Citation count, Pagerank, RAM, and Attrank. For BIP4COVID19, however, these indicators were calculated specifically on the citation network constructed from the COVID-19 literature collection. To construct the citation network needed for these calculations, the team used NCBI's eLink tool to gather citation links between articles in the collection.
There was, however, an important limitation to citation-based indicators in this particular setting: citation lag. The BIP4COVID19 corpus was dominated by very recent publications. Even highly influential papers might take months or years to accumulate citations, making traditional citation indicators less informative for identifying emerging research during a rapidly evolving pandemic. To complement citation-based measures, BIP4COVID19 therefore also incorporated an altmetric signal based on Twitter (now X):
-
Tweet Count, a measure of social media attention derived from a large hydrated collection of pandemic-related tweets.
This provided a more immediate signal of attention, helping capture interest in newly published research before that interest could be reflected in conventional citation counts.
An openly available and continuously updated resource
The dataset was released openly on Zenodo and updated regularly throughout the pandemic. This made it useful not only through the BIP4COVID19 web application but also as a resource that other developers and researchers could reuse directly. Researchers could use the data for bibliometric and scientometric studies, while developers could build their own applications and analyses on top of the dataset. In this sense, BIP4COVID19 was more than a search interface: it was an attempt to create reusable research infrastructure around a rapidly evolving body of scientific literature.
Where it stands today
BIP4COVID19 was very much a product of its moment: a rapid response to an urgent and fast-evolving information crisis.
The original web application at bip.covid19.athenarc.gr has since been deprecated, but the project’s legacy remains in the openly archived dataset. Across the lifetime of the project, 118 versions of the dataset were released, receiving more than 75,000 downloads on Zenodo. The final version was released in January 2023 and the web application became inactive sometime afterwards, having fulfilled the purpose for which it was created.
Looking back, BIP4COVID19 offers a nice example of how quickly parts of the scientometrics and research-infrastructure community mobilized during 2020. The goal was not to conduct COVID-19 research itself, but to build some of the connective tissue around that research: infrastructure that could help people find, assess, and prioritize an overwhelming volume of scientific knowledge when doing so mattered most.
And while the original service is no longer active, the dataset remains as a record of that extraordinary period, and of a small, focused effort to help the research community navigate the pandemic's other explosion: its explosion of papers.
|
Blog post DOI: https://doi.org/10.5281/zenodo.22142270 |