Skip to main content

What if all bioscience knowledge was stored in one universally accessible place?


LA JOLLA, CA—Putting all the world’s bioscience knowledge into one open-access, uniformly structured repository could turbo-charge bioscience research, enabling better insights into the causes of diseases and faster discoveries of new treatments, argues an international and interdisciplinary team of researchers led by bioinformatics scientists at Scripps Research.

The researchers’ call to action, published March 17 in the journal eLife, is an attempt to solve the problem that life sciences data today are spread throughout thousands of bioscience journals and differently formatted datasets maintained by different groups, often with restricted access.

The authors urge that bioscience researchers publish their data in a more organized way that allows open access without licensing restrictions—access that would permit automatic software programs, or “bots,” to gather those data and add them to a central online repository called Wikidata.

“We’re essentially asking our fellow scientists to consider putting their scientific findings out there in a way that will allow others to make better use of them and combine them more easily with other datasets to provide new scientific insights for the benefit of all,” says senior author Andrew Su, PhD, a professor in the Department of Integrative Structural and Computational Biology at Scripps Research.

Wikidata was started in 2012 by the Wikimedia Foundation—which also runs wikipedia.org—as an open-access, document-oriented database, assembled by thousands of collaborative editors and programmers. Unlike Wikipedia, Wikidata uses a standard structure for the information it contains, which allows that information to be queried automatically by computer programs and artificial intelligence systems, and also permits specially designed bots to build new Wikidata entries from accessible non-Wikidata information sources. As of September 2019, Wikidata comprised over 750 million information blocks, or “statements,” concerning 61million items.

There is no limit to the categories of information stored in Wikidata, but participants say the platform holds particular value for scientific information, and it already includes many entries spanning genomics, proteomics, biochemistry, pharmaceutical chemistry and other fields in the life sciences. Proponents view it as a good candidate for an ultimate, universal repository of bioscience and other data.

“The idea is that, for example, if I were to integrate database A and database B in the context of Wikidata, that’s work that nobody else would have to do again,” Su says.

The authors emphasize that life sciences data that are incorporated into Wikidata—and are thereby accessible to researchers and their analytical tools—can be used in ways that benefit not only scientists, but also the wider public. The data may enable the assembly of easily accessible models of biological pathways and processes; AI-type programs that suggest disease diagnoses based on patients’ symptoms and signs; and programs that predict new uses (“repurposing”) for existing drugs—a function that would be particularly useful in medical crises such as the ongoing COVID-19 pandemic.

Su points out that scientists should be able to draw more and more connections and insights from that information base as it grows larger.

“Nowadays, scientific data are still fragmented across many different data ‘silos’, which makes it hard to integrate and query and make use of those data—and that’s the problem Wikidata is trying to solve,” he says.

Su and his co-authors appeal to fellow life scientists to present their information in a Wikidata-accessible way, which perhaps most importantly includes licensing the use of the data according to the Creative Commons Zero concept—putting it in the public domain, essentially. “CC0” licensing declarations are already used widely to allow open access and re-use of data. In the Wikidata context, they ensure that data can be accessed by Wikidata-mining bots and added to Wikidata’s knowledge base without legal hurdles.

Such hurdles often exist because scientists or their sponsors want to keep data private for commercial, competitive reasons. Su and his colleagues urge a relaxation of that attitude.

“Knowledge about such things as genes, proteins and biological activities of existing drug compounds is fundamental scientific information that we really should be collaborating to assemble and share,” Su says. “We can always compete downstream, for example in mining that information to develop drugs and other commercial products.”

The open-access and CC0 movement has advanced to a great degree already. In the life sciences, hundreds of journals—including eLife and the Public Library of Science (PLoS) series—already participate in making information easily accessible.

“There is definitely a trend and a growing appreciation for the value of the open-access and open-data approach, and we just want to accelerate along that path,” Su says.

Thirty scientists from multiple institutions co-authored “Wikidata as a knowledge graph for the life sciences.”

Groundbreaking Science.
Life-changing Medicine.