CORD-19_bioRxiv_medRxiv_subset | | The bioRxiv/medRxiv subset of the CORD-19 dataset: pre-prints that are not peer reviewed.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT.
| 0 | 2023-11-29 | Released | |
bionlp-st-ge-2016-coref | | Coreference annotation to the benchmark data set (reference and test) of BioNLP-ST 2016 GE task.
For detailed information, please refer to the benchmark reference data set (bionlp-st-ge-2016-reference) and benchmark test data set (bionlp-st-ge-2016-test). | 853 | 2024-06-17 | Released | |
GlyCosmos600-docs | | A random collection of 600 PubMed abstracts from 6 glycobiology-related journals: Glycobiology, Glycoconjugate journal, The Journal of biological chemistry, Journal of proteome research, Journal of proteomics, and Carbohydrate research. The whole PMIDs were collected on June 11, 2019. From each journal, 100 PMIDs were randomly sampled. | 0 | 2023-11-29 | Released | |
bionlp-st-ge-2016-test-proteins | | Protein annotations to the benchmark test data set of the BioNLP-ST 2016 GE task.
A participant of the GE task may import the documents and annotations of this project to his/her own project, to begin with producing event annotations.
For more details, please refer to the benchmark test data set (bionlp-st-ge-2016-test).
| 4.34 K | 2023-11-27 | Released | |
CORD-19-PD-UBERON | | PubDictionaries annotation for UBERON terms - updated at 2020-04-30
It is disease term annotation based on Uberon.
The terms in Uberon are uploaded in PubDictionaries
(Uberon), with which the annotations in this project are produced.
The parameter configuration used for this project is
here.
Note that it is an automatically generated dictionary-based annotation. It will be updated periodically, as the documents are increased, and the dictionary is improved. | 1.42 M | 2023-11-24 | Released | |
CORD-19_Non-commercial_use_subset | | The Non commercial use subset of the CORD-19 dataset.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT. | 0 | 2023-11-29 | Released | |
LitCovid-sentences-v1 | | Sentence segmentation of all the texts in the LitCovid literature. The segmentation is automatically obtained using the TextSentencer annotation service developed and maintained by DBCLS. | 16.5 K | 2023-11-27 | Released | |
CORD-19_Custom_license_subset | | The Custom license subset of the CORD-19 dataset.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT. | 5.08 M | 2023-11-24 | Released | |
CORD-19_Commercial_use_subset | | The Commercial use subset of the CORD-19 dataset.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT. | 0 | 2023-11-29 | Released | |
CORD-19_All_docs | | All the documents in the whole CORD-19 dataset.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT. | 0 | 2023-11-29 | Released | |