> top > projects

Projects

NameTDescription# Ann.AuthorMaintainerUpdated_atStatus

1-20 / 115 show all
FSU-PRGE A new broad-coverage corpus composed of 3,306 MEDLINE abstracts dealing with gene and protein mentions. The annotation process was semi-automatic. Publication: http://aclweb.org/anthology/W/W10/W10-1838.pdf59.5 KCALBC ProjectYue Wang2017-03-08Released
spacy-test Random set of articles used for testing in the development of the RESTful spaCy parsing web service. Since development is now finished, they are released for the community to use.137 KNico ColicNico Colic2019-03-16Released
GlyCosmos600-docs A random collection of 600 PubMed abstracts from 6 glycobiology-related journals: Glycobiology, Glycoconjugate journal, The Journal of biological chemistry, Journal of proteome research, Journal of proteomics, and Carbohydrate research. The whole PMIDs were collected on June 11, 2019. From each journal, 100 PMIDs were randomly sampled.0Jin-Dong Kim2019-06-11Released
DisGeNET5_variant_disease The file contains variant-disease associations obtained by text mining MEDLINE abstracts using the BeFree system, including the variant and disease off sets. 144 KIBI GroupYue Wang2020-02-01Released
DisGeNET5_gene_disease The file contains gene-disease associations obtained by text mining MEDLINE abstracts using the BeFree system including the gene and disease off sets.2.04 MIBI GroupYue Wang2020-02-02Released
CORD-19_All_docs All the documents in the whole CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
CORD-19_bioRxiv_medRxiv_subset The bioRxiv/medRxiv subset of the CORD-19 dataset: pre-prints that are not peer reviewed. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT. 0Jin-Dong Kim2020-03-23Released
CORD-19_Commercial_use_subset The Commercial use subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
CORD-19_Non-commercial_use_subset The Non commercial use subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
PubMed_ArguminSci Predictions for PubMed automatically extracted with the ArguminSci tool (https://github.com/anlausch/ArguminSci).777 Kzebet2020-03-31Released
LitCovid-PubTatorCentral Named-entities for the documents in the LitCovid dataset. Annotations were automatically predicted by the PubTatorCentral tool (https://www.ncbi.nlm.nih.gov/research/pubtator/)4.64 Kzebet2020-04-01Released
LitCovid-OGER Using OGER (http://www.ontogene.org/resources/oger) to detect entities from 10 different vocabularies9.31 KFabio RinaldiNico Colic2020-04-02Released
CORD-19_Custom_license_subset The Custom license subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.5.08 MJin-Dong Kim2020-04-10Released
LitCovid-sentences Sentence segmentation of all the texts in the LitCovid literature. The segmentation is automatically obtained using the TextSentencer annotation service developed and maintained by DBCLS.16.5 KJin-Dong Kim2020-04-14Released
CORD-19-PD-MONDO PubDictionaries annotation for MONDO terms - updated at 2020-04-30 It is disease term annotation based on MONDO. Version 2020-04-20. The terms in MONDO are loaded in PubDictionaries, with which the annotations in this project are produced. The parameter configuration used for this project is here. Note that it is an automatically generated dictionary-based annotation. It will be updated periodically, as the documents are increased, and the dictionary is improved.6.32 MJin-Dong Kim2020-04-30Released
CORD-19-PD-UBERON PubDictionaries annotation for UBERON terms - updated at 2020-04-30 It is disease term annotation based on Uberon. The terms in Uberon are uploaded in PubDictionaries (Uberon), with which the annotations in this project are produced. The parameter configuration used for this project is here. Note that it is an automatically generated dictionary-based annotation. It will be updated periodically, as the documents are increased, and the dictionary is improved.1.42 MJin-Dong Kim2020-04-30Released
CORD-19-PD-HP PubDictionaries annotation for HP terms - updated at 2020-04-30 Disease term annotation based on HP. Version 2020-04-20. The terms in HP are loaded in PubDictionaries, with which the annotations in this project are produced. The parameter configuration used for this project is here. Note that it is an automatically generated dictionary-based annotation. It will be updated periodically, as the documents are increased, and the dictionary is improved.1.15 MJin-Dong Kim2020-05-12Released
LitCovid-OGER-BB Using OGER (www.ontogene.com) and Biobert to obtain annotations for 10 different vocabularies.308 KFabio RinaldiNico Colic2020-06-04Released
bionlp-st-ge-2016-reference-tees NER and event extraction produced by TEES (with the default GE11 model) for the 20 full papers used in the BioNLP 2016 GE task reference corpus.14.6 KNico Colic Nico Colic2020-09-13Released
bionlp-st-ge-2016-spacy-parsed Dependency parses produced by spaCy parser, and part-of-speech tags produced by Stanford tagger (with the wsj-0-18-left3words-nodistsim model). The exact procedure is described here. Data set contains the 34 full paper articles used in the BioNLP 2016 GE task. 225 KNico ColicNico Colic2020-10-02Released
NameT# Ann.AuthorMaintainerUpdated_atStatus

1-20 / 115 show all
FSU-PRGE 59.5 KCALBC ProjectYue Wang2017-03-08Released
spacy-test 137 KNico ColicNico Colic2019-03-16Released
GlyCosmos600-docs 0Jin-Dong Kim2019-06-11Released
DisGeNET5_variant_disease 144 KIBI GroupYue Wang2020-02-01Released
DisGeNET5_gene_disease 2.04 MIBI GroupYue Wang2020-02-02Released
CORD-19_All_docs 0Jin-Dong Kim2020-03-23Released
CORD-19_bioRxiv_medRxiv_subset 0Jin-Dong Kim2020-03-23Released
CORD-19_Commercial_use_subset 0Jin-Dong Kim2020-03-23Released
CORD-19_Non-commercial_use_subset 0Jin-Dong Kim2020-03-23Released
PubMed_ArguminSci 777 Kzebet2020-03-31Released
LitCovid-PubTatorCentral 4.64 Kzebet2020-04-01Released
LitCovid-OGER 9.31 KFabio RinaldiNico Colic2020-04-02Released
CORD-19_Custom_license_subset 5.08 MJin-Dong Kim2020-04-10Released
LitCovid-sentences 16.5 KJin-Dong Kim2020-04-14Released
CORD-19-PD-MONDO 6.32 MJin-Dong Kim2020-04-30Released
CORD-19-PD-UBERON 1.42 MJin-Dong Kim2020-04-30Released
CORD-19-PD-HP 1.15 MJin-Dong Kim2020-05-12Released
LitCovid-OGER-BB 308 KFabio RinaldiNico Colic2020-06-04Released
bionlp-st-ge-2016-reference-tees 14.6 KNico Colic Nico Colic2020-09-13Released
bionlp-st-ge-2016-spacy-parsed 225 KNico ColicNico Colic2020-10-02Released