> top > projects

Projects

NameTDescription# Ann.AuthorMaintainerUpdated_atStatus

1-20 / 119 show all
bionlp-st-ge-2016-spacy-parsed Dependency parses produced by spaCy parser, and part-of-speech tags produced by Stanford tagger (with the wsj-0-18-left3words-nodistsim model). The exact procedure is described here. Data set contains the 34 full paper articles used in the BioNLP 2016 GE task. 226 KNico ColicNico Colic2016-05-25Released
bionlp-st-ge-2016-test-tees NER and event extraction produced by TEES (with the default GE11 model) for the 14 full papers used in the BioNLP 2016 GE task test corpus.9.17 KNico ColicNico Colic2016-05-25Released
bionlp-st-ge-2016-reference-tees NER and event extraction produced by TEES (with the default GE11 model) for the 20 full papers used in the BioNLP 2016 GE task reference corpus.14.6 KNico Colic Nico Colic2016-05-25Released
FSU-PRGE A new broad-coverage corpus composed of 3,306 MEDLINE abstracts dealing with gene and protein mentions. The annotation process was semi-automatic. Publication: http://aclweb.org/anthology/W/W10/W10-1838.pdf59.5 KCALBC ProjectYue Wang2017-03-08Released
spacy-test Random set of articles used for testing in the development of the RESTful spaCy parsing web service. Since development is now finished, they are released for the community to use.137 KNico ColicNico Colic2019-03-16Released
craft-sa-dev Development data for CRAFT SA shared task. This project contains the development (training) annotations for the Structural Annotation task of the CRAFT Shared Task 2019. This particular set contains token and sentence annotations with tokens linked via dependency relations. These dependency relations were automatically generated using the manually curated CRAFT constituency treebank files as input.512 KUniversity of Colorado Anschutz Medical Campuscraft-st2019-03-25Released
GlyCosmos600-docs A random collection of 600 PubMed abstracts from 6 glycobiology-related journals: Glycobiology, Glycoconjugate journal, The Journal of biological chemistry, Journal of proteome research, Journal of proteomics, and Carbohydrate research. The whole PMIDs were collected on June 11, 2019. From each journal, 100 PMIDs were randomly sampled.0Jin-Dong Kim2019-06-11Released
DisGeNET5_variant_disease The file contains variant-disease associations obtained by text mining MEDLINE abstracts using the BeFree system, including the variant and disease off sets. 144 KIBI GroupYue Wang2020-02-01Released
DisGeNET5_gene_disease The file contains gene-disease associations obtained by text mining MEDLINE abstracts using the BeFree system including the gene and disease off sets.2.04 MIBI GroupYue Wang2020-02-02Released
CORD-19_All_docs All the documents in the whole CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
CORD-19_bioRxiv_medRxiv_subset The bioRxiv/medRxiv subset of the CORD-19 dataset: pre-prints that are not peer reviewed. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT. 0Jin-Dong Kim2020-03-23Released
CORD-19_Commercial_use_subset The Commercial use subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
CORD-19_Non-commercial_use_subset The Non commercial use subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.0Jin-Dong Kim2020-03-23Released
PubMed_ArguminSci Predictions for PubMed automatically extracted with the ArguminSci tool (https://github.com/anlausch/ArguminSci).766 Kzebet2020-03-31Released
LitCovid-PubTatorCentral Named-entities for the documents in the LitCovid dataset. Annotations were automatically predicted by the PubTatorCentral tool (https://www.ncbi.nlm.nih.gov/research/pubtator/)4.64 Kzebet2020-04-01Released
LitCovid-OGER Using OGER (http://www.ontogene.org/resources/oger) to detect entities from 10 different vocabularies9.31 KFabio RinaldiNico Colic2020-04-02Released
LitCovid-OGER-BioBert Using OGER (http://www.ontogene.org/resources/oger) in conjunction with BioBert as described here (https://arxiv.org/pdf/2003.07424.pdf)4.03 KFabio RinaldiNico Colic2020-04-02Released
CORD-19_Custom_license_subset The Custom license subset of the CORD-19 dataset. The documents in this project will be updated as the CORD-19 dataset grows. See the COVID DATASET LICENSE AGREEMENT.5.08 MJin-Dong Kim2020-04-10Released
LitCovid-sentences Sentence segmentation of all the texts in the LitCovid literature. The segmentation is automatically obtained using the TextSentencer annotation service developed and maintained by DBCLS.16.5 KJin-Dong Kim2020-04-14Released
CORD-19-PD-MONDO PubDictionaries annotation for MONDO terms - updated at 2020-04-30 It is disease term annotation based on MONDO. Version 2020-04-20. The terms in MONDO are loaded in PubDictionaries, with which the annotations in this project are produced. The parameter configuration used for this project is here. Note that it is an automatically generated dictionary-based annotation. It will be updated periodically, as the documents are increased, and the dictionary is improved.6.32 MJin-Dong Kim2020-04-30Released
NameT# Ann.AuthorMaintainerUpdated_atStatus

1-20 / 119 show all
bionlp-st-ge-2016-spacy-parsed 226 KNico ColicNico Colic2016-05-25Released
bionlp-st-ge-2016-test-tees 9.17 KNico ColicNico Colic2016-05-25Released
bionlp-st-ge-2016-reference-tees 14.6 KNico Colic Nico Colic2016-05-25Released
FSU-PRGE 59.5 KCALBC ProjectYue Wang2017-03-08Released
spacy-test 137 KNico ColicNico Colic2019-03-16Released
craft-sa-dev 512 KUniversity of Colorado Anschutz Medical Campuscraft-st2019-03-25Released
GlyCosmos600-docs 0Jin-Dong Kim2019-06-11Released
DisGeNET5_variant_disease 144 KIBI GroupYue Wang2020-02-01Released
DisGeNET5_gene_disease 2.04 MIBI GroupYue Wang2020-02-02Released
CORD-19_All_docs 0Jin-Dong Kim2020-03-23Released
CORD-19_bioRxiv_medRxiv_subset 0Jin-Dong Kim2020-03-23Released
CORD-19_Commercial_use_subset 0Jin-Dong Kim2020-03-23Released
CORD-19_Non-commercial_use_subset 0Jin-Dong Kim2020-03-23Released
PubMed_ArguminSci 766 Kzebet2020-03-31Released
LitCovid-PubTatorCentral 4.64 Kzebet2020-04-01Released
LitCovid-OGER 9.31 KFabio RinaldiNico Colic2020-04-02Released
LitCovid-OGER-BioBert 4.03 KFabio RinaldiNico Colic2020-04-02Released
CORD-19_Custom_license_subset 5.08 MJin-Dong Kim2020-04-10Released
LitCovid-sentences 16.5 KJin-Dong Kim2020-04-14Released
CORD-19-PD-MONDO 6.32 MJin-Dong Kim2020-04-30Released