GlycoBiology-Motifs | | | 4.15 K | | Jin-Dong Kim | 2023-11-29 | | |
GlyCosmos6-docs | | | 0 | | Jin-Dong Kim | 2023-11-29 | Developing | |
LitCoin-entities-OrganismTaxon-PD | | | 949 | | Jin-Dong Kim | 2023-11-27 | Testing | |
PubMed-2000 | | abstracts published in 2000. | 0 | | Jin-Dong Kim | 2023-11-29 | Developing | |
LitCoin-Chemical-MeSH-CHEBI | | ChemicalEntity:
Annotated by PD-MeSH2022_CHEBI_tuned-B | 3.84 K | | yucca | 2023-11-29 | Testing | |
TEST-ChemicalEntity | | ChemicalEntity : Annotated by PD-MeSH2022_CHEBI_tuned-B | 827 | | yucca | 2023-11-29 | Beta | |
CORD-19_Commercial_use_subset | | The Commercial use subset of the CORD-19 dataset.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT. | 0 | | Jin-Dong Kim | 2023-11-29 | Released | |
GlyTouCan-IUPAC | | | 399 K | | kiyoko | 2023-11-29 | Testing | |
GlycoGenes | | annotation for glyco-genes based on GGDB | 1.01 K | | Jin-Dong Kim | 2023-11-29 | Developing | |
CORD-19_bioRxiv_medRxiv_subset | | The bioRxiv/medRxiv subset of the CORD-19 dataset: pre-prints that are not peer reviewed.
The documents in this project will be updated as the CORD-19 dataset grows.
See the COVID DATASET LICENSE AGREEMENT.
| 0 | | Jin-Dong Kim | 2023-11-29 | Released | |
OryzaGP_2022 | | | 41.3 K | | larmande | 2023-11-24 | | |
hydroxychloroquine | | | 2.59 K | | Jin-Dong Kim | 2023-11-29 | Developing | |
SPECIES800_autotagged | | This project comprises the SPECIES800 corpus documents automatically annotated by the Jensenlab tagger.
Annotated entity types are:
Genes/proteins from the mentioned organisms (and any human ones)
PubChem Compound identifiers
NCBI Taxonomy entries
Gene Ontology cellular component terms
BRENDA Tissue Ontology terms
Disease Ontology terms
Environment Ontology terms
The SPECIES 800 (S800) comprises 800 PubMed abstracts. In its original form species mentions were manually identified and mapped to the corresponding NCBI Taxonomy identifiers.
Described in:
The SPECIES and ORGANISMS Resources for Fast and Accurate Identification of Taxonomic Names in Text.
Pafilis E, Frankild SP, Fanini L, Faulwetter S, Pavloudi C, et al. (2013). PLoS ONE, 2013, 8(6): e65390. doi:10.1371/journal.pone.0065390.
The manually annotated corpus is also available as a PubAnnotation project (see here).
| 0 | Evangelos Pafilis, Sampo Pyysalo, Lars Juhl Jensen | evangelos | 2015-11-20 | Testing | |
DisGeNET5_variant_disease | | The file contains variant-disease associations obtained by text mining MEDLINE abstracts using the BeFree system, including the variant and disease off sets. | 144 K | IBI Group | Yue Wang | 2023-11-24 | Released | |
OryzaGP | | A dataset for Named Entity Recognition for rice gene | 29.1 K | Huy Do and Pierre Larmande | Yue Wang | 2023-11-24 | Uploading | |
AlvisNLP-Async-Test | | Test for the asynchronous AlvisNLP/ML annotator family. | 0 | Robert Bossy | rbossy | 2023-11-26 | Testing | |
GlyCosmos600-GlycoEpitope | | | 277 | | Jin-Dong Kim | 2023-11-27 | Testing | |
LitCovid-PubTator | | | 5.88 M | | Jin-Dong Kim | 2023-11-24 | Beta | |
Parkinson | | | 54 | | Jin-Dong Kim | 2023-11-28 | Testing | |
preeclampsia_genes | | | 17.8 K | | Jin-Dong Kim | 2023-11-29 | Developing | |