CyanoBase | | Cyanobacteria are prokaryotic organisms that have served as important model organisms for studying oxygenic photosynthesis and have played a significant role in the Earthfs history as primary producers of atmospheric oxygen.
Publication: http://www.aclweb.org/anthology/W12-2430 | 1.1 K | 2016-05-17 | Released | |
0mytest | | | 144 | 2021-12-03 | | |
AIMed | | The AIMed corpus is one of the most widely used corpora for protein-protein interaction extraction. The protein annotations are either parts of the protein interaction annotations, or are uninvolved in any protein interaction annotation.
Publication: http://www.cs.utexas.edu/~ml/papers/bionlp-aimed-04.pdf | 4.04 K | 2017-04-14 | Testing | |
bionlp-st-pc-2013-training | | The training dataset from the pathway curation (PC) task in the BioNLP Shared Task 2013.
The entity types defined in the PC task are simple chemical, gene or gene product, complex and cellular component. | 7.86 K | 2017-08-28 | Released | |
FSU-PRGE | | A new broad-coverage corpus composed of 3,306 MEDLINE abstracts dealing with gene and protein mentions.
The annotation process was semi-automatic.
Publication: http://aclweb.org/anthology/W/W10/W10-1838.pdf | 59.5 K | 2017-03-08 | Released | |
PIR-corpus2 | | The protein tag was used to tag proteins, or protein-associated or -related objects, such as domains, pathways, expression of gene.
Annotation guideline: http://pir.georgetown.edu/pirwww/about/doc/manietal.pdf | 5.52 K | 2020-02-01 | Released | |
SCAI-Test | | A small corpus for the evaluation of dictionaries containing chemical entities.
Publication: http://www.scai.fraunhofer.de/fileadmin/images/bio/data_mining/paper/kolarik2008.pdf
Original source: https://www.scai.fraunhofer.de/en/business-research-areas/bioinformatics/downloads/corpora-for-chemical-entity-recognition.html | 1.21 K | 2017-04-03 | Released | |
bionlp-st-epi-2011-training | | The training dataset from the Epigenetics and Post-translational Modifications (EPI) task in the BioNLP Shared Task 2011.
The core entities of the task are genes and gene products (RNA and proteins), identified in the data simply as "Protein" annotations. | 7.59 K | 2021-03-10 | Released | |
OryzaGP | | A dataset for Named Entity Recognition for rice gene | 29.1 K | 2021-03-11 | Uploading | |
DisGeNET5_variant_disease | | The file contains variant-disease associations obtained by text mining MEDLINE abstracts using the BeFree system, including the variant and disease off sets. | 144 K | 2020-02-01 | Released | |