PubAnnotation

> top > users > Yue Wang

Yue Wang

User info

Collections

Name		Description	Updated at

1-3 / 3
DisGeNET5		Associations obtained by text mining MEDLINE abstracts using the BeFree system	2019-03-11
PIR		Protein Information Resource (PIR)	2019-03-12
AnEM		the largest manually annotated corpus on anatomical entities	2019-04-03

Projects

Name	T	Description	# Ann.	Updated at	Status

« 1 2 3 » 11-20 / 25 show all
TEST0			3.37 M	2023-11-24
funRiceGenes-all			1.51 K	2023-11-29	Developing
PennBioIE		The PennBioIE corpus (0.9) covers two domains of biomedical knowledge. One is the inhibition of the cytochrome P450 family of enzymes (CYP450 or CYP for short) , and the other domain is the molecular genetics of dance (oncology or onco for short).	23.8 K	2023-11-26	Released
DisGeNET5_gene_disease		The file contains gene-disease associations obtained by text mining MEDLINE abstracts using the BeFree system including the gene and disease off sets.	2.04 M	2023-11-24	Released
0_colil			781 K	2023-11-24
bionlp-st-cg-2013-training		The training dataset from the cancer genetics task in the BioNLP Shared Task 2013. Composed of anatomical and molecular entities.	10.9 K	2023-11-28	Released
AnEM_full-texts		250 documents selected randomly from full-text papers Entity types: organism subdivision, anatomical system, organ, multi-tissue structure, tissue, cell, developing anatomical structure, cellular component, organism substance, immaterial anatomical entity and pathological formation Together with AnEM_abstracts, it is probably the largest manually annotated corpus on anatomical entities.	687	2023-11-29	Uploading
jnlpba-st-training		The training data used in the task came from the GENIA version 3.02 corpus, This was formed from a controlled search on MEDLINE using the MeSH terms "human", "blood cells" and "transcription factors". From this search, 1,999 abstracts were selected and hand annotated according to a small taxonomy of 48 classes based on a chemical classification. Among the classes, 36 terminal classes were used to annotate the GENIA corpus. For the shared task only the classes protein, DNA, RNA, cell line and cell type were used. The first three incorporate several subclasses from the original taxonomy while the last two are interesting in order to make the task realistic for post-processing by a potential template filling application. The publication year of the training set ranges over 1990~1999.	51.1 K	2023-11-26	Released
CyanoBase		Cyanobacteria are prokaryotic organisms that have served as important model organisms for studying oxygenic photosynthesis and have played a significant role in the Earthfs history as primary producers of atmospheric oxygen. Publication: http://www.aclweb.org/anthology/W12-2430	1.1 K	2023-11-26	Released
FSU-PRGE		A new broad-coverage corpus composed of 3,306 MEDLINE abstracts dealing with gene and protein mentions. The annotation process was semi-automatic. Publication: http://aclweb.org/anthology/W/W10/W10-1838.pdf	59.5 K	2023-11-26	Released

Automatic annotators

Name	Description

1-2 / 2
PTO-all
PTO-exact

Editors

none