PubAnnotation

> top > users > Yue Wang

Yue Wang

User info

Collections

Name		Description	Updated at

1-3 / 3
DisGeNET5		Associations obtained by text mining MEDLINE abstracts using the BeFree system	2019-03-11
PIR		Protein Information Resource (PIR)	2019-03-12
AnEM		the largest manually annotated corpus on anatomical entities	2019-04-03

Projects

Name	T	Description	# Ann.	Updated at	Status

« 1 2 3 » 11-20 / 25 show all
GENIAcorpus		multi_cell (1,782) mono_cell (222) virus (2,136) protein_family_or_group (8,002) protein_complex (2,394) protein_molecule (21,290) protein_subunit (942) protein_substructure (129) protein_domain_or_region (1,044) protein_other (97) peptide (521) amino_acid_monomer (784) DNA_family_or_group (332) DNA_molecule (664) DNA_substructure (2) DNA_domain_or_region (39) DNA_other (16) RNA_family_or_group (1,545) RNA_molecule (554) RNA_substructure (106) RNA_domain_or_region (8,237) RNA_other (48) polynucleotide (259) nucleotide (243) lipid (2,375) carbohydrate (99) other_organic_compound (4,113) body_part (461) tissue (706) cell_type (7,473) cell_component (679) cell_line (4,129) other_artificial_source (211) inorganic (258) atom (342) other (21,056)	78.9 K	2023-11-29	Released
SCAI-Test		A small corpus for the evaluation of dictionaries containing chemical entities. Publication: http://www.scai.fraunhofer.de/fileadmin/images/bio/data_mining/paper/kolarik2008.pdf Original source: https://www.scai.fraunhofer.de/en/business-research-areas/bioinformatics/downloads/corpora-for-chemical-entity-recognition.html	1.21 K	2023-11-28	Released
bionlp-st-epi-2011-training		The training dataset from the Epigenetics and Post-translational Modifications (EPI) task in the BioNLP Shared Task 2011. The core entities of the task are genes and gene products (RNA and proteins), identified in the data simply as "Protein" annotations.	7.59 K	2023-11-29	Released
DisGeNET5_variant_disease		The file contains variant-disease associations obtained by text mining MEDLINE abstracts using the BeFree system, including the variant and disease off sets.	144 K	2023-11-24	Released
PIR-corpus1		The Protein Information Resource (PIR) is not biased towards any particular biomedical domain, and is expected to provide more diverse protein names in a given sample size. Annotation category: protein, compound-protein, acronym.	4.44 K	2023-11-27	Released
PennBioIE		The PennBioIE corpus (0.9) covers two domains of biomedical knowledge. One is the inhibition of the cytochrome P450 family of enzymes (CYP450 or CYP for short) , and the other domain is the molecular genetics of dance (oncology or onco for short).	23.8 K	2023-11-26	Released
OryzaGP		A dataset for Named Entity Recognition for rice gene	29.1 K	2023-11-24	Uploading
AnEM_full-texts		250 documents selected randomly from full-text papers Entity types: organism subdivision, anatomical system, organ, multi-tissue structure, tissue, cell, developing anatomical structure, cellular component, organism substance, immaterial anatomical entity and pathological formation Together with AnEM_abstracts, it is probably the largest manually annotated corpus on anatomical entities.	687	2023-11-29	Uploading
funRiceGenes-all			1.51 K	2023-11-29	Developing
funRiceGenes-exact			841	2023-11-28	Developing

Automatic annotators

Name	Description

1-2 / 2
PTO-all
PTO-exact

Editors

none