PubMed_Structured_Abstracts | | Sections (zones) as retrieved from PubMed. | 131 K | | zebet | 2023-11-28 | Released | |
pubmed-sentences-benchmark | | A benchmark data for text segmentation into sentences.
The source of annotation is the GENIA treebank v1.0.
Following is the process taken.
began with the GENIA treebank v1.0.
sentence annotations were extracted and converted to PubAnnotation JSON.
uploaded. 12 abstracts met alignment failure.
among the 12 failure cases, 4 had a dot('.') character where there should be colon (':'). They were manually fixed then successfully uploaded: 7903907, 8053950, 8508358, 9415639.
among the 12 failed abstracts, 8 were "250 word truncation" cases. They were manually fixed and successfully uploaded. During the fixing, manual annotations were added for the missing pieces of text.
30 abstracts had extra text in the end, indicating copyright statement, e.g., "Copyright 1998 Academic Press." They were annotated as a sentence in GTB. However, the text did not exist anymore in PubMed. Therefore, the extra texts were removed, together with the sentence annotation to them.
| 18.4 K | GENIA project | Jin-Dong Kim | 2023-11-28 | Released | |
PubmedHPO | | Human phenotype annotation to PubMed abstracts, based on the HPO ontology | 12.4 M | Tudor Groza | tudor | 2023-11-24 | Beta | |
PubMed-German-test | | A collection of PubMed abstracts which are written in German | 0 | | Jin-Dong Kim | 2023-11-24 | Developing | |
PubMed-French-test | | A collection of PubMed abstract written in French | 0 | | Jin-Dong Kim | 2023-11-29 | Developing | |
pubmed-enju-pas | | Annotating PubMed abstracts for predicate-argument structure (PAS). Enju 2.4.2 is used to automatically compute PAS. | 19.1 M | Enju | Jin-Dong Kim | 2023-11-24 | Developing | |
PubMed_ArguminSci | | Predictions for PubMed automatically extracted with the ArguminSci tool (https://github.com/anlausch/ArguminSci). | 777 K | | zebet | 2023-11-24 | Released | |
PubMed-2017 | | abstracts published in 2017. | 0 | | Jin-Dong Kim | 2023-11-24 | Developing | |
pubmed-2016 | | abstracts published in 2016 | 0 | | Jin-Dong Kim | 2023-11-28 | | |
PubMed-2000 | | abstracts published in 2000. | 0 | | Jin-Dong Kim | 2023-11-29 | Developing | |
PubCasesORDO | | ORDO annotation in PubCases | 865 K | | Toyofumi Fujiwara | 2023-11-24 | Beta | |
PubCasesHPO | | HPO annotation in PubCases | 3.18 M | | Toyofumi Fujiwara | 2023-11-24 | Beta | |
PubCasesCollection | | abstracts in PubCases | 0 | | Jin-Dong Kim | 2023-11-29 | | |
PT_NER_NEL_pruas | | | 334 | Pedro Ruas | pruas_18 | 2023-11-30 | Uploading | |
PT_NER_NEL_mabarros | | | 328 | | mabarros | 2023-11-30 | Developing | |
PT_NER_NEL_Diana | | | 318 | | dpavot | 2023-11-24 | Developing | |
PT_NER_NEL_CONSENSUS | | | 354 | | dpavot | 2023-11-27 | | |
PT_NER_NEL | | Annotations in Portuguese COVID-19 related abstracts from MeSH terminology | 245 | LASIGE-DeST | pruas_18 | 2023-11-29 | Developing | |
proj_h_1 | | | 6.7 K | | | 2023-11-24 | | |
preeclampsia_genes | | | 17.8 K | | Jin-Dong Kim | 2023-11-29 | Developing | |