OpenNLP chunker Biomed data #10

khituras · 2017-10-10T12:18:28Z

Our chunker data is derived from the GENIA treebank corpus. However, this corpus has complete nested constituencies instead of just chunks. So we use an algorithm to create the chunks out of the treebank. For this there are currently two algorithms in the jcore-base version of the opennlp chunker. I think the newer one works better than the old one but it is still not perfect.
Now I found these data in our internal file system: /archives/alumni_homes/tomanek/coling/corpora/Genia/chunks/genia_new.chunks.gz

This appears to be the GENIA conversion used originally within the JULIE Lab. We should do crossevaluations on both corpora to see if there tagging differences and also just a plain comparison. Perhaps the old data is better.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

OpenNLP chunker Biomed data #10

OpenNLP chunker Biomed data #10

khituras commented Oct 10, 2017

OpenNLP chunker Biomed data #10

OpenNLP chunker Biomed data #10

Comments

khituras commented Oct 10, 2017