Xiaoshi Zhong
Extracting Time Expressions and Named Entities with Constituent-Based Tagging Schemes
Zhong, Xiaoshi; Cambria, Erik; Hussain, Amir
Abstract
Time expressions and named entities play important roles in data mining, information retrieval, and natural language processing. However, the conventional position-based tagging schemes (e.g., the BIO and BILOU schemes) that previous research used to model time expressions and named entities suffer from the problem of inconsistent tag assignment. To overcome the problem of inconsistent tag assignment, we designed a new type of tagging schemes to model time expressions and named entities based on their constituents. Specifically, to model time expressions, we defined a constituent-based tagging scheme termed TOMN scheme with four tags, namely T, O, M, and N, indicating the defined constituents of time expressions, namely time token, modifier, numeral, and the words outside time expressions. To model named entities, we defined a constituent-based tagging scheme termed UGTO scheme with four tags, namely U, G, T, and O, indicating the defined constituents of named entities, namely uncommon word, general modifier, trigger word, and the words outside named entities. In modeling, our TOMN and UGTO schemes model time expressions and named entities under conditional random fields with minimal features according to an in-depth analysis for the characteristics of time expressions and named entities. Experiments on diverse datasets demonstrate that our proposed methods perform equally with or more effectively than representative state-of-the-art methods on both time expression extraction and named entity extraction.
Journal Article Type | Article |
---|---|
Acceptance Date | Jan 21, 2020 |
Online Publication Date | May 10, 2020 |
Publication Date | 2020-07 |
Deposit Date | Aug 14, 2020 |
Journal | Cognitive Computation |
Print ISSN | 1866-9956 |
Electronic ISSN | 1866-9964 |
Publisher | Springer |
Peer Reviewed | Peer Reviewed |
Volume | 12 |
Pages | 844-862 |
DOI | https://doi.org/10.1007/s12559-020-09714-8 |
Keywords | Inconsistent tag assignment, Position-based tagging scheme, Constituent-based tagging scheme, Named entities, Time expressions, Intrinsic characteristics |
Public URL | http://researchrepository.napier.ac.uk/Output/2665369 |
You might also like
Applications of Deep Learning and Reinforcement Learning to Biological Data
(2018)
Journal Article
Guided Policy Search for Sequential Multitask Learning
(2018)
Journal Article
Learning Latent Features With Infinite Nonnegative Binary Matrix Trifactorization
(2018)
Journal Article
Cross-modality interactive attention network for multispectral pedestrian detection
(2018)
Journal Article
Downloadable Citations
About Edinburgh Napier Research Repository
Administrator e-mail: repository@napier.ac.uk
This application uses the following open-source libraries:
SheetJS Community Edition
Apache License Version 2.0 (http://www.apache.org/licenses/)
PDF.js
Apache License Version 2.0 (http://www.apache.org/licenses/)
Font Awesome
SIL OFL 1.1 (http://scripts.sil.org/OFL)
MIT License (http://opensource.org/licenses/mit-license.html)
CC BY 3.0 ( http://creativecommons.org/licenses/by/3.0/)
Powered by Worktribe © 2024
Advanced Search