ClusType
Citations Over TimeTop 1% of 2015 papers
Abstract
Entity recognition is an important but challenging research problem. In reality, many text collections are from specific, dynamic, or emerging domains, which poses significant new challenges for entity recognition with increase in name ambiguity and context sparsity, requiring entity detection without domain restriction. In this paper, we investigate entity recognition (ER) with distant-supervision and propose a novel relation phrase-based ER framework, called ClusType, that runs data-driven phrase mining to generate entity mention candidates and relation phrases, and enforces the principle that relation phrases should be softly clustered when propagating type information between their argument entities. Then we predict the type of each entity mention based on the type signatures of its co-occurring relation phrases and the type indicators of its surface name, as computed over the corpus. Specifically, we formulate a joint optimization problem for two tasks, type propagation with relation phrases and multi-view relation phrase clustering. Our experiments on multiple genres-news, Yelp reviews and tweets-demonstrate the effectiveness and robustness of ClusType, with an average of 37% improvement in F1 score over the best compared method.
Related Papers
- → Types of Ambiguity(2006)4 cited
- → Finding and Using Ambiguity to Search for Innovation Opportunities(2018)2 cited
- → Decision‐making for others: Ambiguity attitudes(2024)1 cited
- On the Treatment of Intentional Ambiguity and Unintentional Ambiguity in English(2005)
- Study on the Pragmatic Value of English Intentional Ambiguity(2014)