The Proposition Bank: An Annotated Corpus of Semantic Roles
Martha Palmer,Daniel Gildea,Paul Kingsbury
DOI: https://doi.org/10.1162/0891201053630264
IF: 7.778
2005-03-01
Computational Linguistics
Abstract:The Proposition Bank project takes a practical approach to semantic representation, adding a layer of predicate-argument information, or semantic role labels, to the syntactic structures of the Penn Treebank. The resulting resource can be thought of as shallow, in that it does not represent coreference, quantification, and many other higher-order phenomena, but also broad, in that it covers every instance of every verb in the corpus and allows representative statistics to be calculated. We discuss the criteria used to define the sets of semantic roles used in the annotation process and to analyze the frequency of syntactic/semantic alternations in the corpus. We describe an automatic system for semantic role tagging trained on the corpus and discuss the effect on its performance of various types of information, including a comparison of full syntactic parsing with a flat representation and the contribution of the empty “trace” categories of the treebank.
computer science, artificial intelligence, interdisciplinary applications,linguistics