Sentiment classification of online {C}antonese reviews by supervised machine learning approaches

Ziqiong Zhang,Qiang Ye,Yijun Li,Rob Law
DOI: https://doi.org/10.1504/IJWET.2009.032254
2009-01-01
International Journal of Web Engineering and Technology
Abstract:Cantonese is an important Chinese dialect spoken in some regions of Southern China. Local online users often represent their opinions and experiences with written Cantonese on the web. With two supervised machine learning approaches, this paper conducts a series of experiments to explore appropriate methods for automatic sentiment classification in the very noisy domain of online Cantonese-written reviews. Findings indicate that the support vector machine classifier based on a Mandarin Chinese word segmentation tool performs surprisingly well. The accuracy, precision and recall respectively for positive and negative reviews all reach above 85% when the training corpus contains 5,000 or more reviews.
What problem does this paper attempt to address?