Unconstrained Language Identification Using AShape Codebook
Title | Unconstrained Language Identification Using AShape Codebook |
Publication Type | Conference Papers |
Year of Publication | 2008 |
Authors | Zhu G, Yu X, Li Y, Doermann D |
Conference Name | The 11th International Conference on Frontiers in Handwritting Recognition (ICFHR 2008) |
Date Published | 2008/// |
Conference Location | Montreal, Canada |
Abstract | We propose a novel approach to language identification in document images containing handwriting and machine printed text using image descriptors constructed from a codebook of shape features. We encode local text structures using scale and rotation invariant codewords, each representing a characteristic shape feature that is generic enough to appear repeatably. We learn a concise, structurally indexed shape codebook from training data by clustering similar features and partitioning the feature space by graph cuts. Our approach is segmentation free and easily extensible. We quantitatively evaluate our approach using a large real-world document image collection, which consists of more than 1,500 documents in 8 languages (Arabic, Chinese, English, Hindi, Japanese, Korean, Russian, and Thai) and contains a complex mixture of handwritten and machine printed content. Experimental results demonstrate the robustness and flexibility of our approach, and show exceptional language identification performance that exceeds the state of art. |