Abstract
In a system with a large database, there always has been a problem that names may not be spelled well or might not be spelled in a way that one expected. So, data in the database gets degraded. In this case it is required to search the duplicates and merge them in the single entity. In doing so, one problem is that the way in which the strings would be compared. In such cases rather than looking for exact match, approximate string matching would be appreciable. One of the string matching techniques is Phonetic matching which is used to compare the name based on the pronunciation of the words. The similar sounding words could be retrieved from the large database using different phonetic matching algorithm and best known algorithm is Soundex algorithm. Phonetic matching is needed when many people from different culture come together. They either speak with different pronunciation or their writing habits are different. This scenario is very common in India, as we have many different languages like Hindi, Gujarati, Marathi, Tamil etc. In this research work Soundex algorithm is used for Hindi and Gujarati language and applied on the names along with their variations in order to retrieve the output with minimum false hits.
Cite
CITATION STYLE
Shah, R., & Kumar Singh, D. (2014). Improvement of Soundex Algorithm for Indian Language Based on Phonetic Matching. International Journal of Computer Science, Engineering and Applications, 4(3), 31–39. https://doi.org/10.5121/ijcsea.2014.4303
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.