Graph based feature selection investigating boundary region of rough set for language identification. (15th November 2020)
- Record Type:
- Journal Article
- Title:
- Graph based feature selection investigating boundary region of rough set for language identification. (15th November 2020)
- Main Title:
- Graph based feature selection investigating boundary region of rough set for language identification
- Authors:
- Yasmin, Ghazaala
Das, Asit Kumar
Nayak, Janmenjoy
Pelusi, Danilo
Ding, Weiping - Abstract:
- Highlights: Suitable features are extracted based on the nature of speech. Positive region of RST can measure the similarity if the given information is precise. Boundary region of RST is also explored as audio speech dataset may contain uncertainty. A MST is generated and a feature selection method is devised to select the relevant features. The method is compared with popular diverse feature selection algorithms. Abstract: Language can be chosen to be a species where maximum information can be extracted. In the world, there are many countries, some of which are of numerous types and flavours of regions based on their languages. The challenge is to make the spoken language recognition to be automated through machine learning. The proposed language identification system extracts various features from speech of different languages and constructs a complete weighted graph with extracted features as nodes and similarity among the features as weights of the edges. Similarity values are computed using the concepts of positive region and boundary region of rough set theory and a graph based feature selection algorithm is devised to select only the minimal subset of features relevant to language identification. It is observed that, investigating the boundary region together with the positive region, more valuable information is extracted which helps in selection of more relevant features for language identification. The constructed complete weighted graph is made sparse using GiniHighlights: Suitable features are extracted based on the nature of speech. Positive region of RST can measure the similarity if the given information is precise. Boundary region of RST is also explored as audio speech dataset may contain uncertainty. A MST is generated and a feature selection method is devised to select the relevant features. The method is compared with popular diverse feature selection algorithms. Abstract: Language can be chosen to be a species where maximum information can be extracted. In the world, there are many countries, some of which are of numerous types and flavours of regions based on their languages. The challenge is to make the spoken language recognition to be automated through machine learning. The proposed language identification system extracts various features from speech of different languages and constructs a complete weighted graph with extracted features as nodes and similarity among the features as weights of the edges. Similarity values are computed using the concepts of positive region and boundary region of rough set theory and a graph based feature selection algorithm is devised to select only the minimal subset of features relevant to language identification. It is observed that, investigating the boundary region together with the positive region, more valuable information is extracted which helps in selection of more relevant features for language identification. The constructed complete weighted graph is made sparse using Gini index based sparsity measure. As a result, the graph contains only the edges whose terminal nodes are highly similar. Next, a maximal spanning tree of the graph is generated using Prim's algorithm. This tree is a basic structure that provides the maximal similarity among the nodes in the graph. Finally, score of each node is computed based on weights of the edges in the tree and a node with the highest score is selected and removed from the spanning tree. This process of selection and removal of nodes is continued until the graph becomes null. The resultant set of selected nodes is considered as the important feature subset of the audio speeches used for language identification. Experimental results show the effectiveness of the proposed rough set theory based feature selection method. The results also demonstrate the usefulness of investigation of boundary region of rough sets. … (more)
- Is Part Of:
- Expert systems with applications. Volume 158(2020)
- Journal:
- Expert systems with applications
- Issue:
- Volume 158(2020)
- Issue Display:
- Volume 158, Issue 2020 (2020)
- Year:
- 2020
- Volume:
- 158
- Issue:
- 2020
- Issue Sort Value:
- 2020-0158-2020-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-11-15
- Subjects:
- Language identification -- Feature selection -- Relative indiscernibility relation -- Attribute dependency -- Boundary region exploration
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2020.113575 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 14015.xml