Learning variable-length representation of words. (July 2020)
- Record Type:
- Journal Article
- Title:
- Learning variable-length representation of words. (July 2020)
- Main Title:
- Learning variable-length representation of words
- Authors:
- Ganguly, Debasis
- Abstract:
- Highlights: A variable-length representation learning (embedding) of words. Allows provision for compressing the word vectors. Proposed algorithm uses a smaller number of dimensions for words with consistent contexts (words with specific meanings). Variable length embedding potentially helps removing bias (over-fitting) on certain datasets. Proposed approach outperforms fixed-length embedding, and also transformation-based approaches based on regularization and binarization, on standard word-semantics datasets. Abstract: A standard word embedding algorithm, such as 'word2vec', embeds each word as a dense vector of a preset dimensionality, the components of which are learned by maximizing the likelihood of predicting the context around it. However, as an inherent linguistic phenomenon, it is evident that there is a varying degree of difficulty in identifying words from their contexts. This suggests that a variable granularity in word vector representation may be useful to obtain sparser and more compressed word representations, requiring less storage space. To that end, in this paper, we propose a word vector training algorithm that uses a variable number of components to represent words. Given a text collection of documents, our algorithm, similar to the skip-gram approach of word2vec, learns to predict the context of a word given the current instance of a word. However, in contrast to skip-gram, which uses a static number of dimensions for each word vector, we propose toHighlights: A variable-length representation learning (embedding) of words. Allows provision for compressing the word vectors. Proposed algorithm uses a smaller number of dimensions for words with consistent contexts (words with specific meanings). Variable length embedding potentially helps removing bias (over-fitting) on certain datasets. Proposed approach outperforms fixed-length embedding, and also transformation-based approaches based on regularization and binarization, on standard word-semantics datasets. Abstract: A standard word embedding algorithm, such as 'word2vec', embeds each word as a dense vector of a preset dimensionality, the components of which are learned by maximizing the likelihood of predicting the context around it. However, as an inherent linguistic phenomenon, it is evident that there is a varying degree of difficulty in identifying words from their contexts. This suggests that a variable granularity in word vector representation may be useful to obtain sparser and more compressed word representations, requiring less storage space. To that end, in this paper, we propose a word vector training algorithm that uses a variable number of components to represent words. Given a text collection of documents, our algorithm, similar to the skip-gram approach of word2vec, learns to predict the context of a word given the current instance of a word. However, in contrast to skip-gram, which uses a static number of dimensions for each word vector, we propose to dynamically increase the dimensionality as a stochastic function of the prediction error. Our experiments with standard test collections demonstrate that our word representation method is able to achieve comparable (and sometimes even better) effectiveness than skip-gram word2vec, using a significantly smaller number of parameters (achieving compression ratio of around 65%). … (more)
- Is Part Of:
- Pattern recognition. Volume 103(2020:Jul.)
- Journal:
- Pattern recognition
- Issue:
- Volume 103(2020:Jul.)
- Issue Display:
- Volume 103 (2020)
- Year:
- 2020
- Volume:
- 103
- Issue Sort Value:
- 2020-0103-0000-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-07
- Subjects:
- Word embedding -- Compression and sparsity -- Lexical semantics
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2020.107306 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15150.xml