Compact variant‐rich customized sequence database and a fast and sensitive database search for efficient proteogenomic analyses. Issue 23 (19th November 2014)
- Record Type:
- Journal Article
- Title:
- Compact variant‐rich customized sequence database and a fast and sensitive database search for efficient proteogenomic analyses. Issue 23 (19th November 2014)
- Main Title:
- Compact variant‐rich customized sequence database and a fast and sensitive database search for efficient proteogenomic analyses
- Authors:
- Park, Heejin
Bae, Junwoo
Kim, Hyunwoo
Kim, Sangok
Kim, Hokeun
Mun, Dong‐Gi
Joh, Yoonsung
Lee, Wonyeop
Chae, Sehyun
Lee, Sanghyuk
Kim, Hark Kyun
Hwang, Daehee
Lee, Sang‐Won
Paek, Eunok
Pandey, Akhilesh
Pevzner, Pavel A. - Abstract:
- <abstract abstract-type="main"> <title> <x xml:space="preserve">Abstract</x> </title> <p>In proteogenomic analysis, construction of a compact, customized database from mRNA‐seq data and a sensitive search of both reference and customized databases are essential to accurately determine protein abundances and structural variations at the protein level. However, these tasks have not been systematically explored, but rather performed in an <italic>ad‐hoc</italic> fashion. Here, we present an effective method for constructing a compact database containing comprehensive sequences of sample‐specific variants—single nucleotide variants, insertions/deletions, and stop‐codon mutations derived from Exome‐seq and RNA‐seq data. It, however, occupies less space by storing variant peptides, not variant proteins. We also present an efficient search method for both customized and reference databases. The separate searches of the two databases increase the search time, and a unified search is less sensitive to identify variant peptides due to the smaller size of the customized database, compared to the reference database, in the target‐decoy setting. Our method searches the unified database once, but performs target‐decoy validations separately. Experimental results show that our approach is as fast as the unified search and as sensitive as the separate searches. Our customized database includes mutation information in the headers of variant peptides, thereby facilitating the inspection of<abstract abstract-type="main"> <title> <x xml:space="preserve">Abstract</x> </title> <p>In proteogenomic analysis, construction of a compact, customized database from mRNA‐seq data and a sensitive search of both reference and customized databases are essential to accurately determine protein abundances and structural variations at the protein level. However, these tasks have not been systematically explored, but rather performed in an <italic>ad‐hoc</italic> fashion. Here, we present an effective method for constructing a compact database containing comprehensive sequences of sample‐specific variants—single nucleotide variants, insertions/deletions, and stop‐codon mutations derived from Exome‐seq and RNA‐seq data. It, however, occupies less space by storing variant peptides, not variant proteins. We also present an efficient search method for both customized and reference databases. The separate searches of the two databases increase the search time, and a unified search is less sensitive to identify variant peptides due to the smaller size of the customized database, compared to the reference database, in the target‐decoy setting. Our method searches the unified database once, but performs target‐decoy validations separately. Experimental results show that our approach is as fast as the unified search and as sensitive as the separate searches. Our customized database includes mutation information in the headers of variant peptides, thereby facilitating the inspection of peptide‐spectrum matches.</p> </abstract> … (more)
- Is Part Of:
- Proteomics. Volume 14:Issue 23/24(2014)
- Journal:
- Proteomics
- Issue:
- Volume 14:Issue 23/24(2014)
- Issue Display:
- Volume 14, Issue 23/24 (2014)
- Year:
- 2014
- Volume:
- 14
- Issue:
- 23/24
- Issue Sort Value:
- 2014-0014-NaN-0000
- Page Start:
- 2742
- Page End:
- 2749
- Publication Date:
- 2014-11-19
- Subjects:
- Proteins -- Separation -- Periodicals
Bioinformatics -- Periodicals
Proteomics -- Periodicals
Genomes -- Periodicals
Molecular genetics -- Periodicals
572.605 - Journal URLs:
- http://onlinelibrary.wiley.com/journal/10.1002/(ISSN)1615-9861 ↗
http://onlinelibrary.wiley.com/ ↗ - DOI:
- 10.1002/pmic.201400225 ↗
- Languages:
- English
- ISSNs:
- 1615-9853
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 6936.178000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 3582.xml