Deep CNNs Along the Time Axis With Intermap Pooling for Robustness to Spectral Variations

Cited 5 time in webofscience Cited 0 time in scopus
  • Hit : 714
  • Download : 0
DC FieldValueLanguage
dc.contributor.authorHwaran, Leeko
dc.contributor.authorKim, Geonminko
dc.contributor.authorKim, Ho-Gyeongko
dc.contributor.authorOh, Sang-Hoonko
dc.contributor.authorLee, Soo-Youngko
dc.date.accessioned2016-12-01T04:53:28Z-
dc.date.available2016-12-01T04:53:28Z-
dc.date.created2016-11-17-
dc.date.created2016-11-17-
dc.date.created2016-11-17-
dc.date.issued2016-10-
dc.identifier.citationIEEE SIGNAL PROCESSING LETTERS, v.23, no.10, pp.1310 - 1314-
dc.identifier.issn1070-9908-
dc.identifier.urihttp://hdl.handle.net/10203/214457-
dc.description.abstractConvolutional neural networks (CNNs) with convolutional and pooling operations along the frequency axis have been proposed to attain invariance to frequency shifts of features. However, this is inappropriate with regard to the fact that acoustic features vary in frequency. In this paper, we contend that convolution along the time axis is more effective. We also propose the addition of an intermap pooling (IMP) layer to deep CNNs. In this layer, filters in each group extract common but spectrally variant features, then the layer pools the feature maps of each group. As a result, the proposed IMP CNN can achieve insensitivity to spectral variations characteristic of different speakers and utterances. The effectiveness of the IMP CNN architecture is demonstrated on several LVCSR tasks. Even without speaker adaptation techniques, the architecture achieved a WER of 12.7% on the SWB part of the Hub5'2000 evaluation test set, which is competitive with other state-of-the-art methods.-
dc.languageEnglish-
dc.publisherIEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC-
dc.titleDeep CNNs Along the Time Axis With Intermap Pooling for Robustness to Spectral Variations-
dc.typeArticle-
dc.identifier.wosid000395029000003-
dc.identifier.scopusid2-s2.0-84986188619-
dc.type.rimsART-
dc.citation.volume23-
dc.citation.issue10-
dc.citation.beginningpage1310-
dc.citation.endingpage1314-
dc.citation.publicationnameIEEE SIGNAL PROCESSING LETTERS-
dc.identifier.doi10.1109/LSP.2016.2589962-
dc.contributor.localauthorLee, Soo-Young-
dc.contributor.nonIdAuthorOh, Sang-Hoon-
dc.description.isOpenAccessN-
dc.type.journalArticleArticle-
dc.subject.keywordAuthorAcoustic modeling-
dc.subject.keywordAuthorconvolutional neural networks (CNNs)-
dc.subject.keywordAuthorintermap pooling (IMP) layer-
dc.subject.keywordPlusCONVOLUTIONAL NEURAL-NETWORKS-
dc.subject.keywordPlusSPEECH RECOGNITION-
dc.subject.keywordPlusINVARIANT FEATURES-
dc.subject.keywordPlusAUDITORY-CORTEX-
dc.subject.keywordPlusMAPS-
Appears in Collection
EE-Journal Papers(저널논문)
Files in This Item
There are no files associated with this item.
This item is cited by other documents in WoS
⊙ Detail Information in WoSⓡ Click to see webofscience_button
⊙ Cited 5 items in WoS Click to see citing articles in records_button

qr_code

  • mendeley

    citeulike


rss_1.0 rss_2.0 atom_1.0