1,065 40 98MB
English Pages 897 [909] Year 2021
![Intelligent Systems and Applications: Proceedings of the 2021 Intelligent Systems Conference (IntelliSys) Volume 1: 294 (Lecture Notes in Networks and Systems) [1st ed. 2022]
3030821927, 9783030821920](https://ebin.pub/img/200x200/intelligent-systems-and-applications-proceedings-of-the-2021-intelligent-systems-conference-intellisys-volume-1-294-lecture-notes-in-networks-and-systems-1st-ed-2022-3030821927-9783030821920.jpg)
Lecture Notes in Networks and Systems 294
Kohei Arai Editor
Intelligent Systems and Applications Proceedings of the 2021 Intelligent Systems Conference (IntelliSys) Volume 1
Lecture Notes in Networks and Systems Volume 294
Series Editor Janusz Kacprzyk, Systems Research Institute, Polish Academy of Sciences, Warsaw, Poland Advisory Editors Fernando Gomide, Department of Computer Engineering and Automation—DCA, School of Electrical and Computer Engineering—FEEC, University of Campinas— UNICAMP, São Paulo, Brazil Okyay Kaynak, Department of Electrical and Electronic Engineering, Bogazici University, Istanbul, Turkey Derong Liu, Department of Electrical and Computer Engineering, University of Illinois at Chicago, Chicago, USA; Institute of Automation, Chinese Academy of Sciences, Beijing, China Witold Pedrycz, Department of Electrical and Computer Engineering, University of Alberta, Alberta, Canada; Systems Research Institute, Polish Academy of Sciences, Warsaw, Poland Marios M. Polycarpou, Department of Electrical and Computer Engineering, KIOS Research Center for Intelligent Systems and Networks, University of Cyprus, Nicosia, Cyprus Imre J. Rudas, Óbuda University, Budapest, Hungary Jun Wang, Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong
The series “Lecture Notes in Networks and Systems” publishes the latest developments in Networks and Systems—quickly, informally and with high quality. Original research reported in proceedings and post-proceedings represents the core of LNNS. Volumes published in LNNS embrace all aspects and subfields of, as well as new challenges in, Networks and Systems. The series contains proceedings and edited volumes in systems and networks, spanning the areas of Cyber-Physical Systems, Autonomous Systems, Sensor Networks, Control Systems, Energy Systems, Automotive Systems, Biological Systems, Vehicular Networking and Connected Vehicles, Aerospace Systems, Automation, Manufacturing, Smart Grids, Nonlinear Systems, Power Systems, Robotics, Social Systems, Economic Systems and other. Of particular value to both the contributors and the readership are the short publication timeframe and the world-wide distribution and exposure which enable both a wide and rapid dissemination of research output. The series covers the theory, applications, and perspectives on the state of the art and future developments relevant to systems and networks, decision making, control, complex processes and related areas, as embedded in the fields of interdisciplinary and applied sciences, engineering, computer science, physics, economics, social, and life sciences, as well as the paradigms and methodologies behind them. Indexed by SCOPUS, INSPEC, WTI Frankfurt eG, zbMATH, SCImago. All books published in the series are submitted for consideration in Web of Science.
More information about this series at http://www.springer.com/series/15179
Kohei Arai Editor
Intelligent Systems and Applications Proceedings of the 2021 Intelligent Systems Conference (IntelliSys) Volume 1
123
Editor Kohei Arai Faculty of Science and Engineering Saga University Saga, Japan
ISSN 2367-3370 ISSN 2367-3389 (electronic) Lecture Notes in Networks and Systems ISBN 978-3-030-82192-0 ISBN 978-3-030-82193-7 (eBook) https://doi.org/10.1007/978-3-030-82193-7 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer Nature Switzerland AG 2022 This work is subject to copyright. All rights are solely and exclusively licensed by the Publisher, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed. The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective laws and regulations and therefore free for general use. The publisher, the authors and the editors are safe to assume that the advice and information in this book are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the editors give a warranty, expressed or implied, with respect to the material contained herein or for any errors or omissions that may have been made. The publisher remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. This Springer imprint is published by the registered company Springer Nature Switzerland AG The registered company address is: Gewerbestrasse 11, 6330 Cham, Switzerland
Editor’s Preface
We are very pleased to introduce the Proceedings of Intelligent Systems Conference (IntelliSys) 2021 which was held on September 2 and 3, 2021. The entire world was affected by COVID-19 and our conference was not an exception. To provide a safe conference environment, IntelliSys 2021, which was planned to be held in Amsterdam, Netherlands, was changed to be held fully online. The Intelligent Systems Conference is a prestigious annual conference on areas of intelligent systems and artificial intelligence and their applications to the real world. This conference not only presented the state-of-the-art methods and valuable experience, but also provided the audience with a vision of further development in the fields. One of the meaningful and valuable dimensions of this conference is the way it brings together researchers, scientists, academics, and engineers in the field from different countries. The aim was to further increase the body of knowledge in this specific area by providing a forum to exchange ideas and discuss results, and to build international links. The Program Committee of IntelliSys 2021 represented 25 countries, and authors from 50+ countries submitted a total of 496 papers. This certainly attests to the widespread, international importance of the theme of the conference. Each paper was reviewed on the basis of originality, novelty, and rigorousness. After the reviews, 195 were accepted for presentation, out of which 180 (including 7 posters) papers are finally being published in the proceedings. These papers provide good examples of current research on relevant topics, covering deep learning, data mining, data processing, human–computer interactions, natural language processing, expert systems, robotics, ambient intelligence to name a few. The conference would truly not function without the contributions and support received from authors, participants, keynote speakers, program committee members, session chairs, organizing committee members, steering committee members, and others in their various roles. Their valuable support, suggestions, dedicated commitment, and hard work have made IntelliSys 2021 successful. We warmly thank and greatly appreciate the contributions, and we kindly invite all to continue to contribute to future IntelliSys. v
vi
Editor’s Preface
We believe this event will certainly help further disseminate new ideas and inspire more international collaborations. Kind Regards, Kohei Arai
Contents
Late Fusion of Convolutional Neural Network with Wavelet-Based Ensemble Classifier for Acoustic Scene Classification . . . . . . . . . . . . . . . Cheng Siong Chin and Jianhua Zhang Deep Learning and Social Media for Managing Disaster: Survey . . . . . Zair Bouzidi, Abdelmalek Boudries, and Mourad Amad
1 12
A Framework for Adaptive Mobile Ecological Momentary Assessments Using Reinforcement Learning . . . . . . . . . . . . . . . . . . . . . . Lihua Cai, Laura E. Barnes, and Mehdi Boukhechba
31
Reputation Analysis Based on Weakly-Supervised Bi-LSTM-Attention Network . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Kun Xiang and Akihiro Fujii
51
Multi-GPU-based Convolutional Neural Networks Training for Text Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Imen Ferjani, Minyar Sassi Hidri, and Ali Frihida
72
Performance Analysis of Data-Driven Techniques for Solving Inverse Kinematics Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Vijay Bhaskar Semwal and Yash Gupta
85
Machine Learning Based H2 Norm Minimization for Maglev Vibration Isolation Platform . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 Ahmet Fevzi Bozkurt, Barış Can Yalçın, and Kadir Erkan A Vision Based Deep Reinforcement Learning Algorithm for UAV Obstacle Avoidance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 Jeremy Roghair, Amir Niaraki, Kyungtae Ko, and Ali Jannesari Detecting and Fixing Nonidiomatic Snippets in Python Source Code with Deep Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 Balázs Szalontai, András Vadász, Zsolt Richárd Borsi, Teréz A. Várkonyi, Balázs Pintér, and Tibor Gregorics vii
viii
Contents
BreakingBED: Breaking Binary and Efficient Deep Neural Networks by Adversarial Attacks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 Manoj-Rohit Vemparala, Alexander Frickenstein, Nael Fasfous, Lukas Frickenstein, Qi Zhao, Sabine Kuhn, Daniel Ehrhardt, Yuankai Wu, Christian Unger, Naveen-Shankar Nagaraja, and Walter Stechele Parallel Dilated CNN for Detecting and Classifying Defects in Surface Steel Strips in Real-Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168 Khaled R. Ahmed Selective Information Control and Network Compression in Multi-layered Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184 Ryotaro Kamimura DAC–Deep Autoencoder-Based Clustering: A General Deep Learning Framework of Representation Learning . . . . . . . . . . . . . . . . . 205 Si Lu and Ruisi Li Enhancing LSTM Models with Self-attention and Stateful Training . . . 217 Alexander Katrompas and Vangelis Metsis Domain Generalization Using Ensemble Learning . . . . . . . . . . . . . . . . . 236 Yusuf Mesbah, Youssef Youssry Ibrahim, and Adil Mehood Khan Research on Text Classification Modeling Strategy Based on Pre-trained Language Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248 Yiou Lin, Hang Lei, Xiaoyu Li, and Yu Deng Discovering Nonlinear Dynamics Through Scientific Machine Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261 Lei Huang, Daniel Vrinceanu, Yunjiao Wang, Nalinda Kulathunga, and Nishath Ranasinghe Tensor Data Scattering and the Impossibility of Slicing Theorem . . . . . 280 Wuming Pan Scope and Sense of Explainability for AI-Systems . . . . . . . . . . . . . . . . . 291 A.-M. Leventi-Peetz, T. Östreich, W. Lennartz, and K. Weber Use Case Prediction Using Deep Learning . . . . . . . . . . . . . . . . . . . . . . . 309 Tinashe Wamambo, Cristina Luca, Arooj Fatima, and Mahdi Maktab-Dar-Oghaz VAMDLE: Visitor and Asset Management Using Deep Learning and ElasticSearch . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318 Viswanathsingh Seenundun, Balkrishansingh Purmah, and Zahra Mungloo-Dilmohamud Wind Speed Time Series Prediction with Deep Learning and Data Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 330 Anibal Flores, Hugo Tito-Chura, and Victor Yana-Mamani
Contents
ix
Evaluation for Angular Distortion of Welding Plate . . . . . . . . . . . . . . . 344 Shigeru Kato, Shunsaku Kume, Takanori Hino, Fujioka Shota, Tomomichi Kagawa, Hironori Kumeno, and Hajime Nobuhara A Framework for Testing and Evaluation of Operational Performance of Multi-UAV Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355 Mrinmoy Sarkar, Xuyang Yan, Shamila Nateghi, Bruce J. Holmes, Kyriakos G. Vamvoudakis, and Abdollah Homaifar Addressing Consumer Demands: A Manufacturing Collaboration Process Using Blockchain for Knowledge Representation . . . . . . . . . . . 375 Ricardo Barbosa, Ricardo Santos, and Paulo Novais Cellular Formation Maintenance and Collision Avoidance Using Centroid-Based Point Set Registration in a Swarm of Drones . . . . . . . . 391 Jawad N. Yasin, Huma Mahboob, Mohammad-Hashem Haghbayan, Muhammad Mehboob Yasin, and Juha Plosila The Simulation with New Opinion Dynamics Using Five Adopter Categories . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409 Makoto Fujii and Akira Ishii Intrinsic Rewards for Reinforcement Learning Within Complex 2D Environments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 425 Nathaniel Grabaskas and Zhizhen Wang Analysis of Divided Society at the Standpoint of In-Group and Out-Group Using Opinion Dynamics . . . . . . . . . . . . . . . . . . . . . . . 438 Nozomi Okano and Akira Ishii Simulation of Intragroup Alignment Using a New Model of Opinion Dynamics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453 Nozomi Okano, Hitoshi Yamamoto, Masaru Nishikawa, and Akira Ishii Random Forest Classification with MapReduce in Holonic Multiagent Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 464 Michéle Cullinan and Duncan Coulter Monitoring Goal Driven Autonomy Agent’s Expectations Generated from Durative Effects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484 Noah Reifsnyder and Hector Munoz-Avila Sublinear Regret with Barzilai-Borwein Step Sizes . . . . . . . . . . . . . . . . 499 Iyanuoluwa Emiola Fluid Dynamics of a Pandemic in a Spatial Social Network: A Reflective Measure of the Spreading . . . . . . . . . . . . . . . . . . . . . . . . . . 513 Saad Alqithami
x
Contents
Affective Story-Morphing: Manipulating Shelley’s Frankenstein under Program Control using Emotionally Intelligent Agents . . . . . . . . 526 Clark Elliott Dynamic Strategies and Opponent Hands Estimation for Reinforcement Learning in Gin Rummy Game . . . . . . . . . . . . . . . . 543 Yuexing Hao and Mark Vaysiberg Wireless Sensor Network Smart Environment for Precision Agriculture: An Agent-Based Architecture . . . . . . . . . . . . . . . . . . . . . . . 556 AbdulMutalib Wahaishi and Raafat Aburukba Autonomy Reconsidered: Towards Developing Multi-agent Systems . . . 573 Michael A. Goodrich, Julie A. Adams, and Matthias Scheutz A Real-Time Intelligent Intra-vehicular Temperature Control Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 593 Daniel Jacuinde-Alvarez, James Dols, and Shahab Tayeb Intelligent Control of a Semi-autonomous Assistive Vehicle . . . . . . . . . . 613 David Sanders, Giles Tewkesbury, Malik Haddad, Ya Huang, and Boriana Vatchova One Shot Learning Approach to Identify Drivers . . . . . . . . . . . . . . . . . 622 Malik Haddad, David Sanders, Martin Langner, and Giles Tewkesbury Facial Recognition Software for Identification of Powered Wheelchair Users . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 630 Giles Tewkesbury, Samuel Lifton, Malik Haddad, David Sanders, and Alex Gegov Intelligent User Interface to Control a Powered Wheelchair Using Infrared Sensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 640 Malik Haddad, David Sanders, Giles Tewkesbury, Martin Langner, and Sarinova Simandjuntak A Classification Based Ensemble Pruning Framework with Multi-metric Consideration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 650 Ya-Lin Zhang, Qitao Shi, Meng Li, Xinxing Yang, Longfei Li, and Jun Zhou Customs Risk Assessment Based on Unsupervised Anomaly Detection Using Autoencoders . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668 Dion T. Oosterman, Wouter H. Langenkamp, and Ellen L. van Bergen Best Next Preference Prediction Based on LSTM and Multi-level Interactions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 682 Ivett Fuentes, Gonzalo Nápoles, Leticia Arco, and Koen Vanhoof
Contents
xi
Achieving Trust in Future Human Interactions with Omnipresent AI: Some Postulates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 700 Peer Sathikh, Zong Rui Dexter Fang, and Guan Yi Tan A Decentralized Explanatory System for Intelligent Cyber-Physical Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 719 Étienne Houzé, Jean-Louis Dessalles, Ada Diaconescu, David Menga, and Mathieu Schumann Construction Control Organization with Use of Computer and Information Technologies in Context of Sustainable Development Providing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 739 Zalina Ruslanovna Tuskaeva and Zaurbek Valerievich Albegov Computational Rational Engineering and Development: Synergies and Opportunities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 744 Ramses Sala QPSetter: An Artificial Intelligence-Based Web Enabled, Personalized Service Application for Educators . . . . . . . . . . . . . . . . . . . 764 Mohammad Ali Kadampur and Sulaiman Al Riyaee Is It Possible to Recognize a Philosophical Zombie and How to Do It . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778 R. V. Dushkin Dynamic Analysis of Bitcoin Fluctuations by Means of a Fractal Predictor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 791 Jesús Jaime Moreno Escobar, Oswaldo Morales Matamoros, Ana Lilia Coria Páez, and Ricardo Tejeida Padilla Are Human Drivers a Liability or an Asset? . . . . . . . . . . . . . . . . . . . . . 805 David Sanders, Malik Haddad, Giles Tewkesbury, Alex Gegov, and Mo Adda Negative Emotions Induced by Non-verbal Video Clips . . . . . . . . . . . . . 817 Flavia De Simone, Simona Collina, and Manuela Nuzzo Automatic Recognition of Key Modulations in Symbolic Musical Pieces Using Information Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 823 Michele Della Ventura Increasing Robustness for Machine Learning Services in Challenging Environments: Limited Resources and No Label Feedback . . . . . . . . . . 837 Lucas Baier, Niklas Kühl, and Jörg Schmitt Development Support for Intelligent Systems: Test, Evaluation, and Analysis of Microservices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 857 Charline von Perbandt, Matthias Tyca, Arne Koschel, and Irina Astrova
xii
Contents
An Analysis with Dynamics Between Human Motivation and Messaging on Social Networking Services . . . . . . . . . . . . . . . . . . . . 876 Hidehiro Matsumoto and Akira Ishii Author Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 895
Late Fusion of Convolutional Neural Network with Wavelet-Based Ensemble Classifier for Acoustic Scene Classification Cheng Siong Chin1(B) and Jianhua Zhang2 1 Faculty of Science, Agriculture, and Engineering, Newcastle University Singapore, Singapore 599493, Singapore [email protected] 2 School of Information and Control Engineering, Qingdao University of Technology, Qingdao 266525, Shandong, China
Abstract. Log-Mel spectrogram for the convolutional neural network (CNN) and wavelet time scattering for Ensemble of subspace discriminant classifiers is used for classifying acoustic scenes with human speech. The Tampere University of Technology (TUT) Acoustic Scenes dataset is used to demonstrate the feasibility of the proposed model. Comparisons are performed with the baseline model in the TUT 2017 dataset used for Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge-Task 1. The fused model shows good acoustic classification accuracy of 79.43%. The proposed late fusion of multi-model using CNN and ensemble classifiers exhibits 18.4% higher accuracy than the baseline model with just CNN. Keywords: Acoustic scene classification · Time scattering · Acoustic classification accuracy · Convolutional Neural Network · Wavelet multi-model late fusion system
1 Introduction Acoustic scene classification (ASC) [1–4] classifies audio signals into a pre-selected list of scene types such as car parks, parks, meeting rooms, etc. The problem can resemble speech recognition. The main difference is the target classes are more diversified. They are various applications of ASC. For example, it can be used for acoustic event recognition using the mobile device that detects an individual is having a meeting. It would trigger the device into silent mode automatically. ASC has been used in robots [5, 6], mobile devices [7–9], traffic [10, 11], and medical systems [12, 13]. One of the standard scientific challenges in ASC is Detection and Classification of Acoustic Scenes and Events (DCASE) Challenge. The primary scope of ASC involved obtaining the best acoustic classification accuracy in assigning audio recordings to a specific recorded environment.
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2022 K. Arai (Ed.): IntelliSys 2021, LNNS 294, pp. 1–11, 2022. https://doi.org/10.1007/978-3-030-82193-7_1
2
C. S. Chin and J. Zhang
Many ASC used Convolutional Neural Network (CNN) [14], Recurrent Neural Networks (RNN) [15], Support Vector Machines (SVM) [16], Gaussian Mixture Models [17], and Multilayer Perceptron [18, 19]. Recurrent network architectures such as convolutional (CRNN), Bi-Long Short Term Memory (LSTM), and LSTM [20] were also used. However, LSTM has inherently gradient vanishing and exploding problems. As observed in DCASE Challenge, the best-performing systems used CNN. An ensemble of neural networks [21] and ensemble classifiers [22] were used. The former approach using CNN has outperformed other ASC task approaches [23–25]. The latter has also demonstrated good acoustic classification accuracy with shorter computation time than CNN. To improve the generalization, Mel-frequency cepstral coefficients (MFCCs) [26] and other signal representations such as Constant Q Transform (CQT) [27] and wavelet time scattering [28] to extract the acoustic features of the raw data. Multiple spectrograms [29, 30] such as MFCC, short-time Fourier transform (STFT), and CQT were also utilized to increase the number of features for training. It has shown positive results as more timefrequency characteristics could be extracted. However, the computation time increases as more features are required to be processed. In this paper, a multi-model late fusion is used for ASC. CNN seems to be a reasonable choice. They are provided a time-frequency representation of audio to capture spectro-temporal modulation patterns for identifying various acoustic scenes. The timefrequency representation is used for CNN. It relates to the width and height dimensions of the convolutional filters, respectively. To reduce the overfitting of the dataset in CNN, a Mixup algorithm [31] is used. The original and mixed datasets are combined to train a CNN [14] using the log-Mel spectrograms. It is followed by an ensemble random subspace discriminant classifier using wavelet scattering [28]. The Tampere University of Technology (TUT) 2017 dataset [32] used for DCASE2017 Challenge-Task 1 will be used for both training and evaluation.
2 Proposed Methodology There are 4680 and 1620 labeled audio files for training and evaluation, respectively. The original TUT-2017 datasets are obtained from different environmental scenes at other recording locations with some human speech recorded. There are not more than 5-min audio recordings at each site. The original recordings are split into 10s segments where each audio segment is included as sound files. The details of the outdoor (both open or enclosed areas) and indoor acoustic scenes can be seen in the TUT-2017 dataset [32]. The following 15 acoustic scenes are as follows. • • • • • • •
bus forest path home city center cafe lakeside beach (outdoor) library (indoor)
Late Fusion of CNN with Wavelet-Based Ensemble Classifier
• • • • • • • •
3
car grocery store urban park (outdoor) office: multiple persons (indoor) metro station (indoor) train (traveling, vehicle) residential area (outdoor) tram (traveling, vehicle)
The recordings are recorded from different streets, homes, and parks [32]. Sound recordings were performed via different devices at 24-bit resolution and 44100 Hz sampling rate. The microphones are worn during recording. 2.1 Pre-processing and Feature Extraction The brief descriptions of the pre-processing steps for log-scale Mel-spectrogram can be seen below. • The acoustic signal is sampled at 44100 Hz. The audio clips are then normalized. • The audio is converted to mid-side encoded [14] data to obtain good spatial information for CNN to detect moving sources. • The signal is then divided into 1s segments with an overlap of 0.5 s. It helps to train the network easier and reduces overfitting for certain acoustic events. The overlap increases the data for subsequent data augmentation. • The window size of 2048 samples using short-time Fourier transform with a hop size of 1024 samples are used. The samples overlap is 1024. The spectrogram has 128 bin mel-scale. The Mel-spectrogram is then converted into a logarithmic scale. • The log-Mel spectrogram data is reshaped before they are used as an image for training CNN. The first two dimensions are the height and width of the image, followed by the channels and the segments. For example, the size is 128 × 42 × 2 × 19. • The training labels are replicated to correspond with the 19 segments. • The dataset is augmented via Mixup [31]. It mixes the features of two different classes in equal proportion, as shown. x˜ = λxi + (1 − λ)xj
(1)
y˜ = λyi + (1 − λ)yj
(2)
where xi and xj are from dissimilar classes. The corresponding class labels are denoted by yi and yj , respectively. The mixing value of λ = 0.5 is used. 2.2 Convolutional Neural Network The CNN’s architecture can be seen in Table 1. The Batch Normalization (BN) and rectified linear unit (ReLU) [33] are used. The ReLU increases the non-linearity in
4
C. S. Chin and J. Zhang
the images. The batch normalization learning is used as a regularization to prevent overfitting. The activation function and BN are located before the convolution layer to improve the acoustic classification accuracy. The max-pooling layers come after the convolution process. The feature map that includes a prominent feature is obtained from the output of the max-pooling layer. The average pooling reduces the activation by combining the non-maximal activations. The last few layers consist of a dropout layer that removes 50% of the visible and hidden units to reduce overfitting. The fully connected layer is compiled the data to form the output for the last second layer that uses the softmax activation function to obtain probabilities of the input from the 15 classes. Lastly, the last classification layer produces the final classification. Table 1. CNN architecture. Description of each layer imageInputLayer- 128 × 42 × 2 batchNormalizationLayer convolution2dLayer- 32 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer convolution2dLayer- 32 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer maxPooling2dLayer- pool size 3 × 3, stride 2 × 2 and zero padding convolution2dLayer- 32 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer convolution2dLayer- 32 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer maxPooling2dLayer- pool size 3 × 3, stride 2 × 2 and zero padding convolution2dLayer- 128 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer convolution2dLayer- 128 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer maxPooling2dLayer- pool size 3 × 3, stride 2 × 2 and zero padding (continued)
Late Fusion of CNN with Wavelet-Based Ensemble Classifier
5
Table 1. (continued) Description of each layer convolution2dLayer- 256 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer convolution2dLayer- 256 filters (3 × 3) and zero padding batchNormalizationLayer reluLayer averagePooling2dLayer-pool size 16 × 6 dropoutLayer(0.5) fullyConnectedLayer(15) softmaxLayer classificationLayer
2.3 Wavelet Scattering The next step involves feature extraction using wavelet scattering for subsequent ensemble classifiers. It provides a good representation [28] of the time-frequency content of a signal. The first and second-order coefficients are used as most of the signal energy can be captured. The parameters of the transform are the filter-bank (using 1D Morlet wavelets) resolutions Q1 = 1 and Q2 = 4. The duration 0.75s of the averaging filter (or invariance scale) is used for the modulation structure duration. The sampling frequency is 44100 Hz. The first filter bank has a resolution of 4, and the second filter bank has a resolution of 1. 2.4 Ensemble Classifiers The proposed ensemble classifiers include different discriminant analysis learners, such as linear discriminant analysis (LDA), Quadratic discriminant analysis (QDA), and Regularized linear discriminant analysis (RDA) with other predictors covariance treatments. The random subspace learning method is used to increase the acoustic classification accuracy. In the random subspace, the feature subspaces are chosen randomly from the original feature space. The final prediction of these individual classifiers is then obtained using majority voting. 2.5 Fusion of CNN and Classifiers The fusion of the CNN and classifier prediction results indicates the relative confidence of their prediction. Multiplying the responses and determining the maximum response creates a late fusion system that inherent in the merits of each method. (3) class_pred i = argmax probiCNN , probiensem_class
6
C. S. Chin and J. Zhang
where probiCNN and probiensem_class are the probabilities of sound recording i from CNN and ensemble classifiers, respectively.
3 Results and Discussion The configurations of the proposed model are as follows. • • • • • • •
Stochastic gradient descent with momentum optimizer with a learning rate: 0.05 s Size of the mini-batch for each training iteration: 128 Momentum: 0.9 Maximum number of epochs: 8 Factor for L2 regularization: 0.005 Number of epochs for dropping the learning rate: 2 Multiplicative factor applied to the learning rate for each epoch: 0.2
The training data are shuffled before each training epoch. The entire experiment, including the pre-processing, took not more than three hour. The short audio segments (see Fig. 1) provide less information, thus making ASC difficult. A segment sample of the extracted Mel-spectrograms audio clip for the "lakeside beach" scene is shown in Fig. 1. The frequency along the y-axis, time is displayed along the x-axis, and the signal’s energy at a particular time and frequency is shown as the color map. The Intel® Core i7 CPU and Geforce RTX 2060 are used.
Fig. 1. Example of segments of Mel-spectrogram from the lakeside beach scene.
Late Fusion of CNN with Wavelet-Based Ensemble Classifier
7
The acoustic classification accuracy can be seen in Table 2. The ensemble classifiers have a higher acoustic classification accuracy than CNN. Compared to the baseline model (consists of 2 layers × 50 hidden units, 20% dropout), the fused model exhibits 18.4% higher accuracy. The details of the baseline model can be found in the reference [32]. Table 2. Acoustic classification accuracy of different models. Scenes
Acoustic classification accuracy (%) Baseline model [32]
CNN model
Ensemble classifiers model
Fused model
Beach
40.7
73.1
37.9
50.9
Bus
38.9
58.3
92.5
87.9
cafe/restaurant
43.5
74.0
82.4
82.4
Car
64.8
100
76.8
88.8
city-center
79.6
88.8
91.6
93.5
forest path
85.2
97.2
96.2
98.1
grocery store
49.1
70.3
79.6
79.6
Home
76.9
89.8
76.8
91.6
Library
30.6
49.0
36.1
40.7
metro station
93.5
100
95.3
100
Office
73.1
80.5
83.3
84.2
Park
32.4
20.3
68.5
60.1
residential area
77.8
63.8
77.7
81.4
Train
72.2
76.8
85.1
82.4
Tram
57.4
57.4
64.8
69.4
Average
61.0
73.3
76.3
79.4
Although the result of the scene (i.e., beach) using ensemble classifiers (37.96%) is quite poor as compared to CNN (73.14%), the fused model managed to increase the acoustic classification accuracy to 50.92%. Conversely, the scene (i.e. park) using CNN model is relatively low compared to the ensemble classifiers. The fused model increases it to 60.18%. The confusion matrix of CNN, ensemble classifiers, and the fused model are shown in Fig. 2. The confusion chart of the multi-model late fusion system shows better acoustic classification accuracy for city-center, forest path, and metro station than other scenes. The average acoustic classification accuracy of the fused model is computed as 79.43%. The false-negative for the residential area is around 51.6% with false discovery rate of 18.5%. The false negative is quite negligible for the classes.
8
C. S. Chin and J. Zhang
Fig. 2. Confusion charts of CNN (top), Ensemble classifiers (Middle), and Multi-model late fusion model (Bottom).
Late Fusion of CNN with Wavelet-Based Ensemble Classifier
9
4 Conclusion A multi-model late fusion system model consisting of the log-Mel spectrogram for convolutional neural network and wavelet time scattering for ensemble of subspace discriminant classifiers was proposed. The acoustic scene classification aims to classify the acoustic scenes in a different environment such as the park, car park, beach, citycenter, etc. Based on the dataset from the TUT Acoustic Scenes, it demonstrated that the fused model gives good acoustic classification accuracy of 79.43%. The proposed multi-model late fusion system exhibits 18.4% higher acoustic classification accuracy than the baseline model despite relatively low performance in a few scenes such as the beach and library. Nevertheless, the multi-model late fusion system shows good acoustic classification accuracy for most of the scenes. For future works, an adaptive type of hyperparameter tuning and advanced feature extraction methods will be used to improve the performance further.
References 1. Mesaros, A., et al.: Detection and classification of acoustic scenes and events: outcome of the DCASE 2016 challenge. IEEE/ACM Trans. Audio Speech Lang. Process. 26(2), 379–393 (2018) 2. Mesaros, A., Diment, B., Elizalde, T., Heittola, E., Vincent, B., Raj, T.: Virtanen, sound event detection in the DCASE 2017 challenge. IEEE/ACM Trans. Audio, Speech Lang. Process. 27(6), 992–1006 (2019) 3. Rakotomamonjy, A.: Supervised representation learning for audio scene classification. IEEE/ACM Trans. Audio Speech Lang. Process. 25(6), 1253–1265 (2017) 4. Trowitzsch, I., Mohr, J., Kashef, Y., Obermayer, K.: Robust detection of environmental sounds in binaural auditory scenes. IEEE/ACM Trans. Audio Speech Lang. Process. 25(6), 1344– 1356 (2017) 5. Ribeiro, P.O.C.S., et al.: Underwater place recognition in unknown environments with triplet based acoustic image retrieval. In: 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), Orlando, FL, pp. 524–529 (2018) 6. Aziz, S., Awais, M., Akram, T., Khan, U., Alhussein, M., Aurangzeb,