Now showing 1 - 3 of 3
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Real-Time Zero-Phase Digital Filter Using Recurrent Neural Network
    (2023-01-01)
    Sinjanakhom, Tantep
    ;
    This paper proposes a method to design and implement a zero-phase digital filter that can run in a real-time system. Generally, zero-phase filters are designed for non-causal systems only as the time-reversal operations are required. Thus, the typical usage of these filters is for offline applications. For this reason, we propose a real-time zero-phase digital filter that is designed based on a recurrent neural network model, particularly the gated recurrent units. The model learns to perform zero-phase filtering by using training data made from the filtered signals that are generated by using the conventionally designed zero-phase filter. The original digital filter used to create the dataset is an IIR filter performing forward-backward filtering. The best trained model yields the mean absolute loss values at approximately 0.001 and can process at least 30 times faster than real-time. Furthermore, the trained model was implemented as a 3-band zero-phase graphic equalizer to exhibit one of its applications.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Neural Networks for Real-Time Digital Emulation of Guitar Speaker Cabinet Impulse Response
    (2022-01-01)
    Sinjanakhom, Tantep
    ;
    This This paper presents a real-time signal processing system in which a neural network generates the impulse response (IR) of a Marshall 1960A guitar cabinet with 25W Celestion speakers based on user-specified parameters. The parameters include the microphone type, position of the speaker on which the microphone is mounted, distance between the microphone and the cabinet, and off-axis tilting angle. The trained model of neural network can generate the impulse response for a speaker cabinet, as well as the sound of settings not included in training set. Cross-correlation, error-to-signal ratio, power spectral density error, and magnitude-squared coherence were all utilized to assess the model's output. Mean Opinion Score listening tests were performed to determine the similarity of the convolved guitar signals. According to the results, the emulated cabinet sounds were perceived to be nearly identical to the original sounds. The performance of the real-time audio plugin implementation is proved to be computationally efficient. Because raw IR data for each microphone configuration does not need to be saved directly to the PC's memory, utilizing it in music production work can be more convenient, allowing the user to modify the parameters while hearing the differences without having to repeat the IR file loading procedure.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Musical Key Classification Using Convolutional Neural Network Based on Extended Constant-Q Chromagram
    (2022-01-01) ;
    Sinjanakhom, Tantep
    ;
    In the field of music information retrieval, musical key classification is one of the challenges. This paper illustrates the advantages of the proposed system with relevant experimental results, starting with diverse audio datasets for feature extraction used for training and testing a classification model which is based on a convolutional neural network (CNN). The goal is to develop a feature that can improve the neural network's performance. To compare the effect of input features on efficiency, a basic CNN is trained from the ground up and utilized as an image classification tool. The Chromagram-24, an augmented version of the input chroma feature, is proposed to improve the accuracy of musical key detection. In terms of weighted score, the model using Chromagram-24 as an input feature outperforms the model trained using a conventional 12-dimensional chromagram by 12.77% and achieves the highest score of 85.63% when classifying full-length songs. Chromagrams are generated using audio excerpts ranging in length from 15 to 60 seconds for local key estimation, whereas, for global key estimation, a full-length audio set is used. The results indicate that, given the different lengths of training audio input, executing the model using a chromagram of a 60-second audio excerpt yields the best results.