A Predictive Dual-Stage Neural Framework for Phase-Coherent Auditory Synthesis on Edge Devices

dc.contributor.authorPairoch, Sathit
dc.contributor.authorPhasukkit, Pattarapong
dc.contributor.authorSuteewong, Teeraporn
dc.date.accessioned2026-08-06T10:55:37Z
dc.date.available2026-08-06T10:55:37Z
dc.date.issued2026-06-01
dc.description.abstractReal-time binaural beat synthesis in dynamic acoustic environments is challenged by carrier non-stationarity, interaural phase discontinuities, and processing delay in conventional digital signal processing pipelines. This study proposes a predictive dual-stage neural framework for phase-coherent auditory synthesis under non-stationary acoustic conditions. The framework decouples real-time carrier estimation from phase-coherent signal generation through two specialized modules. An intelligent acoustic sensing module (AI-1) estimates time-varying carrier information across harmonic, fluctuating, and broadband acoustic profiles using a causal neural front-end with an adaptive confidence-driven strategy. A predictive phase-coherent generator (AI-2) then forecasts short-horizon carrier trajectories and drives a discrete-time phase accumulator to maintain continuous phase evolution during binaural beat embedding. Objective evaluation under multiple acoustic profiles and noise conditions shows that the proposed framework maintains strong phase continuity, with a Phase Coherence Factor greater than 0.91, and low artifact levels, with a Signal-to-Artifact Ratio greater than 39.8 dB, under the evaluated conditions. Additional comparisons with conventional DSP baselines, stronger classical F0 estimators, a lightweight neural F0 tracker, and component-wise ablation variants further demonstrate that the performance improvement arises from the combination of adaptive carrier estimation and predictive phase-coherent actuation, rather than from carrier estimation alone. Hardware profiling shows a combined INT8 inference time of 2.4 ms per frame on a resource-constrained Raspberry Pi Zero 2W-class edge device. Importantly, this inference time and the sub-millisecond phase-accumulator resolution should not be interpreted as sub-millisecond end-to-end physical audio latency. The complete system still includes buffering, framing, neural inference, and output processing delay; the proposed method instead reduces effective phase-boundary misalignment through short-horizon predictive compensation. These results support the proposed framework as a lightweight engineering solution for real-time phase-continuous auditory synthesis in dynamic listening environments. The reported PCF and SAR values should be interpreted as signal-level indicators of phase continuity and artifact suppression, rather than as evidence of listener comfort, perceptual preference, or neurophysiological efficacy.
dc.identifier.citationSensors, 26(11), 2026
dc.identifier.doi10.3390/s26113344
dc.identifier.issn14248220
dc.identifier.other2-s2.0-105041481533
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/18130
dc.sourceSensors
dc.subjectauditory neuro-stimulation
dc.subjectbinaural beats
dc.subjectedge AI
dc.subjectphase-coherent synthesis
dc.subjectpredictive signal processing
dc.subjectreal-time processing
dc.titleA Predictive Dual-Stage Neural Framework for Phase-Coherent Auditory Synthesis on Edge Devices
dc.typeArticle

Files

Collections