ICSPCC2026 Conference Program
SCHEDULE AT A GLANCE
Friday, July  17, 2026 (Chuyun Hall)
16:00-18:00 Registration
18:00-20:00 Cocktail Reception
Saturday, July  18, 2026 (Chuyun Hall)
08:30-18:30 Registration
08:30-09:00 Opening Ceremony
09:00-10:00 Keynote Speech 1: Deep Learning Meets Beamforming: Bridging Statistical Signal Processing and Data-Driven Microphone Array Processing
Prof. Sharon Gannot, Bar-Ilan University
10:00-10:20 Coffee/Tea Break
10:20-11:20 Keynote Speech 2: Randomization and Hybrid Learning Based Deep and Shallow Networks for Classification and Forecasting
Prof. Ponnuthurai Nagaratnam Suganthan, Qatar University
11:20-11:50 Invited Talk 1: Minor Manipulations, Major Threat: An Overview of Partially Fake Speech
Dr. Lin Zhang, Johns Hopkins University
11:50-12:20 Invited Talk 2: Developing Controllable, Efficient, and Versatile Diffusion Models via Prior Exploitation
Dr. Zehua Chen, Tsinghua University
12:20-14:00   Lunch Break
  Room 5, Level B Room 6, Level B Room 7, Level B
14:00-15:30 BSPC SPGT 01 COMM 01
15:30-16:00 Coffee/Tea Break 
16:00-17:30 BSPC SPGT 02 CPT 01
Sunday, July 19, 2026
08:30-18:30 Registration
  Room 5, Level B Room 6, Level B Room 7, Level B
09:00-10:30 COMM 02 SPGT 03 CPT 02
10:30-11:00 Coffee/Tea Break 
11:00-12:30 COMM 03 SPGT 04 CPT 03
12:30-14:00 Lunch Break
14:00-15:30 COMM 04 SPGT 05 SPSS 01
15:30-16:00   Coffee/Tea Break
16:00-17:30 COMM 05 SPGT 06 SPSS 02
18:30-20:30 Conference Banquet and BSPC Award Ceremony
Monday, July 20, 2026
  Room 5, Level B Room 6, Level B Room 7, Level B
08:30-10:00 SPSS 03 SPGT 07  
10:00-10:30 Coffee/Tea Break
10:30-12:00 SPGT 08 SPGT 09  
SPGT: Oral Presentation for Signal Processing track
CMPT: Oral Presentation for Computing track
COMM: Oral Presentation for Communication track
SPSS: Signal Processing: Special Session
BSPC: Best Student Paper Contest



Instructions for Oral Presentation
Instructions for Oral Presentation
Oral Presentation enable author to share their knowledge and experience with experts and researchers.  Authors should relax themselves, just summarize the key points of the Paper and give 10-15 talk.  It is not another exam for student author, just like talking to your colleagues will do.
The allocated presentation time for each paper is 20 minutes which includes the time for questions and answer from the audiences.  In genral, each PPT slide presentation takes 30 seconds, so your PPT should not have more than 25 slides.
Presenters are required to report to their Session Chairs at least 10 minutes prior to the start of their Session.   All oral presentations must be loaded from USB into the note-book and tested before the session.
Microsoft Power Point file are recommended.  Movies or animations in MPEG, Windows Media, and etc., should be tested before the session.  


<>
Oral Presentation Progrm 
BSPC Session 01
Time: 14:00~15:30, Saturday, July 18, 2026
Session Chair: Wenxing Yang, University of Shanghai for Science and Technology
  
Oral Session: BSPC Session 01 (Paper No.6418)
Title: A Behavior-Oriented Feature Time-Series Processing Framework for Long-Term Anomalous Spectrum Behavior Detection
Author: Qin Li1, Xiang Wang1, Tao Zhi Huang1, Fa Yi Zhang1 and Ya Shu Cao1
Abstract: Long-term spectrum monitoring is an important foundation for electromagnetic domain security, spectrum situational awareness, and stable wireless system operation. One key task is to detect anomalous spectrum behaviors in continuous observations. Existing spectrum anomaly detection studies mainly focus on signal-level or image-level representations. They usually treat anomalies as abnormal components in local spectrum segments and provide limited support for modeling the persistence and temporal evolution of anomalous spectrum behaviors under long-term non-stationary monitoring conditions. Meanwhile, continuous monitoring produces massive complex baseband I/Q data. End-to-end modeling based on raw I/Q data or high-dimensional spectrum images leads to high storage and computational costs. To address these issues, this paper proposes a behavior-oriented feature time-series processing framework for long-term anomalous spectrum behavior detection. The framework converts continuous I/Q monitoring segments into a multivariate feature time series. It extracts time-domain statistical features, frequency-domain structural features, and cross-time spectral variation features to characterize spectrum states and their behavioral evolution during long-term monitoring. The constructed feature time series compresses the data volume while preserving behavior-level information. It provides a scalable and interpretable representation for downstream anomaly detection models. Experiments on a long-term wideband spectrum monitoring dataset show that the proposed framework improves point-level detection performance and event-level alarm performance compared with power spectral density image and time-frequency image representations.
Oral Session: BSPC Session 01 (Paper No.6356)
Title: Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
Author: Tianyan Deng1, Yanxiong Li2, Rui Gao2 and Jiahao Du2
Abstract: Few-shot Open-set audio classification requires classifying query samples from known classes with a few labeled support samples while rejecting query samples from unknown classes. Transductive inference jointly observes the full unlabeled query set to improve prototype estimation, yet standard transductive updates do not distinguish known from unknown query samples, leaving prototypes vulnerable to open-set contamination. Drawing on latent-inlierness weighting and decoupled scoring for unknown-class samples, we propose a two-phase transductive method operating over a frozen audio encoder. First, each query sample is assigned a latent inlierness score that down-weights likely unknown-class samples, so that prototype refinement is driven primarily by known-class evidence. The refined prototypes are then directly optimized on a transductive loss combining support cross-entropy, inlierness-weighted conditional entropy minimization, and inlierness-weighted marginal entropy maximization, while open-set rejection uses a prior-adaptive free-energy score that adjusts its threshold with the prior proportion of unknown-class samples, decoupling detection from classification. Experiments on three audio datasets show our method achieves state-of-the-art results for few-shot open-set audio classification under multiple experimental conditions.
Oral Session: BSPC Session 01 (Paper No.6349)
Title: Link-Aware GNN-Based Critical Node Identification for Underwater Acoustic Sensor Networks
Author: Ziyi Ding1, Yihao Zhao1, Zheyang Chen1, Shenao Tu2, Yougan Chen2 and Xiaomei Xu1
Abstract: Underwater acoustic sensor networks (UASNs) are indispensable for marine exploration, yet the failure of critical nodes can easily precipitate network fragmentation and routing voids. Existing node identification methods rely primarily on static topological connections, neglecting the dynamic influence of acoustic links. To address this, we propose a novel critical node identification algorithm based on a Link-Aware Graph Neural Network (LAGNN), which explicitly embedding node attributes and link stability metrics into the graph structure. By designing a joint loss function driven by network Quality of Service (QoS) degradation and leveraging a message-passing mechanism, LAGNN deeply extracts the complex correlations between local link condition and global network topology. Simulation results verify that LAGNN outperforms baselines with a 26% higher Top-k hit rate. Removing its identified nodes inflicts catastrophic QoS degradation, plunging the PDR to merely 20% of the benchmark level and spiking the delay by 30%, which proves its precision in isolating critical bottlenecks.
BSPC Session 02
Time: 16:00~17:30, Saturday, July 18, 2026
Session Chair: Wenxing Yang, University of Shanghai for Science and Technology
  
Oral Session: BSPC Session 02 (Paper No.6325)
Title: LiteMeter: Attention-Enhanced Keypoint Detection and Edge Inference for Industrial Gauge Reading
Author: Huasong Li1, Dongliang Fu2, Jiongmin Yu2, Jiakai Li1, Minghui Ouyang1 and Wei Gao1
Abstract: Automatic reading of industrial analog gauges remains challenging in terms of both detection accuracy and deployment flexibility. This paper presents LiteMeter, a unified system that integrates an enhanced keypoint detector, a cross-platform frontend, and an edge inference service. The proposed detector incorporates the convolutional block attention module (CBAM) into the shallow stages of YOLOv11 and introduces a P2 high-resolution prediction head to improve fine-grained feature representation. Gauge values are estimated via a four-keypoint polar-coordinate mapping strategy. Extensive experiments conducted on a Hard Test benchmark with 11 types of industrial degradation demonstrate that the combined use of P2 and CBAM constrains the performance drop to 37.58 pp, limits the generalization gap to 0.22 pp, and achieves a degraded precision of 85.02%, along with a Grad-CAM high-response ratio of 0.1921. For deployment, a frontend based on Tauri 2.0 and ONNX Runtime supports both desktop and mobile platforms, enabling seamless switching between local and remote inference with automatic fallback. Furthermore, a C++17-based inference server deployed on Jetson Orin Nano achieves a per-frame latency of 58.6 ms using TensorRT FP16 (85.1 ms end-to-end, corresponding to 11.7 FPS). An EMA-adaptive RTSP streaming pipeline is further designed to support long-term, unattended operation.
Oral Session: BSPC Session 02 (Paper No.6317)
Title: Lightweight Spatiotemporal Attention-Based Driver Unsafe Behavior Recognition
Author: Mengqi Liu1, Yang Lai1, Xinzi Wang2, Zhiyuan Xue1, Ben Yang1 and Xuetao Zhang1
Abstract: Unsafe driver behaviors remain a major cause of traffic accidents. However, existing methods are often limited by insufficient data diversity and inadequate temporal modeling. To address these issues, we propose a lightweight spatiotemporal modeling framework with dynamic attention (LSDA) for driver behavior recognition. In addition, a driver behavior dataset is constructed by simulating realistic bus and truck driving scenarios, enabling the representation of typical operational conditions while maintaining controllability. Specifically, a lightweight architecture with dense temporal sampling is designed to effectively capture short-term motion dynamics. Furthermore, a prototype-driven dynamic self-attention mechanism is introduced to adaptively emphasize discriminative temporal features, thereby improving the separability of similar behavior classes. In addition, Mixup and CutMix-based feature augmentation are employed to enhance data diversity and generalization. Extensive experiments on the constructed dataset and the SynDD1 dataset demonstrate that LSDA consistently outperforms state-of-the-art methods, particularly in fine-grained short-term behavior recognition, validating its effectiveness and efficiency.
Oral Session: BSPC Session 02 (Paper No.6472)
Title: Deep Learning-Driven DOA Estimation for Distributed MIMO Radar in UAV-Assisted Sensing Systems
Author: Chuang Han1, Ke Ding1, Yan dong Sun1, Yue xian Wang1 and Cheng yan He1
Abstract: Unmanned aerial vehicle (UAV) auxiliary sensing system has been widely used in the fields of national security, emergency response and intelligent transportation. Direction-of-arrival (DOA) estimation is necessary for target localization and trend sensing. However, traditional algorithms such as MUSIC and SBL often suffer from performance degradation under low SNR conditions. The spatial-temporal data transmitted by distributed nodes are difficult to fuse, and the computational complexity is very high. Aiming at these problems, this paper proposes a deep learning-driven DOA estimation algorithm for UAV-assisted distributed multiple-input multiple-output (MIMO) radar, and establishes a modular deep learning model. An end-to-end multi-branch fusion deep neural network is designed. Each branch processes each subarray data independently, and then realizes the collaborative processing of multi-node information through feature fusion, so as to achieve efficient data fusion and high-precision angle estimation. Its performance is verified by a large number of simulations. The simulation results show that compared with the traditional algorithm, the proposed algorithm achieves excellent results in RMSE, accuracy and real-time performance under low SNR.
COMM Session 01
Time: 14:00~15:30, Saturday, July 18, 2026
Session Chair: Long Shi, Southwestern University of Finance and Economics
  
Oral Session: COMM Session 01 (Paper No.6352)
Title: Q-Learning-Based Routing and Caching Optimization for Underwater Information-Centric Networks
Author: Zheyang Chen1, Yihao Zhao1, Shenao Tu2, Ziyi Ding1, Yougan Chen2 and Xu Xiaomei2
Abstract: Underwater acoustic networks encounter substantial challenges in long-distance horizontal transmission, including high propagation delay and excessive energy consumption, which hinder real-time and energy-efficient operation in the underwater Internet of Things. To overcome the broadcast storms, excessive caching overhead, and lack of globally optimal routing in existing underwater information-centric networking schemes, this paper proposes QL-ICN, a Q-learning-based information-centric networking framework that integrates hybrid electromagnetic communication among surface sink nodes with acoustic links to underwater nodes. The framework models request forwarding as a Markov decision process and employs Q-learning to select optimal unicast multi-hop paths, thereby eliminating broadcast overhead. In addition, a trust score based on the count of node selection as delivery targets is introduced as a priority metric to guide intelligent cache node selection during the data return phase. Simulation results demonstrate that as the number of consumers increases, the network latency decreases significantly, while the total energy consumption exhibits only sublinear growth. In both ablation studies and comparisons with baseline algorithms, QL-ICN exhibits superior performance: it maintains the lowest energy consumption while achieving an average latency that is only 75.85% of the benchmark algorithm RAOH's.
Oral Session: COMM Session 01 (Paper No.6367)
Title: DDPG-Based Energy-Efficient Resource Scheduling in Underwater Acoustic Communications
Author: Tong Li1, xuefeng zhong1, Yuting Xiao1, Hao Zhao2 and Zilong Jiang3
Abstract: Abstract—In underwater acoustic communications (UAC),
achieving energy-efficient transmission is critical due to the harsh
marine environment and limited battery capacity of underwater
nodes. To overcome this limitation, this paper proposes an
Adaptive Continuous-Action Deep Deterministic Policy Gradient
(AC-DDPG) algorithm. This approach enables decision-making
directly within a continuous action space, achieving fine-grained
joint optimization of transmission frequency and power without
discretization. Different from the DQN-based reinforcement
learning algorithms, AC-DDPG adopts an Actor–Critic architecture
and incorporates an adaptive Ornstein–Uhlenbeck (OU)
noise mechanism to effectively balance exploration and exploitation
during training. Experimental results demonstrate that ACDDPG
achieves energy efficiency close to the optimal solution in
both small and large action spaces. Its performance significantly
outperforms current reinforcement learning algorithms such as
Q-learning, DQN, and TAS-DQN, exhibiting superior optimization
capability and stability, particularly in large-scale continuous
search spaces.
Oral Session: COMM Session 01 (Paper No.6382)
Title: A Traffic-Aware MAC Protocol for Underwater Acoustic Sensor Networks Based on DSATUR Graph Coloring
Author: Ruisi Gou1, Xiaohong Shen1, Weiliang Xie1 and Haiyan Wang1
Abstract: Underwater acoustic sensor networks (UASNs) are characterized by long propagation delays, limited bandwidth, and strong spatiotemporal variability, which make it difficult for medium access control (MAC) protocols to simultaneously achieve collision-free access and efficient resource utilization. Conventional TDMA suffers from limited spatial reuse, contention-based protocols experience severe collisions under heavy traffic, and existing graph-coloring-based MAC protocols adapt poorly to non-uniform traffic loads. To address these issues, this paper proposes a DSATUR-based traffic-aware MAC protocol, termed TAGC-MAC. First, a two-hop interference conflict graph is constructed according to the network topology, and the DSATUR algorithm is applied to obtain a baseline collision-free slot assignment, thereby reducing the number of required slots and improving spatial reuse. Based on this baseline, per-node slot demand is estimated from the packet generation rate of each node, and additional slots are then iteratively allocated using a greedy maximal independent-set strategy to alleviate queue buildup at high-load nodes. Simulation results show that the proposed protocol achieves good performance in throughput, fairness, and delay, and further improves system throughput under non-uniform traffic loads, effectively enhancing channel access efficiency in UASNs.
Oral Session: COMM Session 01 (Paper No.6430)
Title: A Medium Access Control Protocol for Underwater Acoustic Communication Networks Based on Deep Reinforcement Learning
Author: Jianmin Yang1, Zhihong Peng1, Yan Lin1, Ming Gong1, Huang Ying1, Hongli Wang2 and Pengyu Du3
Abstract: Abstract—Underwater Acoustic Communication Networks
(UACNs) play an irreplaceable role in applications such as
marine environmental monitoring, resource exploration, and
disaster warning systems. The performance of the Medium
Access Control (MAC) protocol directly affects network
communication quality, making the design of efficient MAC
protocols a major research focus in UACNs.
This paper proposes a Parameter-Sharing Independent QLearning
(PS-IQL) MAC protocol for UACNs based on a Deep
Q-Network (DQN), referred to as the DQN-MAC protocol. In a
distributed homogeneous network, an external agent is
introduced to conduct centralized training using the DQN
algorithm. During training, all nodes share the same network
parameters, while their learning experiences are stored in a
shared replay buffer. Mini-batch samples are then drawn from
the buffer to update the neural network through the loss
function. After convergence, the learned policy is deployed to
each node and learning is terminated, leaving only inference
during execution.
When a data packet enters the MAC queue, a node
performs inference based on the observed network state and
the learned policy to determine the transmission frame length.
When packet retransmission is triggered by backoff
mechanisms, the node further optimizes its backoff duration
through policy inference, thereby improving the overall
protocol performance. Experimental results demonstrate that
the DQN-MAC protocol achieves stable and effective
convergence during training and performs robustly under
different testing conditions. Moreover, it shows clear
advantages in Packet Delivery Ratio (PDR) and network
throughput.
Oral Session: COMM Session 01 (Paper No.6434)
Title: Q-Learning and Bayesian Optimization for Energy-Efficient and Self-Adaptive Routing Protocol in Underwater Wireless Sensor Networks
Author: Jianmin Yang1, Yue Ma1, Yonghui Zhong1, Haohua He1, Qing Liang1 and Jiajing Chen1
Abstract: Abstract—Underwater wireless sensor networks (UWSNs)
suffer from limited node energy, unreliable acoustic channels,
and dynamic topologies, which together cause unbalanced
energy consumption and shortened network lifetime. In this
paper, we propose QBEAR, an energy-efficient clustering
routing protocol that combines Q-learning with Bayesian
optimization for UWSNs. In the inner layer, a Q-learning
mechanism drives cluster head (CH) election and inter-cluster
routing decisions: each node perceives multi-dimensional state
features including residual energy, distance to sink, and
neighbor density, and uses a reward function jointly
considering energy retention, topology quality, and link
reliability to update its Q-values. In the outer layer, a Bayesian
optimization framework automatically tunes the Q-learning
hyper-parameters and the cluster-control weight parameters,
allowing the system to autonomously approach globally
balanced performance without manual tuning. Simulation
results show that the proposed QBEAR protocol outperforms
LEACH, EECRAP, and QHUC in terms of packet delivery
ratio, energy-consumption balance, and network lifetime in
large-scale UWSN deployments.
Oral Session: COMM Session 01 (Paper No.6443)
Title: Graph-Aware MAPPO for Event-Coverage-Oriented Underwater Sensor Deployment under Meandering Currents
Author: Jianmin Yang1, Ying Huang2, Can Wang3, Ming Gong1, Yan Lin1 and Wenwei Chen1
Abstract: Event-coverage-oriented underwater acoustic sensor deployment aims to adapt mobile sensor positions to nonuniform event distributions. In current-affected underwater environments, this task is complicated by the coupling between controlled movement and ocean-current-induced drift. Existing heuristic methods usually rely on handcrafted movement rules or scenario-specific search, which limits their ability to provide reusable decision policies. This paper presents a graph-aware multi-agent proximal policy optimization framework for underwater sensor deployment under meandering currents. The environment is modeled as a three-dimensional monitoring region affected by the Meandering Current Mobility model, where each sensor node selects a discrete movement action. For each agent, a sensor-event graph is constructed to encode acoustic communication neighborhoods among sensors and distance-dependent sensing relations between sensors and events. A graph-aware actor extracts the encoded representation of the current sensor node for decentralized action selection, while training is performed with decentralized actors and a shared critic. Simulation results show improved coverage-efficiency performance during training and better final performance than the compared baseline.
Oral Session: COMM Session 01 (Paper No.6446)
Title: Q-Learning Routing with Incremental Topology Maintenance for Multi-Sink Underwater Acoustic Sensor Networks under Ocean-Current Mobility
Author: Jianmin Yang1, Muzi Cui1, Ying Huang1, Haiquan Shi2, Jinwang Luo2, Yonghui Zhong1 and Yue Ma1
Abstract: Underwater acoustic sensor networks (UASNs) suffer from severe routing degradation caused by continuous node displacement driven by ocean currents, which rapidly invalidates cached topology and link quality. Existing protocols either ignore current-induced mobility or rely on costly full topology recomputation, leading to high packet loss and excessive energy consumption. This paper proposes IR2P-DLMH, integrating four coordinated mechanisms. For a start, an event-driven incremental topology maintenance scheme updates only affected neighbor entries and Q-table values when cumulative displacement exceeds a threshold, avoiding global recomputation while staying synchronized with ocean-current-driven topology changes. Additionally, a link stability predictor evaluates candidate links at drift-projected future positions, proactively bypassing paths about to fail under current influence. Next, a capability-aware reward coefficient steers relay traffic toward hardware-superior nodes without centralized control. Ultimately, a multi-sink load balancing strategy jointly minimizes propagation distance and load deviation across sinks. Simulation results demonstrate that IR2P-DLMH significantly reduces packet loss ratio and energy consumption per packet while achieving lower end-to-end delay compared to VBF and HH-VBF under varying ocean current velocities and network densities.
Oral Session: COMM Session 01 (Paper No.6447)
Title: A Delay-Aware Q-Learning-Based Medium Access Control Protocol for Underwater Acoustic Networks
Author: Jianmin Yang1, Longsen Du2, Zhuoqian Wu1, Ma Yue1, Yonghui Zhong1 and Pengyu Du3
Abstract: Underwater acoustic networks suffer from long propagation delay, limited bandwidth, high bit error rate, and severe packet collisions caused by multi-node channel contention. Conventional ALOHA-based medium access control (MAC) protocols are simple and easy to deploy, but their random access nature leads to rapid performance degradation under medium and high traffic loads. Although TDMA-based protocols can reduce collisions through scheduled slot allocation, they usually require strict time synchronization, centralized coordination, or prior topology knowledge, which limits their adaptability in dynamic underwater environments. To address these problems, this paper proposes DA-ALOHA-Q, a delay-aware Q-learning-based MAC protocol for underwater acoustic networks. The proposed protocol introduces a lightweight reinforcement learning mechanism into slotted ALOHA and simplifies the MAC contention problem into a single-state multi-action slot selection problem. Each node only maintains a local one-dimensional Q-table and updates its slot preference according to ACK reception or timeout feedback. In this way, distributed nodes can gradually form a pseudo-orthogonal slot allocation pattern without centralized scheduling or explicit signaling exchange. The protocol is implemented in NS-3 with the Aqua-Sim-NG underwater network module. Simulation results under both ring and three-dimensional random topologies show that DA-ALOHA-Q can reduce packet collisions, improve packet delivery ratio, and maintain more stable throughput under high traffic loads compared with conventional ALOHA and adaptive backoff ALOHA.
COMM Session 02
Time: 09:00~10:30, Sunday, July 19, 2026
Session Chair: Yingke Zhao, Shaanxi University of Science and Technology
  
Oral Session: COMM Session 02 (Paper No.6435)
Title: Change-Point Aware Spatiotemporal Hawkes Process for Dynamic Topology Inference in Non-Cooperative Underwater Acoustic Networks
Author: Gaoyue Ma1, Xiaohong Shen1, Yuwen Yan1, Bo Geng1 and Haiyan Wang1
Abstract: The coexistence of multiple underwater acoustic networks (UANs) in ocean environments leads to highly congested and uncoordinated communication over bandwidth-limited channels. In such scenarios, inferring communication topology from passive observations is essential for spectrum sharing and mission coordination. However, long and uncertain acoustic propagation delays, together with time-varying communication patterns, make this problem particularly challenging. In this paper, we propose a Hawkes process model for passive topology inference in UANs that explicitly accounts for acoustic propagation and piecewise stationarity. The model incorporates an excitation kernel governed by a Gamma distribution to characterize propagation delays and temporal uncertainty in acoustic signal transmission. To model topology evolution over time, we formulate dynamic topology inference as a change-point detection problem and develop a candidate generation and dynamic programming segmentation framework that jointly estimates change points and interaction structures. Experimental results demonstrate that the proposed method achieves accurate and robust topology inference under varying observation durations and arrival-time uncertainty levels, and reliably detects topology changes, outperforming existing baseline approaches.
Oral Session: COMM Session 02 (Paper No.6439)
Title: Doppler-Scale Domain Turbo Equalization for Wideband Time-Varying Channels
Author: Yuanyuan Ou1, Zhehan Guo2, Jianchun Xu1 and Yaokun Liang2
Abstract: Wideband time-varying channels, particularly underwater acoustic channels, present significant challenges for reliable communication due to frequency-dependent non-uniform Doppler shifts. While orthogonal time frequency space (OTFS) has been adopted for time-varying systems, it suffers from performance degradation in these wideband scenarios. In this paper, an encoded orthogonal delay scale space (ODSS) scheme integrated with Turbo equalization is proposed. The proposed scheme leverages the inherent characteristics of the delay-scale domain in ODSS to suppress wideband Doppler interference and enhance equalization accuracy. Moreover, the ODSS modulation is extended to higher-order modulation schemes to improve spectral efficiency for wideband transmission. To validate the superiority of the proposed Turbo equalization algorithm, simulations are carried out under typical wideband time-varying conditions, comparing with existing equalization algorithms under the ODSS framework. The results demonstrate a BER performance improvement over existing ODSS-based algorithms, with performance enhancement with the increasing number of Turbo iterations.
Oral Session: COMM Session 02 (Paper No.6440)
Title: Transmit Beamforming based on Temporally-Stable Channel Parameters for Underwater Acoustic OFDM Communications
Author: Pengxiang Zhang1 and Jun Tao1
Abstract: Due to the long feedback delay, conventional transmit beamforming approaches relying on instantaneous channel state information (CSI) are generally inapplicable over underwater acoustic channels. This paper proposed a robust beamforming method for underwater acoustic orthogonal frequency division multiplexing (OFDM) systems, by relying on temporallystable channel parameters in form of direction of arrival (DoA) and large-scale path gain. A path identification (PI) algorithm is employed to estimate aforementioned channel parameters, based on which beamforming weights are designed under the minimum variance distortionless response (MVDR) criterion. To improve numerical stability, diagonal loading technique is further incorporated. Simulation results demonstrate that the proposed parametric MVDR (P-MVDR) method outperforms existing conventional beamforming (CBF) and null-steering (NS) beamforming both utilizing the DoA knowledge of principal path only. Moreover, it is less sensitive to DoA estimation error.
Oral Session: COMM Session 02 (Paper No.6460)
Title: Robust Adaptive Beamforming Based on Sensitivity-Adaptive Gauss-Chebyshev Quadrature
Author: Han Chuang1, Li Xiang1, Sun Yandong1, Wang Yuexian1 and He Chengyan1
Abstract: Robust adaptive beamforming (RAB) techniques based on interference-plus-noise covariance matrix (INCM) reconstruction have demonstrated strong performance under steering vector mismatch. However, existing methods often suffer from high computational complexity and sensitivity to inaccurate direction-of-arrival (DOA) estimation. To address these issues, this paper proposes a novel RAB algorithm termed sensitivity-adaptive Gauss-Chebyshev quadrature (SA-AGCQ). The proposed method integrates three key components. First, a centroid-based DOA estimation approach with power-squared weighting is developed to improve robustness against noise and angular mismatch. Second, Gauss-Chebyshev quadrature is employed to efficiently reconstruct the INCM, significantly reducing the number of integration points compared with conventional uniform sampling. Third, a sensitivity-adaptive weighting mechanism is introduced to emphasize dominant interference regions while suppressing noise amplification during covariance reconstruction. Simulation results demonstrate that the proposed method performance over a wide range of SNRs, maintains robustness under steering vector mismatches, and performs effectively even with limited snapshots. In addition, the proposed algorithm reduces computational complexity compared with conventional INCM reconstruction approaches, making it suitable for practical applications.
Oral Session: COMM Session 02 (Paper No.6473)
Title: DFT-spread OTFS Waveform Design for ISAC-enabled Underwater Acoustic Communications
Author: Jiaqi Yang1, Tonghui Zheng1, Mingqi Jin1, Chenming Zhu1 and Chengbing He1
Abstract: Integrated Sensing and Communications (ISAC) is a key trend for future underwater acoustic (UWA) systems. While Orthogonal Time Frequency Space (OTFS) modulation is robust against doubly selective fading, its high peak-to-average power ratio (PAPR) and pilot overhead limit its practical use. This paper proposes a DFT-spread OTFS (DFT-s-OTFS) waveform design for ISAC-enabled UWA communications. We employ DFT precoding to suppress PAPR and integrate a superimposed pilot scheme to enhance spectral efficiency by eliminating dedicated guard bands. Furthermore, an advanced iterative receiver is developed to facilitate joint channel estimation, interference cancellation, and equalization. Simulation results validate the efficacy of the proposed framework, demonstrating superior error performance and robustness in challenging UWA channels.
Oral Session: COMM Session 02 (Paper No.6485)
Title: Joint CFO and Channel Estimation for Underwater Acoustic OCDM Communications
Author: Lingling Zhang1, Hanyu Guo1, Chengkai Tang1 and Lu Ma2
Abstract: OCDM, by leveraging the diagonal distribution characteristic of Chirp signals in the time-frequency domain, modulation and demodulation are achieved via the Discrete Fresnel Transformation (DFnT), endowing the signal with full diversity gain and inherent resistance to multipath and Doppler interference. This paper present the joint CFO and channel estimation for underwater acoustic OCDM communications. It decouples the pilot and data sub-blocks by adopting the multicarrier OCDM mode, and estimate the CFO and channel multipath effect with the pilots. Simulation result shows the proposed method could decrease the channel estimation error and improve the robustness of OCDM communication over underwater acoustic channel.
COMM Session 03
Time: 11:00~12:30, Sunday, July 19, 2026
Session Chair: Yingke Zhao, Shaanxi University of Science and Technology
  
Oral Session: COMM Session 03 (Paper No.6377)
Title: A Combinational Neural Network MIMO-OFDM Uplink Receiver for LEO Satellite Constellations
Author: Shan Lu1, Chengkai tang1, Yi Zhang1, Lingling Zhang1, Yangyang Liu1 and Zesheng Dan1
Abstract: Low-Earth-orbit (LEO) mega-constellations are being deployed as a key infrastructure for future space-air-ground integrated networks. The conventional single-TT&C-station/single-satellite uplink, however, is not well suited to scenarios that require wide coverage, reliable access, and high-rate transmission. Severe path loss, multipath fading, limited pilot overhead, and nonlinear interference further complicate channel estimation and signal recovery in the uplink receiver. To address this problem, we develop MIMO-CN, a model-and-data-driven MIMO-OFDM receiver for the single-station multi-satellite TT&C uplink. The receiver consists of a model-initialized channel-estimation subnetwork followed by a signal-detection subnetwork. LS-based channel estimation and ZF/MMSE-based detection are used as communication-domain initialization, and the neural subnetworks are trained to refine the residual channel and detection errors. Simulation results over a Saleh multipath fading channel with additive white Gaussian noise show that MIMO-CN achieves a lower bit-error rate than fully connected DNN and LMMSE-MMSE baselines in both sufficient-pilot and pilot-limited cases. For the same BER target, MIMO-CN reduces the required SNR by approximately 3-5 dB and requires much less memory than a conventional fully connected DNN receiver.
Oral Session: COMM Session 03 (Paper No.6449)
Title: Research on PMF-FFT Capture Algorithm with Dual segment Adjustable Parameter Window
Author: Zhang Hang1, Tang Cheng kai2 and Zhang Ling ling3
Abstract: In response to the problems of severe main lobe attenuation and significant scallop loss in the PMF-FFT acquisition algorithm for low orbit satellite signals under conditions of large Doppler frequency shift range and high capture search dimension, this paper proposes a PMF-FFT acquisition algorithm based on a dual segment adjustable parameter window. This algorithm divides the capture process into two stages: PMF and FFT. In each stage, adjustable parameter windows are introduced for weighted processing, and the window parameters are adjusted to achieve optimized matching in different stages, thereby achieving joint suppression of main lobe attenuation and scallop loss. The simulation results show that this method can significantly suppress main lobe attenuation and scallop loss, and compared with existing methods, the capture time and detection probability are significantly improved.
Oral Session: COMM Session 03 (Paper No.6475)
Title: Research on Anti-Jamming Dynamic Routing and Link Scheduling Algorithms for Heterogeneous
Author: Jing Cao1, Baowang Lian1 and Xiaoqin Xue2
Abstract: Abstract—The Internet of Things composed of low-Earthorbit
satellites, high-altitude platform stations, UAV relays,
and ground tactical networks is typically characterized by wide
communication coverage, heterogeneous network nodes, and
dynamically varying network topology. When external UAVs
deploy electromagnetic suppression or localized link jamming,
communication links are further degraded by node mobility,
multipath fading, and transient link disruptions. Traditional
routing protocols relying on static topology information or
shortest-path criteria suffer from escalating end-to-end delay,
declining packet delivery ratio, and frequent route
reconfiguration under such conditions. To address these
challenges, this paper incorporates link quality indicators,
channel availability probability, predicted link expiration time,
navigation positioning uncertainty, and jamming risk into the
routing decision state space, and constructs a four-layer
heterogeneous network model consisting of LEO satellites,
HAPS, UAV relays, and ground mesh nodes. A reward
function jointly optimizing end-to-end delay, link reliability,
hop count, and scheduling conflicts is designed. On this basis, a
Navigation-aided Graph Proximal Policy Optimization with
Link Scheduling algorithm, termed Navi-GPPO-LS, is
proposed. The algorithm introduces anti-jamming navigation
positioning and timing information into the reinforcement
learning state space, extracts dynamic topology features
through graph feature aggregation, and employs jamming risk
prediction and link scheduling conflict penalties to achieve
cooperative interference avoidance for multiple traffic flows.
Simulation results with LQI-Dijkstra, prediction-aided Qrouting,
TD3-like scheduling, and conventional GNN-PPO as
baselines demonstrate that Navi-GPPO-LS achieves superior
comprehensive performance in terms of packet delivery ratio,
end-to-end delay, throughput, outage ratio, and scheduling
conflict control. This work provides algorithmic reference for
enhancing the anti-jamming communication capability of
ground unmanned systems in UAV and counter-UAV
communication countermeasure scenarios.
Oral Session: COMM Session 03 (Paper No.6478)
Title: Full-Link Simulation of Physical Layer for LEO Satellites
Author: Xinrui Hu1, Wenjin Hou2, Hong Wen3, Ruixiang Yao3, Wendi Ma3 and Yuhui Wu3
Abstract: Abstract-With the advantages of low latency, wide coverage and large bandwidth, low-orbit satellite communication has become a core support for the space-earth integrated information network, and is widely applied in remote area communication, aviation navigation, emergency disaster relief and other fields. High dynamic Doppler frequency shift caused by high-speed movement of low-orbit satellites and additive white Gaussian noise interference seriously affect the reliable transmission of LEO satellite signals, and existing simulation platforms are difficult to achieve accurate full-process simulation. Therefore, this paper constructs a full-link simulation platform for LEO satellite signal physical layer based on MATLAB, which completely covers signal generation at the transmitter, high dynamic channel transmission, as well as synchronization, equalization and decoding processing at the receiver. Experimental results show that QPSK modulation has outstanding robustness with bit error rate reaching zero at 7 dB; 16QAM modulation possesses higher spectral efficiency and achieves zero bit error rate at 11 dB, and LDPC coding brings remarkable error correction gain. This platform can provide effective support for algorithm verification and performance evaluation of LEO satellite physical layer.
Oral Session: COMM Session 03 (Paper No.6480)
Title: LDPC Coding Adaptation and Anti-Jamming Optimization for LEO Satellite Air-Ground Links
Author: Xiangwei Ren1, Xiayu Chen1, Wenjin Hou2 and Haowen Deng1
Abstract: Abstract-To support low-Earth-orbit (LEO) satellite communications in space-based broadband coverage, ubiquitous information access, and highly reliable dedicated communications, physical-layer coding design for different mission scenarios has become a key factor affecting air-ground link performance. Given the significant differences between commercial broadband and dedicate LEO systems in throughput, reliability, and anti-jamming capability, this paper focuses on the uplink and downlink air-ground links between user terminals and satellites, and analyzes their differences in LDPC code construction, parameter configuration, and encoding/decoding mechanisms. A unified link simulation platform is built to quantitatively compare bit error rate, block error rate, coding gain, throughput, and decoding complexity under additive white Gaussian noise, impulsive interference, and broadband intentional jamming. Results show that the commercial LDPC code emphasizes high spectral efficiency and low complexity, achieving over 35% higher throughput than the dedicated code under ideal channels, whereas the dedicated LDPC code emphasizes reliable transmission under low-SNR and strong-interference conditions, improving coding gain by 1.2-2.4 dB and reducing block error rate by one to two orders of magnitude. The findings provide a reference for air-ground link coding optimization, anti-jamming transmission design, and heterogeneous-constellation compatibility.
Oral Session: COMM Session 03 (Paper No.6481)
Title: Routing Optimization Models and Enhanced Intelligent Algorithms for Inter-Satellite Laser Links in LEO Mega-Constellations
Author: Xiangwei Ren1, Wenjin Hou2, Wendi Ma3 and Xingyun Wei1
Abstract: Abstract-To address the highly dynamic topology of LEO mega-constellations and the high ATP overhead and frequent handovers of laser inter-satellite links (LISLs)-which cause routing reconstruction, traffic imbalance, and latency fluctuations-this paper establishes a multi-objective routing model constrained by constellation scale and laser physics. The model jointly optimizes end-to-end latency, link reliability, node load balancing, and handover frequency, while incorporating constraints such as laser terminal count, line-of-sight connectivity, minimum link survival time, and handover cost, thereby overcoming conventional algorithms' weak coupling with laser link characteristics. We propose an Improved Non-dominated Sorting Genetic Algorithm (INSGA-III) for optimal routing decisions. Co-simulation on a 24×66 LEO constellation using STK 11.0 and MATLAB R2022b shows that, compared with Dijkstra, MOPSO, and standard NSGA-III, the proposed method reduces average latency by 19.7%, improves reliability by 13.2%, optimizes the load-balancing factor by 28.4%, and cuts handovers by 32.6%, with faster convergence and a more uniform Pareto front.
COMM Session 04
Time: 14:00~15:30, Sunday, July 19, 2026
Session Chair: Dongyuan Shi, Northwestern Polytechnical University
  
Oral Session: COMM Session 04 (Paper No.6320)
Title: UWB Non-Contact Respiratory and Heart Rate Monitoring System for 6G ISAC
Author: Yuanhui Cao1, Xingguang Geng1, Fei Yao1, Shuaishuai Hou1, Yitao Zhang1 and Yunfeng Wang1
Abstract: Monitoring of respiratory signal and heartbeat signal provides essential insights into an individual's physiological and psychological conditions. However, current methods involving direct contact present challenges including user discomfort and compromised accuracy. This study proposes an advanced non-contact respiratory and heart rate monitoring system based on channel impulse response (CIR) using ultra-wideband (UWB) radar technology operating at 6.5 GHz. The system employs a joint signal processing algorithm that integrates time-domain coherent accumulation (TDCA) and variational mode decomposition (VMD), followed by multi-algorithm fusion for high-precision estimation of vital sign frequencies. Experimental results demonstrate that the root mean square error (RMSE) between the estimated respiratory rate and heart rate and those measured by a polygraph reference device are 2.98% and 1.54%, respectively. This approach thus demonstrates significant potential as a reliable, non-invasive method for accurate detection of physiological vital signs in both clinical and home environments, which may provide reference value for future Integrated Sensing and Communication (ISAC) technology in 6G.
Oral Session: COMM Session 04 (Paper No.6344)
Title: STAR-RIS and RSMA for ISAC: A Joint Optimization Framework
Author: Zhentao Wang1, Wenbin Sun1, Xin Yang1, Lili Chen1, Qian Xu1 and Lin Wang1
Abstract: This paper proposes a novel integrated sensing and communication system assisted by a simultaneous transmitting and reflecting reconfigurable intelligent surface. The framework addresses challenges in serving users with diverse channel conditions while performing sensing tasks in obstructed environments. By incorporating rate-splitting multiple access at the base station, the system manages multi-user interference, improving communication rates and user fairness. The STARRIS dynamically reshapes the wireless environment through its transmission and reflection capabilities, creating optimized propagation paths for both functionalities. We establish a comprehensive system model and formulate a joint optimization problem for active precoding, passive beamforming, and rate-splitting parameters to maximize the weighted sum-rate. An efficient alternating optimization algorithm using fractional programming and semi-definite relaxation is developed to solve this complex problem. Numerical results demonstrate the scheme's substantial improvements in spectral efficiency and sensing accuracy over conventional systems, achieving notable performance gains under
practical constraints.
Oral Session: COMM Session 04 (Paper No.6345)
Title: Integrated Communication and Jamming Waveform Design for OTFS-Based Comb-Spectrum Jamming with Spatial Index Modulation
Author: Tianyu Kou1, Wen-Bin Sun1, Qian Xu1, Zhaolin Zhang1 and Ling Wang1
Abstract: Abstract-We propose a novel integrated communication and jamming (ICAJ) waveform utilizing a Multiple-Input Single-Output (MISO) Orthogonal Time Frequency Space (OTFS) framework combined with Spatial Index Modulation (SIM). By exploiting the delay-Doppler (DD) domain sparsity of wideband comb-spectrum jamming, communication data are orthogonally embedded into low-energy valleys. To mitigate the inherent capacity-covertness trade-off, spatial antenna indexing conveys additional information bits without increasing communication power allocation. We derive closed-form expressions for the effective signal-to-interference-plus-noise ratio (SINR), system throughput, and computational complexity under imperfect channel state information (CSI). Simulations demonstrate that the MISO-SIM architecture significantly enhances throughput compared to the Single-Input Single-Output (SISO) baseline, while strictly preserving macroscopic waveform covertness in terms of time-domain Pearson Correlation Coefficient (PCC) and in-band Power Spectral Density (PSD) metrics.
Oral Session: COMM Session 04 (Paper No.6346)
Title: A 140GHz Array Antenna Based on 4-Bit RF-MEMS Phase Shifters
Author: Haozhe Hou1, Jianming Huang1, Naibo Zhang1, Zilai Wang1, Yansong Cui1, Weizheng Ren1, Yiran Zhang1, Xinyue Sun1 and Guangbin Dong2
Abstract: This paper presents a 140 GHz phased array antenna utilizing a 4-bit distributed micro-electro-mechanical systems (MEMS) transmission line (DMTL) phase shifter. The antenna element employs a Wilkinson power divider with a 90° phase difference to feed adjacent edges of a square patch, achieving circular polarization. Sector-shaped cut corners and rectangular stubs are introduced to widen the axial ratio and impedance bandwidths, respectively. The phase shifter adopts a cascaded architecture of 16 unit cells, each loaded with a MEMS bridge featuring metal-air-metal (MAM) contacts to enhance phase shift per unit length. Simulation results show the element operates from 133 to 150 GHz with S11 < -15 dB and AR < 3 dB. The 4-bit phase shifter provides a 360° phase range with <4 dB insertion loss and <5° phase error. The 2×8 array achieves beam steering from -45° to +45° with a peak gain >13 dB, validating a low-loss beamforming approach for terahertz applications.
Oral Session: COMM Session 04 (Paper No.6371)
Title: An analysis of diurnal VLF propagation field versus local time using modified IRI data
Author: Junke Wang1, Zhewen Chen1, Shitian Zhang2, Kuisong Zheng1, Songming Zou1 and Qiang Wu1
Abstract: To precisely obtain the VLF propagation field during 24 hours in the Earth-ionosphere waveguide, an EM-FDTD method of combing the FDTD and the electron momentum equation is developed. The absent data of electron density in the lower height region is supplemented with the log-linear fitting algorithm. According to the Yee cell, the iterative formulas of the electron momentum equation are deduced in detail. The long-distance propagation field from VTX to Gwalior are calculated. By comparing the numerical results and the measured results, a good agreement is reached. It is shown that the proposed method is validated and corrected.
COMM Session 05
Time: 16:00~17:30, Sunday, July 19, 2026
Session Chair: Dongyuan Shi, Northwestern Polytechnical University
  
Oral Session: COMM Session 05 (Paper No.6366)
Title: A Network Traffic Prediction Method Based on STL-Informer-DLinear
Author: Yuze Su1, Haifeng Yang1 and Yong Wang1
Abstract: Network traffic prediction is a key technology for efficient network management and optimization. To improve the accuracy of network traffic prediction, a network traffic prediction method based on STL-Informer-DLinear is proposed. Firstly, according to the trend, periodicity and randomness characteristics of network traffic, the Seasonal and Trend Decomposition using Loess (STL) algorithm is introduced to decompose the network traffic, obtaining trend, seasonal, and remainder components. Then, the Informer model is used to predict the trend and remainder components, while the DLinear model is used to predict the seasonal component. Finally, the predicted results of all components are combined to obtain the final traffic prediction results. Experimental results show that the proposed STL-Informer-DLinear method achieves better prediction performance than other traffic prediction methods on real network traffic datasets. Specifically, it achieves a 61.5%-96.3% reduction in MSE, a 21.3%-80.7% reduction in RMSE, a 31.2%-83.6% reduction in MAE, and a 26%-89.1% reduction in MAPE, which verifies the excellent performance of the STL-Informer-DLinear network traffic prediction method.
Oral Session: COMM Session 05 (Paper No.6414)
Title: A High-speed Visible Light Coherent Communication System based on Balanced Conjugate Framing
Author: Zijian Zhou1, Wenting Ju1, Yuhan Hu1, Zengyi Xu1 and Nan Chi1
Abstract: The rapid growth in data traffic in today's information age places high demands on throughput of communication systems. Against this background, visible light communication (VLC) systems have attracted great attention for their ample spectrum resources, and strong resistance to electromagnetic interference. However, the performance of intensity-modulation direct-detection (IM/DD) based VLC systems is restricted by limited device bandwidth and channel nonlinearity. To address this issue, this paper proposes a balanced conjugate framing (BCF) scheme for high-speed VLC system based on self-homodyne coherent detection. The proposed scheme offers a new frame structure, which contains the original signal block and its conjugated counterpart. Subtraction is applied to the two signal blocks in the received signal to eliminate the common-mode noise. Experimental results show that the proposed BCF scheme can effectively improve the received signal quality and achieve 12 Gb/s visible light coherent transmission. This scheme provides a new perspective for mitigating noise in visible light coherent communication (VLCC) systems at the coding level.
Oral Session: COMM Session 05 (Paper No.6421)
Title: On the SOR Sampling for Discrete Symbol Detection: An Energy-Based Model Perspective
Author: Zhiheng Zhang1, Le Yang2 and Jun Tao3
Abstract: This paper revisits the successive over-relaxation (SOR) sampling for discrete symbol detection from the perspective of energy-based models (EBMs). We first derive the complex-valued realization of the discrete unadjusted Langevin algorithm (DULA) using the Wirtinger calculus. This allows us to analytically establish the equivalence between the SOR sampler and symbol-wise DULA with a certain stepsize schedule. We further show that the annealed SOR (ASOR) sampling, a recently proposed enhancement of the SOR sampling, can be viewed as the simulated annealing (SA) DULA. Simulation experiments corroborate the theoretical developments.
Oral Session: COMM Session 05 (Paper No.6459)
Title: Channel Estimation for UAV Swarms Based on Space-Time Autoregressive Models
Author: wang pp1, Lian Baowang1, Guo Huachang2, Zhao Na2, Ma Xingbing2 and Hou Yu2
Abstract: To address the challenges posed by the high mobility and dynamic topology of unmanned aerial vehicle (UAV) swarms, which result in complex and highly variable wireless channels, and the difficulty of model-based channel estimation methods in accurately describing actual channel fading characteristics, this paper proposes a Space-Time Autoregressive (ST-AR) channel estimation method suitable for UAV swarm communication scenarios. First, an ST-AR state-space model is constructed based on the acquired channel space-time sequence features, second, the AR coefficients are updated online in an adaptive manner by combining real-time channel autocorrelation functions with relative geometric topology; subsequently, a model goodness-of-fit test based on prediction residuals is introduced to achieve model mismatch detection and order adaptation; finally, the calibrated ST-AR model is used for channel prediction and recursive estimation. Simulation results demonstrate that, in a typical UAV swarm scenario, the proposed method achieves an end-to-end bit error rate (BER) improvement of approximately 2 dB compared to the MMSE method while reducing pilot overhead by 50%, validating the method's effectiveness and robustness in high-dynamic, low-overhead scenarios.
Oral Session: COMM Session 05 (Paper No.6476)
Title: A WFRFT Secure Communication Method Based on Chaos-Driven Constellation Phase and Polarity Flipping Hybrid Encryption
Author: Jiayue Li1, Qingwei Meng1, Yanzhi Yun1, Han Wang1, Dan Wang1, Liangsi Zhou1 and Linghua Su1
Abstract: To address the issues of limited key space and insufficient constellation concealment in physical layer encryption, a Weighted Fractional Fourier Transform (WFRFT) secure communication method based on chaos-driven constellation phase and polarity flipping hybrid encryption is proposed. This method utilizes a four-dimensional (4D) hyperchaotic system to generate four chaotic sequences, which respectively drive row-column scrambling, dynamic phase shift, and polarity flipping. Finally, the WFRFT is applied to achieve constellation confusion and diffusion, rendering the encrypted signal with Gaussian noise-like characteristics. Computer simulation results demonstrate that while ensuring communication reliability for legitimate users, the proposed method significantly enhances system security. Under a fixed WFRFT order, unauthorized users cannot accurately decrypt the encrypted information even if the key error is merely at the level. The system bit error rate (BER) approach.
Oral Session: COMM Session 05 (Paper No.6477)
Title: Enhanced Cross-Representation Domain Black-Box Transfer Attack Method for Automatic Modulation Recognition
Author: xing Yun Wei1, Xiangwei Ren1, Wenjin Hou2, Hong Wen3, Ruonan Jing3 and Wenqi Tang3
Abstract: Abstract-As a core technology in cognitive radio and spectrum supervision, Automatic Modulation Recognition (AMR) is facing severe security challenges arising from the inherent vulnerability of deep learning models. Existing adversarial attack studies are mostly limited to a single representation domain, and the transferability of adversarial samples degrades significantly when confronted with "representation domain mismatch" in black-box scenarios. To this end, this paper proposes an attack enhancement method based on Multi-Representation Loss Ensemble (MRLE). By constructing a unified time-domain physical perturbation space and introducing a differentiable Fourier transform layer, the method establishes a gradient backpropagation path between heterogeneous representations, achieving deep coupling of time-frequency dual-domain gradients and end-to-end collaborative optimization under physical constraints. Experimental results demonstrate that the proposed method can significantly improve the success rate of cross-representation domain transfer attacks, effectively overcome the transfer asymmetry bottleneck, and control the adversarial distortion. It provides an important reference for the security evaluation of AMR systems in black-box scenarios.
SPSS Session 01
Time: 14:00~15:30, Sunday, July 19, 2026
Session Chair: Yi Yu, Southwest University of Science and Technology
  
Oral Session: SPSS Session 01 (Paper No.6375)
Title: Performance Assessment of Galileo HAS PPP Time Transfer
Author: Xiaofei Liang1
Abstract: With the rapid evolution of GNSS Precise Point Positioning (PPP) technology, high-precision remote time transfer using real-time State-Space Representative (SSR) corrections has become a critical research focus. This study investigates the comprehensive performance assessment of the PPP time transfer model utilizing Galileo High Accuracy Service (HAS) real-time corrections.Firstly, an ionosphere-free (IF) PPP time transfer mathematical model based on Galileo HAS was constructed. Secondly,Using global IGS station data connected to external Hydrogen Masers, systematic experiments were conducted across ultra-short, short, and intercontinental long baselines. Finally, Taking post-processed CODE products as the reference, the system was assessed using Root Mean Square Error (RMSE) for external accuracy and Modified Allan deviation (MDEV) for frequency stability. Results show that Galileo HAS-based time transfer is generally consistent with post-processed precision series. The average external RMS across all links is better than 0.3 ns, reaching 0.088 ns for ultra-short baselines. In terms of frequency stability, the average performance at $10,000$ seconds is better than the $1 \times 10^{-14}$ for all links, effectively approaching the level of post-processed products in long-baseline scenarios.The study indicates that Galileo HAS is fully capable of supporting global-scale real-time precision time transfer. This provides a robust theoretical and technical foundation for building low-cost, high-efficiency, and reliable global real-time time transfer systems.
Oral Session: SPSS Session 01 (Paper No.6389)
Title: An Autonomous Rescue USV System with Hierarchical Search and Exact-Penalty Planning
Author: Longyun Yuan1, Chengkai Tang1, Lingling Zhang1, Yangyang Liu1, Zesheng Dan1 and Ding Yuan1
Abstract: Autonomous man-overboard rescue on water surfaces with Unmanned surface vehicles requires reliable long-range target acquisition, robust close-range identification, and safe motion planning under wave-induced disturbances and dynamic obstacles. This paper presents a closed-loop autonomous USV system for search-and-rescue that integrates a radio-based long-range search with near-range multimodal target confirmation and a collision-aware exact-penalty trajectory planner. In the long-range stage, a VHF radio transceiver returns an ID and coarse target information, enabling the USV to navigate towards the target. At closer range, the system switches to sensing using vision, thermal imagery, and mmWave radar to produce refined target state for the planner. The exact-penalty planner incorporates obstacle-avoidance constraints as penalty terms and balances route directness and safety precautions. Field experiments were conducted in both lake and marine environments. Radio-guided acquisition together with near-range recognition achieved <2m positioning precision on most samples, and the repeated tests proved that approach precision reached 2m in real marine scenario. Compared with A*, DWA, and ACO, the proposed planner generates smoother trajectories, provides larger safety margins, and achieves higher path efficiency in mixed static and dynamic obstacle scenarios.
Oral Session: SPSS Session 01 (Paper No.6415)
Title: Multi-Sensor Fusion Positioning Method for UAV Swarms with Improved CPInformer
Author: Ruihan Shen1, Lingling Zhang1, Siyuan Luo2, Yexuan Bai1, Yiheng Chen1 and Zengrui Zhou1
Abstract: With the rapid development of the low-altitude economy, UAV swarms have become a core application platform, and UAV positioning technology is the foundation supporting swarm applications. To address the issues of poor stability and difficult scalability in existing combined positioning methods, this paper proposes a multi-sensor fusion positioning method for UAV swarms based on an improved Cooperative Positioning Informer(CPInformer). This method first converts distributed heterogeneous navigation sources into an information probability model, achieving a unified format for navigation information parameters; it then designs a lightweight improved CPInformer neural network to complete the fusion positioning of the UAV swarm. Simulation analysis under both normal and high-noise conditions, along with validation on an actual UAV swarm, shows that this method can effectively improve the accuracy and stability of UAV swarm fusion positioning.
Oral Session: SPSS Session 01 (Paper No.6425)
Title: UWB Based 3D High Precision Cooperative Positioning System Integrating Information Geometry
Author: Weiming Chen1 and Yangyang Liu2
Abstract: Ultra-wideband technology has become the mainstream solution to achieve indoor high-precision positioning by virtue of nanosecond narrow pulses and centimeter-level ranging accuracy. However, multipath interference, non-line-of-sight (NLOS) propagation and dynamic noise will significantly reduce its positioning accuracy. In this paper, a three-dimensional high-precision cooperative positioning system based on UWB is proposed, which adopts a five-base station redundant topology. In order to further improve the positioning optimization performance, this paper proposes a multi-base station information fusion algorithm based on information geometry, and develops a NLOS detection model based on residual analysis and a clock synchronization error correction mechanism using geometric distance back projection. The simulation results show that the NLOS detection accuracy is more than 75% under both motion conditions. The results show that after the information geometry combined with the extended Kalman filter fusion algorithm proposed in this paper, the average positioning error under the simple motion model is reduced to 0.113m, and the positioning error under the complex motion model is reduced to 0.173m, which is 76.7% and 76.6% higher than the independent Kalman filter, respectively.
Oral Session: SPSS Session 01 (Paper No.6436)
Title: Fingerprint-Based Visible Light Positioning Using Improved WKNN Algorithm With Selected Reference Position
Author: Zijian Bai1, Zhonghua Liang1, Wenqian Jiang1, Dongxin Bai1 and Zhuo Zhou1
Abstract: In indoor visible light positioning (VLP) scenarios, fingerprint-based localization has garnered considerable attention due to its high accuracy and broad applicability. Most existing fingerprint-based VLP systems directly utilize received signal strength (RSS) values from a fixed number of reference nodes (RNs) and employ the weighted K-nearest neighbor (WKNN) algorithm for positioning. Such approaches failed to take into account the region-specific requirements for the number and selection of RNs across different spatial areas, thereby degrading the overall localization accuracy. To address this issue, this paper leverages the relationship between RSS and physical distance to develop an improved weighted K-nearest neighbor algorithm based on selected RN position information (SRP-WKNN). In the proposed algorithm, the nearest RNs are adaptively selected by considering both signal similarity and spatial proximity, with the selection strategy differing for edge and central regions. These selected RNs are then used to derive the mapping between RSS similarity and physical distance. Finally, the coordinates of the selected RNs are weighted to obtain the precise location estimate. Software simulation results show that, at a signal-to-noise ratio (SNR) of 20 dB and an RN interval of 0.25 m, the average positioning error (APE) is significantly reduced by at least 11.0\% and at most 41.6\% compared with traditional WKNN and its representative improved versions.
Oral Session: SPSS Session 01 (Paper No.6451)
Title: A Loosely Integrated GBNS/INS Navigation Method for Weak Vertical Geometry Scenarios
Author: Chengyan He1, Jiawei Liu1, Mingliang Tao1, Zhaolin Zhang1 and Ling Wang1
Abstract: Ground-Based Navigation Systems (GBNS) offer reliable positioning in Global Navigation Satellite System (GNSS)-degraded or GNSS-denied environments. However, due to the near-ground deployment of GBNS transmitters, the user-transmitter geometry often yields strong horizontal constraints but extremely weak vertical observability, leading to unreliable height estimates in GBNS standalone positioning. When such three-dimensional GBNS solutions are directly fed into a loosely coupled GBNS/inertial navigation system (INS) integration filter, the contaminated height information inevitably degrades the inertial error correction and compromises overall navigation stability. This paper proposes a height-constrained loosely coupled GBNS/INS integration method for weak vertical geometry. Unlike conventional approaches that feed unconstrained GBNS positions into the Extended Kalman Filter (EKF), our method incorporates short-term INS-predicted height as a constraint during GBNS positioning. The height-constrained GBNS solution-preserving horizontal geometry while suppressing vertical uncertainty-is then used as the EKF measurement input. Simulations under typical GBNS deployment (good HDOP, poor VDOP) show that, compared with conventional loosely coupled integration, the proposed method reduces height 1σ error from 42.21?m to 27.49?m, while maintaining comparable horizontal accuracy (0.16?m north, 0.32?m east vs. 0.14?m and 0.33?m). These results demonstrate that the proposed INS-height-constrained scheme significantly improves vertical stability without sacrificing the horizontal correction benefit of GBNS. The method provides a practical and effective solution for reliable GBNS/INS integration in weak height geometry scenarios, with potential extensions to other ground-based ranging systems.
SPSS Session 02
Time: 16:00~17:30, Sunday, July 19, 2026
Session Chair: Zhengqiao Zhao, Northwestern Polytechnical University
  
Oral Session: SPSS Session 02 (Paper No.6306)
Title: An Improved YOLOv11-Based Model for Object Detection in Sonar Images
Author: Zeng Li Liu1 and Chao Yang Li1
Abstract: Abstract: Underwater target detection serves as a pivotal technology for domains such as marine exploration, national defense, and underwater engineering. Leveraging its long-range detection capabilities, sonar imaging has emerged as a critical technological modality. However, sonar imagery is intrinsically characterized by severe speckle noise, low contrast, and blurred target edges, which impose significant challenges on deep learning-based detection algorithms. To address these issues, this paper proposes YOLOCS, an enhanced architecture based on YOLOv11. First, a C3CFB module is introduced to strengthen the extraction of multi-scale and multi-directional features from sonar targets. Subsequently, Space-to-Depth Convolution (SPDConv) is incorporated for downsampling to mitigate performance degradation caused by the loss of critical features. Experimental results on the SCTD dataset demonstrate that, compared with the baseline YOLOv11n, YOLOCS achieves improvements of 0.5%, 2.9%, and 2.6% in Precision, Recall, and mAP50-95, respectively.
Oral Session: SPSS Session 02 (Paper No.6315)
Title: A Method for Depth Discrimination of Shallow Water Sound Sources Based on Dynamic Threshold Decision for Broadband Phase Fluctuation
Author: Zhejian Hu1 and Xuanjie Wei1
Abstract: In shallow water environments, the interference effect of normal modes results in unique broadband phase fluctuations in the acoustic field. Under typical negative gradient hydrographic conditions, by dividing the vertical array into upper and lower arrays for summation and cross spectral analysis, the source depth exhibits distinct separability. Since this phase fluctuation characteristic correlates with horizontal distance, fixed decision thresholds struggle to achieve target depth resolution across varying horizontal distances. Therefore, this study improves upon fixed decision thresholds by introducing a horizontal dynamic threshold for phase fluctuation feature discrimination, enabling source depth resolution at different distances. Based on simulation data analysis, the depth resolution performance of the horizontal dynamic threshold selection method is evaluated.
Oral Session: SPSS Session 02 (Paper No.6372)
Title: An Acoustic-Based Side-Scan Sonar Image Simulation Method Incorporating AUV Attitude Parameters
Author: Zizhuo Liang1, Jianfeng Chen1, Fen Liu1, Yifan Zhang1, Hongyu Wei1 and Yaohui Wen1
Abstract: Recent advances in underwater robotics have accelerated the development of automatic underwater target detection. However, due to environmental complexity and cost constraints, it remains difficult to obtain sufficient underwater target images, limiting the performance of detection algorithms. To address this issue, simulation-based methods provide an effective alternative. In this work, we propose a Side-Scan Sonar (SSS) image simulation method based on underwater acoustic propagation modeling. The proposed framework simulates the complete SSS imaging process, including transmit waveform generation, acoustic propagation, target backscattering, echo reception, and array directivity, enabling SSS image simulation under various parameter settings. Meanwhile, the depth and velocity of the Autonomous Underwater Vehicle (AUV), particularly its attitude variations during navigation, are incorporated to reproduce SSS image distortions, enabling more realistic sonar images. To validate the effectiveness of the proposed method, cross-domain transfer experiments are conducted. Results show that models pre-trained on the simulated dataset consistently outperform those pre-trained on COCO in real-world tasks, achieving improvements of 4.7% in mAP@0.5, 8.6% in mAP@0.5:0.95, and 6.6% in Recall.
Oral Session: SPSS Session 02 (Paper No.6409)
Title: Simulation of Elastic and Geometric Acoustic Scattering from Typical Underwater Shell Targets
Author: Zihao Shu1, Jianjun Zhu1, Peihong Wang1, Tian Zhou1, B.A Tarasov2 and V. I. KOROCHENTSEV3
Abstract: The acoustic scattering from spherical and hemispherical-capped cylindrical shells is investigated through numerical simulations of steel and aluminum shells, followed by analysis of their acoustic scattering fields. The broadband acoustic fingerprints and target strength curves reveal how geometry, size, material, and incident angle dominate the spatial texture, orientation, and frequency-domain characteristics of elastic scattering regions. For spherical shells, increasing size produces denser interference fringes and richer modal structure at low-to-mid frequencies. For cylindrical shells, the acoustic fingerprint shows strong azimuthal dependence, enabling shape and orientation identification. For same-sized targets of different materials, backscattering target strength differs more significantly at low-frequency bands but converges at high-frequency bands. These acoustic fingerprints and scattering field patterns provide physical insights and fundamental data for feature extraction and intelligent classification of small underwater targets.
Oral Session: SPSS Session 02 (Paper No.6448)
Title: A Propagation Acoustic Path Error Correction Method for the Overlapping Phase Center Algorithm in High-Frequency Synthetic Aperture Sonar
Author: Guijuan Han1, Xiangrui Zeng1 and Yuhang Gao2
Abstract: When utilizing the Displaced Phase Center (DPC) algorithm for motion error estimation in high-frequency, large-aperture synthetic aperture sonar (SAS) systems, the propagation acoustic path error of the displaced phase center pairs and the correlation peak shift induced by sonar motion errors are of a comparable order of magnitude. This comparability leads to a significant degradation in the motion error estimation performance of the DPC algorithm. To address this issue, this paper analyzes the impact of the equivalent phase center approximation on the DPC algorithm and proposes a propagation acoustic path difference correction method for displaced phase center pairs in dual-transmitter, multi-receiver SAS systems. By employing a dual-transmitter architecture and optimizing the spatial layout of the transceiver array and the pulse repetition interval (PRI), the proposed method ensures that the propagation acoustic path lengths of displaced phase center pairs across adjacent pings are strictly identical. Consequently, the proposed approach effectively eliminates the adverse effects of the equivalent phase center approximation on DPC-based motion error estimation. Simulation results validate the efficacy of the proposed method.
Oral Session: SPSS Session 02 (Paper No.6465)
Title: A Novel Method for Improving the Efficiency of CFAR Detection of Small Targets in Side-scan Sonar Images
Author: Qiuju Li1, Yifan Wu1, Peng Xiao2 and Gang Xiao2,3
Abstract: Abstract-To address the issue of low efficiency in detecting small targets in side-scan sonar images using CFAR (Constant False Alarm Rate) detection, this paper proposes a novel method to improve the efficiency of CFAR detection for small targets in side-scan sonar images. This method employs morphological reconstruction to generate an indexed sonar image, which guides the subsequent CFAR detection algorithm and improves its efficiency in detecting small targets. Experimental results demonstrate that this method can effectively improve the efficiency of CFAR detection for small targets in side-scan sonar images.
SPSS Session 03
Time: 08:30~10:00, Monday, July 20, 2026
Session Chair: Zhongxin Bai, Harbin Engineering University
  
Oral Session: SPSS Session 03 (Paper No.6321)
Title: Physics-Guided Side-Scan Sonar Image Simulation for Enhanced Object Detection
Author: Guoqing Xie1, Ju He1, Haoran Hu1, Jinpeng Xu1, Hu Xu1 and Yang Yu1
Abstract: Side-scan sonar (SSS) plays an important role in underwater perception for applications such as seabed mapping, marine inspection, and object detection. However, the performance of deep learning-based detection methods is often limited by the scarcity of annotated SSS datasets, as acquiring large-scale underwater sonar data is expensive and time-consuming. To address this challenge, this paper proposes a physics-aware side-scan sonar image simulation framework for data augmentation in object detection tasks. The proposed method integrates geometry-driven scene construction and acoustic-aware echo modeling to generate realistic SSS images. Specifically, a geometry-driven mechanism is designed to construct target structures and their corresponding acoustic shadows based on the geometric relationship between sonar beams and object shapes. Experimental results demonstrate that incorporating the proposed synthetic data significantly improves detection performance compared with training using real data alone. The proposed framework provides an efficient solution for alleviating the data scarcity problem in underwater sonar perception tasks.
Oral Session: SPSS Session 03 (Paper No.6362)
Title: LAT: A Lightweight Anti-occlusion Tracker for Hyperspectral Videos
Author: Xu Li1, Fuyuan Ge1, Qing Zhang1, Baoguo Wei1, Zhendong Li1 and Junyin Yu1
Abstract: Hyperspectral object tracking (HOT) has attracted increasing attention due to its powerful discriminative capabilities derived from rich spectral information in challenging scenarios. However, existing HOT methods often face two critical limitations: poor tracking robustness against occlusions and low computational efficiency caused by high-dimensional spectral data and complex architectures. To address these limitations, we propose a Lightweight Anti-occlusion Tracker (LAT) for hyperspectral videos. Firstly, we design a discriminative band selection (DBS) module based on the minimum redundancy maximum relevance (mRMR) criterion, and select few spectral bands to reduce redundant computations while maintaining object-background separability. Secondly, we introduce an occlusion-aware branch (OB) that fuses local fine-grained features and global statistical features to accurately predict the degree of occlusion. Thirdly, we adopt a lightweight backbone and design a Kalman filtering-based motion prediction (MP) module to form a hierarchical anti-occlusion tracking strategy. The extensive experiments conducted on the HOT2024 dataset demonstrate that LAT has advantages in both overall tracking accuracy and anti-occlusion tracking accuracy. In addition, LAT only contains 0.69 million parameters, achieving an excellent balance between tracking accuracy and model complexity.
Oral Session: SPSS Session 03 (Paper No.6363)
Title: Hyperspectral Small Object Tracking Method With Target Region Perception
Author: Xu Li1, Qing Zhang1, Yuchao Wang1, Fuyuan Ge1, Zhendong Li1 and Baoguo Wei1
Abstract: Hyperspectral Small Object Tracking holds significant application value in fields such as remote sensing and autonomous driving. However, challenges such as low resolution of small object, complex background interference, and redundant hyperspectral data hinder the performance of existing tracking methods. This study introduces a novel hyperspectral small object tracking method based on target region perception (HTRP), which enhances tracking accuracy and robustness. Firstly, a Band Grouping Module (BGM) is designed to calculate band contributions, reorder and group hyperspectral bands, and then generate multiple sets of false-color images in order to reduce band redundancy and preserve discriminative spectral information. Secondly, a Target Region Perception Module (TRPM) is proposed to generate a saliency map by enhancing the resolution of the object area, thereby improving the network's ability to perceive the positions of small object and producing region of interest to guide tracking. Finally, numerous experiments show that the HTRP tracker maintains superior performance on the HOT2022 dataset.
Oral Session: SPSS Session 03 (Paper No.6379)
Title: SAR-HyperNet: Spectral-Aware Residual Hyperprior Network for Hyperspectral Image Compression
Author: Fahad Saeed1, Liu Shumin1, Jie Chen1 and Muhammad Salman Khan1
Abstract: Hyperspectral images (HSIs) provide rich spectral information for remote sensing applications, but their high dimensionality creates substantial storage and transmission burdens. Existing deep learning-based compression methods achieve promising rate-distortion performance, yet they often insufficiently model spectral dependencies and adaptive redundancy reduction. To address this issue, we propose \textbf{Spectral-Aware Residual Hyperprior Network (SAR-HyperNet)} for lossy HSI compression. SAR-HyperNet integrates \textit{Hybrid Residual Blocks (HRBs)} and \textit{Spectral Attention Blocks (SABs)} into a hyperprior-based variational autoencoder to improve feature propagation, training stability, and spectral channel recalibration. This design enables compact latent representations while preserving spatial structure and spectral consistency. Experiments on public datasets show that SAR-HyperNet consistently outperforms representative methods. It achieves higher PSNR and MS-SSIM with lower SAM across various bitrates. Ablation results further verify the complementary contributions of HRBs and SABs, demonstrating the effectiveness of spectral-aware design for deep HSI compression.
Oral Session: SPSS Session 03 (Paper No.6380)
Title: Spectrally Constrained Feature Matching for UAV-Borne Hyperspectral Images
Author: Yuhe Liu1, Liu Shumin1, Yelin Liu2 and Jie Chen1
Abstract: Feature matching is essential for UAV-borne hyperspectral image registration, but most existing matchers rely mainly on grayscale or spatial texture information and underuse spectral signatures. This may lead to mismatches in regions with similar spatial structures but different material properties. To address this problem, this paper proposes a plug-and-play spectral consistency filtering method for hyperspectral feature matching. Initial correspondences are first obtained from grayscale representations using existing matchers, and local spectral descriptors are then extracted from the original hyperspectral data around matched keypoints. Spectral Angle Mapper (SAM) is used as the primary consistency measure to reject spectrally inconsistent matches, while Spectral Information Divergence (SID) is evaluated for comparison. The refined matches are finally used for RANSAC-based homography estimation. Experiments on the WHU-Hi-LongKou dataset show that the proposed spectral filtering strategy effectively improves correspondence quality, reduces geometric errors, and enhances the robustness of hyperspectral image matching.
Oral Session: SPSS Session 03 (Paper No.6385)
Title: Motion-Aware OSTrack: Latency Compensation for One-Stream Visual Tracking
Author: Shiduo Zhang1, Cunle Zhang1, Chengkai Tang1, Baowang Lian1 and Dongjia Wang2
Abstract: Online visual tracking on unmanned aerial vehicle (UAV) platforms is sensitive to inference latency, since the tracker output may correspond to an earlier processed frame rather than the target state at output time. Although one-stream Transformer trackers such as OSTrack achieve efficient target-oriented representation by jointly modeling template and search features, they do not explicitly compensate for latency-induced target displacement. This issue becomes more challenging in UAV scenarios, where apparent motion is affected by both target movement and platform-induced camera jitter. To address this problem, we propose a jitter-aware latency compensation method built upon OSTrack. The proposed framework keeps the original one-stream tracker unchanged and introduces two lightweight modules: a camera jitter compensation branch that estimates camera-induced shift from recent target states and produces stabilized motion history, and a temporal motion predictor that estimates the target-center displacement during the inference interval. The latest raw OSTrack output is then refined through a residual center update to obtain a latency-compensated prediction. Experiments are conducted on UAV123, UAVDT, and DTB70 under both offline evaluation and a fixed-latency online evaluation setting. The results demonstrate that the proposed method consistently improves online tracking performance with only limited computational overhead, indicating its effectiveness for latency-aware UAV tracking under challenging motion conditions.
SPGT Session 01
Time: 14:00~15:30, Saturday, July 18, 2026
Session Chair: Fan Zhang, Wuhan University
  
Oral Session: SPGT Session 01 (Paper No.6305)
Title: ARMANET: AUTO-REGRESSIVE-MOVING-AVERAGE INSPIRED NEURAL NETWORKS FOR SOURCE SEPARATION USING MICROPHONE ARRAYS
Author: Yucong Liu1, Chao Pan1, Zhuo Liu1, Jacob Benesty1 and Jingdong Chen1
Abstract: In sequential signal processing, moving-average (MA) and auto-regressive (AR) operations are often cascaded construct efficient systems. The AR stage emphasizes signals from specific frequency bands, while the MA block suppresses interferences from others. Motivated by this principle, we introduce the auto-regressive moving-average network (ARMAnet), a neural architecture that integrates AR and MA functionalities for sequential signal processing. AR processing is implemented via a recurrent neural network, specifically the recently developed Mamba network, where as MA processing is realized through a self-attention mechanism that computes weighted sums over the input sequence. To assess its effectiveness, we combine ARMAnet with the recently proposed SpatialNet and apply it to microphone array source separation. Experimental results demonstrate the great potential of ARMAnet for this task.
Oral Session: SPGT Session 01 (Paper No.6330)
Title: Predictive Step-size Selection Time-frequency-domain Hybrid Filtered-x Normalized Least Mean Square Algorithm for ANC Systems
Author: Zhiyuan Li1, Yi Yu1, Hongsen He1, Yuyu Zhu1 and Rodrigo C de Lamare1
Abstract: The filtered-x normalized least mean square (FxNLMS) algorithm is popular in active noise control systems, but its computational burden increases with the filter length. To enhance the computational efficiency, the frequency-domain FxNLMS (FD-FxNLMS) algorithm was also proposed. However, it still suffers from a trade-off between convergence rate and steady-state residual noise. Therefore, to overcome this problem, this paper proposes a predictive step-size selection strategy for the FD-FxNLMS algorithm based on a delayless time-frequency-domain structure, which adaptively selects an appropriate step-size according to the time-domain mean-square deviation recursion model. Simulation results demonstrate that the proposed algorithm achieves faster convergence and lower steady-state residual noise while preserving computational efficiency.
Oral Session: SPGT Session 01 (Paper No.6395)
Title: On the Maximum Likelihood-Based Wiener Postfilter in the Presence of Unmodeled Interference
Author: Fan Zhang1, Kang Chen1 and Gongping Huang1
Abstract: The cascaded structure of minimum variance distortionless response (MVDR) beamforming followed by single-channel Wiener postfiltering is widely used for microphone array speech enhancement. In practical implementations, interference sources are often not explicitly modeled due to the difficulty of direction estimation, resulting in model mismatch whose impact on the Wiener postfilter remains unclear.In this letter, we analyze the behavior of maximum likelihood (ML)-based Wiener postfilter in the presence of an unmodeled interferer. We first derive closed-form bias expressions for the ML estimators of the desired signal and noise variances. Based on this result, we derive the expected value of the estimated postfilter and analyze its monotonicity with respect to the spatial correlation coefficient (SCC) between sources, input signal-to-noise ratio (SNR), and interference-to-noise ratio (INR). Our analysis reveals that the postfilter gain is a monotonic function of both SNR and SCC, while its monotonicity with respect to INR depends on the SCC. Moreover, a condition on the SCC is established under which effective interference suppression is guaranteed. The theoretical findings are validated through numerical examples and speech enhancement experiments.
Oral Session: SPGT Session 01 (Paper No.6411)
Title: Cross-modal Speech Separation Based on Inter-modal and Intra-modal Consistency
Author: Ke Lv1 and Ying Wei1,2
Abstract: In complex noisy environments, many studies utilized visual information to assist speech separation task. However, audio and visual signals belong to two different modalities and have inherent heterogeneity that is difficult to mitigate. Directly applying visual information into speech separation may introduce irrelevant or redundant features that degrade model performance. Moreover, poor-quality visual information can significantly affect model performance in real-world scenarios. We propose a cross-modal speech separation method based on inter-modal and intra-modal consistency. It conducts a more fine-grained pre-training of the inter-modal consistency feature extraction network to mitigate modality heterogeneity. Using the prior knowledge learned during pre-training, it imposes intra-modal consistency constraints among speech signals to reduce reliance on visual quality. We evaluate the proposed method on VoxCeleb2 dataset. The results show that the method outperforms state-of-the-art methods on several evaluation metrics.
Oral Session: SPGT Session 01 (Paper No.6423)
Title: Robust and Computationally Efficient WPE Dereverberation Using Forward and Backward Beamformers
Author: Kang Chen1, Yujie Zhu1, Gongping Huang1, Jingdong Chen1 and Jacob Benesty2
Abstract: Reverberation severely degrades speech quality and intelligibility. Although the weighted prediction error (WPE) method has been widely adopted for speech dereverberation, it remains sensitive to additive noise and entails high computational complexity. To address these limitations, this paper presents a robust and computationally efficient WPE framework based on forward and backward beamformers. Specifically, a forward beamformer steered toward the target direction is employed to produce a cleaner reference signal, while a backward beamformer steered toward the opposite direction is introduced to estimate the residual late reverberation remaining in the forward beamformer output. In this way, the proposed framework improves dereverberation robustness in noisy environments while substantially reducing the computational complexity. Simulation results demonstrate the effectiveness of the proposed method.
Oral Session: SPGT Session 01 (Paper No.6445)
Title: SUBBAND ADAPTIVE DIFFERENTIAL BEAMFORMING WITH FFT AND MEL FILTER BANKS
Author: Li Du1, Fan Zhang2 and Lijun Zhang1
Abstract: Due to the limited order, differential beamforming has restricted
ability to suppress multiple noise sources, even when
these sources are spectrally non-overlapping. To address
this limitation, this paper presents subband extensions of our
recently developed adaptive differential (AD) beamforming
framework. Two types of filter banks are investigated: the
fast Fourier transform (FFT) and the Mel filter bank. For both
cases, we derive closed-form solutions for the subband AD
beamformer, with particular emphasis on the implementation
of the Mel-based design. Simulation results with two spectrally
distinct noise sources demonstrate that the proposed
approaches achieve superior noise reduction compared to
the original AD beamformer. Furthermore, the Mel-based
design provides a flexible tradeoff between noise reduction
performance and computational complexity.
SPGT Session 02
Time: 16:00~17:30, Saturday, July 18, 2026
Session Chair: Long Shi, Southwestern University of Finance and Economics
  
Oral Session: SPGT Session 02 (Paper No.6340)
Title: Denoising Weak Dolphin Whistles Using a Scale-Normalized Improved Wavelet Thresholding Function
Author: Jiali Chen1, Zhenquan Hu1, Ru Wu2, Peibin Zhu1, Wen Chen1 and Yougan Chen3
Abstract: To mitigate the Pseudo-Gibbs phenomenon and the constant-bias effect in conventional wavelet denoising, we propose a scale-normalized wavelet thresholding function. The proposed method incorporates a scale-dependent adjustment coefficient and a constant gain factor. It also utilizes the Suppression Impulsive and Autocorrelation Function (SI-ACF) to enable adaptive threshold optimization. Simulation experiments using measured background noise from Xiamen Bay and chirp-like synthetic whistle signals show that, for input signal-to-noise ratios from -15 dB to -5 dB, the proposed method effectively suppresses local oscillations and mitigates amplitude distortion. Results show that the proposed method improves the output signal-to-noise ratio by up to 0.22 dB and reduces the normalized root mean square error compared with conventional methods. Additionally, the autocorrelation peak remains above 0.80. These results demonstrate the method's effectiveness in high-fidelity reconstruction and robust feature extraction within complex, low-SNR underwater environments.
Oral Session: SPGT Session 02 (Paper No.6419)
Title: Source-Level Evaluation of Harmonic-Aware Fusion for Acoustic UAV Detection
Author: Fuyu Tao1, Xu Yang1 and Qian Zhao1
Abstract: Random clip-level splits can make acoustic unmanned aerial vehicle (UAV) detection look easier than it is, especially when many short clips come from the same recording. In this paper we ask whether a simple harmonic-aware representation is still useful after clips are grouped by inferred source recording. The method is deliberately lightweight: local spectral contrast enhances narrowband structures in the log spectrogram, and a second convolutional neural network branch keeps the original spectrogram view. We test entropy-based fusion and a quality-informed fusion variant on two public drone audio sources containing drone, background, helicopter, and unknown non-drone recordings. Under three-fold source-level cross-validation, entropy fusion gives the best mean accuracy and precision, reaching 88.38 percent accuracy, while quality-informed fusion gives the best recall and a slightly higher mean F1-score of 85.00 percent. The results are not uniformly positive across metrics. They suggest that the harmonic view is a useful auxiliary cue, but helicopter-like periodic negatives can still make a harmonic branch confidently wrong.
Oral Session: SPGT Session 02 (Paper No.6431)
Title: A Multimodal Interface for Enhancing Communication Accessibility with Bi-Directional Interaction Among Individuals with Hearing and Speech Impairments Using Deep Learning
Author: K SHILPA1 and Dr M. HEMALATHA1
Abstract: The problem of communication barriers lingers particularly in the case of the hearing and speech impaired, especially in daily communication between the two people where they do not both adopt a common communication channel. To overcome this drawback, this paper suggests a multimodal two-way communication system that incorporates the sign language recognition system, speech recognition system, and text-based communication into one assistive interface. The system takes visual input sign (captured by a camera with a resolution of 640?480 pixels) and interprets it through a deep neural network which is trained on 5,240 labeled gesture images of alphabetic hand configurations. The identified characters are translated into written text as well as audio signals. Speech recognition are then represented as a sequence of sign representations that are stored in a repository of 1,320 GIF-based sign representations. Spoken input received on a microphone with a sampling rate of 16,000 Hz is processed as a sequence of speech recognition entries which, in turn, produce text and find related sign representations in a repository. An experimental assessment of 4,860 trials of bidirectional interaction generated 4,417 accurate translation of communication. It has shown to be a stable interface that allows real-time communication with hearing and speech impaired people through the proposed multimodal interface. 
Oral Session: SPGT Session 02 (Paper No.6441)
Title: Approach to Closed-Form DOA Estimation with Covariance Matrix Residual Decomposition
Author: Zhixin Ma1, Chao Pan1, Jing Zhao1 and Jingdong Chen1
Abstract: Direction-of-arrival (DOA) estimation using spherical microphone arrays is sensitive to directional interference, which distorts the observation covariance and degrades the performance of closed-form DOA estimators. This paper proposes a robust closed-form DOA estimation method based on covariance matrix residual decomposition. The observation covariance matrix is modeled as the sum of target, interference, and diffuse noise. Given the spatial coherence priors of interference and diffuse noise, the target covariance is recovered through residual decomposition, while variances are estimated to construct a reliability weight for further fusion. The recovered target covariance is incorporated into a spherical harmonic domain closed-form DOA estimator without spatial spectrum search. Simulation results demonstrate that the proposed method achieves improved DOA estimation accuracy and robustness under directional interference.
Oral Session: SPGT Session 02 (Paper No.6444)
Title: A Two-Stage Collaborative Sorting Framework for Radar Signals in Complex Electromagnetic Environments
Author: Liu Junming1, Li Tao1, Fan Yifei1, Su Jia1 and Liu Xiangyang1
Abstract: Abstract-To address the challenges of high-density interleaved pulse streams and severe parameter overlap in radar signal pre-sorting, this paper proposes a parameter-adaptive hybrid clustering method that integrates Robust Density Kernel Estimation (RDKE) and Fuzzy C-Means (FCM). A two-stage "coarse-to-fine" sorting framework is developed, in which RDKE models the pulse feature distribution and adaptively determines both the number of clusters and initial centers based on local sample density, overcoming the limitations of traditional FCM. Subsequently, a fuzzy membership mechanism is introduced for fine sorting of signals with ambiguous parameter boundaries. Simulations conducted in complex scenarios with pulse loss and parameter overlapping demonstrate a sorting accuracy of 96.8%, outperforming several typical clustering-based methods and validating the effectiveness of the proposed approach. 
Oral Session: SPGT Session 02 (Paper No.6471)
Title: Convergence Analysis of a Robust Distributed Adaptive Strategy for Active Noise Control
Author: Sauravjyoti Senchowa1 and Bijit Kumar Das2
Abstract: An analysis of a robust diffusion strategy is addressed in this paper to achieve improved performance in a distributed active noise control (ANC) system. The system model is considered to consist of multiple microphones and multiple loudspeakers. We define a robust diffusion strategy for distributed ANC based on the exponential hyperbolic cosine cost function proven to be robust to both Gaussian and non-Gaussian observation noise, referred to as the diffusion filtered-x leaky exponential hyperbolic cosine adaptive filter (d-FxLEHCAF). The mathematical analysis shows the convergence criteria of the algorithm. The simulation results illustrate the efficiency of the algorithm in impulsive noise environments.
SPGT Session 03
Time: 09:00~10:30, Sunday, July 19, 2026
Session Chair: Mou Wang, Institute of Acoustics, Chinese Academy of Sciences
  
Oral Session: SPGT Session 03 (Paper No.6355)
Title: A Signal Acquisition Method Based on PMF-FFT and PCA-SVM for High-Dynamic and Low-SNR Environments
Author: Xinyao Wei1, Wenbin Sun1, Xin Yang1, Lili Chen1, Qian Xu1 and Ling Wang1
Abstract: This paper addresses the problems of difficulty in setting the acquisition threshold, high miss rate, and degraded detection performance of the traditional PMF-FFT algorithm in low-SNR and high-dynamic environments. A signal acquisition method based on PMF-FFT and PCA-SVM is proposed. The proposed method retains PMF-FFT as the front-end search framework, extracts statistical features from the acquisition correlation results, employs principal component analysis (PCA) for dimensionality reduction, and utilizes support vector machine (SVM) to perform classification, thereby determining whether the signal has been successfully acquired. Simulation results show that, compared with the traditional threshold-based decision method, the proposed method achieves a lower miss rate, a higher F1-score, and improved acquisition accuracy in low-SNR and high-dynamic environments.
Oral Session: SPGT Session 03 (Paper No.6398)
Title: Direct Localization of Coherent Sources with Moving Arrays Based on Cyclic Cross-Correlation Tensor and Sparse Bayesian Learning
Author: Yiquan Zhang1, Yuexian Wang1, Chuang Han1, Yandong Sun1 and Ling Wang1
Abstract: To address the challenges of limited localization accuracy and strong co-channel interference in localizing coherent radiation sources within complex electromagnetic environments, this paper proposes a direct position determination algorithm combining cyclostationary features, joint Doppler modeling, and sparse Bayesian learning. First, a spatio-temporal geometric model for a moving array is established, which accurately maps the Doppler shift as a nonlinear function of the target coordinates and the array velocity. Second, by exploiting the cyclostationary characteristics of the target signals, the cyclic cross-correlation matrix is calculated. This step not only suppresses stationary noise and non-homogeneous interference but also constructs a multi-delay spatio-temporal third-order tensor. Subsequently, the spatial subspace is extracted via higher-order singular value decomposition, and the forward-backward averaging technique is introduced to restore the rank of the coherent signal subspace without sacrificing array aperture. Finally, a hierarchical sparse Bayesian model is constructed to achieve high-precision, high-resolution direct localization through adaptive hyperparameter updates. Simulation results demonstrate that the proposed method exhibits excellent localization robustness in environments characterized by low signal-to-noise ratios and strong interference.
Oral Session: SPGT Session 03 (Paper No.6399)
Title: Angle Estimation for IRS-Aided Bistatic EMVS-MIMO Array in the Presence of Mutual Coupling
Author: Zijian Luo1, Yuexian Wang1, Yandong Sun1, Chuang Han1 and Ling Wang1
Abstract: To address the problem of angle estimation in bistatic polarized MIMO arrays assisted by Intelligent Reflecting Surfaces (IRS) under mutual coupling effects, this paper proposes a Reduced-Dimension Rank-Reduction (RD-RARE) estimation algorithm based on tensor decomposition. First, the factor matrices are recovered using Tensor Train Decomposition (TTD). Next, the inter-mediate matrix is reconstructed through a transformation in the angle domain and the properties of the Kronecker product, enabling angle estimation using the recovered imperfect array manifolds. The mutual coupling coefficients are then obtained based on the properties of eigenvectors and eigenvalues. This method transforms a two-dimensional search into two one-dimensional searches, significantly reducing the computational complexity of the algorithm. Furthermore, simulation results show that after self-correction of the mutual coupling coefficients, the accuracy of the proposed algorithm is comparable to that of the self-corrected Rank-Reduction (RARE) algorithm.
Oral Session: SPGT Session 03 (Paper No.6400)
Title: Distributed Multi-Dimensional Interference Optimization Method Based on Game Theory
Author: Yuquan Shu1, Xunbin Zhou2, Fangli Tian1, Lingling Zhang3, Peilin Chen1 and Chengkai Tang4
Abstract: As Low-Earth Orbit (LEO) satellite constellations are increasingly utilised in key sectors such as military communications, navigation and positioning, and intelligence and reconnaissance, their strategic importance is becoming increasingly evident. To address the issues of low coordination efficiency, insufficient resource utilisation and poor dynamic adaptability in traditional distributed jamming methods for low-Earth orbit satellite communications, this paper proposes a game-theory-based distributed multi-dimensional jamming optimisation method. By treating each primary jammer as a non-cooperative game participant and the allocation of relative jammers as a pure strategy, we construct a composite utility function that integrates delay, power and Doppler shift, and introduce a benefit-sharing coefficient to address resource-sharing conflicts. We employ mixed-strategy games and the Newton iteration method to solve for the Nash equilibrium, thereby obtaining the optimal allocation scheme for relative jammers. Experimental results demonstrate that this method achieves 100% single-path coverage under multi-path interference conditions, with a bit error rate (BER) consistently maintained at the 10?? level and an average utility value of 3.40, significantly outperforming genetic algorithms, dynamic programming and gradient descent algorithms.
Oral Session: SPGT Session 03 (Paper No.6413)
Title: Noncircular_Source_Enumeration_via_Pseudo_Covariance_Lifting_and_Ritz-Gap_Detection
Author: Chuang Han1, Feixiang Li1, Yandong Sun1 and Yuexian Wang1
Abstract: This paper addresses the problem of estimating the number of noncircular sources from complex-valued array observations. Unlike conventional covariance-based enumeration, extending the Lanczos/Ritz framework to pseudo-covariance is nontrivial, since the pseudo-covariance matrix is complex symmetric rather than Hermitian. To overcome this difficulty, a structured pseudo-covariance lifting method is proposed. By exploiting the anti-diagonal structure of the pseudo-covariance matrix under a uniform linear array, the matrix is first converted into a smoothed one-dimensional sequence through anti-diagonal averaging and then lifted into a Hankel matrix to form a Hermitian positive semidefinite matrix. This transformation converts non-circular source enumeration into a structured rank-determination problem compatible with Lanczos tridiago-nalization and a practical logarithmic Ritz-gap detector. Under mild conditions, the rank of the lifted matrix is shown to equal the number of noncircular sources. Simulation results demonstrate that the proposed method achieves high detection probability in mixed-signal scenarios with significantly reduced runtime compared to full-decomposition-based alternatives.
Oral Session: SPGT Session 03 (Paper No.6428)
Title: Direct Calibration of Sensor Position and Clock Errors Based on Frequency-Weighted Cooperative Signals
Author: XU Jun Nian1, Liu Xiang Yang1, Guo Zi Xun1, Fan Yi Fei1, Xie Jian1 and Tao Ming Liang1
Abstract: To address the degradation in localization performance caused by receiver position errors and clock synchronization biases in multi-station passive localization systems, a direct calibration method based on cooperative signals is proposed. Different from conventional two-step methods that first extract time-difference-of-arrival measurements and then estimate the error parameters, the proposed method directly exploits the received signals at all observation stations. After matched filtering, the cross-spectrum between the reference station and each secondary station is constructed. Based on the maximum likelihood criterion, an objective function with respect to the receiver position errors and clock biases is established, and the joint estimation of these two types of error parameters is achieved by optimizing the objective function. Simulation results show that the proposed direct calibration method can effectively improve the calibration accuracy. Especially under low-SNR conditions, compared with conventional two-step methods, the calibration accuracy can be improved by more than 10 times.
SPGT Session 04
Time: 11:00~12:30, Sunday, July 19, 2026
Session Chair: Mou Wang, Institute of Acoustics, Chinese Academy of Sciences
  
Oral Session: SPGT Session 04 (Paper No.6319)
Title: Parameter Optimization Design for Combined Frequency Modulation Signals
Author: Bo You1, Baozhu Chen2, Fangyong Wang1 and Wenbo Ma1
Abstract: This paper derives the wideband ambiguity function ridge-line equations of combined frequency modulation signals and their linear approximations under multi-target conditions. A method for optimizing waveform parameters to avoid spurious intersections is investigated and validated through simulations. Based on an analysis of the anti-reverberation performance and time-delay resolution, the signal combination is extended to non-equal-period N-pulse pairs.
Oral Session: SPGT Session 04 (Paper No.6339)
Title: Joint Optimization of Waveform, Mismatched Filter, and Classifier for Radar Target Recognition
Author: Jiahang Wang1, Wentao Zhu1, Yifan Wu1, Tao Wang1 and Junli Liang1
Abstract: The matched illumination theory suggests that designing waveforms matched to target scattering characteristics can actively enhance discriminative features beneficial for classification from the returned echoes. However, classification-driven waveforms often compromise the low-sidelobe properties essential for pulse compression, creating a conflict between radar detection and recognition. To address this, this paper proposes a joint waveform-filter-classifier (JWFC) optimization framework that integrates the transmit waveform, the mismatched filter (MMF), and a classifier into a differentiable end-to-end architecture. Within this framework, the waveform acts as a learnable feature encoder, jointly optimized with the classifier to align with its inductive bias. Concurrently, an MMF is co-optimized to suppress range sidelobes induced by the waveform, preserving detection performance. A knowledge-guided two-stage training strategy is employed to ensure stable convergence. Experiments on electromagnetic simulation data demonstrate that the proposed JWFC framework achieves higher classification accuracy than traditional waveforms across various code lengths, while maintaining peak and integrated sidelobe levels better than -30 dB. The results confirm that the collaborative "transmitter-receiver" design achieves functional compatibility between recognition enhancement and detection requirements.
Oral Session: SPGT Session 04 (Paper No.6342)
Title: Domino Group Sparsity Subarray Beampattern Synthesis
Author: Yifan Wu1, Yubin Zhang1, Yitong Li1, Ziyi Liu1, Junli Liang1 and Jiahang Wang1
Abstract: Subarray level array beampattern synthesis is an active area of research. To ensure beampattern performance while enabling the centralized arrangement of subarrays to save antenna space, this paper proposes a domino-inspired sparse pattern synthesis method. First, the centralized arrangement of subarrays is equivalently formulated as a specially designed sparse optimization problem, where the indices of subarrays are sorted according to their distances from the center of the panel. Then, inspired by the domino toppling chain reaction, a composite norm objective function is constructed. By imposing appropriate constraints, the sparsity mechanism of "toppling the first domino" is triggered, leading to an NP-hard optimization problem. The alternating direction method of multipliers (ADMM) is employed to solve this problem. Theoretical analysis and solution steps demonstrate that the proposed model can effectively force the selected subarrays to concentrate near the panel center, achieving a domino-like centralized sparse distribution. As a result, it achieves significant savings in antenna layout space while maintaining the desired beampattern performance.
Oral Session: SPGT Session 04 (Paper No.6368)
Title: A method of adaptive beamforming based on the cross correlation convex combination optimization
Author: Bingjie Yin1, Zhongyong Li2, Xi Zhang2 and Ji Xu2
Abstract: The performance of adaptive beamforming method is affected by the low signal-to-noise ratio and signal steering vector mismatch. A method of adaptive beamforming based on the cross correlation convex combination optimization is proposed in this paper. This method exploits the influence of the conjugate diagonal loading level to the signal to obtain the signal power of the different mutual conjugate power vector in the closed range. Then, the maximum value of the matrix is obtained, and the diagonal loading level corresponding to maximum power is found. Finally, the robustness performance of the algorithm in the complex environment is improved via the diagonal loading technique. This method requires no user parameters, which can greatly save the operation cost in the closed range, and has better anti-interference ability in the lost environment.
Oral Session: SPGT Session 04 (Paper No.6383)
Title: A Fourth-Order Cumulant-Based Method for Joint Angle-Doppler Resolution Enhancement
Author: Mengyang Shui1, Yubin Zhang1, Yitong Li1, Qi Guo1 and Junli Liang1
Abstract: This paper proposes a fourth-order-cumulant-based (FOC-based) space--slow-time virtual extension method to improve the resolution of joint angle--Doppler estimation in radar systems. Different from conventional processing that directly operates on the original array--pulse data, the proposed method employs a one-step FOC operation to simultaneously synthesize virtual spatial and slow-time degrees-of-freedom. By introducing the reference transmit waveform into the cumulant computation, the proposed method incorporates matched-filter-consistent waveform matching and enables joint array--pulse phase combination. The resulting cumulant outputs are rearranged into a structured 2D virtual angle--Doppler matrix with enlarged effective apertures in both dimensions. Under the assumption that targets are separable in the range domain, the proposed method yields a separable virtual manifold for angle--Doppler representation. Therefore, additional structured degrees-of-freedom can be obtained without increasing the number of physical sensors or transmitted pulses. Simulation results show that the proposed method improves joint angle--Doppler resolution compared with conventional processing on the original data.
Oral Session: SPGT Session 04 (Paper No.6412)
Title: Surrogate-Assisted Multi-Objective Optimization of Distributed Arrays Based on Incremental Stratified Sampling
Author: Chuang Han1, Chi Cheng1, Dong Yan Sun1 and Xian Yue Wang1
Abstract: Distributed array geometry optimization involves high-dimensional non-convex search, multi-objective performance trade-offs, and expensive sidelobe evaluation. To address these challenges, this paper proposes a surrogate-assisted multi-objective optimization method based on incremental stratified sampling. To reduce the cost of generating training data for the surrogate model, an incremental data generation strategy is developed that combines stratified uniform sampling with local subarray perturbations. This approach effectively limits the number of expensive radiation pattern evaluations while maintaining uniform coverage of the sample space. A gradient boosting decision tree (GBDT) surrogate model is then constructed to approximate the peak sidelobe level and integrated into the multi-objective optimization framework to accelerate the search process. Simulation results demonstrate that the proposed method significantly reduces computational time while achieving optimization performance comparable to that of the high-fidelity pattern model, validating its effectiveness for distributed array optimization problems.
SPGT Session 05
Time: 14:00~15:30, Sunday, July 19, 2026
Session Chair: Xuehan Wang, Guangzhou University
  
Oral Session: SPGT Session 05 (Paper No.6384)
Title: Interleaved Radar Pulse Sorting via Physics-Guided PA/DTOA Dual-Task Modeling
Author: Huaifeng Tang1, Weixuan Wang1, Hong Luo1, Ming Zhang1, Taoran Qi1 and Junfeng Qiu1
Abstract: In dense electromagnetic scenes, intercepted radar pulses rarely appear as clean single-emitter sequences. Pulses from different emitters are mixed along the time axis, while jitter, missing pulses, spurious pulses, and measurement errors further weaken the regular PRI structure. This paper treats radar pulse sorting as two physics-related recognition problems instead of one general PDW classification task. Pulse amplitude (PA) is used to describe envelope activity and beam-related energy changes, whereas difference of time of arrival (DTOA) is used to describe PRI timing rhythm. Based on this separation, PA-envelope classes and DTOA/PRI rhythm classes are defined according to radar operating mechanisms.
Oral Session: SPGT Session 05 (Paper No.6408)
Title: Enhancing Ant-Jamming Perfomance Against DRFM Jammers in Mobile Surveillance Radars Using Chebyshev Chaotic Encryption-Based Pulse Authentication
Author: Denis Betwell Temba1, Qingwei Meng1 and Wang Han1
Abstract: Mobile surveillance radars working in contested environment are increasingly vulnerable to Digital Radio Frequency Memory (DRFM) threats, which produce coherent deceptive signal by presenting a false target or conceal a real one. Traditional Anti-jamming techniques are insufficient because DRFM systems can replicate and retransmit them, leaving the radar unable to distinguish from authorized echoes and deception echoes which are copied. This paper proposes Chebyshev chaotic encryption based-pulse Authentication technique to enhance Radar Anti-jamming performance against DRFM deception. The method exploits Chebyshev polynomial chaotic map which offers strong ergodicity, sensitivity to initial condition and excellence correlation properties, we generated dynamic pseudo random phase code and embedded it on each transmitted pulse. At the receiver, synchronization using a regenerated expected chaotic code and verifies pulse authenticity through high correlation with genuine echoes while rejecting DRFM retransmitted signal. Simulation results in mobile radar scenario show improved probability of detection under deception jamming reduced false alarms from repeater attack over 85% to below 5% and rejects over 95% of deceptive pulse. Compared to non-chaotic and other map alternatives, Chebyshev approach offers superior autocorrelations, large space and robust performance suitable for mobile platform including Unmanned Aerial Vehicle (UAV) and ground based radars.
Oral Session: SPGT Session 05 (Paper No.6429)
Title: Distributed Gated Generalized Likelihood Ratio Detection Fusion for Underwater Active Detection
Author: Jinhu Zhao1, Xi'an Feng2 and Lu Qiao3
Abstract: Underwater active detection is critical for long-range surveillance, yet single-sensor performance is often degraded by propagation loss, noise uncertainty, multipath, and target scattering fluctuation. Distributed detection fusion mitigates these effects by combining information from spatially separated sensors. Classical hard-decision fusion has low communication requirements, but its optimal sensor weights depend on detection probabilities that are determined by unknown target-present signal amplitudes, making weight specification difficult for underwater active echoes. To remove this dependence, this paper proposes a gated generalized likelihood ratio test (GLRT) detection-fusion criterion for unknown single-sensor normalized signal amplitudes. Under a normalized Gaussian model, the GLRT estimates the unknown signal amplitude from the current observation, yielding a statistic in which stronger target-consistent observations contribute more without requiring preassigned per-sensor signal-to-noise ratios. Each sensor gate is set through a prescribed sensor false-alarm probability, equivalently a noise-only transmission probability, so that only sensors whose statistics exceed the resulting threshold transmit real-valued contributions to the fusion center. The mixed distribution of the gated statistic is derived in closed form, and the global distribution is computed by fast Fourier transform based convolution for Neyman-Pearson threshold design under a false-alarm constraint. Simulations show accurate global false-alarm control and demonstrate that, at a detection probability of 0.85, the proposed method reduces the required input signal-to-noise ratio by about one decibel and increases the detection-range gain factor from 1.19 to 1.26 relative to hard-decision fusion, without using preassigned sensor signal-to-noise ratios in fusion weighting.
Oral Session: SPGT Session 05 (Paper No.6437)
Title: Distributed MVI-CFAR detection method based on fuzzy logic
Author: Yujun Hou1 and Qunfei Zhang1
Abstract: To address the limitations of traditional CA-CFAR detection, which performs poorly under multi-target interference and clutter edge backgrounds, a multi-strategy CFAR detection method based on fuzzy logic with adaptive threshold selection is proposed. This method incorporates fuzzy logic and combines the advantages of the CA, GO, and ACCA detection algorithms, enabling the selection of the most appropriate detection strategy by effectively discriminating the current background. Building on this, a distributed CFAR detection system employing this method as the local detector is further developed, with the fusion center utilizing three fusion rules: fuzzy algebraic product, fuzzy maximum, and fuzzy minimum. Simulation results demonstrate that, compared to existing methods, the proposed approach achieves superior detection performance in homogeneous environments, multi-target interference scenarios, and clutter edge backgrounds. Among the three fusion rules, the distributed system based on the fuzzy algebraic product fusion rule delivers the best detection performance.
Oral Session: SPGT Session 05 (Paper No.6469)
Title: Method of Helicopter Target Detection for Anti-rotor False Flicker Jamming
Author: Jie Jiang1, Junqiang Wu1, Shuangshuang Li1, Xuezhou Zhao1, Shuwen Wang1 and Zhenkai Wu1
Abstract: Abstract-The detection and identification of helicopter rotor signals has always been an important research issue in the field of low altitude defense. In this paper, the method of helicopter target detection is proposed, and the target segmentation, data smoothing, and quadratic clustering algorithm are used to detect the target of the suspension helicopter in the complex clutter background. The experimental results show that the method has greatly improved the accuracy of the target detection.
Oral Session: SPGT Session 05 (Paper No.6483)
Title: DOA estimation via acoustic vector sensor array under time-varying element position errors
Author: Yu Chen1, Xinshi Zhang1, Weidong Wang1, Hui Li1, Lingling Zhang2 and Wentao Shi2
Abstract: To mitigate the degradation in direction of arrival (DOA) estimation performance caused by time-varying element position errors in an acoustic vector sensor array (AVSA), a joint calibration iterative minimization algorithm (JCIMA) is proposed in this paper. Initially, a correction source is introduced, and eigenvalue decomposition is applied to derive an initial correction matrix. The corrected received data are subsequently utilized to provide a preliminary estimate of the signal azimuth. Following this, a cost function is constructed based on both the corrected received data and the ideal array signal derived from the estimated azimuth. Through iterative updates of the correction matrix and the DOA aimed at minimizing this cost function, the algorithm achieves an accurate DOA estimation. Simulation results indicate that, in comparison to existing methods, the proposed JCIMA algorithm significantly enhances the DOA estimation performance under the time-varying element position errors in AVSA.
SPGT Session 06
Time: 16:00~17:30, Sunday, July 19, 2026
Session Chair: Xuehan Wang, Guangzhou University
  
Oral Session: SPGT Session 06 (Paper No.6350)
Title: Multi-Layer Coupled Transition Network based on SSA and Inter-Layer Association Representation for Underwater Weak Target Signal Detection
Author: Zihang Xue1, Haiyan Wang1, Bo Geng1 and Xiaohong Shen1
Abstract: Complex network-based nonlinear time series analysis can directly reveal intrinsic target characteristics, thereby enabling underwater target detection without prior target information. However, most existing complex network-based methods are primarily designed for ideal noise-free time series and cannot effectively characterize target signals in complex ocean environments. Regarding this issue, we propose a multi-layer coupled transition network for weak target detection based on singular spectrum analysis and inter-layer association representation (SSA-ILAR-MLCTN). First, a progressive residual embedding matrix is constructed through singular spectrum analysis to enhance the representation of subtle dynamic differences under low signal-to-noise ratio (SNR) conditions. Next, multi-channel time series are mapped into a multi-layer coupled network through an inter-layer association representation, and a multi-layer weighted cross clustering coefficient entropy is proposed to distinguish the noisy target signal from the background noise. Comparative experiments were conducted using two types of data: simulated Chen chaotic signals under different SNR conditions and measured ocean trial data at different ranges. The results show that the proposed method is superior to the traditional transition network method and the narrowband cross spectrum detection method. The proposed method provides a new approach to weak target detection using vector hydrophones when no prior target information is available.
Oral Session: SPGT Session 06 (Paper No.6391)
Title: Space-Time Division Cluster-Based DOA Estimation for Underwater Acoustic Spread-Spectrum Signals
Author: Biao Kuang1, Yibo Ma2, Yuxi Liu3 and Feng Zhou4
Abstract: Aiming at the problems of insufficient direction-of-arrival (DOA) estimation accuracy, dependence on prior knowledge of source number, and poor wideband direction-finding performance of traditional algorithms for underwater acoustic spread-spectrum communication signals in multipath, noise and interference environments, this paper conducts a systematic study on the space-time division cluster (STDC) DOA estimation algorithm. The algorithm integrates spatial multi-beamforming and time-domain matched filtering, directly utilizes the excellent autocorrelation characteristics of spread-spectrum signals, and can realize joint estimation of incident angle and propagation delay of multipath signals without knowing the number of sources in advance, which can effectively suppress noise interference and improve coherent multipath resolution. Simulation experiments are carried out to compare the proposed algorithm with spatial smoothing MUSIC, generalized almost-cyclostationary MUSIC and the space-time division cluster algorithm. The results show that the STDC algorithm has higher direction-finding accuracy and robustness under low signal-to-noise ratio, strong interference and dense multipath conditions, and is naturally compatible with the underwater acoustic spread-spectrum communication system. The research results can provide a reliable direction-finding method and theoretical support for array signal processing of underwater acoustic spread-spectrum communication systems.
Oral Session: SPGT Session 06 (Paper No.6396)
Title: Optimal Detection Range Analysis of a Rectangular Parametric Acoustic Array
Author: Peihong Wang1, Jianjun Zhu1, Zihao Shu1, Chuan Liu1, Gouqing Zhao1 and ?.?. Tarasov2
Abstract: Parametric acoustic array can generate low-frequency narrow beams with a compact aperture, but the difference-frequency sound pressure level and beamwidth exhibit distinct range-dependent behaviors, making the detection performance non-monotonic along the propagation axis. This paper establishes a numerical model for a rectangular-aperture parametric array based on the KZK equation and calculate the axial evolution of the difference-frequency sound pressure level and ?3 dB beamwidth. A figure of merit (FoM) combining signal strength and beam footprint area is proposed to quantify the optimal detection range, defined as the range maximizing FoM, with the effective detection zone bounded by a 3 dB drop. Simulation results show that the FoM curve exhibits a clear single-peak structure, with the optimum located near both the sound pressure level peak and the beamwidth convergence region. The proposed method provides a quantitative basis for working-range selection in underwater target detection and other applications such as sub-bottom profiling.
Oral Session: SPGT Session 06 (Paper No.6417)
Title: Acoustic Signal Characteristic Analysis of Marine Targets Based on Distributed Acoustic Sensing
Author: Wenlong Wei1, Haitao Wang1 and Xiangyang Zeng1
Abstract: This paper investigates the application of Distributed Acoustic Sensing (DAS) for wide-area marine target identification. Leveraging the Ocean Observatories Initiative (OOI) deep-sea dataset, we effectively suppress common-mode system noise and random noise through two-dimensional (2D) median filtering and bi-directional median subtraction. Our research demonstrates that distinct target characteristics can be identified using spatio-temporal maps and frequency-wavenumber (f-k) spectra. Specifically, whale vocalizations manifest as clear V-shaped trajectories, whereas vessel signals present as broadband energy shadows characterized by distinctive displacement patterns. Furthermore, f-k spectral analysis confirms that differences in axial phase velocity, represented by the slope of energy distribution, serve as the critical physical basis for sound source identification. These results highlight the significant advantages of DAS in wide-area acoustic field reconstruction, suggesting that its complementary integration with traditional hydrophone arrays will provide robust support for digitized marine monitoring.
Oral Session: SPGT Session 06 (Paper No.6426)
Title: VMD-Assisted Adaptive Frequency-Band Partitioning and Cross-Band Attention Fusion for Underwater Acoustic Target Recognition
Author: Lin Zhiyang1, Wang Depeng2 and L? Yaohui1
Abstract: To fully exploit the multi-band discriminative information of ship-radiated noise and improve the recognition accuracy and noise robustness of underwater acoustic targets, this paper proposes a VMD-assisted adaptive frequency-band partitioning and cross-band attention fusion method. The proposed method employs variational mode decomposition to obtain the center frequency and energy proportion of each modal component, and then uses a clustering algorithm to adaptively divide the modal components into low-, middle-, and high-frequency sub-bands. According to the time-frequency structural differences among different frequency bands, Mel spectrograms with different window lengths are constructed as multi-resolution inputs. Furthermore, a Dual-path CBA Network is designed, in which the main path extracts global multi-band features, while the CBA auxiliary path models the correlations among different frequency-band features, thereby achieving multi-band information fusion. Experimental results on the DeepShip dataset show that the proposed method achieves a recognition accuracy of 99.30%, outperforming the compared methods.
Oral Session: SPGT Session 06 (Paper No.6438)
Title: A Multi-Frame 2D Autoregressive Prewhitening Method with Adaptive Order Selection for Reverberation Suppression
Author: Yuan Cao1, Tianjun Zhou1 and Qunfei Zhang2
Abstract: Detection performance in active sonar is severely degraded by spatially and Doppler-spread reverberation under limited sample support. To address this issue, a multi-frame two-dimensional autoregressive (2DAR) prewhitening approach is proposed. 2DAR parameters are jointly estimated across adjacent data frames. In addition, a multi-frame averaging-based model order selection scheme is developed, whose advantage lies in improved robustness under limited snapshots. Based on this framework, a generalized likelihood ratio test detector is derived. Simulation results demonstrate that the method achieves superior whitening and improves detection performance compared with conventional 1DAR and single-frame 2DAR approaches, particularly in low signal-to-reverberation ratio conditions.
Oral Session: SPGT Session 06 (Paper No.6466)
Title: Acoustic-Band Guided Multi-Feature Fusion for Underwater Acoustic Target Recognition
Author: Weijie Ning1, Zhe Jiang1, Bingbing Zheng1 and Bo Geng1
Abstract: Underwater acoustic target recognition (UATR) relies on discriminative acoustic structures in target-radiated noise, such as narrowband tonal components, spectral envelopes, and broadband energy distributions. However, these acoustic patterns are unevenly distributed across frequency bands and can vary with target types and observation conditions. Existing multi-feature methods usually stack different acoustic features in a fixed order, which may cause the model to over-rely on a specific feature position and underuse complementary acoustic information. To address this problem, this paper proposes an acoustic-band guided multi-feature fusion method. Acoustic Band Adaptive Weighting (ABAW) adaptively enhances informative acoustic bands and suppresses noise-dominated bands according to band content and band position. Feature De-correlation Mixing (FDM) further reduces fixed feature-order dependence in a shared latent feature space through random feature permutation and cross-feature local exchange. Experiments on the ShipsEar and DeepShip datasets show that the proposed method achieves recognition accuracies of 96.3% and 88.5%, respectively. Ablation studies and qualitative analyses further verify the effectiveness of ABAW and FDM.
SPGT Session 07
Time: 08:30~10:00, Monday, July 20, 2026
Session Chair: Fan Zhang, Wuhan University
  
Oral Session: SPGT Session 07 (Paper No.6335)
Title: IoT-Based Multimodal Early Stroke Detection System Using Edge Computing
Author: Pedamala Bharath1
Abstract: Stroke is a significant cause of death and long term disability in most sections of the world with timely response within the golden hour being a game changer on patient survival. However, the standard diagnostics such as the computed tomography (CT) and magnetic resonance imaging (MRI) is reactive, infrastructural and cannot be used in nonclinical environments where continuous monitoring is needed. The present paper suggests an IoT-based, multimodal, early stroke detection platform, which functions on the Raspberry Pi edge computing platform. The proposed system is a combination of gait analysis, facial asymmetry detection, and speech impairment analysis to monitor the early neurological symptoms. A camera, microphone, and inertial measurement unit (IMU) are used to capture the data and are then processed locally with lightweight machine learning models that do not need a high latency, which ensures data privacy and provides them with the ability to work offline. The gait, facial, and speech features are fused on a feature-level to provide a composite stroke risk score. The proposed multimodal approach, according to the results of the experiments, has accuracy of 93% t and it is superior to single-modality systems. It is not invasive and is affordable, and may be applicable in home-based, resource constrained health care settings.
Oral Session: SPGT Session 07 (Paper No.6348)
Title: DNCDSEM:A novel detection network with color dithering, squeeze and excitation modules
Author: Yilong Niu1, Xiao Yuan1, Tianhao Shi1, Qi Wang1, Yingmin Wang1 and Yi Wang1
Abstract: Object detection in remote sensing images aims to identify and localize targets accurately. With the rapid development of deep learning, numerous detection networks have been proposed for this task. Although these methods have achieved satisfactory results, some useful information may still be irretrievably lost during feature extraction, which can lead to false positives and false negatives. To address these issues, this paper proposes a novel Detection Network with Color Dithering and Squeeze-and-Excitation Modules (DNCDSEM) based on Real-Time Models for Object Detection (RTMDet). Specifically, the color dithering module is introduced into the data augmentation stage, while the squeeze-and-excitation module is integrated into the neck to enhance feature representation. These two modules enable the network to capture more discriminative and fine-grained feature information, thereby improving its feature extraction capability and detection performance. Extensive qualitative and quantitative experiments demonstrate the effectiveness of the proposed network and verify that it achieves promising performance.
Oral Session: SPGT Session 07 (Paper No.6357)
Title: Robust Surround-View Extrinsic Calibration for Water-Surface Platforms via ArUco-Constrained Bundle Adjustment
Author: Changsong Pang1, Xiaomin Zhang1 and Yang Yu1
Abstract: Accurate multi-camera extrinsic calibration is essential for surround-view perception in intelligent waterborne systems, yet remains challenging in water-surface environments due to weak textures, strong reflections, dynamic illumination, and vessel motion. To address this problem, this paper proposes an ArUco bundle adjustment framework for surround-view extrinsic calibration on marine platforms. Four floating ArUco boards are placed around the vessel as shared geometric references, and a global optimization problem is constructed to jointly refine camera extrinsics and board poses. Besides the reprojection term, a board-coordinate prior constraint and a coplanarity constraint are introduced to improve robustness and enforce the physical structure of the water-surface deployment. Experiments on a real platform with four fisheye cameras show that the proposed method yields geometrically consistent stitched top-view results in representative near-shore scenes and remains stable across different vessel poses and surrounding layouts. The proposed framework provides an effective and practical solution for surround-view calibration in realistic maritime environments.
Oral Session: SPGT Session 07 (Paper No.6361)
Title: Al-Driven Innovations to Enhance Tea Production and Quality in Sri Lanka's Mid Country
Author: Nipuna Abeysekara1, Dinithi Wijerathne1, Harsha Waduthanthri1 and Nimesh Bandaranayake1
Abstract: The Sri Lankan tea industry faces several operational challenges, including inconsistent leaf grading, inefficient yield forecasting, unpredictable auction demand, and delayed disease/pest detection. This research presents a novel multi-modal AI framework aimed at addressing these challenges by automating key decision-making processes. The system leverages deep learning techniques for four tasks: (1) freshness grading and tea-type classification from mobilecaptured leaf images, (2) yield forecasting using estate records and scanned logbooks, (3) auction demand prediction based on historical transaction data, and (4) disease/pest detection from images with seasonal risk prediction linked to climate data. 
Oral Session: SPGT Session 07 (Paper No.6452)
Title: Low-Light Object Detection via Feature Refinement and Semantic Enhancement
Author: Baoguo Wei1, Xinyu Wang1 and Xu Li1
Abstract: Object detection in low-light conditions remains a challenging problem due to the loss of fine details and global semantics caused by poor illumination. In this paper, we analyze two key factors limiting detection performance: the scarcity of local detail in shallow feature maps and the degradation of global semantic information in deep feature maps. To address these issues, we propose a unified detection framework with two plug-and-play modules: a Feature Refinement Encoder (FRE) and a Semantic Enhancement Module (SEM). The FRE bridges adjacent feature maps through a cross-feature attention mechanism, compensating for missing edge and texture information in early stages. The SEM employs grouped dilated convolutions with adaptive pooling to expand the receptive field, thereby recovering contextual cues lost due to dark occlusion. We evaluate our method on the ExDark dataset. Without any illumination priors or pretrained weights, our approach improves mAP by 5.8% over the YOLOX baseline, demonstrating the complementary benefits of synergizing local detail restoration and global semantic aggregation.
Oral Session: SPGT Session 07 (Paper No.6454)
Title: Multi-temporal Frame Difference Fusion and EfficientNet Ensemble Model for Man Overboard Action Recognition
Author: Kaixuan Tian1
Abstract: In complex marine scenarios, man overboard action recognition faces challenges such as insufficient motion feature representation, missed detection and false alarms caused by imbalanced sample categories. This paper proposes a man overboard action recognition method combining multi-temporal frame difference fusion and EfficientNet ensemble model. Firstly, YOLOv8 is adopted to detect and track personnel in surveillance videos, and continuous image sequences containing floating, struggling and swimming actions are obtained by cropping detection bounding boxes. Then, four temporal intervals of 3, 6, 9 and 12 are selected to extract multi-scale motion features via inter-frame difference. After eliminating the redundant model with poor performance, the three optimal frame difference features are mapped to RGB three channels to construct fused motion feature maps. Subsequently, multiple EfficientNet-B4 sub-models are trained respectively, and a weighted voting ensemble strategy is applied to complete sequence-level action classification. Experimental results show that the proposed fusion method greatly improves the recall rate of struggling actions to 84.44%, effectively reduces false alarms of swimming recognition. The overall accuracy exceeds 90% with an obvious improvement of average F1-score. The proposed method has strong robustness in complex marine environments, and it can provide technical support for intelligent maritime search and rescue.
SPGT Session 08
Time: 10:30~12:00, Monday, July 20, 2026
Session Chair: Qian Wang, North Electro-Optic Co., Ltd.
  
Oral Session: SPGT Session 08 (Paper No.6334)
Title: Fine-Grained Radio Frequency Fingerprint Open-Set Recognition using Discriminative Loss
Author: Zitai Liu1, Yulan Zhang1, Weishi Chen2, Jun Hu1 and Zengping Chen1
Abstract: Wireless devices recognition based on radio frequency (RF) fingerprint is promising for physical layer authentication. Many existing methods mostly adhere to the closed-set assumption, and may be unsuitable for complex and dynamic electromagnetic environments. In this letter, we propose a deep learning based open-set recognition method to extract fine-grained RF fingerprints, and design a discriminative loss to reduce the intra-class variation and increase the inter-class variation to expand the decision boundary. Experimental results on real-world RF datasets demonstrate that our proposed method can significantly improve the open-set recognition performance of RF fingerprints, while the proposed loss function does not introduce additional parameters.
Oral Session: SPGT Session 08 (Paper No.6406)
Title: Amplitude-Aware Dual Encoding in Ordinal Partition Transition Networks for Robust Signal Classification
Author: Bo Geng1, Haiyan Wang1, Gaoyue Ma1, Weijie Ning1 and Xiaohong Shen1
Abstract: Ordinal partition transition networks (OPTNs) map continuous time series into discrete topological graphs but fundamentally discard relative amplitude information. This structural blindness inherently causes state degeneracy, where distinct local fluctuations are mapped to identical ordinal patterns, severely limiting their discriminative capacity for complex signals. In this paper, we propose Amplitude-Aware Ordinal Partition Transition Networks (AOPTNs), a novel dual-encoding framework that systematically integrates quantized relative amplitude differences into the topological permutation process. By enriching the state space symbols, AOPTN seamlessly captures both the ordinal structure and the local metric dynamics of the underlying attractor. This provides a richer symbolic state representation than traditional methods, effectively resolving structural blindness. Through extensive experiments on synthetic chaotic systems, we demonstrate that AOPTN significantly enhances bifurcation detection sensitivity and provides higher resolution in distinguishing dynamical regimes. Furthermore, applying AOPTN to a highly non-stationary real-world 5-class underwater acoustic dataset yields an exceptional peak classification accuracy of 94.33%, substantially outperforming standard OPTN and permutation entropy methods. Crucially, by unfolding entangled feature spaces and resolving state degeneracy, AOPTN accelerates optimal hyperplane convergence in downstream support vector machines. This enables the proposed method to boost classification accuracy and stability without compromising computational efficiency.
Oral Session: SPGT Session 08 (Paper No.6442)
Title: VLM-NCD: Words as Anchors - Retrieval-Augmented Multimodal Discovery of Novel Classes
Author: Baoguo Wei1, Yuetong Su1, Xinyu Wang1, Xu Li1 and Lixin Li1
Abstract: Novel Class Discovery (NCD) aims to transfer knowledge from labeled known classes to partition unlabeled data that may contain both known and novel categories. Existing vision-only methods suffer from limited feature discriminability and severe degradation under long-tailed distributions. In this paper, we propose VLM-NCD, a multimodal framework that breaks through these bottlenecks by fusing visual and textual semantics for prototype-guided clustering. The core innovations include: (1) a joint representation that aligns image features with retrieved text descriptions to construct semantic prototypes for known classes, and (2) a dual-phase discovery mechanism that separates known and novel samples via semantic affinity thresholds, followed by adaptive semi-supervised k-means clustering. Experiments on CIFAR-100 and ImageNet-100 show that VLM-NCD improves novel-class accuracy by up to 25.3% over the state of the art while exhibiting strong robustness to long-tailed data, a rarely addressed issue in previous NCD literature. With only minimal fine-tuning of a frozen CLIP backbone, our framework provides an efficient and effective solution for open-world recognition.
Oral Session: SPGT Session 08 (Paper No.6450)
Title: STAR-Beat: A Spatio-Temporal Attention Recurrent Framework for Lightweight and Interpretable Atrial Fibrillation Detection using Wearable Photoplethysmography
Author: Mengcheng Hu1, Jiarong Chen2, Guoxing Wang2 and Cheng Chen2
Abstract: Wearable Photoplethysmography (PPG) enables continuous, non-invasive, and low-cost screening for real-time Atrial Fibrillation (AF) detection. However, deploying robust AF detection on edge devices remains challenging. While existing deep learning methods attempt to mitigate motion artifacts, they fundamentally treat time-series as local spatial data, lacking the capability to capture the long-range temporal rhythm dependencies crucial for identifying AF. Furthermore, their severe over-parameterization and limited black-box interpretability hinder real-world edge deployment and clinical trust. To address these critical gaps, this paper proposes STAR-Beat, an ultra-lightweight, temporally-aware, and highly interpretable multi-task AF detection framework. Building upon a convolutional denoising autoencoder (CDAE) for morphological prior extraction, STAR-Beat introduces Squeeze-and-Excitation (SE) attention to adaptively suppress noise channels, and innovatively utilizes a Bidirectional GRU (BiGRU) to explicitly capture long-range irregular rhythm dependencies. Validated on the Stanford and MIMIC PERform datasets, STAR-Beat achieves outstanding F1 scores of 0.962 and 0.946, respectively. Moreover, extensive post-hoc interpretability analyses verify that the model's attention accurately targets pathological hemodynamic characteristics, effectively bridging the trust gap for clinical application. Finally, to overcome hardware constraints, an edge-oriented structural ablation strategy compresses model parameters to merely 8.6% of the baseline, reducing inference delay to an ultra-low 1.885ms. The code of this article is available at https://github.com/04-07-08/STAR-Beat.
Oral Session: SPGT Session 08 (Paper No.6455)
Title: Fuzzy Measure Enhanced Unscented Kalman Filtering for Robust Range-Bearing Maneuvering Target Tracking
Author: Junkai Wang1, Yongsheng Yan2 and Zhuying Wang1
Abstract: This paper proposes a fuzzy-measure-enhanced unscented Kalman filter (FMUKF) for robust range-bearing maneuvering target tracking. FMUKF embeds fuzzy innovation reliability into model-conditioned unscented Kalman filtering, so unreliable measurements are downweighted before they distort the posterior state estimate. The implementation combines interacting model mixing, analytic prediction, an unscented range-bearing update, fuzzy reliability evaluation, and reliability-driven covariance inflation for glint-contaminated observations. A fixed-rate interacting multiple model unscented Kalman filter (IMM-UKF) and a posterior Cramer--Rao lower bound (PCRLB) provide the comparison baselines. Monte Carlo results show that FMUKF maintains competitive Gaussian accuracy while improving robustness under non-Gaussian glint noise.
Oral Session: SPGT Session 08 (Paper No.6464)
Title: A Fast Variational Dirichlet Process Mixture Tracking Filter for Online Multi-Target Tracking with Unknown Target Births
Author: Sihang Zhang1, Yongsheng Yan2 and Wang Xiong1
Abstract: Multi-target tracking (MTT) requires joint estimation of the number and kinematic states of targets from noisy measurements with unknown data association. Generalized labeled multi-Bernoulli (GLMB) filtering provides a principled random-finite-set solution, but its performance and computational cost can be affected by the birth model when target-birth times and locations are not available a priori. This paper presents a fast online variational Dirichlet process mixture model (DPMM) tracking filter based on a truncated stick-breaking representation and mean-field variational inference. At each scan, existing tracks are represented by active mixture components, while additional free components serve as a measurement-driven birth pool. The resulting soft association probabilities are used for Kalman state update, Bernoulli existence recursion, candidate birth confirmation, and track management. The proposed method is compared with an adaptive-birth GLMB filter under the same dynamic and measurement models. Simulation experiments show that the variational DPMM tracking filter achieves competitive tracking accuracy and improved cardinality estimation while substantially reducing computational cost.
Oral Session: SPGT Session 08 (Paper No.6470)
Title: CIAIC_MED8: A Dual-Stage Multi-Modal Emotion Database
Author: Qian wang1, Jun Zhu2, Mou Wang3, Jingdong Chen4, Yan Yang4 and Jieling Liu1
Abstract: This paper presents a newly developed multi-modal emotion database, namely CIAIC Multi-modal Emotion Database (CIAIC_MED8).This database is designed to serve as a standardized reference for affective computing research and will be released to the public. It contains multi-modal emotional data from 100 participants recorded in an anechoic chamber. Four target emotions, i.e., neutral, sad,angry and happy are induced by carefully selected, evaluated and validated video clips. Signals are captured in two successive stages: emotion perception while watching clips, and emotion expression while reading scripts aligned with target emotions. Synchronously recorded modalities include facial videos, ElectroEncephalo-Graph(EEG), electrocardiogram(ECG), electrodermalactivity(EDA), photo-plethysmography(PPG) and respiration(RSP). Inaddition, air-conduction and bone-conduction speech are collected during the expression stage. Subjective ratings of emotional experience and expression are annotated under both categorical andvalence-arousal dimensional models.This database aims to support fair evaluation and innovation in multi-modal emotion recognition research.
SPGT Session 09
Time: 10:30~12:00, Monday, July 20, 2026
Session Chair: Yingwei Tian, Wuhan University
  
Oral Session: SPGT Session 09 (Paper No.6390)
Title: Denoising and Prediction of Time-Varying Underwater Acoustic Channels in the Delay-Doppler Domain
Author: Yuxin Liu1, Yudong Chen1 and Fujun Lin1
Abstract: The complex time-space-frequency variability of underwater acoustic (UWA) channels poses a fundamental challenge to reliable underwater communications. To address the severe channel state information (CSI) outdating caused by the low speed of sound in adaptive modulation systems, this paper proposes a comprehensive processing and prediction framework for UWA channels in the delay-Doppler (DD) domain. First, measured channel impulse response (CIR) data are amplitude-normalized and segmented via a dual-layer sliding window, then transformed to the DD domain via fast Fourier transform (FFT); the 32 most energetic multipath/Doppler components are retained for dimensionality reduction. Second, a deep denoising convolutional neural network (DnCNN) tailored to the highly sparse DD-domain representation is proposed, incorporating per-sample RMS normalization, complex orthogonal decomposition, and an L1+MSE hybrid loss to suppress background noise while preserving sparse path peaks. Finally, a Multi-Scale Convolutional LSTM (MultiScale-ConvLSTM) prediction network with a parallel multi-branch architecture captures channel evolution across different temporal scales, enabling DD-domain channel prediction and inverse mapping to the time domain. Experiments on the Trex04, Unet06, and AUVFest07 datasets demonstrate that the denoising module reduces the noise floor to approximately 41.7% of the noisy energy, and an output SNR of 8.4-10.2 dB; the prediction module converges to a training loss below 0.002, validating the effectiveness of the proposed approach. 
Oral Session: SPGT Session 09 (Paper No.6393)
Title: Direction of Arrival Estimation Techniques for Vector Arrays Embedded in Viscoelastic Materials
Author: Pengfu Ding1, Yu Zhang1, Ke Li2, Chenglong Xia2 and Jun Fan1
Abstract: Vector sensors have incomparable advantages over traditional acoustic pressure sensors in detection applications. The ideal working environment for vector sensors is a free field, but they typically need to be installed on platforms, operating under non-free field conditions. Therefore, many scholars both domestically and internationally have successively begun researching vector sensor applications under non-free field conditions. However, the non-free field conditions studied here are all based on rigid materials, with very little research involving the effects of other material properties on vector sensors. Addressing the problem that vector sensors are affected by absorption and scattering from viscoelastic materials, which causes phase shifts leading to inaccurate azimuth estimation, this paper conducts research on robust azimuth estimation for vector arrays under these conditions. This paper applies an adaptive phase correction maximum likelihood azimuth estimation method for vector arrays to adaptively compensate for the phase of vector sensor arrays under viscoelastic material conditions, thereby achieving target azimuth estimation. Through comparative analysis of simulation experiments and lake trials, the excellent azimuth estimation performance of this method under viscoelastic material conditions has been verified, laying a technical foundation for enhancing the application of vector sensors in broader scenarios.
Oral Session: SPGT Session 09 (Paper No.6401)
Title: Design and Analysis of a High-Frequency LLC Resonant DC-DC Converter
Author: Ankita Ramanna Katapur1, Dr.Prashant V Joshi1 and Dr. Sudharshan K M1
Abstract: This paper discusses the design and analysis of an LLC resonant half-bridge converter. It operates under both open-loop and closed-loop control for a wide input voltage range of 370 to 430 volts, delivering a regulated 48 volts output at 1 kW. First, the converter is analyzed in open-loop configuration to examine the resonant tank characteristics, operating regions, and voltage gain using the First Harmonic Approximation (FHA) method. The converter is set to operate in the inductive region to achieve Zero-Voltage Switching (ZVS). This approach reduces switching losses and improves efficiency. To maintain stable output regulation, a closed-loop control scheme with a PI controller is used. In this scheme, the output voltage is compared to a reference voltage, and the error signal is processed to adjust the switching frequency. The resulting switching frequency generates gate signals with a fixed 50% duty cycle for the half-bridge switches. Simulation results show improved voltage regulation, reduced switching losses, and stable operation across the wide input voltage range, confirming the effectiveness of the proposed LLC resonant converter.
Oral Session: SPGT Session 09 (Paper No.6427)
Title: Fusing Time Frequency Representations and Orbital Parameters via Self Attention for LEO Satellite Interference Optimization
Author: Chengkai Tang1, Aomi Chen2, Zesheng Dan2 and Yangyang Liu2
Abstract: Low?Earth orbit (LEO) satellites have become a cornerstone of integrated space?ground communication networks, yet their high mobility, narrow beams and rapid frequency hopping pose severe challenges to military communication security. Traditional jamming methods struggle to maintain accurate frequency tracking and rapid synchronization in such dynamic scenarios. To address this, we propose an optimized interference method based on a multimodal learning transformer model (OI?MLT). The framework fuses time?frequency spectrum images with numerical signal?state parameters and introduces a composite positional encoding that captures both global orbital periodicity and local fast variations. A standard transformer encoder then models cross?modal dependencies through its symmetric self?attention mechanism, enabling end?to?end prediction of interference effectiveness indicators. Extensive experiments under simulated LEO environments demonstrate that OI?MLT achieves frequency tracking within ±12?kHz in 97% of trials, a mean synchronization time of 0.85?ms, and an interference success rate of 95%, representing an average improvement of 27.52% over existing methods. The results validate the superiority of multimodal fusion and transformer?based optimization for LEO satellite countermeasures.
Oral Session: SPGT Session 09 (Paper No.6453)
Title: Design of an Analog Front-End using Differential Current Signal Technic for Electric Air-Soft Gun Shooting Detection
Author: Pattapong Sripho1 and Mr.Jedsada Kraikhow1
Abstract: This research presents the design of an analog front-end circuit using differential current signal processing techniques for shooting detection in an electric air-soft gun system. A Hall-effect current sensor was used to measure the motor current waveform, while analog filtering and differential processing were applied to extract automatic firing transient characteristics. 
Oral Session: SPGT Session 09 (Paper No.6457)
Title: A New Method for Forward Scanning Imaging of Missile-Borne Radar
Author: Yuanyuan Sha1, Liling Niu2, Di Wu*3, Daiyin Zhu3 and Xiangjun Xu3
Abstract: Missile-borne radar can obtain high-resolution images of wide-area scenes through scanning and imaging signal processing. However, in the forward-looking region, the resolution of Doppler beam sharpening (DBS) degrades rapidly because of the reduced Doppler gradient. Array super-resolution provides higher spatial resolution in the forward-looking region, but its performance in the squint forward-looking region is inferior to that of DBS. To overcome these limitations, this paper proposes a scanning imaging method of missile-borne radar that improves the resolution of both the forward-looking and squint forward-looking regions. By decomposing the echo signal, the high-resolution imaging regions of DBS and array super-resolution are partitioned according to the effective array aperture, and Doppler super-resolution is introduced to further enhance DBS resolution. Simulation and measured results demonstrate that the proposed method satisfies the requirements of forward-looking scanning imaging and significantly improves image clarity in complex scenes.
CPT Session 01
Time: 16:00~17:30, Saturday, July 18, 2026
Session Chair: Kaili Yin, Qingdao University of Science and Technology
  
Oral Session: CPT Session 01 (Paper No.6351)
Title: Genetic Algorithm and Q-Learning Based Channel-Aware Cluster Routing Protocol for Underwater Acoustic Networks
Author: Yihao Zhao1, Zheyang Chen1, Ziyi Ding1, Shenao Tu2, Yougan Chen2 and Xiaomei Xu1
Abstract: To address the challenges of energy consumption constraints and channel dynamicity in underwater acoustic networks (UANs), this paper proposes a Genetic Algorithm and Q-Learning Based Channel-Aware Cluster Routing Protocol (GA-QL-RP). The protocol establishes a hybrid cluster routing framework featuring centralized decision-making and distributed adjustment: the Sink node employs genetic algorithms to periodically perform global cluster head (CH) selection, base on node residual energy, local link quality, and Sink link quality; meanwhile, each node conducts online channel-state learning and dynamic Q-value updates via stateless Q-learning, enabling distributed adaptive fine-tuning of the clustering structure based on real-time Q-values. By synergistically combining the global optimization capability of genetic algorithms with the local dynamic adaptability of Q-learning, the protocol effectively balances routing reliability and energy efficiency. Simulation results demonstrate that, compared with the LEACH protocol and the genetic-algorithm-only clustering scheme GA-RP, GA-QL-RP achieves average improvements of approximately 172.20% and 23.00% in packet delivery ratio, respectively, and average reductions of approximately 72.06% and 25.95% in the energy tax metric, respectively, thereby verifying the significant performance advantages of the proposed protocol in dynamic underwater acoustic environments.
Oral Session: CPT Session 01 (Paper No.6369)
Title: Numerical and Experimental Study of a Built-in Pendulum-based Wave Energy Converter
Author: Shi Yihong1, Yangyang Cui2, Tao Wang3, Dongyang Chen3 and Siya Jin3
Abstract: This paper investigates a built-in pendulum-based wave energy converter (PWEC) to address the power constraints of marine Internet of Things (IoT) nodes, through theoretical modeling, numerical simulation and experimental testing. A unified mathematical model considering coupled-decoupled state switching is established and validated experimentally. Experiments are performed on a 1:4 scale PWEC model mounted on a six-degree-of freedom shaking table to reproduce surge motions representative of real ocean waves. The full-scale device is designed with a resonance pitch period of approximately 3 s (i.e., 1.5 s in the scaled model). Results indicate that increasing the arm length of the pendulum increases the device's resonance period: a 0.335 m arm length yields a resonance period of 1.25 s in the scaled model. Additionally, the system achieves a maximum output power of 0.992 W under an optimal load matching condition. These findings advance the understanding of PWEC dynamics and support the design and application of self-powered marine equipment using wave energy.
Oral Session: CPT Session 01 (Paper No.6370)
Title: UAVs-PPSim: A Simulator for Cooperative Path Planning of UAV Swarms
Author: Cunle Zhang1, Shiduo Zhang1, Haonan Wang1, Chengkai Tang1, Baowang Lian1 and Lingling Zhang1
Abstract: Reinforcement learning (RL) has demonstrated immense potential in UAV swarm path planning, yet its training heavily relies on efficient simulation environments. Based on the underlying framework of Isaac Sim, this paper proposes a simulation platform specifically tailored for UAV swarm path planning tasks. By decoupling low-level flight control from high-level decision-making, it facilitates the easy design and experimentation of various path planning scenarios on top of GPU-parallel simulations. It is equipped with 2 drone models, 5 sensor modalities, 4 control modes, and a selection of widely-used RL baselines. To showcase the capabilities of the platform, we provide preliminary results on several benchmark planning tasks. We hope this platform will facilitate future Sim-to-Real research on complex UAV swarms
Oral Session: CPT Session 01 (Paper No.6386)
Title: Performance assessment of a 6-float M4 wave energy converter for offshore autonomous power supply in the East China Sea
Author: Wei Li1, Dongyang Chen1 and Siya Jin1
Abstract: The 6-float M4 wave energy converter (WEC) is a representative multi-body system with potential for broadband energy capture. In this study, a time-domain numerical model is developed in WEC-Sim and validated against experimental data to investigate its hydrodynamic response and power performance. The results show that the relative hinge response exhibits two distinct resonance peaks, corresponding to different coupled motion modes. Consistently, two high-power regions are identified, indicating that energy capture is governed by the relative hinge motion. The maximum average power of 2.38 W is achieved at a wave period of 1.15 s and wave height of 0.04 m with a power take-off (PTO) damping coefficient of 3 N·m·s/rad. Using Froude similarity, the results are extrapolated to full scale and applied to the East China Sea. A tuned 6-float M4 device (~50 m) is predicted to achieve about 100 kW under the representative wave condition of the East China Sea, showing reasonable frequency compatibility with the regional wave climate. These findings highlight the importance of multi-body coupling and resonance tuning in enhancing energy capture and adapting WEC systems to site-specific environments and provide a useful reference for offshore autonomous power supply applications, such as ocean monitoring and Internet of Things (IoT) systems.
Oral Session: CPT Session 01 (Paper No.6388)
Title: Multi-modal Fusion based UAV Data Acquisition System for Maritime Search and Rescue
Author: Xiaodong Zheng1, Longyun Yuan2, Lingling Zhang2, Peilin Chen1, Hongjun Ye1 and Fangli Tian1
Abstract: Maritime person-overboard search and rescue faces critical challenges including low visibility, small target size, and difficulty in assessing survivor physiological states. This paper presents an integrated UAV-based perception and positioning system that fuses visible and thermal infrared imagery. The system uses a cascaded detection framework combining YOLOv26 and YOLO-World-S that achieves 79.5% mAP@0.5 with only 3.98M parameters. Vital sign assessment is performed through a six-dimensional fusion model that integrates thermal, pose, motion, physiological, consistency, and signal-to-noise features via entropy weighting, The assessment accuracy reaches 97.5%, a 23.5% improvement over single-modality baselines in lake trials and maintains 92.1% under complex sea state. Relative positioning between the UAV and a moving vessel is realized through an RTK-based dynamic scheme with centimeter-level accuracy. Lake and sea trial experiments validate the system's effectiveness and robustness for real-world maritime search and rescue operations.
Oral Session: CPT Session 01 (Paper No.6432)
Title: A Double Deep Reinforcement Learning Approach for Adaptive Access Control in Underwater Acoustic Networks
Author: Jianmin Yang1, Yan Lin2, Can Wang3, Zhihong Peng4, Ming Gong5 and Ying Huang6
Abstract: Abstract-To address challenges such as extreme propagation delays, high bit-error rates (BER), and stringent bandwidth constraints in Underwater Acoustic Sensor Networks (UASNs), this paper proposes DDQN-CA, an intelligent adaptive MAC protocol based on Double Deep Q-Networks (DDQN) and multi-dimensional feature perception. By integrating non-stationary features-including transmission success rates, collision frequencies, and node queue status-the protocol achieves precise prediction and proactive regulation of channel contention, breaking through the limitations of the traditional Binary Exponential Backoff (BEB) mechanism. To mitigate the stochasticity of underwater environments, the protocol introduces experience replay, target network, and safety circuit-breaker mechanisms, significantly enhancing policy convergence speed and robustness in complex dynamic scenarios. Simulation results demonstrate that, compared to the Pure ALOHA protocol, DDQN-CA maintains a stable Packet Delivery Ratio (PDR) exceeding 90%, reduces end-to-end latency by 18.7%-34.2%, and achieves more than a two-fold leap in peak throughput. Future work will focus on evaluating the protocol's performance in large-scale networks to further verify its robustness in complex oceanic environments.
Oral Session: CPT Session 01 (Paper No.6484)
Title: SitePrompt-PV: Metadata-Aware Cross-Site Few-Shot Forecasting for Cold-Start Photovoltaic Power
Author: Qing Wang1, Guohong Li2, Yun Wang3, Luoxiao Yang1 and Yue Zhao4
Abstract: Short-term photovoltaic (PV) power forecasting is commonly trained for each site independently, yet newly deployed sites often have only a few days of reliable local observations while related sites already contain transferable PV generation patterns. This paper formulates this setting as cross-site cold start PV forecasting and proposes SitePrompt-PV, a metadata-aware few-shot adaptation framework. SitePrompt-PV first maps site outputs to a per-unit capacity scale, learns a shared PV prior from pooled non-target sites, represents transferable site attributes as a metadata-conditioning vector, and then performs shot-aware lightweight calibration using 0-14 days of local target history. The framework is evaluated under a strict six target leave-one-site-out protocol on a two-month UNISOLAR slice containing 41 PV sites across five campuses, and is further examined on a power-only SKIPP'D external validation slice. On UNISOLAR, SitePrompt-PV Adaptive reduces daylight nRMSE/nMAE from 15.56%/10.08% with zero target days to 11.79%/7.95% with 7 days and 11.59%/7.94% with 14 days. At 7 days, it outperforms persistence, which obtains 16.44%/10.69%, and target-only PatchTST, which obtains 15.84%/10.19%. On SKIPP'D, daylight nRMSE decreases from 18.63% to 13.03% after 14 days of adaptation despite the external domain shift. Ablation results show that the source-site prior is the largest contributor, while metadata conditioning, forecast-horizon non-power covariates, target calibration, and shot-aware calibration weighting each provide additional gains. These results indicate that a capacity-normalized source prior, metadata conditioning, and shot-aware target calibration provide a practical solution for cold-start PV forecasting.
CPT Session 02
Time: 09:00~10:30, Sunday, July 19, 2026
Session Chair: Ningning Pan, Southwestern University of Finance and Economics
  
Oral Session: CPT Session 02 (Paper No.6359)
Title: YOLO-DFG: A Lightweight Drone Vehicle Target Detection Algorithm for Remote Sensing Images
Author: Congjing Wang1 and Yongqing Qian2
Abstract: Aiming at the problems of small vehicle target size, complex background, susceptibility to occlusion in remote sensing images, as well as the large number of parameters of existing detection models and the difficulty of deployment on resource-constrained platforms, a lightweight and high-precision remote sensing vehicle target detection model based on an improved YOLOv11 is proposed. In this paper, a DualPoolDown downsampling module is designed to replace the original structure and combined with C3k2-PFDConv (Partial Frequency Convolution C3k2 module) to construct an efficient feature extraction backbone, enhancing the model's ability to extract features of rotationally varying targets. Finally, an LSCDECA Gate (Long-Short Convolution Enhanced Coordinated Attention Gating Mechanism) is introduced in the detection head, which effectively fuses high-level semantic information and low-level spatial details, suppresses complex background noise, and improves the localization accuracy of dense small targets. Experimental results show that, compared with the baseline YOLOv11n model, the proposed algorithm reduces computational cost while maintaining detection performance, providing a new solution for real-time vehicle target detection in drone imagery.
Oral Session: CPT Session 02 (Paper No.6360)
Title: TEFoley: Text-Semantically Enhanced Text-Video-to-Audio Generation
Author: Zhi Cheng1, Cien Fan1 and Lehui Wei1
Abstract: Recent advances in text-video-to-audio (TV2A) generation have made it possible to synthesize high-fidelity audio from multimodal inputs. However, the generated audio aligns with textual semantics remains a significant challenge. Specifically, existing methods heavily rely on synchronization and semantic information provided by visual cues, with text content serving only as a secondary reference, which compromises the overall user experience. To overcome this limitation, we propose TEClip, a visual encoder for text semantic enhancement that captures text-related semantic information from videos. We also construct the VGGSound-TE dataset to train this encoder. Building upon this encoder, we introduce TEFoley, a multimodal audio generation framework that significantly enhances text-based semantic guidance in audio synthesis. Experiments on the VGGSound test set demonstrate that our method achieves state-of-the-art performance in audio generation tasks, including improvements in audio quality, semantic alignment, and temporal synchronization.
Oral Session: CPT Session 02 (Paper No.6364)
Title: Improving Instruction Encoding for Vision-and-Language Navigation via Similarity-Aware Contrastive Learning
Author: Chengyan He1, Jia'ao sun1, Zhaolin Zhang1, Yandong Sun1 and Ling Wang1
Abstract: Vision-and-Language Navigation (VLN) requires agents to understand natural language instructions and align them with vision environments for effective decision-making. However, existing methods often struggle to capture semantic consistency under diverse linguistic expressions, limiting retrieval performance. To address this issue, this paper proposes a BERT-based contrastive learning framework for semantic similarity modeling in VLN. A dual-encoder architecture is adopted to obtain contextualized sentence representations, and a pseudo-label generation strategy combining paraphrase augmentation and in-batch pseudo-labeling is introduced to enhance data diversity without additional manual annotation. Furthermore, a contrastive loss is designed to jointly encourage positive alignment and negative separation, enabling the model to learn more discriminative embeddings. Experiments on the R2R dataset demonstrate that the proposed method achieves superior retrieval performance compared with conventional approaches, validating its effectiveness for semantic representation and instruction retrieval in VLN tasks.
Oral Session: CPT Session 02 (Paper No.6376)
Title: CNN-Transformer Super-resolution Network for Hyperspectral Image Reconstruction
Author: Yuxiao Li1, Yukun Zhang1, Jinhang Yan1, Chang Liu1, Shuo Li1 and Yifan Zhang1
Abstract: To address the issues of insufficient spatial detail recovery, spectral distortion, and low computational efficiency in hyperspectral image super-resolution reconstruction, CNN-Transformer Super-resolution Network (CTSN) integrating CNN and Transformer architectures, is proposed in this paper. The model employs a multi-layer CNN for initial feature extraction, followed by a spatial self-attention mechanism for deep feature extraction, and the high spatial resolution image is then reconstructed through feature fusion. Reconstruction performance employing CNNs with 1 to 5 layers are compared. It is illustrated that a 3-layer CNN exhibits the best overall performance and strongest generalization ability, while a 1-layer CNN performs better in certain scenarios and would be a preferable option for practical applications. Furthermore, a hybrid loss function with weighted L1 loss with spectral angle loss (SAM) is employed to achieve better balance between spatial detail recovery and spectral fidelity. Finally, comparison with current mainstream models illustrate that, the newly proposed model has a smaller reconstruction parameter scale than the mainstream methods and significantly improves computational efficiency, providing an efficient and feasible solution for hyperspectral image super-resolution applications aimed at real-time inference.
Oral Session: CPT Session 02 (Paper No.6378)
Title: Lite-RP-DETR: Lightweight Transformer-Based Rice Pest Detection via Pruning and Distillation
Author: Rongfu Chen1, Canyang Zhou2, Fanlong Zhang2, Caifeng Zou3 and Jianqi Liu3
Abstract: We propose Lite-RP-DETR, a lightweight transformer-based detector for rice pest detection in resource-constrained environments. The framework combines dependency-aware structured pruning with multi-level knowledge distillation to reduce model complexity while preserving accuracy. Experiments show that the proposed method significantly improves efficiency with minimal performance loss, making it suitable for real-time agricultural applications.
Oral Session: CPT Session 02 (Paper No.6416)
Title: Training-Time SAM Knowledge Injection for Deployment-Efficient Metallic Surface Defect Segmentation
Author: Jie Xu1, Zongfang Ma1 and Yun Wang2
Abstract: Metal surface defect segmentation in industrial scenarios faces the challenge of learning stable and generalizable defect representations under limited annotation. Different from existing methods that improve model recognition ability by expanding datasets, introducing transferable segmentation knowledge from vision foundation models and injecting it into classical domain-specific segmentation models provides a promising and low-cost alternative. To this end, this paper proposes Prompt-Consensus Teacher Distillation (PCTD), which leverages the generic region, shape, and boundary priors embedded in SAM to enhance the training of metal surface defect segmentation models. Specifically, PCTD integrates multiple promptconditioned SAM predictions into a consensus teacher map through learnable prompt fusion, and transfers reliable region and contour information to the student model via confidenceaware distillation and boundary prior learning. Experiments on FSSD-12 and Surface Defects-4i show that PCTD improves the strongest plain baseline by 3.11/7.89 and 7.08/9.94 percentage points in IoU/B-F1, respectively, while keeping SAM only in the training stage.
Oral Session: CPT Session 02 (Paper No.6458)
Title: SAR-ShipDetNet: A Multi-Scale Speckle-Robust Oriented Detection Network for Ship Detection in SAR Images
Author: Yida Yang1, Jiewen Tian2 and Yifei Zhang1
Abstract: Ship detection in synthetic aperture radar (SAR) imagery is an important task for maritime surveillance, port monitoring, traffic management, and ocean security. Compared with optical ship detection, SAR ship detection is robust to illumination and weather changes, but it remains challenging due to speckle noise, complex sea--land clutter, small target size, dense harbor distributions, and arbitrary ship orientations. To address these issues, this paper proposes \textit{SAR-ShipDetNet}, a multi-scale speckle-robust oriented detection network for SAR ship detection. The proposed network contains four key components: a SAR-aware stem for shallow speckle-resistant feature extraction, a large-kernel convolutional backbone for elongated ship modeling, a P2--P5 multi-scale feature pyramid for small and medium ship representation, and a decoupled anchor-free oriented detection head for predicting ship center, size, angle, objectness, and category confidence. An optional auxiliary foreground mask branch is introduced to enhance localization and suppress sea--land background interference when segmentation annotations are available. The overall loss combines focal classification loss, objectness loss, rotated-box regression loss, angle regularization, and optional mask supervision.
CPT Session 03
Time: 11:00~12:30, Sunday, July 19, 2026
Session Chair: Zhengqiao Zhao, Northwestern Polytechnical University
  
Oral Session: CPT Session 03 (Paper No.6310)
Title: AI-Driven IoT Framework for Real-Time Predictive Degradation Modeling in Solar Photovoltaic Systems
Author: Atul Anand1, Sudhakar Kumar1, Sunil K. Singh1, Varsha Arya2, Kwok Tai Chui2 and Brij B. Gupta3
Abstract: Solar Photovoltaic (PV) systems experience inefficiency due to environmental factors such as dust, thermal stress, and degradation of materials. This paper develops an IoT-based machine learning prediction framework for the reliability of solar PV systems. By leveraging a dataset of 68,378 records from 2020, including DC power output, solar irradiance, module temperature, and ambient temperature, the framework uses the Decision Tree, Random Forest, and Gradient Boosting techniques provided by the scikit-learn library. It has achieved 96.2% accuracy, 95.8% precision, 94.7% recall, 95.2% F1 score, and 0.99 AUC-ROC values for the Random Forest model. Unlike other studies, which have utilized a limited number of data records, the proposed framework has the ability for real-time monitoring of standalone and grid-connected solar PV systems, reducing the cost of maintenance for the sustainable management of energy resources.
Oral Session: CPT Session 03 (Paper No.6365)
Title: Design, Verification, and ASIC Implementation of a Five-Stage Pipelined RV32I RISC-V Processor with Dynamic Branch Prediction
Author: Hoi Lam Huang1, Hao Yang1, Muhammad Irfan2 and Ray Chak-Chung Cheung1
Abstract: This paper presents the design, verification, and ASIC implementation of a five-stage pipelined RV32I RISC-V processor through an open-source RTL-to-GDSII flow. The architecture incorporates a dynamic branch prediction unit featuring a 16-entry direct-mapped branch target buffer (BTB) and a 2-bit saturating counter branch history table (BHT). Functional verification was conducted using Verilator-based RTL simulation, including directed tests and CoreMark benchmarking, which demonstrated a performance improvement from 0.83 to 1.02 CoreMark/MHz relative to a non-predicting pipeline. The design was implemented to GDSII using OpenLane 1.0.2 and the SkyWater 130 nm (SKY130) process node, integrating two 1 KB SRAM macros for instruction and data memory. Physical implementation achieved timing closure at 100 MHz with a signoff area of 1.34 mm? and an estimated power of 50.1 mW. This work demonstrates that open-source RTL-to-GDSII flows are viable for implementing pipelined processors with microarchitectural enhancements, while providing physical signoff PPA metrics that are absent from prior FPGA-based open-source processor works.
Oral Session: CPT Session 03 (Paper No.6402)
Title: Cooperative Dual-UAV Target Tracking via Hybrid Supervised Pretraining and Residual MADDPG Fine-Tuning
Author: Lina Zeng1, Zhe Zhang1, Haoyu Wang1 and Hang Xu1
Abstract: This paper presents a three-phase hybrid learning framework for cooperative dual-UAV target tracking. The task requires two unmanned aerial vehicles to maintain a prescribed right-angle observation geometry relative to a moving ground target while satisfying inter-UAV safety constraints. Behavior cloning suffers from distribution shift, whereas conventional reinforcement learning suffers from high sample complexity and unstable convergence in multi-agent settings. To address these limitations, the proposed framework first uses a nonlinear model predictive controller to automatically generate expert demonstration data. A Transformer encoder is then pretrained through behavior cloning to provide a stable baseline policy. Finally, a lightweight residual multi-agent reinforcement learning module fine-tunes the baseline policy by learning corrective actions rather than replacing the pretrained controller. Three complementary smoothness mechanisms, including action-rate reward shaping, actor temporal regularization, and test-time exponential moving average filtering, are introduced to suppress high-frequency control jitter without sacrificing tracking responsiveness. Experiments over ten randomized episodes indicate that the proposed method achieves a 100 percent success rate and a mean steady-state angle error of 2.18 degrees, outperforming behavior-cloning-only control, which achieves 40 percent success with a 32.2 degree mean error, and multi-agent reinforcement learning trained from scratch, which achieves 80 percent success with a 6.2 degree mean error using twice the training budget. Ablation studies further confirm the individual contribution of each smoothness mechanism.
Oral Session: CPT Session 03 (Paper No.6420)
Title: NADST-Net: Noise-Adaptive Dual-Domain Shrinkage Transformer Network for Robust HRRP Target Recognition
Author: Songyan XSY XIE1 and Meng YM Yu1
Abstract: Abstract—High-resolution range profile (HRRP) recognition
is an important technique for radar automatic target recognition
because HRRP preserves the one-dimensional scattering
distribution of a target along the radar line of sight. However,
recognition performance degrades significantly in low signal-tonoise
ratio (SNR) environments, where weak scattering centers
are easily submerged by noise and discriminative HRRP structures
are distorted. To address this problem, this paper proposes
a noise-adaptive dual-domain shrinkage transformer network,
termed NADST-Net, for robust HRRP target recognition under
complex noise conditions. The proposed method contains four
core components. First, a lightweight noise estimation head
extracts a noise-state embedding from the input HRRP. Second, a
noise-conditioned multi-scale shrinkage front-end adaptively suppresses
irrelevant components while preserving salient scattering
structures. Third, a dual-domain encoder jointly models local
scattering patterns in the range domain and global dependencies
in the frequency domain. Finally, a noise-aware fusion module
dynamically balances the two domains according to the noise
state. To further enhance noise invariance, a multi-objective
loss is introduced by combining classification loss, reconstruction
loss, supervised contrastive loss, and prediction-consistency
regularization. The experiments evaluate evaluates not only
conventional recognition accuracy under multiple SNR levels, but
also robustness to mismatched noise types, module effectiveness,
and computational efficiency. The proposed framework provides
a unified solution that integrates adaptive denoising, dual-domain
feature learning, and noise-invariant representation learning for
robust HRRP recognition.
Oral Session: CPT Session 03 (Paper No.6422)
Title: Adaptive State Evolution for Ionospheric TEC Forecasting in Emergency Communication Planning
Author: Rongchang Lu1, Guoming Yuan2, Yongjun Cheng3 and jiayu Sheng4
Abstract: Emergency communication planning after earthquakes, floods, or other sudden disasters often depends on GNSS, satellite messaging, and high frequency radio as fallback channels. These links still traverse the ionosphere, where Total Electron Content, or TEC, affects group delay, phase advance, positioning accuracy, and link reliability during space weather disturbances. This paper presents the Adaptive State Evolution (ASE) recurrent forecasting framework for TEC spatial sequence prediction under this combined terrestrial and ionospheric risk. ASE treats hidden state propagation as an input dependent discretization of a continuous dynamic system, allocating short memory steps to storm transients and longer steps to quiet intervals. For deployment, the evaluated pipeline exports forecasts to PostgreSQL/PostGIS as a spatial database and GIS service layer. The evaluation covers Quiet, Storm, Solar Wind Forced, and IONEX Regional TEC scenarios. After discarding models whose accuracy is below 60 percent and Pearson correlation is below 0.5, ASE obtains the best retained OMNI forced result, with 97.40 percent accuracy, 0.9979 Pearson correlation, 0.4668 RMSE, 36.93 dB PSNR, and 0.9893 SSIM. These results indicate that adaptive temporal evolution can support fast TEC nowcasting for emergency communication planning.
Oral Session: CPT Session 03 (Paper No.6424)
Title: Knowledge Distillation for Efficient Underwater Acoustic Classification
Author: Yintao Li1, Zhengqiao Zhao1, Haoxiang Wu1 and Jie Chen1
Abstract: Underwater acoustic monitoring systems require models that achieve high accuracy while remaining lightweight enough for deployment on edge devices. This paper investigates knowledge distillation (KD) from convolutional neural network (CNN)-based teachers to a MobileNetV2 student using a five-class underwater target recognition dataset, ShipsEar. Experimental results show that a lightweight MobileNetV2 model trained from scratch on this dataset attains only 42.94% mean test accuracy, limiting its practical applicability. To improve performance, we evaluate three CNN-based teacher models for knowledge distillation, namely, CNN14, CNN10, and CNN6, and show that teacher selection plays a critical role in knowledge transfer effectiveness. Although CNN14 achieves the highest mean validation accuracy, CNN6 yields the best test performance and produces the most effective distilled student model. Using the response-based knowledge distillation, the CNN6-to-MobileNetV2 framework achieves a mean test accuracy of 55.40% and outperforms all other teacher-student configurations. Performance gains are consistently observed across all five folds, together with a notable reduction in cross-fold variance. These results suggest that a shallower yet architecturally compatible teacher transfers structural knowledge more effectively to a compact student model, demonstrating the potential of MobileNetV2 for real-time underwater acoustic monitoring on resource-constrained devices.
Oral Session: CPT Session 03 (Paper No.6456)
Title: Dual-Branch Image Transmission Network with Nested Vector Quantization for Extremely Low Bandwidth
Author: Shaosong Cao1 and Xuechen Chen1
Abstract: High-fidelity image transmission under extremely low channel bandwidth ratio remains challenging, as digital schemes suffer from cliff effects, while existing deep JSCC methods are limited in preserving fine-grained textures under severe bandwidth constraints. This paper proposes a dual-branch image transmission network with nested vector quantization for low-bandwidth wireless channels. The base branch transmits a low-resolution image to preserve global structure and provide graceful degradation. The detail branch extracts multi-scale residual features and quantizes them with a coarse-to-fine nested vector quantization module, enabling compact transmission of high-frequency information. At the receiver, a CRC-guided multi-stage fusion decoder adaptively incorporates reliable residual features and suppresses corrupted digital information to prevent error propagation. Experiments on the Kodak dataset over AWGN and Rayleigh fading channels demonstrate that the proposed method outperforms DeepJSCC and SwinJSCC under extremely low bandwidth while maintaining robustness when the digital branch fails.
Oral Session: CPT Session 03 (Paper No.6474)
Title: SolarPatchTST: A Linear-Anchored Patch Transformer for Short-Term Photovoltaic Power Forecasting
Author: Ting Kang1, Yue Zhao2 and Luoxiao Yang3
Abstract: Short-term photovoltaic (PV) power forecasting must handle daily regularity, irradiance-driven ramps, and physically nonnegative outputs. Generic Transformer forecasters can model nonlinear temporal patterns, but direct use on compact PV slices may be unstable. This paper proposes SolarPatchTST, a lightweight PV-aware framework that combines a linear PV anchor, a gated PatchTST residual branch, and a nonnegative power projection. A conservative SafeEns inference rule averages the nonnegative anchor stream and residual stream to improve robustness. On public PVDAQ and UNISOLAR datasets with chronological splits, SolarPatchTST-SafeEns reduces MSE by 20.58% and 42.01%, respectively, relative to the strongest compared neural baseline.
  Conference Organizing Committeetee
 
  Steering Committee Co-Chair
  Jianguo HUANG, NWPU
  Jingdong CHEN, NWPU
 
  General Co-Chair
  Deshi Li, Wuhan Univ.
  Ray CHEUNG, City Univ. of HK
  Gongping HUANG, Wuhan Univ. 
 
  Technical Program Co-Chair
  Yulin Hu, Wuhan Univ.
  Chengbing He, NWPU
  Peter Tam, IEEE HK Section
  Technical Pprogram Committee
  Zhongxin Bai, HEU
  Oliver Choi, Chinese Univ. of HK
  Yuzhou Li, HUST 
  Chao Pan, NWPU 
  Kai Pan Mark, Cityu, HK
  Ningning Pan, SWUFE 
  Dongyuan Shi, NWPU 
  Long Shi, SWUFE
  Yingwei Tian, Wuhan Univ.
  Xuehan Wang, GZHU
  Mou Wang, IOA, CAS
  Wenxing Yang, SHU
  Kaili Yin, QUST
  Haijian Zhang, Wuhan Univ.
  Yingke Zhao, SUST
  Xudong Zhao, KCL
 
  Finance Co-Chair
  Yat-Wing LIU, IEEE HK Section
 
  Registration Chair
  Chengkai TANG, NWPU
 
  Publication Chair
  Edward CHEUNG, HK PolyU.
 
  Publicity Co-Chair
  Yuan Gao, Wuhan Univ.
  Lingling Zhang, NWPU  
 
  Student Activity Co-Chair
  Jing Han, NWPU
  Xiaodong CUI, NWPU
 
  Special Section Co-chair
  Wentao Shi, NWPU
  Lianyou Jing, NWPU
 
  Local Arrangement Chair
  Huai Yu, Wuhan Univ.
 
  Secretary
  Si ZHAO, NWPU
 
  Sponsors
  IEEE Xi’an Section
  IEEE Hong Kong Section
  IEEE Wuhan Section
 
  Supporters
  Wuhan University
  Northwestern Polytechnical University
  The City University of Hong Kong
  The HK Polytechnic University.
 
Contact us 
Email: icspcc@163.com 
备案号  
陕ICP备12013028号  

Contact us