Latest Articles

Articles in press have been peer-reviewed and accepted, which are not yet assigned to volumes/issues, but are citable by Digital Object Identifier (DOI).
Display Method:
STAVFT: Spatio-Temporal Contrastive Learning for AIS-Video Fusion Tracking
WU You, ZHANG Mingxuan, CHANG Ren, YANG Kang, TAO Shifei
Available online  , doi: 10.11999/JEIT260236
Abstract:
  Objective  As maritime ship density increases and the demand for intelligent maritime supervision rises, utilizing multi-source heterogeneous sensors to achieve continuous and robust target tracking has become a critical means of enhancing navigational perception. In nearshore ship monitoring, visual modalities are highly susceptible to environmental interference such as low visibility and occlusions. Furthermore, traditional tracking methods struggle with the severe temporal asynchrony between Automatic Identification System (AIS) data and video trajectories, as well as the instability of identity association in multi-target tracking. This research aims to address these challenges by proposing a ship tracking algorithm based on the deep fusion of AIS and video information combined with spatio-temporal contrastive learning, referred to as STAVFT. The primary goal is to leverage AIS spatial priors to guide visual detectors in locking onto targets under low-visibility conditions and to ensure stable identity consistency in complex multi-target scenarios.  Methods  The STAVFT algorithm integrates perception-level enhancement with trajectory-level association to create a unified tracking pipeline. At the detection stage, the algorithm employs a Soft Mask feature enhancement strategy and an ROI re-detection compensation strategy. The Soft Mask strategy maps AIS spatial priors into Gaussian-weighted guidance maps that are injected into the feature extraction and fusion stages of the YOLOv11 detector, forcing the network to focus its representational power on high-probability target regions. For targets missed by the initial detector, the ROI re-detection strategy generates high-confidence candidate windows based on AIS priors to perform secondary detection and recover lost tracks. In the association stage, the algorithm constructs a Spatio-Temporal Aware Sequence Transformer (SAST) encoder to handle the asynchronous nature of AIS and video data. By introducing a continuous time-aware self-attention mechanism, the encoder explicitly models the dynamic intensity of heterogeneous trajectories and maps them into a unified embedding space. Additionally, a Spatio-Temporal Negative-sample Contrastive Estimation (STNCE) loss function is designed to strengthen the discrimination between physically adjacent and confusing targets. This loss function incorporates physical constraints, such as spatial distance and motion similarity, to penalize hard negative samples in the feature space.  Results and Discussions  Experimental results on the FVessel dataset demonstrate that the STAVFT algorithm significantly improves detection performance, raising the recall rate from 65.6% to 79.4% and the mAP50 to 0.792. The SAST encoder handles asynchronous AIS data effectively, achieving a Multi-Object Fusion Accuracy (MOFA) of 0.934, while the STNCE loss successfully mitigates identity swaps in dense traffic by widening the decision boundaries between confusing targets. Overall, the algorithm maintains a stable MOFA above 95.8% and an identity consistency (IDF1) exceeding 96.8%. Although the integration of AIS guidance and ROI re-detection increases the average processing time to 0.334 seconds per video second, the system remains well within real-time requirements, and its adaptive logic provides superior robustness against clutter and track fragmentation.  Conclusions  This paper presents the STAVFT algorithm, which successfully fuses AIS and video information through spatio-temporal contrastive learning to achieve robust ship tracking. By addressing the issues of environmental interference, data asynchrony, and identity confusion, the algorithm significantly improves the recall rate and identity consistency compared to current mainstream methods. The experimental results on the FVessel dataset validate the scientific validity of the "perception-guided association" logic. Future research will focus on model lightweighting for edge-cloud collaboration to further optimize the deployment performance of the system on embedded terminals.
Design of a Low-Cost High-Isolation RF Switch for X-Band Active Phased Array T/R Modules
LI Pan, DUAN Chuangwei, ZHU Wenbo, YANG Lixia, LIAO Guisheng
Available online  , doi: 10.11999/JEIT260326
Abstract:
  Objective  Radio frequency (RF) switches are key front-end components in communication and radar systems, especially in active phased array transmit/receive (T/R) modules. Positive-intrinsic-negative (PIN) diodes are widely used in high-power microwave switches because of their high power-handling capability, good linearity, and mature fabrication process. However, low-cost plastic-packaged PIN diodes exhibit significant parasitic effects at X-band frequencies. Junction capacitance, package inductance, and grounding-via inductance cause shunt branches to deviate from ideal short- or open-circuit conditions, resulting in degraded isolation and increased insertion loss. To address this problem, a low-cost high-isolation X-band single-pole double-throw (SPDT) RF switch design method based on parasitic-parameter compensation is proposed.  Methods  A step-wise synergistic matching strategy, termed “OFF-state prioritized determination and ON-state structural compensation”, is developed. OFF-state isolation and ON-state transmission conditions are treated separately to reduce coupling between isolation and transmission objectives. For the isolation branch, OFF-state diode parasitic parameters, a compensation stub, and grounding-via inductance are combined to form a series-resonant network. By compensating parasitic reactance, the branch satisfies the required impedance condition and suppresses signal leakage. After the compensation-stub parameters are determined, ON-state compensation is realized by introducing equivalent structural capacitance through T-junction edge effects and a stepped-impedance main line. This capacitance, together with the total equivalent inductance of the shunt branch, forms a parallel-resonant network, making the branch approach an open-circuit condition at the operating frequency. Thus, signal shunting is reduced and low-insertion-loss transmission is achieved without additional lumped compensation components.  Results and Discussions  The proposed switch topology is presented, and a full-wave simulation model is established (Fig. 3, Fig. 4(a)). Simulation results show that return loss is better than 20 dB, isolation is higher than 20 dB, and insertion loss remains around 0.2 dB over 8.57–9.34 GHz (Fig. 4(b)). A prototype is fabricated and measured (Fig. 5). Measurement results show that, over 9.5–10.4 GHz, return loss is better than 10 dB, insertion loss is lower than 2 dB, and isolation is higher than 20 dB (Fig. 6). Although the measured operating band shifts upward by about 1 GHz relative to initial simulation results, measured trends remain consistent with theoretical analysis. To explain this frequency shift, OFF-state junction capacitance and ON-state parasitic inductance are further analyzed. As OFF-state junction capacitance decreases from 0.24 pF to 0.195 pF, the high-isolation band shifts to higher frequencies (Fig. 7(a)). With OFF-state capacitance fixed at 0.195 pF, reducing ON-state parasitic inductance from 0.7 nH to 0.2 nH shifts the return-loss response to higher frequencies (Fig. 7(b)). After parameter correction, simulated return loss, insertion loss, and isolation agree more closely with measured results (Fig. 6). Compared with existing SPDT RF switch designs, the proposed design uses low-cost plastic-packaged PIN diodes and a simple microstrip compensation structure, requires no additional lumped compensation components, reduces dependence on the high intrinsic OFF-state impedance of PIN diodes, and achieves a good balance between cost and high-frequency switch performance (Table 1).  Conclusions  A low-cost high-isolation X-band SPDT RF switch based on parasitic-parameter compensation is presented. By using OFF-state series resonance and ON-state structural capacitance compensation, high isolation and low insertion loss are realized simultaneously. The fabricated switch achieves isolation higher than 20 dB and insertion loss lower than 2 dB over 9.5-10.4 GHz. The switch features a simple structure, ease of integration, and low cost, and provides a practical reference for low-cost high-isolation T/R module RF switches in active phased array systems.
PUF-driven Secure Anonymous Authentication Protocol for Internet of Vehicles
WU Linjun, DING Lin
Available online  , doi: 10.11999/JEIT260606
Abstract:
  Objective  In the Internet of Vehicles (IoV), vehicles often send Basic Safety Messages through open wireless links to support road safety and traffic control. However, these messages may expose vehicle identity and travel data. An attacker may collect messages over time and link them to the same vehicle. Thus, vehicle authentication must protect both network security and user privacy. Many current anonymous authentication schemes use complex cryptographic operations, which may cause high computation and communication costs for on-board units (OBUs) with limited resources. Some schemes also store long-term secret keys in vehicle devices. Such keys may be exposed by physical access or side-channel attacks. A Physical Unclonable Function (PUF) uses small process changes formed during chip production to create a unique hardware feature. It can produce device-related responses without storing a secret key directly. Based on this feature, this paper proposes a PUF-driven secure anonymous authentication protocol for IoV. The main goal is to reduce the risk of long-term key storage, protect vehicle identity, and keep the cost of vehicle authentication low.  Methods  The proposed protocol uses a response-feedback-based lightweight anti-machine-learning-attack PUF (FLAM-PUF) as its hardware root of trust. FLAM-PUF combines an Arbiter PUF with a reconfigurable Linear Feedback Shift Register (LFSR). The PUF response is fed back to change the LFSR state, which hides the link between Challenge–Response Pairs (CRPs) and makes machine-learning modeling attacks more difficult (Fig. 1). The system includes a Center Management (CM), Local Management entities (LMs), Roadside Units (RSUs), and vehicle OBUs (Fig. 2). During setup, reliable CRPs are collected in a trusted environment and sent to the CM through a secure channel. The main protocol parameters are listed in Table 1. A PUF challenge is used to form a vehicle pseudonym, while its response is used as authentication or session-key data. The OBU can recover the response from its local PUF when needed, so it does not need to store a long-term secret key. The setup process is shown in Fig. 3. The protocol uses two types of pseudonyms. Long-term pseudonyms are used for network access and identity management, while temporary pseudonyms are used for communication within an LM area. The local access process is shown in Fig. 4. Temporary pseudonyms and session keys are updated with unused CRPs (Fig. 5), and long-term pseudonyms are renewed before they expire (Fig. 6). Security is studied under the Dolev–Yao model by using BAN logic, ProVerif formal verification and informal analysis.  Results and Discussions  The proposed protocol combines PUF-based key generation with dynamic pseudonym management. First, FLAM-PUF uses response feedback and a reconfigurable LFSR to hide the internal CRP relation, which raises the cost of machine-learning modeling attacks (Fig. 1). Second, the use of long-term and temporary pseudonyms separates network access from local vehicle communication. CRPs, pseudonyms, and session keys are updated together, which reduces the chance that an attacker can link vehicle identities across different authentication periods (Fig. 4; Fig. 5; Fig. 6). BAN logic analysis shows that the vehicle can confirm the link between the temporary pseudonym and the PUF-based session key, while the LM can confirm that this link is approved by the CM. ProVerif verification shows that the secrecy queries for the vehicle identity and PUF responses hold, and all three injective authentication correspondences are satisfied. Under the assumed freshness and one-time CRP usage conditions, the results also support the protocol's resistance to replay and impersonation attacks. Informal security analysis shows that the protocol supports mutual authentication, anonymity, traceability, revocability, and protocol-level identity unlinkability. It can also resist replay, impersonation, man-in-the-middle, and false-message attacks (Table 3). In addition, the OBU does not need to keep a long-term secret key in nonvolatile memory because the required PUF response can be recovered when needed. This reduces the risk of direct key extraction. For performance tests, the running times of the main cryptographic operations are measured (Table4), and the operations performed by the OBU, CM, and LM in each protocol stage are listed in Table 4. For one complete online authentication and session-key agreement, the OBU computation time of the proposed scheme is about 1365.99 μs, and its communication cost is 992 bits (Table 6). Under the same test rules, its computation time is close to that of Ref. [14] and lower than those of Refs. [15]-[17]. Its communication cost is also lower than those of all four compared schemes. These results show that the protocol can provide more security functions while keeping the vehicle-side cost low.  Conclusions  This paper proposes a PUF-driven secure anonymous authentication protocol for IoV. FLAM-PUF is used as a hardware root of trust to recover authentication and session-key data when needed, so long-term secret keys do not need to be stored directly in the OBU. The use of long-term and temporary pseudonyms supports anonymous access, identity tracing, revocation, and protocol-level unlinkability. BAN logic, ProVerif formal verification, and informal security analysis show that the protocol meets its main security goals and can resist common network attacks. Performance results show that one online authentication and session-key agreement needs about 1365.99 μs of OBU computation and 992 bits of communication (Table 6). The proposed protocol therefore provides a useful balance among security, privacy, and low cost, and is suitable for resource-limited IoV devices.
COMPASS: An Integrated Computing-Network Routing Mechanism for Spatiotemporal Mismatch in Computing Power Network
XIAO Wei, ZHAO Baokang, SU Jinshu, HUANG Xuefeng, SHI Weijia
Available online  , doi: 10.11999/JEIT260555
Abstract:
  Objective  Global data growth challenges computing infrastructure, while Computing Power Networks (CPN) face a critical limitation: separating network routing from computing offloading yields suboptimal scheduling due to inconsistent computing and network state dimensions. This spatiotemporal resource mismatch causes nearly 20% higher task failure rates in imbalanced versus uniform task arrival scenarios. Furthermore, existing DRL-based scheduling schemes lack topology generalization and flexible action spaces for arbitrary CPN nodes. This work aims to design an end-to-end integrated computing-network routing mechanism to resolve this mismatch, jointly optimize task failure rates and completion times, and enhance CPN dynamic adaptability and service quality.  Methods  First, we construct a centralized software-defined CPN architecture, mathematically modeling task arrivals, network bandwidth, computing resources, and service times, alongside a multi-objective function minimizing task failures and overall service time. Next, we develop the DDQN-based COMPASS algorithm featuring: (1) a 3D state representation integrating network topology, real-time computing status, and task attributes; (2) a joint, truncated action space coupling computing node selection and routing paths to reduce exploration complexity; and (3) a multi-objective weighted reward with heavy penalties for task failure to prevent poor strategy convergence. A tailored training workflow addresses experience delays caused by execution latency, incorporating routing and DRL training algorithms. To validate performance, we developed the open-source CNRSim platform based on EdgeCloudSim, integrating three real backbone topologies, four task types, and mainstream benchmark schemes.  Results and Discussions  Parameter tuning identified an optimal action space dimension of 20 (Fig. 5) and a learning rate of 5.0E-5, effectively balancing convergence speed and final performance (Fig. 6). Evaluated on three real backbone topologies, COMPASS outperformed traditional benchmarks, improving the average reward by 48% (Fig. 7) and achieving the lowest cumulative task failure rate (Fig. 8). Notably, under imbalanced task distributions, it significantly reduced the failure rate from 27.0% to 9.5% (Fig. 9). In-depth analysis on the GEANT2 topology revealed that COMPASS ensures superior computing load balancing across servers (Fig. 10) and processes more tasks with lower bandwidth occupancy (Fig. 11). Failure analysis indicated that jointly optimizing resources mitigates failures caused by bandwidth shortages and node overloads (Fig. 9). Furthermore, COMPASS maintained peak average rewards under dynamic workloads (Table 3).  Conclusions  This paper proposes COMPASS, a DRL-based integrated computing-network routing algorithm, to address CPN bottlenecks caused by spatiotemporal mismatches between resource distribution and task arrivals. By utilizing a 3D state representation, joint action design, and a multi-objective reward mechanism, COMPASS enables global collaborative scheduling. Extensive experiments on the CNRSim platform confirmed its superiority over benchmarks in reward maximization, failure reduction, and load balancing, particularly in imbalanced scenarios. This work offers a framework for CPN optimization and large-scale distributed resource scheduling. Future work will explore task fault tolerance and retransmission mechanisms, and investigate the potential of collaborative optimization by combining COMPASS with emerging intelligent technologies.
Rotatable-Antenna-Array-Enhanced Direction Sensing for Low-Altitude Communication Networks: Method and Performance Analysis
JIANG Jinbing, SHU Feng, ZHENG Weihai, DENG Bin, LI Maolin, BAI Jiatong, WANG Yan, JIANG Hao, WANG Jiangzhou
Available online  , doi: 10.11999/JEIT260580
Abstract:
  Objective  In practical multi-antenna receiving systems, antenna elements usually exhibit anisotropic radiation patterns. However, the impact of such pattern characteristics on direction sensing remains insufficiently explored in both academia and industry. Particularly in extreme scenarios, when the emitter direction significantly deviates from the array boresight or is close to a null of the antenna pattern, the received signal power at the array will be seriously attenuated. Consequently, traditional fixed antenna arrays struggle to meet the dynamic sensing and coverage demands of low-altitude networks. To address this issue, a rotatable antenna array system is developed by accounting for the directional radiation pattern of each antenna element. This research offers a solution for high-precision Direction of Arrival (DOA) estimation of Unmanned Aerial Vehicles (UAVs).  Methods  To achieve high-precision direction sensing, a Recursive Rotation Root-MUltiple SIgnal Classification (RR-Root-MUSIC) algorithm based on a rotatable array architecture is proposed. Specifically, the Root-MUSIC method is first utilized to obtain an initial DOA estimate of the target. Subsequently, the array is rotated toward this estimated direction, and the Root-MUSIC algorithm is applied once more to obtain an updated estimate. This sensing-rotation-resensing loop is repeated iteratively until the difference between adjacent estimates falls below a predefined convergence threshold. Furthermore, to evaluate the performance of the algorithm, the closed-form Cramér-Rao Lower Bound (CRLB) accounting for the directional antenna gain is derived as a theoretical benchmark.  Results and Discussions  To evaluate the proposed RR-Root-MUSIC algorithm, its sensing performance is compared with the derived CRLB under various Signal-to-Noise Ratios (SNRs) and incident angles. Simulation results demonstrate that the proposed method achieves a convergence probability of at least 99.8%, and the required number of iterations decreases as the SNR increases (Figs. 4-6). When the target direction significantly deviates from the array boresight, the proposed method effectively improves the subsequent observation conditions through iterative array rotation. Additional simulations using a non-ideal microstrip patch antenna pattern show that the proposed method maintains relatively stable estimation performance under antenna-pattern variations (Fig. 7). Along the considered UAV flight trajectories, the estimation error remains low and varies smoothly, indicating that array-orientation adjustment can mitigate the effect of UAV position variations on direction-sensing performance (Fig. 8).  Conclusions  This paper develops a rotatable antenna array system for direction sensing of a single UAV in low-altitude communication networks. Based on a sensing-rotation-resensing process, the proposed RR-Root-MUSIC method iteratively adjusts the array orientation according to the current DOA estimate, thereby improving the subsequent observation conditions. Simulation results show that the proposed method alleviates the DOA estimation performance degradation caused by directional-gain attenuation, especially when the target direction significantly deviates from the array boresight. Additional simulations using a non-ideal microstrip patch antenna pattern further verify the effectiveness of the proposed method under the considered non-ideal antenna pattern. The proposed method also exhibits stable convergence performance and maintains relatively low estimation errors along the considered UAV trajectories. These results indicate that rotatable antenna arrays provide a promising approach for low-altitude UAV direction sensing.
Space-Time Joint Clutter Suppression Technology for ISAC Sensing Echo Signals
LIU Yuhan, LIU Suning, ZHANG Haiying, MENG Weixiao
Available online  , doi: 10.11999/JEIT260591
Abstract:
  Objective  The rapid development of low-altitude logistics, urban air mobility, and security monitoring requires Integrated Sensing and Communication (ISAC) systems to detect weak uncooperative “low, slow, and small” (LSS) targets while reusing communication waveforms and hardware. In practical low-altitude environments, echoes from buildings, towers, and ground facilities may be much stronger than target echoes and mask targets during range, angle, and Doppler processing. Static clutter exhibits stable directions of arrival and near-zero-Doppler slow-time components. Conventional moving target indication (MTI) suppresses zero-frequency clutter but may attenuate low-speed targets and distort micro-Doppler modulation. Spatial nulling based on clutter angles preserves low-Doppler information, but it may remove a target together with clutter when their angles are close. Neither approach alone can simultaneously ensure strong clutter rejection, low-speed target preservation, and robustness to angular overlap. Therefore, this study proposes a joint spatial–slow-time clutter suppression method that exploits environmental angular prior information while preserving weak low-speed target signatures.  Methods  After reception of the MIMO-OFDM sensing echo, down-conversion, digitization, cyclic-prefix removal, and FFT processing are performed. Element-wise division by the transmitted communication symbols removes data modulation, and measurements over array elements, subcarriers, and slow-time symbols are stacked into a three-dimensional sensing data matrix. The proposed method comprises offline clutter-angle-map construction and online two-stage suppression. During the offline stage, an inverse FFT along the subcarrier dimension separates range cells in pure clutter measurements. Strong clutter cells are selected, their slow-time snapshots construct spatial covariance matrices, and two-dimensional MUSIC estimates the azimuth and elevation angles of dominant static scatterers. These angles are stored in a clutter angle map (CLAM) as reusable environmental prior information. During online sensing, steering vectors associated with the stored angles form a clutter spatial manifold matrix, from which an orthogonal projection operator is derived to remove spatially separable static clutter. To maintain model consistency, steering vectors for subsequent angle estimation are projected using the same operator. When a target lies close to a stored clutter direction, the CLAM angles are examined individually. For each candidate direction, the remaining CLAM directions are treated as interference and suppressed using a normalized local orthogonal projection, while the current direction is retained. This traversal avoids prematurely removing an angularly overlapped target and provides a candidate slow-time sequence for further discrimination. The residual sequence of each subcarrier is mapped onto the complex I/Q plane, where residual static clutter appears as an approximately fixed DC offset and a moving target forms a circular arc due to slow-time phase evolution. Because low target speeds and finite observation intervals may produce only short arcs, Taubin circle fitting estimates the circle center through a normalized generalized eigenvalue problem. The estimated center represents the residual clutter bias and is subtracted from the complex slow-time samples. Thus, CLAM-guided global projection removes spatially separable clutter, while local angular traversal and circle fitting handle angular overlap and residual slow-time clutter without modifying the target phase trajectory.  Results and Discussions  MIMO-OFDM simulations under the ISAC architecture verify the proposed method in several representative scenarios. When target and clutter angles are separated, the method preserves distinct target peaks in the two-dimensional angular spectrum, and the root mean square errors (RMSEs) of azimuth and elevation estimation are substantially lower than those of conventional MTI (Fig. 4). When target and clutter angles are close, spatial filtering alone may suppress the target together with the clutter and cause missed detection (Fig. 5). The proposed local traversal and circle-fitting stage resolves this limitation and, at a low target speed of 2 m/s, reduces the angle-estimation RMSE to a level close to the ideal clutter-free result, outperforming both MTI and CLAM+FFT (Fig. 6). It also exhibits lower sensitivity to changes in subcarrier spacing than MTI, indicating stronger robustness to communication-system bandwidth configurations (Fig. 7). Time-frequency results further show that MTI causes nonlinear distortion and energy loss in the target echo, whereas the proposed method retains the principal translational component and the periodic rotor-induced micro-Doppler structures required for fine-grained target characterization (Fig. 8).  Conclusions  The proposed joint spatial–slow-time method suppresses strong static clutter while retaining weak LSS target information. It provides accurate angle estimation in both angularly separated and closely spaced target–clutter scenarios, with a particularly clear advantage for low-speed targets. By estimating and removing the residual DC bias through circle fitting instead of Doppler-domain high-pass filtering, it avoids unnecessary target-energy attenuation and preserves micro-Doppler structures needed for subsequent classification and recognition. The method therefore provides purified, high-fidelity echoes for downstream low-altitude target perception. Future work will investigate dynamic clutter broadening in measured environments and intelligent target classification using real-world echo data.
WiFi RFFID: A Lightweight Temporal Convolutional Network Integrating Multi-scale and Channel Attention
TIAN Xinyu, ZHANG Xianshou, ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260599
Abstract:
  Objective  With the widespread deployment of WiFi in the industrial Internet of Things (IoT), enterprise wireless local area networks, and smart homes, its security vulnerabilities have become increasingly prominent. The radio frequency fingerprint identification (RFFID) leverages hardware-level intrinsic variations introduced during manufacturing to authenticate devices, offering a robust solution for physical-layer security. However, existing RFFID methods often struggle with the trade-off between recognition robustness in dynamic channel environments and computational efficiency for deployment on resource-constrained platforms. This paper aims to develop a lightweight and efficient RFFID method that achieves high accuracy under varying channel conditions and device scales while maintaining low computational complexity suitable for embedded systems.  Methods  This paper proposes WiFi-RF-LTCN, a lightweight temporal convolutional network designed for WiFi RFFID. The method processes the legacy long training field (L-LTF) sequence from WiFi physical layer preamble. First, multi-dimensional features, including I/Q components, magnitude, and phase, are extracted to form a four-channel input. The core network architecture comprises two parallel branches: a temporal dilated convolution branch with residual connections to capture the long-range temporal dependencies across the signal, and a multi-scale learnable convolution branch utilizing kernels of varying sizes to enhance sensitivity to short-term local waveform distortions. Subsequently, a channel attention mechanism based on the squeeze-and-excitation block adaptively fuses features from both branches, highlighting discriminative fingerprints while suppressing redundant information. Finally, the model is trained and evaluated on a comprehensively simulated WiFi dataset incorporating diverse transmitter impairments and multipath channel effects.  Results and Discussions  Extensive experiments demonstrate that WiFi-RF-LTCN achieves the superior performance across various conditions. It attains high recognition accuracies of 93.12% at 0 dB SNR and 99.27% at 20 dB SNR. The model maintains robust performance even with limited training data (e.g., 82.21% accuracy with only 50 samples per device) and scales effectively as the number of devices increases. Ablation studies confirm the necessity of each component, with the complete model achieving average accuracy of 96.3% with the inference speed of 0.41 ms, significantly outperforming configurations missing any single module. Crucially, WiFi-RF-LTCN surpasses Transformer, ResNet50, TCN, 1D-CNN, and LSTM by 2.84%, 2.36%, 1.40%, 2.62%, and 3.61%, respectively. Moreover, it accomplishes this with only 0.17 M parameters and a remarkably low inference time, substantially reducing the computational overhead compared to larger models like ResNet50 (23.54 M) and Transformer (0.85 M) with inference time of 3.9381 ms and 0.8953 ms, respectively.  Conclusions  This paper presents a novel lightweight temporal convolutional network, WiFi-RF-LTCN, for WiFi device RFFID. By integrating multi-scale feature extraction, temporal modeling, and channel attention, the method effectively captures subtle hardware-induced distortions while demonstrating the strong robustness to noise and channel variations. The experimental results validate that WiFi-RF-LTCN achieves an optimal balance between high recognition accuracy and low computational cost, significantly outperforming existing deep learning methods. Its minimal parameter count and fast inference time make it highly suitable for real-time deployment on resource-constrained embedded platforms, offering the promising solution for enhancing physical-layer security in IoT and other wireless applications. Future work can further combine channel characteristics with adaptive optimization strategies to enhance the model’s generalization ability and deployment adaptability.
A Modulation Recognition Method Based on Gated Recurrent Network Integrating Adaptive Denoising and Aggregation Attention
ZHENG Qinghe, LI Binglin, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260213
Abstract:
  Objective  In wireless communication environments under low signal-to-noise ratio (SNR) conditions, the separability of signal features deteriorates, and nonlinear distortion caused by multipath effects poses significant challenges for automatic modulation classification (AMC). Existing deep learning models often suffer from insufficient robustness and generalization under low SNR, while Transformer-based models incur high computational complexity. To address these issues, this paper proposes a gated recurrent network for modulation recognition that integrates adaptive denoising and aggregation attention, aiming to enhance recognition accuracy in low-SNR conditions while maintaining low model complexity.  Methods  The proposed AMC method consists of three key components. Firstly, an adaptive denoising unit based on a convolutional Kolmogorov-Arnold network (CKAN) is designed to decouple the nonlinear superposition of signals and noise, providing purified input for subsequent feature extraction. The CKAN combines the nonlinear approximation capability of the Kolmogorov-Arnold representation theorem with the local feature extraction ability of convolutional operations. Then an aggregation attention mechanism is introduced to reduce the computational complexity of traditional attention by employing aggregation matrix. Based on this mechanism, a talk-attention module and an attention-gated recurrent unit (Att-GRU) are constructed to enhance cross-head information interaction and long-term temporal dependency modeling, thereby mitigating the impact of multipath interference. Finally, an orthogonal dual-channel strategy and multi-scale feature fusion architecture are adopted to integrate globally denoised features with local time-frequency features of I/Q components, improving the discriminative ability of essential signal characteristics.  Results and Discussions  The proposed method achieves average classification accuracies of 64.36% and 71.52% on the RadioML 2016.10a and RML22 datasets, respectively, with only 0.23M parameters and an average inference time of 9.1 ms (Table 2). Comparative experiments with existing deep learning models demonstrate the superiority of the proposed approach. It outperforms AWN, MCDformer, FE-SKVIT, AMC-NET, and MCLDNN on RadioML 2016.10a by 2.08%, 1.83%, 1.03%, 1.96%, and 2.15% in average accuracy, respectively (Table 2). When SNR > 8 dB, the accuracy exceeds 90% on both datasets, and reaches nearly 100% on RML22 when SNR > 16 dB. Even at low SNR (–10 dB), the model maintains robust accuracies of 48.5% and 33.0% on the two datasets (Figure 5). The confusion matrix at SNR = 0 dB reveals that classification difficulties primarily occur between QPSK and 8PSK, AM-DSB and WBFM, and among high-order QAM modulations, due to phase distortion, spectral similarity during silent carrier periods, and reduced symbol spacing (Figure 6). Significant improvements are also observed under low SNR conditions (Figure 7). Ablation studies confirm the contribution of key components. Removing the adaptive denoising unit reduces accuracy by 2.15–4.30% on RadioML 2016.10a and 3.46–6.02% on RML22, while increasing inference time only marginally (Table 3). The CKAN configuration with 16/4 channels achieves the best trade-off between accuracy and efficiency (Figure 8).  Conclusions  In this paper, we present a gated recurrent network that integrates adaptive denoising and aggregation attention for AMC. The method effectively addresses performance degradation in low-SNR wireless communication environments, enhances adaptability to complex channel conditions, and reduces deployment costs. Experimental results on two public datasets validate the superiority of the proposed approach in terms of accuracy, model size, and inference efficiency, especially in challenging low-SNR scenarios. Future work will focus on optimizing the computational efficiency of CKAN-based components to further accelerate inference.
Topology Optimization Method for UAV Swarm SAR Using Illuminators of Opportunity from LEO Communication Satellites
ZHOU Song, DU Xiankang, WANG Yuhao, WEN Pin, YANG Lei, XING Mengdao
Available online  , doi: 10.11999/JEIT260676
Abstract:
  Objective  Low Earth Orbit Communication Satellite Constellations (LEO-COS) provide globally distributed, continuously available communication signals, which offer a promising opportunity for passive Synthetic Aperture Radar (SAR) imaging. However, when LEO-COS signals are used as opportunistic illuminators, the limited effective bandwidth and relatively low Signal-to-Noise Ratio (SNR) restrict the imaging capability of conventional single-receiver SAR systems. To overcome these limitations, this paper proposes a novel “constellation–swarm” cooperative SAR imaging mode that integrates LEO-COS illuminators with an Unmanned Aerial Vehicle (UAV) swarm. In this mode, multiple UAV receivers cooperatively collect echo signals, enabling spectrum extension and coherent energy accumulation. Since the spatial configuration of the UAV swarm directly determines the wavenumber spectrum distribution and ultimately affects SAR imaging quality, it is necessary to optimize the UAV swarm configuration under dynamic and heterogeneous bistatic observation geometries. This study aims to establish a configuration optimization method for UAV swarm SAR using LEO-COS opportunistic illumination, thereby improving spectrum utilization, imaging resolution, and robustness in complex observation scenarios.  Methods  The relationship between UAV swarm spatial configuration and SAR imaging performance is first analyzed from the perspective of wavenumber spectrum geometry. Based on the echo model of the “constellation–swarm” SAR system, the influence of UAV positions and motion parameters on spectrum distribution is derived. The configuration design problem is then formulated as a multi-objective optimization problem by considering spectrum orthogonality, spectrum parallelism, spectrum misalignment suppression, and minimum imaging SNR. To solve this high-dimensional and strongly coupled optimization problem, an Associated-Structure-Encoding Nondominated Sorting Genetic Algorithm II (AS-NSGA-II) is proposed. In the proposed algorithm, the associated structure encoding mechanism is used to preserve the geometric coupling relationship among configuration variables during crossover and mutation. Layered Cubic chaotic mapping is introduced to initialize the population, which improves the uniformity and diversity of candidate solutions in the search space. In addition, an adaptive mutation strategy is designed to dynamically adjust the mutation probability during evolution, so that the algorithm can balance global exploration and local exploitation while maintaining configuration feasibility.  Results and Discussions  Simulation results show that the optimized UAV swarm configurations can effectively rearrange the wavenumber spectrum support areas of different receiving nodes, reduce spectrum overlap and misalignment, and improve spectrum stitching quality(Fig. 5). In comparison with conventional Genetic Algorithm (GA) and Particle Swarm Optimization (PSO) methods, the proposed AS-NSGA-II algorithm shows faster and more stable convergence. The convergence curves of the four objective functions demonstrate that AS-NSGA-II reaches a stable optimization state after approximately 60 generations, while GA and PSO exhibit slower convergence and larger fluctuations(Fig. 6). Area target imaging results further show that the proposed method suppresses false grating lobes more effectively and preserves the target edge and structural features more clearly(Fig. 7). Under low-SNR conditions, the optimized configuration maintains reliable imaging quality through multi-aperture coherent integration, demonstrating better robustness than the comparison methods.  Conclusions  This paper investigates the UAV swarm configuration optimization problem for “constellation–swarm” cooperative SAR imaging using LEO-COS signals as opportunistic illuminators. By analyzing the coupling relationship between spatial configuration and wavenumber spectrum distribution, a multi-objective optimization model is established to jointly consider geometric spectrum constraints and imaging SNR. The proposed AS-NSGA-II algorithm improves the optimization process by incorporating associated structure encoding, Layered Cubic chaotic initialization, and adaptive mutation. Simulation results demonstrate that the proposed method outperforms conventional GA and PSO methods in convergence speed, optimization stability, spectrum stitching performance, and imaging quality. The optimized UAV swarm configuration can effectively compensate for the bandwidth and SNR limitations of LEO-COS opportunistic illumination and enhance the resolution and robustness of cooperative SAR imaging. This work provides a feasible technical approach for future wide-area, high-resolution, and robust SAR imaging based on integrated satellite–UAV swarm systems.
Research on Multimodal Sentiment Analysis Method Based on Adversarial Optimization and Triplet Soft Contrastive Learning
KANG Shouqiang, ZHANG Bohao, XIE Jinbao
Available online  , doi: 10.11999/JEIT260255
Abstract:
  Objective  Multimodal sentiment analysis has become an important research topic in the field of human–computer interaction, as it aims to understand human emotions by integrating information from multiple modalities such as text, audio, and video. However, existing methods often suffer from two key limitations. First, many approaches neglect the temporal dependency relationships among multimodal features, which leads to insufficient utilization of sequential information. Second, conventional multimodal fusion strategies usually rely on simple feature concatenation, which limits cross-modal interaction and may introduce redundant features. Moreover, due to the continuous and ambiguous nature of sentiment expressions, traditional hard contrastive learning strategies tend to treat sample relationships in a binary manner, which fails to capture the subtle semantic differences between samples. To address these challenges, paper proposes a multimodal sentiment analysis method based on adversarial optimization and triple soft contrastive learning. The proposed framework aims to enhance multimodal feature representation by modeling temporal dependencies and by introducing a soft contrastive learning mechanism that assigns continuous similarity weights between samples. Meanwhile, adversarial optimization is incorporated to improve the robustness and discriminative capability of the learned representations.  Methods  The proposed approach constructs a triple soft contrastive learning framework integrated with adversarial optimization. The overall architecture consists of three main components: multimodal feature extraction, temporal contrastive representation learning, and adversarial optimization. First, modality-specific encoders are used to extract representations from textual, acoustic, and visual inputs. These features are then aligned in a shared representation space to enable multimodal interaction. To capture temporal dependencies within each modality, a temporal contrastive learning strategy is introduced to model relationships between sequential features and encourage temporally consistent representations. Second, a soft contrastive learning mechanism is adopted to overcome the limitations of conventional hard contrastive learning. Instead of using binary labels to distinguish positive and negative pairs, the proposed method assigns continuous similarity weights according to the sentiment distance between samples. Design allows the model to better capture gradual emotional transitions and reduces the impact of ambiguous sentiment boundaries. Third, an adversarial optimization strategy is introduced to further enhance the robustness of the representation learning process. By constructing a dynamic adversarial training mechanism between representation learning and feature perturbation, the model is able to learn more discriminative and stable multimodal features. The overall training objective integrates the soft contrastive loss, temporal contrastive loss, and adversarial optimization process to achieve more effective multimodal representation learning. The overall framework of the proposed model is illustrated in 图 2.  Results and Discussions  To evaluate the effectiveness of the proposed method, extensive experiments are conducted on widely used multimodal sentiment analysis datasets. The experimental results demonstrate that the proposed adversarial optimization based triple soft contrastive learning framework significantly improves the performance of multimodal sentiment prediction. Compared with several representative baseline models, the proposed approach achieves competitive results in terms of classification accuracy, F1-score, and regression metrics. The experimental comparisons with existing multimodal methods are summarized in表 and表. The results show that incorporating soft contrastive learning enables the model to capture fine-grained emotional relationships between samples, while adversarial optimization enhances the robustness of feature representations against noisy inputs. Furthermore, ablation studies indicate that each component of the proposed framework contributes to performance improvement. In particular, the soft contrastive learning mechanism effectively alleviates the problem of hard boundary assumptions in traditional contrastive learning, while temporal contrastive modeling helps preserve sequential semantic information. The adversarial optimization module further strengthens the discriminative ability of the learned features. These findings confirm the effectiveness of integrating adversarial learning with contrastive representation learning for multimodal sentiment analysis.  Conclusions  TPaper proposes a multimodal sentiment analysis method based on adversarial optimization and triple soft contrastive learning. The proposed approach introduces a soft contrastive learning strategy to model continuous sentiment similarity between samples, while temporal contrastive learning captures sequential dependencies in multimodal data. In addition, adversarial optimization is incorporated to enhance representation robustness and improve generalization ability. Experimental results on benchmark datasets demonstrate that the proposed method achieves competitive performance and effectively improves multimodal sentiment representation learning. The proposed framework provides a promising direction for future research on robust and fine-grained multimodal sentiment analysis.
A Multi-Agent Active-Inference Collaborative Decision-Making Method for Space-Air-Ground Integrated Networks
LIN Jiaqi①②, WANG Yong③, LIN Xin④, YAN Shi①②, XU Xin④, GU Jiangchun④, ZHANG Senbai⑤
Available online  , doi: 10.11999/JEIT260727
Abstract:
  Objective  Space-Air-Ground Integrated Networks (SAGINs) combine satellites, aerial platforms, terrestrial access points, and edge nodes to support heterogeneous services in highly dynamic domains. Collaboration is difficult under local, delayed observations, time-varying topology, restricted raw-data exchange, and coupled spectrum, power, and computing resources. Existing multi-agent reinforcement learning (MARL) and game methods do not jointly address uncertainty-aware semantic consensus, hierarchical resource allocation, cross-domain adaptation, and lightweight deployment. This study develops a Collective Active Inference (CAI)-based intelligence-generation/transmission loop in which agents exchange inferred beliefs rather than raw observations, consensus beliefs drive resource decisions, and learned models are adapted and compressed for constrained satellite and edge nodes. The scope is a time-varying, partially observable SAGIN evaluated in an Open Radio Access Network (O-RAN) system-level simulation, not an operational deployment.  Methods  The problem is formulated as a time-varying decentralized partially observable Markov decision process over channel, queue, congestion, interference, and node states. The formulation explicitly captures decentralized decisions under incomplete information and intermittent connectivity. Multi-agent cognition is modeled by minimizing Collective Variational Free Energy (CVFE): a likelihood term preserves local observation fidelity, while a Kullback-Leibler term penalizes disagreement between neighboring beliefs. A Gaussian precision-weighted update favors reliable beliefs. Symmetric Metropolis weights yield a doubly stochastic mixing matrix; consensus error decreases geometrically with the second-largest eigenvalue modulus and extends to intermittent links under B-joint connectivity. Agents exchange only belief means/covariances and act on the converged belief (Algorithm 1). A vertical-horizontal game maps consensus to resources. An operator serves as the Stackelberg leader, while service slices respond to prices; horizontally, slices form a generalized Nash game under shared spectrum, power, and service-level constraints. Cross-domain evolution combines differentially private federated learning, Top-10% update sparsification, Model-Agnostic Meta-Learning, and knowledge distillation. The distillation temperature is 4, with supervised and distillation weights of 0.3 and 0.7. Evaluation uses an OMNeT++ O-RAN simulator with 12 coverage nodes, 19 underlay nodes, 72 links, a 28 GHz carrier, and 100 MHz bandwidth. Concurrent Ultra-Reliable and Low-Latency Communications, enhanced Mobile Broadband, and massive Machine-Type Communications traffic is simulated for 10 000 control periods. The network-load factor rises from 0.2 to 1.0, while frequency-sweeping interference and traffic bursts test out-of-distribution recovery. Baselines use the same observations, action space, constraints, training budget, and random-seed set; values are averaged over at least 10 independent runs with 95% confidence intervals.  Results and Discussions  At full load, the method reaches 93.2 Mb/s throughput, versus 81.8, 88.5, and 78.4 Mb/s for symmetry-informed MARL (SI-MARL), quantum MARL (QMARL), and graph-attention-network-based MARL (GAT-MARL). Its task-scheduling success rate is 0.92, its 95th-percentile latency is 20.5 ms rather than 31.2, 26.4, and 33.0 ms, and its service-level-agreement violation rate is approximately 0.081 (Fig. 3). After an out-of-distribution disturbance, the retained key-performance-indicator level is about 0.83 and recovery to the 95% threshold requires 23 control periods, compared with 35 and 41 for QMARL and GAT-MARL (Fig. 4). In resource competition, the mean scheduling success rate is 0.92 with a 95% confidence interval of [0.90, 0.94], exceeding three hierarchical-game or federated-game baselines ranging from 0.80 to 0.85 (Fig. 5). With 50 agents and an update dimension of 80, Top-10% sparsification reduces transmitted updates from 321.79×103 to 34.58×103 scalars, an 89.3% reduction, while convergence rounds rise from 58.84 to 71.26, or 21.1% (Table 1). For few-shot relation inference, F1 scores are 0.71 with 20 labels and 0.98 with 80 labels; about 48 labels reach 0.90, compared with 75, 90, and 120 for the baselines (Fig. 6). Distillation reduces model size from 1 050 KB to 160 KB and inference latency from 21 ms to 15 ms. The student retains 91% task success and 93% path efficiency, while the teacher achieves 95% and 97%; removing environmental-domain information causes the largest path-efficiency loss (Table 2).  Conclusions  The proposed CAI method connects precision-weighted semantic consensus, hierarchical resource coordination, privacy-preserving model evolution, few-shot adaptation, and lightweight deployment in one closed loop. Theory relates consensus speed to graph connectivity, while the O-RAN simulation shows improvements in throughput, tail latency, reliability, disturbance recovery, communication efficiency, and deployment cost under the stated settings. Limitations include simulated channels, idealized observation quality, synchronous updates, and modeled prior distributions that may differ from operational conditions. These findings are not performance guarantees for arbitrary operational networks; validation with measured channels, hardware-in-the-loop tests, and field experiments remains necessary.
Dynamic Multi-objective Optimization for Rumor Control in Time-varying Social Networks
LI Jiting, SUN Yi, GUO Hao, SONG Yanjie, OU Junwei, JIA Jun
Available online  , doi: 10.11999/JEIT260646
Abstract:
  Objective  The rapid dissemination of rumors in online social networks (OSNs) poses significant threats to societal stability. Counter-rumor strategy, which combats misinformation by proactively spreading factual information, is a promising non-intrusive countermeasure. Its essence is a resource-constrained optimization problem: selecting an optimal set of seed users to initiate the counter-rumor, aiming to simultaneously minimize the rumor's final influence and the intervention cost. However, existing research predominantly relies on a static network assumption, neglecting the intrinsic dynamic nature of OSNs where user influence, activity, and susceptibility evolve over time and are perturbed by external events. Strategies optimized for a static snapshot of the network often fail when the environment changes, highlighting a critical gap. This work addresses this gap by formally modeling the problem as a Dynamic Multi-objective Optimization Problem (DMOP). The necessity lies in developing an adaptive framework that can track the time-varying optimal trade-off between control effectiveness and resource expenditure, which is essential for robust and practical rumor governance systems.  Methods  First, a Dynamic Belief-driven Competitive Cascade (DyBCC) model is proposed. Its core innovation is the time-varying user belief, B(v,t), which synthesizes static structural influence with a dynamic component driven by natural decay and external event shocks (Definition 2.1), providing a realistic stochastic simulation framework for competitive propagation under dynamic environments. Second, based on DyBCC, the counter-rumor seed selection is formalized as a Dynamic Multi-objective Optimization Problem for Rumor Control (DMOP-RC), with objectives to minimize expected rumor impact f1 and seed cost f2, where both the objective function and the Pareto-optimal set change over time. Third, to solve DMOP-RC, a Prediction-Memory guided Multi-population Evolutionary Algorithm (PM3EA) is designed. It maintains three collaborative populations: a main population Pm using MOEA/D for local exploitation, a memory archive Pa storing historical elites, and an exploration population Pe using Differential Evolution for global search. A key mechanism is its adaptive response to detected environmental changes (Algorithm 1).  Results and Discussions  The proposed framework is evaluated in a dynamic environment constructed from a real-world OSN dataset. Algorithm Performance: PM3EA is compared against three dynamic multi-objective optimizers. In terms of the comprehensive Dynamic Inverted Generational Distance (DIGD), PM3EA achieves a significantly superior value of 0.0996±0.0569 (Table 2, Fig. 3b). Visually, the final Pareto front obtained by PM3EA is the closest to the reference front and shows the best distribution (Fig. 3a). Its DIGD value remains consistently the lowest and most stable across all time windows, demonstrating robust tracking capability (Fig. 3c). Model and Strategy Validation: A detailed case study provides micro- and macro-level insights. The evolution of individual node beliefs visually demonstrates the dynamic competition captured by the DyBCC model, showing patterns like belief oscillation in rumor seeds and sudden "clearance" in some nodes upon effective counter-rumor exposure (Fig. 5). At the macro level, the effectiveness of the rumor control method is demonstrated on three datasets of different scales (Fig. 6), quantitatively proving the suppression effect of the derived seed set.  Conclusions  This work systematically tackles adaptive rumor control in dynamic social networks. The primary contributions are: 1) The DyBCC model effectively captures the core dynamic of information competition by integrating time-varying user belief, providing a more realistic foundation than static models. 2) The DMOP-RC formulation correctly frames the seed selection as a dynamic trade-off, aligning with practical needs. 3) The PM3EA algorithm, with its tri-population synergy and adaptive change-response mechanisms, demonstrates superior performance in tracking the time-varying Pareto front, outperforming established counterparts. The experiments, from algorithm comparison to case study, holistically validate the effectiveness and superiority of the proposed framework. This work provides a complete methodology for developing intelligent, self-adaptive online rumor governance systems. Future work may explore integrating more complex user behavior models and online learning mechanisms.
Dynamic Frequency Guidance and Semantic Purification Network for Infrared Dim and Small Target Detection
SHEN Xiaoru, CHANG Xia, WEI Wenjie
Available online  , doi: 10.11999/JEIT260582
Abstract:
  Objective  Infrared dim and small target detection remains challenging because targets are small, have low contrast, and are easily obscured by complex background clutter. Although existing methods improve detection performance, several limitations remain, including feature loss during downsampling, insufficient use of physical priors, and contamination of semantic features by background noise. These limitations can result in missed detections and false alarms. To address these problems, a detection network combining dynamic frequency guidance and semantic purification is proposed. Deep semantic features are dynamically coupled with adaptive frequency-domain priors to strengthen target and edge representations, improve detection accuracy, reduce false alarms, and enhance robustness in complex scenes.  Methods  A framework consisting of a shared encoder and a dual-branch decoder is constructed for infrared dim and small target detection (Fig. 1). First, a Local Target Enhancement (LTA) module is used to preprocess the original infrared image, strengthen small-target boundaries, and suppress background noise. The Sobel operator is then applied to extract edge information and provide a structural prior for subsequent feature learning. During feature extraction, a Learnable Downsampling (LD) module replaces conventional max pooling. Stride-2 convolution adaptively compresses the feature maps while preserving target details and reducing information loss. Residual blocks and progressive multiscale feature fusion are also used to strengthen semantic representations at different levels. A Dynamic Frequency Guidance Module (DFGM) is then introduced to enhance high-frequency details and target-edge information. Deep semantic features are used to dynamically predict the cutoff radius of a high-pass filter. A content-adaptive high-pass filter mask is constructed in the frequency domain to suppress low-frequency components. The filtered features are subsequently transformed back to the spatial domain to generate a frequency guidance map. To suppress spurious background responses introduced during frequency-domain enhancement, a Semantic Purification Module (SPM) is further designed. The deepest semantic features are used to generate a spatial attention map for top-down unidirectional filtering of the frequency guidance features. Background interference is thereby suppressed without contaminating high-level semantic representations with low-level noise. Finally, the dual-branch decoder reconstructs the target region and edge structure, thereby improving segmentation consistency, target boundary localization, and target-shape recovery.  Results and Discussions  Experimental results on the SIRST-Aug and IRSTD-1k datasets show that the proposed method outperforms mainstream comparison methods. On SIRST-Aug (Table 1), the method achieves the best Iintersection over Union (IoU), normalized Intersection over Union (nIoU), area under the receiver operating characteristic curve (AUC), and probability of detection (Pd), reaching 76.66%, 73.20%, 94.45%, and 99.17%, respectively. The false-alarm rate (Fa) is maintained at 33.96 × 10–6. On IRSTD-1k (Table 1), the method achieves the best IoU, AUC, and Fa values of 67.17%, 89.52%, and 10.62 × 10–6, respectively, while obtaining an nIoU of 67.67% and a Pd of 92.59%. Qualitative comparisons further show that the proposed method provides more accurate target localization and more complete target-shape segmentation under complex background interference (Figs. 4 and 5). The model contains 2.94 M parameters, requires 14.68 G FLoating-point OPerations (FLOPs), and has an average inference time of 7.54 ms, providing a favorable balance between detection performance and computational cost (Table 2). Ablation experiments further confirm the contributions of the key modules (Table 3). Adding DFGM increases AUC from 93.17% to 94.60%. Introducing LD increases nIoU to 73.32%, and adding SPM further increases IoU to 76.66%. The loss-function ablation results also show that the combined mask and edge losses provide a better balance among segmentation accuracy, target boundary localization, and false-alarm suppression (Table 4). Overall, the proposed method provides strong detection performance and effective false-alarm suppression.  Conclusions  An infrared dim and small target detection network based on dynamic frequency guidance and semantic purification is proposed. Through the coordinated use of LD, DFGM, and SPM, target details are effectively preserved, high-frequency structural information is strengthened, and semantic representations are purified. Background interference is consequently suppressed, and target-boundary recovery is improved. Experiments on SIRST-Aug and IRSTD-1k demonstrate that the proposed method outperforms existing mainstream methods across multiple evaluation metrics and maintains strong robustness in complex scenes. Future work can incorporate richer physical priors and self-supervised learning strategies to further improve model generalization and robustness under challenging conditions.
Secrecy Performance Analysis of Multi-tag Bistatic Backscatter Communication Systems With Outdated CSI and Link Correlation
LIU Yingting, TANG Yong, LI Xingwang
Available online  , doi: 10.11999/JEIT260823
Abstract:
  Objective  Due to feedback delay, the Channel State Information (CSI) used during tag selection may become outdated before data transmission, causing a mismatch between the selected tag and the tag with the largest backscatter-link channel gain during transmission. Most existing studies of outdated CSI assume independent and identically distributed (i.i.d.) channels, which may not adequately reflect the heterogeneous characteristics of practical links caused by different propagation distances. In addition, when the eavesdropper is close to the destination, the legitimate and eavesdropping links may experience correlated fading. Neglecting these factors may cause theoretical results to deviate from actual system performance. Accordingly, under independent and non-identically distributed (i.n.i.d.) channel conditions, the secrecy performance of a multi-tag Bistatic Backscatter Communication (BBC) system is studied under the joint effects of outdated CSI and correlation between the legitimate and eavesdropping links.  Methods  Candidate tags are ranked according to their backscatter-link channel gains, and the tag with the largest backscatter-link channel gain is selected for transmission. An outdated CSI model is introduced to characterize the mismatch between the CSI used during tag selection and the CSI during data transmission. Because tag selection depends only on the backscatter-link CSI, CSI aging is modeled only for this link. A correlated Rayleigh fading model is used to characterize the statistical dependence between the legitimate and eavesdropping links. Under i.n.i.d. Rayleigh fading, order statistics are used to derive the Probability Density Function (PDF) of the selected tag’s outdated backscatter-link channel gain, whereas the joint PDF of the legitimate- and eavesdropping-link channel gains is derived by incorporating their correlated fading relationship. Based on these results, closed-form and high-transmit-power asymptotic expressions for the Secrecy Outage Probability (SOP) are derived. The secrecy performance is further analyzed in terms of the legitimate-to-eavesdropping channel-gain ratio.  Results and Discussions  Monte Carlo simulations validate the analytical and asymptotic results. The results show that outdated CSI significantly degrades secrecy performance because feedback delay causes the CSI used during tag selection to differ from the CSI during transmission, so the selected tag may no longer provide the largest backscatter-link channel gain. For a fixed legitimate-to-eavesdropping channel-gain ratio, a secrecy outage floor emerges at high transmit power (Fig. 2), indicating that increasing transmit power alone cannot eliminate this performance bottleneck. In contrast, increasing the legitimate-to-eavesdropping channel-gain ratio effectively mitigates the outage floor and yields a secrecy diversity order of 1 with respect to this ratio (Fig. 4). Under the considered system model, correlation between the legitimate and eavesdropping links also improves secrecy performance (Fig. 3) by reducing the probability that the legitimate link experiences severe fading while the eavesdropping link remains strong. Moreover, despite outdated CSI, the proposed tag selection scheme based on backscatter-link channel-gain ranking remains effective and consistently outperforms random tag selection (Fig. 3).  Conclusions  A secrecy-performance analysis framework is developed for multi-tag BBC systems with outdated CSI and correlated legitimate and eavesdropping links. Closed-form SOP and high-transmit-power asymptotic expressions characterize the effects of outdated CSI, link correlation, and the legitimate-to-eavesdropping channel-gain ratio. The results identify outdated CSI as a major source of secrecy degradation and indicate that low-latency feedback is beneficial. Increasing the legitimate-link gain advantage over the eavesdropping link and selecting tags according to backscatter-link channel-gain ranking effectively improve secrecy performance.
Resource Allocation and Node Deployment for Multi-UAV Integrated Localization and Communication with Differentiated Requirements
ZHAO Yicheng, WANG Hai, QIN Zhen, SUN Weihao
Available online  , doi: 10.11999/JEIT260819
Abstract:
  Objective  In areas with weak ground-network coverage or limited Global Navigation Satellite System (GNSS) availability, multiple dual-functional Unmanned Aerial Vehicles (UAVs) can provide uplink communication and cooperative localization services. Ground terminals have different requirements for data volume, minimum communication rate, localization-accuracy threshold, service priority, and communication and localization service weights. Maximizing communication rate, localization accuracy, or a weighted sum of the two cannot directly reflect these differentiated requirements. A service below its activation threshold may be unusable, whereas resources allocated after demand saturation provide little additional value. A quasi-static service period with prior terminal positions is therefore considered. A topology-dependent Value of Service (VoS) is formulated, and Physical Resource Block (PRB) scheduling and UAV positions are jointly optimized.  Methods  Data or Sounding Reference Signals (SRSs) are transmitted by terminals over assigned PRBs. Time Difference of Arrival (TDoA) measurements are formed by synchronized UAVs, and a probabilistic air-to-ground channel model determines communication access and valid localization anchors. Communication VoS combines a sigmoidal rate utility with data completeness. Position Error Bound (PEB), derived from the accumulated Fisher Information Matrix (FIM), is mapped to an exponential utility to quantify localization VoS. Terminal priorities and communication and localization service weights are used to aggregate the two service values. The resulting problem P0 couples binary scheduling, non-concave utilities, time-frequency resources, service relationships, localization geometry, and UAV positions. For a fixed topology, the average number of allocated PRBs and active slots are used to construct a continuous service-level resource profile. This profile approximates the original schedule but does not provide an upper bound. Bandwidth responses are obtained by deterministic one-dimensional branch-and-bound search with damped Newton refinement and shadow-price bisection. Slot responses are obtained by deterministic comparison of a finite set of service-boundary, integer-slot, and feasible-domain-boundary candidates. Adjacent-integer recovery and deterministic slot packing are then used to construct a feasible integer schedule under resource-capacity and terminal-power constraints. If packing fails, the upward-rounded component with the smallest unit VoS loss is rolled back, and the schedule is repacked. VoS is then recomputed from the recovered integer schedule. Thus, the reported schedules remain feasible for P0 without any claim of global optimality. The fixed-topology resource-allocation procedure produces a complete resource response, including the recovered feasible schedule and its VoS, which is subsequently used to drive deployment. Coverage-repair, localization-geometry-improvement, and resource-saving candidates are generated according to service deficits, service relationships, and resource consumption. Normalized proxy scores, per-UAV candidate truncation, and beam search reduce the number of complete resource-response evaluations. For each retained deployment, service relationships are rebuilt, including communication access and localization anchors, and the FIM and resource allocation are recomputed. A candidate is accepted only when the VoS gain obtained from its complete resource response exceeds the preset threshold. The resulting outer sequence is monotonic and bounded and converges to a stable point within the generated candidate set. Resource allocation and node deployment remain coupled throughout the procedure. For each retained topology, communication access, localization anchors, the FIM, and the feasible integer schedule are recomputed before VoS is evaluated. Proxy scores are used only to rank candidates, whereas final acceptance is always based on the complete resource response. The next topology is therefore not selected from distance or localization geometry alone. The accepted update reflects terminal demand, resource scarcity, and localization geometry under the same feasible scheduling constraints used in the objective. This design keeps the optimization objective consistent with the final deployment decision and avoids a geometry-only selection rule.  Results and Discussions  Simulations are conducted in a 700 m × 700 m area with four UAVs at a baseline deployment height of 150 m. Communication-dominant, localization-dominant, and balanced terminals are included in equal proportions. At each load, all methods share 30 independent scenarios and identical random seeds. The deployment-height experiment reports a 95% confidence interval based on 1,000 scenario-level bootstrap resamples. For fair comparisons, all resource-allocation methods use the same topology, resource pool, and random scenarios. All deployment methods use the same proposed joint resource-allocation response, are evaluated from a cold start, and are subject to the same cap on complete resource-response evaluations, with initialization and training costs included. VoS-driven allocation yields smaller communication-rate and localization-accuracy demand deviations than the communication-priority and average-utility metrics because service saturation redirects resources from overprovisioned requests to insufficiently served requests (Fig. 2). In the convergence test, three feasible initializations converge to similar system VoS values. Damping suppresses oscillations near transitions between non-concave segments, and the small-scale benchmark indicates limited empirical loss from the continuous response and integer recovery (Fig. 3). With increasing terminal load, the proposed joint allocator outperforms a Particle Swarm Optimization (PSO)-based slot-response variant, as well as non-joint, learning-based, weight-driven, and random allocation methods (Fig. 4). Communication service is more sensitive to terminal load because both communication rate and data completeness must be maintained, whereas localization service accumulates information across anchors and slots. Resource-pool tests further show that additional bandwidth provides little benefit when slots are scarce, whereas increasing SRS bandwidth more directly improves localization. The structured deployment search achieves higher VoS with shorter end-to-end deployment time than PSO, Differential Evolution (DE), and Bayesian Optimization (BO) by focusing complete resource-response evaluations on service-aware candidates (Fig. 5). VoS first increases and then decreases with UAV deployment height because improved line-of-sight probability and multi-anchor visibility compete with increased propagation distance and path loss. Horizontal deployment updates balance communication-link quality and localization geometry according to terminal requirements (Fig. 6).  Conclusions  The proposed framework evaluates communication and localization services according to demand satisfaction and coordinates time-frequency resource allocation with iterative UAV deployment updates through complete resource responses and structured candidate search. It improves demand matching, system VoS, deployment efficiency, and interpretability while preserving feasibility under the original scheduling constraints. Prior terminal positions and a quasi-static service period are assumed. Future work will address dynamic terminal movement and changing demands over longer service periods through dynamic resource allocation and continuous UAV trajectory optimization.
Impact of Wireless Priors on the Computation and Energy Cost of MU-MIMO Precoding Learning
CONG Pengyu, HAN Shengqian, DENG Mingyu, LIU Shengjie, YANG Chenyang, SHEN Songhui
Available online  , doi: 10.11999/JEIT260388
Abstract:
  Objective  This paper investigates downlink Multi-User Multi-Input Multi-Output (MU-MIMO) precoding policy learning from the perspectives of computational complexity and energy consumption. Traditional numerical optimization algorithms achieve high performance but exhibit rapidly increasing computational complexity as the numbers of base station antennas and served users increase, leading to high inference latency and energy consumption. In recent years, deep learning has been widely adopted to reduce online computational cost; however, existing evaluations generally rely on training or inference time and FLoating-point OPerations (FLOPs), without direct measurements of energy consumption and power. More importantly, the computational cost of a deep learning model is closely related to network architecture design, which should effectively exploit the wireless prior knowledge of the precoding policy. Therefore, this paper develops a network architecture that matches the multidimensional permutation properties of the precoding policy and systematically investigates how wireless priors affect computational complexity and energy consumption through comprehensive hardware-based measurements and simulations.  Methods  The MU-MIMO precoding policy is formulated as a mapping from multiuser channel information to the optimal precoding matrix under a transmit power constraint. The optimal policy satisfies multidimensional joint permutation equivariance and invariance with respect to user indices, receive antenna indices, and base station antenna indices. To exploit these wireless priors, an Attention-based Graph Neural Network (AGNN) is proposed based on a hypergraph structure, in which the update and aggregation operations satisfy the required equivariance and invariance properties. An attention mechanism is incorporated to model inter-user interference and improve generalization across different numbers of users. For broadband precoding, multi-subcarrier channel information is aggregated at the input layer to construct an expanded feature representation. To quantify computational energy cost, a hardware measurement platform is developed to collect energy consumption and power for the GPU, CPU, and DRAM during both training and inference. Simulations are conducted using 3GPP TR 38.901 Urban Macrocell (UMa) channel datasets with different antenna array sizes and bandwidth configurations. The proposed AGNN is compared with a traditional numerical optimization algorithm based on Zero-Forcing Block Diagonalization (ZFBD) with greedy user pairing and two Transformer-based architectures that satisfy only one-dimensional permutation equivariance.  Results and Discussions  Two major findings are obtained. First, incomplete exploitation of wireless priors results in inferior performance and higher computational cost. In the MU-MISO scenario, the Transformer-based architectures achieve lower spectral efficiency than the ZFBD+Greedy baseline while requiring substantially larger models and higher inference FLOPs than AGNN. By matching the multidimensional permutation properties of the precoding policy, AGNN achieves higher spectral efficiency while reducing inference FLOPs by approximately one order of magnitude. Hardware measurements further demonstrate that AGNN reduces inference energy consumption and power on both the CPU and GPU. Second, in small- and large-scale MU-MIMO scenarios, ZFBD+Greedy increases the system sum rate by 10.9×, whereas inference FLOPs, inference time, and inference energy increase by 79.0×, 21.3×, and 38.6×, respectively. In contrast, AGNN increases the system sum rate by 11.5×, while inference FLOPs increase by only 1.89×. Meanwhile, inference time and inference energy are reduced to 0.03× and 0.20×, respectively. These results demonstrate that exploiting the multidimensional permutation properties of the precoding policy provides an effective approach for reducing computational complexity, inference latency, and energy consumption in large-scale 6G MU-MIMO systems.  Conclusions  This paper investigates the effect of wireless priors on the computational complexity and energy consumption of MU-MIMO precoding learning. By exploiting the multidimensional joint permutation equivariance and invariance of the optimal precoding policy, an AGNN is developed that is well matched to these properties. A hardware-aware measurement platform is established to obtain direct measurements of energy consumption and power for the GPU, CPU, and DRAM during training and inference. Simulations based on 3GPP TR 38.901 UMa channel datasets demonstrate that Transformer-based architectures satisfying only one-dimensional permutation equivariance achieve lower spectral efficiency while incurring substantially higher computational and energy costs. In contrast, AGNN achieves higher spectral efficiency while substantially reducing inference FLOPs, inference time, inference energy consumption, and training complexity. As system size increases, traditional numerical optimization algorithms exhibit much faster growth in computational and energy costs than in system sum rate, whereas the proposed learning method based on wireless priors maintains low inference latency and energy consumption. Overall, exploiting the wireless prior knowledge of the MU-MIMO precoding policy in network architecture design provides an effective solution for computationally efficient and energy-efficient high-dimensional precoding optimization in future 6G networks.
Determination of Key Geometric Parameters for Spaceborne Dual-Beam Along-Track Interferometric SAR under Asymmetric Geometry
SHEN Qingyuan, ZHANG Yangyang, LAI Tao, ZHU Yuting, XIE Zhifeng, WANG Xiaoqing
Available online  , doi: 10.11999/JEIT260870
Abstract:
  Objective  Spaceborne Dual-Beam Along-Track Interferometric Synthetic Aperture Radar (DBATI-SAR) acquires fore- and aft-looking radial velocities to support two-dimensional ocean surface current retrieval. The stability of the inversion depends on the relative directions of the two Radar Line-Of-Sight (RLOS) projections on the target local tangent plane. Conventional flat-Earth geometry models generally assume symmetric squint angles or time offsets. However, orbital curvature, Earth curvature, Earth rotation, and local projection nonlinearity produce asymmetric spaceborne geometries. Therefore, symmetric squint angles do not necessarily guarantee orthogonal ground-projected RLOS directions. A method is developed to determine the fore- and aft-looking observation times, squint angles, and down-looking angles directly under the ground-projected RLOS orthogonality constraint.  Methods  A complete Earth-Centered Earth-Fixed (ECEF) geometry model is established from the satellite state and target position. The RLOS is projected onto the target local tangent plane, and a signed ground-projected RLOS angle is defined relative to the zero-squint reference direction. The squint and down-looking angles are calculated from the observation times rather than prescribed independently. A two-dimensional observation geometry matrix is constructed to relate the horizontal current components to the two radial velocities. Under an ideal symmetric geometry used for the analytical derivation, the singular values and condition number show that a one-sided ground-projected RLOS angle of \begin{document}$ {45}^{{^{\circ}}} $\end{document} provides the best inversion conditioning. A normalized geometric amplification factor is introduced to quantify the additional error amplification caused by nonorthogonal projections. Near the zero-squint reference point, a closed-form analytical leading term is derived to relate the observation-time offset to the ground-projected RLOS angle. This expression reveals the effects of the reference slant range, down-looking angle, equivalent along-track velocity, and second-order range-history curvature. A local polynomial inverse mapping is then constructed from complete ECEF forward-geometry samples. Target ground-projected RLOS angles of \begin{document}$ {-45}^{{^{\circ}}} $\end{document} and \begin{document}$ +{45}^{{^{\circ}}} $\end{document} are substituted into the inverse mapping to obtain the fore- and aft-looking observation times, after which the corresponding squint and down-looking angles are calculated.  Results and Discussions  Numerical experiments are conducted using 500 random orbit-parameter sets. Fifth- and sixth-order polynomials produce relatively large inverse-mapping errors, whereas a seventh-order model substantially improves the accuracy. With 20 samples, the mean ground-projected RLOS angle error is 0.0392°, and further increases in polynomial order or sample number provide limited improvement (Fig. 2, Table 2). The analytical leading term agrees well with the complete ECEF geometry model. For a representative case, the angular root-mean-square error and maximum deviation are approximately \begin{document}$ {0.11}^{\circ } $\end{document} and \begin{document}$ {0.31}^{\circ } $\end{document}, respectively. Across the 500 cases, the root-mean-square errors of the fore- and aft-looking observation times predicted by the analytical leading term are 0.95 s and 0.57 s, respectively. These results indicate that the analytical leading term captures the dominant observation-time scale, while the numerical inverse mapping accounts for asymmetric higher-order effects (Fig. 3). The conventional flat-Earth geometry model produces a squint-angle correction of up to approximately \begin{document}$ {3}^{\circ } $\end{document} and a ground-projected RLOS orthogonality error of approximately \begin{document}$ {6.5}^{\circ } $\end{document}. The proposed method reduces the mean ground-projected RLOS angle error to approximately \begin{document}$ {0.04}^{\circ } $\end{document} and decreases the normalized geometric amplification factor from approximately 1.12 to 1.000 7 (Fig. 4). The semimajor axis has the strongest effect on the aft-looking observation time, which varies from approximately 34 s to 76 s, while variations caused by other orbital parameters remain within 7 s (Fig. 5). Under combined squint- and down-looking-angle perturbations of up to 0.2°, most samples retain ground-projected RLOS angle errors below \begin{document}$ {1}^{\circ } $\end{document} (Fig. 6).  Conclusions  A method is proposed to determine key geometric parameters for spaceborne DBATI-SAR under asymmetric geometry. The analytical leading term explains the dominant observation-time scale, whereas the complete ECEF numerical inverse mapping accurately determines the fore- and aft-looking observation times and corresponding squint and down-looking angles. The proposed method provides substantially higher geometric accuracy than the conventional flat-Earth geometry approach. Practical mission design should also consider pulse repetition frequency, azimuth ambiguity, Doppler bandwidth, beam-steering range, and along-track interferometric coherence.
Adaptive Fusion Detection for Dual-radar with Non-identical Clutter Distributions
ZHOU Baoyi, YANG Yong, YANG boyu
Available online  , doi: 10.11999/JEIT260616
Abstract:
  Objective  Small sea-surface targets have low Radar Cross Sections (RCSs), and their echoes are easily masked by intense sea clutter, resulting in extremely low Signal-to-Clutter Ratios (SCRs). Single-radar systems are constrained by a single operating frequency band and fixed observation angles, which limits their detection performance. Multi-radar collaborative detection provides an effective approach to overcoming this limitation. Existing multi-radar fusion detection methods are developed for different information fusion levels. However, feature-level and signal-level fusion methods generally assume identical clutter distributions, whereas decision-level fusion is less dependent on the specific clutter distribution. In practical multi-radar detection, differences in radar parameters, including frequency band, range resolution, and grazing angle, result in non-identical statistical characteristics of sea clutter. Furthermore, target RCS varies with observation azimuth and operating frequency, resulting in different SCRs for the same target observed by different radars and causing model mismatch in conventional fusion detectors. To address the coexistence of non-identical clutter distributions and different SCRs, a Neyman-Pearson (NP) criterion-based dual-radar Adaptive Fusion Detection method, termed NP-AFD, is proposed for Rayleigh and Weibull sea clutter.  Methods  Amplitude distribution fitting is performed on measured S-band and X-band sea clutter data using five commonly used models: Rayleigh, lognormal, Weibull, Gamma, and K distributions. Fitting accuracy is evaluated using the Mean Square Error (MSE) of the Probability Density Function (PDF) and Complementary Cumulative Distribution Function (CCDF), as listed in Table 1. Based on the fitting results, local optimal test statistics are derived for Rayleigh and Weibull sea clutter. Because the two radar observations are independent, the joint likelihood ratio is given by the product of their individual likelihood ratios. The fusion test statistic is formulated as a weighted sum of the two local test statistics. The fusion weights are adaptively determined from the estimated SCRs, with larger weights assigned to radars with higher estimated SCRs. Closed-form analytical expressions relating the false alarm probability and detection probability to the decision threshold are also derived.  Results and Discussions  Monte Carlo simulations with 105 independent trials verify the derived closed-form expressions. Under identical SCRs for the two radars, the simulated detection curves of the individual radars and NP-AFD agree well with the theoretical curves (Fig. 3). Compared with decision-level OR and AND fusion, NP-AFD consistently achieves the highest detection probability, followed by OR fusion, whereas AND fusion performs worst. Performance gain analysis shows that NP-AFD maintains a positive performance gain over the better-performing single radar across all tested Weibull shape parameters. In contrast, OR fusion exhibits negative performance gain at low SCRs when the Weibull shape parameter is small, whereas AND fusion maintains a negative performance gain across the full SCR range (Fig. 4). The simulated false alarm probability is well controlled around the preset value of 10–3 (Fig. 5). Furthermore, a two-dimensional joint evaluation is performed by independently varying the SCRs of the two radars, and detection probability contours are plotted with the SCRs of the two radars as the coordinate axes (Fig. 6). The results show that NP-AFD requires lower SCR combinations than OR and AND fusion to achieve the same detection probability. Experiments with measured sea clutter data further demonstrate that NP-AFD achieves the highest detection probability among all compared methods (Figs. 7 and 8). Its detection performance agrees well with the theoretical results (Figs. 9 and 10).  Conclusions  The challenges posed by non-identical clutter distributions and different SCRs in multi-radar collaborative detection are addressed. Based on the Neyman-Pearson criterion, a dual-radar adaptive fusion detection method, NP-AFD, is derived for Rayleigh and Weibull sea clutter. Closed-form expressions are obtained for the fusion weights, decision threshold, and detection probability. Theoretical derivations, simulations, and experiments with measured sea clutter data consistently demonstrate that NP-AFD outperforms OR fusion and AND fusion under arbitrary SCR combinations, providing superior detection performance and robustness to different SCR combinations. The method, however, remains sensitive to non-uniform sea clutter, as reflected by degraded false alarm control in the presence of sea spikes.
Resource Allocation for Multi-UAV Relay Networks in 6G Semantic Communication
XIAO Liming, GUAN Zheng, LIU Jie, YU Jihong, CHEN Liyuan
Available online  , doi: 10.11999/JEIT260520
Abstract:
  Objective  Sixth-Generation (6G) mobile networks aim to achieve global seamless coverage through space-air-ground integrated architectures. In this context, Unmanned Aerial Vehicles (UAVs) serve as mobile aerial relay nodes to support massive ground-user access in complex environments. However, traditional data-oriented communication paradigms incur substantial bandwidth overhead, limiting their applicability in spectrum-constrained UAV networks. In addition, conventional centralized resource allocation methods are difficult to implement in real time because of highly dynamic network topologies and the strong coupling among multidimensional resources. To address these challenges, semantic communication has emerged as a communication paradigm that extracts semantic information at the transmitter and reconstructs it at the receiver, thereby reducing redundant data transmission. Existing semantic-driven resource allocation methods, however, primarily focus on static terrestrial networks or single-UAV scenarios and do not adequately address coverage limitations and co-channel interference in multi-UAV relay networks. Therefore, a joint resource optimization model and a distributed resource allocation framework are proposed to improve semantic transmission efficiency and long-term user fairness through the joint optimization of multidimensional resources, thereby supporting intelligent resource scheduling in future 6G integrated networks.  Methods  A joint resource optimization model is formulated for multi-UAV relay networks under the semantic communication paradigm (Fig. 1), jointly optimizing the number of semantic symbols, UAV trajectories, power control, and channel allocation. To evaluate semantic communication performance, a Semantic Communication Quality of Service (SC-QoS) metric is proposed by combining Semantic Quantization Efficiency (SQE) with normalized transmission delay. The optimization objective is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem that maximizes the weighted sum of system-wide SC-QoS and long-term user fairness measured by Jain’s fairness index. To solve this problem, a Two-Stage Hybrid Reinforcement Learning (TS-HRL) framework is proposed (Fig. 2). In the first stage, a Capacity-Aware K-means (CA-K-means) algorithm performs heuristic UAV pre-deployment. By introducing a dynamic distance compensation term based on residual capacity, edge users are guided toward lightly loaded UAVs, thereby achieving load balancing while preserving spatial proximity. In the second stage, the dynamic scheduling problem is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and solved using a Recurrent Independent Proximal Policy Optimization with Parameter Sharing (R-IPPO-PS) algorithm (Algorithm 1). The algorithm employs Long Short-Term Memory (LSTM) networks to aggregate historical observations and action trajectories, enabling the inference of hidden environmental states. Furthermore, the parameter-sharing mechanism improves training efficiency as the number of agents increases, while the association preference decoupling strategy transforms discrete channel allocation decisions into continuous association preference variables to facilitate policy optimization.  Results and Discussions  The proposed TS-HRL framework is evaluated in a dynamic environment with randomly moving ground users. Convergence analysis shows that the proposed method achieves a higher initial reward and converges with fewer iterations than the random deployment, memoryless resource allocation, and Enhanced Independent Soft Actor-Critic (EI-SAC) baselines (Fig. 3). By combining LSTM-based temporal modeling with CA-K-means pre-deployment, the proposed framework reduces ineffective exploration and improves convergence stability. Compared with conventional bit-based communication, the semantic communication framework increases Semantic Spectral Efficiency (S-SE) by 3.6-fold, reduces transmission delay by 92.1%, and improves fairness by 14.3% (Fig. 4). Within the semantic communication framework, compared with the memoryless and heuristic schemes, the proposed method improves S-SE by 92.7% and 51.9%, reduces transmission delay by 13.5% and 57.7%, and improves fairness by 17.3% and 39.7%, respectively. Although the EI-SAC baseline achieves a fairness index of 0.93, its S-SE remains relatively low. In contrast, the proposed TS-HRL framework maintains a high level of fairness while achieving an S-SE approximately 4.3 times that of EI-SAC. Compared with random deployment, the proposed method improves S-SE by 11.3% with only a 1.1% decrease in fairness, demonstrating a better balance between transmission efficiency and fairness. As the number of users increases from 10 to 40, most baseline methods exhibit decreases in S-SE and fairness because of intensified co-channel interference and spectrum limitations (Fig. 5). In contrast, the adaptive scheduling strategy and global fairness reward mechanism mitigate performance degradation and maintain the average transmission delay below 0.1 ms. These results demonstrate that the proposed method improves overall system performance while ensuring long-term user fairness.  Conclusions  This paper investigates joint resource allocation and trajectory optimization for dynamic 6G multi-UAV relay networks under the semantic communication paradigm. A joint optimization model that couples the number of semantic symbols, UAV trajectories, power control, and channel allocation is formulated to maximize SC-QoS and long-term user fairness. To address the high-dimensional coupling of the optimization problem, a TS-HRL framework integrating CA-K-means pre-deployment with the R-IPPO-PS algorithm is proposed for multidimensional resource scheduling under partial observability. Simulation results demonstrate the convergence, stability, and scalability of the proposed method under different user densities. Through distributed multi-agent cooperation among UAVs, the proposed method improves semantic transmission efficiency and reduces transmission delay while maintaining long-term service fairness for ground users. These findings provide an effective approach to intelligent resource allocation for UAV-assisted semantic communication in future space-air-ground integrated networks.
Space-Time-Coding Metasurface-Enabled Integrated Design of Radar Communication and Electromagnetic Stealth
ZHANG Ming, WANG Zhe, WANG Boya, YANG Lin, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260536
Abstract:
  Objective  To address the complexity and limited modulation capabilities of existing reconfigurable metasurfaces that integrate radiation and stealth functions, a space-time-coding metasurface is proposed for dynamic switching between beam scanning and Radar Cross Section (RCS) reduction. By periodically modulating the states of the meta-atoms in the time domain, differentiated phase distributions are generated for the incident wave and its harmonics. This enables integrated radiation and electromagnetic stealth without complex Transmit-Receive (T/R) components or multilayer structures. The proposed design provides a simple and low-cost approach for integrated radar communication and electromagnetic stealth.  Methods  The metasurface adopts a metal-dielectric-metal structure, with each meta-atom integrating a PIN diode for 1-bit reflection-phase modulation. Electromagnetic simulations are performed using CST Microwave Studio. At the center frequency of 10.0 GHz, the two diode states provide a reflection phase difference close to 180°, with a near-lossless co-polarized reflection amplitude. Binary Particle Swarm Optimization (BPSO) is used to optimize the space-time-coding sequences for beam scanning and RCS reduction. An 8 × 8 prototype is fabricated using Printed Circuit Board (PCB) technology. A Vector Network Analyzer (VNA) equipped with an S97082A option is used to measure the radiation and scattering characteristics of the prototype.  Results and Discussions  Simulations and measurements confirm that the proposed space-time-coding metasurface provides beam scanning and RCS reduction. In the radiation mode, the fundamental, +1st, +2nd, +3rd, –1st, –2nd, and –3rd harmonics are steered to +2°, +14°, +32°, +46°, –15°, –28°, and –44°, respectively. The average error between the measured and target beam directions is 1.3°. The average sidelobe levels of the ±1st, ±2nd, and fundamental harmonics are 10.59 dB lower than the corresponding main lobes. In the scattering mode, the peak echo gain at 10.0 GHz is reduced by 11.56 dB relative to a copper plate. Over 9.9~10.1 GHz, the echo gain is reduced by approximately 10 dB, except at 9.93 GHz and 10.04 GHz, where the reductions are 8.07 dB and 7.89 dB, respectively. The maximum reduction reaches 14.83 dB at 9.95 GHz. These results verify the integrated radiation and scattering-control capabilities of the proposed metasurface.  Conclusions  A space-time-coding metasurface is proposed for integrated radiation and electromagnetic stealth. By integrating PIN diodes into the meta-atoms, the reflection states are periodically modulated on a planar metasurface to generate differentiated equivalent phase distributions at different harmonic frequencies. This enables multi-angle beam scanning without T/R components and reduces the echo gain through spatial and temporal coding. The proposed design simplifies the system architecture and reduces hardware requirements. The demonstrated beam scanning and RCS reduction indicate its potential for integrated radar communication and electromagnetic stealth.
Communication Performance and Fault Degradation Analysis of Boundary-interface Configurations in 2.5D Chiplet Systems
HOU Shuaikang, LIU Qinrang, LV Ping, LIU Zhengyu, XU Yuhang, LI Peijie, GUO Wei
Available online  , doi: 10.11999/JEIT260633
Abstract:
  Objective  Chiplet-based integration provides an important approach for constructing large-scale heterogeneous systems. In a 2.5D chiplet system, inter-chiplet packets travel from a source node to a boundary interface in the source chiplet, traverse an interposer network, and enter the destination chiplet through a destination-side interface. Because packaging resources, micro-bump count, routing resources, and power and area budgets are limited, interfaces can be deployed only at a subset of boundary nodes. Their number and spatial distribution affect access distance, service-region formation, interface-load distribution, and traffic remapping after failures. Existing studies have addressed die-to-die standards, interposer architectures, placement, routing, simulation, and fault tolerance. However, the independent structural effects of boundary-interface configuration remain insufficiently characterized. This study investigates how interface count, location, service-region partitioning, and failures affect communication performance and fault-induced degradation in 2.5D chiplet–interposer networks.  Methods  A graph-level structural model is developed for a 2.5D chiplet–interposer network. Each chiplet is modeled as an n × n mesh, and selected boundary nodes serve as active interfaces. Each node is mapped to its nearest active interface, and nodes mapped to the same interface form a service region. Average interface-access distance, abstract end-to-end hop count, squared coefficient of variation of interface load, maximum load ratio, and fault-induced degradation are evaluated. Three chiplet scales, n = 4, 8, and 16, are considered, with corresponding interface budgets of k = 2, 4, 6; k = 4, 8, 12; and k = 8, 16, 24, respectively. Five layouts are analyzed: Uniform-corner, Balanced-edge, Two-edge-centered, Single-edge-clustered, and Greedy-DL. Uniform, Hotspot, Transpose, Tornado, and Neighbor traffic patterns are considered. For fault analysis, each active interface is removed in turn, and affected nodes are remapped to their nearest healthy interfaces. A cycle-accurate gem5 model with Ruby/GARNET, comprising four 4 × 4 mesh chiplets and a 4 × 4 interposer mesh, is constructed. Uniform and Hotspot traffic are evaluated at injection rates from 0.01 to 0.15 flit/node/cycle. The latency saturation injection rate is defined as the first sampled rate at which the average packet latency exceeds 50 cycle. This two-level evaluation separates structural effects from cycle-level network behavior and enables interface count, placement, traffic pattern, and fault location to be compared under consistent topology and routing assumptions.  Results and Discussions  Graph-level results show that increasing the number of boundary interfaces reduces average interface-access distance and abstract end-to-end hop count, whereas the marginal benefit decreases as boundary coverage becomes sufficient (Fig. 3). Under a fixed interface budget, different interface locations induce different service-region partitions and interface-load distributions (Table 2, Fig. 4). For an 8 × 8 chiplet with k = 8 under Uniform traffic, Greedy-DL achieves the smallest average interface-access distance of 3.125, whereas Balanced-edge provides a favorable compromise, with an access distance of 3.500 and a maximum load ratio of 1.250. Single-edge-clustered produces the largest access distance and abstract end-to-end hop count because its interfaces are concentrated along one boundary. Under Hotspot traffic, Two-edge-centered reduces the squared coefficient of variation of interface load to 0.016, but its access distance remains higher than those of Balanced-edge and Greedy-DL, indicating a trade-off between distance and load balance (Table 2, Fig. 5). Interface failures cause only moderate increases in average access distance but substantial load concentration at the remaining healthy interfaces. Under Hotspot traffic, the worst-case maximum-load-ratio degradation is 23.94% for Balanced-edge and 93.75% for both Two-edge-centered and Single-edge-clustered (Table 3), indicating that fault sensitivity is mainly associated with service-region migration and load reconcentration. The gem5 results further support the graph-level observations. When the interface count increases from k = 1 to k = 4, low-load latency decreases from 24.50/24.44 cycle to 18.05/18.09 cycle under Uniform and Hotspot traffic, respectively, whereas the latency saturation injection rate increases from 0.06 to 0.14 (Fig. 6). Across the k = 4 layouts, Balanced-edge achieves low latency and hop count, whereas Single-edge-clustered has the highest baseline communication cost. Under Hotspot traffic, Balanced-edge reduces low-load latency by about 14.3% and average hop count by about 18.7% compared with Single-edge-clustered. Greedy-DL provides short communication paths under low load, but its latency saturation injection rate is 0.13, lower than the 0.14 achieved by Balanced-edge and several other layouts. This result indicates that minimizing access distance alone does not guarantee better medium- and high-load performance (Fig. 7). For fault validation, all 16 single-interface failure scenarios are evaluated for each of the three layouts, yielding 48 scenarios in total. The results show increased latency and hop count, together with an earlier onset of congestion. The worst-case latency saturation injection rate decreases to 0.10~0.11, and Balanced-edge exhibits smaller average and worst-case degradation than the more concentrated layouts (Fig. 8, Table 5).  Conclusions  Boundary-interface configuration is a key structural parameter in 2.5D chiplet interconnect design. It affects interface-access distance, service-region formation, interface-load distribution, and fault-induced traffic remapping. Increasing the interface count improves communication efficiency, but the benefit diminishes as boundary coverage increases. Under the same interface budget, interface placement determines whether traffic remains balanced or becomes concentrated after mapping and remapping. The graph-level metrics support low-cost structural screening and mechanism analysis, whereas gem5/Ruby GARNET simulations provide cycle-accurate validation. Therefore, interface count, location, load balance, and fault degradation should be evaluated jointly in 2.5D chiplet–interposer network design.
Efficient Hyperdimensional Computing Accelerator Design for Chinese Text Classification
YU Tianyang, WU Bi, LIU Weiqiang
Available online  , doi: 10.11999/JEIT260556
Abstract:
  Objective  With the proliferation of edge computing in the Internet of Things (IoT), smart wearables, and offline terminals, low-latency and privacy-preserving Chinese text analysis on local devices has become a core requirement. Although neural network-based and Transformer-based language models achieve high accuracy, their large parameter sizes and computational demands make them difficult to deploy on power- and storage-constrained edge devices. HyperDimensional Computing (HDC), an emerging brain-inspired computing paradigm, represents text using 2 k~10 k-dimensional hypervectors and employs lightweight encoding and querying mechanisms instead of complex multilayer networks, providing a hardware-friendly approach for edge-side text processing. However, existing HDC studies have mainly focused on alphabetic writing systems such as English, in which a limited set of letters is sufficient for N-gram encoding. For Chinese, which contains thousands of commonly used characters, directly applying these methods would require a large number of base hypervectors, resulting in substantial storage overhead and weakening the lightweight advantage of HDC. This paper aims to address this limitation by proposing an efficient character encoding method and a dedicated HDC framework with hardware acceleration for Chinese text classification.  Methods  Based on the glyph structure of Chinese characters, a hyperdimensional encoding method based on character glyph structure is proposed. Specifically, the encoding process of the Wubi input method is reverse-engineered to decompose each Chinese character into an equivalent Wubi letter sequence for efficient hyperdimensional encoding. A retrieval library covering 3 500 commonly used Chinese characters, as defined in the General Standard Chinese Characters Table issued by the Ministry of Education of the People’s Republic of China, is constructed. For each character, the Wubi library is queried to obtain a letter sequence of length 3 or 4, with each letter mapped to a base hypervector. Cyclic shift and binding operations are then applied to the base hypervectors of consecutive letters. This produces a character-level hypervector that captures both letter identity and positional order information, thereby avoiding confusion caused by different letter permutations. Compared with directly assigning an independent hypervector to each Chinese character, the proposed method reduces the storage requirement for base hypervectors by more than 95% (Table 1). Building on this character encoding method, the HDChinese framework is developed to support inference, training, and retraining. During inference, the sentence-level hypervector obtained by bundling all character hypervectors queries the class hypervectors using cosine similarity. Because the hypervectors are binary, the inner-product operation is efficiently implemented using XNOR logic. During training, class hypervectors are generated by bundling sentence hypervectors belonging to the same class and applying sign-based binarization. During retraining, the nonbinary class hypervectors are iteratively updated using misclassified samples with a learning rate of 0.25, thereby reducing the effect of outlier hypervectors on the cluster center. Furthermore, a dedicated hardware accelerator architecture is designed for HDChinese. The architecture comprises three main modules: an encoding module, a querying module, and a class hypervector update module. These modules can be selectively activated to support inference, training, and retraining modes, respectively (Fig. 2). The accelerator is prototyped on an Ultra96v2 FPGA development board equipped with a Xilinx ZYNQ System on Chip (SoC).  Results and Discussions  Four open-source Chinese text classification datasets covering binary and multiclass tasks are used for evaluation (Table 2). The Wubi-based encoding method achieves higher accuracy than the Pinyin-based method on all four datasets and reduces the storage requirement by 31.46% (Table 4). Ignoring uncommon Chinese characters, which occur at an average frequency of less than 0.5%, has a negligible effect on accuracy (Table 4). A hypervector dimension of 2 k is selected because increasing the dimension beyond 2 k provides no significant accuracy improvement (Fig. 3). FPGA measurements show that HDChinese achieves a model size of approximately 16 kB, representing a reduction of more than 99% compared with k-Nearest Neighbor (kNN), Support Vector Machine (SVM), and random forest models (Table 7). Compared with SVM and random forest, the proposed accelerator reduces inference time by 6.67%~47.25% and total training time by 4.91%~77.36%. After retraining, the classification accuracy reaches levels comparable to those of the comparison methods (Fig. 4, Table 7). The kB-scale model also enables operation without external memory chips, whereas the comparison methods require more than 400 MB of runtime memory. At 100 MHz, the FPGA implementation consumes 0.278 W, providing the most favorable trade-off between throughput and power consumption (Table 6).  Conclusions  A hyperdimensional encoding method based on Chinese character glyph structure is proposed to address the incompatibility between existing HDC methods and Chinese text. The HDChinese framework and its dedicated hardware accelerator reduce model complexity by more than 99% while maintaining competitive classification accuracy. Training time is reduced by 4.91%~77.36%, and inference time is reduced by 6.67%~47.25%. The proposed approach achieves a favorable balance between accuracy and computational and storage overhead, providing a hardware solution for Chinese text analysis on edge devices.
SkipSync: Accelerating Instruction Sampling for LLM Workloads
CAI Luoshan, ZHOU Yaoyang, WANG Kaifan, LIU Tianyi, SUN Ninghui, BAO Yungang
Available online  , doi: 10.11999/JEIT260397
Abstract:
  Objective   The rapid evolution of Large Language Models (LLMs) has increased the demand for efficient design and evaluation of domain-specific accelerators. Sampling-based performance evaluation methods reduce the cost of cycle-accurate simulation but rely on Instruction Set Simulators (ISS) to execute complete workloads for profiling and checkpoint generation. In LLM inference scenarios, frequent changes in model architectures, inference frameworks, and accelerator instruction sets prevent checkpoint reuse and substantially increase the overhead of ISS-based functional simulation. This overhead has become a major bottleneck in the evaluation workflow. Existing ISS acceleration methods, such as hardware virtualization and Dynamic Binary Translation (DBT), either require target and host systems to share the same Instruction Set Architecture (ISA) or have high implementation complexity, making them difficult to apply to rapidly evolving LLM workloads. Therefore, a flexible and efficient ISS acceleration method is needed to support fast and accurate sampling-based performance evaluation of LLM accelerators.  Methods   An ISS acceleration method, SkipSync, is proposed for instruction sampling of LLM inference workloads. The key observation is that the control flow of most LLM operators is independent of runtime tensor values and is determined by static parameters. Based on this observation, the Skip mechanism is introduced to bypass the execution stage of accelerator instructions within selected operators while preserving instruction fetch and decode to maintain sampling accuracy. To ensure correct subsequent execution, a lightweight Host-Simulator Synchronization mechanism is further introduced to synchronize host-computed results back to the ISS. Lightweight synchronization primitives and custom instructions are designed to integrate SkipSync into existing sampling-based performance evaluation workflows with minimal implementation effort.  Results and Discussions   SkipSync is implemented on QEMU and supports RISC-V vector and matrix extensions as accelerator instruction sets. Experimental results show that SkipSync substantially reduces ISS functional simulation overhead while preserving sampling accuracy. Compared with the baseline, average speedups of 4.64× and 5.51× are achieved in the Prefill and Decode stages, respectively (Fig. 5), by eliminating the dominant execution cost of vector and matrix instructions, which accounts for 78.67% of runtime in the Prefill stage and 81.95% in the Decode stage (Table 1). SkipSync outperforms SIMD_DBT (Fig. 6) and provides better support for newly introduced accelerator instructions. The synchronization mechanism introduces limited overhead, with an average runtime increase of only 8.5% relative to ideal Skip-only execution (Fig. 8). Thus, synchronization does not negate the performance gains from the Skip mechanism. SkipSync also maintains high sampling accuracy, with an average sampling error of 2.55%, and the sampling results are nearly identical to those of baseline QEMU execution (Fig. 9).  Conclusions  An extensible ISS acceleration method, SkipSync, is presented for sampling-based performance evaluation of LLM inference workloads. By combining the Skip mechanism with Host-Simulator Synchronization, SkipSync effectively alleviates the ISS bottleneck in LLM sampling workflows while preserving execution correctness and sampling accuracy. Experimental results demonstrate substantial speedups, low synchronization overhead, and negligible effects on sampling accuracy. The proposed method provides a practical solution for efficient accelerator evaluation in rapidly evolving LLM inference systems. Future work will extend SkipSync to more complex LLM inference optimization scenarios and explore automatic identification of skippable operators and insertion of synchronization primitives to further improve usability.
Reynolds Decomposition Motion-Guided Texture Learning for Scarred Myocardium Phenotyping
RUAN Dongsheng, YANG Daiguo, ZHANG Xiaolin, MA Jianhua, JIANG Mingfeng, WANG Yaming
Available online  , doi: 10.11999/JEIT260330
Abstract:
  Objective  Scarred myocardium is a key imaging marker of multiple cardiovascular diseases, and accurate phenotyping is clinically valuable for treatment planning and prognosis. Cine-MRI noninvasively provides both cardiac motion and anatomical texture information. However, existing classification methods have two major limitations. First, motion representation is often overly discretized and insensitive to subtle abnormalities. Second, multimodal fusion commonly relies on simple feature concatenation, which limits the ability to capture the deep pathological association between motion impairment and texture alteration. A motion-guided texture learning framework is therefore developed to improve the accuracy of non-invasive scarred myocardium classification.  Methods  A Motion-Guided Texture fusion Network (MGTNet) is proposed to enable deep interaction between motion and texture features. First, inspired by Reynolds decomposition in fluid dynamics, a Reynolds Decomposition Motion Network (RDMNet) is designed within a diffeomorphic registration framework to decompose the myocardial motion field into a regular mean motion component and an abnormal pulsatile motion component. This decomposition provides a refined motion representation that is sensitive to subtle abnormalities. Second, a Motion-Guided Cross-Attention (MGCA) module is designed, in which refined motion features serve as query features to dynamically modulate texture features and enhance the perception of suspected lesion regions. In addition, an inter-frame motion interaction module is used to capture temporal dependencies across cardiac frames, while an intra-frame texture extraction module learns multi-scale spatial texture patterns within each frame. Through a serial pipeline of motion field estimation, motion-guided texture enhancement, and feature fusion, end-to-end scarred myocardium classification is achieved.  Results and Discussions  Experiments on the CMRD and ACDC datasets show that MGTNet consistently outperforms existing single-modality and multimodal methods (Tables 1 and 2). It achieves accuracies of 95.6% on CMRD and 97.3% on ACDC, with AUC values of 96.6% and 96.0%, respectively. Compared with the baseline MTNet, MGTNet improves accuracy by up to 1.4 percentage points and the F1-score by up to 2.3 percentage points. Further comparisons with different motion field estimation methods (Tables 3 and 4) show that RDMNet provides more discriminative motion priors and achieves the best overall performance on both datasets. These results indicate that fine-grained motion modeling and motion-guided texture enhancement effectively capture complementary pathological information from Cine-MRI. The ROC curves of the nine methods on both datasets further support the superior classification performance of the proposed method.  Conclusions  A scarred myocardium classification method based on motion-guided texture learning is presented. Reynolds decomposition is incorporated into motion field estimation to separate regular mean motion from abnormal pulsatile motion, and cross-attention is used to guide texture extraction with motion priors. This strategy addresses the limitations of discrete motion representation and shallow multimodal fusion. The results confirm that deep motion-texture interaction improves the accuracy and robustness of non-invasive scarred myocardium classification and provides an effective approach for Cine-MRI-based myocardial phenotyping.
Survey of Satellite Covert Communications: Status, Key Technologies, and Future Challenges
DENG Hao, SUN Weiyuan, ZHU Zhengyu, PAN Gaofeng, SUN Gangcan
Available online  , doi: 10.11999/JEIT260177
Abstract:
  Objective  This survey systematically reviews the theoretical foundations, key technologies, and future challenges of Satellite Covert Communications. Based on the classical Alice-Bob-Willie model, the effects of satellite channel characteristics on covert communication capacity are analyzed to establish the theoretical basis. The network architecture of Satellite Covert Communications under the space-air-ground three-layer framework (Fig. 1) is summarized. Core technologies and optimization methods, including signal camouflage coding, beamforming, spectrum agility, Quantum Key Distribution (QKD), and Artificial Intelligence (AI)-assisted techniques, are systematically reviewed. Major security threats and corresponding multi-layer defense strategies, including Physical-Layer Security (PLS) and intelligent collaborative defense, are also summarized. This survey provides a theoretical foundation and technical guidance for developing highly secure and intelligent Satellite Covert Communications systems.  Significance   The significance of this survey lies in its systematic integration of the theoretical framework and key technologies for Satellite Covert Communications. To address the threats of detection, interference, and eavesdropping in open satellite communication environments, representative space-air-ground integrated architectures reported in the literature are reviewed. These architectures overcome the limitation of conventional encryption techniques, which protect only information content, by reducing statistical distinguishability at the physical layer to achieve a low probability of detection. The constraints imposed by satellite channels on covert communication capacity are clarified through the modified Square Root Law. Enhancement strategies based on adaptive coding, beamforming, spectrum agility, and related techniques are reviewed to establish a comprehensive technical framework for Satellite Covert Communications. These advances provide theoretical support and technical guidance for constructing highly survivable and secure space-air-ground integrated communication networks, with important applications in national defense, emergency communications, and Sixth-Generation (6G) Non-Terrestrial Networks (NTNs).  Progress   Existing studies demonstrate that Doppler spread has a dual effect on covert communication capacity. It increases the missed detection probability while introducing signal distortion, making adaptive coding necessary to maintain reliable transmission. The differences in detection capability among terrestrial, aerial, and orbital wardens (Willie) are quantified (Table 1), providing a theoretical basis for hierarchical defense design. Existing studies have also proposed multi-level covertness enhancement strategies. AI-assisted dynamic camouflage combined with sparse coding exploits background noise, inter-satellite links, and dynamic beamforming to improve covert throughput. At the network level, cooperative Unmanned Aerial Vehicle (UAV)-assisted transmission and dynamic spectrum coordination are identified as representative enhancement approaches (Fig. 3). Furthermore, a hierarchical defense framework is summarized from representative studies. This framework combines Reconfigurable Intelligent Surface (RIS)-assisted signal control, Stackelberg game-based incentives for cooperative jamming, XOR network coding, and federated learning for cross-domain threat feature sharing. These advances provide effective solutions for improving the covertness and security of Satellite Covert Communications.  Conclusions  This survey systematically reviews the theoretical foundations and recent advances in Satellite Covert Communications. The integration of multi-layer satellite constellations, dynamic aerial relay platforms, and software-defined networks supported by Quantum Key Distribution (QKD) enables resilient global covert communication. Extending the Alice-Bob-Willie model to practical satellite channels with non-ideal propagation characteristics provides guidance for covert throughput optimization and secure transmission. AI-assisted coding and waveform design further enable adaptive transmission strategies that respond to dynamic channel conditions. Future research should focus on robust covert transmission in dynamic Low Earth Orbit (LEO) environments, scalable constellation management, and the deep integration of AI and quantum technologies into 6G NTNs. The convergence of programmable satellites, Reconfigurable Intelligent Surfaces (RIS), and advanced machine learning is expected to further improve secure space communications.  Prospects   Future research should focus on four major directions: robust covert transmission under non-ideal channels, AI-assisted intelligent decision-making, integrated 6G NTN networking, and the integration of quantum communication technologies (Fig. 4). High-precision Doppler compensation techniques should be developed to mitigate rapid channel variations in LEO satellite systems. Robust covert transmission schemes should also be developed for imperfect Channel State Information (CSI), with deep reinforcement learning providing real-time resource optimization. Future studies should strengthen the integration of AI and quantum technologies by combining cross-layer QKD with covert transmission protocols and exploiting Software-Defined Satellite (SDS) capabilities for adaptive strategy deployment. Additional opportunities include using RIS to enhance spatial-domain signal control and applying blockchain technology to address trust management in multi-node cooperative networks. Efficient lightweight onboard algorithms and coordinated international frameworks for spectrum and orbital resource management should also be developed to support future Satellite Covert Communications systems.
A Novel TDMOSFET and Its Neural Network Modeling for Ternary Logic Applications
LU Bin, LU Haoran, ZHAO Xiaohong, DI Jiayu, XING Linlin
Available online  , doi: 10.11999/JEIT260413
Abstract:
  Objective  Complementary Metal-Oxide-Semiconductor (CMOS) technology continues to advance toward smaller device dimensions and higher integration. As circuit integration increases, short-channel effects and other phenomena increase leakage current in MOSFET devices, resulting in higher static power consumption. With the rapid development of artificial intelligence, traditional binary logic faces limitations in processing and storing massive amounts of data. Ternary logic has therefore attracted increasing attention because it provides higher information density and lower system complexity than binary logic. However, current ternary logic circuits still face challenges, including the use of multiple components, passive elements, and poor compatibility with conventional CMOS processes.  Methods  A novel Tunneling-Drift-Diffusion Metal-Oxide-Semiconductor Field-Effect Transistor (TDMOSFET) that combines quantum tunneling and drift-diffusion mechanisms is proposed. The device exhibits a constant off-state current characteristic, making it suitable for ternary logic applications. Its operating principle is analyzed, and an Artificial Neural Network (ANN) is used to model its electrical characteristics. The ANN model is trained using Technology Computer-Aided Design (TCAD) simulation data to predict the current-voltage (I-V) and capacitance-voltage (C-V) characteristics. The trained ANN model is further converted into a Verilog-A model and integrated into HSPICE to simulate basic ternary logic circuits, including the Standard Ternary Inverter (STI), Negative Ternary Inverter (NTI), Positive Ternary Inverter (PTI), Ternary NOT-AND gate (T-NAND), and Ternary NOT-OR gate (T-NOR).  Results and Discussions  The trained ANN model accurately predicts the I-V and C-V characteristics of the TDMOSFET. Compared with the TCAD results, the maximum relative errors for the drain current (IDS), gate-drain capacitance (CGD), and gate-source capacitance (CGS) are 39.43%, 5.05%, and 14.19%, respectively, whereas the corresponding average relative errors are 0.46%, 0.69%, and 0.51%. The ANN model is successfully converted into a Verilog-A model and integrated into HSPICE for circuit-level simulation. The STI, NTI, PTI, T-NAND, and T-NOR circuits are successfully simulated. The TDMOSFET-based ternary logic circuits do not require passive elements and are compatible with conventional CMOS processes.  Conclusions  A novel TDMOSFET with dual conduction mechanisms is proposed. When the gate-source voltage is below the turning voltage (Vturn), band-to-band tunneling is the dominant conduction mechanism, and the device operates similarly to a reverse-biased tunneling diode. When the gate-source voltage exceeds Vturn, the drift-diffusion mechanism becomes dominant, and the device exhibits characteristics similar to those of a conventional MOSFET. The proposed device maintains compatibility with conventional CMOS processes, simplifying the manufacturing process and reducing cost and integration complexity. TDMOSFET-based ternary logic circuits realize ternary operation without increasing the number of transistors, using passive elements, or requiring multivalued supply voltages. The proposed TCAD simulation → ANN modeling → Verilog-A modeling → HSPICE simulation framework can also be applied to the study of other emerging semiconductor devices.
An Adaptive Kalman Speech Enhancement Method Driven by Burst Noise Suppression and Dual-time-scale Perception
CHEN Bo, ZHENG ZeRui, SUN Chao, WANG ZheMing, SHEN Ying
Available online  , doi: 10.11999/JEIT260636
Abstract:
  Objective  Traditional Auto-Regressive (AR) Kalman speech enhancement algorithms face three critical challenges under non-stationary noise: limited adaptability of noise model updates, model mismatch caused by burst noise, and a lack of environment-aware covariance adjustment. These limitations degrade enhancement performance and restrict their use in practical speech communication. An improved adaptive Kalman speech enhancement method is therefore proposed to improve speech quality and intelligibility under complex noise conditions.  Methods  First, an environmental deviation measure based on the Energy Entropy Ratio (EER) is constructed to quantify the statistical deviation between the current frame and the background environment. A dual-time-scale EER tracking mechanism is then established to capture instantaneous variations and stable background statistics. Their difference is used to generate an adaptive deviation intensity factor for the joint adjustment of the process-noise and observation-noise covariances. Second, a two-stage burst noise discrimination scheme is developed based on the characteristics of burst noise. Frame energy changes and the Spectral Flatness Measure (SFM) are jointly used for preliminary burst noise detection. A Speech Modeling Metric (SMM) based on the linear prediction residual variance ratio is further used to distinguish speech from weak-colored burst noise and prevent contamination of the speech AR model.  Results and Discussions  Experiments are conducted on the NOIZEUS database. The proposed method outperforms the conventional AR-based Kalman filter and the sensitivity-based Augmented Kalman Filter (AKF) in terms of Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ), and Segmental Signal-to-Noise Ratio (SegSNR). It shows improved adaptability and stability under non-stationary and burst noise conditions. Dual-time-scale environmental perception and burst noise discrimination facilitate rapid adaptation to environmental changes and reduce speech distortion.  Conclusions  The proposed burst noise suppression and dual-time-scale adaptive Kalman method addresses the limitations of conventional AR-based Kalman speech enhancement under complex noise conditions. EER-based tracking enables environment-aware covariance adjustment, whereas the two-stage discrimination scheme improves burst noise detection and reduces contamination of the speech AR model. The experimental results demonstrate the robust performance of the proposed method and support its application to speech enhancement under non-stationary noise.
Intelligent Detection of DSSS Signals Under False-Alarm Rate Constraints Based on a Noise Score-Pool Threshold Calibration Mechanism
ZHANG Tao, TANG Xiaomei, SUN Guangfu
Available online  , doi: 10.11999/JEIT260414
Abstract:
  Objective  To address the degradation in Global Navigation Satellite System (GNSS) signal detection performance under weak-signal conditions and the limited control of the probability of false alarm (Pfa) in existing Deep Learning (DL) models, a DSSS signal detection method with a prescribed Pfa is investigated. The method enables the DL detector to be evaluated within a Constant False Alarm Rate (CFAR) framework.  Methods  DSSS signal detection is formulated as a binary classification problem, and a DL-based detection framework is developed. An improved one-dimensional ResNet-18 is designed for I/Q sampled time-series data. The input convolution kernel is set to 1×7, and the initial maximum pooling layer is removed to preserve weak-signal temporal features. The first residual layer is also configured without downsampling. A noise score-pool-based threshold calibration mechanism is developed to impose Pfa constraints on the detection decision. A large number of pure-noise samples are processed by the trained network to obtain the empirical distribution of confidence scores for the signal-present class. The decision threshold is then calibrated according to the quantile corresponding to the preset Pfa. In addition, the effect of signal normalization on detection performance is evaluated. The proposed method is validated using a simulated GPS L1 C/A signal dataset under different Signal-to-Noise Ratio (SNR) and Pfa settings and in non-ideal colored-noise environments.  Results and Discussions  The proposed method achieves high detection performance under different Pfa constraints. At a Pfa of 0.01, the detection probability reaches 100% at an SNR of –8 dB. When the Pfa decreases from 0.01 to 0.001 and 0.000 1, the detection curve shifts toward higher SNRs, but the decrease in detection performance remains limited. The unnormalized preprocessing strategy consistently outperforms Root-Mean-Square (RMS) normalization, providing a performance gain of approximately 1 dB. Compared with the traditional autocorrelation detection method, the proposed DL-based detector provides a detection performance gain of approximately 3~4 dB across the tested Pfa settings. In colored-noise environments not used for network training, the proposed method maintains effective detection performance and demonstrates robustness to noise mismatch. The structural ablation results further show that removing the maximum pooling layer and retaining the temporal resolution of the first residual layer improve detection performance under low-SNR conditions.  Conclusions  A DL-based DSSS signal detection method with noise score-pool-based threshold calibration is proposed. The empirical distribution of pure-noise confidence scores is used to calibrate the decision threshold, thereby incorporating the DL detector into a CFAR-based detection framework. The improved one-dimensional ResNet-18 effectively extracts features from I/Q time-series data, whereas the unnormalized preprocessing strategy preserves useful signal-amplitude information. The proposed method improves detection sensitivity while maintaining effective Pfa control and exhibits robustness under non-ideal colored-noise conditions.
A WiFi Multi-link Collaborative Human Tracking Method for Smart Home
PAN Houcheng, CAI Yushuang, YAO Junmei, ZHANG Tingting
Available online  , doi: 10.11999/JEIT260267
Abstract:
  Objective  WiFi-based indoor passive human tracking has attracted increasing attention for smart home applications because Internet of Things (IoT) devices are widely interconnected through existing WiFi infrastructures. Existing approaches estimate the Angle of Arrival (AoA) or Time of Flight (ToF) of the target-reflected path to achieve tracking. However, these methods are fundamentally limited by the difficulty of separating weak target-reflected signals from dense multipath propagation when commercial WiFi devices provide only a small antenna array and limited bandwidth under existing communication protocols. Therefore, Doppler Frequency Shift (DFS)-based approaches that exploit multiple links have become more practical. Dead Reckoning (DR) is widely adopted in these methods, but dynamic link selection for signal fusion remains challenging. Furthermore, the performance of DR-based methods degrades substantially when device-location errors are present. Existing methods also generally employ a single-transmitter-multi-receiver architecture, in which Access Point (AP) broadcasts downlink WiFi signals to STAtions (STA). It complicates data aggregation and limits the use of multiple receive antennas of AP’s. To address the challenges of multi-link signal fusion and device-location uncertainty, a Particle Filter (PF)-based method is proposed. Furthermore, a multi-transmitter-single-receiver architecture is adopted, in which multiple STAs transmit uplink packets to a single AP. This architecture simplifies data aggregation while exploiting the AP’s multi-antenna capability.  Methods  In the IEEE 802.11 standard, the wireless channel is estimated at the receiver using pilot signals embedded in WiFi packets, and Channel State Information (CSI) is continuously obtained. Human motion perturbs multipath propagation and produces time-varying changes in CSI. Therefore, CSI serves as the primary sensing signal because it implicitly captures target motion. In practice, raw CSI is affected by amplitude and phase impairments. Accordingly, signal preprocessing is first performed before Doppler extraction. Subsequently, multi-link DFS measurements are fused using PF, in which the target state is sequentially propagated and updated according to likelihoods derived from DFS measurements. This probabilistic framework naturally enables dynamic multi-link fusion because unreliable links receive lower weights rather than being deterministically discarded. Moreover, the effects of device-location errors are incorporated into the measurement-noise model, improving robustness to device-location uncertainty. When prior knowledge of device locations is unavailable, device self-localization is first performed by jointly estimating the Line-of-Sight (LoS) AoA and ToF between the AP and each STA. Preliminary device self-localization experiments (Fig. 5) demonstrate a median STA localization error of approximately 0.56 m (Table 1), providing reliable initialization for the subsequent tracking algorithm.  Results and Discussions  A prototype system is implemented using a multi-transmitter-single-receiver architecture with commercial Intel AX200/AX201 WiFi cards. CSI is collected over a 20 MHz channel centered at 5.24 GHz using PicoScenes, with IEEE 802.11ac packets transmitted at 100 or 200 Hz. Each CSI sample contains measurements from 57 subcarriers. Meanwhile, Fine Time Measurement (FTM) is performed on Channel 11 at 2.4 GHz using iw and hostapd. In the prototype system, each STA transmitter is equipped with a single antenna, while the AP receiver employs two antennas. Experiments are conducted using one AP and three or four STAs (Fig. 7). WiTraj and PITrack, both based on DR, are used as baseline methods. When device self-localization is required, the proposed method achieves a median tracking error of 0.47 m, compared with 1.78 m for WiTraj and 1.77 m for PITrack (Fig. 9), representing an accuracy improvement of approximately 70% over conventional DR-based methods. Error-injection experiments with random device-location perturbations show that, unlike conventional DR-based methods, which are highly sensitive to device-location errors, the proposed method remains robust to device-location uncertainty (Fig. 10). Additionally, computational complexity analysis shows that, although execution time increases with the particle count (Table 3), tracking performance reaches a stable level beyond a moderate particle count (Fig. 11). Therefore, real-time operation can be achieved by selecting an appropriate particle count. Finally, experiments in complex environments demonstrate that the proposed method consistently achieves higher tracking accuracy than the baseline methods (Fig. 12).  Conclusions  A PF-based method is proposed for multi-link collaborative passive human tracking using commercial WiFi devices. By probabilistically fusing measurements from multiple links, robust tracking is achieved even when device-location errors are present. When device locations are unavailable, device self-localization is achieved by combining CSI-based estimation with FTM measurements, providing the geometric information required for tracking. A prototype system based on a multi-transmitter-single-receiver architecture is developed to simplify data aggregation while exploiting the AP’s multi-antenna capability. Experimental results demonstrate that the proposed method achieves sub-0.5 m median tracking error when device-location errors are present, representing an improvement of approximately 70% over conventional DR-based methods. The proposed method also exhibits strong robustness to device-location uncertainty and supports real-time implementation when an appropriate particle count is selected. Future work will extend the framework to multi-person tracking and evaluate its performance under more realistic deployment conditions.
Cross-Domain Collaborative Enhancement for Tiny Object Detection in Remote Sensing Images
ZHANG Tianyang, ZHANG Xiangrong, WANG Guanchun, TANG Xu
Available online  , doi: 10.11999/JEIT260317
Abstract:
  Objective  Deep learning has substantially advanced object detection in Remote Sensing Images (RSIs). However, because of imaging conditions and the inherently small size of many objects, a large proportion of targets in RSIs occupy fewer than 16 × 16 pixels. Therefore, current object detection methods achieve substantially lower detection accuracy for tiny objects than for normal-scale objects. This limitation primarily arises from two critical factors: insufficient positive sample assignment and weak feature representation. To address these challenges, a Cross-Domain Collaborative Enhancement Detector (CDCEDet) is proposed. CDCEDet jointly optimizes label assignment in the spatial domain and enhances feature representation in the frequency domain, thereby improving the accuracy and robustness of tiny object detection in RSIs.  Methods  The overall framework of CDCEDet is illustrated in Fig. 2 and consists of three major components. First, a Scale-Adaptive Anchor Generator (SAAG) is designed to dynamically generate anchors that match the scales of ground-truth (GT) objects, thereby effectively alleviating the scale mismatch between anchors and tiny objects that has been largely overlooked in previous studies. Compared with conventional uniformly distributed anchor generators, SAAG substantially increases the number of positive samples assigned to tiny objects, even under Intersection over Union (IoU)-based label assignment. Second, a Quantile-based Adaptive Label Assignment (QALA) mechanism is developed to replace the conventional fixed IoU threshold-based label assignment. QALA models the IoU distribution between each GT object and its matched anchors to generate an adaptive label assignment threshold for each object, thereby further increasing the number of positive samples assigned to tiny objects. Third, a Frequency-Adaptive Fusion (FAF) module is developed to enhance feature representation from a frequency-domain perspective. An adaptive high-pass filter is used to strengthen high-frequency details and compensate for information loss caused by channel compression, whereas an adaptive low-pass filter preserves semantic consistency during feature upsampling, thereby reducing semantic inconsistency within upsampled objects.  Results and Discussions  Extensive experiments are conducted on two public remote sensing tiny object detection datasets, AI-TODv2 and AI-TOD-R. The proposed method is compared with several state-of-the-art methods, including RFLA, DCNet, and DCFL. On the AI-TODv2 dataset (Table 1), CDCEDet improves AP50 and AP50–95 by 1.8% and 0.7%, respectively, compared with the best existing method. On the AI-TOD-R dataset (Table 2), AP50 and AP50–95 are improved by 2.6% and 0.7%, respectively. These results demonstrate that CDCEDet achieves superior detection performance and strong generalization capability for tiny object detection in RSIs. Ablation studies and parameter analyses of the proposed modules (Tables 36) further verify the effectiveness of each component and their complementary contributions. Qualitative results on both datasets (Fig. 3) show that the proposed method accurately detects tiny objects in both sparse and dense scenes. As illustrated in Fig. 4, SAAG generates scale-matched anchors for individual objects and assigns substantially more positive samples to tiny objects than the conventional uniformly distributed anchor generator. Furthermore, visual comparisons with RFLA and DCNet (Fig. 5) demonstrate that CDCEDet achieves higher detection accuracy while substantially reducing missed detections.  Conclusions  A CDCEDet is proposed to address insufficient positive sample assignment and weak feature representation in remote sensing tiny object detection. Specifically, SAAG dynamically generates anchors that match the scales of GT objects, substantially increasing the number of positive samples assigned to tiny objects. QALA further improves label assignment by modeling the IoU distribution between GT objects and their matched anchors to adaptively determine the label assignment threshold, thereby effectively reducing the scale bias introduced by fixed IoU thresholds. In addition, FAF enhances feature representation from a frequency-domain perspective through an adaptive high-pass filter and an adaptive low-pass filter. Experimental results on two benchmark datasets demonstrate the superior detection performance and strong generalization capability of CDCEDet. Future work will focus on improving model efficiency and real-time performance to facilitate practical deployment in remote sensing applications.
DroneRFc-MM: Anti-UAV Multimodal Detection Measured Dataset
YU Taosong, YANG Qianqian, HU Zhuo, LI Mingkai, WU Jiajun, SU Yufan, PAN Junyu, SHI Zhiguo, CHEN Jiming
Available online  , doi: 10.11999/JEIT260889
Abstract:
Objective: A comprehensive multimodal benchmark is developed for Anti-Unmanned Aerial Vehicle (UAV) detection in low-altitude urban environments. Existing datasets generally provide limited sensing modalities and UAV models, with relatively coarse annotations that constrain tasks requiring spatial, motion, and cross-modal information. DroneRFc-MM addresses these limitations by providing synchronized multimodal data, broader coverage of consumer-grade DJI UAV models, and fine-grained annotations for target detection, UAV model recognition, trajectory analysis, flight-direction reasoning, and multimodal fusion evaluation. Methods: DroneRFc-MM is synchronously collected using six heterogeneous sensor types: a Pan-Tilt-Zoom (PTZ) camera, a fisheye camera, a Radio Frequency (RF) antenna, LiDAR, millimeter-wave radar, and a microphone array. Data are acquired on an open rooftop at a university in Zhejiang Province, representing a typical urban low-altitude environment. The dataset contains recordings of six consumer-grade DJI UAV models. All devices are synchronized using a common network time reference, with inter-device timestamp discrepancies of approximately 0.3 s. The UAVs fly in “H”-shaped and vertical reciprocating trajectories at distances of 20–60 m from the sensor array. Fine-grained annotations, including UAV model, position, attitude, and velocity, are derived from flight logs. For the flight-direction reasoning task, approximately 5-s multimodal clips are generated, including camera videos, RF spectrogram videos, microphone audio, and coordinate-based text representations of radar point-cloud data. Zero-shot inference is conducted using Qwen 3.6-Plus and Qwen 3.5-Omni-Plus with unified prompts. Prediction accuracy and inference time are evaluated by comparing predicted directions with ground-truth directions calculated from UAV positioning data. Results and Discussions: The DroneRFc-MM dataset provides multimodal data from six sensor types and six consumer-grade DJI UAV models, together with fine-grained annotations and sample extraction tools. In the flight-direction reasoning task, the Qwen-series multimodal large language models (MLLMs) achieve accuracies ranging from 20% to 30% across the different input modalities. The inference time is also relatively long, with the mean response time exceeding 40 s for most sensor inputs. These results indicate that current general-purpose MLLMs can capture weak motion-related information from UAV videos, audio, RF spectrograms, and point-cloud data, but their accuracy and response speed remain insufficient for practical real-time Anti-UAV detection. Conclusions: DroneRFc-MM provides a multimodal benchmark for Anti-UAV detection, UAV model recognition, flight-direction reasoning, and multimodal model evaluation. The dataset integrates six sensor types, six consumer-grade DJI UAV models, and fine-grained annotations within a common measurement framework. The experimental results show that current general-purpose MLLMs remain limited in flight-direction reasoning and real-time inference in Anti-UAV scenarios. Domain-specific pre-training, supervised fine-tuning, knowledge augmentation, and lightweight inference are therefore needed to improve their practical utility. Future work will expand the dataset scale and application scenarios to support intelligent and efficient low-altitude airspace management systems.
Dynamic Data Mapping and Co-Optimization Method for TSVs in 3D-Integrated MoE Accelerators
YANG Jialin, XIA Chenjie, WU Huiming, LI Ningyuan, SONG Yuan, LIU Bo
Available online  , doi: 10.11999/JEIT260565
Abstract:
  Objective  The rapid progress of large-scale intelligent computing, especially Mixture-of-Experts (MoE), has positioned Three-Dimensional Integrated Circuit (3D IC) based on Through-Silicon Via (TSV) as a key solution to memory-wall bottlenecks via high bandwidth and density. As a core 3D IC technology, TSVs enable vertical inter-chip connections, reducing path length, parasitic delays, power, and boosting data rates. MoE-specific accelerators, characterized by high data density and strong fault tolerance, introduce new challenges and opportunities for TSV layout. These include aggravated signal integrity and reliability issues in dense arrays, and the inadequacy of static TSV allocation for dynamic, bursty MoE traffic. Conversely, their inherent fault tolerance permits optimization design spaces for employing fault-tolerance mechanisms. This paper exploits MoE dataflow characteristics and hardware fault tolerance to devise a data mapping strategy for high-density TSV arrays based on fault-tolerance mechanisms, targeting improved performance and reliability.  Methods  This paper investigates cluster partitioning schemes and data mapping strategies to enhance the reliability of TSV data transmission. To address the high complexity of global optimization in large-scale TSV arrays, a cluster size partitioning scheme is proposed. By structurally partitioning a large-scale TSV array into several small-scale TSV clusters, the global optimization problem is decomposed into local, scalable subproblems, thereby improving optimization efficiency and flexibility while ensuring optimization moderation. Through comprehensive consideration of multiple metrics and simulation-based evaluation, the cluster size is finally determined to be 6×6. In response to the varying dataflow characteristics and load distribution across different computational stages, this paper proposes a Phase- and Load-Aware Dynamic Data Mapping (PLDM) strategy. The strategy pre-partitions the TSV array into multiple fixed-size clusters and classifies them into critical clusters and general clusters based on metrics such as coupling strength, bandwidth, and latency. At runtime, the PLDM strategy dynamically adjusts data mapping according to the characteristics of different computational stages. Furthermore, this paper achieves a co-optimization design of PLDM with the encoding circuit. The load monitoring module and the error monitoring module share certain data buffers and control status registers, enabling hardware resource reuse. Meanwhile, the error monitoring results provide real-time feedback on the reliability level of each TSV cluster, based on which the mapping controller preferentially allocates data transmission to clusters with lighter loads and lower bit error rates. This approach realizes resource sharing and load balancing, thereby improving data transmission reliability and link utilization efficiency for high-density TSV arrays.  Results and Discussions  This paper analyzes the bandwidth utilization and load balancing performance of three mapping schemes: random, static, and dynamic. The results show that both the dynamic and random mapping schemes achieve average bandwidth utilization close to the theoretical maximum. However, the random mapping scheme maps approximately 37.52% of critical data into general clusters with relatively high bit error rates, thereby increasing unreliability. Compared with static mapping, the dynamic mapping scheme improves average bandwidth utilization from 0.7982 to 0.8984, a relative increase of about 12.6%, reduces inter-cluster load fluctuation by 54.6%, and correspondingly improves load balancing by a factor of 2.2 (Fig. 4). Compared with random mapping, the dynamic mapping scheme reduces inter-cluster load fluctuation by about 8.6%, and reduces the latency of critical data and non-critical data by 36.6% and 34.7%, respectively (Table 2). To further evaluate the optimization effects of the proposed PLDM strategy on metrics such as load balancing and bandwidth utilization, four comparative schemes are configured: (1) Baseline scheme; (2) Static mapping scheme; (3) Sparse TSV layout using TSV-Aware Adaptive Fault-Tolerant Coding (TSV-AFTC) and PLDM; (4) High-density TSV layout based on scheme (3). Taking the Qwen3-30B-A3B model as an example, the normalized loads of 16 clusters in the Multi-Head Attention (MHA) and Feed-Forward Network (FFN) stages are compared across the four schemes. The results indicate that the proposed dynamic data mapping scheme achieves balanced load distribution across clusters in both the MHA and FFN stages, ranging from 0.48 to 0.52, while ensuring that all critical data are mapped to critical clusters. The high-density TSV scheme further reduces the load per cluster to approximately 0.34–0.37, demonstrating that dynamic mapping can effectively suppress stage-wise hot spots and improve load balancing (Fig. 7). Subsequently, system-level fault injection is applied to the transmitted data to simulate data reliability under extreme conditions for different schemes. The results show that for the proposed scheme (sparse), the degradation in perplexity (PPL) compared to the ideal case is controlled within 0.02, while the average bandwidth utilization is improved by approximately 15% and cluster load balancing is enhanced by a factor of 3.4. Under the high-density scheme, the PPL increase is controlled within 0.05, the average bandwidth utilization reaches about 71.7%, and the cluster load balancing is improved by a factor of 2 (Table 4).  Conclusions  This paper investigates TSV data mapping for 3D MoE accelerators and proposes a PLDM strategy based on TSV-AFTC, which allocates data from different computational stages to reliable and lightly loaded TSV clusters according to cluster-level bit error rates and load conditions. Through circuit co-design, approximately 12% of hardware resources can be saved. Compared with static mapping, the proposed scheme improves average bandwidth utilization by about 12.6% and enhances cluster load balancing by a factor of 2.2. Under system-level fault injection, the scheme limits the degradation of model inference perplexity to within 0.02, while achieving approximately 15% improvement in bandwidth utilization and a 3.4× enhancement in cluster load balancing.
A Dual-Trellis Message-Passing Decoding for Non-Binary LDPC Codes
XX XX
Available online  , doi: 10.11999/JEIT260958
Abstract:
  Objective  Due to their capacity approaching performance, Low Density Parity-Check (LDPC) codes have been widely applied to wireless communication and data storage systems. Compared to their binary counterparts, Non-Binary LDPC (NB-LDPC) codes with short or moderate code lengths have been demonstrated to achieve superior error performance under non-binary Belief Propagation (BP) decoding. However, the computational complexity of Check Node (CN) update of the optimal BP decoding is too complex for practical applications. Recently, many works have been presented to perform updates of CNs based on truncated messages, rather than full-length reliability messages, to significantly reduce the computational complexity of CN updates. Most of them construct the trellis of a CN based on the truncated input vectors, called truncated-trellis, such that CN updates are efficiently processed in parallel based on the selected candidate paths. These paths generally contain only a small number of deviation nodes, and such deviation nodes usually have high reliability. However, the Variable Node (VN) update in most decoding algorithms based on CN truncated-trellis still sequentially processes each element in the input vectors of each VN by the elementary steps. When the CN update is simplified, the complexity of the VN update may primarily determine the overall computational complexity. To address the above issues, this paper proposes the Dual-Trellis Min-Sum (DTMS) decoding algorithm. By further introducing truncated-trellises for VNs and updating the output messages of CNs and VNs in parallel, respectively, it further improves the decoding efficiency, while maintaining the similar decoding performance.  Methods  The different contributions of nodes in the CN truncated-trellis of the Pruning path Min-Sum (PMS) decoding algorithm on the selected highly reliable candidate paths are first analyzed, and it reveals that the selected highly reliable paths are primarily determined by the deviation nodes from the first few rows of the trellis of a CN, especially the second row. Thereby, it is not critical to update and sort every element of each output vector of one VN during the VN update. Next, a new trellis of one VN is constructed, and highly reliable elements over this trellis shared by all the output vectors of this VN are searched using a row-wise pruning strategy, such that the conventional element-wise VN updating procedure is transformed into a trellis-based parallel updating process based on an extra column in the trellis. In this basis, the unequal protection for the reliability values of each VN output vector is conducted, e.g., only the first few elements in each output vector of VN are updated and arranged, and the rest elements of each output vector are directly set to a compensation value. As a result, the computational complexity required for less reliable elements during each VN update can be significantly reduced, while retaining the crucial messages.  Results and Discussions  Experimental results show that compared with the PMS decoding algorithm using the original VN updating procedure, the proposed DTMS decoding algorithm maintains almost the same Bit Error Rate (BER) performance and convergence speed for decoding NB-LDPC codes under different finite fields, code lengths, and code construction methods (Figs. 37). Meanwhile, the number of real-domain operations required for the proposed simplified VN updates is reduced by approximately 71.8% on average (Table 2). In addition, the error-correction performance and convergence speed of the proposed DTMS decoding algorithm are close to those of the sub-optimal BP decoding algorithms (Figs. 37) with relatively low computational complexity (Table 3). The average performance gap of the DTMS decoding algorithm from the optimal BP decoding algorithm is only about 0.11 dB (Figs. 37). Thus, optimizing the VN updating is an effective way to further reduce the decoding complexity of truncated-trellis-based message-passing decoding algorithms.  Conclusions  This paper proposes a DTMS decoding algorithm to reduce the computational complexity of VN update in truncated-trellis-based decoding algorithms for NB-LDPC codes. Based on the CN updating process of the PMS decoding algorithm, the proposed algorithm further constructs a truncated-trellis and introduces the unequal protection scheme for VN update, such that the output vectors of each VN can be efficiently updated in parallel. Experimental results show that, under the same CN trellis-based update, the proposed parallel VN updating method significantly reduces the computational complexity compared with the original VN updating method, while maintaining similar decoding performance. Moreover, the proposed DTMS decoding algorithm performs closely to the sub-optimal BP decoding algorithms with similar convergence speed and lower complexity. In future studies, it will be interesting to further exploit the adaptive pruning strategies for the VN parallel updates. Based on the distribution of field elements from different iterations, less reliable field elements can be adaptively eliminated to reduce the set of candidate field elements, which may further reduce the complexity of VN update with negligible performance loss.
Study on Deployment Optimization of Reconfigurable Intelligent Surface for Troposcatter Communications
ZHAO Ziyan, SONG Zhiqun, LIU Lizhe, LI Yong, LI Xingjian, WANG Bin
Available online  , doi: 10.11999/JEIT260922
Abstract:
  Objective   Troposcatter communication serves as a valuable complement to satellite communication and thus is still quite promising in scenarios such as military long-distance communication. However, when a troposcatter communication system is deployed in mountainous environments, it is often faced with a prevalent and challenging engineering problem known as the “Line-of-Sight (LoS) obstruction”. Traditional solutions to this issue are still confronted with engineering difficulties. Increasing the antenna elevation angle to cross obstacles makes the scattering angle increase sharply and consequently lead to transmission loss surging beyond acceptable link budget limits; alternatively, building tall towers to raise antenna height preserves low-angle transmission but introduces construction difficulties and sacrifices the advantage of terrain concealment. Reconfigurable Intelligent Surface (RIS) has emerged as a disruptive technology in wireless communications, with the capability of reconstructing the wireless environment and artificially altering channel characteristics. It has been successfully applied in various civilian mobile communication systems. Obviously, it also provides an alternative to address the problem of LoS obstruction in troposcatter communications. Unfortunately, it has never been reported that RIS had been applied in such scenarios. Herein, to solve the LoS obstruction problem in troposcatter communications, RIS is involved for the first time in this field, a conceptual architecture of RIS-assisted troposcatter communication is set up, and then the problem of optimal RIS deployment is systematically investigated.  Methods   Based on the proposed framework of RIS-assisted troposcatter communication system, the deployment optimization of RIS is addressed step by step:Firstly, a three-dimensional model of feasible deployment region is established under four types of practical engineering constraints, i.e., the intrinsic constraint of obstacle-crossing, the optimal constraint of engineering upper bound, the antenna radiation constraint of Fresnel near-field region, and the hardware constraint of RIS effective angle.Secondly, the optimization problem of RIS deployment is formulated as minimizing the comprehensive system gain loss. The overall loss consists of three major components, namely, troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles. The first two parts are easily computed according to corresponding engineering knowledge of troposcatter communication and typical antenna theory, respectively. Then to calculate the dynamic gain attenuation of Cassegrain antennas, a quantitative model is developed based on cantilever beam bending theory. The model quantifies the pointing errors of a Cassegrain antenna caused by dynamic over-compensation with its elevation angle adjustment, and then the gain loss is calculated with the assistance of Taylor radiation pattern.Thirdly, through theoretical analysis and numerical verification via sectional slicing heatmaps, a dimensionality reduction property of the objective function is observed and validated. Within the feasible region, the first-order partial derivative of the objective function with respect to deployment height is always negative, implying that the global optimal deployment position necessarily lies on the upper boundary surface of the feasible region. The dimensionality reduction property is rigorously validated through slice analysis across the entire feasible domain, with more than 46,100 verification points confirming that the optimal position always resides on the upper boundary surface. This finding reduces the intractable three-dimensional constrained optimization problem to a much simpler two-dimensional manifold optimization, which significantly reduces the computational complexity of the optimization.Finally, based on this dimensionality reduction property, an improved gradient descent algorithm with momentum and adaptive backtracking line search (IGD-M&ABLS) is put forward. The algorithm introduces momentum gradient updates to suppress zigzag oscillations and accelerate convergence; it also incorporates an adaptive backtracking line search strategy to dynamically adjust step sizes, balancing iterative stability with computational efficiency.  Results and Discussions   A series of simulation experiments are conducted under typical engineering parameters, i.e., a 3-meter aperture Cassegrain antenna; 5 GHz signal frequency; mountain heights of 150 m, 200 m and 250 m, representing medium-high hills, the dividing line between hills and mountains, and relatively mountainous terrain, respectively. The proposed IGD-M&ABLS algorithm is benchmarked against Grid Search (GS, a classic deterministic exhaustive-search method) and Particle Swarm Optimization (PSO, a typical efficient heuristic algorithm). The results demonstrate that IGD-M&ABLS consistently converges to the global optimal solution with less gain losses than both benchmarks. Specifically, for the typical case with a 200 m mountain height, in 50 independent runs, IGD-M&ABLS always achieves the best objective function value of 11.6094 dB, much more steadily than PSO does, and it also outperforms GS's 11.6127 dB. As far as time consumption is concerned, IGD-M&ABLS exhibits remarkable advantages of computational efficiency. Its average runtime is approximately 0.005 seconds, compared with 0.07 seconds of PSO and more than one hour of GS. This order-of-magnitude improvement in computational speed is consistent with the theoretical complexity analysis. IGD-M&ABLS optimizes two independent variables on a 2D manifold, whereas PSO and GS handle three independent variables in the 3D feasible domain. Robustness tests under varying terrain conditions (H = 150 m and H = 250 m) confirm that IGD-M&ABLS reliably obtains the best results across different scenarios. In all simulation tests, IGD-M&ABLS demonstrates excellent stability and reproducibility, producing consistent results across multiple independent runs, while PSO exhibits randomness-induced variations and GS remains limited by its discretization step size.  Conclusions   This paper pioneers the application of RIS technology in troposcatter communication, providing a new technical solution to address the LoS obstruction problem in mountainous environments. A conceptual framework of RIS-assisted troposcatter communication system is established, incorporating a three-dimensional feasible RIS deployment region model with four practical engineering constraints. Then the optimization problem of RIS deployment is formulated as minimizing the overall system gain loss including troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles, and the computation method of its third term is also developed for the first time based on cantilever beam bending theory. More interestingly, the objective function is found and verified with dimensionality reduction property, i.e., its minimum value always resides on the upper boundary manifold surface. That property effectively transforms the complex 3D optimization into an equivalent 2D manifold problem. Finally, a new algorithm IGD-M&ABLS is proposed by introducing momentum and adaptive backtracking line search into the traditional gradient descent framework. Simulation results show that compared with benchmarks, IGD-M&ABLS algorithm achieves the best deployment positions with order-of-magnitude faster computation, while maintaining excellent stability and reproducibility.
Non-Orthogonal PSWFs Signal Detection Method Based on Adaptive Temporal-Spatial Feature Fusion
CHEN Wenhua, MAO Zhongyang, LU Faping, SUN Ye, GAO Yixuan
Available online  , doi: 10.11999/JEIT260024
Abstract:
  Objective   To address the demands of B5G/6G systems for high spectral efficiency and transmission reliability, Prolate Spheroidal Wave Functions (PSWFs)-based non-orthogonal modulation has attracted extensive research interest because of its strong time-frequency energy concentration. However, severe mutual interference among multiplexed PSWF signals degrades the performance of conventional detection methods in complex channel environments. Existing methods are limited by ideal channel assumptions or single-modal feature extraction and therefore cannot fully exploit the temporal and spatial information of PSWF signals or adapt to dynamic interference. An Adaptive Temporal-Spatial Feature Fusion (ATSFF) architecture is proposed for accurate and robust detection of non-orthogonal PSWF signals.  Method   A dual-path parallel framework is constructed to extract complementary temporal and spatial features. A Gated Recurrent Unit (GRU) network extracts deep temporal features and captures long-term dependencies from one-dimensional received signals. In the other path, one-dimensional signals are transformed into two-dimensional representations using the Gramian Angular Difference Field (GADF), and hierarchical spatial features are extracted using ResNet50. An adaptive probability-weighted fusion mechanism dynamically adjusts the contributions of the two feature branches according to their prediction uncertainty, thereby integrating complementary temporal and spatial information and improving detection robustness.  Results and Discussion   Simulations on a 32-class non-orthogonal PSWF signal dataset (Fig. 2) show that the proposed ATSFF method outperforms coherent detection, cross-term detection, Approximate Message Passing-Interleave Division Multiple Access (AMP-IDMA), and Temporal Multiple Sparse Bayesian Learning-Least Squares (TMSBL-LS) over the full Signal-to-Noise Ratio (SNR) range. t-SNE visualization (Fig. 4) shows that the fused features achieve better inter-class separation and greater intra-class compactness. At a bit error rate of 4 × 10–5, the proposed method achieves a gain of approximately 0.2 dB over cross-term detection (Fig. 6). Although ATSFF has higher computational overhead and lower real-time performance than conventional methods, its single-sample inference cost remains fixed after the network architecture is established, and GPU-based batch processing is supported. The method is therefore suitable for communication scenarios with high detection-accuracy requirements.  Conclusions   An adaptive temporal-spatial feature fusion method is proposed for non-orthogonal PSWF signal detection under severe mutual interference. Dual-path feature extraction is achieved using GRU and ResNet50, and a prediction-uncertainty-based adaptive probability-weighted fusion mechanism is used to integrate complementary temporal and spatial features. The simulation results demonstrate improved detection accuracy and robustness under complex channel conditions. The proposed method provides a feasible approach for high-accuracy detection of non-orthogonal PSWF signals.
UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos
LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian
Available online  , doi: 10.11999/JEIT260221
Abstract:
  Objective  With the rapid development and extensive application of unmanned aerial vehicle (UAV) observation platforms, remote sensing video data with high spatiotemporal resolution has witnessed explosive growth. Understanding dynamic relationships in UAV videos is recognized as a pressing research challenge in intelligent remote sensing analysis. Current research on video scene graph generation is mainly focused on natural videos, while studies on remote sensing videos remain in the exploratory stage, with a lack of benchmark datasets annotated with high-order dynamic relationships. Traditional visual models are difficult to be directly adapted to remote sensing scenes, which are characterized by dynamic view changes, extreme object scale variations, and dense target distributions. To address these limitations, a UAV video relationship dataset (UAVREL) with high-order dynamic relationship annotations is constructed, a hypergraph-enhanced transformer method tailored to the characteristics of UAV videos is proposed, and a unified benchmark evaluation system for video scene graph generation in the field of UAV remote sensing is established in this study. These efforts promote the transformation of intelligent remote sensing interpretation technology from static object cognition to global dynamic scene understanding.  Methods  This study is mainly composed of three core parts: dataset construction, model design, and comprehensive experimental validation. First, in the dataset construction stage, the Unmanned Aerial Vehicle Benchmark for Object Detection and Tracking (UAVDT) dataset is selected as the basic data source. A semi-automatic annotation strategy is proposed, integrating object tracking, multi-person collaboration, and consensus verification to mitigate subjective bias and improve efficiency. The relationships are divided into three levels according to the requirements of cross-frame inference, among which high-order relationships are annotated with high priority. Second, A Hypergraph-Enhanced Transformer model (STHG) is proposed in this work. It includes six core functional modules: basic feature extraction, pairwise context encoding, spatial-temporal dual hypergraph enhancement, multi-source feature fusion, global temporal modeling and relation prediction. On this basis, a multi-label classifier generates the final dynamic scene graphs. In the experimental design, to verify the effectiveness of the dataset and the model, an object detection task and three video scene graph generation subtasks, namely predicate classification (PreCls), scene graph classification (SGCls), and scene graph detection (SGDet), are conducted on the UAVREL dataset. Representative object detection models and classic scene graph generation models are selected as baselines. Mean Average Precision (mAP) and mAP@50 are adopted as evaluation metrics for detection, while Recall@K and meanRecall@K (mR@K) are used for scene graph generation to comprehensively evaluate the performance of the model in relation recognition and graph construction.  Results and Discussions  In the object detection task, the impact of model architectures on detection performance is systematically verified through comparative experiments on eight models with different architectures (Table 2). The results show that the selection of model architecture is of crucial importance to the mean Average Precision (mAP), a core evaluation metric. Single-stage anchor-free models represented by VFNet and DDOD exhibit significant advantages in comprehensive detection performance. From the perspective of category characteristics, all models perform poorly in detecting small-scale and easily occluded target categories, which reflects the common technical challenge faced by current general object detectors in small target detection tasks. In the video scene graph generation task, five methods are tested on three subtasks respectively (Table 3, Table 4, Table 5). The STHG method proposed in this paper shows significant performance advantages in all three core tasks. Meanwhile, experimental data indicate that the value of the average recall metric is consistently significantly lower than that of the traditional recall metric. This phenomenon clearly shows that the dataset poses great modeling challenges in task scenarios with low-frequency object relationships, and implicitly reflects that relationship prediction in such complex scenarios remains a key challenge to be solved urgently in the field of drone video scene graph generation.  Conclusions  This paper focuses on the critical theme of understanding dynamic relationships in drone videos. It constructs the UAVREL benchmark dataset, providing data support with deeper semantic relationships for this field. And it proposes a Hypergraph-Enhanced Transformer approach for Remote Sensing Videos. Experimental results demonstrate that this approach achieves superior performance across multiple evaluation metrics, thus validating its practical applicability in remote sensing dynamic relationship prediction tasks. Through dataset construction and algorithmic innovation, this paper not only lays a solid data foundation for understanding drone video relationships, but also facilitates a leap from low-level semantic analysis to high-level dynamic semantic cognitive modeling in remote sensing video analysis. Future research will focus on the following two directions: Firstly, deepening the temporal dimension modeling of dynamic remote sensing scene graph generation to enhance the ability to capture long-term evolutionary events and complex relationships; secondly, expanding the scene coverage and diversity of relationship categories in the dataset, continuously improving algorithm benchmarks, and promoting technological iteration and industry application implementation.
Random-Linear-Network-Coding-based Cooperative Reliable Transmission Protocol for Underwater Acoustic Communication Networks
ZHANG Zhilin, PU Zhanqing, ZHU Yunan, LI Xueying, TIAN Jie, HUANG Haining
Available online  , doi: 10.11999/JEIT260648
Abstract:
  Objective  Reliable data delivery in underwater acoustic communication networks is challenged by high packet error rates, long propagation delays, limited bandwidth, and topology variations. In single-source dual-destination multi-hop transmission, the same data generation must be reliably delivered to two destination nodes. Packet losses at individual hops can accumulate during multi-hop forwarding and joint recovery at the two destinations, further complicating reliable delivery. Existing reliability-enhancement mechanisms, including retransmission, redundant forwarding, forward error correction, and multipath redundant transmission, generally rely on predetermined forwarding structures or fixed redundancy configurations. They have limited capability to exploit complementary coded information distributed among multiple relay nodes, resulting in insufficient joint recovery capability and high redundant transmission overhead. To address these limitations, a Network-Coded Cooperative Reliable Transmission Protocol for underwater acoustic communication networks (NCCRTP) is proposed.  Methods  NCCRTP operates on a generation basis and employs Random Linear Network Coding (RLNC) over the Galois field \begin{document}$ \text{GF}({2}^{8}) $\end{document}. To reduce coding overhead, each packet carries a Code IDentifier (CodeID) rather than the complete coding vector. The corresponding coding vector is recovered from a shared coding-vector dictionary at the relay node. During hop-by-hop forwarding, NCCRTP generates forward candidate structures subject to a residual-hop decreasing constraint and adaptively selects among three transmission modes: SINGLE, COOP, and BRANCH. SINGLE maintains a shared forwarding process toward the two destinations. COOP enables two relay nodes to jointly utilize linearly independent coded packets received at different nodes. BRANCH divides the transmission into two branches toward the different destinations. For each candidate structure, NCCRTP estimates the link success rate, calculates the required transmission budget, and evaluates the two-hop structural utility. The forwarding mode is then selected according to the tradeoff between recovery capability and transmission overhead.  Results and Discussions  Simulation results show that NCCRTP achieves the highest Joint Packet Delivery Ratio (JPDR) under both regular and random topologies. In the controlled comparison with Cooperative Uncoded transmission (CU), Single-branch Uncoded transmission (SU), and Single-branch Coded transmission (SC), NCCRTP consistently outperforms schemes using only cooperative forwarding or only RLNC. This result indicates that the reliability gain is jointly provided by distributed relay cooperation and joint utilization of linearly independent coded packets (Fig. 5). As the packet error rate increases or the end-to-end transmission depth increases from 3 to 7 hops, NCCRTP maintains a higher JPDR, demonstrating stronger robustness under lossy multi-hop conditions (Figs. 5(a) and 5(b)). In random topologies, Vector-Based Forwarding (VBF) and Focused Beam Routing (FBR) are separately combined with packet REPlication (REP) or RLNC to form the VBF+REP, VBF+RLNC, FBR+REP, and FBR+RLNC schemes (Figs. 6 and 7). Under medium-to-high packet error rate or multi-hop transmission conditions, NCCRTP improves the JPDR by up to approximately 50%, while reducing the equivalent transmission overhead per successful joint delivery by up to approximately 40% (Fig. 6). These results indicate that NCCRTP improves dual-destination reliability through adaptive forwarding-structure selection, link-quality-based transmission-budget control, and joint utilization of linearly independent coded packets rather than simply increasing redundant transmissions.  Conclusions  The reliability and redundant transmission overhead challenges in single-source dual-destination underwater acoustic multi-hop transmission are addressed by designing a cooperative transmission structure that enables RLNC to exploit distributed reception and complementary coded information among relay nodes. The proposed NCCRTP protocol adaptively selects the SINGLE, COOP, and BRANCH transmission modes according to residual-hop constraints, link-quality-based transmission-budget control, and two-hop structural utility evaluation. A lightweight coding-vector representation based on CodeID is also adopted to reduce the header overhead associated with carrying complete coding vectors. The protocol is evaluated under both regular and random topologies. The results show that: (1) NCCRTP achieves the highest JPDR among all compared schemes, demonstrating stronger joint recovery capability at the two destination nodes; (2) under medium-to-high packet error rates or multi-hop transmission conditions, NCCRTP improves the JPDR by up to approximately 50%; and (3) the equivalent transmission overhead per successful joint delivery is reduced by up to approximately 40%, indicating that the reliability gain mainly comes from adaptive structure selection, transmission-budget control, and joint utilization of linearly independent coded packets rather than excessive redundant transmissions. Future work will extend NCCRTP to more complex multi-source, multi-destination, multi-hop scenarios and further investigate its implementation and performance under node mobility and realistic underwater acoustic channel dynamics.
LEO Satellite Multi-beam Multicast Precoding and User Grouping Joint Optimization Algorithm
GUO Lili, FENG Yimeng, YUAN Peihong, GAO Yue
Available online  , doi: 10.11999/JEIT260375
Abstract:
  Objective  In Sixth-Generation (6G) Low Earth Orbit (LEO) satellite communication systems, multicast precoding is adopted to mitigate severe inter-beam interference caused by Full Frequency Reuse (FFR). However, conventional precoding algorithms exhibit cubic computational complexity, limiting their applicability to massive Multiple-Input Multiple-Output (MIMO) systems. Existing user grouping methods also fail to satisfy the fixed group-size requirement specified by the DVB-S2X standard. To address these limitations, a joint optimization framework is proposed that combines a low-complexity unsupervised deep learning-based precoding model with improved user grouping algorithms to improve the system sum rate and fairness.  Methods  An unsupervised deep learning model based on a hybrid Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) architecture is proposed for precoding (Fig. 2). The Convolutional Neural Network (CNN) extracts spatial features from Channel State Information (CSI), while the Long Short-Term Memory (LSTM) network captures high-level feature correlations. The model is trained by directly maximizing the system sum rate while satisfying the Per-Antenna power Constraint (PAC), without requiring supervised labels. For user grouping, two algorithms compatible with the DVB-S2X standard are developed. First, the CK-means algorithm extends conventional K-means clustering to ensure an equal number of users in each group while preserving high intra-group channel similarity. Second, the Fairness-Aware MAUG (FA-MAUG) algorithm prioritizes users with poor channel conditions during grouping, thereby improving system robustness and fairness.  Results and Discussions  The intra-group similarity metric is used to evaluate user grouping performance. The results show that the CK-means algorithm achieves an average similarity approximately 0.1 higher than that of the MAUG algorithm and nearly 0.5 higher than that of random grouping across different group sizes (Fig. 3), resulting in improved beamforming gain. In terms of system sum rate, the proposed CNN-LSTM precoding combined with CK-means grouping consistently outperforms the conventional Minimum Mean Square Error (MMSE) algorithm under different Signal-to-Noise Ratios (SNRs) and total transmit power levels (Fig. 4 and Fig. 5). Under different SNR conditions, the proposed CNN-LSTM precoding scheme improves the average system sum rate by 48.59% compared with the MMSE algorithm, whereas CK-means grouping increases the average system sum rate by 30.12% relative to random grouping. The effects of the number of users per group, the number of groups, and the number of antennas are further evaluated (Fig. 6-Fig. 8), demonstrating that the proposed framework maintains superior performance across systems of different scales. Complexity analysis further shows that the proposed precoding method reduces the online computational complexity from the cubic complexity of the conventional MMSE algorithm to linear complexity with respect to the number of antennas, making it well suited for real-time deployment in large-scale LEO satellite communication systems.  Conclusions  A joint optimization framework is proposed for LEO satellite multicast communication systems to address the high computational complexity of precoding and the limited fairness of conventional user grouping methods. Simulation results demonstrate that the proposed framework substantially improves the system sum rate, achieving an average gain of 48.59% over the conventional MMSE algorithm under different SNR conditions while simultaneously reducing online computational complexity. The proposed framework provides an effective and scalable solution for multi-beam interference mitigation and resource optimization in future 6G LEO satellite communication systems.
Image Classification Network Based on Complementary Decay Learning
YUAN Heng, TIAN Wenyue, ZHANG Shengchong
Available online  , doi: 10.11999/JEIT260751
Abstract:
  Objective  Image classification depends on complete and discriminative feature representations. Existing convolutional neural networks usually enhance positive and high-amplitude responses through activation functions and attention mechanisms. However, excessive reliance on dominant positive responses may shift attention from the whole object to local salient regions, while negative and low-amplitude responses containing edge, texture, and foreground-background transition information are often weakened. To address this problem, a Complementary Decay Learning Network (CDLNet) is proposed to suppress high-response dominance and preserve complementary feature information.  Methods  Inspired by the signal attenuation mechanism of the biological visual system, CDLNet introduces complementary decay learning into a residual network. The Spatial Complementary Decay (SCD) module divides features into positive-response, negative-response, and global-response branches, and applies differentiated decay to preserve salient regions, boundary details, and contextual information (Fig.2). The Channel Complementary Decay (CCD) module attenuates high-response channels while retaining middle- and low-response channels, thereby reducing channel dominance and promoting cooperative channel representation (Fig.3). The Complementary Decay Attention (CDA) module integrates SCD and CCD in parallel and is embedded into ResNet residual blocks to jointly regulate spatial structures and channel semantics (Fig.4, Fig.6).  Results and Discussions  Experiments are conducted on CIFAR-10, CIFAR-100, SVHN, Imagenette, Imagewoof, and ImageNet datasets. CDLNet achieves classification accuracies of 96.68%, 81.81%, 97.38%, 92.45%, 85.46%, and 62.88%, respectively. Compared with ResNet-34 and representative classification networks, CDLNet obtains higher accuracy on multiple datasets (Table 5, Table 6). Ablation experiments demonstrate that removing either SCD or CCD reduces classification accuracy, indicating that spatial and channel complementary decay both contribute to feature regulation (Table 4, Table 5). Visualization results show that CDLNet can enhance target regions, preserve structural details, and suppress irrelevant background responses (Fig.1, Fig.12). Although CDA increases parameters and computation, the accuracy improvement shows a reasonable balance between performance and complexity (Table 3).  Conclusions  CDLNet introduces spatial and channel complementary decay mechanisms into residual networks. By suppressing excessive high-response dominance and preserving negative-response and low-amplitude information, the proposed method improves the completeness and balance of feature representations, alleviates attention drift, and enhances image classification performance. Future work will further optimize the decay strategy and model complexity to improve efficiency and generalization.
A Heterogeneous Multi-View Semantic Fusion Training Method for Text Classification in Low-Resource Scenarios
XU Sen, DENG Jubiao, XU Xiufang, YAO Shanliang, BIAN Xuesheng, BEN Xianye
Available online  , doi: 10.11999/JEIT260365
Abstract:
  Objective  The increasing demand for text classification in specialized domains, such as finance, healthcare, and social media, is often hindered by the scarcity of labeled data. Low-resource or data-scarce scenarios significantly limit the effectiveness of conventional supervised learning and pre-trained language models, such as BERT, due to insufficient task-specific supervision. Existing methods that rely on single-view weak supervision or clustering often generate noisy pseudo-labels and fail to fully exploit the discriminative potential of hard-to-classify samples. Therefore, it is essential to develop a robust training method that leverages unlabeled data, mitigates noise sensitivity, and enhances semantic representation in low-label conditions. The primary objective of this research is to propose a text classification training framework specifically designed for data-scarce scenarios.  Methods  To address these challenges, a Heterogeneous Multi-View Semantic Fusion (HMVSF) training framework is introduced. HMVSF consists of three main stages: (1) Dual Clustering Consistency Screening (DCCS): Text samples are first represented in the word-frequency view using TF-IDF vectors. Two heterogeneous clustering algorithms, sIB and an improved K-means initialized via a genetic algorithm, are applied to the same feature space. Samples that are consistently assigned to intersecting clusters are selected as high-confidence pseudo-labeled instances. These samples are subsequently used for a lightweight supervised update of the shared encoder, ensuring reliable initialization for subsequent training. (2) Latent Dirichlet Allocation (LDA) Guided Hard-Sample Retraining: The remaining hard examples, which are not selected in the first stage, typically exhibit low discriminability in the word-frequency view but carry global semantic and topic-level information. LDA is employed to model the topic distribution of these samples. The optimal number of topics is adaptively determined by a combination of semantic coherence and statistical fitting criteria. Hard and soft pseudo-labels are generated based on the document-topic distribution, and the model undergoes targeted retraining using a combined cross-entropy and symmetric KL divergence loss to integrate topic-level semantic cues. (3) Downstream Fine-Tuning: Finally, the encoder is fine-tuned on a small labeled set by replacing temporary heads with a task-specific classification head. Standard cross-entropy is employed as the objective. The two-stage intermediate training is independent of the pre-training initialization, allowing HMVSF to be applied to both general-purpose BERT and masked language model (MLM)-based initializations, making it widely applicable.  Results and Discussions  The effectiveness of HMVSF is evaluated on six publicly available datasets, including DBpedia, AGNews, ISEAR, SMS Spam, Subjectivity, and Polarity, which cover both topic-based and non-topic-based text classification tasks. Two initialization routes are tested: Route A (BERT/BERTIT:CLUST) and Route B (BERTIT:MLM/BERTIT:MLM+CLUST). Under extremely low label budgets (e.g., 64 labeled samples), CCS-BERTIT:CLUST and CCS-BERTIT:MLM+CLUST outperform baseline models and single-stage variants across datasets (Figs.23, Tables 23). For instance, in Route A, CCS-BERTIT:CLUST achieves up to 9% accuracy improvement on DBpedia and ISEAR compared with BERTIT:CLUST, while error rates decrease by 10–54% (Table 2). Normalized Mutual Information (NMI) and intra-cluster Euclidean distance analyses confirm that the CCS selection produces stable and compact clusters, ensuring reliable pseudo-labels (Tables 45). Ablation experiments show that the two-stage design combining CCS and LDA is superior to single-stage variants, demonstrating complementary advantages: CCS provides high-confidence core samples, and LDA captures latent semantic structures of hard examples (Table 7). Parameter sensitivity analysis reveals that the model is robust under a wide range of cluster numbers and LDA weight coefficients, achieving stable performance when the LDA weighting coefficient is set to 0.3 (Figs.45). HMVSF also maintains competitive performance as labeled data increase. Computational experiments on an NVIDIA A40 GPU demonstrate that the two-stage intermediate training introduces acceptable additional overhead while enhancing representation quality and overall accuracy.  Conclusions  This study presents a novel heterogeneous multi-view semantic fusion training method (HMVSF) for data-scarce text classification. By integrating cluster-consistent sample selection in the word-frequency view and LDA-guided hard-sample retraining, HMVSF effectively leverages unlabeled data, mitigates pseudo-label noise, and enriches semantic representations. Extensive experiments across multiple datasets confirm that HMVSF significantly improves classification accuracy under low-label conditions while remaining robust to pre-training initialization. The framework is generalizable to various pre-trained language models and represents an effective and practical solution for text classification in low-resource scenarios.
Bayesian-Optimized Neural Network Rapid Solver for HEMP Waveform Distribution
WANG Jinjin, ZHAO Mo, JING Jing, WANG Wenbing, LIU Zheng, LI Jinxi, WU Wei, LIU Tieming
Available online  , doi: 10.11999/JEIT260877
Abstract:
  Objective  High-altitude Electromagnetic Pulse (HEMP) has large area, strong fields. It is difficult for HEMP's damage of critical infrastructures which are crucial to country's electromagnetic security. Especially under E1 environment, the waveform distribution characteristic will have direct influence on the effectiveness of protection design. But for cases which need a lot of waveforms calculation like system level effects simulation and protection scheme optimization etc., traditional calculation method of waveform parameter in coverage region is time consuming, which can not realize the real-time calculation of massive ground waveforms on the field distribution. The rapidly and accurately calculation of other parameters of the waveform function is still an issue for the electromagnetic environment field, limiting the useful application of computed fast HEMP environment results for a particular type of explosion when using these results in another calculation. This approach solves a problem that conventional computations cannot be linked into software to perform on-line computation, which provides the basis for the simulation calculation as well as HEMP damage evaluation, enables a fast comparison between different large scale scenarios, enabling us to calculate waveform functions for hundreds of different cases within a few minutes and thus constitutes an efficient tool for rapid analyses as part of the HEMP robust design and evaluation.  Methods  In this paper, an approach based on Bayesian optimization for deciding how many nodes should be included within each hidden layer of a multilayer feed forward NN is proposed. Reducing searching times for hidden layer node and ensure model precision, in fast solution modeling of HEMP waveform function. And a neural network is firstly utilized for abstracting all of the physical calculation process as a function and builds up a multilayer feedforward neural network model between input and output. However, in such networks, the choice of hidden layer node count affects network performance. In order to achieve both high prediction accuracy and low model complexity, Bayesian optimization finds the best solution in few iterations. While building the surrogate model, an initial data set is formed, and a gaussian process is utilized for the surrogate model which represents the distribution of the objective function. Updating the acquisition function, the best fit parameters, and For the HEMP waveform function, models for Emax, \begin{document}$ \alpha $\end{document}, \begin{document}$ \beta $\end{document}, k, and t0 are sequentially established to predict waveform parameters under different conditions. By calculating the waveforms along the north south axis and incorporating the angle between the burst point projection and the observation point, the HEMP waveform parameters at any location are further derived and computed.  Results and Discussions  The above method uses both numerical calculation and Bayesian optimization based neural network algorithm, in order to build up an artificial intelligence model of predicting the ground HEMP waveform parameters, covering arbitrary height of bursts, yield of gammas, position in a certain range. Bayesian optimization neural network method reduces searching times on the number of nodes in hidden layers and improve the predicting precision. In Bayesian optimization process, the searching space of hidden layer as [8, 50] is defined to avoid model over fitting, and gives a best network structure of quite small number of hidden layer nodes. To control algorithm training time, the maximum number of Bayesian evaluations is set to 20, approximately half the search space. Through Bayesian optimization, the optimal hidden layer node counts for Emax, \begin{document}$ \alpha $\end{document}, \begin{document}$ \beta $\end{document}, k, and t0 are found to be 19, 16, 11, 13, and 10, respectively. The simulation results indicate that the error of waveform parameter predicted by this method compared to the calculation result of simulation is lower than 3.12%, which has a better performance in comparison with several other methods. It works best on each metric. Experimental comparisons show the stability and generalization ability of the proposed algorithm for predicting different parameters. Analysis shows that Bayesian optimization, using a probabilistic surrogate model and an acquisition function, defines a probabilistic map between the number of nodes and the error on the validation set, to guide the following sampling steps with uncertainty estimation on predictions, which reduces the original calculation time from hours to seconds and the time complexity from O(n5) to O(n2), supporting large scale real time computation for the parameters in a HEMP waveform, given different scenarios.  Conclusions  This paper proposes a Bayesian optimization based multilayer feedforward neural network method to model the simulation computation process of HEMP ground waveforms, enabling rapid calculation of standard waveform functions for all points within the ground field distribution of HEMP over a certain range. The method uses Bayesian optimization to optimize the number of hidden layer nodes in the multilayer feedforward neural network, reducing search iterations and improving model prediction accuracy. Compared with five other artificial intelligence methods, the FBEMP method performs best in terms of all metrics. By combining the neural network with numerical derivation, all waveform functions within the field distribution coverage area can be calculated for different burst heights, gamma yields, longitudes and latitudes. This approach lowers the order of operation count from O(n5) in conventional numerical computation to O(n2), and reduces the calculation time of waveform functions for a given field distribution from hour scale to second scale, with errors on each parameter less than 3.12%. It realizes real time calculation of waveform function in HEMP field distribution, and had become applied to large-scale real-time HEMP waveform function calculations, solving the longstanding technical challenge of time-consuming HEMP environment computations that previously prevented real-time calculation, thereby providing an online computational environmental foundation for digital simulation and assessment in HEMP experiments.
Design and Performance Evaluation of Ultrasonic Nebulization Glow Discharge Detector for High-sensitivity and Rapid Detection of Metal Elements in Liquids
DING Yu, XU Jianan, PANG Maoyuan, WANG Yuhang, LI Jinyi, YU Weiye, HE Yihua, LI Xiangchu, TAN Qiang, LIU Xinxin, ZHOU Wangping
Available online  , doi: 10.11999/JEIT260436
Abstract:
  Objective  Water is essential for all organisms and ecological systems, and the composition and content of dissolved metal elements, especially copper (Cu), sodium (Na), and potassium (K), are crucial for maintaining ecological balance and biological health. Cu is an essential human trace element that forms enzymes with functional proteins, participating in antioxidant, energy supply, and immune processes. However, excessive Cu in water—from industrial wastewater, feed additives, pipeline corrosion, and electronic waste leakage—harms human health, causing vomiting, hypotension, jaundice, and hemolytic anemia. Aquatic organisms and plants, lacking effective detoxification systems, are more vulnerable to Cu pollution, which damages roots, induces oxidative stress, and inhibits growth. Na and K are vital for nerve conduction, muscle movement, and fluid balance, but excessive intake endangers those with hypertension or renal/cardiac insufficiency, and their imbalance in irrigation water causes soil salinization. Conventional detection methods (ICP-MS/AES, AAS, AFS) have excellent sensitivity but are limited by large size, complex pretreatment, high cost, and professional operation. Atmospheric pressure glow discharge (APGD) shows potential for miniaturization, but existing APGD-based technologies (PN-APGD, SCGD) require expensive equipment or complex pretreatment. Thus, a highly sensitive, rapid detector for on-site real-time detection of Cu, Na, K without additional pretreatment or driving equipment is urgently needed.  Methods  A Ultrasonic Nebulization Glow Discharge (UNGD) detector was designed, consisting of an ultrasonic nebulization unit, a plasma excitation unit, and a spectral signal collection unit. The ultrasonic nebulization unit adopted a self-designed centrifuge tube-based diversion chamber with a detachable microporous atomizing sheet, argon inlet/outlet, and a space-constrained transfer tube to form a short gas path, reducing aerosol loss. The plasma excitation unit used 180° coaxial tungsten needle electrodes (cathode/anode) with a double-layer fixing sleeve, clamped on a 3D platform for precise spacing adjustment, powered by a high-voltage DC power supply with a 20 kΩ ballast resistor. The spectral unit included an optical fiber probe and a three-channel AvaSpec spectrometer for high-resolution signal collection. Performance was evaluated using Cu/Na/K mixed solutions: single-factor experiments optimized parameters; characteristic spectra for qualitative analysis; recovery rates for anti-interference assessment; standard curves, LOD, RSD, and CRM detection for quantitative verification; and comparison with similar technologies.  Results and Discussions  Qualitative analysis of a water sample (Cu: 5.28 mg/L, Na: 5.23 mg/L, K: 3.89 mg/L) showed OH (281.1~309.0 nm) and N2 (315.0~406.0 nm) molecular bands, with obvious Cu (324.7 nm), Na (589.0 nm), and K (766.5 nm) characteristic peaks (Fig. 2); 324.7 nm was selected as Cu’s analytical line. Parameter optimization determined optimal conditions: discharge current 38 mA (Fig. 3), argon flow rate 0.5 L/min (Fig. 4), electrode spacing 1 mm (Fig. 5), sampling distance 22 mm (Fig. 6). Under these conditions, Cu, Na, K showed good linearity (R2: 0.9942, 0.9968, 0.9973), with LODs of 141.29 μg/L, 9.54 μg/L, 12.05 μg/L, and RSDs of 6.8%, 6.1%, 5.5% (n=11) (Table 1). Anti-interference tests showed 90%~110% recovery rates with 500 mg/L interfering cations (Fig. 7). CRM detection showed 89%~110% recovery rates, consistent with standard values (Table 2).  Conclusions  The UNGD detector achieves accurate quantitative analysis of Cu, Na, K in liquids without additional pretreatment or driving equipment. Its detachable atomizing sheet and short gas path improve sampling efficiency, while coaxial electrodes concentrate excitation energy. With excellent sensitivity, precision, and anti-interference ability, it provides reliable technical support for on-site real-time monitoring of water metal elements and has potential for extending to other metal detections, contributing to water ecological protection and biological health.
Patch-Sinusoidally Modulated SSPPs Leaky-Wave Antenna and Its Random Forest-Assisted Optimization Design
TANG Luping, CHENG Yonghao, CHEN Yibo, LIAO Chen
Available online  , doi: 10.11999/JEIT260651
Abstract:
  Objective  Spoof surface plasmon polaritons (SSPPs) leaky-wave antennas feature low-profile configuration and inherent frequency-scanning capability, making them promising for modern radar, communication, and intelligent sensing systems. However, strong nonlinear coupling among geometric parameters makes traditional optimization computationally costly, as full-wave simulations require thousands of evaluations and often converge to suboptimal local solutions due to landscape complexity. The leakage dynamics in SSPPs—slow-wave propagation, spatial harmonic coupling, and leakage rate distribution—adds complexity beyond conventional designs. To address this, we propose a patch-sinusoidally modulated SSPPs leaky-wave antenna and a machine learning framework integrating a random forest surrogate with particle swarm optimization (PSO) for efficient high-dimensional global optimization with reduced cost.  Methods  Unlike conventional groove-depth modulation, our antenna maintains uniform groove depth and a complete metal ground. Patch arrays with sinusoidally varying widths are loaded on both sides of the transmission line, with envelope functions \begin{document}$ Y=A\sin (Tx) $\end{document} and \begin{document}$ Y=A\sin (Tx+\pi ) $\end{document}, enabling flexible leakage control. The antenna is fully described by a nine-dimensional continuous parameter vector: groove width g, depth s, period p, six transition lengths g1–g6, port dimensions l and w, and modulation A, T. Using Latin hypercube sampling, 450 parameter samples are generated to ensure uniform coverage. Full-wave frequency-domain simulations (COMSOL, 9 GHz, approximately 27 min each) extract gain, S11, S21, scanning angle, side lobe level (SLL), and total efficiency. Four regression models—MLP, SVR, random forest (RF), and GPR—are systematically trained. RF employs bootstrap resampling with hyperparameters optimized via random search cross-validation (trees: 500–1200, depth: 12–24, min samples per split: 2–5). The trained RF surrogate is then embedded into PSO (40 particles, 120 iterations, inertia 0.72, c1=c2=1.5) with a weighted fitness function (G:1.5, η:0.7, SLL:0.55, S11:0.25, S21:0.7, θscan:0.18).  Results and Discussions  RF achieves the highest average R2 of 0.9554 across six outputs, outperforming GPR (0.9489), SVR (0.9212), and MLP (0.8952). For key radiation indicators, RF attains gain MAE of 0.032 dBi, SLL MAE of 0.168 dB, and efficiency MAE of 0.015. Scatter plots of predicted versus simulated values cluster tightly around the diagonal, and residual histograms show means near zero with no systematic bias, confirming excellent prediction accuracy and generalization. After RF-PSO optimization, full-wave simulation confirms substantial improvements: gain rises from 13.84 dBi to 14.52 dBi, SLL drops from –17 dB to –19 dB,peak total efficiency increases from 84.9% to 92.9%, S11 improves from –23.26 dB to –27.92 dB (4.66 dB), and scanning range expands from 57.2° to 61.5°. The scanning angle versus frequency curve exhibits good linearity across the operating band, and the two-dimensional far-field patterns show improved symmetry. The decrease in S21 (from –3.76 dB to –5.50 dB) together with the gain/efficiency increase indicates that more energy is effectively converted into radiation rather than being dissipated or reflected. Sensitivity analysis with ±2% perturbations (50 samples) shows all coefficients of variation (CV) below 1.3%: gain CV 0.21% (<±0.1 dBi), SLL CV 1.26% (±0.4 dB), efficiency CV 0.88% (±0.01), S11 CV 0.97%, S21 CV 0.91%, scan CV 0.37%. These fluctuations are far smaller than optimization gains, confirming excellent robustness under typical fabrication tolerances. Comparison with recent leaky-wave antennas (both SSPP-based and SIW) demonstrates superior SLL (−19 dB), competitive efficiency (89% vs. 94.95% and90% in prior SSPP work), and scanning range (61.5°) outperforming most single-port SSPP antennas (e.g., 20°, 16°, 33°, 13°). The number of full-wave simulations is reduced by approximately 90% (450 training + 1 validation vs. 4,800 simulations for conventional PSO).  Conclusions  This paper proposes a patch-sinusoidally modulated SSPPs leaky-wave antenna and an RF-assisted PSO framework for synergistic optimization in nine dimensions. The RF surrogate achieves an average R2 of 0.9554. The optimized antenna shows significantly improved performance across all metrics: gain by 0.68 dBi, SLL by 2 dB, efficiency by 8%, S11 by 4.66 dB, and scanning range by 4.29°, while maintaining compact dimensions. Sensitivity analysis confirms robustness under typical fabrication tolerances. The proposed methodology reduces the number of full-wave simulations by approximately 90%. This work marks a methodological advancement by introducing machine learning surrogate modeling into SSPPs leaky-wave antenna design for efficient high-dimensional optimization. Future work includes fabrication, experimental validation, extension to millimeter-wave bands, and reconfigurable antenna designs.
Physical-layer Network Coding Aided Polar Slotted Random Access Algorithm
SHAO Caiping, QIU Yuping, XIE Zhaopeng, SONG Dan, CHEN Jian, CHEN Pingping
Available online  , doi: 10.11999/JEIT260867
Abstract:
  Objective  Massive Machine-Type Communications (mMTC) constitutes a fundamental pillar of 5G and emerging 6G wireless networks, dedicated to supporting massive connectivity for the Internet of Things (IoT). In grant-free random access scenarios, sporadic and uncoordinated transmissions by massive terminals inevitably induce severe packet collisions under heavy traffic loads. Traditional scheduled access protocols become inefficient due to prohibitive signaling overhead. Consequently, grant-free slotted ALOHA protocols based on Successive Interference Cancellation (SIC)—such as Contention Resolution Diversity Slotted ALOHA (CRDSA), Irregular Repetition Slotted ALOHA (IRSA), and Coded Slotted ALOHA (CSA)—have garnered widespread attention. However, these classical schemes fundamentally rely on the presence of collision-free (degree-1) slots to trigger and sustain the iterative graph-peeling decoding process. Under practical constraints of finite frame lengths and heavy traffic, collision-free slots are drastically depleted, causing severe decoding stalling and throughput degradation. Furthermore, pure SIC mechanisms are intrinsically vulnerable to error propagation. Although Polar Slotted ALOHA (PSA) introduces polarization transforms across time slots to enhance packet recovery, its collision resolution remains constrained by the initial SIC condition. To resolve these challenges, a joint decoding algorithm combining Polar Slotted ALOHA with Physical-Layer Network Coding (PSA-PNC) is proposed over the Slot Erasure Channel (SEC). The objective is to transform destructive multi-user packet collisions into algebraically solvable linear network coding equations, thereby eliminating the strict reliance on collision-free slots and significantly elevating concurrent multi-user detection capability and throughput performance under heavy traffic loads.  Methods  A joint physical-layer and MAC-layer random access framework is established over the Slot Erasure Channel (Fig. 1). At the transmitter side, active users independently select transmission time slots based on an irregular degree distribution polynomial without inter-user coordination or channel collision feedback. The sender remains blind to multi-user collision patterns in the channel. At the base station receiver, the superimposed signals across slots are equivalently modeled as a sparse global input matrix over a binary finite field, followed by packet-level polar encoding. To resolve dense collisions without relying on clean slots, a closed-loop iterative receiver architecture is developed (Fig. 2). In each iteration, channel observation sequences are initially processed by a packet-level Successive Cancellation (pSC) or packet-level Successive Cancellation List (pSCL) decoder to extract reliable equivalent combined packets from information slots. Instead of being discarded, the collided slots corresponding to these reliable packets are utilized to construct a local sparse linear Network Coding (NC) equation system. A Generalized Matrix Inversion (GMI) criterion is subsequently executed to analyze the column-rank characteristics of the access pattern matrix and achieve global algebraic multi-user decoupling. Solvable user packets are directly recovered without requiring full-rank matrix conditions or degree-1 slots. The algebraically decoupled user packets are then utilized as prior information to reconstruct physical-layer codewords and subtracted from the observation buffer via iterative SIC, continuously reducing the dimensionality of the unresolved collision space. High-reliability packets output by the pSCL decoding paths are leveraged in the iterative loop to effectively suppress error propagation and guarantee the linear independence of residual equations. Furthermore, the polarization evolution process of equivalent multi-user packets over the SEC is theoretically proved to be equivalent to that over a scalar Binary Erasure Channel (BEC), enabling rigorous calculation of frame error rate bounds via Bhattacharyya parameters (Fig. 3).  Results and Discussions  Extensive theoretical analyses and Monte Carlo simulations are conducted to evaluate the performance of the proposed PSA-PNC scheme over the Slot Erasure Channel. The theoretical polarization bounds of packet-level polar decoding are verified under various slot erasure probabilities (\begin{document}$ \epsilon \in \left\{0.1,0.2,0.3\right\} $\end{document}), exhibiting precise consistency with simulation curves and validating the polarization threshold effect (Fig. 3). In terms of system throughput, simulation results demonstrate that for a frame length of \begin{document}$ N=1024 $\end{document}, the proposed PSA-PNC scheme with pSCL (\begin{document}$ L=8 $\end{document}) achieves a peak normalized throughput of approximately 0.87 packets/slot at a normalized load of \begin{document}$ G\approx 0.90 $\end{document}, yielding an approximate 15% throughput improvement over baseline PSA and outperforming Coded Slotted ALOHA under identical finite-length configurations (Fig. 4(a)). In the low-load region, all evaluated schemes exhibit near-identical linear throughput growth due to the abundance of collision-free slots (Fig. 4(a)). When the slot erasure rate increases to \begin{document}$ \epsilon =0.35 $\end{document}, the number of recoverable reliable equivalent packets decreases, leading to insufficient NC equations and observable throughput degradation in high-load regions, which confirms the operational boundary of the algorithm (Fig. 4(a)). For a short frame length of \begin{document}$ N=64 $\end{document}, a consistent throughput gain ranging from 0.12 to 0.18 is maintained by PSA-PNC, demonstrating strong robustness against finite-length decoding stalling in short-packet scenarios (Fig. 4(b)). In terms of transmission reliability, under a target Packet Loss Rate (PLR) of \begin{document}$ {10}^{-2} $\end{document} at \begin{document}$ N=1024 $\end{document}, the supportable normalized load upper bound is extended from \begin{document}$ G\approx 0.75 $\end{document} in baseline PSA to \begin{document}$ G\approx 0.84 $\end{document} in PSA-PNC (Fig. 5(a)). Furthermore, steeper waterfall regions and significantly lower error floors are consistently maintained across various frame lengths from \begin{document}$ N=64 $\end{document} to \begin{document}$ N=1024 $\end{document} (Fig. 5(b)).  Conclusions  A joint decoding scheme combining Polar Slotted ALOHA with Physical-Layer Network Coding (PSA-PNC) is established to resolve the severe decoding stalling problem in grant-free random access. By constructing a closed-loop iterative receiver integrating packet-level polar decoding, GMI-based algebraic equation solving, and iterative SIC cancellation, destructive multi-user collisions are converted into solvable linear equations. The dependence on collision-free slots is effectively eliminated, and the supportable load threshold, peak normalized throughput, and packet recovery reliability are substantially enhanced under heavy traffic loads. Future research will be directed toward non-ideal multipath fading channels, such as Rayleigh fading, and the design of low-complexity sparse receiver architectures for practical massive access implementations.
A knowledge distillation framework for hypergraph neural networks with rapid inference capabilities
LI Junzheng, YU Hongtao, HUANG Ruiyang, JIANG Haocong, LIU Shuo, YANG Suchang
Available online  , doi: 10.11999/JEIT260694
Abstract:
  Objective  Hypergraph Neural Networks (HGNNs) have gained widespread attention for their strong ability to model high-order correlations among entities, but their computational complexity and memory consumption grow exponentially as the hypergraph scale expands, severely restricting their deployment in large-scale industrial scenarios. Existing knowledge distillation methods that distill HGNNs into Multi-Layer Perceptrons (MLPs) are troubled by poor interpretability, low accuracy, severe information loss caused by Softmax-based soft labels, and neglect of node reliability heterogeneity. To address these critical challenges, this paper proposes a novel hypergraph knowledge distillation framework named DH2KAN (Distill Hypergraph Neural Network to Kolmogorov-Arnold Network) for fast inference, which breaks through the bottlenecks of traditional hypergraph knowledge distillation.  Methods  This paper designs a three-module knowledge distillation framework DH2KAN(图2). Firstly, we replace the traditional MLP with KAN as the student model, which uses learnable spline-based univariate functions instead of fixed activation functions and linear weights to improve the fitting ability and interpretability of the student model. Secondly, we propose a representation similarity distillation mechanism, which directly aligns the pre-logits representations of HGNN and KAN to avoid information loss caused by the Softmax normalization layer and completely retain the high-order structural knowledge of hypergraphs. Thirdly, we introduce a high-reliable node-aware distillation method(图3), which quantifies the node reliability by information entropy variation, screens out robust nodes with strong anti-noise ability, and takes their soft labels as the core supervision signal to improve the purity of distilled knowledge.  Results and Discussions  The DH2KAN algorithm achieves remarkable performance under both transductive learning (表2) and production learning (表3) settings. Quantitative experimental results reveal that DH2KAN obtains an accuracy improvement of approximately 10.53% over the vanilla KAN student model, 1.73% over the teacher HGNN model, and 1.2% over existing MLP-based distillation methods. Such results verify the effectiveness of knowledge transfer from HGNN to KAN, and demonstrate that the proposed method outperforms conventional MLP-oriented distillation schemes in inference performance. In addition, DH2KAN achieves optimal performance on feature-dominated hypergraph datasets, and possesses strong robustness when handling sparse structures and noisy node samples.  Conclusions  This paper proposes DH2KAN to accelerate HGNN inference for large-scale low-latency applications. Via knowledge distillation, it bridges the performance gap between KAN and HGNN and eliminates structural dependence for efficient reasoning. With representation similarity and reliable node-aware distillation, it transfers effective task knowledge via pre-logit features and soft labels, showing great practical application potential.
Research on Covert Communication Transmission Scheme Combining Relay Selection and Mode Selection over Nakagami-m Fading Channels
HUANG Haiyan, HUANG Yi, ZHANG Ning, LIANG Linlin, ZHANG Xuejun
Available online  , doi: 10.11999/JEIT260287
Abstract:
  Objective  Covert communication enhances the security of wireless communication systems by concealing both transmitted information and communication activities from unauthorized detection. However, practical wireless channels exhibit random and uncertain propagation conditions. The Nakagami-m fading channel, which can characterize a wide range of channel conditions, provides a realistic framework for evaluating the performance of covert communication. Relay-assisted transmission has attracted considerable attention because it improves transmission reliability over fading channels. Moreover, relay selection and transmission mode selection substantially affect system performance. Therefore, investigating their combined effect on covert communication over Nakagami-m fading channels is of both theoretical and practical significance for the design of next-generation secure wireless communication systems.  Methods  This paper proposes a covert communication system incorporating relay selection and transmission mode selection. The source node transmits covert information to the destination node through multiple relays, while a warden monitors transmissions from both the source and relay nodes. A friendly jammer transmits interference signals to degrade the warden’s detection capability. Four transmission schemes are considered: optimal relay selection with fixed Half-Duplex (HD) or Full-Duplex (FD) operation, optimal relay selection with random transmission mode selection, random relay selection with optimal transmission mode selection, and joint optimal relay and transmission mode selection. Closed-form expressions for the warden’s detection error probability under both HD and FD optimal relay selection are derived over Nakagami-m fading channels. Closed-form expressions for the transmission outage probability, asymptotic transmission outage probability, and covert rate are also derived for all transmission schemes. The theoretical analysis is validated through MATLAB simulations.  Results and Discussions  Simulation results demonstrate that an optimal detection threshold exists that minimizes the detection error probability (Fig. 2). As the detection threshold or jamming power increases, the warden’s ability to detect covert communication decreases, causing the detection error probability to approach one (Figs. 2 and 3). Under the same target transmission rate and high Signal-to-Noise Ratio (SNR) conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability, thereby providing the highest transmission reliability (Figs. 4 and 5). At a target transmission rate of \begin{document}$ \text{6.5 bit/(s}\cdot \text{Hz)} $\end{document}, the transmission outage probability of the joint relay and transmission mode selection scheme is 6.9% lower than that of the FD transmission scheme (Fig. 4). As the transmit power and the number of relays increase, the covert rate gradually approaches a constant value. Among all transmission schemes, the joint relay and transmission mode selection scheme consistently achieves the highest covert rate (Figs. 6 and 7).  Conclusions  This paper proposes a covert communication system based on relay selection and transmission mode selection over Nakagami-m fading channels. Closed-form expressions for the warden’s detection error probability and the system’s transmission outage probability are derived under different relay selection and transmission mode selection strategies. The asymptotic transmission outage probability and covert rate are then analyzed. Simulation results show that increasing the detection threshold or jamming power weakens the warden’s ability to detect covert communication, causing the detection error probability to approach one. Under identical target transmission rates and high SNR conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability. These results indicate that appropriate relay selection and transmission mode selection not only reduce the warden’s detection capability and protect covert communication, but also improve both transmission reliability and covertness. Future work will consider practical factors, including imperfect channel state information, residual self-interference, and incomplete knowledge of the warden’s channel.
Research on Channel Multipath Prediction Based on an Environmental Graph
ZHANG Zhaoling, JIN Jing, ZHAO Jingbo, YU Li, CAI Yichen, MA Liang, ZHANG Jianhua
Available online  , doi: 10.11999/JEIT260416
Abstract:
  Objective  Environment-driven channel prediction requires a structured representation that connects physical objects in a propagation environment with the resulting multipath topology. However, conventional data-driven methods generally treat environmental information as unstructured global features and have difficulty representing the interactions among the transmitter (Tx), receiver (Rx), and surrounding scatterers. This study investigates whether an environmental graph can provide an effective intermediate representation for identifying effective scatterers and predicting candidate propagation paths.  Methods  An environmental graph is constructed by representing the Tx, Rx, and scatterers as graph nodes. Spatial distance and visibility between nodes are encoded as edge features to describe their geometric relationships. A ScatterGNN based on the Edge-aware Graph Isomorphism Network (EGIN) is developed to extract structural features and identify effective scatterers involved in signal propagation. Candidate single- and two-bounce paths are subsequently generated from the detected scatterers. A path-ranking network, PathRankingNet, is then designed to estimate the validity scores of candidate paths and rank them using a Listwise Ranking Loss. The proposed framework is evaluated in a controlled indoor Industrial Internet of Things scenario generated using Wireless InSite. The scenario covers an area of 150 m × 63 m × 22 m and contains 3,381 Rx sampling locations.  Results and Discussions  For effective scatterer detection, the proposed method achieves an average precision of 0.957 9, while 77.1% of the test samples achieve complete detection of all effective scatterers. For candidate propagation-path prediction, the model converges stably after approximately 30 training epochs. Precision@3 ranges from 0.60 to 0.64, Recall@3 ranges from 0.42 to 0.45, and Hit@3 ranges from 0.90 to 0.92. These results indicate that the environmental graph preserves useful structural information related to propagation-path topology. In particular, the model retains at least one ground-truth propagation path among the three highest-ranked candidates for more than 90% of the test samples.  Conclusions  The proposed framework provides a graph-based approach for transforming environmental geometry into structured representations of effective scatterers and candidate propagation paths. Rather than replacing ray tracing or channel measurements, the method is intended to reduce the candidate search space before detailed path-parameter calculation or channel reconstruction. The current results demonstrate its feasibility within a single simulated environment and primarily reflect its ability to approximate the propagation-path topology labels generated by Wireless InSite. Further validation using independent environments, measured channel data, and path-level parameters such as power, delay, and complex gain is required before its cross-scenario generalization and practical applicability can be established.
An Incremental Density-Based Clustering Method with TDOA Prior for Mobile Multi-Station Radar Signal Sorting
CHEN Jinli, FAN Yu, WANG Yanjie, ZHANG Jindong
Available online  , doi: 10.11999/JEIT260151
Abstract:
  Objective  Radar signal sorting is a core technology in electronic reconnaissance that aims to deinterleave and classify pulses from multiple radar emitters within dense, overlapping pulse streams. In practical reconnaissance missions, especially those employing mobile platforms such as aircraft, observation stations continuously change position, whereas radar emitters are typically stationary. The resulting variation in the relative geometry causes the Time Difference of Arrival (TDOA) of intercepted signals to evolve over time. Conventional multi-station radar signal sorting methods are generally developed under a static TDOA assumption and cannot effectively characterize this temporal evolution. Therefore, pulses from the same radar emitter are easily split into multiple clusters, resulting in cluster proliferation and degraded sorting performance. To address this problem, an Incremental Density-Based Clustering (ICDC) method with TDOA prior information is proposed for mobile multi-station radar signal sorting. The proposed method exploits the temporal evolution of TDOA to improve sorting stability and accuracy in mobile multi-station scenarios.  Methods  The temporal evolution of TDOA in mobile multi-station scenarios is first analyzed, and the TDOA trajectory is approximated as a linear function of time to establish a linear prior state model. Online micro-clusters are then constructed from multi-station TDOA observations and aggregated into macro-clusters according to their spatiotemporal intersection relationships. For each macro-cluster, an independent Kalman Filter (KF) model is established for each TDOA dimension. The state vector consists of the TDOA value and its rate of change, and recursive state estimation is performed to provide dynamic TDOA priors for newly arriving samples. During incremental clustering, a spatiotemporal joint scoring function is developed by incorporating KF prediction residuals as dynamic constraints. The matching criterion therefore evolves from a conventional density-based rule into a joint spatiotemporal consistency criterion, enabling more accurate assignment of newly arriving TDOA observations. To suppress cluster proliferation caused by TDOA evolution, a concept drift detection strategy based on macro-cluster center evolution is further employed. When concept drift is detected, posterior state estimates and covariance matrices generated by the KF are used to construct a dual Mahalanobis distance criterion that jointly evaluates state-distribution overlap and predicted TDOA overlap. Radar clusters produced by erroneous splitting are then adaptively merged under a 95% confidence threshold, effectively suppressing cluster proliferation caused by concept drift.  Results and Discussions  Simulation data are generated according to the radar parameters listed in Table 1. The multi-station TDOA corresponding to the same radar emitter exhibits an approximately linear evolution with respect to the Time of Arrival (TOA) at the observation station (Fig. 5), consistent with the proposed linear prior model. The ability of the proposed method to suppress radar cluster proliferation is first evaluated (Fig. 6). Radar emitters E4 and E9 are selected from the nine simulated emitters as representative cases because their TDOA trajectories are closely spaced and difficult to separate. Compared with Density-Based Spatial Clustering of Applications with Noise (DBSCAN), ICDC, cloud model-based sorting, and PointNet++ sorting, the proposed method more effectively suppresses radar cluster proliferation and maintains greater cluster stability. The overall sorting performance is further compared with the histogram method, grid-based clustering, DBSCAN, ICDC, cloud model-based sorting, and PointNet++ sorting. When the TOA measurement error ranges from 50 to 300 ns (Fig. 7), the proposed method consistently achieves a sorting accuracy above 96%, demonstrating strong robustness to measurement errors. Under different pulse interference rates (Fig. 8), the sorting accuracy also remains above 96%, indicating excellent interference robustness. The performance under different observation station velocities is further evaluated (Fig. 9). The proposed method maintains high sorting accuracy over the entire velocity range and still achieves approximately 94% accuracy at relatively high observation station velocities, demonstrating strong robustness under dynamic observation conditions. Radar cluster proliferation probability, missed-cluster probability (Fig. 10), and computational complexity (Table 2) are also analyzed. The results demonstrate that the proposed method achieves a favorable balance among sorting accuracy, cluster proliferation suppression, missed-cluster control, and computational complexity.  Conclusions  An ICDC method with TDOA prior information is proposed to address the degradation of radar signal sorting performance caused by TDOA evolution in mobile multi-station scenarios. By incorporating observation station motion into a dynamic TDOA state model and applying KF-based recursive prediction, the proposed method effectively suppresses radar cluster proliferation and erroneous cluster splitting caused by concept drift. The spatiotemporal joint criterion and the adaptive cluster merging strategy further improve robustness in complex dynamic environments. Simulation results demonstrate the effectiveness and stability of the proposed method for mobile multi-station cooperative reconnaissance. Future work will focus on real-time multi-parameter fusion-based sorting in complex electromagnetic environments. Furthermore, adaptive estimation and adjustment of the micro-cluster spatial intersection threshold, process noise covariance matrix, and measurement noise variance will be investigated to further improve performance in complex scenarios.
Difference-aware Adaptive Prompt Learning and Dense Alignment for Weakly Supervised Building Change Detection
CHEN Yanxia, MA Longlong, CHEN Yanhua, HUANG Yuchun
Available online  , doi: 10.11999/JEIT260595
Abstract:
  Objective  Building change detection from bi-temporal high-resolution remote sensing images is important for urban planning, land resource management, illegal construction monitoring, and disaster damage assessment. Existing fully supervised change detection methods generally achieve high detection accuracy but require pixel-level annotations. However, obtaining pixel-level labels for large-scale remote sensing images is labor-intensive and time-consuming, which limits their application to large-scale monitoring scenarios. Image-level weakly supervised change detection reduces annotation costs by using only image-level labels indicating whether an image pair contains changes. However, the lack of spatial supervision makes accurate localization of changed regions difficult. Existing weakly supervised methods generally rely on Class Activation Maps (CAMs) to generate pseudo labels. CAMs tend to highlight only the most discriminative regions, resulting in incomplete coverage of changed areas or background noise. Vision-language models provide semantic priors for weakly supervised learning. However, directly applying Contrastive Language-Image Pre-training (CLIP) to change detection remains challenging. Fixed text prompts are difficult to adapt to the difference semantics of bi-temporal images, and the original CLIP objective mainly focuses on global image-text alignment rather than local pixel-level localization. To address these problems, a Difference-aware Adaptive Prompt Learning and Dense Alignment method for weakly supervised building change detection, termed DAPL-CD, is proposed.  Methods  The proposed framework introduces CLIP-based cross-modal semantic knowledge into image-level weakly supervised building change detection. For a pair of bi-temporal remote sensing images, a shared CLIP visual encoder is first used to extract visual representations from the two temporal images. The local visual features are fused along the channel dimension to obtain bi-temporal difference features containing semantic information related to changed and unchanged regions. Based on the difference characteristics of building change detection, a difference-aware adaptive prompt learning strategy is designed. Instead of using manually designed fixed text templates, learnable context vectors are inserted into the text prompts while preserving category-related semantic words. The resulting foreground and background text embeddings are used as foreground and background text prototypes to provide adaptive semantic guidance for change localization. Furthermore, a pixel-text dense alignment mechanism is introduced to extend CLIP’s global image-text alignment capability to local feature matching. The initial CAM generated by the classification branch is used to obtain preliminary foreground and background regions. Visual-text positive and negative sample pairs are then constructed between local difference features and the foreground and background text prototypes. An InfoNCE-based dense alignment loss is used to pull matched visual and textual features closer and push mismatched features apart. Finally, the classification and segmentation branches are jointly optimized using the classification loss, global alignment loss, dense alignment loss, and segmentation loss, with the segmentation loss introduced only after the quality of the generated CAMs has stabilized.  Results and Discussions  Experiments are conducted on WHU-CD and LEVIR-CD, two public benchmark datasets for building change detection. Only image-level labels are used during training, whereas pixel-level annotations are used only for evaluation. Overall Accuracy (OA), F1-score, and Intersection over Union (IoU) are adopted as the main evaluation metrics. Because changed buildings usually occupy a small proportion of remote sensing images, OA can be strongly affected by the dominant unchanged background pixels. Therefore, F1-score and IoU are emphasized for evaluating the detection quality of changed regions. Quantitative comparisons show that DAPL-CD achieves an OA of 94.7%, an F1-score of 82.8%, and an IoU of 70.6% on WHU-CD, and an OA of 92.3%, an F1-score of 68.0%, and an IoU of 51.5% on LEVIR-CD. The method achieves the best F1-score and IoU among the compared weakly supervised change detection methods (Table 1). Visual comparisons further show that the proposed method produces more complete responses for large-scale building changes and more continuous predictions for small and scattered changed buildings (Figs. 3 and 4). Ablation experiments verify the effectiveness of difference-aware adaptive prompt learning and pixel-text dense alignment. The baseline model using fixed text prompts without foreground or background alignment achieves an F1-score of 63.2% and an IoU of 46.2%. Introducing both foreground and background alignment increases these metrics to 66.3% and 49.6%, respectively, indicating that dense semantic matching between local visual features and text prototypes improves the discrimination of changed regions. After difference-aware adaptive prompt learning is incorporated, the F1-score and IoU increase to 68.0% and 51.5%, respectively, indicating that learnable context vectors reduce the semantic mismatch between fixed text descriptions and bi-temporal difference features. Under the adaptive-prompt setting, foreground alignment alone achieves an F1-score of 64.4% and an IoU of 47.5%, whereas background alignment alone achieves 65.3% and 48.5%, respectively. Combining the two alignment branches yields the best F1-score and IoU, indicating that foreground and background semantic constraints provide complementary guidance for change localization (Table 2). The CAM results further show that pixel-text dense alignment produces stronger and more complete responses over actual changed regions while suppressing irrelevant background activations (Fig. 5).  Conclusions  A weakly supervised building change detection framework based on difference-aware adaptive prompt learning and pixel-text dense alignment is proposed. By introducing CLIP-based cross-modal semantic priors, the proposed method converts text-level semantic knowledge into local change localization capability. The difference-aware adaptive prompt learning strategy improves the representation of change-related semantic descriptions, whereas the pixel-text dense alignment mechanism establishes direct correspondence between local difference features and foreground and background text prototypes. Experimental results on WHU-CD and LEVIR-CD demonstrate that DAPL-CD achieves high performance under image-level supervision and improves the completeness and accuracy of changed building localization. The proposed framework provides an effective approach for reducing annotation requirements in large-scale remote sensing change detection. Future research will focus on improving pseudo-label reliability, reducing dependence on large-scale pre-trained models, and extending the method to multi-temporal and multi-spectral remote sensing data.
A Multi-Dimensional Scenario-Based Evaluation Method for Deep Learning Side-Channel Analysis Using a Multi-Attribute Decision Model
GU Zepeng, CHEN Lin, CAI Juesong, YAN Yingjian
Available online  , doi: 10.11999/JEIT260198
Abstract:
  Objective  Deep Learning Side-Channel Analysis (DL-SCA) has substantially improved the effectiveness of attacks against protected cryptographic implementations. However, the transition of DL-SCA models from research to practical deployment is limited by the lack of systematic, fair, and scenario-specific evaluation methods. Existing evaluations mainly rely on Guessing Entropy (GE) and Success Rate (SR), while overlooking practical factors such as resource overhead and environmental adaptability. Moreover, inconsistent hyperparameter optimization leads to unfair model comparisons and provides limited quantitative guidance for model selection under different deployment constraints, including resource-constrained devices, high-noise environments, and real-time applications. This paper proposes a systems engineering-based evaluation framework that enables comprehensive, quantitative, and scenario-specific assessment of DL-SCA models.  Methods  A multi-dimensional, scenario-based evaluation framework is developed using systems engineering principles. First, a hierarchical evaluation index system is established, comprising three criteria—attack effectiveness, resource overhead, and environmental adaptability—and six evaluation metrics: GE, SR, training time (TC), peak memory consumption (MC), model complexity (MoC), and noise robustness (Rob). Second, a standardized evaluation process based on the V-model is designed to ensure fair comparison. Each candidate model, including a Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), and CNN-LSTM hybrid model, undergoes independent hyperparameter optimization using grid search before multi-dimensional performance evaluation. Third, a hybrid Criteria Importance Through Intercriteria Correlation-Analytic Hierarchy Process (CRITIC-AHP) Multi-Attribute Decision-Making (MADM) framework is developed. The CRITIC method derives objective weights from the statistical characteristics of the evaluation data, whereas the AHP method incorporates scenario-specific preferences through pairwise comparison matrices. The objective and subjective weights are fused to generate scenario-specific weights. Finally, a Multi-dimensional Attack Performance Metric (MAPM) is defined as the weighted sum of normalized evaluation metrics using the fused weights, providing a composite score for each model under a specific deployment scenario.  Results and Discussions  The proposed framework is validated using the ASCAD fixed-key dataset. After independent hyperparameter optimization, the three model architectures are evaluated using all six metrics. The CRITIC method produces the objective weight vector W critic = [0.17, 0.19, 0.15, 0.21, 0.14, 0.14]. Four representative deployment scenarios—Resource-Constrained, High-Performance, High-Noise, and Real-Time—are then defined, and the corresponding AHP preference weights are fused with the objective weights to generate the final scenario-specific weights. For example, MC receives the highest weight (0.52) in the Resource-Constrained scenario, whereas Rob dominates the High-Noise scenario with a weight of 0.57. The resulting MAPM scores (Table 9, Fig. 9, and Fig. 10) clearly differentiate the strengths of the evaluated models and demonstrate the scenario-specific decision capability of the proposed framework. CNN achieves the highest score in the High-Performance scenario (0.894), MLP ranks first in the Real-Time scenario (0.758) because of its shortest training time, and the CNN-LSTM hybrid model performs best in the High-Noise scenario (0.863) because of its superior noise robustness despite higher resource overhead. These results demonstrate that no single model is optimal across all deployment scenarios and that MAPM provides a clear and quantitative basis for model selection under specific deployment constraints.  Conclusions  This paper proposes a systems engineering-based, multi-dimensional evaluation framework to address the major limitations of current DL-SCA model assessment. By integrating a hierarchical evaluation index system, a standardized V-model evaluation process, and a hybrid CRITIC-AHP Multi-Attribute Decision-Making (MADM) framework, the proposed method quantitatively balances the trade-offs among attack effectiveness, resource overhead, and environmental adaptability. Experimental results obtained using the ASCAD benchmark demonstrate that the framework provides clear, quantitative, and scenario-specific guidance for model selection. The proposed Multi-dimensional Attack Performance Metric (MAPM) provides a practical decision basis for selecting DL-SCA models under diverse deployment constraints, narrowing the gap between academic attack development and practical model deployment. Future work will extend the framework to additional model architectures and datasets, improve evaluation automation, and validate its effectiveness in practical deployment environments.
Research on LEO Constellation Interference Prediction and Detection Algorithm Driven by Collaborative Spatial Feature Mapping and Temporal Transformer
YANG Boyu, QIU Kun, CHEN Zhe, ZHAO Jin, GAO Yue
Available online  , doi: 10.11999/JEIT260368
Abstract:
  Objective  The rapid deployment of large-scale Low Earth Orbit (LEO) satellite constellations has intensified competition for orbital and spectrum resources. To improve spectrum utilization, different LEO satellite constellations commonly employ co-frequency reuse, which increases the risk of inter-system co-frequency interference. The high-speed motion of LEO satellites and their highly dynamic, heterogeneous topology further cause rapid variations in interference power, resulting in severe co-frequency interference. Existing interference assessment methods mainly rely on the regulations of the International Telecommunication Union (ITU), with the Interference-to-Noise Ratio (I/Noise) as a key evaluation metric. However, conventional ITU-based physical-iteration methods require repeated calculations of satellite positions, link attenuation and interference contributions for all visible satellites, causing computational cost to increase rapidly with constellation size. A complete interference assessment for a single ground station in a constellation of approximately 10 000 satellites can require more than 30 h. Recent deep-learning-based methods still commonly use all visible satellites as inputs and therefore do not adequately exploit the spatial sparsity of interference features. To address these limitations, an LEO constellation interference prediction and detection method driven by collaborative spatial feature mapping and temporal Transformer is proposed to reduce computational cost while maintaining accurate interference prediction and detection.  Methods  A spatiotemporal framework is developed to model dynamic co-frequency interference in large-scale heterogeneous multi-constellation LEO networks. Based on the ITU physical interference model, three key properties of the interference function are identified: permutation invariance, spatial sparsity and temporal continuity. The conventional physical-iteration process is therefore reformulated as a spatiotemporally decoupled feature-mapping problem. A spatial feature mapping module with adaptive attention and symmetric pooling is constructed to compress unordered and variable-length satellite interference-source features into fixed-dimensional, permutation-invariant spatial features. The attention mechanism adaptively focuses on dominant interference sources, whereas symmetric pooling using maximum and mean pooling captures extreme and global statistical characteristics while suppressing redundant satellite nodes. A temporal Transformer is then employed to model the long-range evolution of interference trajectories. The future I/Noise trajectory is predicted to support rapid interference detection under dynamic constellation configurations.  Results and Discussions  A large-scale heterogeneous LEO constellation scenario consisting of 6 800 Starlink satellites and 650 OneWeb satellites is simulated according to ITU regulations. Real Two-Line Element (TLE) data are used to propagate satellite orbits, and 100,000 time-series samples are generated at 0.5 s intervals. Parameter analysis shows that appropriate feature dimensions and encoder depths provide a favorable balance between feature extraction accuracy and computational cost (Figs. 4 and 5). The sampling interval and sliding-window length are further optimized to balance prediction accuracy and real-time inference performance (Figs. 6 and 7). Compared with the baseline methods, the proposed method produces interference trajectories that closely follow the physical-iteration ground truth (Fig. 8). At a cumulative probability of 90%, the absolute prediction error is maintained within 0.5 dB (Fig. 9). The method also maintains a high interference recall under a low false-alarm-rate constraint (Fig. 10). For a 20.0 s prediction horizon, the Root Mean Square Error (RMSE) remains at 0.45 dB, substantially lower than those of the baseline models. At a 20 s prediction step, the RMSE values of the Multi-Layer Perceptron (MLP) and Long Short-Term Memory (LSTM) network increase to 5.53 dB and 4.82 dB, respectively (Fig. 11). In terms of computational efficiency, the proposed method requires 14.2 ms for a single inference, which is only 4.1% of the computational time required by the ITU-based physical-iteration method (Table 3). The computational cost is therefore effectively decoupled from constellation size.  Conclusions  To address the high computational cost caused by the highly dynamic topology of LEO satellite networks, a collaborative spatial feature mapping and temporal Transformer-based interference prediction and detection method is proposed. The spatial feature mapping module compresses variable-length interference-source features into fixed-dimensional representations, whereas the temporal Transformer captures the long-range evolution of interference trajectories. Simulation results show that the proposed method provides accurate long-term trajectory tracking while remaining consistent with ITU-based interference assessment. Parameter-sensitivity experiments demonstrate that the proposed method can balance feature extraction accuracy and computational cost under different configurations. With its long-range dependency modeling capability, the temporal Transformer maintains an RMSE of 0.45 dB over a 20 s prediction horizon. By filtering redundant satellite nodes and decoupling computational cost from constellation size, the method reduces single-inference time to 14.2 ms. Comprehensive evaluations of prediction error distributions, detection performance and model parameters demonstrate that the proposed method achieves high prediction accuracy and substantially reduced computational complexity, providing a flexible engineering solution for interference monitoring in large-scale LEO constellation deployments.
A Parametric Architecture Description Framework for Embedded FPGAs and Multi-objective QoR-driven Architecture Exploration
ZHOU Jing, ZHANG Shengbing, CHEN Lei, FENG Hanxu, WANG Shuo, TIAN Chunsheng
Available online  , doi: 10.11999/JEIT260609
Abstract:
  Objective  Architecture parameters of commercial off-the-shelf Field-Programmable Gate Arrays (FPGAs) are fixed by vendors and reused across products. Embedded FPGAs (eFPGAs), in contrast, allow architects to select architecture parameters according to specific application requirements. The LUT input count K, the number of LUTs per cluster N, the interconnect topology, and the types of heterogeneous tiles can therefore be configured for the target application. Architecture design space exploration thus becomes an engineering task in which tens to hundreds of architectures may need to be generated and evaluated before a suitable configuration is selected. Existing architecture description practices rely largely on manually written architecture files and batch scripts. After each parameter change, shared fields across multiple backend toolchains must be updated and aligned manually, making large-scale architecture exploration difficult to support. To address this problem, a parametric architecture description framework is proposed for unified description across multiple backend toolchains. Architecture parameters are organized into three layers according to their independence: process invariants, coupled parameters, and independent parameters. Architecture descriptions for different backends are automatically derived from the same source object through independent derivation functions. The framework currently supports VPR, OpenFPGA, and Yosys and has been extended to COFFE. Its operation is validated across the complete toolchain. Based on the framework, a parameter-sweep design space exploration method is developed, and an open Quality of Results (QoR) dataset covering five application domains and 64 benchmark circuits, including homogeneous and heterogeneous architectures, is released as a public benchmark.  Methods  HorizonArch, the proposed parametric architecture description framework, organizes architecture parameters into three layers according to parameter independence (Table 1). L0 contains process invariants that are fixed once the technology node is determined. L1 contains coupled parameters, including K, N, tier, segment length, switch block type, and channel connectivity, for which a single parameter change can trigger updates across multiple fields and backend architecture descriptions. L2 contains independent parameters that can be specified separately. Architecture construction is formalized by an operator B that maps a parameter vector p to a complete architecture object (Fig. 4). Five formal rules are imposed: parameter completeness (R1), fragment independence (R2), type compatibility (R3), explicit coupling (R4), and static checkability (R5). Each backend architecture description is then derived by an independent derivation function from the same source object. Thus, adding a new backend requires only an additional view rather than modifications throughout the existing description structure. Field-level validation and cross-field validation are performed when the architecture object is loaded, before any backend tool is invoked. The class structure (Fig. 3) divides the architecture description into synthesis, circuit, and layout views, with each semantic element declared only once. Three extension levels are defined: G1 adds a black-box model, G2 extends the value set of an existing coupled parameter, and G3 adds a new coupled parameter together with its constrained value set. Based on this framework, a parameter-sweep design space exploration method is developed to scan the (K, N) parameter grid and heterogeneous tile configurations. Each configuration is evaluated using three QoR metrics: area, Critical-Path Delay (CPD), and Area-Delay Product (ADP).  Results and Discussions  End-to-end validation shows that a single source description consistently generates architecture descriptions for VPR, OpenFPGA, and Yosys. COFFE is connected and verified at the interface layer, including SPICE simulation startup (Table 4, Table 5). The G1, G2, and G3 extension experiments pass all cross-field checks. The design space exploration results show different preferences among area, CPD, and ADP across the (K, N) parameter space (Figs. 5 and 6). Area favors smaller K values, with K=4 and N=4 providing favorable area and ADP performance for a large proportion of circuits. CPD, in contrast, favors larger K and N values, with optimal configurations concentrated near (K, N)=(8, 10) and (7, 10). Across the twenty (K, N) configurations, the relative-range distribution shows that parameter selection has a much greater effect on area and ADP than on CPD (Table 6). The mean relative ranges are 50.9% for CPD, 675.1% for area, and 516.3% for ADP. A comparison of default configurations (Table 7) shows that K=4 and N=4 achieves the minimum ADP for 67.9% of the circuits and has an average ADP deviation of 5.8%, although its average CPD deviation reaches 47.9%. In contrast, K=8 and N=10 reduces the average CPD deviation to 5.9%, with 30.8% of the circuits achieving the CPD optimum, but increases the average area and ADP deviations to 673.2% and 487.0%, respectively. A random-forest cross-domain surrogate achieves a top-5 accuracy of approximately 65%. Therefore, parameter sweeping remains necessary when strict design targets are imposed.  Conclusions  HorizonArch, a parametric architecture description framework for eFPGA exploration, is developed and validated. The framework generates VPR, OpenFPGA, and Yosys backend architecture descriptions from a single source object and provides an extensible interface for COFFE. The parameter-sweep exploration shows that area and CPD favor opposite regions of the (K, N) parameter space. Therefore, eFPGA architecture parameters should be selected according to explicit design targets rather than fixed default values. An open QoR dataset covering five application domains and 64 benchmark circuits is also released as a reusable benchmark for eFPGA architecture design space exploration. Future work will complete the COFFE SPICE topology-rewriting component, refit the routing-area coefficient using measured data, and explore more efficient design space exploration strategies.
A Reconfigurable Parallelized Coprocessor Design for the RISC-V-based Grain Cryptographic Algorithm
NAN Longmei, WANG Haoyu, DU Yiran, LI Wei, CHEN Tao
Available online  , doi: 10.11999/JEIT260391
Abstract:
  Objective  To address the performance limitations of Grain cryptographic algorithms on General-Purpose Processors (GPPs), as well as the inflexibility and high hardware overhead of Application-Specific Integrated Circuit (ASIC) implementations, a dedicated cryptographic hardware accelerator is integrated into an RISC-V coprocessor through a custom instruction extension mechanism. A reconfigurable parallelized architecture is proposed for the Grain algorithm family based on the RISC-V coprocessor interface. Corresponding custom instructions are designed to support the flexible and efficient execution of Grain-80, Grain-128, Grain-128a, and Grain-128AEAD on a unified hardware platform. The proposed architecture provides a favorable balance between processing efficiency, design flexibility, and hardware resource overhead, making it suitable for resource-constrained embedded systems.  Methods  A unified shift-register architecture is adopted to support flexible switching among Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, with a configurable parallelization degree of 1 to 8. To implement the custom instructions for the Grain cryptographic algorithms, the software and hardware functions are analyzed, and the cryptographic process is divided between the processor and coprocessor to achieve efficient execution. The proposed coprocessor uses a streamlined architecture that reuses common hardware resources across the four algorithms. Combined with a configurable feedback tap selection network and feedback logic tailored to a predefined set of algorithms, the architecture enables reconfigurable parallel execution with limited additional hardware resource overhead.  Results and Discussions  The proposed custom instructions enable flexible implementation of four Grain cryptographic algorithms on the same hardware platform. Compared with implementations without instruction extensions, the proposed custom instructions reduce the number of clock cycles and instructions required for cryptographic processing while improving throughput (Table 5). On the Hummingbird E203 platform with a parallelization degree of 4, the proposed implementation requires 105, 129, and 183 clock cycles for Grain-80, Grain-128/Grain-128a, and Grain-128AEAD, respectively, with corresponding throughputs of 220.67, 179.74, and 126.70 Mbit/(s·Hz). Compared with purely software-based implementations, the proposed approach substantially reduces both the number of executed instructions and the number of clock cycles (Table 5). Synthesis results further demonstrate that the coprocessor occupies 7 252.55 μm² in a 65 nm process and achieves a throughput of 1.56 Gbit/(s·Hz) at a parallelization degree of 4. Although the reconfigurable architecture requires slightly more area than a dedicated single-algorithm implementation, it supports four Grain algorithms on a unified hardware platform and provides improved hardware resource reuse and design flexibility.  Conclusions  A hardware-software cooperative reconfigurable parallelization scheme is designed to accelerate the Grain cryptographic algorithm family in lightweight embedded systems. The scheme exploits the RISC-V custom instruction extension mechanism and combines a unified shift-register architecture, a configurable feedback tap selection network, and feedback logic tailored to a predefined set of algorithms. Four algorithms, namely Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, can therefore be flexibly selected and processed in parallel on a single hardware platform. This design improves processing efficiency while maintaining design flexibility and limiting hardware resource overhead. The present work focuses on reconfigurable parallelization for the Grain algorithm family. Future research will investigate more general reconfigurable architectures for nonlinear Boolean functions by incorporating configurable units, such as LookUp Tables (LUTs) or programmable logic arrays. Such architectures may further improve compatibility with multiple stream cipher algorithms while maintaining high throughput and low hardware resource overhead.
Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng
Available online  , doi: 10.11999/JEIT260442
Abstract:
  Objective  Accurate grasp detection from language instructions is essential for service robots to achieve natural human-robot interaction. Existing methods primarily rely on large-scale data-driven training or hierarchical feature fusion to align visual perception with textual instructions. However, they generally overlook the strong coupling between target objects and background clutter in low-level visual features, leading to degraded compositional generalization under cross-view and unseen-scene conditions. To address this limitation, a dual-view cross-scene grasp detection and object localization dataset is constructed to systematically evaluate and improve the compositional generalization of existing models. Based on this benchmark, a Simultaneous Grasp detection and object Localization Network (SGL-Net) is proposed to jointly predict object locations and optimal grasp poses. The proposed framework enables service robots to manipulate objects according to natural language instructions in real-world dynamic environments, providing technical support for embodied intelligence.  Methods  The proposed SGL-Net is illustrated in Fig. 1. First, a Cross-modal Global Context Modulation Module (CGCMM) is proposed to exploit semantic priors from language instructions for adaptive viewpoint correction and background suppression during the early stage of visual feature extraction. Second, a Word-Pixel Cross-modal Alignment Module (WPCAM) is designed to achieve fine-grained semantic decoupling through a flattening-based cross-modal attention mechanism, thereby improving semantic understanding in complex dynamic scenes. Finally, a unified decoder jointly predicts object locations and optimal grasp poses from the fused multimodal features.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on the reconstructed dual-view cross-scene dataset containing bottom-view and top-view scenes and on a real-world robotic grasping platform. Comparative results demonstrate that SGL-Net consistently outperforms mainstream CNN-based and CLIP-based methods in both grasp detection and object localization (Tables 2 and 3). Ablation studies further verify the effectiveness of CGCMM and WPCAM in improving fine-grained semantic alignment and semantic decoupling (Tables 4 and 5). Furthermore, qualitative results (Figs. 46) and real-world robotic experiments (Fig. 7) demonstrate that SGL-Net can be reliably deployed in complex physical environments. Overall, the proposed network exhibits strong generalization capability and excellent potential for practical robotic applications.  Conclusions  To improve cross-view and cross-scene generalization in language-conditioned robotic grasp detection, this paper constructs a dedicated validation dataset and proposes SGL-Net, which jointly performs grasp detection and object localization. By integrating CGCMM and WPCAM, the proposed network accurately localizes instruction-specified objects and predicts optimal grasp poses. Experimental results obtained on multiple benchmark scenarios and a real-world robotic platform demonstrate the superior performance and practical applicability of the proposed method. Future work will focus on integrating Large Multimodal Models (LMMs) and adapting the proposed framework through fine-tuning to further improve zero-shot robotic grasp detection.
Physics-Aware Reconstruction for MilliMeter-Wave Radar Gait Recognition Under Complex Wearing Scenarios
HUANG Ling, QIU Liying, WANG Jiacheng, HAN Penglin, ZHOU Qingdi, YAN Huimei
Available online  , doi: 10.11999/JEIT260522
Abstract:
  Objective  MilliMeter-Wave Radar (MMW Radar) gait recognition has demonstrated considerable potential for non-contact biometric identification because of its inherent advantages in privacy preservation and robustness to variable lighting conditions. However, practical deployment remains challenging under complex wearing conditions, such as long coats or backpacks. These external factors introduce non-stationary high-frequency interference, resulting in spectral aliasing and masking of intrinsic micro-Doppler (m-D) features. Conventional deep learning methods generally treat m-D spectrograms as generic images and overlook the physical relationship between Doppler frequency and human motion. Therefore, clothing-induced interference is easily confused with motion-related features, resulting in reduced recognition performance. This study proposes a physics-aware framework that integrates radar signal physics with human biomechanics to achieve frequency-domain decoupling and adaptive interference suppression for robust radar-based gait recognition under complex wearing conditions.  Methods  To reduce clothing-induced interference, this paper proposes PRISM-Net, a physics-aware gait recognition framework for MMW Radar. The framework is built on the biomechanical characteristics of human motion, where the torso, representing the primary body mass, generates relatively stable low-frequency Doppler components, whereas limb motion produces higher-frequency periodic components. (1) Physics-aware Frequency Structural Reconstruction: Instead of uniformly processing the entire m-D spectrogram, the proposed method performs Physics-aware Frequency Structural Reconstruction by exploiting the velocity distribution associated with different body parts. The original aliased m-D spectrogram is reconstructed into torso- and limb-related frequency components. Low-frequency components preserve stable identity-discriminative features, whereas high-frequency components characterize limb motion. This frequency-domain structural reconstruction isolates spectral regions that are most susceptible to clothing-induced interference, thereby reducing interference propagation. (2) Weighted Attention Mechanism (WAM): The WAM adaptively reweights feature responses according to the reliability of different frequency components. Because clothing-induced interference predominantly affects high-frequency regions, the WAM suppresses interference-contaminated high-frequency responses while enhancing stable torso features, thereby improving feature fusion and recognition robustness. (3) Experimental Configuration: The proposed method is evaluated on the MMRGait-1.0 dataset under a subject-independent evaluation protocol. The training set contains data from 74 subjects, while the remaining 47 unseen subjects are used for testing. All m-D spectrograms are resized to 224 × 224 pixels. The network is optimized using the AdamW optimizer with a joint loss comprising Cross-Entropy Loss, Triplet Loss, and Center Loss to improve both classification performance and feature discriminability.  Results and Discussions  Experimental results demonstrate the effectiveness of incorporating biomechanical priors into gait recognition. As shown in Table 2, PRISM-Net achieves an average Rank-1 accuracy of 89.7% under the 90° side-view condition. In the coat (CT) scenario, the proposed method maintains a Rank-1 accuracy of 85.1%, representing an 18.1% improvement over ShuffleNetV2 and a 7.4% improvement over the ResNet-18 baseline. Model stability is verified through ten independent trials. As illustrated in Fig. 1, PRISM-Net achieves a standard deviation of ±0.55%, compared with ±1.75% for the baseline model. An independent-samples t-test yields p<0.001, confirming that the improvement is statistically significant. Ablation results in Table 2 further verify the contribution of each component. Removing Physics-aware Frequency Structural Reconstruction reduces the CT Rank-1 accuracy to 75.5%, demonstrating the importance of physics-aware frequency-domain decoupling for preventing feature distortion. The WAM further improves CT accuracy by 4.2% through adaptive suppression of high-frequency interference. Regarding computational complexity, Table 3 shows that PRISM-Net contains 11.33 M parameters and requires only 0.75 GFLOPs. Compared with computationally intensive 3D convolution-based models requiring more than 10 GFLOPs, the proposed method achieves superior recognition performance with significantly lower computational complexity. Furthermore, the t-SNE visualization in Fig. 5 shows more compact intra-class distributions and clearer inter-class separation, demonstrating improved feature discriminability.  Conclusions  The proposed PRISM-Net demonstrates that incorporating biomechanical priors into deep learning improves the robustness of MMW Radar gait recognition under complex wearing conditions. By combining Physics-aware Frequency Structural Reconstruction with the Weighted Attention Mechanism, the proposed framework effectively performs frequency-domain decoupling and suppresses clothing-induced high-frequency interference. Experimental results on the MMRGait-1.0 dataset demonstrate that the proposed method achieves high Rank-1 accuracy with low computational complexity, indicating its potential for real-time security applications on edge computing devices.
THz Ultra-Massive MIMO Channel Estimation via a Noise-Conditioned Fixed-Point Network
ZHANG Huawei, NIU Yaning, JIANG Zhanjun, LIU Yingting
Available online  , doi: 10.11999/JEIT260420
Abstract:
  Objective  Terahertz (THz) Ultra-Massive Multiple-Input Multiple-Output (UM-MIMO) systems are expected to support future high-capacity wireless communications. However, accurate channel estimation remains challenging under hybrid near-/far-field propagation and Array-of-SubArrays (AoSA) architectures, where limited Radio-Frequency (RF) chains, low Signal-to-Noise Ratio (SNR), noise uncertainty, and structural perturbations degrade compressed observations. Existing compressed sensing, Bayesian inference, and deep unfolding methods generally rely on fixed statistical assumptions, which limit their cross-SNR generalization under varying noise conditions and statistical mismatches. To address these limitations, this paper proposes a Noise-Conditioned Fixed-Point Network (FPN-NCAS) for robust THz UM-MIMO channel estimation. The proposed method aims to improve estimation accuracy, cross-SNR generalization, robustness, and iterative stability by incorporating noise-aware nonlinear recovery.  Methods  FPN-NCAS is developed within an Orthogonal Approximate Message Passing (OAMP)-based fixed-point unfolding framework. A coarse noise power estimate is obtained from repeated pilot differences and injected into the nonlinear recovery module as an explicit conditioning variable. After each linear update, the vector-domain estimate is reshaped into an AoSA-aligned tensor to exploit subarray-level structural priors. The nonlinear recovery chain consists of three components. Token-Gate performs lightweight subarray-level reliability pre-calibration to suppress unreliable responses under low-SNR and structurally inconsistent conditions. Block Shrink serves as the core noise-conditioned block-sparse proximal operator, in which a smooth dual-threshold mechanism balances strong denoising at low SNR with structural preservation at medium-to-high SNR. G-HMTD(Gated Hybrid Multi-scale Transformer Denoiser) further refines residual errors by combining Local Multi-scale Enhancement and Global Context Modeling. Bridge relaxation and nonlinear residual scaling are also incorporated to improve inter-stage stability.  Results and Discussions  Simulation results demonstrate that FPN-NCAS consistently outperforms LS, OAMP, ISTA-Net+, FPN-OAMP, and FPN-OTFN over the 0~20 dB SNR range (Fig. 6). At SNR = 0 dB, FPN-NCAS achieves Normalized Mean Square Error (NMSE) gains of approximately 3.0 dB and 1.9 dB over FPN-OAMP and FPN-OTFN, respectively. At SNR = 20 dB, the gains increase to approximately 4.5 dB and 2.6 dB (Fig. 6(a)). The convergence curves show that FPN-NCAS reaches a stable plateau after approximately three layers at SNR = 5 dB and five layers at SNR = 15 dB, demonstrating stable fixed-point iterative behavior (Fig. 6(b) and Fig. 6(c)). Analysis of repeated pilots shows that four repeated pilot pairs introduce only 3.13% additional pilot overhead while reducing the relative standard deviation of the coarse noise power estimate to 25.00%. Under moderate noise power mismatch, NMSE degradation remains within 0.3 dB. The proposed method also maintains strong robustness under colored Gaussian noise, impulsive noise, near-/far-field distribution shifts, variation in the number of propagation paths, AoSA subarray shuffling, and RF amplitude/phase mismatch (Fig. 7 and Fig. 8). Ablation studies indicate that Block Shrink provides the largest performance gain, whereas Token-Gate and G-HMTD further improve performance through structural calibration and residual refinement (Fig. 9).  Conclusions  This paper proposes FPN-NCAS for noise-conditioned fixed-point channel estimation in THz UM-MIMO systems. By integrating repeated-pilot-based noise conditioning, AoSA-aware feature reshaping, Token-Gate calibration, Block Shrink recovery, and G-HMTD refinement, the proposed method improves NMSE performance, robustness, and iterative stability under different SNR conditions and non-ideal scenarios. The improved performance is achieved at the cost of higher inference complexity. Future work will focus on lightweight implementations and extensions to wideband, multi-user, and hardware-impaired THz communication systems.
A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming
Available online  , doi: 10.11999/JEIT260512
Abstract:
  Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is essential for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often reduce the stability of target region segmentation, leading to centroid deviation, trajectory jitter, missed detections, and trajectory discontinuity. To address these challenges, a Spatial-Temporal Collaborative Optimization Method is proposed for stable and continuous grab trajectory extraction. While maintaining high inference speed, the proposed method improves both trajectory extraction accuracy and trajectory stability, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates Spatial Representation Enhancement (SRE) and Temporal CONSistency constraint (TCONS). First, CBAM, BiFPN-lite, and shallow feature aggregation are incorporated to improve target-background separability, enhance multi-scale feature representation, and preserve boundary details. TCONS is then imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss to suppress the temporal accumulation of local errors. In addition, a stage-wise training strategy with warm-up epochs and a joint optimization objective is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The proposed method achieves the best segmentation performance on the real portal crane grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also improves performance on DAVIS2016 while maintaining comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, and decreases the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further demonstrate that, under different occlusion ratios, the proposed method maintains good region integrity and continuous target extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementary effects of SRE and TCONS, whereas parameter analysis shows that a TCONS weight of 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  A Spatial-Temporal Collaborative Optimization Method is proposed to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves high segmentation accuracy and stable trajectory extraction on DAVIS2016 and the real portal crane grab dataset, while maintaining comparable segmentation performance on SegTrackV2. It also exhibits strong continuous target extraction capability under occlusion without significantly sacrificing inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further improve the robustness and applicability of the proposed method in real-world environments.
Hierarchical Prototype Learning with Shared Subspace Factorization for Generalizable Deepfake Detection
PENG Shufan, LU Tianliang, HE Chunhao, ZHANG Lu, ZHAO Kai
Available online  , doi: 10.11999/JEIT260426
Abstract:
  Objective  Deepfake detectors often exhibit performance degradation when applied to unseen manipulation methods, cross-dataset distribution shifts, diffusion-generated faces, and common image degradations. In forensic applications, the generation process, data source, and post-processing conditions of a questioned sample are usually unknown. Existing methods often represent the fake class with a single feature center, although different generation methods, data sources, and processing conditions produce heterogeneous patterns. Transferable forensic cues may therefore be mixed with mode-specific artifacts, limiting generalization to unknown domains. To address this problem, a hierarchical prototype learning framework with shared subspace factorization (HPL-SF) is proposed. Within-class diversity is first used to observe latent fake modes, followed by estimation of a low-rank structure shared across these modes and sample-level exploitation of the shared structure during inference.  Methods  A pretrained DINOv2 Vision Transformer (ViT-L/14) is used as the backbone. Its original parameters are frozen, and Low-Rank Adaptation (LoRA) modules are inserted into the query and value mappings of the self-attention layers for parameter-efficient training. All features and prototypes are L2-normalized, and cosine similarity is used for prototype assignment, binary classification, and test-time representation adaptation. HPL-SF consists of three successive stages (Fig. 1). First, one real prototype and multiple mode-specific fake prototypes are maintained in the normalized feature space. Each training sample is softly assigned to the fake prototypes according to its cosine similarity to each prototype, and the weighted prototypes are aggregated to obtain a sample-adaptive fake representation. The response distribution thus provides an observable representation of latent within-class modes without using forgery-source labels. A binary classification loss, a sample–prototype contrastive loss, and a prototype diversity loss are jointly optimized to separate real and fake samples, improve sample–prototype alignment, and prevent prototype collapse. Second, the normalized mode-specific fake prototypes are arranged into a prototype matrix. Singular Value Decomposition (SVD) is applied to this matrix, and the largest gap between consecutive singular values is used to determine the dimension of the shared fake subspace. Each fake prototype is then decomposed into a projection onto the shared fake subspace and a mode-specific residual. Only the shared projections are aggregated to update the shared fake prototype, whereas the residuals retain mode-specific information and are excluded from this update. The real prototype and shared fake prototype are updated using Exponential Moving Average (EMA). Gradients are not propagated through subspace construction, dimension selection, or semantic prototype updates. Third, the backbone, LoRA modules, prototypes, and subspace basis are fixed during inference. A test feature is compared with the real and shared fake prototypes, and their relative responses determine a sample-specific adaptation weight. The original feature is blended with its projection onto the shared fake subspace and then L2-normalized. The resulting feature is classified according to its cosine similarities to the two semantic prototypes. Thus, the estimated shared structure is exploited on a sample-by-sample basis without updating model parameters during testing.  Results and Discussions  The method is evaluated under cross-dataset, cross-forgery-type, diffusion-forgery, repeated-run, image-degradation, ablation, and mechanism-analysis protocols. Across seven unseen datasets, HPL-SF achieves the highest area under the receiver operating characteristic curve (AUC) on all seven datasets, with an average AUC of 91.67%, which is 2.10 percentage points higher than the second-highest average (Table 1). When trained on FaceForensics++ and evaluated on the Diffusion Facial Forgery dataset, HPL-SF achieves the highest AUC on the text-to-image, image-to-image, face-swapping, and face-editing subsets, with an average AUC of 84.39% (Table 2). Across four cross-forgery-type settings on FaceForensics++, the average accuracy and AUC reach 85.29% and 92.41%, respectively, both ranking first among the compared methods. When DeepFakes under high-quality compression is held out for testing, HPL-SF trails the best-performing method by only 0.91 percentage points in accuracy and 0.39 percentage points in AUC (Table 3). Repeated experiments with multiple random seeds yield the highest mean values for all four aggregate metrics, with standard deviations ranging from 0.38 to 0.62 percentage points (Table 4). Under five severity levels of compression, blur, and noise on the Deepfake Detection Challenge Preview dataset, HPL-SF achieves the best or joint-best AUC at most severity levels and remains relatively stable under moderate and severe degradations (Fig. 2). All six ablated variants perform worse than the complete HPL-SF model on the four unseen test sets, indicating complementary contributions from mode-specific prototype learning, sample–prototype contrastive loss, prototype diversity loss, shared subspace factorization, and test-time representation adaptation (Fig. 3). Performance generally increases as the number of mode-specific fake prototypes increases and plateaus when the number reaches 10. The largest spectral gap occurs between the fifth and sixth singular values, yielding a shared subspace dimension of 5. With this setting, HPL-SF achieves a diffusion-forgery AUC of 84.39% ± 0.62%, compared with 79.68% for direct mean aggregation and 81.56% ± 0.80% without test-time representation adaptation (Figs. 4(a)–4(c)). Linear-probe analysis further shows that the shared component provides stronger discrimination in unknown domains and weaker domain-identifying capability than the mode-specific residual, supporting the separation of transferable and domain-related cues (Fig. 4(d)). The t-distributed stochastic neighbor embedding (t-SNE) visualization shows clearer real–fake separation and closer distributions of same-class samples from different data sources (Fig. 5). Prototype allocation statistics show differentiated responses across fake types, with prototype usage remaining close to the uniform baseline. The adaptation weights are higher for fake samples, particularly for correctly classified fake samples, whereas misclassified samples show responses closer to the balanced-response line (Fig. 6). False positives are mainly associated with low resolution, compression, filters, or occlusion, whereas false negatives usually contain high-quality forgeries or weak forgery traces (Fig. 7).  Conclusions  HPL-SF organizes generalizable deepfake detection into three successive stages: observation of within-class diversity, estimation of shared structure, and sample-level exploitation of that structure. Experimental results show that separating shared projections from mode-specific residuals provides transferable discriminative cues under dataset shifts, unseen forgery types, diffusion-generated manipulations, and common image degradations. The framework requires no parameter updates during testing and adaptively exploits the shared structure according to each sample’s relative responses to the real and shared fake prototypes. Errors remain for degraded real images whose artifacts resemble forgery traces and for high-quality forgeries with weak traces. Future work will extend the framework to temporal prototypes for video, multimodal forensic evidence, and adaptive discrimination in broader open-world settings.
Dual-Domain Differentiated Feature Extraction Network for MRI Reconstruction
XUE Nan, QIAO Han, WANG Peng
Available online  , doi: 10.11999/JEIT251093
Abstract:
  Objective  In magnetic resonance imaging (MRI), undersampling of k-space data is an effective approach to accelerate image acquisition. Reconstructing high-quality MR images from undersampled data is of great clinical significance for diagnostic efficiency. Currently, dual-domain reconstruction methods that jointly exploit spatial- and frequency-domain features have become mainstream. However, existing dual-domain MRI reconstruction methods fail to design differentiated feature extraction strategies for the two domains, and their physical prior constraints are insufficient, leading to loss of original information during the reconstruction process.  Methods  To address these issues, this study proposes a Dual-Domain Differentiated Feature Extraction Network (DDF-Net) for MRI reconstruction. In the spatial domain, an interlaced row-column self-attention mechanism combined with depthwise convolution is designed to accurately capture anisotropic structures and texture details, thereby addressing the limitations of conventional convolutional feature extraction. In the frequency domain, amplitude and phase characteristics are independently modeled to capture intensity and structural positional information, respectively, and a frequency-domain feature enhancement module is introduced to fully exploit spectral information for improving reconstruction fidelity. Finally, a data consistency layer and a cross-domain adjustment module are incorporated to integrate physical priors and cross-domain information, thereby reinforcing measurement consistency and stabilizing the reconstruction process.  Results and Discussions  The proposed DDF-Net was evaluated on two publicly available MRI datasets, CC359 and IXI, and compared with six representative reconstruction algorithms, including DAGAN, KIKI-Net, MD-Recon-Net, SwinMR, Reconmer, and KTMR. As shown in Fig. 6 and Fig. 7, under a Gaussian 1D 30% undersampling mask, DDF-Net achieves the most faithful anatomical restoration with clearer cortical edges and finer texture details, while suppressing aliasing artifacts effectively. The error maps exhibit more uniform residual distributions, indicating better consistency between the reconstructed and reference images. Quantitative comparisons (Table 1 and Table 2) show that DDF-Net attains the highest PSNR and SSIM values across all sampling patterns. Specifically, it improves PSNR by 0.17 dB and 0.23 dB, and SSIM by 0.0063 and 0.0019 on the CC359 and IXI datasets, respectively, achieving the best overall performance over the second-best competing method. These results demonstrate that the differentiated spatial-frequency feature extraction enables DDF-Net to leverage complementary information between domains for more precise recovery. To further verify robustness, a noise experiment was conducted by adding Gaussian noise of varying intensity levels to the k-space data, following the procedure in SwinMR. As illustrated in Fig. 8 and summarized in Table 3, DDF-Net exhibits stronger robustness to noise perturbations, maintaining higher PSNR and SSIM values than SwinMR under all noise conditions. Even at NL = 80%, DDF-Net effectively preserves most structural details, confirming that the amplitude-phase separation and multi-level data consistency jointly improve noise resilience and stability.  Conclusions  This study proposes an end-to-end Dual-Domain Differentiated Feature Extraction Network to address the limitations of existing dual-domain MRI reconstruction methods, which often lack domain-specific feature extraction strategies and sufficient physical prior constraints, leading to potential information loss during reconstruction. In the spatial domain, an interlaced row-column self-attention mechanism is designed to more effectively model local structures and texture details. In the frequency domain, amplitude-phase separation and a frequency-domain feature enhancement module are introduced to effectively improve the decoupling and utilization of spectral information. Furthermore, by integrating a cross-domain adjustment module with multiple data consistency layers, DDF-Net enhances cross-domain interaction and measurement fidelity. Experimental results on the CC359 and IXI datasets demonstrate that the proposed method achieves superior performance in both quantitative metrics and visual reconstruction quality, successfully recovering fine anatomical details. In addition, noise experiments confirm the robustness of DDF-Net under complex conditions, showing that the network can effectively preserve image details across various noise levels. Future work will focus on extending DDF-Net to unsupervised or semi-supervised learning frameworks to reduce dependence on labeled data and further enhance its potential for clinical applications in fast MRI reconstruction.
Multi-RAT Fusion Architecture and Intelligent Routing for Marine Heterogeneous Wireless Networks
CHEN Jin, ZHOU Xuan, LIN Haitao, YU Huagang, LI Yun
Available online  , doi: 10.11999/JEIT260482
Abstract:
  Objective  Marine Heterogeneous Wireless Networks (MHWNs), which deeply integrate multi-dimensional resources spanning space, air, and sea, must accommodate multiple coexisting communication systems and thus face the dual challenges of interconnection and resource coordination. Meanwhile, existing routing algorithms based on Deep Reinforcement Learning (DRL) exhibit insufficient representation capability for dynamic topologies, making it difficult to make efficient routing decisions when the network topology changes frequently. This is mainly because mainstream frameworks typically adopt standard Graph Neural Networks (GNNs) or fully connected networks for state encoding, which cannot effectively capture the structural features of highly dynamic topologies.  Methods  This paper designs a modular Multi-Radio Access Technology (Multi-RAT) gateway supporting the fusion of 5G and ad hoc (Mesh) networks, and proposes an intelligent routing method combining a Contrastive Message Passing Graph Neural Network with Deep Reinforcement Learning (CMPGNN-DRL), so as to improve Quality of Service (QoS) and forwarding efficiency and achieve optimized resource allocation. In the gateway, service data are IP-encapsulated, protocol-identified, and semantically converted by communication interface modules before being forwarded to the target interface. The gateway periodically collects node features and link states to construct the input graph for routing decisions, and monitors key indicators such as link bandwidth utilization, queue depth, packet loss rate, and end-to-end delay. To compensate for the information delay introduced by periodic reporting, short-term trend terms of key indicators are incorporated into the state vector, and an asynchronous decision-execution architecture is adopted. Under a centralized Software-Defined Networking (SDN) control plane, the network is modeled as a graph with continuously monitored node and link features. For each node, the CMPGNN module synchronously constructs homophily and heterophily views along two message-passing paths and constrains the resulting embeddings with a contrastive loss, yielding discriminative node representations that are robust to edge perturbations. The learned representations are then fed into a Double Deep Q-Network (DDQN) agent that makes hop-by-hop routing decisions with an ε-greedy exploration strategy. Specifically, the topology prior values predicted by CMPGNN for neighboring nodes are fused with the DDQN Q-value estimates through weighted summation to form joint action values, which can correct inaccurate Q-value estimates when training is insufficient or observations are noisy. A normalized multi-objective reward function is designed to be negatively correlated with latency, packet loss rate, and link load, and positively correlated with throughput, while explicitly penalizing routing loops.  Results and Discussions  The proposed solution was validated through hardware prototype measurements and extensive simulations. Prototype tests showed that the average CPU utilization of the fusion gateway was 14%, 25%, and 37% in Mesh-only, 5G-only, and dual-mode operation, respectively; the dual-mode aggregate throughput reached 108 Mbps, compared with 32 Mbps for Mesh-only and 84 Mbps for 5G-only (uplink); and the average ping latencies between the gateway and the application server were 6 ms for Mesh and 16 ms for 5G. CMPGNN-DRL was compared with six baseline methods, namely OSPF, AODV, GNN, DQN, MPNN-DQN, and DAR-DRL, on the GEANT2, GBN, Germany, and Synth50 topologies, covering dynamic traffic, random link failures with failure rates of 3%–24%, and large-scale topology variations. The training reward increased rapidly and then stabilized, and ablation experiments verified the effectiveness of the contrastive learning mechanism. Compared with the optimal baseline, the proposed method reduces the average end-to-end delay by 20.8%–47.7%, reduces the packet loss rate by 0.3%–5.3%, and increases the average throughput by 5.2%–14.2%. In maritime heterogeneous wireless network scenarios constructed according to the environmental constraints of the Maritime Internet of Things (MIoT), i.e., a 1500 m × 1500 m area with 50–100 randomly deployed nodes evaluated through repeated Monte Carlo simulations, the method adapted stably to variations in network scale and node mobility in terms of Packet Delivery Ratio (PDR) and packet loss rate. As the load rate increased from 20% to 50%, it improved PDR by 4.2%–17.3% over MPNN-DQN and DAR-DRL while maintaining lower latency, higher bandwidth utilization, and a lower packet retransmission ratio under medium-to-high loads.  Conclusions  Aiming at sea-air cross-domain heterogeneous networks, this paper designed a Multi-RAT fusion gateway supporting ad hoc networks and 5G, and proposed an intelligent multi-path routing algorithm integrating CMPGNN with DRL. The contrastive learning mechanism strengthens the topology representation capability of the graph neural network and improves the robustness of routing policies. Experimental results show that the proposed method outperforms existing mainstream algorithms in key performance indicators such as PDR, end-to-end delay, and throughput, and exhibits good cross-topology generalization capability. Future work will focus on verification in real maritime environments and optimization of training efficiency, so as to support the practical deployment and application of integrated sea-air communication systems.
Cooperative Search and Tracking of Moving Ships Using Constellation Multi-Functional Payloads Based on Dynamic Information Gain
ZHANG Yumo, ZHAO Fuhai, LI Xiaobin, FAN Shenghua, QU Tao
Available online  , doi: 10.11999/JEIT260500
Abstract:
  Objective  Wide-area maritime surveillance requires satellite constellations to search for and revisit non-cooperative maneuvering ships whose positions become uncertain after missed observations. Meanwhile, multi-functional payloads are subject to coupled constraints on observation timing, attitude maneuvering, payload mode, and energy consumption. To address dynamic target uncertainty and executable constellation scheduling, a cooperative search-and-tracking method based on dynamic information gain is proposed.  Methods  A closed-loop rolling-horizon framework inspired by Model Predictive Control is constructed to perform prediction, optimization, execution, and feedback. At each decision epoch, candidate atomic tasks are generated over a planning horizon, while only tasks within the current execution window are committed. Each task specifies the executing satellite, target, candidate pointing grid, start/end times, and payload mode. Target uncertainty is represented by a parameterized probabilistic grid derived from the latest confirmed state, speed and heading perturbations, and elapsed time since the last successful observation. Hit/Miss feedback updates the uncertainty baseline for the next rolling step, where the probability grid, candidate tasks, and observation plan are regenerated. A state-driven dual-mode benefit model is established according to target information entropy and consecutive successful observations. In the robust tracking state, narrow-field tasks are evaluated by the prior capture probability, namely the probability mass covered by the task footprint. In the lost-search state, wide-field tasks are evaluated by the binary entropy of Hit/Miss events as an approximation of search information value. This approximation is motivated by Kullback-Leibler divergence and avoids explicit posterior reconstruction for every candidate task. A dynamic priority coefficient increases scheduling urgency for long-unobserved targets and moderately down-weights repeatedly confirmed targets. The resulting multi-objective model maximizes weighted task benefit and information gain while minimizing energy consumption, subject to hard constraints on single-satellite temporal exclusivity, attitude-transition stabilization time, and available energy. Based on NSGA-II, the Cooperative Evolutionary Planning-Multi-Objective (CEP-MO) algorithm employs global integer-index encoding, constraint-aware Top-K heuristic initialization, satellite-group crossover, and adaptive repair to improve feasible-solution generation. Feasible Pareto solutions are normalized, and the solution closest to the ideal point (1,1,0) is selected for execution.  Results and Discussions  Simulations with a 48-satellite Walker constellation demonstrate the effectiveness of the proposed method. In the 200-target scenario, Standard NSGA-II obtains an average revisit interval of 53.5 min and a weighted coverage of 13.9%, whereas CEP-MO achieves 19.9 min and 41.7%, respectively, reducing the average revisit interval by 62.8% (Fig. 7). Removing the binary-event-entropy benefit increases system-average uncertainty and revisit interval, while replacing constraint-aware initialization with random initialization degrades early convergence and weighted coverage. As the target number increases from 100 to 200, CEP-MO maintains acceptable scalability (Fig. 8). At 200 targets, its weighted coverage is 16.6 percentage points higher than that of CEP-MO w/o Entropy, the computation time per rolling decision is approximately 22 s, and the Gini coefficient of remaining satellite energy stays below 0.3, indicating that energy consumption is not excessively concentrated on a small subset of satellites.  Conclusions  The proposed framework integrates probabilistic-grid uncertainty representation, state-driven search/tracking benefit evaluation, rolling feedback, and constraint-aware multi-objective evolutionary planning. CEP-MO improves revisit and weighted-coverage performance while maintaining temporal, attitude, and energy feasibility. The method provides an effective approach for large-scale resource-constrained maritime surveillance and a basis for future extensions involving identification errors, communication delays, and constrained inter-satellite links.
Preamble-Referenced Cyclic Cross-Correlation Chirp Spread Spectrum Communication Technology in Complex Multipath Environments
YE Yun, ZHANG Chengyu, PANG Haodong, MA Wenfeng, LI Xuejiao, ZHANG Xiaokai
Available online  , doi: 10.11999/JEIT260702
Abstract:
  Objective  Ground unmanned platforms operating in urban streets, industrial parks, and underground passages require short-burst reliable command-and-control communication. These complex near-ground environments simultaneously impose strong multipath fading, large Carrier Frequency Offset (CFO), residual Timing Offset (TO), Sampling Frequency Offset (SFO), and in-band interference from coexisting wireless systems. Conventional Chirp Spread Spectrum (CSS) receivers based on single-peak decisions in the dechirp–Discrete Fourier Transform (DFT) domain suffer from multipath-induced spectral splitting and interference-induced bin masking, while Direct-Sequence Spread Spectrum (DSSS) baselines exhibit synchronization fragility under combined offsets and degraded energy efficiency under multipath. Simultaneously addressing these impairments is essential for enabling robust low-latency control links and for the coexistence of unmanned platforms with legacy wireless infrastructure in dense deployments.  Methods  This paper proposes a Preamble-Referenced Cyclic Cross-Correlation CSS (PRCC-CSS) scheme that jointly designs frame structure, synchronization estimation, and payload detection. Each frame comprises multiple identical up-chirp preamble symbols, a down-chirp Start Frame Delimiter (SFD), and CSS-modulated payload symbols. The complementary frequency-domain indices at the dechirp-DFT outputs of the up-chirp preamble and the down-chirp SFD are exploited to jointly estimate integer CFO and TO via closed-form linear combinations. Fractional CFO is recovered from inter-symbol phase differences across adjacent preamble symbols; fractional TO is extracted from the centroid of the main-peak neighborhood of the averaged preamble spectrum; and under a common oscillator-reference assumption, SFO-induced bin drift is compensated using the CFO-derived clock-offset relation. After offset compensation, the averaged preamble spectrum serves as a frame-specific reference spectrum that captures the instantaneous multipath fingerprint of the channel. Payload symbols are then detected by computing the cyclic cross-correlation between this reference spectrum and each candidate-shifted payload spectrum, taking the maximum-correlation index as the demodulated symbol. This formulation converts the multipath-induced frequency-domain structure from an adverse perturbation into a matchable intra-frame reference feature, thereby enabling multipath-robust structure-matched payload detection without explicit path-by-path channel estimation.  Results and Discussions  PRCC-CSS is evaluated under identical bandwidth, sampling rate, and processing gain at a target Bit Error Rate (BER) of 10–4. First, it is compared against three DSSS baselines using Binary Phase-Shift Keying (BPSK), Quadrature Phase-Shift Keying (QPSK), and 16-ary Quadrature Amplitude Modulation (16QAM) under three channel conditions. Under Additive White Gaussian Noise (AWGN, Fig. 1), PRCC-CSS reaches the target at approximately 5 dB Eb/N0 versus 8.5–9 dB for the best DSSS baseline, indicating that chirp index modulation combined with preamble-referenced correlation provides inherent frequency-domain energy aggregation independent of any specific multipath profile. Under the Extended Typical Urban (ETU) channel without interference (Fig. 2), PRCC-CSS requires approximately 6.5 dB versus approximately 10 dB, as the preamble reference spectrum captures and reuses the per-frame multipath structure that finite-finger DSSS-RAKE cannot fully exploit due to path-capture and code-synchronization errors. Under ETU with in-band interference at Jamming-to-Signal Ratio (JSR) = 5 dB (Fig. 3), PRCC-CSS requires approximately 8.1 dB versus 10.8–11.0 dB, since the cyclic cross-correlation preserves decision separability—reference and payload spectra share nearly identical channel structure within the same frame—whereas DSSS-RAKE accumulates interference residue across all combining fingers. Second, against Dechirp Non-Coherent (DNC) and coherent peak detection at Spreading Factor (SF) = 7 and SF = 9 under ETU with interference (Fig. 4), DNC fails to reach the target within the tested Signal-to-Interference-plus-Noise Ratio (SINR) range at SF = 7 and requires approximately 4–5 dB at SF = 9, whereas PRCC-CSS reaches the target at approximately –5 dB SINR at SF = 7 and –11 dB at SF = 9, yielding a 10–15 dB SINR threshold improvement and confirming that the gain stems from frame-wide cyclic matching rather than phase compensation alone. Third, a Software-Defined Radio (SDR) prototype on an ETU-emulated channel at JSR = 5 dB (Fig. 5Fig. 6) retains 300 of 332 received frames as reliable (90.4%). In the retained reliable frames, no symbol errors were observed among 6,000 payload symbols and no bit errors were observed among 54,000 payload bits, corresponding to a one-sided 95% upper confidence bound of 5.56 × 10-5 on the retained-frame conditional BER.  Conclusions  This paper proposes the PRCC-CSS scheme that jointly integrates integer/fractional CFO–TO and SFO estimation with cyclic cross-correlation payload detection. Results demonstrate that: (1) at BER = 10–4, PRCC-CSS lowers the required Eb/N0 by approximately 2.7–4.0 dB relative to the best DSSS-BPSK/QPSK/16QAM baseline across AWGN, ETU, and ETU with JSR = 5 dB cases; (2) under ETU with interference, PRCC-CSS lowers the required SINR by approximately 10–15 dB relative to DNC at SF = 7 and SF = 9, with a further consistent margin over coherent peak detection; (3) the SDR prototype retains 90.4% of received frames, and the retained-frame conditional BER has a one-sided 95% upper confidence bound of 5.56 × 10–5. By exploiting the multipath-induced frequency-domain structure as a matchable intra-frame reference feature rather than as a perturbation, PRCC-CSS provides a candidate physical-layer solution for short-burst reliable communication in complex near-ground environments. Future work will extend the scheme to higher-order CSS modulations, multi-antenna diversity reception, and adaptive reference-spectrum updating for time-varying channels.
Energy Efficiency Analysis of Discrete Phase-shifted Active RIS Enhanced Communication Systems
SHU Feng, LIN Zhiyuan, ZHENG Weihai, WANG Yan, JIANG Hao, WANG Jiangzhou
Available online  , doi: 10.11999/JEIT260462
Abstract:
  Objective  Active Reconfigurable Intelligent Surface (RIS) enhances wireless communication performance by integrating radio frequency amplifiers to mitigate the multiplicative fading inherent to passive RIS. However, amplification noise and additional power consumption are introduced. Furthermore, high-precision digital phase control at the base station incurs considerable communication overhead. Employing low-precision phase shifters is therefore an effective approach for practical RIS deployment. Therefore, characterizing the Energy Efficiency (EE) performance of active RIS-assisted communication systems and quantifying the effect of finite-bit phase quantization errors on EE are essential for system design and practical implementation. To this end, a discrete phase-shifted active RIS-assisted communication system over Rayleigh fading channels is investigated. The EE loss caused by phase quantization errors is analyzed, approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE are derived, and the relationship between RIS EE and user EE is established, providing theoretical guidance for the practical deployment of active RIS.  Methods  Based on the law of large numbers and Taylor series expansion, closed-form expressions for the user EE loss and its approximation are derived. The effects of system parameters on EE are investigated by expressing EE as explicit univariate functions. Ferrari’s method and the Lambert W function are then employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE. Finally, the relationship between RIS EE and user EE is established using the law of large numbers and the Lambert W function.  Results and Discussions  User EE is expressed as a function of six parameters: the number of quantization bit (\begin{document}$ k $\end{document}), power allocation factor (\begin{document}$ \beta $\end{document}), the number of RIS elements (\begin{document}$ N $\end{document}), the total power sum of base station and active RIS (\begin{document}$ {P}_{\text{t}} $\end{document}), the noise at active RIS (\begin{document}$ \sigma _{\text{r}}^{2} $\end{document}), and the noise at user (\begin{document}$ \sigma _{\text{u}}^{2} $\end{document}). First, the EE loss decreases as \begin{document}$ k $\end{document} increases. When \begin{document}$ k $\end{document}=3, the difference between the approximate EE loss and the lossless case is less than 0.026 8 Mbit/J, while the difference between the EE loss and the lossless case is less than 0.026 5 Mbit/J (Fig. 3). Therefore, 3-bit to 4-bit discrete phase shifters achieve performance close to that of continuous phase shifters. Second, user EE exhibits a unimodal dependence on both \begin{document}$ \beta $\end{document} and \begin{document}$ N $\end{document}. The approximate optimal solution for \begin{document}$ \beta $\end{document} differs from the exact optimal solution obtained by the Dinkelbach algorithm by less than 0.01 (Fig. 4), whereas the approximate and exact optimal solutions for \begin{document}$ N $\end{document} are identical (Fig. 5), demonstrating the high accuracy of the proposed approximations. Third, user EE exhibits a unimodal trend as \begin{document}$ {P}_{\text{t}} $\end{document} increases. Higher phase quantization precision produces a higher EE peak while requiring a lower optimal \begin{document}$ {P}_{\text{t}} $\end{document} to achieve the maximum EE (Fig. 6). In addition, user EE decreases as both \begin{document}$ \sigma _{\text{r}}^{2} $\end{document} and \begin{document}$ \sigma _{\text{u}}^{2} $\end{document} increase. User EE is more sensitive to the amplification noise introduced at the RIS, indicating that reducing the RIS noise power yields a greater EE improvement (Fig. 7). Finally, user EE first increases and then decreases sharply to zero as RIS EE increases. The signal-to-noise ratio at the RIS is identified as the key factor governing the relationship between RIS EE and user EE (Fig. 8).  Conclusions  The EE performance of active RIS-assisted wireless networks employing discrete phase shifters over Rayleigh fading channels is investigated. First, closed-form expressions are derived for the user EE in the lossless case, the lossy case, and the approximate-loss case. Simulation results demonstrate that 3-bit to 4-bit discrete phase shifters closely approach the performance of continuous phase shifters. Next, explicit functions describing the effects of key system parameters on user EE are established. Ferrari’s method and the Lambert W function are employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE, and both exhibit negligible errors relative to the exact solutions. Finally, the relationship between RIS EE and user EE is established, demonstrating that user EE initially increases and subsequently decreases to zero as RIS EE increases.
Multi-Frequency Feature Interaction and Adaptive Fusion for cooperative Spacecraft 6D Pose Estimation
TONG Wei, LIN Xi, YAN Ying, LIN Jinxing, LI Tao, WU Qi
Available online  , doi: 10.11999/JEIT260452
Abstract:
Spacecraft 6D pose estimation aims to determine the relative pose between the target spacecraft and the service spacecraft in the spatial coordinate system, which is a critical step for a series of close-range operation tasks, such as failed satellite cleaning, space debris capture, on-orbit spacecraft manipulation, and space station rendezvous and docking. In recent years, CNN-based 6D pose estimation methods have received widespread attention. However, their over-reliance on convolutional network architectures makes them sensitive to image textures and limits their capacity to effectively model long-range contextual information. Moreover, current mainstream methods typically adopt a pipeline design consisting of object detection followed by pose estimation, which suffers from limited diversity in feature extraction and is sensitive to low-light conditions and complex background interference. To address these issues, this paper proposes a spacecraft pose estimation network based on multi-spectral feature interaction and dynamic fusion. Specifically, the network first leverages backbone networks with different receptive fields to separately extract spatial high-frequency features (such as semantic and edge details) and spatial low-frequency features (such as global structural information). On this basis, a Transformer-based feature matching mechanism is employed to perform self-attention and cross-attention feature interactions, thereby aggregating long-range contextual information. To further exploit the rich frequency-domain feature representations, a frequency-guided feature module is introduced to dynamically fuse multi-spectral features. Finally, extensive experiments on spacecraft pose estimation datasets demonstrate that the proposed method achieves competitive performance and strong generalization ability, showing advantages over existing methods.  Objective  Due to the limitation of local receptive field of CNN convolution operator, the mining of remote correlation information is often not ideal. In contrast, Transformer with multi-head intra-attention mechanism is more effective in globally modeling long-range context information and is good at processing multi-scale objects and small spacecraft at a long distance. Therefore, enhancing the processing ability of low-frequency and high-frequency features is of great significance for improving the accuracy.  Methods  This work proposes a 6D pose estimation network with Transformer-based multi-frequency feature interaction and fusion. The framework consists of two parts. One is to extract high-frequency features such as image semantics and edges via ResNet18 and leverage them for spacecraft semantic segmentation directly. The other is to extract global low-frequency features such as image texture through DarkNet53. Then Transformer-based feature interaction module is designed to perform attention between high-frequency and low-frequency features, enabling the aggregation of long-range contextual information, which can enhance the diversity of texture-less spacecraft image features and promote network optimization.  Results and Discussions  The proposed method estimates spacecraft 6D pose via multi-frequency feature interaction and adaptive fusion. On the SwissCube dataset, it achieves an overall ADI-0.1d accuracy of 82.31%, outperforming CA-SpaceNet (79.39%) and WDR* (78.78%). On the SPEED dataset, the combined error eq+et is 0.0299, which is 22.3% lower than CA-SpaceNet (0.0385) and 25.2% lower than WDR* (0.0400). Qualitative results in Fig. 5 and Fig. 7 show that predicted key points are significantly closer to the ground truth, especially under challenging low-frequency conditions. Ablation studies confirm that the feature interaction module alone increases overall accuracy from 79.39% to 80.67%, and adding frequency-guided fusion further raises it to 82.31%. These results demonstrate that the proposed framework enhances pose estimation robustness and meets the stringent requirements for on-orbit servicing and space rendezvous.  Conclusions  To enhance the accuracy of spacecraft 6D pose estimation under extreme atmospheric environments, this work innovatively designs a sub-branch based on multi-frequency feature interaction and additional semantic edge segmentation, and overcomes the limitation of low feature extraction efficiency of existing methods by aggregating long-range multi-frequency features. In addition, a frequency-guided dynamic feature fusion module is introduced to fully leverage the rich frequency-domain feature representation. Comprehensive experimental comparison with mainstream methods on SwissCube and SPEED datasets verifies that the proposed work can enhance the representation of feature information and improve the robustness of spacecraft pose estimation.
Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets
LIU Minglong, JIANG Tingyao, LI Yulan
Available online  , doi: 10.11999/JEIT260847
Abstract:
  Objective  Infrared dim and small target detection exploits thermal-radiation differences between targets and backgrounds for passive sensing in air-ground inspection, maritime surveillance, and wide-area warning. Under long-range imaging, low signal-to-noise ratios, and complex backgrounds, targets occupy few pixels and exhibit weak textures, blurred contours, and low contrast. Bounding-box detection is well suited to edge-deployed rapid detection and subsequent tracking initialization by directly predicting class confidence and location. Although the Real-Time Detection Transformer (RT-DETR) provides an efficient end-to-end framework, its application faces three limitations: repeated downsampling weakens local details and initial localization cues; intra-scale global interaction insufficiently distinguishes high-frequency target responses from low-frequency background context; and cross-scale fusion may propagate homogeneous clutter without explicit spatial constraints. Therefore, a Frequency-Domain Decoupling and Spatial-Prior-Constrained DETR (FSP-DETR) is proposed to coordinate shallow detail preservation, target-background decoupling, and localization constraints.  Methods  FSP-DETR is built on RT-DETR-R18 and follows backbone feature extraction, intra-scale interaction, cross-scale fusion, and query-based decoding (Fig.1). The three modules operate at successive stages to address the three identified limitations. First, a Subspace Progressive Attention Modulated Inverted Residual Block (SPA-MIRB) is embedded in the Residual Network 18 (ResNet18) backbone to preserve shallow details (Fig.2). Its inverted-residual branch combines pointwise and depthwise convolutions with a residual connection to retain edges, spots, and weak textures; its attention branch partitions channels into subspaces and progressively transfers interactions to strengthen weak targets. Second, a Frequency-Domain Asymmetrically Decoupled Attention-based Intra-scale Feature Interaction module (FD-AIFI) replaces the Attention-Based Intra-Scale Feature Interaction module (AIFI) (Fig.3). Unified global interaction is divided into high-pass local-attention and low-pass global-attention paths. The former uses local-window attention to preserve compact target responses and limit distant clutter. The latter retains query resolution but pools key and value features, compressing low-frequency context while preserving spatial queries and target-position indices. Their outputs are concatenated channel-wise and linearly projected. Third, a Spatial-Prior-Constrained CNN-based Cross-scale Feature-fusion Module (SPC-CCFM) retains the upsampling, concatenation, and convolution path while introducing a High-Order Spatial Representation Module (HSRM) and an Overall-Level Mapping Injection Mechanism (OMIM). HSRM aligns multilevel features and constructs spatial passthrough, high-order spatial-relation, and low-order fidelity branches (Fig.4). The branches preserve original spatial responses, model spatially separated yet response-similar clutter through hypergraph aggregation, and supplement local edges and weak textures. OMIM adapts and injects the prior into each fusion node to constrain cross-scale refinement.  Results and Discussions  Experiments are conducted on IRSTD-1k and NUAA-SIRST, with pixel-level masks converted into single-class minimum bounding rectangles. Component ablation shows that SPA-MIRB, FD-AIFI, and SPC-CCFM each improve mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@0.5) on both datasets (Table 1). FSP-DETR achieves 89.03% and 98.75% mAP@0.5 on IRSTD-1k and NUAA-SIRST, exceeding RT-DETR-R18 by 3.39 and 3.08 percentage points. Parameters decrease from 19.87 million to 17.24 million and computational cost from 56.9 to 50.4 billion floating-point operations, by 13.2% and 11.4%. Although some variants yield higher precision or recall at one operating point, the complete model obtains the highest mAP@0.5 with lower complexity. Replacement experiments support the structural choices: SPA-MIRB and FD-AIFI achieve the highest mAP@0.5 among their counterparts, while SPC-CCFM exceeds both fusion alternatives by 0.48 and 0.26 percentage points, respectively, highlighting spatial constraints over complexity reduction (Table 2). SPA-MIRB reaches 96.75% mAP@0.5 with 15.28 million parameters and 46.8 billion floating-point operations, providing a balanced accuracy-complexity trade-off. A balanced path-allocation factor of 0.50 performs best, reaching 97.27% mAP@0.5 and 53.04% mAP@0.5:0.95 on NUAA-SIRST (Table 3). Attention maps show concentrated target-neighborhood responses in the high-pass path and broader, smoother background responses in the low-pass path (Fig.5). Under unified settings, FSP-DETR obtains the highest mAP@0.5 on both datasets, with 172 frames/s and an average per-image forward time of 5.81 ms (Table 4). It also surpasses deeper RT-DETR variants with fewer parameters, less computation, and shorter forward time, showing the advantage of task-oriented adaptation over backbone enlargement. Detection visualization shows better agreement between predicted boxes and target regions in cluttered scenes (Fig.6). Instance-level analysis shows missed-detection rates decrease from 15.70% to 13.22% on IRSTD-1k and from 10.71% to 5.36% on NUAA-SIRST, while average false positives per image decrease from 0.490 to 0.380 and from 0.209 to 0.070 (Table 5). Center-offset errors also decrease on both datasets. However, the scale error on NUAA-SIRST increases slightly from 10.63% to 11.05%, indicating that scale regression for extremely small targets remains challenging.  Conclusions  FSP-DETR coordinates shallow detail preservation, frequency-domain decoupled intra-scale interaction, and spatial-prior-constrained cross-scale fusion for end-to-end bounding-box detection. It improves accuracy, missed-detection control, false-alarm suppression, and center localization while reducing complexity and maintaining efficient inference. An accuracy-efficiency balance is achieved in complex infrared scenes. Future work will investigate multi-frame spatiotemporal modeling for thermal-crossover backgrounds and dense dim-target scenes.
Lightweight image-to-image Steganography Based on Improved Emd and Dual-domain Graph Convolutional Network
DUAN Xintao, CHEN Rusheng, LI Sen, QIN Chuan
Available online  , doi: 10.11999/JEIT260857
Abstract:
  Objective  Image steganography embeds secret information into a cover image to achieve secure transmission, and it serves as an important technique in confidential communication and privacy protection. With the development of deep learning, learning-based steganography has notably improved embedding capacity and reconstruction quality. However, three requirements, namely high steganographic performance, strong resistance to steganalysis, and lightweight design, are difficult to satisfy simultaneously, and this trade-off has become the main bottleneck for practical deployment. Single-domain processing methods cannot balance the three objectives, whereas existing high-performance models are structurally complex and hard to deploy in resource-constrained environments. Moreover, mainstream dual-domain schemes usually assume that the embedding distortion follows a continuous distribution in both the spatial domain and the frequency domain, yet actual steganographic modifications tend to concentrate in the discontinuous and complex texture regions of the cover image. To address these problems, a lightweight image steganography network is designed in this study to jointly optimize the embedding and extraction paths, so that the stego-image quality and the anti-steganalysis ability are improved while a low computational cost and a low inference latency are maintained.  Methods  A lightweight steganographic network named GISNet is proposed, in which an encoder-decoder architecture with skip connections is adopted and the hiding network and the extraction network are made structurally symmetric without weight sharing. First, an Improved Bidimensional Empirical Mode Decomposition (IBEMD) module is applied, by which the secret image is adaptively decomposed into several intrinsic mode components and one residual component, so that the hidden information is dispersed hierarchically among components of different morphology. Gaussian blur is used to replace extremum interpolation for estimating the local mean envelope, the separability of the Gaussian kernel is exploited to reduce the computational complexity, and a parameter-binding mechanism is introduced to guarantee deterministic and reversible reconstruction. Next, a multi-scale spatial-frequency block (IMFB) is designed, in which multi-branch dilated convolutions, an attention mechanism, dynamic gating, and a frequency-domain perception unit are integrated, so that the spatial features and the frequency features are deeply fused, the anti-detection ability is enhanced, and the high-fidelity extraction of the secret information is ensured. Finally, an improved graph fusion neural network (GFNN) is employed as the bottleneck layer, in which a sparse graph is constructed in the feature space through the K-nearest-neighbor algorithm and messages are propagated only among non-local nodes with high similarity, so that the long-range pixel associations are explicitly modeled, the modification patterns in discontinuous texture regions are characterized, and the model complexity is substantially reduced. The hiding network and the extraction network are jointly optimized by a four-term loss function that combines the hiding loss, the restriction loss, the Laplacian pyramid loss, and the perceptual loss.  Results and Discussions  GISNet achieves the best overall image-hiding and recovery performance on DIV2K, COCO, and ImageNet, demonstrating high reconstruction quality and stable cross-dataset generalization (表1). On DIV2K, the PSNR values reach 58.07 dB for cover/stego image pairs and 60.10 dB for secret/recovered-secret image pairs. The cover and stego images are visually indistinguishable, and the residual maps remain nearly black after 30-fold magnification (图5). Ablation experiments show that replacing DWT with IBEMD improves the PSNR values by 8.04 dB and 5.57 dB, respectively (表2). The complete combination of IBEMD, IMFB, and GFNN provides the best results, confirming the complementary effects of hierarchical information dispersion, multiscale spatial-frequency mapping, and nonlocal feature association (表3). Multi-dilation-rate branches and joint spatial-frequency processing further improve hiding quality and recovery accuracy (表4,表5). The detection accuracies under four steganalysis methods range from 49.35% to 49.85%, indicating that the stego images are difficult to distinguish from natural cover images (表6). GISNet requires only 8.00 M parameters and 7.88 GFLOPs, with an inference time of 76 ms (表7). These results demonstrate that GISNet effectively balances visual quality, recovery accuracy, resistance to steganalysis, and computational efficiency.  Conclusions  A lightweight dual-domain graph convolutional network for image steganography, named GISNet, is proposed in this paper. The experimental results demonstrate the following. (1) The improved bidimensional empirical mode decomposition disperses the secret information hierarchically and supports deterministic and reversible reconstruction, by which a reliable basis is provided for high-quality hiding and recovery. (2) The multi-scale spatial-frequency block and the graph fusion neural network jointly improve the stego-image quality and the anti-steganalysis ability, so that the stego images can resist detection by multiple steganalysis tools. (3) The lightweight design, which is based on depth-wise separable convolutions and sparse graph construction, significantly reduces the model complexity and the computational cost, by which the model is made suitable for deployment in resource-constrained environments. Future work will focus on the robustness against channel interference and the extension to multi-image and cross-modal steganography, so that the security and applicability of the scheme are further enhanced.
A Structure-Preserving Semantic Transmission Method for Low-Bandwidth Networks
HU Tianwei, ZHANG Xiangrui, CHEN Jian, DUAN Haodong, JIA Jie
Available online  , doi: 10.11999/JEIT260525
Abstract:
  Objective  Low-bandwidth visual transmission is essential for edge-intelligence applications such as disaster inspection, underwater exploration, and remote assistance. These scenarios require visual communication systems to simultaneously achieve low bitrate, high structural fidelity, stable color reconstruction, and low end-to-end latency. However, conventional image coding methods suffer from severe quality degradation at extremely low bitrates due to the digital-cliff effect and compression artifacts. Although semantic communication provides a promising solution by transmitting task-relevant representations rather than pixel-level information, existing approaches still face challenges in balancing bitrate efficiency, reconstruction fidelity, and decoding complexity. Therefore, a structure-preserving semantic transmission method, termed Edge-Link, is proposed for low-bandwidth networks.  Methods  Edge-Link adopts a multimodal decoupled representation framework that separates image information into three complementary streams: a semantic stream, a structural stream, and a low-frequency appearance stream. The semantic stream is extracted using a frozen CLIP encoder to provide global semantic guidance, while the structural stream is obtained from Canny edge information to explicitly preserve object boundaries. A low-resolution color map is introduced as the appearance stream to maintain global color distribution and illumination characteristics with minimal transmission overhead. Furthermore, a FastSAM-based semantic gateway is developed to distinguish foreground objects from background regions, and an object-aware bitrate allocation strategy is designed to prioritize important semantic regions under bandwidth constraints. At the receiver, a deterministic dual-stream generation network based on SPADE is proposed, where structural information provides spatial constraints and semantic features guide the reconstruction process, avoiding the high latency caused by iterative diffusion sampling. A real LoRa communication prototype based on GNU Radio and USRP is also implemented to validate transmission feasibility and robustness under practical wireless conditions.  Results and Discussions  Experiments are conducted on Set14 and DIV2K to evaluate transmission efficiency, reconstruction quality, perceptual naturalness, and latency. The proposed method achieves an average payload of 10.96 KB for 512×512 images, corresponding to about 0.31 bpp, while keeping the end-to-end inference latency below 60 ms; under a LoRa narrowband link of about 8 kbps, the airtime of a single frame is about 11 s, verifying its feasibility in bandwidth-limited environments (Table 1). The ablation study shows that the object-aware coding strategy improves the bitrate-quality trade-off by reducing the average payload on DIV2K from 126.09 KB to 93.55 KB while preserving better foreground reconstruction quality (Table 2). Compared with JPEG, the proposed method also shows better structural preservation and perceptual quality at low bitrates; for example, on image 0843 with a payload of 15.54 KB, the global SSIM is improved from 0.783 to 0.963, and the subject-region LPIPS is reduced from 0.561 to 0.307 (Table 4). Under the same bitrate constraint, the proposed method further outperforms VQGAN and ControlNet+Canny, achieving PSNR, SSIM, LPIPS, and NIQE values of 23.738, 0.622, 0.223, and 4.570, respectively, indicating a better balance between fidelity and perceptual quality in low-bandwidth semantic reconstruction (Table 5).  Conclusions  A structure-preserving semantic transmission framework for low-bandwidth networks is presented. By combining multimodal decoupled representation, object-aware bitrate allocation, and deterministic dual-stream reconstruction, the framework balances transmission efficiency, structural fidelity, perceptual quality, and decoding latency. The reported results show that the proposed method is well suited to high-reliability edge visual communication scenarios in which accurate contours, stable color appearance, and efficient inference are simultaneously required. The real-link validation on a LoRa prototype further suggests its practical potential for bandwidth-constrained edge networks.
Anomaly Detection on Irregular Signals in Adaptive Decay Reservoir Network Model Space
CHEN Ao, LI Wenpeng, XIE Xiaoyan, CHEN Pengpeng
Available online  , doi: 10.11999/JEIT260423
Abstract:
  Objective  Signals acquired from industrial systems often exhibit irregular sampling due to sensor instability, intermittent operation, communication dropout, and multi-source asynchronous acquisition. This irregularity violates the uniform sampling assumption of most signal analysis methods and challenges anomaly detection. Interpolation and resampling may distort the underlying temporal dynamics, while continuous-time models based on neural ordinary differential equations as well as time-aware Transformers suffer from high training costs and strong dependence on large training sets, making them impractical under limited training resources. Model space learning offers an alternative by fitting each signal with a dynamic model and analyzing the fitted models instead of the raw signals. However, existing model space methods for irregular sampling rely on fixed reservoir configurations and lack adaptive optimization of the model space. This paper aims to develop an anomaly detection framework for irregularly sampled signals that is simultaneously robust to non-uniform intervals and efficient to train.  Methods  An adaptive model space learning framework based on the Adaptive Decay Reservoir Network (ADRN) is proposed (Fig. 1). ADRN extends the echo state network by introducing an exponential decay mechanism derived from a linear ordinary differential equation (Fig. 2). Between consecutive observations, the hidden state decays according to a learnable decay rate over the actual elapsed time, so that the state update naturally adapts to non-uniform sampling intervals without numerical ordinary differential equation solvers. Each new observation then updates the decayed state through a nonlinear activation. Ridge regression with a closed-form solution then fits a readout model mapping hidden states to the original signal, and the fitted readout weights serve as a compact fixed-dimensional representation regardless of signal length. The model space is further optimized by two complementary losses. A time-interval-weighted reconstruction loss assigns higher weights to larger intervals, preventing densely sampled segments from dominating the optimization and improving fitting quality under non-uniform sampling. A separability loss inspired by Fisher discriminant analysis acts through a learnable projection matrix to minimize intra-class scatter and maximize inter-class separation in the projected model space. The two losses are combined into a joint objective that simultaneously updates the reservoir parameters, the decay rate, and the projection matrix, and gradients propagate through the differentiable closed-form ridge regression to enable end-to-end optimization. A downstream classifier, namely a support vector machine on CWRU and SU and a random forest on the higher-dimensional TEP model space, performs the final detection. The echo state property of ADRN is formally established, and the resulting spectral-norm condition is more relaxed than the classical one, allowing richer reservoir dynamics.  Results and Discussions  Experiments cover the CWRU bearing dataset (five subsets, 50% missing rate), the SU gearbox dataset (30%, 50%, and 70% missing rates), and the Tennessee Eastman Process (TEP) chemical dataset (19 classes, three missing rates), with only 200 labeled signals per subset for training on CWRU and SU. The proposed method achieves the highest accuracy in 7 of the 8 CWRU and SU settings (Table 1), with accuracies ranging from 87.3% to 93.8% on CWRU. On SU, it maintains 94.4% accuracy even at a 70% missing rate, and the fluctuation across missing rates is only 2.8%, in contrast to 18.7% for ODE-RNN, while Neural CDE drops from 86.8% to 62.6%. Ablation studies confirm the contribution of each component (Table 1, Table 2). Removing the exponential decay reduces accuracy by up to 28.4 percentage points, and the interval weighting and the separability loss contribute complementary gains of 2.0 and 3.1 percentage points on SU at the 70% missing rate. t-SNE visualization shows that the optimized model space exhibits compact and clearly separated classes (Fig. 3). Training on a CWRU dataset completes in about 50 seconds, over two orders of magnitude faster than neural ordinary differential equation methods, which require 3000 to 7000 seconds (Table 3). Hyperparameter analysis indicates that a loss balance coefficient between 0.1 and 0.3 performs well and that a small reservoir suffices (Table 4). On TEP, a non-rotating-machinery industrial object whose faults manifest as changes in process dynamics, the proposed method attains the highest accuracy among all compared methods at every missing rate, with the largest margin of 7.0 percentage points at the highest missing rate (Table 5).  Conclusions  The ADRN based model space learning framework provides an effective solution for anomaly detection on irregularly sampled signals. The exponential decay mechanism enables interval-aware state updates without numerical ordinary differential equation solvers, ridge regression yields efficient closed-form readout fitting, and the joint optimization of the reconstruction and separability losses produces a model space with both high fitting quality and strong discriminative structure. The framework requires only a small reservoir of 10 to 50 dimensions and completes training within one minute, making it well suited to scenarios with limited training resources and irregular sampling. Future work includes adaptive reservoir sizing, extension to multivariate joint modeling, and validation in further domains such as structural health monitoring and biomedical or meteorological time series.
An ECO Repair Method for Max Transition Violations in Multi-Load Nets
FAN Lingyan, XU Xinchen, HUANG Cankun, SHEN Zhengnuo, DENG Jiangxia, LIU Hailuan
Available online  , doi: 10.11999/JEIT260600
Abstract:
  Objective   Max transition violations in digital integrated circuit physical design may increase gate delay, reduce timing margin, and introduce additional power-consumption and signal-integrity risks. With technology scaling and the increasing complexity of high-performance SoC and CPU designs, long interconnects and heavy effective loads make such violations more prominent in post-routing optimization and ECO stages. Conventional repair methods usually rely on manual analysis or global heuristics, such as total-wire-length-ratio-based buffer insertion, which may lead to low repair efficiency, inaccurate repair targeting, and redundant buffer insertion in multi-load nets. Although these approaches can alleviate some violations in simple cases, they often fail to accurately identify the real violating branches in multi-load nets. As a result, repair targeting becomes insufficient and redundant buffer insertion is likely to occur. To address these limitations, a path-level buffer insertion method is proposed for max transition violation repair in multi-load nets.  Methods   The proposed method first reconstructs the physical topology of the target net from routing information extracted from the Design Exchange Format (DEF) file. Routing endpoints, turning points, and via connection points are abstracted as physical nodes with coordinate and metal-layer attributes, and a weighted undirected graph is established to preserve branch structures and cross-layer connectivity (Fig. 3). To reduce the influence of small coordinate deviations, a spatial-tolerance-based node merging strategy is introduced during graph construction. Since the violation coordinates reported by timing analysis tools may not lie exactly on valid routed segments, a vector-projection-based coordinate snapping strategy is then adopted to align logical violation coordinates with the actual physical topology (Fig. 4). After endpoint binding, the actual physical propagation path from the driver to each violating load is recovered by Dijkstra shortest-path search. Based on the Elmore-model intuition that inserting a buffer near the midpoint of a long interconnect can effectively segment the distributed RC load, the midpoint of each recovered path is selected as the initial candidate insertion point. The exact insertion coordinate is obtained by accumulating segment lengths along the path and interpolating on the segment where half of the total path length is reached. To improve robustness, the initial candidate point is further expanded into an effective candidate interval with a spatial tolerance factor. For multi-load nets, different violating branches may share long common physical segments. Therefore, an interval-intersection-based shared-buffer optimization strategy is introduced to merge overlapping candidate intervals into shared insertion regions, thereby reducing redundant buffer insertion (Fig. 5-Fig. 7).  Results and Discussions   Experiments are conducted on five designs, including CPU, SAS, RAID, PCIe, and HBA, implemented in the UMC 28 nm process. Synopsys IC Compiler II is used for physical implementation, and PrimeTime is used for timing analysis. A violation-margin threshold of -8 ps is adopted, and only paths below this threshold are included in the repair and evaluation. The complete ECO flow retains the existing repair method for single-load nets and applies the proposed path-level method to multi-load nets. To illustrate the repair mechanism, a representative multi-load violating net is selected for detailed analysis. The net contains one driver and fourteen loads, among which thirteen violating load paths share a long common routed trunk and exhibit max transition violation margins ranging from -75.2 ps to -38.3 ps. If these paths are repaired independently, thirteen nearby buffer insertion demands are generated on the shared trunk. After interval-intersection-based merging, one shared buffer is sufficient to repair all thirteen violating paths jointly (Fig. 8). This case study indicates that path recovery improves repair targeting, while shared-buffer optimization improves resource efficiency. Across the five designs, the complete ECO flow reduces the total number of max transition violations from 10,616 to 108, corresponding to an overall repair rate of 98.98% (Table 1). The CPU, SAS, and RAID designs are completely repaired, while only 17 and 91 violations remain in PCIe and HBA, respectively. Among the 8,216 multi-load violating paths, 8,166 are successfully repaired, corresponding to a multi-load violation repair rate of 99.39%. If these successfully repaired paths were handled independently, 8,166 buffer insertions would be required. After shared-buffer optimization, only 1,548 buffers are inserted, giving an overall compression ratio of 81.04%. The setup worst negative slack, total negative slack, and number of violating paths remain generally stable before and after repair, with only minor fluctuations observed in individual designs (Table 2). These results show that the repair flow does not cause obvious degradation in setup timing quality. Further analysis indicates that the residual violations in PCIe and HBA mainly occur when candidate buffer locations fall inside standard-cell placement blockages around hard macros or IP cores. Metal routing is allowed through these regions, but buffers cannot be legally placed, revealing a current limitation in the physical-feasibility handling of candidate insertion locations.  Conclusions   A path-level buffer insertion method is proposed for max transition violation repair in multi-load nets. By combining physical topology reconstruction, violation-coordinate snapping, shortest-path-based path recovery, midpoint-guided candidate generation, and interval-intersection-based shared-buffer optimization, the proposed method improves repair targeting and reduces redundant buffer insertion. Experimental results on five designs show that the complete ECO flow reduces the number of max transition violations from 10,616 to 108, achieving an overall repair rate of 98.98%. Among the 8,216 multi-load violating paths, 8,166 are successfully repaired, corresponding to a repair rate of 99.39%. For these successfully repaired paths, shared-buffer optimization reduces the number of buffer insertion demands from 8,166 to 1,548, corresponding to a compression ratio of 81.04%, without causing obvious degradation in setup timing quality. The remaining violations are mainly associated with standard-cell placement blockages around hard macros or IP cores, indicating that the physical-feasibility handling of candidate insertion locations should be further improved.
Design of a Channel-Adaptive Denoiser for Digital Semantic Communications
WU Yanjun, LIU Zhangyuhang, YANG Wenxin, YAN Mubiao, ZHOU Hao, ZHAO Yajun, XIE Zhuochen, LIANG Xuwen
Available online  , doi: 10.11999/JEIT260523
Abstract:
  Objective  Practical semantic communication should simultaneously satisfy two requirements: compatibility with existing digital communication infrastructures and robustness under varying channel conditions. Semantic-oriented modulation (SOM) provides a feasible way to map continuous semantic features into layered digital constellation symbols, thereby making semantic transmission compatible with conventional digital systems. However, the digitization process also introduces structured quantization distortion, which makes receiver-side recovery more difficult than in continuous semantic transmission. Although diffusion models have shown strong capability in channel-adaptive semantic recovery, directly applying them to SOM-based digital semantic communication is still limited by the structured distortion introduced by SOM. Therefore, this paper focuses on channel-adaptive receiver design for digital semantic communication and investigates how to compensate SOM-induced structured distortion before subsequent recovery.  Methods  An SOM-based digital semantic communication system for image transmission over an additive white Gaussian noise (AWGN) channel is considered. The proposed receiver adopts a two-stage structure composed of a Quantization Noise Predictor (QNP) and a diffusion recovery module. In the first stage, QNP estimates and compensates the structured quantization distortion introduced by SOM from the layer-wise soft received symbols. In the second stage, the compensated semantic representation is further refined by a diffusion denoiser, whose inference step number is adaptively selected according to the estimated signal-to-noise ratio (SNR). The QNP includes a shared feature extraction frontend, a classification branch exploiting discrete SOM constellation priors, and a regression branch performing fine-grained continuous distortion compensation. A Feature-wise Linear Modulation (FiLM) mechanism is used to incorporate SOM parameters and channel-state information, so that the same QNP can adapt to different modulation configurations and channel conditions. In addition, a composite loss with classification loss, regression loss, and distribution regularization is designed to improve the statistical properties of the compensated residual noise.  Results and Discussions  Experiments are conducted on the CLIC dataset using PSNR and MS-SSIM. First, the proposed method is compared with VAE, VAE+Diff, VAE+SOM, VAE+SOM+QNP, VAE+SOM+Diff, and JCM. The results show that direct SOM-based digitization causes noticeable performance degradation, while the proposed method consistently improves reconstruction quality over digital semantic baselines. In particular, QNP alone already provides stable gains over the SOM-only receiver, indicating that its effectiveness does not rely on diffusion recovery itself. Moreover, VAE+SOM+QNP achieves performance close to VAE+SOM+Diff while requiring much lower computational cost, and combining QNP with diffusion yields the best overall performance. Second, two training strategies, namely independent QNP training and diffusion-assisted fine-tuning, are compared. The results show that diffusion-assisted fine-tuning provides only limited additional gains but significantly increases training cost and complexity, so independent training offers a more practical balance. Third, experiments under different SOM configurations and different SNR conditions verify that QNP provides stable gains across different modulation orders and SOM layer settings. Latency analysis further shows that QNP introduces only a small fixed overhead, whereas the diffusion module dominates the total inference time; therefore, the adaptive diffusion-step schedule is selected according to the measured latency-PSNR trade-off. Fourth, Gaussianity analysis based on the Kullback-Leibler divergence and Wasserstein distance shows that QNP compensation significantly improves the Gaussianity of the residual noise, while the version with distribution regularization achieves the best statistical consistency.  Conclusions  This paper proposes a channel-adaptive receiver for digital semantic communication, in which QNP-based front-end compensation is combined with diffusion-based semantic recovery. The main contribution lies in introducing a lightweight and independently effective QNP module to compensate SOM-induced structured quantization distortion before subsequent recovery. Experimental results show that QNP alone can already stably improve digital semantic reconstruction under different SNR conditions and different SOM configurations, while its combination with diffusion recovery yields the best overall performance. Therefore, the proposed method provides an effective way to improve semantic reconstruction quality and channel adaptability while preserving compatibility with existing digital communication infrastructures.
Intelligent Resource Allocation Algorithm Based on Outdated CSI for Multi-Node URLLC
ZHAO Yizhen, GAO Wei, HU Yulin, ZHU Yao
Available online  , doi: 10.11999/JEIT260216
Abstract:
  Objective  Ultra-Reliable and Low-Latency Communications (URLLC) is widely used in Industrial Internet of Things (IIoT) systems. However, in mobile industrial scenarios such as transportation and inspection, instantaneous Channel State Information (CSI) is difficult to obtain because of feedback overhead. Resource allocation decisions therefore need to be made using outdated CSI. This mismatch restricts system energy efficiency. Traditional convex optimization methods have difficulty addressing this problem. Classical Deep Reinforcement Learning (DRL) algorithms also have limited convergence stability and policy performance under the stringent latency and reliability constraints of URLLC. To address these challenges, this paper considers a multi-node URLLC system under outdated CSI in dynamic scenarios. An energy-efficiency maximization problem is formulated under the Finite BlockLength (FBL) regime, with communication latency and reliability constraints. An efficient and stable algorithm is then designed for joint power and blocklength allocation.  Methods  A Successive Convex Approximation (SCA)-assisted DRL framework is proposed to maximize energy efficiency under outdated CSI. First, an SCA-based algorithm is developed to obtain a pre-allocation solution for transmit power and blocklength. This solution is feasible and physically interpretable, but relatively conservative. Based on this baseline, a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is used for incremental refinement through interaction with the dynamic environment. This process reduces the conservatism of SCA. The SCA solution is used as prior knowledge in the state representation. Node location information is also incorporated into the state space. These designs narrow the policy search space and enable the DRL agent to better capture large-scale channel characteristics and system dynamics under outdated CSI. Learning efficiency and stability are therefore improved.  Results and Discussions  The proposed algorithm is evaluated through simulations and compared with three benchmark algorithms: an SCA-based optimization algorithm, a TD3 algorithm without SCA guidance, and a TD3 algorithm without node location information. The results show that the proposed method outperforms all benchmarks in convergence stability and system energy efficiency. In the training phase (Fig. 3), the average reward of the proposed algorithm increases steadily and converges stably. By contrast, removing node location information leads to lower rewards and stronger fluctuations. Removing SCA guidance causes the algorithm to converge to a much lower reward level. These results confirm the roles of SCA-based prior guidance and location-aware state representation in improving training stability. In the actual operation stage (Fig. 4), the proposed algorithm achieves high and stable energy efficiency and outperforms all comparison algorithms. Under outdated CSI, DRL-based methods can obtain higher energy efficiency than conservative optimization methods when transmission succeeds. However, removing node location information reduces energy efficiency, and removing SCA guidance increases transmission failures. These results verify the effectiveness of both designs in improving energy efficiency and maintaining policy feasibility. The effects of key system parameters are also examined. For basic resource parameters, a moderate increase in the blocklength budget (Fig. 5) or power budget (Fig. 6) improves system energy efficiency. For reliability constraints (Fig. 7), the reliability requirement should be set according to service requirements to avoid resource waste. Finally, the average energy efficiency under different numbers of nodes and different numbers of neurons in the TD3 network is analyzed (Fig. 8). The results provide guidance for algorithm configuration and network-scale design.  Conclusions  This paper addresses energy-efficient resource allocation for multi-node URLLC systems with outdated CSI by integrating SCA and DRL. In the proposed framework, a TD3-based DRL algorithm is guided by an SCA reference solution, and node location information is incorporated into the state representation. This optimization-learning dual-driven framework combines the interpretability and feasibility of model-based optimization with the adaptivity of data-driven learning. Simulation results show that the proposed method achieves higher energy efficiency than SCA-based optimization and conventional TD3 while satisfying URLLC latency and reliability constraints. The SCA reference solution improves policy stability and effectiveness under outdated CSI. Node location information further supports efficient decision-making. This work focuses on a single-cell multi-node scenario under Time Division Multiple Access (TDMA). Practical issues such as multi-cell interference, cooperative scheduling among multiple base stations, and more complex mobility patterns are not considered. Future work will extend the proposed framework to multi-cell and multi-agent scenarios and test its applicability under more severe CSI imperfections.
A Cryptographic Side-Channel Security Modeling and Formal Verification Method
WANG Xingxin, HU Wei, HUANG Xuan, LIN Chenyu, ZHOU Yi
Available online  , doi: 10.11999/JEIT260631
Abstract:
  Objective  Compared with post-silicon side-channel security analysis, pre-silicon side-channel security verification during the design phase enables the earlier identification of potential side-channel security vulnerabilities in cryptographic core designs, thereby effectively reducing the cost and time of post-silicon remediation. However, most existing pre-silicon side-channel security assessment approaches rely on data-driven statistical analysis or artificial intelligence techniques and require complex calculations on large amounts of data to mitigate the impact of insufficient coverage on the assessment results. In addition, existing methods typically adopt independent modeling strategies for different types of side channels, lacking a unified side-channel security modeling approach. A cryptographic side-channel security modeling and formal verification method is proposed, supporting unified and automated modeling of different types of side channels by constructing a side-channel security model. The method can identify potential timing side-channel, power side-channel and fault injection vulnerabilities in cryptographic core designs, and analyze the effectiveness of side-channel countermeasures based on side-channel security property checking.  Methods  The proposed cryptographic side-channel security modeling and formal verification method includes side-channel security model construction, side-channel security property extraction, and side-channel security verification. The side-channel security model uses information flow analysis to characterize timing side-channel leakage, power side-channel leakage, and fault propagation behavior in cryptographic core designs, providing an effective mathematical model for side-channel security verification. Specifically, the side-channel security model utilizes changes in signal labels to analyze information flows during the encryption process by assigning a label to a signal bit and defining label propagation rules. Side-channel security properties formally describe the behavioral characteristics of side-channel leakage, including timing properties, power properties, and fault properties, providing theoretical support for side-channel security verification. Side-channel security verification uses the extracted security properties as verification constraints and employs formal verification tools to identify potential timing side-channel, power side-channel, and fault injection vulnerabilities in cryptographic core designs. Furthermore, the method can analyze the effectiveness of masking and fault injection countermeasures against side-channel vulnerabilities.  Results and Discussions  The proposed side-channel security verification method utilizes formal verification techniques to accurately identify potential side-channel security vulnerabilities in various block cipher core designs, and evaluate the effectiveness of side-channel countermeasures based on side-channel security property constraints. The timing side-channel verification results demonstrate that the proposed method can accurately identify timing side-channel security vulnerabilities in AES, SM4, LED, PRESENT, and IDEA core designs within 20s (Table 2, Fig. 7). No timing side-channel vulnerabilities are identified in the other cryptographic core designs, except for IDEA, which exhibits timing side-channel vulnerabilities caused by modular multiplication operations. The power side-channel security verification results show that formal checks based on controllability property, key–power distinguishability coupling property and key–power nonlinear coupling property can accurately identify target modules with potential power side-channel security vulnerabilities in AES, SM4, LED and PRESENT within 1 minute (Table 3). The key expansion module in cryptographic core designs does not cause key leakage through key-dependent power consumption, as it fails to satisfy the controllability property. In addition, the experimental results indicate that the masking protection in the RSM core design can prevent the correct key from being distinguished through random masking (Fig. 8). The fault injection security verification results for three AES core designs with infective countermeasures demonstrate that the proposed method can analyze the effectiveness of fault infection countermeasures. The results show that the infection countermeasure requires not only altering the fault propagation path but also disrupting the algebraic relationships among faults (Table 4, Fig. 10).  Conclusions  This paper proposes a cryptographic side-channel security modeling and formal verification method to address the lack of formal mathematical models and the limited completeness of existing data-driven side-channel security assessment methods. The proposed method first achieves unified modeling of timing leakage, power leakage, and fault propagation behaviors from the perspective of information flow analysis. Based on the constructed side-channel security model, side-channel security properties are extracted to formally characterize the behavioral features of side-channel information leakage and propagation during the encryption process. Potential side-channel security vulnerabilities in cryptographic core designs are then identified through formal checking using the extracted security properties as constraints. The proposed method provides an effective solution for the unified modeling and formal security verification of different types of side channels. Experimental results obtained from the side-channel security verification of various block cryptographic core designs demonstrate that: (1) the proposed method can uniformly model timing side-channel leakage, power side-channel leakage, and fault propagation behaviors in cryptographic core designs; (2) the proposed method can accurately identify timing side-channel vulnerabilities, power side-channel vulnerabilities, and fault injection vulnerabilities in cryptographic core designs, including AES, SM4, IDEA, LED and PRESENT; (3) the proposed method can analyze the effectiveness of masking and fault infection countermeasures. However, this study only qualitatively identifies side-channel security vulnerabilities in cryptographic core designs; pre-silicon quantitative assessment of side-channel leakage should be investigated in future work.
A Lightweight Dual-Stream Convolutional Network Feature Fusion Method for UAV RF Recognition
DONG Pengyu, XIANG Xin, LV Siting, LIANG Yuan, WANG Rui, MAO Hu
Available online  , doi: 10.11999/JEIT260464
Abstract:
  Objective  With the rapid proliferation of Unmanned Aerial Vehicles (UAVs) and the escalating demand for airspace security, radio frequency (RF) fingerprint recognition has emerged as a pivotal technology for identifying non-cooperative UAVs. However, existing methods grapple with significant challenges, including poor robustness in low signal-to-noise ratio (SNR) environments and prohibitive computational complexity, which severely hinder their deployment on resource-constrained tactical edge devices. To address these critical limitations, this paper proposes a novel lightweight dual-stream convolutional network tailored for UAV RF recognition. This network is designed to extract static spectral texture features and dynamic temporal gradient features in parallel, complemented by a meticulously crafted lightweight feature fusion strategy.  Methods  The proposed network architecture is ingeniously designed to process RF signals. The input signal undergoes a Short-Time Fourier Transform (STFT) to generate a two-dimensional spectrogram, which serves as the primary input. The network is bifurcated into two parallel streams: a static stream and a dynamic stream. The static stream is engineered to capture the inherent static spectral patterns and energy distributions within the STFT spectrogram. It comprises a series of stacked convolutional blocks, each integrating convolutional layers, batch normalization, and ReLU activation functions, followed by max-pooling layers to progressively downsample the feature maps and increase the channel depth. Conversely, the dynamic stream is dedicated to enhancing feature discriminability, particularly in low-SNR scenarios. It begins by computing the temporal gradient of the input spectrogram, effectively suppressing static background noise and accentuating dynamic signal variations. This gradient map is then processed by a symmetric set of convolutional blocks, mirroring the structure of the static stream. To maintain model efficiency, an element-wise addition fusion strategy is employed to integrate the features from both streams, ensuring a balance between feature complementarity and computational overhead. The fused features are subsequently fed into a classification head, consisting of an adaptive average pooling layer, dropout layers for regularization, and fully connected layers to produce the final classification output. Extensive experiments are conducted on the publicly available DroneRF dataset, encompassing ablation studies to dissect the contribution of each component, comparative analyses of various fusion strategies, and rigorous evaluations of the model’s lightweight characteristics.  Results and Discussions  The experimental results unequivocally demonstrate the efficacy of the proposed method. The dual-stream network achieves a remarkable 95.65% accuracy on the test set, representing a substantial 5 percentage point improvement over the best-performing single-stream network. A critical analysis reveals that the temporal gradient operation contributes significantly to this enhancement by improving the average SNR by 1 dB, thereby bolstering feature discriminability in challenging low-SNR environments. Furthermore, the model’s lightweight design is a standout feature, with a mere 0.58 million parameters, making it eminently suitable for deployment on tactical edge devices. Ablation studies and feature visualization analyses provide compelling evidence for the complementary nature of static and dynamic features. The static stream adeptly captures broad spectral contours, while the dynamic stream focuses on fine-grained temporal variations. The element-wise addition fusion strategy proves superior, outperforming other approaches like feature concatenation and attention-based fusion in terms of both performance and computational efficiency, thereby validating the rationale behind the lightweight design.  Conclusions  This paper presents a comprehensive solution to the challenges of UAV RF recognition in complex environments by proposing a lightweight dual-stream convolutional network. The method effectively enhances recognition accuracy and robustness through the synergistic combination of dual-stream feature extraction and the SNR-enhancing properties of temporal gradient features, all while maintaining a lightweight architecture suitable for edge deployment. The proposed approach offers a significant advancement in the field, providing a robust and efficient solution for UAV identification. Future research endeavors will focus on further enhancing the model’s adaptability to complex electromagnetic environments, incorporating the effects of sensor noise, and extending the framework to multi-UAV cooperative scenarios.
Complex-domain Joint Spectrum Sensing Method for UAV Swarms in Complex Electromagnetic Environments
QIAN Hui, CHEN Li, YIN Huarui, WANG Weidong
Available online  , doi: 10.11999/JEIT260499
Abstract:
  Objective  The rapid development of the low-altitude economy is increasing the use of unmanned aerial vehicle (UAV) swarms in emergency communication, urban logistics, reconnaissance, and low-altitude network coverage. These applications require reliable spectrum awareness to support cooperative communication, dynamic spectrum access, and interference avoidance. However, low-altitude electromagnetic environments often contain low-SNR signals, multipath propagation, non-cooperative interference, and multiple coexisting transmissions. Conventional energy, cyclostationary-feature, and matched-filter detectors are sensitive to noise uncertainty, computational cost, or prior waveform knowledge. Learning-based methods can improve robustness, but many rely on power spectral density (PSD) or short-time Fourier transform (STFT) representations. These representations may weaken phase information or require costly two-dimensional time-frequency preprocessing. Existing methods also focus mainly on spectrum occupancy and provide limited information about overlapping transmissions. This study therefore develops a low-latency joint sensing method that preserves magnitude and phase information while estimating spectrum occupancy and interference overlap for each frequency bin. The method is intended for local spectrum sensing at UAV nodes under resource and latency constraints.  Methods  The proposed RadioSEUnet pipeline contains two stages: magnitude-phase feature construction and joint spectrum-state estimation. First, each complex baseband in-phase/quadrature (I/Q) sequence is multiplied by a Hann window and transformed using a one-dimensional fast Fourier transform (FFT). The resulting complex spectrum is decomposed into logarithmic magnitude and phase components. The two components are normalized separately and stacked as a two-channel feature tensor. Binary labels indicate spectrum occupancy and multi-signal overlap at each frequency bin. RadioSEUnet adopts a U-shaped encoder-bottleneck-decoder architecture with four encoder stages containing 64, 128, 256, and 512 channels. Each RadioSEBlock combines a complex-parameterized convolution, squeeze-and-excitation channel attention, and a residual connection. The convolution couples the magnitude and phase feature streams through constrained cross-channel operations. Two prediction heads convert the shared representation into a spectrum-occupancy probability mask and an interference-overlap probability mask. The model is optimized using an equally weighted sum of two binary cross-entropy losses. Training uses AdamW, cosine-annealing learning-rate scheduling, early stopping, a batch size of 64, and at most 200 epochs. The complete data collection contains 72,000 complex I/Q records, including 48,000 simulated records and 24,000 measured records. The simulated subset covers Wi-Fi, BLE, ZigBee, LoRa, QPSK/16QAM, FM, and AM signals. Signal-to-noise ratios range from –15 dB to 10 dB under additive white Gaussian noise and Rayleigh fading. The measured subset was collected using a USRP N310 in the 2.4–2.5 GHz ISM band at 100 MS/s over a 1 ms observation interval. The controlled quantitative evaluation uses an 8:1:1 split of the simulated subset. A separate simulated-to-measured protocol is defined in the main text to examine cross-domain generalization. RadioSEUnet is compared with six PSD- or STFT-based baselines under matched data splits and hardware conditions. Performance is measured using intersection over union (IoU), precision, recall, preprocessing time, inference time, and total sensing latency.  Results and Discussions  The SNR-dependent quantitative results reported here are obtained using the controlled simulated-data protocol. At -15 dB, RadioSEUnet achieves an IoU of 0.768 and a recall of 0.846 for spectrum occupancy detection. Compared with the second-best STFT-RADN baseline, these values correspond to absolute improvements of 0.186 and 0.166, respectively. For interference-overlap detection, RadioSEUnet achieves an IoU of 0.456 and a precision of 0.768 at –15 dB. The corresponding improvements over STFT-RADN are 0.246 and 0.275. The lower IoU for interference-overlap detection indicates that weak overlap boundaries remain difficult to separate from strong-signal sidelobes and background noise. The latency evaluation is conducted on the workstation specified in the main text. Magnitude-phase preprocessing requires 21.04 ms, and network inference requires 4.77 ms, producing a total sensing latency of approximately 25.81 ms. STFT-YOLOv3 requires 98.9 ms under the same hardware setting, so the proposed pipeline is approximately 3.8 times faster in this comparison. Ablation experiments show that magnitude-phase preprocessing, complex-parameterized feature coupling, and channel attention each improve low-SNR sensing performance. Removing the magnitude-phase preprocessing produces the largest degradation. These results indicate that preserving complementary magnitude and phase information is useful for weak-signal and interference-overlap detection. They do not, however, establish performance on airborne hardware or across unreported radio environments.  Conclusions  RadioSEUnet combines a magnitude-phase representation, constrained cross-channel feature coupling, channel attention, multiscale feature fusion, and dual-head prediction. It jointly estimates spectrum occupancy and interference-overlap states while avoiding two-dimensional STFT preprocessing. Under the controlled simulated-data protocol, the method provides higher point estimates than the six evaluated baselines at low SNR and reduces total sensing latency on the evaluated workstation. The present evidence is limited to the reported signal types, channel models, hardware configuration, and the 2.4–2.5 GHz measurement band. Quantitative simulated-to-measured results, tests on wider bands, additional interference types, repeated trials, and deployment on airborne edge hardware are still required. Future work will therefore focus on cross-domain validation, lightweight deployment, boundary-aware interference modeling, and integration with spectrum resource management for UAV networks.
Decision Learning Correction Network: Fusion Classification of Hyperspectral Images and LiDAR Data
WANG Haoyu, LIU Nuofei, CHENG Yuhu, LIU Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260362
Abstract:
  Objective  HyperSpectral Images (HSI) and Light Detection And Ranging (LiDAR) provide complementary information for land-cover classification. HSI captures rich spectral information for material discrimination, while LiDAR provides elevation and structural information for spatial characterization. However, most existing fusion methods treat multimodal fusion as a one-shot static aggregation process, implicitly assuming that a fixed fusion strategy is applicable to all pixels and regions. This assumption is difficult to satisfy in complex remote sensing scenes, where class-boundary and cross-modal heterogeneous regions exhibit high information density but account for only a small proportion of samples (Fig. 1). To address this limitation, this paper proposes a Decision Learning Correction Network (DLCN) that reformulates static HSI-LiDAR fusion as a context-dependent sequential decision-making process.  Methods  The proposed DLCN consists of feature extraction, fusion decision learning, and classification. First, HSI and LiDAR are processed through two parallel branches to extract spectral and spatial features and elevation and structural features, respectively. The extracted features are then concatenated to form the current state and are fed into an Actor-Critic framework. The Actor network generates fusion actions to adaptively adjust modality contributions, while the Critic network evaluates the long-term value of each action for classification. To improve learning from difficult samples, a key-sample-oriented sampling module assigns higher sampling probabilities to samples with larger modal fidelity loss. Meanwhile, a modal fidelity constraint mechanism evaluates spectral fidelity, feature consistency, structural preservation, and resolution matching, and corrects destructive actions during fusion. Through this closed-loop framework, DLCN performs dynamic generation, evaluation, and correction of fusion actions, thereby producing high-quality fusion features for classification (Fig. 2).  Results and Discussions  Experiments are conducted on the Houston2013, Trento, and MUUFL datasets. DLCN achieves the highest Overall Accuracy (OA) of 97.85%, 99.58%, and 94.38% on the three datasets, respectively, outperforming CHNet, DSymFuser, mPMCL, MEDFN, S3F2Net, and MSAF. The classification maps demonstrate that DLCN effectively reduces misclassification in class-boundary, mixed land-cover, and structurally complex regions, producing results that more closely match the ground-truth maps across all three datasets (Figs. 35). Ablation studies further demonstrate that the value-guided policy optimization mechanism, key-sample-oriented sampling module, and modal fidelity constraint mechanism each improve classification performance. Compared with the baseline models, the complete DLCN consistently increases OA on Houston2013, Trento, and MUUFL, validating the effectiveness of the proposed decision-learning-correction framework. Time-step analysis shows that DLCN progressively improves classification accuracy while maintaining stable spectral-angle variation during sequential decision making (Fig. 6). Furthermore, DLCN achieves inference times of 1.32 s, 0.86 s, and 2.23 s on the three datasets, respectively, ranking first among the compared methods. These results indicate that the additional computation introduced by the Actor-Critic decision framework and modal fidelity constraint mechanism is effectively translated into improved classification performance without imposing excessive computational cost.  Conclusions  This paper proposes a DLCN for HSI and LiDAR fusion classification. Unlike conventional static fusion methods, DLCN formulates multimodal fusion as a sequential decision-making process and adaptively adjusts fusion strategies according to the local context. Its closed-loop framework enables fusion actions to be generated, evaluated, and corrected throughout the decision process, thereby producing high-quality fusion features for classification. Experimental results demonstrate that DLCN produces more accurate classification maps in heterogeneous remote sensing scenes, and the time-step analysis further confirms the stability of the sequential decision-making process. Future work will focus on more fine-grained feature representation and more robust policy optimization to improve model generalization in complex remote sensing scenes.
Radiation-Hardened Ga2O3 MOSFET Design Featuring NiO Heterojunction and Comb-Shaped Gate Modulation
GAO Sheng, ZHANG Lin, WU Yanjun, WANG Qi, JING Liang
Available online  , doi: 10.11999/JEIT260396
Abstract:
  Objective  Gallium Oxide Metal-Oxide-Semiconductor Field-Effect Transistor (Ga2O3 MOSFET) is regarded as a promising power device for high-voltage applications, particularly in aerospace and satellite power systems, because of its ultra-wide bandgap and high critical breakdown field. However, the Conventional MOSFET (C-MOSFET) exhibits limited reliability in space radiation environments. Under off-state conditions, the electric field is highly concentrated near the gate edge. Heavy-ion irradiation generates dense electron-hole pairs along the ion track. Driven by the intense electric field, these carriers undergo avalanche multiplication through impact ionization, causing the drain current to increase sharply without recovery and ultimately leading to irreversible Single-Event Burnout (SEB) at relatively low drain bias. This failure mechanism severely limits the application of Ga2O3 MOSFETs in harsh radiation environments. Furthermore, the lack of reliable and efficient p-type doping restricts the implementation of conventional radiation-hardening techniques, including junction termination extension and junction isolation. Therefore, ionization-induced carriers readily accumulate in sensitive regions, increasing susceptibility to Single-Event Effect (SEE). The extremely low thermal conductivity of Ga2O3 further promotes local heat accumulation following heavy-ion irradiation, producing localized hot spots that increase the likelihood of thermal burnout. Existing hardening approaches, including field-plate optimization and dielectric engineering, provide only limited improvement. Moreover, the application of heterojunction structures for radiation hardening has rarely been investigated, and systematic hardening strategies have not yet been established. To address these limitations, this paper proposes a Comb-Shaped Gate Metal-Oxide-Semiconductor Field-Effect Transistor (CSG-MOSFET) incorporating a NiO heterojunction. The proposed structure redistributes the channel electric field, suppresses electric-field crowding at the conventional gate edge, and significantly improves SEB tolerance, providing an effective solution for Ga2O3 power devices operating in harsh radiation environments.  Methods  Technology Computer-Aided Design (TCAD) simulations are performed to evaluate the electrical characteristics and SEB performance of the proposed CSG-MOSFET in comparison with the C-MOSFET. The simulations incorporate high-field mobility, Shockley-Read-Hall recombination, Auger recombination, impact ionization, and heavy-ion models. Based on the charge-compensation effect of the p-NiO/n-Ga2O3 heterojunction, the proposed structure utilizes the extended depletion region formed at the heterointerface to redistribute the channel electric field. This heterojunction-induced depletion region improves electric-field uniformity and enhances SEB tolerance. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field away from the conventional gate edge, suppress local electric-field crowding, and improve device reliability under high-voltage and radiation conditions.  Results and Discussions  Simulation results demonstrate that the optimized Double Comb-Shaped Gate MOSFET (DCSG-MOSFET) significantly improves radiation hardness compared with the C-MOSFET. The SEB Threshold Voltage (VSEB) increases from 240 V to 2 280 V, while the Breakdown Voltage (BV) increases from 2 000 V to 3 500 V. Meanwhile, the specific on-resistance decreases. Therefore, the Baliga Figure of Merit (BFOM) and the SEB-based figure of merit are substantially improved. The NiO heterojunction and comb-shaped gate columns effectively redistribute the electric field, shifting the peak electric field from the conventional gate edge to the outer gate-column edge and suppressing local electric-field crowding. These improvements substantially enhance the radiation hardness of the device.  Conclusions  A radiation-hardened DCSG-MOSFET incorporating a NiO heterojunction is proposed and evaluated using TCAD simulations. The optimized structure significantly improves SEB tolerance while maintaining excellent electrical performance. Compared with the C-MOSFET, both VSEB and BV are substantially increased, demonstrating enhanced blocking capability. Charge compensation at the p-NiO/n-Ga2O3 heterojunction forms an extended depletion region that effectively redistributes the channel electric field and suppresses electric-field crowding near the conventional gate edge. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field to the outermost gate-column edge, thereby suppressing impact ionization induced by heavy-ion irradiation and effectively mitigating SEB. The reduced specific on-resistance further improves the BFOM and the SEB-based figure of merit. These results demonstrate that the proposed DCSG-MOSFET is a promising candidate for power electronic applications in harsh radiation environments, including aerospace and satellite systems.
A Multi-Station Emitter TDOA Deinterleaving Method for Severe Pulse-Loss Environments
LIU Yuchen, ZHAO Yaqin, WU Longwen
Available online  , doi: 10.11999/JEIT260401
Abstract:
  Objective  Modern electronic reconnaissance systems must deinterleave dense and overlapping radar pulse streams in non-cooperative environments. As radar emitters increasingly employ agile waveforms, similar pulse descriptor words, and low-intercept-probability strategies, conventional single-station methods based on carrier frequency, pulse width, and Pulse Repetition Interval (PRI) become less reliable. Multi-station deinterleaving based on Time Difference of Arrival (TDOA) provides a more stable geometric observable, but severe pulse loss still causes sparse cross-station pairing, weak true TDOA peaks, ambiguity-induced spurious peaks, isolated pulses, and fragmented trajectories across time slices. These effects increase false alarms and weaken track continuity. To address these issues, a closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments, with Time of Arrival (TOA) sequences used as the core observables.  Methods  A slice-based framework is developed for continuous reconnaissance. Residual unmatched pulses are carried forward by a sliding window to alleviate cross-slice misalignment. First, candidate pulse pairs satisfying geometric TDOA constraints are generated, and pulse descriptor word constraints on carrier frequency and pulse width are used to remove inconsistent pairs. To reduce the sparsity and binning sensitivity of conventional histograms, multiscale Kernel Density Estimation (KDE) is introduced to reconstruct the TDOA density from sparse candidate differences. Gaussian kernels with different bandwidths are fused, and candidate peaks are adaptively extracted using local statistics and peak widths. Second, a dynamic memory matrix is designed to suppress ambiguity-induced spurious peaks in high pulse repetition frequency scenarios. Since dependent spurious peaks collapse after the dominant peak is extracted and removed, a collapse-rate criterion is defined, and the spurious regions are recorded in a memory mask for subsequent iterations. Third, Dynamic Time Warping (DTW) is used to compare incomplete TOA sequences of isolated pulses with extracted pulse sequences, enabling reassignment of unequal-length and incomplete sequences. Finally, a Kalman-filter-based state-space model tracks multi-baseline TDOA trajectories across successive slices. Predicted and observed TDOA residuals are jointly used for association, so intermittent observations can still be linked to the correct track. In this way, the proposed method forms a closed-loop processing chain that links weak-peak reconstruction, spurious-peak suppression, isolated-pulse reassignment, and trajectory association (Fig. 3).  Results and Discussions  Four simulation scenarios are designed: a high pulse repetition frequency scenario dominated by ambiguity-induced spurious peaks, a parameter-overlapping scenario dominated by isolated pulse reassignment, an ablation scenario for evaluating the memory matrix and DTW modules, and a 10-emitter mixed-regime scenario including fixed PRI, staggered, jittered, frequency-agile, pulse-group frequency-agile, frequency-agile jittered-PRI, linear-sliding, and sinusoidal-sliding PRI signals. In the mixed-regime scenario, the total reconnaissance duration is 1 s and the slice duration is 0.1 s. The environmental pulse loss rate is fixed at 10%, and the receiver-specific loss rate increases from 0% to 40%. Both loss rates are calculated with respect to the initial theoretical number of transmitted pulses; therefore, the total loss rate is their sum, ranging from 10% to 50%. The proposed method is compared with an extended TDOA histogram method under constrained criteria, a cloud-model-based multi-station sorting method, and a Dirichlet Process Mixture Model (DPMM)-based method (Table 5). In the high pulse repetition frequency scenario, the proposed method maintains near-zero false alarms by identifying the collapse of dependent spurious peaks and suppressing them through the memory matrix, whereas the comparison methods show severe false alarms (Figs. 4 and 5). In the parameter-overlapping scenario, DTW-based reassignment improves isolated-pulse recovery, while the memory matrix suppresses spurious TDOA peaks. Their combination improves extraction reliability and reduces false alarms (Figs. 6 and 7). The ablation results verify their complementary roles: at a 50% loss rate, the memory matrix reduces the TDOA false alarm rate from 29.92% to 3.32%, DTW increases pulse extraction accuracy from 66.71% to 91.99%, and the complete method achieves a TDOA detection rate of 99.25% with a false alarm rate of 0.88% (Fig. 8). In the 10-emitter mixed-regime scenario, the proposed method achieves a favorable overall trade-off. At an overall pulse loss rate of 50%, its pulse extraction accuracy remains 92.72%, and the TDOA false alarm rate is limited to 7.75%, lower than 35.39%, 35.05%, and 37.58% for the DPMM, cloud-model, and constrained recursive histogram methods, respectively. After cross-slice trajectory association, the mean number of identity switches decreases to 3.83, compared with 10.64, 9.85, and 12.64 for the three comparison methods (Fig. 9 and Table 5).  Conclusions  A closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments. By integrating multiscale KDE-based weak peak reconstruction, dynamic memory-matrix-based spurious peak suppression, DTW-based isolated pulse reassignment, and Kalman-filter-based trajectory association, the method addresses the coupled failure mechanisms caused by severe pulse loss. Simulation results demonstrate high extraction accuracy, low TDOA false alarm rates, and strong trajectory continuity in high-loss and mixed-regime scenarios. These results demonstrate the effectiveness of the method under the simulated conditions and indicate its application potential for persistent multi-station passive reconnaissance.
LLM-Aided Secure Routing Method in Industrial IoT Against Flooding Attacks
LI Jieling, XIAO Liang, WANG Chengyao, FANG Mingyang, CHEN Chen, LEI Yan
Available online  , doi: 10.11999/JEIT260400
Abstract:
  Objective  Industrial Internet of Things (IIoT) routing forwards and schedules control commands, equipment status information and sensing data to support critical tasks such as collaborative equipment control, safe system operation and environmental monitoring, but the routing process is prone to congestion and resource exhaustion under flooding attacks. Existing intelligent secure routing methods apply reinforcement learning (RL) to optimize next-hop selection based on network topology, but the heterogeneity in queue capacity and link bandwidth of IIoT terminals is often overlooked, leading to load imbalance and local congestion, and limiting performance under high load or malicious traffic attacks. Therefore, we propose a large language model (LLM)-based global situation-aware assisted secure routing method in IIoT against flooding attacks, which applies RL to optimize multi-path selection and achieve load balancing across the network.  Methods  Based on global security awareness, queue congestion of neighboring nodes, queue capacity, link bandwidth, and service types, the proposed secure routing method applies RL to optimize multi-path selection against flooding attacks. The cloud–edge large model infers global security situational awareness including global load distribution and anomalous traffic distribution based on network topology, node resource occupancy and link state information, and feeds the inference result back to IIoT terminals to construct RL states and evaluate routing policies risks. In addition, a risk-aware function is formulated to quantify the routing disruption potential by integrating end-to-end latency, packet delivery ratio and node vulnerability to attacks. An experience replay buffer that incorporates both reward and risk is constructed, where both factors are considered during routing parameter updates to guide routing policy selection, thereby balancing safe path exploration and optimization efficiency.  Results and Discussions  Simulations are conducted using 30 industrial nodes under varying configurations, including bandwidths of 5 MHz, 10 MHz, 20 MHz, and queue capacities ranging from 100 to 500 packets. The global security situational awareness is inferred by the Qwen3.5-27B-AWQ-4bit, which is deployed on a cloud–edge server equipped with dual 24 GB RTX 4090 GPUs. In each time slot, each terminal sends 5 packets of 2 KB each to the industrial gateway. A flooding attacker injects \begin{document}$ y\in \{10,20,30\} $\end{document} packets into neighboring queues per time slot to excessively consume network resources. Compared with the baseline method EEMR, the proposed secure routing method improves 39.4% packet delivery ratio, reduces 48.2% end-to-end latency and 41.1% routing energy consumption. Compared with the baseline method RLMR, the proposed secure routing method improves packet delivery ratio by a factor of 1.48, reduces end-to-end latency by 53.8% and routing energy consumption by 54.5%. This is because the proposed method leverages an LLM to infer global security situational awareness, integrating load distribution and anomalous traffic patterns to assist in selecting low-load nodes while avoiding high-load nodes, potential attack nodes, and abnormal or faulty nodes.  Conclusions  This paper proposes an LLM-based global situation-aware assisted secure routing method for IIoT against flooding attacks, which applies RL to optimize multi-path selection based on global security situational awareness including load distribution and anomalous traffic distribution. A risk assessment network is constructed based on attack behavior characteristics and service requirements to evaluate the risk level of routing performance degradation, thereby enabling risk-aware rerouting. Simulation results show that the proposed method increases the packet delivery ratio by 39.4%, reduces the end-to-end latency by 48.2% and the routing energy consumption by 41.1%.
Accelerated Broadband Electromagnetic Scattering Analysis via ACA-Driven Measurement Matrix Interpolation
WANG Zhonggen, WU Chenggang, NIE Wenyan, SUN Yufa
Available online  , doi: 10.11999/JEIT260392
Abstract:
  Objective  Broadband electromagnetic scattering analysis is widely used in radar target recognition, stealth technology, and microwave imaging. Although the Method of Moments (MoM) provides high computational accuracy, it incurs substantial computational and memory costs for electrically large or geometrically complex targets because full impedance matrices must be constructed and solved. Existing acceleration techniques, including the MultiLevel Fast Multipole Method (MLFMM) and Adaptive Cross Approximation (ACA), reduce the computational burden but still require repeated matrix construction and equation solving at every frequency during wideband analysis. Methods such as Asymptotic Waveform Evaluation (AWE), Model-Based Parameter Estimation (MBPE), and impedance matrix interpolation have been proposed to reduce this redundancy. However, AWE is prone to error accumulation over wide frequency bands, MBPE requires expensive initial sampling, and conventional impedance matrix interpolation still requires the computation of full high-dimensional impedance matrices at the sampling frequencies. More recently, Compressive Sensing Method of Moments (CS-MoM) and its extension, CS-HBFM, have improved wideband analysis by employing Hyper-Basis Functions (HBFs). By constructing Characteristic Mode Basis Functions (CMBFs) only once at the highest frequency, CS-HBFM eliminates repeated basis-function generation. Nevertheless, existing CS-HBFM methods rely on nondeterministic random or uniform sampling, require expensive large-scale matrix-vector products, and repeatedly reconstruct and solve impedance equations throughout the frequency sweep.  Methods  A CS-ACA-MMI framework is proposed for broadband electromagnetic scattering analysis by combining dual ACA decomposition with Measurement Matrix Interpolation (MMI). First, CMBFs are constructed at the highest frequency, and dominant HBFs are selected according to the Modal Significance (MS) criterion. ACA is then applied to the full impedance matrix to extract deterministic row indices corresponding to the dominant Rao-Wilton-Glisson (RWG) basis functions. These indices are reused throughout the frequency band, eliminating nondeterministic sampling and repeated index extraction. Second, four sampling frequencies are selected using Chebyshev-Lobatto nodes. Low-dimensional measurement matrices are constructed directly from the extracted row indices, avoiding the generation of full high-dimensional impedance matrices. The measurement impedance elements at the sampling frequencies are corrected according to the geometric distance, interpolated to the target frequency, and then restored to the actual measurement impedance elements, thereby eliminating repeated construction of measurement matrices during frequency sweeping. Third, ACA is applied to the far-field component of the interpolated measurement matrix, converting large-scale matrix-vector products into low-dimensional matrix multiplications. The near-field sensing matrix is obtained directly by multiplying the measurement matrix by the basis functions, enabling rapid construction of the complete sensing matrix. Finally, the dense linear system is transformed into an overdetermined system under the compressive sensing framework, and the least-squares method is used to reconstruct the current coefficients, from which the broadband Radar Cross Section (RCS) is calculated. The Root Mean Square Error (RMSE) is used to evaluate numerical accuracy. Three representative targets, namely a cylinder, a slotted cone, and an almond, are analyzed. Broadband RCS, numerical accuracy, total computation time, and single-frequency measurement-matrix memory consumption are compared with those obtained using MoM and CS-HBFM to validate the proposed framework.  Results and Discussions  Three numerical examples, including a perfect electric conductor cylinder, a slotted cone, and an almond, are used to validate the proposed CS-ACA-MMI framework. The ACA-extracted row indices are concentrated near geometric boundaries and structural junctions, demonstrating the physical validity of the deterministic sampling strategy (Fig. 2). Parametric studies show that appropriate ACA thresholds and four sampling frequencies provide the best balance between computational efficiency and numerical accuracy (Figs. 35). The broadband RCS predicted by the proposed framework agrees closely with the MoM results over the entire frequency band (Figs. 68), and the RMSE remains low, demonstrating high numerical accuracy. Compared with CS-HBFM, the proposed framework reduces the total computation time by 93.4% for the cylinder, 96.7% for the slotted cone, and 80.9% for the almond (Table 2). These improvements result from deterministic index reuse, MMI, and dual ACA acceleration, which substantially reduce the computational cost of broadband frequency-sweeping analysis.  Conclusions  A CS-ACA-MMI framework is proposed by integrating ACA with MMI for efficient broadband electromagnetic scattering analysis. The proposed framework eliminates repeated matrix construction and equation solving during frequency sweeping while overcoming the nondeterministic sampling strategy and the high computational and memory costs of conventional CS-HBFM. Dominant row indices extracted by ACA at the highest frequency provide a deterministic measurement-matrix construction strategy and a stable physical basis for broadband interpolation. By shifting the interpolation target from full impedance matrices to low-dimensional measurement matrices, the computational complexity and redundant matrix construction are substantially reduced. A second ACA decomposition further accelerates sensing-matrix construction by converting large-scale matrix-vector products into low-dimensional matrix multiplications. Numerical results demonstrate that the proposed framework achieves numerical accuracy comparable to that of MoM while reducing total computation time by more than 80% and decreasing single-frequency measurement-matrix memory consumption by up to 65%. Because only the measurement matrices at four sampling frequencies need to be stored, the overall memory requirement is further reduced.
Construction and Performance Analysis of Optimal Low-Hit-Zone Frequency Hopping Sequence Sets
TIAN Xinyu, CHEN Xiaoyu, ZHANG Jitao
Available online  , doi: 10.11999/JEIT260343
Abstract:
  Objective  ElectroMagnetic Interference (EMI) is a critical factor limiting the reliability of synchronization systems. Existing Fifth-Generation (5G) synchronization schemes extensively employ Zadoff-Chu (ZC) sequences to distinguish users through cyclic shifts. However, finite sequence lengths and limited orthogonal resources create substantial capacity bottlenecks in high-density access scenarios. To address these challenges, this paper investigates the problem from two perspectives. At the system level, a synchronization framework is developed by integrating Frequency Hopping (FH) with ZC sequences. By jointly exploiting code, time, and frequency-domain resources, the proposed framework improves concurrent access capability for local clusters while enhancing robustness against complex EMI through frequency diversity. At the sequence-design level, a class of multi-subset Low-Hit-Zone (LHZ) Frequency Hopping Sequence (FHS) sets is constructed to provide an efficient sequence allocation scheme for local-cluster synchronization.  Methods  Based on the theoretical framework proposed by Cai et al., the sequence mapping mechanism is reconstructed, and a disjoint Cyclic Perfect Mendelsohn Difference Family (CPMDF) is introduced to construct FHS sets that are optimal with respect to the Peng-Fan bound. The generating units are further expanded through Cartesian products, and a column-incoherent partitioning strategy is proposed to construct multi-subset LHZ FHS sets. It is proved that every nonempty subset satisfies the Peng-Fan-Lee bound with equality. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure. At the system level, an FH-ZC synchronization architecture is developed by combining predefined FH patterns with the frequency-domain correlation properties of ZC sequences for subband signal detection. A Peak-to-SideLobe Ratio (PSLR) decision metric and an early-termination strategy are adopted to evaluate synchronization preamble detection under interference. Furthermore, a multi-user simulation model is established to evaluate synchronization detection performance under accumulated co-channel collisions and EMI.  Results and Discussions  The proposed construction generates an FHS set that is optimal with respect to the Peng-Fan bound and a class of multi-subset LHZ FHS sets in which every nonempty subset is optimal with respect to the Peng-Fan-Lee bound. Example 2 demonstrates the construction procedure and the intra-subset and inter-subset Hamming correlation properties of the proposed multi-subset LHZ FHS sets. Table 1 shows that, under the same frequency-resource constraints, the proposed construction generates more sequences than existing methods under the compared parameter settings, indicating higher sequence-resource utilization. Table 2 compares the parameters of the proposed sequence sets with representative constructions reported previously and demonstrates that the proposed multi-subset optimal sequence family provides a new parameter combination. To the best of our knowledge, an optimal sequence family with a multi-subset structure has not been reported previously. Figures 2 and 3 demonstrate that the proposed FH-ZC synchronization architecture achieves a higher synchronization detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI. Figure 4 shows that the synchronization detection probability decreases as the number of active users increases because accumulated co-channel collisions degrade synchronization performance. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure characterized by strong intra-cluster competition and weak inter-cluster coupling.  Conclusions  To satisfy the sequence-capacity requirements of massive-access scenarios, this paper proposes a class of multi-subset LHZ FHS sets. By expanding the generating sequence sets through Cartesian products and partitioning subsets using a column-incoherent strategy, the proposed construction achieves both a large family size and optimal LHZ performance. The proposed multi-subset structure is well suited to local-cluster synchronization and substantially improves sequence family size and sequence-resource utilization, thereby providing a richer sequence resource pool for high-density multi-user systems. Simulation results under the considered physical-layer model demonstrate that the proposed LHZ FHS subsets reduce the effect of frequency collisions during multi-user synchronization detection. Furthermore, the FH-ZC synchronization scheme achieves a higher synchronization preamble detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI.
A Phase Transition Obstacle Avoidance Method for UAV Swarms Driven by Multistable Potential Fields
HE Ming, CHEN QiYang, HAN Wei, PAN Fan, MA YiSong
Available online  , doi: 10.11999/JEIT260357
Abstract:
  Objective  Unmanned Aerial Vehicle (UAV) swarms have demonstrated considerable potential for complex missions, such as search, surveillance, and disaster response, because of their distributed coordination and robustness. However, in dynamic environments with dense obstacles and rapidly changing risks, conventional swarm control methods often exhibit discontinuous behavior switching and control chattering, which reduce system stability and coordination efficiency. Existing approaches, including threshold-based switching and Artificial Potential Field (APF) methods with fixed potential weights, rely on abrupt transitions between behavioral modes, leading to oscillatory responses. To address these limitations, a phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. Swarm behavior evolution is modeled as a continuous phase transition process within a unified potential field framework, enabling smooth and adaptive transitions between formation flight and obstacle avoidance.  Methods  An environmental risk assessment model is first established by integrating static obstacle risk, dynamic obstacle risk, and inter-agent proximity risk. A distributed consensus protocol is then employed to establish global risk consensus. Subsequently, a morphology factor is generated through nonlinear mapping of the global risk consensus and is used as an order parameter to characterize the macroscopic swarm state. A unified time-varying potential field, comprising formation, obstacle avoidance, and navigation potentials, is constructed, and the relative weights of these potentials are continuously adjusted by the morphology factor. When the risk level is low, the system exhibits a monostable structure dominated by the formation and navigation potentials. As the risk increases, the potential field continuously evolves into a multistable structure dominated by the obstacle avoidance potential, thereby enabling distributed obstacle avoidance. A distributed consensus control law based on the negative gradient of the unified potential field is further developed. A damping term is incorporated to dissipate system energy and improve stability, while a dynamic compensation term addresses nonlinear dynamics. The control law depends only on local information, ensuring good scalability. The global uniform ultimate boundedness of the closed-loop system is established using Lyapunov theory.  Results and Discussions  Simulation results demonstrate that the proposed method enables the swarm to maintain a compact solid-phase swarm formation in low-risk regions and to transition smoothly to a dispersed liquid-phase swarm configuration when obstacles are encountered, followed by rapid formation recovery after obstacle avoidance. The pitch and roll angles of each UAV vary smoothly without abrupt changes, and both the UAV-to-obstacle distance and the inter-UAV separation remain above the prescribed safety threshold throughout the flight, ensuring collision-free operation. Statistical results obtained from 20 independent simulation runs show that, compared with the threshold-switching method, the proposed method reduces the rate of control input variation by approximately 26% and decreases the peak control input by approximately 18%. Compared with the bio-inspired diversion method, the average formation recovery time after obstacle avoidance is reduced by approximately 16%. Ablation experiments further demonstrate that removing the morphology-driven phase transition mechanism significantly increases trajectory oscillation and control oscillation, confirming the critical role of the multistable continuous phase transition mechanism in maintaining smooth swarm motion. In complex narrow-channel environments, the proposed method effectively avoids the local minimum problem encountered by conventional APF methods and generates smoother flight trajectories with substantially reduced oscillation.  Conclusions  A phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. By introducing a morphology factor and constructing a unified potential field framework, swarm behavior evolution is represented as a continuous phase transition process. The distributed control law enables smooth behavioral transitions while maintaining system stability and scalability. Simulation results demonstrate that the proposed method achieves better safety, smoother control, and higher coordination efficiency than conventional methods.
An SO(3)-Manifold-Constrained Registration Method for Twin-Fisheye Panoramic Images
WANG Zhuopeng, LIN Shanling, LIN Jianpu, LÜ Shanhong, LIN Zhixian
Available online  , doi: 10.11999/JEIT260798
Abstract:
  Objective  Twin-fisheye cameras provide near-360° coverage with low hardware complexity and are widely used in immersive imaging, surveillance, and mobile robotics. Their panoramic output depends on registration over a narrow overlapping band, so geometric accuracy and temporal consistency directly affect seam quality and video smoothness. After the two fisheye views are unfolded into the Equirectangular Projection (ERP), three coupled problems arise. First, the near-co-centric lens pair is ideally related by a pure rotation R ∈SO(3), whereas a conventional 8-Degree-of-Freedom (DoF) homography introduces five redundant parameters that may couple with matching noise. Second, ERP sampling is nonuniform with latitude, so identical pixel residuals do not represent identical spherical angular errors. Third, the cyclic ±π longitude boundary splits structures that are continuous on the sphere and weakens correspondences around the seam. Existing planar pipelines and generic learned matchers rarely combine these constraints under a unified rotation-referenced evaluation. This study therefore develops a lightweight registration framework that explicitly exploits spherical rotation geometry while addressing ERP boundary discontinuity, temporal fluctuation, and long-tail residuals.  Methods  The proposed framework contains three modules (Fig. 1). First, an overlapping-band Region-of-Interest (ROI) is cropped around the ERP seam and rearranged with modulo-W wrap-around (Fig. 2). A default longitude half-width of ±15° and an approximately 3° margin on each side preserve cross-boundary feature continuity while restricting the search region. XFeat detects, describes, and matches features under a fixed Top-K budget. Second, the two-dimensional matches are restored to global ERP coordinates, mapped to unit-sphere direction vectors, and processed by rotation-only SO(3)-RANSAC. Spherical angular residuals are used as the inlier criterion with a 0.8° threshold, a maximum of 2000 iterations, confidence 0.999, and at least 12 inliers; the iteration bound is updated adaptively. All inliers are then used for Kabsch/SVD closed-form rotation re-estimation, which reduces the randomness of a minimal sample while preserving the SO(3) constraint. Third, a local increment on the Lie algebra so(3) is optimized under a Huber loss by the Levenberg–Marquardt algorithm. The refinement is triggered only when the inlier-residual P95 exceeds 0.90° and the inlier ratio is below 0.58, thereby concentrating nonlinear optimization on difficult image pairs.  Results and Discussions  Experiments are conducted on PanoraMIS Sequences 3 and 4 under a unified relative inter-frame rotation protocol (Table 1). The proposed method achieves a 97.10% success rate, a 0.549° P95 angular residual, and a 0.313° temporal-stability error. Compared with SuperPoint+LightGlue, the P95 and temporal-stability errors are reduced by 22.8% and 77.3%, respectively. Compared with Efficient LoFTR, peak GPU memory and runtime are reduced by 52.7% and 60.9%, although Efficient LoFTR retains the lowest overall P95. Under a unified SO(3)-RANSAC back-end (Table 2), XFeat provides the largest average inlier count of 625.4 and the lowest temporal-stability error of 0.313° at 38.04 ms. The ablation and sensitivity results (Tables 34) show that the ±15° ROI reduces the P95 from 0.720° for the full ERP to 0.600°. Replacing H-RANSAC with SO(3)-RANSAC reduces temporal instability from 0.931° to 0.313°, a 66.3% reduction, while increasing runtime from 24.91 ms to 35.24 ms. Adaptive refinement operates on approximately one third of the image pairs and improves both P95 and temporal stability with lower overhead than always-on refinement; its five-seed mean and median trigger rate are both 36.23%. A Top-K budget of 2048 reaches the saturated accuracy level, because increasing the budget to 4096 yields no further P95 or stability improvement. Five fixed-seed repetitions produce standard deviations no greater than 0.011° for P95 and stability, indicating that the main conclusions are insensitive to RANSAC randomness. In a 5×5 threshold sweep, the maximum changes in P95 and inter-frame rotation jitter within the central neighborhood are 4.40% and 0.037%, respectively, and changing the robust-error truncation from 3° to 2° or 5° does not alter the relative ranking. On the more difficult Sequence 4, characterized by weak texture and unstable overlap, the proposed method obtains 919.7 average inliers and a P95 of 0.383°, the lowest among the evaluated learning-based matchers, although robust-estimation time increases. With an identical standardized stitching back-end, it produces a lower seam-band gradient than H-RANSAC in the small-rotation example (26.90 versus 28.71) and a lower truth-referenced temporal-stability error (0.313° versus 0.931°; Fig. 6). On outdoor Sequence 7-L2, 324 of 346 correspondences are retained, yielding a 93.6% inlier ratio and a 0.524° P95 residual (Fig. 5). Because sequence-specific calibration is unavailable and the fixed inter-lens baseline may cause depth-dependent parallax, this result serves only as a diagnostic consistency check, not as evidence of absolute pose accuracy.  Conclusions  By restoring feature continuity across the ERP boundary, replacing the redundant planar homography with an explicit SO(3) rotation model, and selectively refining difficult image pairs on so(3), the proposed method balances registration accuracy, temporal consistency, and resource cost. It provides a lightweight front-end for twin-fisheye panorama stitching. The current evaluation is limited to pairwise registration on a small number of sequences; future work will address translation compensation, multi-frame global optimization, end-to-end integration with seam finding, exposure compensation, and blending, as well as generalization across additional platforms, dynamic scenes, and illumination conditions.
Co-Frequency Interference Analysis and Dynamic Simulation Validation of Satellite-Direct-to-Device Systems Against Terrestrial IMT Networks in Cross-Border Scenarios
LIU Quan, ZHAO Weisong, XIAO Na, SONG Yanjun, ZHOU Meng, ZHANG Zhili, WANG Jinhai, WANG Lichong
Available online  , doi: 10.11999/JEIT260263
Abstract:
  Objective   Satellite-Direct-to-Device (SD2D) systems that reuse terrestrial IMT spectrum may generate harmful downlink interference to incumbent IMT networks in neighboring administrations, particularly in cross-border deployments where SD2D downlinks overlap the receive bands of both IMT user equipment (UE) and IMT Base Stations (BSs). A practical coexistence methodology is therefore required to (i) translate IMT receiver protection criteria into explicit Power Flux Density (PFD) and Equivalent Power Flux Density (EPFD) constraints and (ii) validate these constraints using a dynamic simulation framework so that they can be converted into enforceable geographic coordination measures, such as minimum isolation distances. This study focuses on the dominant interference path, namely SD2D downlink interference to IMT receivers, and establishes a traceable workflow from deterministic protection limits to dynamic simulation validation and the corresponding minimum isolation distances.  Methods  A cross-border scenario is modeled in which Country A deploys an SD2D system and Country B operates a terrestrial IMT network. Two representative downlink frequencies, 1 995 MHz and 2 190 MHz, are evaluated for two representative Starlink configurations, Starlink-1 and Starlink-2. The IMT network is modeled using ITU-R- and 3GPP-compliant parameters, with an I/N protection threshold of –6 dB and a target percentile κ (baseline κ=99.5%) for both IMT UEs and BSs. Satellite transmit antennas follow the ITU-R S.1528 reference pattern, IMT BS receive antennas follow the ITU-R F.1336 sector pattern, and IMT UEs are modeled with omnidirectional antennas. A back-lobe blockage model is incorporated into both satellite and BS antenna patterns to account for rear-side shielding. Signal propagation follows the ITU-R P.619 model, using free-space path loss as the conservative baseline, while an optional clutter-loss term is incorporated through a clutter-occurrence probability. Deterministic protection limits are derived by calculating the maximum permissible aggregate PFD for IMT UE protection and the maximum permissible aggregate EPFD for IMT BS protection. A dynamic simulation framework then validates these limits and searches for the required minimum isolation distances (Fig. 4). Co-channel beam isolation angles are optimized using the C/I Complementary Cumulative Distribution Function (CCDF), and a segmented search algorithm determines the minimum UE- and BS-side isolation distances together with the corresponding κ-percentile PFD/EPFD statistics.  Results and Discussions  The deterministic analysis yields a maximum permissible aggregate PFD of –102.72 dBW/m2/MHz at 2 190 MHz for IMT UEs and a maximum permissible aggregate EPFD of –129.53 dBW/m2/MHz at 1 995 MHz for IMT BSs (Fig. 3). For Starlink-1, the C/I design criterion yields a minimum co-channel beam isolation angle pair of (12°, 12°) (Fig. 5). Dynamic simulation shows that, under the representative baseline configuration with an I/N threshold of –6 dB and κ=99.5%, the minimum isolation distances are 195 km for UE protection and 290 km for BS protection (Fig. 6, Fig. 7, and Table 4). The resulting coordination isolation distance is therefore 290 km, and the simulated κ-percentile PFD and EPFD agree with the deterministic protection limits, with a residual margin below 0.5 dB. For Starlink-2, the optimized co-channel beam isolation angles increase to (15°, 15°), and the corresponding minimum isolation distances increase to 272 km for UEs and 420 km for BSs under the same baseline configuration (Table 5). These baseline distances should be interpreted as representative values for the specified simulation configuration rather than unique, strictly converged results. Stability verification shows that, under different sampling intervals, simulation durations, and random seeds, the UE- and BS-side minimum isolation distances remain within 195~210 km and 290~300 km, respectively, for Starlink-1, and within 266~290 km and 370~420 km, respectively, for Starlink-2 (Table 7). Sensitivity analysis for Starlink-1 further indicates that the required minimum isolation distance is governed by the upper tail of the aggregate I/N distribution (Table 6). Increasing κ from 99.5% to 100% increases the UE- and BS-side minimum isolation distances from 195/290 km to 304/560 km. Clutter attenuation substantially reduces the UE-side minimum isolation distance, decreasing it to 173 km when the clutter-occurrence probability is 0.5, while producing little change in BS protection. Polarization reuse increases the UE- and BS-side minimum isolation distances to 222 km and 360 km, respectively, whereas increasing the number of co-channel beams to 16 increases the BS-side minimum isolation distance to 330 km. The minimum service elevation angle and the link establishment strategy are identified as the dominant operational factors. Changing the minimum service elevation angle from 10° to 35° changes the required UE- and BS-side minimum isolation distances from 340/460 km to 101/150 km, whereas replacing the Sat-MaxElevation strategy with the UE-MaxElevation strategy reduces them to 80/180 km.  Conclusions   The proposed workflow converts IMT receiver protection criteria into deterministic protection limits expressed as PFD and EPFD constraints and validates them using a dynamic simulation framework. Under an I/N threshold of –6 dB and κ=99.5%, the baseline and stability analyses jointly indicate representative UE- and BS-side minimum isolation-distance ranges of 195~210 km and 290~300 km for Starlink-1 and 266~290 km and 370~420 km for Starlink-2, rather than unique, strictly converged values. Sensitivity analysis further shows that κ only changes the statistical criterion used to extract tail events from the sample set, whereas clutter attenuation primarily benefits IMT UEs. In contrast, the minimum service elevation angle, polarization reuse, the number of co-channel beams, and the link establishment strategy reshape the worst-case interference geometry and can produce substantial, and sometimes non-monotonic, changes in the required minimum isolation distances. The proposed framework establishes a traceable link between IMT receiver protection criteria and enforceable border coordination measures.
MG-MoE: Routed Multi-Granularity Expert Ensemble
XIAN Fengyu, JIAN Haifang, XIE Zihui, DU Jun, ZHANG Yuanyuan, NING Xin, DONG Miaomiao, WANG Hongchang
Available online  , doi: 10.11999/JEIT260219
Abstract:
  Objective  Fine-Grained Image Recognition (FGIR) aims to distinguish visually similar subcategories that differ only in subtle local patterns. It must also remain robust to large intra-class variations caused by pose changes, occlusion, illumination shifts, and complex backgrounds. In real-world scenarios, these challenges are further intensified by long-tailed category distributions. Rare or difficult classes are more likely to overfit spurious contextual cues and suffer from unstable decision boundaries. Therefore, a conditional computation paradigm is needed, in which complementary inductive biases are separated into specialized expert branches and adaptively combined for each sample. This work aims to develop a routed multi-granularity mixture-of-experts framework that improves discriminative performance under controllable inference cost. It also enhances robustness for difficult samples and long-tailed categories through adaptive sparse expert activation.  Methods  A Multi-Granularity Mixture-of-Experts (MG-MoE) model is proposed. It is a routed ensemble architecture composed of a shared backbone, four heterogeneous experts, and a learnable router that predicts input-conditioned expert weights (Fig. 2). The experts are designed with complementary inductive biases to address key factors in FGIR. MPSA emphasizes global structure and contour-level semantics. PMG captures fine local details through multi-granularity part modeling. TransFG focuses on pose and deformation modeling. PIM improves robustness in cluttered backgrounds through background suppression. To limit interference and reduce unnecessary computation, MG-MoE adopts sparse fusion. Only the Top-K experts, with K=2 by default, contribute to the final prediction during inference. To improve routing stability and generalization, a two-stage optimization strategy is designed. In the first stage, dynamic cluster-level training is performed. A cluster-level soft teacher distribution is constructed from validation-set statistics and imposed through Kullback-Leibler (KL) divergence regularization. This process stabilizes routing behavior and promotes effective expert specialization. In the second stage, residual fine-tuning is conducted. The feature-driven routing mechanism is kept unchanged, while the classification heads of the Top-2 experts associated with each cluster are selectively unfrozen. The router and expert heads are then jointly optimized with grouped learning rates. This design reduces fusion bias and strengthens discrimination for difficult samples and long-tailed categories.  Results and Discussions  MG-MoE achieves strong performance on standard FGIR benchmarks. On CUB-200-2011, it obtains 92.89% Top-1 accuracy. This result is higher than those of representative expert backbones used individually, including MPSA (91.23%), PIM (91.17%), and TransFG (90.49%). It also outperforms the multi-granularity baseline PMG (88.32%) (Table 1). On the Bird-1445 sampled set, MG-MoE achieves 96.80% Top-1 accuracy and consistently improves over strong baselines (Table 2). These results indicate that routed multi-expert specialization remains effective in data-limited and highly similar fine-grained scenarios. The efficiency-accuracy trade-off is summarized in Table 3. With Top-2 sparse routing, MG-MoE reaches 92.89% accuracy with a compute budget of 143.9 GFLOPs. It avoids dense expert activation during inference by selecting only the Top-2 experts for each sample, thereby achieving a favorable balance between accuracy and efficiency. Ablation experiments show that increasing K beyond 2 does not yield consistent gains, which suggests that indiscriminate fusion can dilute discriminative evidence. Top-2 fusion produces the best performance, whereas Top-1 fusion is more sensitive to routing errors and larger K values may introduce noise and reduce accuracy (Table 4). The role of expert diversity and composition is also analyzed. Two- and three-expert variants generally underperform the full four-expert configuration, indicating that each inductive bias contributes to different fine-grained difficulty factors. In contrast, adding homogeneous experts without new functional diversity brings diminishing or negative gains, which is consistent with increased routing ambiguity and limited expert complementarity (Table 5). These results support the use of a compact set of heterogeneous experts combined with sparse routing. To interpret the learned specialization, category-wise routing statistics are visualized. The expert-category heatmap shows that MPSA receives dominant routing weights across many categories, reflecting the central role of global structure in fine-grained discrimination. PIM and TransFG show higher activation for specific difficult categories, which is consistent with their roles in background suppression and pose and deformation modeling (Fig. 3). Finally, t-SNE visualizations illustrate the qualitative effect of expert fusion on class separability. Shared backbone features show stronger inter-class entanglement among visually similar subcategories. In contrast, fused outputs form clearer clusters with better between-class separation and within-class compactness, indicating a more reliable decision space shaped by routed expert aggregation (Fig. 4).  Conclusions  MG-MoE is a multi-granularity routed mixture-of-experts framework for fine-grained recognition. By combining four complementary experts, Top-2 sparse fusion, and a two-stage optimization strategy for stable routing and calibrated fusion, MG-MoE improves recognition accuracy on CUB-200-2011 and the Bird-1445 sampled set. It also provides interpretable evidence of expert specialization (Table 1, Table 2, Fig. 3, Fig. 4). Ablation results confirm that controlled Top-2 fusion and heterogeneous expert design are key to the observed performance gains. Overly dense fusion or homogeneous expert expansion provides limited benefit (Table 4, Table 5).
Efficient Non-Orthogonal Multiple Access Scheme Based on Modified Alamouti Code Design
WAN Dehuan, HUANG Ronglan, JI Fei, LIU Jingxian, LIANG Yaokun, LÜ Lu, LI Xingwang, YUE Xinwei
Available online  , doi: 10.11999/JEIT260567
Abstract:
  Objective  Existing Alamouti-coding-based Non-Orthogonal Multiple Access (NOMA) schemes adopt an equal-number symbol transmission mode for both cell-center and cell-edge users, which overlooks the significant channel disparity between the two types of users. This leads to two drawbacks: on one hand, the superior channel condition of cell-center users is not fully exploited, resulting in wasted transmission resources; on the other hand, the equal-number transmission inevitably weakens the performance of Space-Time Block Coding (STBC) in suppressing intra-group multi-user interference. Therefore, a novel design that adapts to user channel differences is urgently needed to improve spectral efficiency and interference mitigation capability.  Methods  In light of the significant channel disparity between cell-center and cell-edge users, this paper proposes an unequal-number symbol transmission scheme for Alamouti coding in NOMA. Specifically, when constructing Alamouti group transmission codes, the number of symbols required for the cell-edge user is made smaller than that for the cell-center user. By reducing the number of transmitted symbols for the cell-edge user, two benefits are achieved: first, the multi-user interference imposed on the cell-center user is directly reduced; second, under a total power constraint, reducing the number of symbols for the cell-edge user equivalently increases the per-symbol transmission power, thereby significantly improving the received signal-to-interference-plus-noise ratio (SINR). Furthermore, the paper derives perfect closed-form solutions for the achievable sum rate and outage probability of the proposed scheme, and validates its effectiveness via numerical simulations.  Results and Discussions  The theoretical derivations yield closed-form analytical expressions for the achievable sum rate and outage probability, providing an accurate basis for system performance evaluation. Numerical simulation results demonstrate that, compared with existing equal-symbol Alamouti coding schemes, the proposed unequal-symbol Alamouti coding scheme effectively reduces the intra-group interference from the cell-edge user to the cell-center user, while significantly enhancing the SINR of the cell-edge user through power reallocation. Under typical channel parameters, the system achieves a notable sum-rate gain and a substantial reduction in outage probability, confirming the high efficiency of the proposed scheme.  Conclusions  The proposed unequal-number symbol transmission scheme based on Alamouti coding for cell-edge and cell-center users fully exploits the potential benefits arising from channel disparity. By reducing the number of symbols transmitted by the cell-edge user, the scheme achieves both suppression of intra-group interference and enhancement of the edge user’s transmission power. The theoretical closed-form solutions and simulation results consistently show that the proposed scheme outperforms conventional equal-number transmission schemes, providing an effective new approach for mitigating multi-user interference and optimizing resource allocation in NOMA systems.
Function-Aware Partitioning Driven Hierarchical Circuit Representation Learning
YE Juyang, CHEN Qilin, WANG Yaohua
Available online  , doi: 10.11999/JEIT260645
Abstract:
  Objective  One of the core challenges in applying machine learning techniques to electronic design automation (EDA) lies in learning high-quality circuit representations from large-scale gate-level netlists. As modern digital integrated circuits scale to tens of millions of gates, existing methods based on graph neural networks (GNNs) and graph transformers (GTs) suffer from excessive computational and memory overhead, rendering them impractical for industrial-scale designs. The fundamental issue is the lack of a proper tokenization mechanism for netlists—unlike natural language, where subword tokenization effectively compresses long sequences, the circuit domain lacks an analogous decomposition strategy that preserves functional semantics while reducing the effective graph size. This work aims to bridge this gap by introducing a hypergraph-partitioning-based circuit tokenizer that decomposes massive netlists into functionally cohesive sub-circuits, termed circuit elements, thereby enabling scalable and fine-grained representation learning.  Methods  This paper proposes a function-aware hypergraph partitioning driven framework for large-scale circuit representation learning. Inspired by the tokenization paradigm of large language models, the framework first models a gate-level netlist as a directed hypergraph, where gates are nodes and signals are hyperedges that can connect multiple gates. A novel optimization objective, the Functional Independence Ratio (FIR), is introduced to guide the partitioning process. FIR incorporates circuit structural priors and a bus recognition correction mechanism that identifies bus-structured signals based on structural similarity (gate type purity across predecessor/successor levels) and spatial similarity (topological distance variance within candidate groups). The bus recognition module corrects the effective interface count, ensuring that functionally cohesive modules are not penalized for using wide buses. An iterative greedy refinement procedure accepts only moves that strictly decrease FIR, converging to a locally optimal partition. On top of the partitioned circuit elements, a two-stage self-supervised pretraining framework is designed. In the first stage, a masked autoencoder with edge prediction tasks is applied to the coarse-grained circuit-element graph, learning global inter-element dependencies. The graph transformer encoder is then frozen and circuit element embeddings are saved. In the second stage, within each circuit element, a contrastive learning scheme is employed at the gate level. Positive pairs are constructed via Boolean equivalence transformations (e.g., associativity, De Morgan's laws), which preserve the Boolean function while altering the gate-level structure. Negative pairs are drawn from functionally different circuit elements within the same batch. The training jointly optimizes node-level and local-global alignment losses. A feature-wise modulation mechanism injects circuit-element-level context into gate-level representations, enabling the same gate type to acquire different embeddings depending on its functional context.  Results and Discussions  Extensive experiments are conducted on circuits collected from multiple sources, including ITC99, EPFL, OpenCores, and three RISC-V SoC designs (Rocket, BOOM, and OpenC910), with the largest design containing 22.3 million gates. For the tokenizer evaluation, FIR-based partitioning is compared against Mt-kahypar using the Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) metrics, with the original module hierarchy serving as ground truth. Across all three RISC-V designs, FIR achieves AMI improvements of 24.6% to 27.1% and ARI improvements of 25.5% to 35.8% over Mt-kahypar, demonstrating consistent and substantial gains in functional coherence. On downstream tasks, the proposed method consistently outperforms state-of-the-art baselines including DeepGate4 and NetTAG across all comparable datasets. For full-circuit function recognition, the proposed method achieves F1 scores of 0.896, 0.861, 0.817, and 0.773 on ITC99, EPFL, OpenCores, and Rocket, respectively, representing improvements of 10.0% to 15.4% over the strongest baseline. Crucially, on BOOM (8.12M gates) and OpenC910 (22.3M gates), the proposed method is the only approach capable of completing training without running out of memory, achieving F1 scores of 0.771 and 0.748 for full-circuit tasks and 0.723 and 0.679 for gate-level tasks, respectively. Ablation studies with six model variants reveal a clear division of labor: the module-level pretraining stage dominates full-circuit performance (20.9% F1 drop when removed), while the gate-level pretraining stage dominates gate-level performance (26.1% F1 drop when removed). The FIR partitioning objective contributes 8.6% and 11.7% F1 improvements at the full-circuit and gate levels, respectively. The bus recognition module and the two-stage architecture also show consistent positive contributions across all metrics.  Conclusions  This paper presents a novel framework that addresses the scalability challenge in circuit representation learning by introducing a function-aware hypergraph partitioning tokenizer and a two-stage self-supervised pretraining architecture. The key insight is that by abstracting the intermediate circuit element level between individual gates and the full circuit, one can achieve both scalability to multi-million-gate designs and fine-grained gate-level discriminative capability. The proposed FIR objective effectively captures functional cohesion during partitioning, and the two-stage pretraining framework decouples global context learning from local representation refinement. Experimental results demonstrate state-of-the-art performance on function recognition tasks and, more importantly, the unique ability to scale to industrial-sized designs where existing methods fail. Future work includes extending the framework to other EDA tasks such as logic synthesis and physical design, exploring more aggressive hierarchical strategies for billion-gate designs, and adapting the bus recognition mechanism to non-standard cell libraries.
A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar
HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang
Available online  , doi: 10.11999/JEIT260695
Abstract:
  Objective  Traditional optimization-based power allocation algorithms for collocated MIMO radar have two fundamental limitations. First, they optimize tracking performance only for the next time step and therefore lack a full-time-horizon view of the power allocation process. This myopic strategy cannot achieve optimal multi-target tracking accuracy over extended periods, particularly when target trajectories vary substantially. Second, these algorithms rely on iterative nonlinear constrained optimization, resulting in high computational complexity. Therefore, they cannot satisfy the real-time requirements of dynamic battlefield environments where target states change rapidly. To address these limitations, this paper proposes a Reinforcement Learning (RL)-driven power allocation algorithm. Unlike conventional methods, the proposed approach formulates the power allocation problem as a Markov Decision Process (MDP) that maximizes long-term cumulative tracking accuracy. The algorithm adaptively allocates limited transmit power among multiple beams according to the current system state, balancing immediate tracking performance with long-term cumulative tracking accuracy.  Methods  The Posterior Cramér-Rao Lower Bound (PCRLB) is employed to quantify the theoretical lower bound of the tracking error for each target. The state space is constructed by combining the motion states (position and velocity) of all targets with the normalized PCRLB from the previous allocation step. The action space consists of discrete transmit power levels for each beam, subject to the total power budget and individual beam power constraints. All feasible power allocation vectors are enumerated and encoded to reduce the action-space dimensionality. The reward function is defined as the negative weighted sum of the normalized PCRLB, encouraging the agent to minimize tracking errors. The power allocation process is formulated as an MDP and solved using the Dueling Double Deep Q-Network (D3QN) algorithm. The D3QN framework incorporates three major enhancements: (1) a double-network architecture comprising an online Q-network and a target Q-network to improve training stability; (2) a dueling architecture that decomposes the Q-value into a state-value function and an action-advantage function to improve action discrimination; and (3) off-policy learning with experience replay to improve the use of historical trajectories. An ε-greedy strategy is adopted for exploration, with ε gradually decreasing during training. After offline training, the learned network directly generates real-time transmit power allocation decisions from the current system state without iterative optimization.  Results and Discussions  Simulations are conducted using three targets following the Constant Velocity (CV) model. Fixed power allocation yields the lowest tracking accuracy because of inefficient resource utilization. The traditional optimization method, which minimizes the instantaneous tracking error, achieves moderate tracking performance but remains myopic. When the discount factor \begin{document}$ \gamma =0 $\end{document}, the D3QN algorithm achieves performance comparable to that of the traditional optimization method because both optimize only immediate rewards. In contrast, when \begin{document}$ \gamma =0.99 $\end{document}, the D3QN algorithm significantly improves full-time-horizon tracking accuracy. The resulting power allocation strategy allocates more transmit power to distant, low-Signal-to-Noise Ratio (SNR) targets at earlier stages while reducing redundant power assigned to nearby high-SNR targets. The training curves show that \begin{document}$ \gamma =0.99 $\end{document} achieves a higher steady-state cumulative reward, although convergence exhibits greater oscillation because of the increased difficulty of estimating long-term returns. Furthermore, the trained D3QN network generates transmit power allocation decisions almost instantaneously, whereas the traditional optimization method must solve a constrained optimization problem at every time step, providing a substantial real-time computational advantage.  Conclusions  This paper proposes an RL-driven power allocation algorithm for collocated MIMO radar multi-target tracking that overcomes the myopic behavior and high computational complexity of conventional optimization methods. The proposed algorithm constructs the state space and reward function using the PCRLB, models the power allocation process as an MDP, and solves it using the D3QN algorithm. Simulation results demonstrate that, with an appropriate discount factor (\begin{document}$ \gamma =0.99 $\end{document}), the proposed approach significantly improves full-time-horizon tracking accuracy. This improvement results from the agent’s ability to learn a long-term optimal policy that proactively allocates transmit power to future distant, low-SNR targets. Furthermore, the trained network enables real-time decision-making through direct forward propagation, substantially reducing computational latency compared with iterative optimization. This work provides a new approach for intelligent radar resource management in complex battlefield environments.
Decoupled Learning for Long-tailed Oracle Bone Character Recognition Based on Adaptive Difficulty Sampling
SUN Junwei, GUAN Suyan, CHEN Xinyu, WANG Kun, CAI Yuanqiang
Available online  , doi: 10.11999/JEIT260327
Abstract:
  Objective  Oracle Bone Character (OBC) recognition is challenged by an extreme long-tailed distribution and substantial intra-class variation. Conventional deep learning methods are often dominated by head classes, whereas existing approaches tend to overfit tail classes or fail to account for differences in learning difficulty across classes. To address these limitations, a two-stage decoupled learning framework is proposed to improve the recognition of tail and difficult classes while preserving the discriminative capability of head classes.  Methods  The proposed framework decouples feature representation learning from classifier optimization. In the first stage, the backbone network is trained using a mixed data augmentation strategy that combines CutMix and RandAugment with Label-Distribution-Aware Margin (LDAM) loss to learn robust feature representations and alleviate the effect of intra-class variation. In the second stage, the backbone network is frozen, and only the classifier is optimized. An adaptive difficulty sampling strategy is proposed to dynamically assign sampling weights according to historical and current class-level training difficulty. The classifier is further optimized using a Class-Balanced LDAM (CBL) loss, which combines class-balanced weighting with LDAM to refine decision boundaries for long-tailed classification.  Results and Discussions  Experiments on the highly imbalanced OBC306 dataset demonstrate that the proposed method achieves an overall accuracy of 94.34% and an average class accuracy of 89.89%. Compared with the Inception-v4 baseline, the proposed method improves the average class accuracy by 19.61%. Comparisons with representative long-tailed OBC recognition methods further demonstrate superior overall performance. Comprehensive ablation studies verify the effectiveness of the mixed data augmentation strategy and the adaptive difficulty sampling strategy in improving the recognition of rare and difficult characters. Parameter sensitivity analysis and qualitative error analysis further confirm the robustness and effectiveness of the proposed framework.  Conclusions  The proposed two-stage decoupled learning framework effectively addresses long-tailed OBC recognition by balancing the learning priorities of head, tail, and difficult classes. The mixed data augmentation strategy improves feature robustness, whereas the adaptive difficulty sampling strategy and the Class-Balanced LDAM loss jointly optimize classifier learning and refine decision boundaries without degrading head-class recognition performance. The proposed framework provides an effective solution for the digital recognition of Oracle Bone Characters and offers technical support for low-resource ancient character recognition.
Construction of a DNA Strand Displacement Memristor and Its Filter Circuit Characteristics
WANG Yanfeng, CHEN Guanzhou, SUN Ce, SUN Junwei
Available online  , doi: 10.11999/JEIT260283
Abstract:
  Objective  Filter circuits are widely used in modern control and signal-processing systems for noise suppression and signal integrity enhancement. Conventional Resistor-Capacitor (RC) filters are widely applied, but their fixed parameters limit adaptability and miniaturization in emerging molecular and nanoscale computing platforms. To address these limitations, DNA Strand Displacement (DSD) technology is integrated with memristor theory to develop tunable multistable molecular filter circuits. This study aims to design and validate first- and second-order low-pass filter circuits based on the dynamic response and state-dependent behavior of a DSD-based memristor. The proposed filters are designed to improve frequency selectivity, parameter adaptability, and system stability compared with traditional filter architectures. This approach is intended for molecular signal processing, integrated biocircuits, and adaptive filtering systems that require compact size and reconfigurability.  Methods  The method consists of four stages. First, core DSD reaction modules, including sine, cosine, integration, addition, and multiplication modules, are designed to construct a programmable multistable memristor model. Second, square-wave and sinusoidal input signals are generated through DSD reactions to evaluate the memristor response under different frequencies and amplitudes. Third, the memristor is embedded into low-pass filter structures to construct first- and second-order DSD-based memristor filter circuits. Fourth, simulations are performed using Visual DSD for molecular dynamics analysis and MATLAB for circuit-level analysis. Circuit performance is evaluated using transfer functions, Nyquist plots, Bode diagrams, and time-domain comparisons with classical RC filters. This combined simulation strategy verifies both molecular feasibility and circuit functionality.  Results and Discussions  The DSD-based memristor exhibits multistable behavior and converges to six stable equilibrium points under different initial conditions (Fig. 8). Its hysteresis characteristics further confirm the state-dependent memory behavior of the designed molecular memristor (Fig. 7). The first-order DSD-based memristor filter circuit provides stable attenuation for square-wave and sinusoidal input signals. Its output amplitudes are consistently higher than those of the traditional RC filter across the tested frequencies (Table 3). The second-order DSD-based memristor filter circuit further reduces signal delay and improves stability, especially under high-frequency inputs (Table 4). Frequency-response analyses show that the cutoff frequency can be dynamically tuned by adjusting DSD reaction rates and initial concentrations (Figs. 9 and 11). Time-domain simulations further confirm the filtering performance of the first- and second-order circuits (Figs. 10 and 12). Reliability analysis indicates that lower initial copy numbers increase stochastic molecular noise, whereas higher initial copy numbers make the output distribution closer to the deterministic response and improve the probability of successful filtering. These results verify the feasibility of DSD-memristor integration for adaptive molecular filtering.  Conclusions  A DSD-based memristor with multistable characteristics and its corresponding first- and second-order low-pass filter circuits are designed and validated. Compared with traditional RC architectures, the proposed filters show improved output stability, parameter tunability, and frequency adaptability. By combining DSD technology with memristor theory, this study provides a reconfigurable molecular-scale filtering framework for signal-processing applications. The results provide a basis for future work on adaptive molecular circuits, intelligent filtering, and nanoelectronic system design. Further studies should focus on experimental validation, real-time tuning strategies, sequence optimization, anti-interference design, signal amplification, and circuit integration.
Spatial-domain Anti-jamming for Unmanned Systems Under Limited Prior Information
PAN Zihao, ZHANG Bangning, ZHEN Pan, ZHU Bowen, WANG Ning, GUO Daoxing
Available online  , doi: 10.11999/JEIT260296
Abstract:
  Objective  Unmanned systems play an increasingly important role in emergency response, public safety, intelligent transportation, and other mission-critical applications. Reliable communications in complex electromagnetic environments are essential for autonomous operation. However, communication links are directly exposed to open, non-cooperative electromagnetic environments and are therefore vulnerable to intentional jamming and unintentional interference. In practical scenarios, prior information regarding the desired signal, jamming sources, and multipath propagation is often unavailable, substantially degrading the performance of conventional spatial-domain anti-jamming methods. To address this challenge, this paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information.  Methods  The proposed method first applies a spatial smoothing algorithm to the received signals to decorrelate coherent multipath components. Capon spatial spectrum estimation is then performed to detect the Direction Of Arrival (DOA) of potential incident signals. Spectrum peaks corresponding to individual incident signals are subsequently identified. A Covariance Matrix Reconstruction (CMR)-based beamforming algorithm is then applied by traversing all detected spectrum peaks to sequentially extract the signal associated with each peak, thereby separating the mixed signals. After signal separation, a signal classification method based on spectral similarity and time delay is employed. Kullback-Leibler (KL) divergence between the spectrum of each separated signal and the reference spectrum is calculated to identify jamming signals. The remaining communication signals are further classified into direct-path and multipath signals according to their relative time delays. Finally, different processing strategies are applied according to the identified signal type. Specifically, multipath signals are either suppressed as interference or coherently combined with the direct-path signal after time-delay and phase alignment.  Results and Discussions  Two simulation scenarios, including jamming only and combined jamming and multipath, are designed to evaluate the proposed method in terms of the output Signal-to-Interference-plus-Noise Ratio (SINR), beam pattern, Bit Error Rate (BER), and Error Vector Magnitude (EVM). Simulation results demonstrate that, under the jamming-only scenario, the proposed method achieves performance close to the theoretical optimum. The output SINR increases with the input Signal-to-Noise Ratio (SNR) at a fixed Jamming-to-Signal Ratio (JSR) (Fig. 3(a)) and remains nearly unchanged as JSR increases at a fixed SNR (Fig. 3(b)), indicating stable jamming suppression capability. The recovered time-domain waveform and spectrum remain highly consistent with the transmitted signal (Fig. 4). The BER curve nearly overlaps that of the optimal beamformer (Fig. 5). At \begin{document}$ {E}_{\rm b}/{N}_{0}=10\;{\mathrm{dB}} $\end{document}, the recovered Quadrature Phase-Shift Keying (QPSK) constellation closely matches the ideal constellation, achieving an EVM of –11.52 dB (Fig. 6). Under simultaneous jamming and multipath conditions, the proposed framework flexibly suppresses or exploits multipath signals. Compared with multipath suppression, multipath utilization further improves both the output SINR and BER (Fig. 7(a) and Fig. 7(b)). The corresponding beam pattern forms a beam toward the multipath direction rather than a null, demonstrating effective multipath exploitation (Fig. 7(c)).  Conclusions  This paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information. Using only the received mixed signals, the proposed framework estimates the directions of arrival, separates incident signals, and classifies them as direct-path, multipath, or jamming signals. Appropriate suppression or preservation strategies are then applied according to the identified signal type. Therefore, the framework flexibly suppresses or exploits multipath signals while preserving the direct-path signal and mitigating jamming. Simulation results demonstrate the effectiveness of the proposed method in terms of output SINR and demodulation accuracy, confirming reliable jamming suppression and communication performance even when prior information regarding the desired signal, jamming sources, and multipath propagation is unavailable. Future work will investigate the effects of array perturbations, intelligent jamming, and heterogeneous communication modes on the proposed framework and extend it to more complex unmanned-system communication environments.
A Nested Multi-scroll Memristive Hopfield Neural Network and Its Hardware Implementation
WANG Zhe, WAN Qiuzhen, ZHOU Pan, RAO Huhui
Available online  , doi: 10.11999/JEIT260516
Abstract:
  Objective  In recent years, memristors have been employed to emulate neuronal synapses with dynamically adjustable synaptic weights, enabling the construction of Memristive Hopfield Neural Networks (HNNs). Compared with conventional HNNs, Memristive HNNs more accurately reproduce the nonlinear dynamical behavior of biological neural systems. Multi-scroll attractors have attracted considerable attention in secure communication because of their complex topological structures and strong state-space ergodicity. However, previous studies have primarily focused on conventional multi-scroll attractors with single structural patterns, whereas multi-scroll attractors with special structures remain largely unexplored. Therefore, this paper proposes a nested multi-scroll Memristive HNN system that generates nested multi-scroll attractors, thereby overcoming the limitations of conventional single-structure multi-scroll attractors.  Methods  A Four-Dimensional (4D) Memristive HNN system is constructed from a three-neuron HNN by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of neuron 2. Equilibrium-point and stability analyses are performed to investigate the regulatory effects of the Memristive self-connected synapse coupling strength and system initial conditions on the system dynamics. The number of multi-scroll attractors is regulated by adjusting the memristor control parameters. Building on this framework, a Multi-level Logic Pulse current (IMLP) is introduced to construct a nested multi-scroll Memristive HNN system. The proposed system generates nested multi-scroll attractors with enhanced dynamical complexity. Finally, the MATLAB numerical simulation results are validated through Multisim circuit simulations and Field-Programmable Gate Array (FPGA)-based hardware experiments.  Results and Discussions  The results demonstrate that regulating the Memristive self-connected synapse coupling strength enables the proposed 4D Memristive HNN system to exhibit period-doubling bifurcations and chaotic behavior, as illustrated by the bifurcation diagrams and Lyapunov exponent spectra (Fig. 3). Various types of coexisting attractors are generated under different coupling strengths (Fig. 4). By adjusting the memristor control parameters, multi-scroll attractors with different numbers of scrolls are generated through one-directional extension (Figs. 58). After the introduction of the IMLP, the proposed nested multi-scroll Memristive HNN system generates nested multi-scroll attractors while preserving the controllable scroll-number extension property (Figs. 810). Spectral Entropy (SE) analysis demonstrates that the IMLP increases the dynamical complexity of the proposed system compared with the original 4D Memristive HNN system (Figs. 9 and 10). The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments confirms the physical realizability of the proposed nested multi-scroll Memristive HNN system (Figs. 1214).  Conclusions  A 4D Memristive HNN system is constructed by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of a three-neuron HNN. Equilibrium-point and stability analyses reveal the regulatory effects of the Memristive self-connected synapse coupling strength and the evolution of coexisting attractors associated with different initial conditions. The results show that the system enters chaos through the period-doubling route to chaos and generates single-scroll and double-scroll chaotic attractors. The number of multi-scroll attractors is continuously increased by adjusting the memristor control parameters. Furthermore, introducing the IMLP produces a nested multi-scroll Memristive HNN system capable of generating nested multi-scroll attractors with increased dynamical complexity. The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments validates the physical realizability of the proposed nested multi-scroll Memristive HNN system.
Research on Ka-band Enhanced Active Load Modulation Ultra-wideband High-efficiency Doherty Power Amplifier
YANG Lin, YAN Chengyu, WANG Yanping, ZHANG Ming, WANG Baozhu, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260514
Abstract:
  Objective  The Ka-band has become a key frequency band for satellite communications, placing stringent requirements on the millimeter-wave power amplifier, a core component of the transmitter, to provide high efficiency, compact size, and broadband operation. To maximize spectral efficiency, millimeter-wave satellite communication signals typically exhibit a high Peak-to-Average Power Ratio (PAPR), making high back-off efficiency particularly important. Although the Doherty Power Amplifier (DPA) is widely adopted because of its high efficiency under power back-off conditions, its operating bandwidth is inherently limited. In addition, both saturated efficiency and back-off efficiency degrade substantially at millimeter-wave frequencies. Therefore, extending the operating bandwidth while maintaining high efficiency remains a major challenge for millimeter-wave DPAs used in satellite communication transmitters.  Methods  An ultra-wideband enhanced active load modulation technique is proposed to overcome the trade-off between bandwidth and back-off efficiency in DPAs. The proposed method achieves optimal load impedance modulation for both the carrier and peaking amplifiers over an ultra-wide frequency range by introducing an Impedance Tunable Bias Network (ITBN) and a dual-drive impedance control mechanism. These techniques improve load modulation while extending the load modulation bandwidth, thereby enhancing both efficiency and bandwidth in the millimeter-wave DPA architecture. Furthermore, a broadband phase-compensation technique is integrated into an unequal power division network to achieve sufficient load modulation across the entire operating band while accurately compensating for the phase difference between the carrier and peaking paths. The proposed ultra-wideband phase-compensated power division network further extends the high-efficiency operating bandwidth while reducing the chip area.  Results and Discussions  To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was designed and fabricated using a 0.15-μm GaN process. Across the 24~33 GHz frequency band, corresponding to a relative bandwidth of 31.6%, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB with a gain flatness of ±1 dB. The measured saturated output power is 28.9~31.0 dBm, with a Power-Added Efficiency (PAE) of 25.2%~33.5% at saturation and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating over similar frequency bands, the proposed design achieves the highest reported small-signal gain and saturated PAE while maintaining high saturated output power and high 6 dB back-off PAE. Furthermore, it occupies the smallest chip area among reported three-stage DPA MMICs.  Conclusions  An ultra-wideband enhanced active load modulation method is proposed to achieve sufficient active load modulation through multi-frequency impedance tuning across the entire operating band. A novel ultra-wideband phase-compensated unequal power division network is also proposed to reduce the chip area while maintaining accurate phase compensation. To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was fabricated using a 0.15-μm GaN process. Measurement results demonstrate that, over the 24~33 GHz frequency band, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB, a saturated output power of 28.9~31.0 dBm, a saturated PAE of 25.2%~33.5%, and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating in similar frequency bands, the proposed design achieves a maximum relative bandwidth of 31.6%, the highest reported small-signal gain and saturated PAE, high saturated output power, and high 6 dB back-off PAE, while occupying the smallest chip area among reported two- and three-stage DPA MMICs. These measurement results validate the proposed method and demonstrate its strong potential for millimeter-wave satellite communication transmitters.
Iterative Parameter Estimation Method for Energy Detection Threshold in Ambient Backscatter
QU Wenfeng, YE Yinghui, SHI Liqin, LU Guangyue
Available online  , doi: 10.11999/JEIT260418
Abstract:
  Objective  In Ambient Backscatter Communication (AmBC) systems, a low-complexity Energy Detector (ED) is commonly employed at the reader to recover symbols transmitted by the tag. The detection performance of ED depends strongly on the accurate setting of the detection threshold, which is determined by the average received signal power corresponding to tag symbols “1” and “0”. Existing parameter estimation methods assume that the two symbols are transmitted with equal probability and therefore divide the sorted received signal power samples into two equal groups. However, because the number of transmitted symbols is finite, the actual numbers of symbols “1” and “0” are generally unequal. Therefore, equal partitioning introduces sample misclassification, causing the estimated threshold to deviate from its optimal value and reducing detection performance. To address this limitation, an iterative threshold parameter estimation method is proposed to reduce the parameter estimation bias caused by sample misclassification and improve the accuracy of detection threshold estimation.  Methods  An iterative threshold parameter estimation method is proposed to overcome the sample misclassification introduced by conventional sorting-based grouping. Because the initial detection threshold obtained by the sorting-based grouping method provides reliable decisions for most received samples, these initial decisions are used as the basis for sample reclassification. The received signal samples are then reclassified to iteratively update the threshold parameters, progressively refining the detection threshold. The proposed method is evaluated through simulations under three representative ambient radio-frequency source conditions: complex Gaussian, Phase-Shift Keying (PSK), and Quadrature Amplitude Modulation (QAM) sources.  Results and Discussions  Simulation results show that, over a wide range of Signal-to-Noise Ratio (SNR) values, the proposed iterative method substantially reduces the Bit Error Rate (BER) compared with the conventional sorting-based grouping method and approaches the theoretical lower bound obtained with perfect parameter estimation. At a given SNR, the proposed method improves BER by approximately 0.1, 1.5, and 1.3 orders of magnitude under complex Gaussian, PSK, and QAM sources, respectively (Fig. 2). These results demonstrate that the proposed iterative method effectively corrects sample misclassification and reduces the performance loss caused by parameter estimation bias. Moreover, most of the performance gain is achieved after only one iteration, indicating rapid convergence with minimal additional computational overhead. Under different numbers of sampling points, BER improvements of approximately 0.5, 1.6, and 1.1 orders of magnitude are achieved under complex Gaussian, PSK, and QAM sources, respectively (Fig. 3). These results indicate that using the initial decisions for sample reclassification effectively reduces the estimation bias introduced by fixed equal partitioning, thereby improving detection performance under limited-sample conditions. Under different Relative Channel Difference (RCD) values, BER improvements of approximately 0.56 and 0.7 orders of magnitude are achieved under complex Gaussian and QAM sources, respectively (Fig. 4). As the RCD increases, the separation between the received signal power distributions becomes more pronounced, improving the accuracy of the initial decisions and enabling more reliable sample reclassification. This positive feedback process further refines the parameter estimates and improves detection performance.  Conclusions  An iterative threshold parameter estimation method is proposed to address the sample misclassification introduced by conventional sorting-based grouping in ED. The proposed method uses the initial decisions to reclassify the received signal samples and iteratively update the threshold parameters. In addition, closed-form expressions for the detection threshold and BER under QAM sources are derived. Simulation results demonstrate that the proposed method effectively reduces parameter estimation bias with only one iteration while maintaining robust performance under limited-sample and varying channel conditions. Significant BER improvements are achieved with minimal additional computational overhead, making the proposed method well suited for practical, high-reliability AmBC systems.
Low-Complexity Phase Ambiguity Resolution DOA EstimationAlgorithm for Composite Hierarchical Receiving Array Structure
CHEN Yiwen, DONG Yangze, CHEN Xiahua, LING Wenchang, XIONG Yiwen
Available online  , doi: 10.11999/JEIT260447
Abstract:
  Objective  Direction Of Arrival (DOA) estimation is a key technique for sonar target localization. As the demand for high-precision DOA estimation in complex environments continues to increase, the number of array elements used for estimation is steadily growing, leading to massive arrays. Although larger arrays improve DOA estimation accuracy and resolution, they also impose a substantial computational burden on conventional DOA estimation algorithms. To address this issue, a low-complexity composite hierarchical receiving array structure is constructed, and two fast phase ambiguity resolution algorithms are proposed: Composite HierArchical Global Nearest-Neighbor Matching (CHA-GNNM) and Composite HierArchical Cross-Correlation Covariance Merging (CHA-CCM).  Methods  The CHA-GNNM algorithm constructs multiple candidate solution sets by exploiting the auto-covariance and cross-covariance relationships among the subarrays within each group. The true solution in each candidate solution set is identified through nearest-neighbor matching based on source consistency, and the final DOA estimate is obtained through multilevel coherent combining. This approach achieves phase ambiguity resolution and angle matching with relatively low computational cost. However, because the correlation information among all array elements is not fully exploited, some estimation performance is sacrificed. To improve DOA estimation performance, the CHA-CCM algorithm reorganizes the composite hierarchical structure into evenly partitioned groups, which are regarded as several large subarrays. Multiple large candidate solution sets are first constructed from the cross-correlation relationships among these groups. Each group is then divided into multiple small subarrays, from which additional candidate solution sets are generated using the corresponding auto-covariance and cross-covariance relationships. A coarse DOA estimate is obtained through coprime clustering, followed by a more accurate initial DOA estimate derived from the small candidate solution sets. This initial DOA estimate is subsequently used to eliminate spurious solutions from the large candidate solution sets, yielding the final DOA estimate. Combined with a low-complexity covariance block-processing strategy, this approach avoids computationally expensive operations while improving DOA estimation accuracy.  Results and Discussions  Simulation results demonstrate that both proposed algorithms substantially reduce the computational burden as the number of array elements increases, while effectively achieving phase ambiguity resolution through the proposed composite hierarchical receiving array structure (Fig. 5). Compared with the conventional Root-MUSIC algorithm, CHA-GNNM achieves coarse DOA estimation with nearly four orders of magnitude lower computational complexity (Fig. 7), making it suitable for applications with stringent real-time requirements. In contrast, CHA-CCM requires only a modest increase in computational cost (Fig. 7) while achieving DOA estimation performance close to the Cramér-Rao Lower Bound (CRLB) above a certain signal-to-noise ratio threshold (Fig. 6). Therefore, a favorable balance is achieved between DOA estimation accuracy and computational complexity.  Conclusions  To address the rapid increase in computational complexity associated with massive arrays, a composite hierarchical receiving array structure is constructed for efficient DOA estimation. By hierarchically grouping the array elements, the proposed structure provides a new framework for low-complexity DOA estimation. Based on this structure, two fast DOA estimation algorithms are developed. Both algorithms achieve effective phase ambiguity resolution with low computational complexity by exploiting the structural differences among array groups and the consistency of observations from the same source across different groups, thereby enabling rapid DOA estimation. CHA-GNNM primarily exploits the phase relationships among subarrays to perform phase ambiguity resolution and angle matching through a simple computational procedure, making it suitable for applications requiring high computational efficiency and real-time processing. Because the cross-correlation information among all array elements is not fully exploited, some estimation performance is reduced under challenging signal conditions. To overcome this limitation, CHA-CCM reorganizes the composite hierarchical receiving array into evenly partitioned groups while preserving the low-complexity advantage of the hierarchical structure. Group-level cross-correlation information is further exploited so that the intrinsic relationships among different groups are more fully utilized. In addition, the signal processing procedure is simplified by eliminating unnecessary computational steps, thereby improving the robustness and accuracy of DOA estimation while maintaining manageable computational complexity. Compared with CHA-GNNM, CHA-CCM incurs only a small increase in computational cost and achieves a better balance between computational complexity and DOA estimation performance. Overall, the proposed composite hierarchical receiving array structure and the two fast DOA estimation algorithms provide an effective solution for efficient DOA estimation in massive arrays. CHA-GNNM is more suitable for applications with stringent real-time requirements, whereas CHA-CCM is better suited for applications requiring higher DOA estimation accuracy and robustness. The proposed structure achieves efficient phase ambiguity resolution and accurate DOA estimation and provides both theoretical significance and practical value for engineering applications of massive array signal processing.
Off-grid Blind Near-Field Integrated Sensing And Communication: Algorithm Design and Lower Bound
YUAN Zhengdao, GUO Qinghua, HUANG Chongwen, GAO Dawei, MEI Fengtong, LIAO Guisheng
Available online  , doi: 10.11999/JEIT260404
Abstract:
  Objective  With the widespread deployment of extra-large-scale antenna arrays in 6G networks, user terminals are increasingly located in the near-field region. Existing Near-Field Integrated Sensing And Communication (NF-ISAC) algorithms face key challenges, including off-grid power leakage, severe model mismatch, and strong pilot dependence. These limitations make them unsuitable for low-overhead, high-performance 6G transmission. This paper aims to design an off-grid blind NF-ISAC algorithm and derive the theoretical performance bound for near-field sensing.  Methods  To overcome the limitations of analytical geometric steering vectors and adapt to more accurate electromagnetic propagation characteristics without closed-form expressions, an amplitude-phase separation method is first proposed. This method decomposes the nonlinear near-field steering vector into amplitude and phase terms, enabling high-precision characterization of the steering vector using a single-hidden-layer neural network. Second, the NF-ISAC problem is formulated as a constrained matrix factorization problem, and a corresponding factor graph model is constructed. The trained neural network is embedded into the factor graph as a function node. Message passing through the embedded neural network is then achieved, enabling joint blind coordinate sensing, channel estimation, and signal detection in a pilot-free manner. Finally, the Cramér-Rao Lower Bound (CRLB) for multi-user near-field joint distance and angle sensing in polar coordinates is derived based on the neural-network-fitted steering vector.  Results and Discussions  Extensive Monte Carlo simulations are conducted to evaluate the performance of the proposed algorithm. The simulation results show that the proposed algorithm achieves millimeter-level position sensing. Compared with existing mainstream algorithms, it improves both communication Bit Error Rate (BER) and sensing accuracy. The proposed algorithm achieves a 2~3 dB gain in sensing accuracy over the state-of-the-art near-field off-grid algorithm, and its performance is closest to the derived theoretical CRLB. These results indicate that the proposed algorithm effectively mitigates off-grid power leakage and model mismatch.  Conclusions  The proposed off-grid blind NF-ISAC algorithm overcomes the pilot dependence and model mismatch of existing NF-ISAC schemes. It achieves integrated high-precision sensing and reliable communication for near-field users in a pilot-free manner. The derived CRLB provides a theoretical benchmark for evaluating the sensing performance of NF-ISAC systems. This work provides technical support for the design of 6G NF-ISAC systems.
Research on Adaptive Hybrid Beamforming Method for Massive MIMO LEO Satellite Communication Systems
XIAN Yongju, HUANG Xiaolong, XING Zhitong, LI Yun
Available online  , doi: 10.11999/JEIT260458
Abstract:
  Objective  With the growing demand for high-capacity and high-spectral-efficiency transmission in LEO satellite communications, massive MIMO has become a promising enabling technology. However, conventional fully digital beamforming is difficult to implement in practice due to the strict constraints on power consumption, hardware complexity, and payload cost of satellite platforms. Although partially connected hybrid beamforming can reduce hardware complexity, the conventional fixed subarray structure lacks flexibility and suffers from performance degradation, especially under low-resolution PSs constraints. To address this issue, this paper investigates adaptive antenna-RF chain mapping for hybrid beamforming design in LEO satellite massive MIMO systems.  Methods  This paper first establishes a system model for LEO satellite multi-user downlink massive MIMO hybrid beamforming and formulates a joint optimization problem with the objective of maximizing spectral efficiency. Considering the constant-modulus discrete phase constraints of low-resolution PSs, antenna-RF chain connection constraints, and transmit power constraint, the resulting problem is highly non-convex. To solve it, the original problem is transformed into an equivalent WMMSE formulation, and auxiliary variables are introduced to decouple the coupled variables. Based on this reformulation, a double-layer iterative optimization framework is developed by combining the BCD method and the PDD method. For adaptive antenna-RF chain mapping, the mapping problem is reformulated as a capacity-constrained linear assignment problem, and an optimal adaptive mapping method based on the Hungarian algorithm is proposed. Furthermore, to reduce the computational burden in large-scale antenna array scenarios, a low-complexity adaptive mapping method based on antenna priority sorting is developed.  Results and Discussions  Simulation results show that the proposed methods exhibit good convergence behavior. Specifically, the spectral efficiency increases rapidly in the initial iterations and then gradually converges, while the constraint violation decreases continuously, confirming the effectiveness of the proposed iterative optimization framework (Fig. 3). In terms of spectral efficiency, the proposed adaptive mapping methods consistently outperform the conventional fixed subarray and the existing greedy dynamic subarray scheme over different transmit powers and antenna scales. Among them, the Hungarian-based method achieving better spectral efficiency, whereas the antenna priority sorting-based method attains near-optimal performance with significantly reduced computational complexity (Figs. 4 and 5). As the number of PSs quantization bits increases, the system performance gradually approaches that of the continuous-PSs case, demonstrating the effectiveness of the proposed low-resolution PSs-based design (Fig. 6). In terms of energy efficiency, the proposed methods also outperform the conventional fully digital beamforming, fully connected, fixed subarray, and greedy dynamic subarray hybrid beamforming under different transmit powers and antenna scales (Figs. 7 and 8).  Conclusions  This paper proposes an adaptive hybrid beamforming design for LEO satellite massive MIMO systems under low-resolution PS constraints. By combining WMMSE reformulation with a PDD-BCD based optimization framework, joint design of digital precoding, analog precoding, and adaptive antenna-RF chain mapping is achieved. Simulation results demonstrate that the proposed methods provide superior performance in both spectral efficiency and energy efficiency. In particular, the Hungarian-based method provides better system performance, while the antenna priority sorting-based method achieves a favorable trade-off between performance and computational complexity. The proposed design provides an effective solution for high-performance hybrid beamforming in hardware-constrained LEO satellite massive MIMO systems.
FedFACO: Personalized Federated Learning Method Based on Fisher Information Matrix for Adaptive Aggregation and Client Collaborative Optimization
JIANG Wei-Jin, LIU Zhi-Hua, CUI Xin-Yu, XU Yu-Sheng, CHEN Shen-You, HU Jia-Long
Available online  , doi: 10.11999/JEIT260344
Abstract:
  Objective  Non-IID data heterogeneity remains one of the major challenges in personalized federated learning, as it often leads to inconsistent local optimization directions, insufficient global knowledge transfer, and degraded model personalization performance. To address these issues, this paper proposes FedFACO, a personalized federated learning method based on Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The proposed method aims to improve the adaptability of federated models to heterogeneous client distributions while maintaining effective knowledge sharing across clients. By introducing an adaptive aggregation mechanism and a collaborative optimization strategy, FedFACO provides a more principled way to balance global generalization and local personalization, which is particularly important in complex Non-IID federated environments.  Methods  FedFACO consists of two key components. First, an adaptive aggregation (AA) mechanism is employed to dynamically adjust fusion weights between global and local models based on client-specific states, generating personalized initializations aligned with local data. Second, a collaborative optimization (CO) mechanism is introduced, combining feature alignment with FIM-based client weighting. This enhances useful global knowledge transfer and suppresses low-quality updates. The FIM is utilized to quantify the information contribution of each update, ensuring reliable aggregation. The method is evaluated on MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet under Non-IID settings, and compared with representative baselines. Convergence behavior, dropout robustness, and sensitivity to low-quality updates are also examined.  Results and Discussions  Experimental results demonstrate that FedFACO consistently outperforms competitive baseline methods across all four benchmark datasets, achieving an average accuracy improvement of approximately 3.1% over mainstream approaches (Fig. 1, Table 2). On the more challenging Tiny-ImageNet dataset, FedFACO reduces the total training time required to reach convergence by approximately 4.8% compared with the best baseline (Table 3). Ablation studies confirm that performance is substantially improved by both AA and CO mechanisms, with their joint application yielding optimal accuracy (Table 6). Furthermore, FIM-guided weighting is shown to accurately quantify contribution quality in client dropout and asynchronous scenarios, significantly enhancing aggregation reliability (Fig. 2Fig. 3). Superior robustness is also demonstrated in malicious client scenarios (Fig. 4).  Conclusions  This paper presents FedFACO, a personalized federated learning method for Non-IID environments via Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The method effectively balances global knowledge sharing and local personalization while enhancing training stability and robustness under heterogeneous client participation. Experimental results validate its effectiveness and superiority in accuracy, convergence efficiency, and robustness. Future work will focus on lightweight Fisher information approximations and adaptive triggering strategies to reduce computational overhead, as well as integration with privacy-preserving and security-defense mechanisms for deployment in resource-constrained and high-security environments.
A High-Parallelism Simulated Adiabatic Bifurcation Processor for Combinatorial Optimization Problems
LI Renlong, HAO Xin, CHEN Zhuojun, DING Ding
Available online  , doi: 10.11999/JEIT260780
Abstract:
  Objective  Combinatorial optimization problems (COPs) are widely encountered in fields such as network optimization, autonomous driving path planning, VLSI design, and computational biology. The solution space of these problems grows exponentially with problem size, making it impossible for traditional von Neumann architectures (e.g., CPUs) to find high-quality solutions within a feasible time. Quantum-inspired Ising machines have emerged as a promising computing paradigm for accelerating COP solving. However, existing CMOS-based Ising machines still face significant challenges in simultaneously achieving high speed, high energy efficiency, and high accuracy. Discrete-time Ising machines often require thousands of iterations and rely heavily on random number generators, leading to large chip area and long solving times. Continuous-time Ising machines suffer from poor solution quality, frequently getting trapped in local minima. To address these limitations, this paper designs and fabricates a high-parallelism simulated adiabatic bifurcation application-specific processor for combinatorial optimization problems in 65 nm CMOS technology.  Methods  The proposed processor adopts the simulated adiabatic bifurcation (SAB) algorithm, which is inspired by quantum adiabatic optimization. Unlike simulated annealing, SAB does not require Gibbs sampling or random number generators to escape local minima. The algorithm models each spin as a nonlinear oscillator and solves a set of ordinary differential equations to simulate the adiabatic evolution of a classical nonlinear Hamiltonian system exhibiting bifurcation phenomena. The processor builds a hardware architecture that supports fully connected spin topologies and leverages the inherent fully parallel spin update characteristic of the SAB algorithm. To achieve bubble-free iterative computation, a three-stage pipelined spin update strategy is proposed, dividing the update process into coupling coefficient access, momentum update, and position update. To reduce the storage overhead introduced by the fully connected coupling matrix, a folded coupling coefficient storage array is designed, exploiting matrix symmetry to eliminate redundant storage. The momentum update unit and position update unit are implemented using 8-bit fixed-point arithmetic (2 bits for integer, 6 bits for fractional part) to achieve SAB evolution with low hardware overhead. The chip is fabricated in 65 nm CMOS technology, occupying an area of 0.47 × 1.25 mm2 and operating at a 200 MHz clock frequency and 1 V supply voltage.  Results and Discussions  The chip achieves a total power consumption of only 14 mW, with the spin evolution module consuming 58% of the total power. The folded coupling coefficient storage array reduces storage area by 52% compared to full matrix storage, and the rectangular restructuring avoids irregular shapes in physical layout, reducing routing congestion and layout voids (Fig.6). The parallel loading access mechanism allows all coupling coefficients to be read and distributed within a single clock cycle, eliminating the memory access bottleneck inherent in serial reading. For predefined Max-Cut problems configured as 8×8 grid structures (64 nodes) with coupling coefficients quantized to 2-bit precision, the chip converges to the global optimum within only 5 computing cycles, achieving a final Ising energy of –4032 (Fig.9). This energy follows the analytical expression (n4−n2)(n4−n2), confirming that the chip solves Max-Cut problems with the shortest solving time. For larger extended Max-Cut problems (108×108 nodes), the simulated adiabatic bifurcation algorithm achieves 100% accuracy relative to the theoretical ground state (Fig.10). Monte Carlo simulations over 1,000 independent trials on randomly generated Max-Cut problems demonstrate that simulated adiabatic bifurcation achieves an average Hamiltonian of –3617.42, significantly outperforming simulated annealing which achieves only –3352.12 (Fig.11). For 3-SAT problems with 40 clauses and 8 variables, the chip solves instances with clause-to-variable ratios of 3, 4, and 5 in approximately 6 μs, 10 μs, and 16 μs, respectively (Fig.12). These results align perfectly with theoretical phase transition predictions, confirming the processor's effectiveness across varying problem complexities.  Conclusions  This work proposes an simulated adiabatic bifurcation machine that enables fully parallel updates without duplicating spin copies. To improve throughput, a three-stage pipeline strategy is designed that integrates coupling-coefficient access, momentum update, and position update, achieving bubble-free parallel updating and low-latency solving. For sparse coupling coefficients, a folded storage scheme is adopted to significantly reduce memory area overhead. Both momentum and position variables are represented in 8-bit fixed-point format, ensuring sufficient computational accuracy while balancing resource efficiency. Compared with previous fully connected Ising machines, the proposed bifurcation machine achieves 100% solving accuracy, along with higher energy efficiency and lower hardware overhead, demonstrating substantial application prospects in edge-side combinatorial optimization.
A State Prediction Method for Long-Endurance Fixed-Wing UAV Propulsion Systems
LI Sicheng, WANG Lianqing, LI Zhiyong, WANG Guochang, GE Kaihua, CHEN Junfeng, TAN Rongqing
Available online  , doi: 10.11999/JEIT260188
Abstract:
  Objective  Accurate single-step prediction of key propulsion-system states is essential for early fault warning and autonomous health management of long-endurance fixed-wing unmanned aerial vehicles (UAVs). During high-altitude missions lasting more than 24 h, electrical and thermal variables in the propulsion system exhibit strong coupling, multi-time-constant dynamics, and pronounced day-night regime shifts. These characteristics cause short-term disturbances and long-term drifts to coexist, and hinder adaptive feature weighting under time-varying variable sensitivities. General-purpose time-series predictors may therefore fail to meet the accuracy and robustness requirements of multivariate propulsion-state prediction. To address these challenges, a Grouped Squeeze-and-Excitation Multi-scale Temporal Convolutional Network (GEMS-TCN) is developed by enhancing a modern pure-convolution forecasting backbone with multi-scale embedding and grouped channel attention. The aim is to obtain accurate 10 s-ahead single-step forecasts for 16 key propulsion states from 72-dimensional flight telemetry while satisfying the real-time inference requirement of the 1 Hz telemetry cycle.  Methods  Real flight telemetry from a representative long-endurance fixed-wing UAV is used for model construction and evaluation. The data are sampled at 1 Hz for 9 consecutive days, yielding 806,629 time steps and 72 variables. (Fig.2) Sixteen propulsion-related key states, including the bus voltage, control-unit temperature, winding temperature, and power-device temperature of four motors, are selected as prediction targets, and all 72 variables are used as inputs. (Table 1) The raw data are processed through time indexing, interquartile range (IQR)-based anomaly handling, and interpolation, and are then chronologically divided into training, validation, and test sets at a ratio of 8:1:1 to avoid information leakage. (Fig.4) (Fig.5) GEMS-TCN uses a multi-scale embedding layer with parallel one-dimensional convolutions to extract temporal patterns at different receptive-field scales. Stacked GEMS-TCN blocks combine depthwise temporal convolution, grouped convolutional feed-forward networks, and Grouped Squeeze-and-Excitation (GroupSE) modules to recalibrate intra-variable and cross-variable channel responses hierarchically. (Fig.1) The models are trained with Adam and mean squared error (MSE) loss, and are evaluated using mean absolute error (MAE), MSE, and symmetric mean absolute percentage error (SMAPE). PatchTST, ModernTCN, FEDformer, DLinear, and TimesNet are used as comparison models, with ablation and robustness experiments conducted for further verification.  Results and Discussions  On the full 16-dimensional target set, GEMS-TCN achieves a test-set MAE of 0.171 and an MSE of 0.069. (Table 4) Compared with TimesNet, the strongest baseline in overall trend tracking, GEMS-TCN reduces MSE by 28.1% while maintaining comparable MAE and SMAPE, indicating stronger suppression of large prediction deviations. (Table 4) Stable accuracy is obtained in both daytime and nighttime segments, with MAE/MSE values of 0.189/0.085 and 0.150/0.051, respectively, demonstrating robustness to diurnal operating-condition changes. (Table 4) The prediction trajectories show tighter alignment at thrust-transition points, reduced overshoot, and fewer spurious spikes, while low-bias tracking is preserved for slowly varying nighttime temperature profiles. (Fig.7) Ablation results show that multi-scale embedding and GroupSE provide complementary improvements, and their combination achieves the best overall performance among the ablation settings. (Table 6) Under 1% and 5% random dropouts and 60 s continuous missing intervals, the MSE remains below 0.07, indicating tolerance to practical data-loss scenarios. (Table 5) In addition, GEMS-TCN contains 27.88 M parameters and achieves an inference latency of 0.90 ms per sample, which is well below the 1 s sampling interval.  Conclusions  GEMS-TCN provides a practical convolution-based solution for multivariate propulsion-state prediction in long-endurance fixed-wing UAVs. By integrating multi-scale temporal embedding with hierarchical group-wise channel recalibration, the proposed method jointly represents rapid fluctuations and slow evolution, and better captures multi-time-constant dynamics and multivariate coupling in propulsion telemetry. Real-flight experiments demonstrate stable prediction performance across diurnal regimes, state categories, and data-missing scenarios. Ablation results further confirm the effectiveness and complementarity of multi-scale embedding and GroupSE. These findings indicate that structure-aware modeling tailored to propulsion-state evolution can support health monitoring, early fault warning, and autonomous health management of long-endurance fixed-wing UAVs.
Customized Preparation of Single Crystal Tungsten Tips and Surface Reconstruction Mechanism
GUO Jiamei, YIN Shengyi, ZHANG Yongqing, SUN Wanzhong
Available online  , doi: 10.11999/JEIT260328
Abstract:
  Objective  Refractory metal tungsten, particularly single crystal tungsten, serves as a critical material for high-performance field emission cathodes, which are core components in advanced electron microscopes and electron beam lithography systems. The fabrication of single crystal tungsten microtips with precisely controlled geometry and surface cleanliness remains a significant challenge, especially for applications requiring tip radii ranging from sub-100nm for cold field emission to 0.3–1.0 μm for Schottky-type thermal field emission. Currently, the domestic development of high-end electron microscopes in China faces a major bottleneck due to heavy reliance on imported field emission cathodes. Electrochemical corrosion, the mainstream method for tip preparation, suffers from limitations such as difficulty in cutoff timing control, susceptibility to tip bending or passivation, residual surface impurities, and low yield rates. Moreover, the anisotropic nature of single crystal tungsten introduces additional complexity in morphology control. This study aims to establish a controllable fabrication method for single crystal tungsten tips, enabling tailored geometry and surface quality to meet demanding requirements of different field emission applications while significantly improving fabrication yield.  Methods  Single crystal tungsten wires (diameter 0.12 mm, purity 99.95%, (100) orientation) were used as the starting material. The tips were first pre-shaped by electrochemical corrosion in 1 mol/L NaOH solution under a 10 V DC voltage with pulsed control (6 kHz, 100 μs pulse width) for 10 min. Following corrosion, the tips underwent surface cleaning by sequential immersion in ultrapure water and anhydrous ethanol, followed by drying with high-purity nitrogen. Subsequently, the samples were placed in an ultra-high vacuum system (base pressure < 1 × 10–6 Pa) and subjected to high-temperature oxygen treatment. Process parameters were systematically varied, including temperature (15001600 °C), oxygen flow rate (0.015–1.0 sccm), and treatment duration (5–1440 min). During treatment, high-purity oxygen was introduced while maintaining vacuum levels better than 10–4 Pa. Tip morphologies were characterized by scanning electron microscopy, and surface compositions were analyzed by energy dispersive X-ray spectroscopy. Geometric parameters, including half-cone angle and tip radius, were measured following standardized protocols.  Results and Discussions  The combination of electrochemical corrosion and high-temperature oxygen treatment enabled both effective surface purification and controlled morphological reconstruction. Scanning electron microscopy characterization revealed distinct evolutionary pathways depending on oxygen partial pressure (Fig. 3, Fig. 4). Under oxygen-rich conditions (1.0 sccm, 15001600 °C), the tips underwent “blunting reconstruction,” evolving from an initial inverted cone into a characteristic “cylindrical segment + hemispherical cap” structure (Samples 2–5, Table 2). For Sample 3 treated at 1500 °C for 10 min, the tip radius increased to 327 nm with a half-cone angle reduction of 21%; for Sample 4 treated for 15 min, further evolution occurred with tip radius decreasing to 275 nm and cylindrical segment height increasing substantially. With increasing temperature from 1500 °C (Sample 2) to 1600 °C (Sample 5), the half-cone angle change rate increased from 11% to 25%, while tip radius decreased from 380 nm to 172 nm, indicating accelerated kinetics. This anisotropic behavior is attributed to orientation-dependent surface energies of tungsten and preferential oxidation along specific crystallographic planes. Under oxygen-lean conditions (0.015 sccm, 1500 °C, 1440 min), “sharpening reconstruction” occurred, with tip radius decreasing dramatically from 143 nm to 28 nm and half-cone angle reducing from 8° to 1° (Sample 6). In contrast, treatment in vacuum at 1500 °C primarily removed surface contaminants without altering tip geometry (Sample 1). EDS analysis confirmed that treated surfaces were free from impurities other than trace carbon contamination (Table 3), demonstrating the dual functionality of the process in achieving both purification and reconstruction. The findings reveal that oxygen partial pressure serves as the key determinant of reconstruction direction, with oxygen-rich conditions favoring blunting and oxygen-lean conditions promoting sharpening.  Conclusions  A combined process of electrochemical corrosion followed by high-temperature oxygen treatment was successfully developed for the controllable fabrication of single crystal tungsten field emitter tips. The process achieves dual functionality: effective surface purification through high-temperature vacuum treatment, which removes residual surface contaminants from electrochemical corrosion, and controlled morphological reconstruction enabled by the introduction of oxygen under precisely regulated conditions. By systematically adjusting process parameters including treatment temperature, oxygen flow rate, and treatment duration, tip morphology and dimensions can be tailored to meet specific application requirements. The oxygen partial pressure plays a decisive role in determining the reconstruction pathway. Under oxygen-rich conditions (e.g., 1500 °C, 1.0 sccm), blunting reconstruction yields a “cylindrical segment + hemispherical cap” structure ideally suited for Schottky-type thermal field emission cathodes requiring tip radii between 0.3 μm and 1.0 μm. Under oxygen-lean conditions (e.g., 1500 °C, 0.015 sccm), sharpening reconstruction produces nanoscale sharp tips with tip radius as low as 28 nm, suitable for cold field emission cathodes and scanning tunneling microscope probes. This anisotropic reconstruction mechanism is explained by the selective reaction of oxygen atoms with different tungsten crystal planes under high-temperature conditions, where the formation of volatile oxides modifies local surface energy distribution and promotes the exposure of specific crystallographic planes. Control experiments confirmed that such morphological reconstruction does not occur under oxygen-lean or oxygen-free conditions at the same temperatures, further demonstrating the critical role of oxygen in triggering this process. Importantly, this approach provides an effective method to correct imperfections from the initial electrochemical corrosion step, significantly improving fabrication yield and process robustness, thereby offering a viable technical pathway for the domestic fabrication of high-performance field emission cathodes.
Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding
ZHOU Tong, CHEN Hongzhi, XU Jialong, SUN Peng, JIANG Dajie, LIU Jiankang
Available online  , doi: 10.11999/JEIT260379
Abstract:
Currently, the 3rd Generation Partnership Project (3GPP) Technical Specification Group Service and System Aspects Working Group 4 (TSG SA Working Group 4, SA4) has begun researching Artificial Intelligence (AI)-based audio codecs. Based on the Descript Audio Codec (DAC) being studied by SA4, this paper presents a joint optimization scheme for DAC source-channel coding and modulation with total power constraints, a joint optimization scheme for DAC source-channel coding and modulation with constant modulus constraints, and a DAC joint source-channel coding scheme. Furthermore, simulation evaluation is conducted using the typical physical layer parameter configuration of 3GPP Geostationary Earth Orbit (GEO) voice scenario. Finally, preliminary research suggestions are made, with the hope of providing useful insights for the subsequent research on GEO voice.  Objective  In recent years, Joint Source-Channel Coding (JSCC) has gained widespread attention in academia and industry. In academia, the sources studied for JSCC include video, image, and audio. However, the industry has not focused on these sources as in academia, but rather on JSCC where channel state information (CSI) is considered as the source. The reason the industry has not considered JSCC for video and images is due to two main factors. First, for video and images, 3GPP generally does not conduct independent research, but instead directly reuses source compression standards developed by external organizations such as the Moving Picture Experts Group (MPEG) and the Joint Photographic Experts Group (JPEG). Second, it is constrained by the challenges of the source-channel coding architecture. For example, in joint source-channel coding at the application layer, the channel information obtained at the application layer is often outdated, which limits the potential gains of source-channel coding. However, the situation is quite different for audio. Audio coding and decoding are designed independently by 3GPP, and currently, 3GPP SA4 is researching AI-based speech coding and decoding. Moreover, in GEO speech scenarios, where satellites remain largely stationary relative to ground stations and the channel changes slowly, the physical layer CSI can be transmitted to the application layer, overcoming the issue of outdated channel information. This makes the research on JSCC for audio more promising within 3GPP, compared to JSCC for video or image sources.  Methods  This paper firstly uses the typical physical layer parameter configurations of 3GPP GEO voice and the DAC encoder, currently being studied by SA4, as an evaluation baseline, ensuring the fairness of the proposed scheme’s gain evaluation. Based on this, the paper proposes the integration of DAC with channel coding and modulation to improve audio quality in low-bitrate scenarios. First, the total power constrained (TPC) DAC-JSCCM scheme is considered, which offers the maximum potential gain due to the joint optimization of source coding, channel coding, and modulation. Then, considering the high PAPR (Peak-to-Average Power Ratio) impact on the power amplifier, the paper proposes a constant-envelope constrained (CEC) DAC-JSCCM scheme. Finally, since the modulation symbols included in RTP payload require significant changes to the existing protocol, a DAC-JSCC scheme that is more compatible with current protocols is proposed, where bit sequences are included in RTP payload.  Results and Discussions  In the evaluation of baseline scheme 1, the link-level simulation at the physical layer is first conducted to obtain the relationship between SNR and BLER for an input length of 64 bits using a 1/3 Turbo code. Then, assuming that the error rate of an audio frame is the same as the BLER, the transmission error rate of an RTP packet containing L audio frames is calculated. According to the existing processing logic, if one audio frame in an RTP packet is transmitted with an error, the entire RTP packet will be discarded. In real-time communication, VoIP, or speech enhancement scenarios, a PESQ score of 2.5 is generally considered the minimum usable threshold, and voice quality below this score is regarded as unsuitable for professional or commercial services. Therefore, in GEO voice calls, we choose PESQ = 2.5 as the reference point. When PESQ = 2.5, the TPC DAC-JSCCM and CEC DAC-JSCCM provide a coverage gain of approximately 4 dB compared to baseline scheme 1 (Fig.4). The DAC-JSCC scheme offers a coverage gain of 2.5 dB compared to baseline scheme 1 (Fig.4). The coverage gains of DAC-JSCCM and DAC-JSCC come from the application layer performing joint source-channel coding, which avoids the cliff effect caused by using Turbo coding at the physical layer. When the Complementary Cumulative Distribution Function (CCDF) of Peak-to-Average Power Ratio (PAPR) is \begin{document}$ {10}^{-3} $\end{document}, the TPC DAC-JSCCM is 4 dB higher than the QPSK modulation, with sharper signal peaks, higher requirements for power amplifier linearity, and is more prone to distortion. In contrast, the CEC DAC-JSCCM is very close to the QPSK modulated PAPR.  Conclusions  This paper proposes the Total Power Constrained DAC-JSCCM, Constant Modulus Constrained DAC-JSCCM, and DAC-JSCCM schemes, based on the DAC codec currently being researched by SA4. Simulations are conducted using the typical physical layer configuration of the existing 3GPP GEO voice scenario. The simulation results show that, when the PESQ is 2.5 (the minimum usable threshold), compared to DAC combined with existing physical layer transmission technologies, the proposed DAC-JSCCM provides a coverage gain of 4 dB, while the proposed DAC-JSCC scheme provides a coverage gain of 2.5 dB. SA4 is currently researching more advanced GEO voice codecs beyond DAC, and we will extend our research by combining better GEO voice codecs in the future.
Lightweight Semantic Communication System Driven by User Personalization in UAV Networks
WEI Yuxuan, CHEN Xiao, CHEN Qiuyu, JIANG Hao, YANG Zhaohui
Available online  , doi: 10.11999/JEIT260370
Abstract:
  Objective  With the rapid development of the low-altitude economy and 6G intelligent networks, Unmanned Aerial Vehicle (UAV) image communication shows strong potential in target reconnaissance, emergency communication, and intelligent inspection. However, conventional pixel-level transmission cannot meet the requirements of efficient, low-latency, and intelligent communication because UAVs are constrained by limited bandwidth, payload capacity, and onboard computational resources. Semantic communication, which transmits only task-relevant information, provides an effective solution for improving communication efficiency in resource-constrained scenarios. However, current studies on UAV image transmission face several challenges. First, fixed network architectures use unified semantic encoding and transmission strategies for all users and cannot adapt to different personalized requirements. Second, new user access usually requires interest pre-training or model fine-tuning, which increases deployment overhead. Third, most models have high computational complexity. To address these issues, this paper proposes the Lightweight Personalized UAV Semantic Communication (LPUSC) system to balance computational cost, transmission bandwidth, and personalized requirements. The system enables personalized transmission through low-overhead semantic index interaction and a lightweight semantic extraction module, without pre-training for new users. A dual-branch end-to-end network is also designed. In this network, the semantic index transmission network works with the semantic image transmission network trained by a weighted hybrid loss function, thereby supporting high-precision and high-quality transmission of personalized semantic images.  Methods  The proposed LPUSC system adopts a dual-branch architecture for accurate task-driven semantic content transmission. In the semantic index interaction branch, the lightweight object detection model YOLO11s is used to perform semantic perception on UAV-captured visual scenes. Complex image information is compressed into low-dimensional semantic index vectors, which reduces transmission redundancy and communication overhead. On this basis, an end-to-end semantic index transmission network is designed to improve the robustness of semantic index transmission under complex wireless channel conditions. Through the semantic index interaction mechanism, the system accurately identifies targets of user interest and provides prior guidance for subsequent semantic content extraction. In the semantic image transmission branch, the lightweight and high-precision MobileSAM model is adopted for semantic region extraction. This branch uses the target bounding boxes returned by the semantic index interaction branch as box-prompt inputs, enabling pixel-accurate segmentation and extraction of specific semantic targets. To further improve semantic image reconstruction quality, a weighted hybrid loss function is designed. This function integrates Mean Squared Error (MSE), L1-norm loss, Structural Similarity Index Measure (SSIM) loss, gradient loss, perceptual loss, and background suppression loss. These losses jointly optimize pixel accuracy, structural preservation, and fine-detail restoration. Through the joint constraints of multiple loss terms, the proposed system improves semantic region reconstruction and achieves high-quality semantic image transmission.  Results and Discussions  Simulation results validate the proposed LPUSC system in semantic extraction and end-to-end transmission. For semantic extraction, three schemes are compared: YOLO11s-seg, YOLO11s + Segment Anything Model (SAM), and YOLO11s + MobileSAM (Fig. 4). The results show that the detection-segmentation decoupled architecture achieves better semantic boundary localization accuracy. Combined with the quantitative analysis in Table 1, the YOLO11s + MobileSAM scheme reduces resource use while maintaining high extraction accuracy. This confirms its suitability for resource-constrained UAV platforms. For end-to-end transmission, the semantic index vector transmission results (Fig. 5) show that the Bit Error Rate (BER) decreases monotonically as the Signal-to-Noise Ratio (SNR) increases in all three channel environments. The rural environment achieves the best performance, followed by the suburban and urban environments. These differences are mainly caused by variations in scatterer density and link blockage across environments. The proposed transmission network maintains stable BER under different Doppler frequencies, demonstrating its robustness under dynamic channel conditions. For semantic image transmission, the proposed weighted hybrid loss function shows good training stability (Fig. 6), and LPUSC consistently outperforms the Deep Joint Source-Channel Coding (DeepJSCC) and JPEG + Low-Density Parity-Check (LDPC) baselines across the full SNR range (Fig. 7). Specifically, LPUSC achieves SSIM and Peak Signal-to-Noise Ratio (PSNR) gains of 1.3% and 4.8% over DeepJSCC, respectively, and gains of 43% and 79.5% over JPEG + LDPC, respectively. These results indicate that the proposed personalized semantic image transmission network achieves high-quality reconstruction and remains robust to channel variations.  Conclusions  To improve the efficiency and flexibility of UAV image communication, this paper proposes LPUSC, a lightweight personalized semantic communication system. The system uses a dual-branch transmission architecture that integrates lightweight, high-precision object detection and semantic segmentation models. It enables personalized content transmission without interest pre-training. This design satisfies personalized user requirements while maintaining low computational and communication overhead. Simulation results show that the LPUSC system achieves stable and reliable semantic index interaction and outperforms the DeepJSCC and JPEG + LDPC baselines in semantic region reconstruction. The proposed system provides a useful reference for efficient UAV image semantic communication in 6G low-altitude intelligent networks.
Design of Lightweight Gated Recurrent Unit Network Model Based on Memristor
HUA Honghu, XU Jia, ZHANG Bohao, WANG Wei, LI Zhiwei, LIU Haijun
Available online  , doi: 10.11999/JEIT260152
Abstract:
  Objective  With the slowdown of Complementary Metal-Oxide-Semiconductor (CMOS) technology scaling and the inherent memory-computation separation of von Neumann architectures, conventional computing systems face increasing challenges in processing large-scale sequential data. Memristors provide a promising solution because of their high integration density, fast switching speed, and synaptic plasticity. Memristor crossbar arrays naturally support Vector-Matrix Multiplication (VMM) in the analog domain, enabling energy-efficient in-memory computing. As a representative recurrent neural network, the Gated Recurrent Unit (GRU) has achieved excellent performance in sequential tasks such as trajectory prediction and urban sound classification. However, conventional hardware implementations of GRU networks require frequent data transfer between memory and processing units, resulting in high energy consumption and limited throughput. Although memristor-based GRU implementations improve computational efficiency, their large parameter size and high weight precision require substantial hardware resources and reduce deployment reliability on resource-constrained memristor arrays. In addition, device non-idealities, such as conductance fluctuations, further reduce inference accuracy. Existing memristor-based GRU methods generally treat weights and activations using the same quantization strategy without considering their different hardware implementation characteristics, and they provide limited robustness against device variations. This paper addresses these issues through a hardware-algorithm co-design strategy.  Methods  This paper proposes a lightweight memristor-based GRU network model. A 1T1R (one-transistor-one-resistor) memristor crossbar array is adopted for weight mapping and analog Multiply-Accumulate (MAC) operations. Signed weights are represented by differential pairs of positive and negative conductance matrices because memristor conductance values are inherently non-negative. A linear transformation is used to map trained network weights to memristor conductance values. To account for the different hardware implementation paths of weights and activations, a device-aware fusion quantization method based on performance analysis is proposed. Symmetric quantization is applied to weights stored in the memristor array because the zero-centered quantization range eliminates zero-point storage and simplifies write-driver circuit design. In contrast, asymmetric quantization is applied to activation values computed in peripheral circuits, thereby preserving the dynamic range and reducing quantization error. To improve robustness against memristor conductance fluctuations, weight noise training is incorporated into Quantization-Aware Training (QAT). Gaussian noise with an intensity determined by the device variation parameter is injected into quantized weights during each forward pass. This strategy acts as a regularizer that guides the model toward flatter loss minima and improves tolerance to weight perturbations. During backpropagation, the straight-through estimator updates the full-precision floating-point weights, whereas noise is dynamically resampled in every forward pass.  Results and Discussions  On the public UrbanSound8K dataset, the proposed full-precision lightweight memristor-based GRU network model achieves a classification accuracy of 93.94%. After applying the device-aware fusion quantization method, the 6-bit quantized model achieves 92.68% accuracy, corresponding to only a 1.26% decrease while reducing weight precision by 81.25% (Table 1). The proposed model outperforms Dilated Convolution (78.00%), LM-MFCC+GRU (92.00%), TFFS-DNN (88.74%), TFCNN (93.10%), and CL-Transformer (92.95%) under their full-precision settings (Table 2). Under noisy input conditions with Signal-to-Noise Ratios (SNRs) ranging from −10 dB to 10 dB, the 6-bit quantized model exhibits robustness comparable to or better than that of the full-precision model, demonstrating the effectiveness of the proposed device-aware fusion quantization strategy (Table 3). From the perspectives of storage, hardware resources, and device feasibility, 6-bit quantization reduces weight storage from 5.6 MB to 1.05 MB, corresponding to a compression ratio of 81.2%, while requiring only 2.8 million memristor cells under the 1T1R mapping scheme. Weight noise training also substantially improves robustness against device non-idealities. When the conductance variation reaches 14%, the classification accuracy increases from 82.97% to 91.14%. At the maximum simulated variation of 28%, the accuracy increases from 54.23% to 87.01% (Fig. 7), demonstrating improved tolerance to memristor device variations. On a self-constructed true-false trajectory dataset, the lightweight memristor-based GRU network model achieves 97.35% accuracy at full precision and 96.51% after 6-bit quantization, with only a 0.84% decrease, outperforming the Dilated Convolution baseline (Table 4). To further verify its applicability to different sequential tasks, the lightweight memristor-based GRU network model is evaluated on lithium-ion battery State-of-Charge (SOC) estimation using a public dataset. The 6-bit quantized model achieves Root Mean Square Errors (RMSEs) of 1.48%, 0.79%, and 0.74% at 0 °C, 25 °C, and 45 °C, respectively, outperforming the existing memristor-based GRU implementation. The proposed model also achieves lower RMSEs than the comparison method at all evaluated quantization precisions of 6 bits and above (Table 5).  Conclusions  This paper presents a lightweight memristor-based GRU network model for hardware deployment. By combining device-aware fusion quantization with weight noise training integrated into Quantization-Aware Training (QAT), the model achieves substantial memory compression while maintaining high classification accuracy and improving robustness to memristor device non-idealities. Experimental results on multiple datasets and sequential tasks demonstrate that the 6-bit quantized model preserves competitive accuracy and stable performance, providing an effective solution for deploying GRU networks on resource-constrained memristor-based edge computing platforms.
Heterogeneous Task Cooperative Scheduling Architecture for Networked Radar in Saturation Attack Air Defense Early Warning
YE Juhang, FANG Yuyuan, WEI Shaopeng, DUAN Jia, ZHANG Lei
Available online  , doi: 10.11999/JEIT260373
Abstract:
  Objective  To address the severe challenges posed by Unmanned Aerial Vehicle (UAV) swarms and intelligent loitering munitions, which generate massive, sudden, and heterogeneous early warning tasks during saturation attacks, a scalable networked-radar cooperative scheduling architecture is proposed. Existing architectures suffer from rigid dynamic coordination and insufficient capability for heterogeneous task scheduling. The non-convex cooperative scheduling problem is therefore decoupled into a multi-stage decision process consisting of multidimensional dynamic resource coordination and adaptive heterogeneous task scheduling. By incorporating a dispatch mechanism and Hierarchical Reinforcement Learning (HRL), a hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling is developed. The architecture is implemented through an execution-redispatch cognitive closed loop, together with a Target Dispatch Algorithm (TDA) and a hierarchical command-and-scheduling method, to address multi-radar, multi-target, and multi-task scheduling in saturation attack scenarios.  Methods  The cooperative scheduling problem is first decoupled into dispatch-oriented network-level multidimensional resource coordination and hierarchical-command-based node-level adaptive heterogeneous task scheduling. An environment perception layer establishes a dynamic uncertainty model based on the Bayesian Cramér-Rao Lower Bound (BCRLB) and radar detection probability to jointly characterize target threat levels and radar operating states. The network coordination layer then adopts an adaptive weighted dispatch model based on comprehensive combat effectiveness to achieve dynamic target allocation while constructing a scalable execution-redispatch cognitive closed loop. To solve the resulting large-scale, nonlinear, multi-constraint generalized bipartite graph matching problem, the proposed TDA, an MMAS-based algorithm incorporating feasibility-rule-based constraint handling, is developed as a constructive solution. At the node level, Hierarchical Q-Learning (HQL) is implemented through a serial dual-Q-table implementation for distributed heterogeneous task scheduling. Using task proportion control as the hierarchical subgoal, the proposed method transforms upper-level operational intent into lower-level beam dwell scheduling, enabling adaptive, long-term, and interpretable execution of heterogeneous tasks with complex dependency relationships.  Results and Discussions  A simulated point-defense scenario against UAV swarm and loitering munition saturation attacks is established using three networked homogeneous S-band medium-range phased-array radars to counter 200 high-speed maneuvering targets. For network-level coordination, the proposed TDA replaces conventional penalty functions with a hierarchical solution framework based on feasibility-rule constraint handling. By exploiting prior model information, TDA achieves higher solution quality and faster convergence than the Max-Min Ant System (MMAS), Artificial Bee Colony (ABC), and Genetic Algorithm (GA) (Fig. 3). Although computational complexity increases, the execution time remains well within the dispatch cycle, improving solution quality with only millisecond-level computational overhead while satisfying real-time operational requirements (Fig. 4). For node-level scheduling, HQL employs hierarchical macro- and micro-level decisions to ensure policy consistency. Supported by an internal dense transfer-reward mechanism, HQL achieves higher learning efficiency and better long-term policy quality than Q-Learning (QL) and the Priority-Based Method (PBM) (Fig. 5). The hierarchical serial dual-Q-table framework maintains balanced task proportions, maximizes comprehensive combat effectiveness, and improves resource utilization (Fig. 6). Furthermore, comparisons of the target track-loss rate and mean tracking error show that HQL achieves the lowest mean tracking error by prioritizing high-quality tracking tasks, despite a moderately higher target track-loss rate, demonstrating superior long-term scheduling capability (Fig. 7).  Conclusions  The proposed hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling effectively addresses the limitations of rigid dynamic coordination and insufficient heterogeneous task scheduling capability in existing networked radar systems. The proposed framework enables real-time cooperative scheduling of search, confirmation, and tracking tasks during large-scale, high-speed saturation attacks, thereby improving the operational capability of the air defense early warning system. Simulation results demonstrate improvements in applicable processing scale, environmental adaptability, long-term scheduling capability, scalability, and interpretability. Future work will explore online learning to reduce the discrepancy between offline training and online deployment and will further extend the architecture by integrating weapon-target assignment to support unified early warning and fire-control systems.
Analysis of Age upon Decisions and Distortion at Decisions in IoT Status Update Systems with Batch Arrivals
LIU Lei, JIN Wenkai, ZHANG Qingqing, LI Yuzhou, JIANG Fan
Available online  , doi: 10.11999/JEIT260359
Abstract:
  Objective  The rapid development of the Internet of Things (IoT) makes the timely transmission and processing of status updates essential for modern systems, where information freshness at decision epochs plays a critical role. In many IoT applications, such as smart grid fault detection and Industrial Internet of Things (IIoT) cluster monitoring, status updates typically arrive in batches rather than individually. However, most existing studies on Age of Information (AoI) assume single-update arrivals and therefore cannot accurately characterize the queueing dynamics caused by batch arrivals. Besides information freshness, distortion at decision epochs is another key factor because it directly affects decision quality. A fundamental tradeoff therefore exists between information freshness and distortion. Waiting for more complete status update information allows more completed status updates to be incorporated into joint estimation, but increases queueing and transmission delays, thereby reducing information freshness. In contrast, triggering decisions earlier reduces delay but increases distortion because fewer completed status updates are available for joint estimation. To address this problem, this paper investigates the tradeoff between information freshness and distortion in an IoT status update system with batch arrivals by adopting Age upon Decisions (AuD) and Distortion at Decisions (DaD) as performance metrics. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. Furthermore, for the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. The results provide theoretical insight and practical guidance for the design of IoT status update systems with batch arrivals.  Methods  Information freshness and distortion at decision epochs are analyzed for an IoT status update system with batch arrivals. AuD and DaD are adopted to quantify information freshness and distortion, respectively. Based on queueing theory, analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. A typical case with geometrically distributed batch sizes is then investigated. An alternating iterative optimization algorithm is further developed to jointly optimize the batch arrival rate, average batch size, and decision threshold to minimize the weighted sum of the average AuD and average DaD.  Results and Discussions  Simulation results validate the theoretical analysis. The average AuD exhibits a nonmonotonic trend as the batch arrival rate increases, first decreasing and then increasing. In addition, the Batch-size Coefficient Of Variation (BCOV) has a significant effect on the average AuD, with a smaller BCOV providing better information freshness performance. Under high-load conditions, queue backlogs become more severe, and stochastic fluctuations in batch arrivals have a greater effect on the queueing process. This increases service-time variability and amplifies the effect of BCOV on the average AuD (Fig. 2). As the average batch size increases, the system queue length and queueing delay increase, leading to a larger average AuD. At the same time, the decision control unit can utilize more completed status updates for joint estimation, thereby reducing the average DaD (Fig. 3). Moreover, the average DaD decreases as the decision threshold increases because more completed status updates are incorporated into the joint estimation process, improving estimation accuracy. A larger BCOV also increases the number of completed status updates available for joint estimation and therefore further reduces the average DaD (Fig. 4). The optimization results show that the solutions obtained by the proposed algorithm lie on the Pareto frontier, demonstrating its effectiveness. By comparison, fixed batch arrival rates and decision thresholds produce performance that is considerably farther from the Pareto frontier, demonstrating the advantage of jointly optimizing system parameters (Fig. 5).  Conclusions  This paper investigates an IoT status update system with batch arrivals by adopting AuD and DaD to quantify information freshness and distortion, respectively. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. For the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. Simulation results verify the theoretical analysis and reveal the effects of the batch arrival rate, average batch size, and decision threshold on the average AuD and average DaD. The results also demonstrate that the proposed low-complexity algorithm effectively identifies Pareto-optimal solutions for the AuD-DaD tradeoff. This study considers only the batch arrival characteristics of status updates. Future work will incorporate batch service mechanisms to further examine their effects on the tradeoff between AuD and DaD. Flexible decision mechanisms can also be developed to achieve adaptive AuD-DaD tradeoffs according to the heterogeneous requirements for information freshness and distortion across applications with different batch characteristics.
Kolmogorov-Arnold Nonlinear Enhancement Method for Aerial-Ground Person Re-IDentification
CHEN Yijun, ZENG Xianxian, LIU Shun, WANG Leijun
Available online  , doi: 10.11999/JEIT260430
Abstract:
  Objective  Aerial-Ground Person Re-IDentification (AG-PReID) aims to match the same person across Unmanned Aerial Vehicle (UAV) and ground-camera views. Compared with conventional same-platform person re-identification, this task faces larger cross-view appearance variation and more severe cross-domain distribution shifts. Under these conditions, identity-consistent cues are often weakened by strong viewpoint asymmetry and cross-domain appearance distortion. Existing methods mainly focus on feature extraction and cross-view representation alignment. However, the classification supervision branch still relies heavily on linear feature transformation, which limits its ability to model complex nonlinear discriminative relationships in high-dimensional feature spaces. A stronger nonlinear supervision mapping is therefore needed to better exploit high-order feature interactions and local discriminative variations. To address this issue, this paper proposes a Kolmogorov-Arnold Nonlinear Enhancement Module (KANEM). KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier. It uses learnable nonlinear mappings to adaptively enhance features for more discriminative cross-view representation learning.  Methods  The backbone follows the View-Decoupled Transformer (VDT), which introduces an additional view token and performs layer-wise view decoupling. This design separates view-related factors from identity features and reduces representation bias between aerial and ground domains. Based on this framework, KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier, thereby providing adaptive nonlinear mappings for feature enhancement. Specifically, KANEM consists of a base activation branch and a spline branch, which are stacked into cascaded function-mapping layers. This design enables more flexible nonlinear modeling than conventional linear or MultiLayer Perceptron (MLP)-based transformations. It allows the model to capture local nonlinear variations and complex correlations among feature dimensions. To improve discriminability and further separate identity and view information, the network is jointly optimized using identity classification loss, view classification loss, triplet loss, and orthogonality loss. KANEM is used only during training and is removed during inference, so no extra inference cost is introduced.  Results and Discussions  Comprehensive evaluations are conducted on the CARGO and AG-ReID datasets. The results show that the proposed method consistently performs better than the baseline model and existing state-of-the-art methods. On CARGO, the proposed method achieves 70.19%/63.16%/51.34% in Rank-1 accuracy, Mean Average Precision (mAP), and Mean Inverse Negative Penalty (mINP), respectively, under the overall ALL retrieval protocol. It also achieves 58.75%/53.27%/41.11% under the most challenging aerial-ground (A↔G) cross-view retrieval protocol (Table 1). On AG-ReID, the proposed method achieves the best performance under both retrieval protocols. It reaches 84.41%/76.21%/53.05% in Rank-1/mAP/mINP for aerial-to-ground (A→G) retrieval and 86.69%/77.99%/52.28% for ground-to-aerial (G→A) retrieval (Table 2). Ablation studies on CARGO further verify the effectiveness of KANEM. They show that KANEM achieves better overall performance than conventional linear transformation and MLP-based alternatives. This result indicates that the proposed nonlinear enhancement strategy is more suitable for supervision mapping in CARGO (Tables 3 and 4). In addition, integrating KANEM into other person re-identification tasks further demonstrates its potential generalization ability across different scenarios (Table 5). Parameter analysis shows that setting λ to 0.001 enables the model to better balance the complexity difference between view classification and identity classification (Fig. 2(a)). When G and P are set to 5 and 3, respectively, the model effectively fits nonlinear variations in the feature space while preserving the smoothness and continuity of spline functions. This setting achieves effective nonlinear feature enhancement (Fig. 2(b)(d)). The two-dimensional t-distributed Stochastic Neighbor Embedding (t-SNE) visualization shows that the enhanced features have higher intra-class compactness and better inter-class separability (Fig. 3). The top-5 retrieval comparisons further provide qualitative evidence that the proposed method improves ranking quality and retrieval robustness under all four retrieval protocols on CARGO. It promotes correct matches to higher positions and returns more relevant samples among the top-ranked results (Fig. 4).  Conclusions  This paper presents KANEM for AG-PReID. The proposed module is motivated by the large discrepancy between UAV and ground-camera views and by the limited capacity of linear feature transformation in the classification branch to capture complex nonlinear discriminative relationships. By replacing the conventional fully connected feature transformation between backbone output features and the linear classifier, KANEM provides a more flexible nonlinear supervision mechanism for cross-view representation learning. Through adaptive nonlinear enhancement, it better models complex feature interactions in high-dimensional spaces and strengthens the representation of cross-view consistency and fine-grained discriminative cues. Experimental results on CARGO and AG-ReID demonstrate the effectiveness of the proposed method, particularly in challenging scenarios with large view discrepancies. Future work will further refine the nonlinear mapping mechanism of KANEM and explore its use in more complex cross-view settings to improve model discriminability and generalization performance.
A Task Prediction-augmented Hierarchical Offloading Method for Space-Air-Ground Integrated Networks
ZHANG Linghao, XU Bo, SUN Jinlong, LAI Haiguang, ZHAO Haitao
Available online  , doi: 10.11999/JEIT260217
Abstract:
  Objective  Space-Air-Ground Integrated Networks (SAGIN) have become key infrastructure for future 6G communications. They support wide-area coverage and flexible deployment through the coordinated operation of Low Earth Orbit (LEO) satellites, Unmanned Aerial Vehicles (UAVs), and Ground Users (GUs). With the rapid growth of Internet of Things (IoT), Internet of Vehicles (IoV), and smart city applications, terminal devices generate increasingly diverse computation-intensive tasks. These tasks impose high requirements on real-time computing and resource scheduling. Mobile Edge Computing (MEC) has been integrated into SAGIN architectures to provide near-user computing services by using UAVs and satellites as edge nodes, thereby reducing task completion latency. However, efficient task offloading remains challenging when average task completion latency and UAV flight energy consumption must be jointly reduced. This difficulty is caused by the strong coupling among UAV trajectory planning, task offloading, and computational resource allocation. It is further intensified by the dynamic and partially observable nature of SAGIN environments. Existing Multi-Agent Reinforcement Learning (MARL) methods mainly rely on reactive decisions based on instantaneous observations. They lack awareness of future task workload changes, which leads to decision lag and limited adaptability under bursty traffic. To address these issues, a task prediction-augmented MARL method is proposed to support forward-looking decisions in dynamic SAGIN environments.  Methods  A three-layer SAGIN-MEC architecture is considered, including one LEO satellite, multiple UAVs, and GUs. Tasks can be processed locally, offloaded to UAVs through Ground-to-Air (G2A) links, or further relayed to the LEO satellite through Air-to-Satellite (A2S) links under a partial offloading mechanism. The joint optimization of UAV trajectory, user association, offloading ratios, and computational resource allocation is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem. The objective is to minimize the weighted sum of average task completion latency and UAV flight energy consumption. Owing to the nonconvexity and high dimensionality of this problem, it is reformulated as a DECentralized Partially Observable Markov Decision Process (DEC-POMDP). A Prediction-Augmented Multi-Agent Proximal Policy Optimization (PA-MAPPO) algorithm is then developed. A lightweight Exponential Smoothing-Autoregressive (ES-AR) prediction module is used to generate multi-step workload forecasts, which are incorporated into the state space of each agent. The algorithm adopts a bilevel structure. In the outer layer, Centralized Training and Decentralized Execution (CTDE)-based PA-MAPPO generates UAV trajectory actions. In the inner layer, Block Coordinate Descent (BCD)-based convex optimization solves the resource allocation and offloading subproblems, and closed-form resource allocation solutions are obtained through Lagrangian analysis. Generalized Advantage Estimation (GAE) and the PPO-Clip objective are used to improve training stability and convergence.  Results and Discussions  Simulations are conducted with one LEO satellite, five UAVs, and 50 GUs in a 1×1 km2 area. PA-MAPPO is compared with MAPPO without prediction and Prediction-Augmented Multi-Agent Deep Deterministic Policy Gradient (PA-MADDPG). The training curves show that PA-MAPPO converges within 500~700 episodes, with the highest average reward and the smallest variance, indicating better stability (Fig. 3). As the number of GUs increases from 20 to 80, PA-MAPPO consistently achieves the lowest system cost. Compared with MAPPO and PA-MADDPG, it reduces the average cost by approximately 12.4% and 18.7%, respectively (Fig. 4). Experiments with different UAV numbers show a U-shaped cost curve for all algorithms. The best configuration is obtained when U=5, where PA-MAPPO achieves the minimum cost (Fig. 5). Sensitivity analysis of the latency-energy tradeoff weight ω confirms that PA-MAPPO remains robust under different optimization preferences (Fig. 6). The prediction horizon H has a nonmonotonic effect on performance. When H=5, PA-MAPPO obtains the best result and reduces the cost by approximately 14.9% compared with the no-prediction case. Longer horizons degrade performance because prediction errors accumulate (Fig. 7).  Conclusions  The PA-MAPPO algorithm is proposed to solve the joint optimization of UAV trajectory planning, user association, task offloading, and computational resource allocation in dynamic SAGIN environments. By integrating a lightweight ES-AR task workload prediction module into the MARL process, PA-MAPPO enables UAV agents to account for future task dynamics. This design reduces the decision lag caused by purely reactive methods. The inner BCD-based convex optimization converges to a Karush-Kuhn-Tucker (KKT)-stationary point, while the outer CTDE-based PPO mechanism improves training stability and scalability. Simulation results show that PA-MAPPO outperforms baseline methods in average task completion latency, UAV flight energy consumption, and overall system cost. It also maintains strong scalability and robustness under different system configurations. Future work will study online prediction and decision co-optimization in multi-satellite cooperative scenarios and examine the effect of dynamic network topology changes on algorithm performance.
A Dual-polarized Magnetoelectric Dipole Antenna Array with Differential Feeding
TANG Li, WANG Zhihui, ZHAO Luyu
Available online  , doi: 10.11999/JEIT260505
Abstract:
  Objective  This study addresses key challenges in Fifth-Generation (5G) millimeter-wave terminal antennas by designing a compact, high-performance dual-polarized array. Existing designs often face trade-offs among bandwidth, beam-scanning range, and integration complexity. To address these limitations, this paper proposes a differentially fed magnetoelectric dipole array. A stacked stripline-slot-stripline balun is used to enable efficient single-ended-to-differential conversion, and the array design is optimized. The objective is to realize an integrated solution with wideband operation, low cross-polarization, wide-angle beam scanning, and high integration density for practical 5G millimeter-wave applications.  Methods  A structured design method is adopted. First, a stacked differential balun based on a stripline-slot-stripline configuration is developed to achieve efficient single-ended-to-differential conversion. A single-polarized magnetoelectric dipole antenna element is then designed and integrated with the balun, and its performance is characterized. The design is further extended by orthogonally integrating two elements to form a dual-polarized unit, which is used to construct a 1×4 linear array. Iterative full-wave electromagnetic simulation and optimization are conducted to balance wideband impedance matching, high port isolation, stable wide-angle beam scanning, grating-lobe suppression, and mutual-coupling reduction.  Results and Discussions  The optimized 1×4 dual-polarized differentially fed magnetoelectric dipole antenna array uses an element spacing of 4.6 mm, corresponding to 0.4 free-space wavelength at 26 GHz. This spacing achieves a favorable balance between grating-lobe suppression and inter-element mutual-coupling reduction. The measured –10 dB reflection coefficient bandwidths are 25~29.4 GHz for the +45° polarization port and 25~27.7 GHz for the –45° polarization port (Fig. 20). The slight matching difference is attributed to the incomplete structural symmetry of the baluns under the two polarization modes (Fig. 13). At 26 GHz, both polarization modes provide a peak gain of 10.7~11 dBi and support ±60° wide-angle beam scanning, with main-lobe gain attenuation no greater than 3 dB (Fig. 21). The measured radiation performance agrees well with the simulated results. Minor deviations are mainly caused by the high dimensional sensitivity of millimeter-wave structures and small errors in fabrication and test assembly. The array also maintains stable low cross-polarization and high port isolation across the operating band. These results are achieved through equal-length feed lines, symmetric layout, ground-pad shielding, and metallized-via electromagnetic isolation (Fig. 16), which suppress mutual coupling and parasitic radiation and ensure consistent dual-polarized radiation performance.  Conclusions  This paper presents a dual-polarized magnetoelectric dipole antenna array with differential feeding for 5G millimeter-wave applications. By using a stacked stripline-slot-stripline balun and optimizing the radiating structure and array layout, the design achieves wide bandwidth, high gain, low cross-polarization, and wide-angle beam scanning. The differential balun enables efficient single-ended-to-differential conversion with good amplitude and phase balance across the target band. The implemented 1×4 array, with an optimized element spacing of 4.6 mm, achieves a simulated peak gain of 11 dBi at 26 GHz and supports ±60° beam scanning, with gain variation below 3 dB. The overall design verifies the feasibility of a differentially fed magnetoelectric dipole architecture for compact, high-performance 5G millimeter-wave terminal antenna modules. Future work may focus on larger array configurations and further integration with BeamForming Integrated Circuits (BFICs).
A Behavioral Economics-Based Game Model for Side-Channel Security Attack and Defense Strategies
CAI Juesong, YAN Yingjian, WANG Jindong
Available online  , doi: 10.11999/JEIT260121
Abstract:
  Objective  The field of side-channel security currently lacks a systematic, quantifiable, and reproducible model for guiding the selection of attack and defense strategies, particularly in real-world engineering contexts where resource constraints necessitate informed cost-benefit trade-offs. The absence of such a framework impedes the practical realization of the “appropriate security” principle, often leading to either over-protection or under-protection of cryptographic modules. Traditional approaches to evaluating attack and defense costs rely heavily on subjective expert judgments, which are inherently arbitrary, difficult to replicate, and lack a structured multi-dimensional assessment. To bridge this critical gap, this research proposes an interdisciplinary model that integrates game theory, behavioral economics, and the Analytic Hierarchy Process (AHP). The primary objective is to establish a holistic decision-support system that not only quantifies the multi-faceted costs of various side-channel strategies but also incorporates the psychological dimensions of decision-making under risk, thereby enabling dynamic and economically rational security strategy selection tailored to specific asset values and security levels.  Methods  This study constructs a multi-layered modeling framework based on a static non-cooperative game with incomplete information. First, the side-channel analyst and the defense designer are formally defined as rational players, each possessing a finite set of strategies: the attacker may choose from non-modeling attacks such as DPA/CPA, modeling-based attacks like template attacks, or emerging deep learning-based side-channel analysis; the defender may adopt countermeasures including time hiding, amplitude hiding, or masking/blinding techniques. To systematically quantify the often-overlooked cost dimension, an AHP-based structured cost model is introduced. Through pairwise comparison matrices, the model decomposes costs into multiple criteria—such as time, data storage, computational resources, expertise, and hardware overhead—and assigns objective weights to each criterion, thereby replacing subjective cost estimates with a reproducible, hierarchical evaluation system. Furthermore, to reflect real-world decision-making behavior, key concepts from behavioral economics are integrated: Prospect Theory models how gains and losses are perceived relative to a reference point, while risk aversion coefficients capture players’ tolerance for uncertainty. These behavioral parameters are explicitly linked to the security level of the cryptographic module, allowing the model to adapt to different operational contexts. The resulting behavioral-augmented Bayesian game is then solved using the concept of Bayes-Nash Equilibrium, wherein each player’s optimal mixed strategy is derived based on their private type (behavioral profile) and beliefs about the opponent. To ensure engineering relevance, the As Low As Reasonably Practicable principle is incorporated as a constraint, enforcing that any selected defense strategy must be justifiable in terms of risk reduction versus cost incurred. Numerical solutions are obtained via a customized sequential quadratic programming algorithm implemented in Python.  Results and Discussions  A comprehensive experimental evaluation was conducted to validate the proposed model’s consistency, sensitivity, and practical utility. The AHP-based cost quantification demonstrated strong internal consistency, with all consistency ratios below the 0.1 threshold, confirming the reliability of the judgment matrices. The derived weight distributions revealed intuitive priorities: attackers placed greater emphasis on technical barriers and computational cost, whereas defenders prioritized design complexity and performance overhead. The behavioral adjustment layer successfully modulated perceived costs according to security levels: under low-security conditions (high risk aversion), costs were perceptually inflated, leading to conservative strategy choices; under high-security conditions, decision-making aligned more closely with objectively quantified costs. Equilibrium analysis across varying asset values and security levels yielded interpretable and rational strategy profiles. For low-value assets, both players exhibited a strong tendency toward low-cost or “no action” strategies, adhering to the lower bound of the ALARP region. As asset value increased, a clear threshold effect was observed, triggering a shift toward high-cost, high-efficacy strategies such as deep learning-based attacks and masking-based defenses. Sensitivity analysis further confirmed that defense strategy probabilities increased monotonically with asset value, validating the model’s ability to capture the non-linear relationship between protection intensity and asset criticality. These findings underscore the model’s capacity to support context-aware, adaptive security decision-making that balances risk, cost, and psychological factors.  Conclusions  This research presents a novel, behaviorally informed game-theoretic model for side-channel security strategy selection, addressing a significant void in existing literature regarding structured cost-benefit assessment. By integrating AHP-based objective cost quantification, behaviorally adjusted subjective valuations, and ALARP-driven engineering constraints, the proposed framework offers a multi-dimensional, reproducible, and context-sensitive tool for analyzing attack-defense interactions. The model advances the field by explicitly linking security levels to behavioral parameters, enabling dynamic strategy adaptation in response to both asset value and decision-makers’ risk perceptions. Although the current implementation relies partially on expert-defined parameters and operates within a static game setting, it establishes a critical foundation for transitioning from heuristic-based security decisions to quantitatively grounded, interdisciplinary analysis. Future work will focus on parameter calibration using real-world attack/defense datasets, extension to multi-stage dynamic games to capture strategic evolution over time, and empirical validation in industrial cryptographic evaluation scenarios. This study contributes the first systematic methodology for cost-aware, behaviorally realistic strategy optimization in side-channel security, offering both theoretical insights and practical guidance toward achieving “appropriate security” in cryptographic engineering.
An Overview of Key Technologies for 6G-Enabled Communication-Computing Integration and Energy-Efficiency Optimization
LIU Guangyi, CAI Qing, WANG Xinyao, CHEN Tianjiao, JIN Jing, XUE Yahui, WANG Ailing, WANG Hanning
Available online  , doi: 10.11999/JEIT260399
Abstract:
  Significance   Constrained by size, power consumption, and cost, emerging intelligent terminals often face excessive energy consumption and limited battery life. These limitations have become major bottlenecks to large-scale deployment. Compared with Fifth-Generation (5G) wireless networks, Sixth-Generation (6G) wireless networks are expected to enhance the Radio Access Network (RAN) architecture and move computing capability toward the RAN side. High-energy and compute-intensive Artificial Intelligence (AI) tasks that are originally executed by end devices can therefore be processed by the network. Through End-Edge Collaboration, emerging intelligent terminals can be upgraded toward lightweight design, low cost, and long battery life, thereby supporting the large-scale deployment of ubiquitous intelligence in 6G networks.  Progress   Current progress in Terminal Energy Consumption Optimization through 6G End-Edge Collaboration is reviewed, with emphasis on local execution, full offloading, and partial offloading. In local execution, User Equipment (UE) processes all tasks locally, which leads to high computing energy consumption. In full offloading, all tasks are transferred to the RAN. This reduces terminal-side computing energy consumption but can increase transmission energy consumption, especially under poor channel conditions. Partial offloading combines the benefits of both modes and optimizes energy consumption according to real-time network conditions. For partial offloading, four representative optimization techniques are summarized. (1) Feature Extraction and Filtering. Semantic encoding and information extraction are performed at the UE, and only task-relevant data are transmitted to the RAN. This reduces redundant data transmission and lowers transmission energy consumption. (2) Split Offloading. A large Deep Neural Network (DNN) is divided into layers according to its structure. Simpler shallow layers are processed at the UE, whereas more complex deep layers are offloaded to the RAN. This method balances terminal-side and RAN-side computational loads through End-Edge Collaborative Inference. (3) Model Lightweighting. Model complexity is reduced through pruning, quantization, and knowledge distillation, which lowers computational overhead while maintaining task performance. (4) Incremental Inference. Only changed data or updated features are processed, while historical computations are reused. This reduces redundant computation. Together, these techniques improve terminal performance and energy efficiency within the 6G End-Edge Collaboration framework.  Conclusions  This paper systematically reviews Terminal Energy Consumption Optimization for 6G End-Edge Collaboration. It summarizes the functional evolution of enhanced RAN, constructs an End-Edge Collaborative service framework for Communication-Computing Integration, and establishes a theoretical model of terminal computing energy consumption and transmission energy consumption. The composition and influencing factors of energy consumption under different offloading modes are clarified. Key energy optimization technologies, including Feature Extraction and Filtering, Split Offloading, Model Lightweighting, and Incremental Inference, are then discussed. To address energy consumption fluctuations caused by dynamic wireless channels, the paper proposes energy optimization mechanisms based on Adaptive Semantic Compression, Dynamic Split Offloading, Adaptive Model Pruning, and Incremental Inference. These mechanisms maintain a dynamic balance between energy optimization and task performance. Using embodied intelligent robot video understanding as a typical application scenario, a test platform is developed to verify the effectiveness of the proposed mechanisms. Current challenges and future research directions are also analyzed.  Prospects   Although End-Edge Collaborative energy-saving technologies have achieved initial progress, practical deployment still faces challenges in real network environments, dynamic wireless channels, and large-scale user access. Future research should examine the trade-off between optimization overhead and system robustness. It should also study dynamic communication-computing resource substitution modeling in stochastic resource environments, multi-user collaboration strategies, and global energy-efficiency optimization. As the technology matures, standardization and engineering implementation of End-Edge Collaborative energy-saving frameworks will become critical to the large-scale adoption of 6G applications. Future studies should further integrate algorithm design with network architecture, enabling practical deployment of low-power and high-efficiency intelligent communication systems.
Gating Adaptive Repeat Query Framework for Reliable Collaborative Inference with Edge Heterogeneous LLMs
WANG Tengsheng, YU Tao, LI Jihong, ZHENG Guhan, ZHANG Shunqing
Available online  , doi: 10.11999/JEIT260218
Abstract:
  Objective  Reliable inference at the network edge is indispensable for 6G-enabled ubiquitous AI, yet the deployment of large language models (LLMs) in such environments remains a cornerstone challenge. In resource-constrained edge settings, single-LLM inference often proves unreliable due to knowledge limitations and inherent biases, severely hampering real-world deployment. Collaborative inference leveraging multiple heterogeneous LLMs emerges as a promising remedy to boost robustness, but it introduces nontrivial hurdles under stringent latency and energy budgets, especially when wireless channel conditions and query content vary unpredictably. These challenges include the need for dynamic sequential decision-making for LLM selection and resource allocation, the fundamental paradigm mismatch between bit-level reliability protocols and semantic-level error correction, and the lack of adaptive mechanisms to align and fuse disparate LLM outputs effectively. To fill these critical gaps, this paper presents a novel framework that fundamentally reinterprets collaborative inference as a semantic-driven, closed-loop process, thereby transitioning from conventional bit-retransmission to semantic-retransmission and offering a practical path toward reliable 6G edge intelligence.  Methods  In response to these critical challenges, we propose the Gating Adaptive Repeat Query (G-ARQ) framework. Its core innovation is the Semantic-Space Alignment and Error-Guided Retransmission (SEMAR) mechanism. SEMAR first aligns the token-level probability distributions from heterogeneous LLMs into a unified semantic space using relative representation, enabling comparable outputs. It then models the collaborative process probabilistically, explicitly capturing error dependencies among models, and uses an Expectation-Maximization (EM) algorithm to infer a latent error direction, which guides the selection of the next LLM for query retransmission, steering it towards outputs orthogonal to previous errors. To jointly optimize the LLM gating and uplink power allocation under communication constraints without requiring explicit system dynamics—often unavailable in practice—we design a black-box trajectory optimizer. This optimizer formulates the sequential decision problem as sampling from a target distribution that encodes dynamic feasibility, optimality, and constraints. It employs a diffusion-based sampling process with a model-guided prior and Monte Carlo estimation to generate near-optimal policy trajectories that satisfy hard latency and energy limits.  Results and Discussions  To evaluate the practical viability of G-ARQ under realistic edge conditions, simulations are conducted in a scenario with five base stations hosting five heterogeneous 7B-parameter LLMs (Mistral-7B, Vicuna-7B, Nous-Capybala-7B, Gemma-7B, and Llama-2-7B). The user equipment (UE) performs a question-answering task evaluated on a mixed SQuAD and TriviaQA dataset. Component-level evaluations, each designed to isolate the contribution of a single innovation, validate the effectiveness of every key element. The error-guided gating of SEMAR, compared to a Top-k gating baseline, improves accuracy by 0.23 % on average, and its dynamic weight ensemble contributes an additional 0.7 % gain (Fig. 3). The black-box trajectory optimizer, which operates without any explicit channel model, achieves accuracy close to that of the unconstrained model-greedy strategy while ensuring strict latency constraints (Fig. 4). The convergence of the optimizer is verified by tracking the evolution of \begin{document}$ J(\boldsymbol{S}) $\end{document} and selection probability over diffusion steps (Fig. 5). System-level performance under varying latency and energy constraints demonstrates that G-ARQ consistently surpasses two baselines: one combining model-greedy selection with Proximal Policy Optimization (PPO) for power optimization, and another combining model-greedy selection with Simulated Annealing for power optimization, both using average output weights. The accuracy improvement is most significant under the most stringent resource limits, reaching up to 2.2 % for \begin{document}$ {K}_{\max }=1 $\end{document} and 1.9 % for \begin{document}$ {K}_{\max }=2 $\end{document} (Fig. 6, Fig. 7). The framework successfully establishes a Pareto boundary that characterizes the inherent trade-off between inference accuracy and communication latency, providing a valuable design guideline for resource-constrained edge systems and offering actionable insights for real-world deployment. The GARQ-S variant is noted to outperform GARQ-E by avoiding the integration of outputs from previously erroneous models.  Conclusions  This paper proposed G-ARQ framework, an innovative closed-loop framework that transforms collaborative edge inference into a semantics-guided retransmission process. By introducing SEMAR for error-based alignment and selection of heterogeneous LLMs, and employing a black-box trajectory optimizer for joint model selection and power allocation, the framework achieves up to a 2.2% accuracy improvement under strict resource constraints. The results validate G-ARQ as an effective and practical approach toward reliable and efficient 6G edge intelligence.
Queue Stability Constrained Robust Secure Beamforming for Low-Altitude UAV-ISAC Systems
OUYANG Jian, REN Wei, XU Ba, LIU Xiaoyu, JIANG Wanmu
Available online  , doi: 10.11999/JEIT260275
Abstract:
  Objective  To address the challenges of antenna array angle errors caused by UAV jitter, transmission instability induced by random data arrivals, and secure transmission guarantee in multi-eavesdropper scenarios for low-altitude UAV-ISAC systems, this paper proposes a robust secure beamforming algorithm based on queue stability constraints. The proposed algorithms aims to minimize long-term transmit power consumption while maintaining data queue stability while enhancing beamforming robustness against UAV jitter.  Methods  To guarantee the stability of the wireless transmission in UAV-ISAC systems, this paper formulates an optimization problem aimed at minimizing the long-term average transmit power, subject to constraints on data queue stability, secure communication rate, sensing performance, and maximum transmit power. Since this long-term optimization problem is intractable, the Lyapunov optimization framework is employed to transform it into a sequence of short-term subproblems. To handle the antenna array angle errors caused by UAV jitter within each short slot, we jointly adopt the second-order Taylor series expansion and the S-Procedure method to approximate the short-term subproblem into a convex form. Consequently, a robust secure beamforming algorithm based on penalty successive convex approximation optimization is proposed.  Results and Discussions  Simulation results demonstrate the impact of the number of antennas, secrecy rate threshold, beam gain threshold, UAV jitter error, and Lyapunov weight factor on the system transmit power. As illustrated by the beam gain pattern in Fig. 3, the communication beamformer facilitates cooperative sensing toward the sensing area, while the sensing beamformer enhances system security by directing interference toward potential eavesdroppers. This validates the effectiveness of the proposed algorithm in simultaneously improving sensing and secure communication performance. Furthermore, leveraging the dual-function characteristics of the communication and sensing beamformers, the proposed integrated sensing and communication scheme achieves significantly higher resource utilization efficiency than the communication-only and sensing-only schemes, as shown in Fig. 5. Additionally, Fig. 6 indicates that, compared with the non-robust scheme, the proposed robust scheme strictly satisfies security requirements under various angle errors. Finally, Fig. 8 shows that the proposed queue-aware scheme can effectively suppress the transmit power fluctuations caused by random data arrivals, exhibiting superior stability compared to the queue-free baseline.  Conclusions  This paper investigates a robust secure beamforming method for UAV-ISAC systems subject to queue stability constraints. First, based on the Lyapunov optimization framework, the long-term stochastic optimization problem is transformed into a sequence of short-term subproblems. Second, to address the issue of UAV jitter, the second-order Taylor series expansion and the S-Procedure method are jointly employed to approximate the non-convex constraints into tractable convex forms. Finally, a robust secure BF optimization algorithm based on penalty successive convex approximation is proposed to efficiently solve the deterministic short-term subproblems. Simulation results demonstrate that the proposed scheme can effectively tackle the challenges posed by random data arrivals, UAV jitter, and eavesdropping threats, thereby ensuring the stability and security of downlink data transmission in low-altitude UAV-ISAC systems.
VT2R: Video and Text-driven Method for Generating Large-scale Millimeter-wave Radar Data
DENG Kaikai, LING Yue, XING Ling, WU Honghai, ZHAO Dong, MA Huahong
Available online  , doi: 10.11999/JEIT260240
Abstract:
  Objective  The lack of large-scale training data impedes progress in developing robust and generalized deep learning models. However, existing millimeter-wave radar data generation methods are ineffective due to a lack of sufficient data sources. To address this gap, this paper proposes a video and text-driven radar data generation method, VT2R, which utilizes video or text data to generate large-scale, realistic radar data, solving the key problem of constructing the mapping relationship between video and text and radar data.  Methods  The proposed method consists of three main components: video feature encoding network, text feature encoding network, radar feature encoding network and data fitting and decoding network. Video feature encoding networks and text feature encoding networks extract temporally consistent visual representations and alignable semantic features, respectively, while the radar encoding network learns the structure and dynamic information of point clouds through hierarchical spatiotemporal modeling. In the data fitting and decoding network based on Variational AutoEncoder (VAE), multi-modal features are mapped to a unified latent distribution space and decoded into radar data through reparameterized sampling. During training, reconstruction loss, Kullback-Leibler (KL) divergence loss, and cross-modal similarity loss are jointly optimized.  Results and Discussions  This paper constructs the first radar point cloud dataset for reclining gesture recognition (Figs. 6 and 7), covering 5 gesture categories, 32 participants, and a total of 14,400 samples. Experimental results based on this dataset show that VT2R achieves a recognition accuracy of 89.2% when trained using only generated radar data, a 33.88% improvement over the representative RFGen (Figs. 9 and 10). When combined with a small amount of real radar data for joint training, the accuracy further improves to 97.62%, a 21.48% improvement over RFGen (Figs. 9 and 11). Furthermore, VT2R still achieves average recognition accuracies of 89.35% and 97.21% under different scenarios and factors (Figs. 16-18). In addition, this paper also verifies the accuracy of VT2R under different postures, achieving average accuracies of 89.98% and 97.55% in the first and third settings, respectively (Fig. 19), which is basically consistent with the result obtained when lying down, demonstrating its robustness under cross-posture conditions.  Conclusions  This paper proposes a radar data generation system, VT2R, which addresses the severe lack of realistic radar training data when users are performing gestures in a lying position. Through a video feature encoding network built on a vision-language pre-trained model, a text encoding network incorporating cue templates, a hierarchical radar encoding network for sparse point clouds, and a VAE-based data fitting and decoding network, these components collaboratively generate large-scale, realistic radar data. It also supports augmented reconstruction based on limited real radar data, providing rich data support for radar perception tasks. Future work will focus on solving multi-modal data generation for more complex gesture scenes, providing better data support for emerging large-scale models.
Multi-task Lightning Nowcasting with Spatio-temporal Focal Perception and Synergistic Weighted Loss
TANG Zhihao, HAN Yuanpeng, ZHANG Hui, SONG Lin, ZHANG Qilin, LIU Yi
Available online  , doi: 10.11999/JEIT260234
Abstract:
  Objective  Lightning nowcasting is essential for early warning systems and for protecting critical infrastructure, including aviation, power grids, and transportation systems. Traditional numerical weather prediction models depend strongly on parameterization schemes and require high computational costs, which limits their use in rapid-update nowcasting. Although deep learning methods have advanced, they still have difficulty handling extreme data sparsity, suffer from serial-computation bottlenecks in recurrent architectures, and mainly focus on binary occurrence prediction rather than the joint optimization of lightning-frequency prediction and regional localization. Moreover, conventional loss functions are easily dominated by extensive non-lightning areas, which biases predictions toward zero or causes excessive false alarms. To address these limitations, a Spatio-Temporal Focal perception and synergistic weighted loss Network (STF-Net) is proposed as a multi-task lightning nowcasting model that jointly predicts lightning frequency and occurrence regions. It integrates three key components: a Lightning Adaptive Attention Module (LAAM) for explicit spatio-temporal dependency modeling, a Spatio-Temporally Weighted Hybrid Loss for data sparsity and imbalance, and a spatio-temporal dual-branch Generative Adversarial Network (GAN) to improve prediction fidelity and temporal coherence.  Methods  STF-Net is built on the SimVP video prediction architecture and adopts an encoder-translator-decoder paradigm. LAAM uses a three-dimensional decoupled attention mechanism along the height, width, and channel dimensions, enabling adaptive focus on convectively sensitive regions while maintaining computational efficiency. The Spatio-Temporally Weighted Hybrid Loss combines Temporally Weighted Mean Squared Error (TW-MSE) for frequency regression and Dual-Weighted Cross-Entropy loss (DWCE) for regional localization. Time-increasing weights are incorporated to improve medium- to long-term forecast robustness (Fig. 5). DWCE integrates static class weights with dynamic grid weights, thereby balancing global class proportions and local lightning-frequency heterogeneity. A spatio-temporal dual-branch GAN, consisting of a spatial PatchGAN discriminator and a temporal three-dimensional convolutional discriminator, is used to improve the textural fidelity and temporal coherence of predicted lightning-frequency fields. The model uses six consecutive historical lightning-frequency frames at 256×256 resolution and 10-min intervals to predict the next six frames, corresponding to a 1-h forecast window. Experiments are conducted on a high-resolution Very Low Frequency Lightning Location Network (VLF-LLN) dataset containing 11,748 images that cover different seasonal and weather conditions. The dataset is split at a ratio of 7:3 for training and testing.  Results and Discussions  Comprehensive evaluation metrics are used, including Mean Squared Error (MSE) and Mean Absolute Error (MAE), which are computed only on lightning pixels; Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) for image fidelity; and Probability Of Detection (POD), False Alarm Rate (FAR), and Critical Success Index (CSI) for regional detection. STF-Net achieves a CSI of 0.663 within the 1-h forecast window, representing a 14.5% improvement over the SimVP baseline (0.579). It also reduces FAR from 0.351 to 0.216, corresponding to a relative reduction of 38.5% (Table 1). Ablation studies validate each component. Adding GAN improves CSI to 0.624 and reduces MSE to 0.109. Further incorporation of LAAM increases CSI to 0.629 and yields the highest POD of 0.894. The complete STF-Net with the hybrid loss achieves the best performance, with a CSI of 0.663 and an MSE of 0.105 (Table 1, Fig. 5). The joint prediction of frequency and region is supported by simultaneous improvements in regression metrics (MSE and MAE) and detection metrics (CSI and FAR). Time-step analysis shows that LAAM reduces long-term performance degradation, with STF-Net maintaining the highest CSI compared with SimVP+GAN and SimVP (Fig. 6). Comparative experiments with ConvLSTM and PredRNN further demonstrate the superiority of STF-Net across all lead times. STF-Net consistently achieves higher CSI and lower FAR, and its advantage becomes more evident as the forecast horizon increases (Fig. 7). Its consistent gains in PSNR and SSIM further indicate that the spatio-temporal GAN helps generate coherent, detail-rich predictions. Visualization results show that STF-Net produces structurally clear and continuous lightning-activity regions centered on high-frequency areas. It accurately tracks dynamic evolution patterns, including movement, merging, and splitting, while generating minimal noise in non-lightning regions. These results demonstrate effective collaborative prediction of both frequency magnitude and spatial distribution.  Conclusions  STF-Net is presented as a deep learning model for joint lightning-frequency prediction and regional localization. It explicitly models long-range spatio-temporal dependencies, focuses on critical convective zones, addresses extreme data sparsity and class imbalance, and jointly optimizes frequency regression and regional localization while suppressing false alarms. Its spatio-temporal dual-branch GAN further improves the spatial structural consistency and temporal coherence of predictions. Experimental results show that STF-Net outperforms baseline and state-of-the-art models, achieving a CSI of 0.663, an FAR of 0.216, and the best MSE and MAE values within a 1-h forecast window. The model effectively reduces long-term performance degradation, captures the evolution trends of lightning regions, and generates physically plausible predictions with minimal background noise. This study provides an efficient end-to-end solution for operational lightning nowcasting systems and offers guidance for model design in sparse meteorological spatio-temporal sequence prediction.
Intelligent Privacy-Aware Computation Offloading Method against Multi-server Joint Inference Attacks
MIN Minghui, LIU Mingcheng, ZHANG Peng, DUAN Jincheng, LI Shiyin, ZHANG Hongliang
Available online  , doi: 10.11999/JEIT260249
Abstract:
  Objective  With the rapid development of the low-altitude economy, services such as intelligent transportation, smart healthcare, and low-altitude logistics have become increasingly common. Their efficient operation depends on the real-time processing of massive sensing data. Mobile Edge Computing (MEC) improves task execution efficiency and reduces device computational burdens by offloading tasks to nearby servers. However, user privacy and security risks have become increasingly severe. In dynamic scenarios where multiple MEC servers jointly process tasks, information sharing can enable multi-server joint inference attacks and greatly increase the risk of user location privacy leakage. Although existing studies have used Differential Privacy (DP) to protect user location privacy, current DP-based solutions remain limited. These methods inject noise into offloading decisions, but unconstrained noise may reduce task allocation accuracy. In addition, user mobility causes continuous changes in channel states during dynamic computation offloading. Privacy leakage risks and attacker behaviors are also uncertain. Traditional optimization methods based on static system models are therefore unsuitable for such dynamic environments. To address these challenges, this paper proposes an Asynchronous Advantage Actor-Critic (A3C)-based Intelligent Privacy-Aware Computation Offloading (AIPCO) scheme against multi-server joint inference attacks. The proposed scheme protects user location privacy while maximizing the overall utility of the MEC system.  Methods  This paper proposes a DP-based task offloading rate perturbation mechanism. By adding controlled noise, the mechanism increases the randomness of user task offloading toward multiple MEC servers. A truncated Laplace mechanism is used to constrain the perturbed offloading rates within valid boundaries. This design satisfies the mathematical guarantees of DP and reduces the accuracy of multi-server joint inference attacks on sensitive user locations. Privacy entropy is then introduced to dynamically evaluate the real-time privacy protection level. Finally, the AIPCO scheme is constructed. Through a multi-threaded asynchronous training mechanism, the scheme interacts with the environment through iterative trial and error and efficiently learns the optimal real-time offloading policy online. The proposed scheme dynamically protects user privacy, reduces computational cost, and maximizes comprehensive system utility.  Results and Discussions  The AIPCO scheme jointly optimizes user privacy and task offloading cost by incorporating multidimensional performance variables into the reinforcement learning reward function. A comprehensive performance analysis (Fig. 4) shows that, when the number of continuous learning iterations reaches 200, the privacy protection level of AIPCO increases by 2.52%, 3.56%, and 22.90% compared with RCLM, JODRL, and DODA-DT, respectively. This advantage is mainly attributed to the DP-based task offloading rate perturbation method, which uses the truncated Laplace mechanism to increase data randomness while strictly constraining the perturbation range. By contrast, RCLM perturbs the task offloading rate through range-limited DP without using the truncated Laplace mechanism. JODRL increases randomness only through network policy optimization, resulting in a lower privacy protection level. DODA-DT focuses on balancing energy consumption and system latency without optimizing user privacy. For the privacy weight parameter (Fig. 5), increasing $\omega$ improves privacy protection. For example, the privacy protection level increases by 5.64% when $\omega$ rises from 0.2 to 0.7, with a clear performance gain at 0.7. As the system agent reduces its focus on computational cost, user utility remains optimal despite increased cost. When the physical distance between users and the MEC server is adjusted (Fig. 6), AIPCO shows stronger privacy protection in long-distance scenarios. A greater distance reduces the number of tasks offloaded to the server. Therefore, attackers obtain less information, and privacy protection improves. Although computational cost increases with distance, AIPCO consistently outperforms competing schemes. These results confirm that AIPCO achieves optimal MEC system utility while protecting user privacy.  Conclusions  To mitigate multi-server joint inference attacks caused by information sharing among collaborative MEC servers, this paper proposes an AIPCO method. A DP-based task offloading rate perturbation scheme is designed to increase randomness, and a truncated Laplace mechanism is used to constrain perturbed rates within reasonable boundaries. The scheme is proven to satisfy strict DP mathematical guarantees, and privacy entropy is introduced to quantitatively evaluate the privacy protection level. In addition, the AIPCO scheme uses a multi-threaded asynchronous training mode, enabling the agent to efficiently learn the optimal perturbed offloading policy in a continuous space and maximize overall system utility. Simulation results show that the proposed scheme outperforms the baselines in both dynamic and average performance. It achieves optimal system utility while protecting user privacy.
Bearing Fault Diagnosis of Roadheader via Cross-modal Kernel Fusion-sphere Space Learning
SU Shuzhi, GUI Yang, MA Tianbing, ZHU Yanmin, WU Kanghui
Available online  , doi: 10.11999/JEIT260494
Abstract:
  Objective  Traditional roadheader bearing fault diagnosis methods often struggle with high-dimensional and nonlinear multi-sensor data. They also fail to effectively perceive cross-modal, multi-scale fault information or integrate local and global structural features. To address these limitations, this paper proposes a Cross-modal Kernel Fusion-sphere Space Learning (CKFSL) method. By perceiving cross-modal multi-scale fault information, CKFSL extracts highly discriminative features from roadheader bearing cross-modal fault samples and improves diagnostic accuracy.  Methods  CKFSL first maps roadheader bearing cross-modal fault samples into a high-dimensional kernel space through implicit transformation. Dual extremal point anchoring and polar neighbor allocation mechanisms are then used to capture fault sample clusters with similar isomorphic information, forming kernel fusion-spheres. An adaptive binary partitioning strategy is designed according to the geometric span of internal fault samples. This strategy tightens isomorphic boundaries, constructs micro-neighbor kernel fusion-spheres, and achieves highly isomorphic manifold aggregation at the microscopic scale. A micro-neighbor kernel fusion-sphere space is further formed to re-evaluate local isomorphism (Fig. 1). To characterize wide-area topological correlations, a wide-area topological isomorphism constraint is proposed. This constraint constructs a wide-area dynamic isomorphism graph among micro-neighbor kernel fusion-spheres (Fig. 1). Finally, an objective optimization function is formulated within the space learning framework. It integrates local manifold isomorphism and wide-area topological correlations of roadheader bearing cross-modal fault samples, as shown in the CKFSL diagnostic flowchart (Fig. 2). The analytical solution for spatial projection is theoretically derived to obtain discriminative cross-modal kernel fusion-sphere space isomorphic features from roadheader bearing cross-modal fault samples.  Results and Discussions  CKFSL is first validated on the self-built AUST roadheader bearing cross-modal fault dataset, with the experimental platform shown in Fig. 3. The average recognition rates obtained with increasing numbers of training fault samples are shown in Fig. 4. On the AUST dataset, CKFSL achieves a recognition rate of 99.49% with only 70 training fault samples and reaches 100% as the number of training fault samples increases. Table 1 summarizes the standard deviations under different training fault sample sizes. The results show that CKFSL has the lowest standard deviation and stronger robustness than the other seven comparison algorithms. Three-dimensional fault feature distributions are shown in Fig. 5. The results confirm that CKFSL effectively separates highly overlapping fault samples into different clusters and reduces the boundary confusion observed in the comparison algorithms. To verify generalization capability, CKFSL is further evaluated on the public Paderborn dataset, with the experimental setup shown in Fig. 6. As shown in Fig. 7 and Fig. 8, CKFSL achieves a 100% average recognition rate across four complex fault categories. It also outperforms the comparison algorithms, which have difficulty exceeding an 85% recognition rate for the F4 fault category.  Conclusions  CKFSL effectively addresses the inability of traditional roadheader bearing fault diagnosis methods to perceive complex multi-scale fault information. By using the wide-area dynamic isomorphism graph learned in the micro-neighbor kernel fusion-sphere space, CKFSL integrates local manifold isomorphism with wide-area topological correlations of roadheader bearing cross-modal fault samples. This process enables CKFSL to extract highly discriminative cross-modal kernel fusion-sphere space isomorphic features. It improves the accuracy of roadheader bearing fault diagnosis and supports the reliability and continuous operation of roadheaders.
A General Evaluation Framework for Mission Planning Algorithms for Remote Sensing Satellite Constellations
LI Jinfei, YU Xiaogang, TIAN Jing, HE Haochen, XING Xiangwei, ZHANG Xiaohan
Available online  , doi: 10.11999/JEIT260335
Abstract:
  Objective  The rapid growth in remote sensing satellite constellations has shifted mission planning from single-satellite static scheduling to large-scale dynamic coordination across heterogeneous constellations. However, evaluation methods have not kept pace with algorithm development. Existing studies often rely on private datasets, simplified metrics centered on Completion Rate, and idealized simulations that ignore realistic constraints, such as attitude maneuvers, illumination conditions, and dynamic task insertion. These limitations prevent fair cross-paper comparison and slow engineering application. To address this gap, this paper proposes the Remote Sensing Constellation Mission Planning Benchmark (RSCMP-Bench), a general, open, and reproducible evaluation framework. It is designed as a unified benchmark for the community, similar to ImageNet in computer vision and General Language Understanding Evaluation (GLUE) in Natural Language Processing (NLP).  Methods  RSCMP-Bench consists of three components. First, the multi-scenario standard task library contains 300 standardized scenarios at three difficulty levels: Low, Medium, and High, with 100 scenarios per level. Satellite numbers range from 30 to 200, and task demands range from 56 to 560. All scenarios are generated from public Two-Line Element (TLE) data and explicitly model realistic constraints. Optical satellites require a minimum solar elevation angle, and Synthetic Aperture Radar (SAR) satellites require incidence angles within specified ranges. General constraints, such as per-orbit maximum on-time, minimum single-operation on-time, attitude maneuver time, and valid execution windows, are also modeled. The scenarios include point tasks, area tasks, static tasks, and dynamically inserted tasks. Second, the multi-dimensional effectiveness evaluation system includes a Basic Performance layer and a Dynamic Adaptability layer. The Basic Performance layer uses Completion Rate, Weighted Completion Rate, Average Response Delay, and Time Utilization. The Dynamic Adaptability layer uses multi-stage rolling evaluation with random dynamic task insertion. The Dynamic Adaptability Score measures the post-insertion Completion Rate relative to the baseline, and Dynamic Response Efficiency measures the performance gain per unit replanning time. A composite RSCMP-Bench Score is also provided. Third, the simulation and evaluation platform uses a client-server architecture. It integrates a Simplified General Perturbations 4 (SGP4) propagator, algorithm adapters, two-stage constraint verification, an intelligent scenario generator, and visualization tools. The platform has been deployed at https://www.tianzhibei.com and has supported a national competition with more than 80 research teams.  Results and Discussions  Baseline experiments comparing Random Scheduler and Priority Greedy validate the feasibility, reproducibility, and discriminative capacity of RSCMP-Bench. Random Scheduler yields very low Completion Rates of 7.3%, 3.8%, and 1.9% on the Low, Medium, and High levels, respectively. These results confirm the extreme sparsity of the feasible solution space. Priority Greedy achieves higher Completion Rates but still degrades as scenario difficulty increases, decreasing from 76.1% at the Low level to 63.7% at the Medium level and 49.2% at the High level. These findings indicate that high-difficulty scenarios remain challenging even for reasonable heuristic methods. They also show considerable room for more advanced algorithms. The dynamic adaptability protocol quantifies algorithm robustness under unexpected dynamic task insertion, which is not captured by static evaluations. The two-stage constraint verification module rejects infeasible plans and generates detailed error reports to support debugging.  Conclusions   RSCMP-Bench provides a unified, fair, and reproducible benchmark for remote sensing constellation mission planning. By combining a public library of 300 standardized scenarios, a multi-dimensional effectiveness evaluation system based on Basic Performance and Dynamic Adaptability, and a simulation and evaluation platform with realistic constraints and automated scenario generation, the framework addresses the long-standing lack of standardized evaluation in this field. Baseline results confirm its discriminative capacity and reveal clear performance bottlenecks in large-scale dynamic scenarios. Inspired by ImageNet and GLUE, RSCMP-Bench can support systematic community evaluation and fair competition. The framework has been deployed at https://www.tianzhibei.com, and its adoption can accelerate progress in intelligent mission planning for next-generation remote sensing constellations.
Dual-MPC-Driven Modeling and Spatiotemporal Evolution of Intelligent Connected Traffic Risk Fields
JIANG Linyuan, DING Fei, FAN Xuan, YANG Xuechao, SONG Aiguo, ZHANG Dengyin
Available online  , doi: 10.11999/JEIT260194
Abstract:
  Objective  With the deployment of intelligent connected vehicle-road-cloud cooperative systems, roadside infrastructure is evolving from traffic-state sensing units into intelligent decision-support platforms for multi-vehicle interaction analysis and dynamic risk inference. In highway and urban freeway scenarios, traffic operation is affected not only by the kinematic responses of individual vehicles but also by lane-changing intentions, car-following competition, and conflict propagation under local interactions. From a roadside perspective, a unified framework is therefore needed to continuously represent traffic risk, reveal its spatiotemporal evolution, and couple risk information with behavior decision-making and trajectory planning. Existing car-following models, such as the Optimal Velocity Model (OVM), Full Velocity Difference (FVD) Model, and Intelligent Driver Model (IDM), can describe speed-spacing evolution. However, these models mainly focus on longitudinal interactions and usually embed risk implicitly in safety-distance or acceleration constraints. They cannot explicitly characterize the coupling between longitudinal following and lateral lane changing, nor can they provide a continuous risk representation suitable for regional traffic assessment. Although Artificial Potential Field (APF) methods and Model Predictive Control (MPC) methods can improve trajectory safety, existing studies still lack a unified mechanism that links risk assessment, behavior decision-making, and motion planning. In addition, discrete behavior choices and continuous control actions are difficult to process efficiently within a single optimization framework.  Methods  An intelligent connected traffic risk-field model oriented toward vehicle-road cooperation is first established. The model integrates vehicle-interaction risk, lane-marking constraint risk, and road-boundary repulsive risk (Fig. 1). In the vehicle-interaction layer, motion-state-induced risk is formulated by considering the relative speed and relative orientation between the ego vehicle and surrounding vehicles. Distance-induced risk is modeled to reflect attenuation as separation distance increases (Fig. 2(a)(b)). To represent the stronger influence of forward hazards than lateral and rear hazards, a directional non-uniformity coefficient is used. This coefficient adjusts the angular attenuation of field strength and enables anisotropic spatial risk representation around the vehicle. In the road-constraint layer, lane markings and road boundaries are modeled separately. The total driving risk field is obtained by weighting and combining the lane-marking field, road-boundary field, and multi-vehicle interaction field (Fig. 2(e)(f)). Based on this representation, a dual-MPC hierarchical decision and motion-planning architecture is designed (Fig. 3). In each control cycle, the upper layer evaluates candidate behavior modes according to the vehicle state, surrounding traffic state, and dynamic risk field, and then outputs a unique behavior mode. The lower layer activates the corresponding control branch. When lane keeping is selected, longitudinal speed-planning MPC is used. When lane changing is selected, lane-change trajectory-planning MPC is activated under road-boundary, lane-marking, and safe-gap constraints.  Results and Discussions  The proposed framework reconstructs microscopic traffic risk evolution under different datasets and time scales. In the HighD highway scenario, when the slicing interval is 0.4 s, the local evolution of following and lane-changing interactions is captured in detail. This includes the process in which the ego vehicle initially follows a preceding vehicle and then starts changing to the adjacent lane (Fig. 4(a)(d)). When the interval is increased to 1.0 s, a wider spatiotemporal interaction range becomes visible, and lane-change completion and the subsequent return maneuver are identified more clearly (Fig. 4(e)(h)). In the NGSIM scenario, a smaller interval provides a finer description of the lane-change disturbance process. By contrast, a larger interval reveals the wider reconstruction of interaction relationships among the original lane, target lane, and surrounding vehicles (Fig. 4(i)(p)). These results indicate that the proposed roadside-oriented risk field can describe both local interaction details and larger-scale evolution trends, depending on the selected reconstruction interval. Sensitivity experiments further confirm the role of the directional non-uniformity coefficient. As this coefficient increases, forward risk concentration becomes stronger, local peak risk increases, and the coverage of high-risk regions decreases (Table 2). This finding shows that the coefficient effectively regulates anisotropic field distribution. Comparative experiments with IDM, OVM, FVD, and APF show that the proposed method performs better in most representative scenarios and error metrics (Fig. 5, Table 3). In the lateral cut-in scenario, its advantage lies in the early representation of lateral intrusion risk, which enables the behavior decision layer to anticipate conflict and the motion-planning layer to generate continuous evasive actions. In congested scenarios, the superposition of forward congestion risk, lateral neighboring-vehicle influence, and road-boundary constraints allows the dual-MPC controller to evaluate safety and feasibility simultaneously in local space.  Conclusions  A unified framework for roadside-oriented traffic-risk modeling and behavior-driven trajectory planning is developed. By integrating multi-vehicle interaction risk, lane-marking constraint risk, and road-boundary repulsive risk into a continuously evolving dynamic risk field, the spatial quantification of multi-vehicle interaction risk is realized. The directional non-uniformity coefficient further enables asymmetric risk perception modeling in forward, lateral, and rear directions. On this basis, a dual-MPC hierarchical architecture is constructed to couple behavior decision-making with motion planning, so that lane-keeping and lane-changing behaviors can be adaptively selected and optimized under a unified risk-driven mechanism. Experiments based on HighD and NGSIM datasets show that the proposed method can effectively characterize the spatiotemporal evolution of traffic-risk fields and outperform representative comparison models in most typical scenarios and error metrics.
A Low-latency Synchronization Header Detection Algorithm and Circuit for the JESD204C Interface
YIN Peng, ZHANG Chao, LEI Changan, HOU Weizhou, SHU Zhou, LIU Shubin, ZHU Zhangming
Available online  , doi: 10.11999/JEIT260163
Abstract:
  Objective  With rapid advances in high-speed electronics, front-end Analog-to-Digital Converters and Digital-to-Analog Converters (ADCs/DACs) continue to increase in sampling rate and resolution. Back-end Field-Programmable Gate Arrays and Application-Specific Integrated Circuits (FPGAs/ASICs) also provide stronger computing capability. These trends impose strict requirements on high-speed data interfaces, including high bandwidth, low latency, low power consumption, and reliable synchronization. As a mainstream high-speed Serializer/Deserializer (SerDes) interface, the JESD204C interface still suffers from long link initialization latency and high synchronization power consumption. These limitations restrict system real-time performance and energy efficiency. To address these issues, this study optimizes the link-layer design of the JESD204C receiver and proposes an efficient Synchronization Header (SH) detection method. The method implements exponential compression of the search set through global observation and iterative convergence. Detection efficiency is improved, fast and accurate SH positioning is achieved, link synchronization latency is reduced, and synchronization stability and energy efficiency are enhanced.  Methods  A typical JESD204C interface uses serial sliding detection for SH detection, which causes high link initialization latency and large delay jitter. To solve these problems, an Iterative Set Screening (ISS)-based SH detection algorithm is proposed. The SH detection task is modeled as the rapid localization of a deterministic pattern in a binary random sequence. A theoretical model based on information theory and stochastic processes is constructed. Expected space utilization and Bit Error Rate (BER) are introduced to support quantitative performance evaluation. In this model, SH candidate positions are defined as a dynamic set. Based on the inherent polarity inversion characteristic of the SH and global observations in each clock cycle, multilevel XOR logic is used to verify all candidate hypotheses in parallel. Non-inverting candidate positions are eliminated, and the search space is dynamically compressed. This design improves synchronization speed and position robustness, providing a low-latency and reliable initialization solution for high-speed SerDes links.  Results and Discussions  The proposed ISS-based SH detection algorithm is validated under harsh conditions, including SH crossing block boundaries and loss of lock caused by burst errors. The results demonstrate strong robustness, with rapid SH locking and link resynchronization under all test conditions (Figures 1116). To evaluate performance, four representative schemes are reproduced: a single-bit serial locking circuit, a 66-bit serial locking architecture, a register-intensive block synchronization method, and a parallel search circuit. A systematic comparison is then conducted between these schemes and the proposed design. The results show that the normalized locking time of the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method varies substantially with SH position (Figure 17(a)), especially at block boundaries (Figure 17(b)). When the SH is located at the Most Significant Bit (MSB), typical sliding detection requires about 1.8 times the time needed at the Least Significant Bit (LSB), indicating strong sensitivity to the starting position and search path. In contrast, the proposed ISS scheme maintains a stable normalized locking time within 1.0 ± 0.05 across all positions, with the standard deviation reduced by more than 70%. By evaluating all candidate positions equally through parallel filtering, the scheme eliminates position dependence. Synchronization can be completed within tens of clock cycles whether the SH is located at the LSB, the MSB, or any other position in the block. The experimental results verify that the ISS algorithm improves synchronization robustness and predictability while accelerating link initialization. Table 3 summarizes the performance metrics. The average locking time is only 24.4 clock cycles, representing an overall improvement of more than 70% compared with the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method. The standard deviation of locking time is only 4.7, indicating a more stable synchronization process. In terms of resource utilization, the design consumes 509 Look-Up Tables (LUTs) and only 2.0 mW, much lower than the 3 503 LUTs and 94.1 mW required by the register-intensive scheme. Its energy efficiency reaches 0.03 mW/bit, which is better than those of the three conventional methods. Compared with the parallel search circuit, the average locking time is reduced by 6.11%, power consumption is reduced by 50.3%, and energy efficiency is improved by 53.8%. Therefore, the proposed JESD204C receiver link shows advantages in SH detection speed, stability, power consumption, and energy efficiency.  Conclusions  An ISS-based SH detection algorithm is proposed for the JESD204C receiver. By screening the data stream in parallel through multilevel XOR logic, dynamically compressing the search space, and efficiently eliminating non-inverting candidate positions, the algorithm converges to the true SH position. This approach improves the conventional serial detection mechanism. The design is verified on the Xilinx KC705 FPGA platform. A Pseudorandom Binary Sequence 31 (PRBS31) is used to emulate the random distribution of polarity transitions, and a high-speed SubMiniature version A (SMA) cable is used for data loopback transmission. The results show that the algorithm achieves an average locking time of only 24.4 clock cycles, with a standard deviation as low as 4.7. High robustness is maintained for the SH at any position within the 66-bit block, and the energy efficiency reaches 0.03 mW/bit. The algorithm is superior to existing typical schemes in locking speed, delay stability, and energy efficiency. It provides a low-latency, reliable, and energy-efficient synchronization initialization approach for high-speed SerDes links.
Survey on Intelligent Semantic Covert Communication
FENG Zhaoxin, XU Yifan, XING Chengwen, XU Yuhua, ZHAO Nan, WANG Jinlong
Available online  , doi: 10.11999/JEIT260184
Abstract:
  Significance   As the Sixth-Generation mobile communication network (6G) evolves from the Internet of Everything to the Intelligent Internet of Everything, the communication paradigm is shifting from reliable bit transmission to effective semantic transmission. Semantic communication extracts and compresses task-related semantics to reduce redundancy and resource use. However, because semantic information is highly structured and task-specific, it is vulnerable to eavesdropping, inference, and attacks. Covert communication addresses this risk by hiding transmission behavior from unauthorized monitoring. With support from Artificial Intelligence (AI), covert communication can use reinforcement learning to adjust power and resource allocation in dynamic environments. Generative models can also conceal transmitted signals by learning and reproducing environmental patterns. However, strict covertness constraints limit the achievable transmission rate and make large-scale information transmission difficult. Intelligent semantic covert communication integrates semantic extraction with covert transmission, providing a reliable approach to secure and efficient 6G communications.  Progress   With the development of AI, especially deep learning for complex feature modeling, semantic communication can support efficient semantic extraction and nonlinear compression of multimodal data. Research on semantic communication has also shifted from Separate Source-Channel Coding (SSCC) to Joint Source-Channel Coding (JSCC), which supports end-to-end training and improved transmission performance. For image transmission, Convolutional Neural Networks (CNNs) use local receptive fields to capture spatial correlations. For sequential data transmission, Long Short-Term Memory (LSTM) networks use gating mechanisms to maintain temporal coherence. In covert communication, Generative Adversarial Networks (GANs) and diffusion models can learn the statistical patterns of environmental noise in the time, frequency, and spatial domains, thereby concealing transmitted signals. These methods reduce the effectiveness of unauthorized monitoring and detection, and improve system adaptability in dynamic environments. AI also improves autonomous decision-making in dynamic covert communication. By modeling covert transmission as a Markov Decision Process (MDP), Deep Reinforcement Learning (DRL) can learn resource allocation strategies through interaction with the environment. This approach reduces computational complexity compared with traditional convex optimization methods. By integrating semantic extraction and covert transmission, intelligent semantic covert communication further supports semantic-driven covert transmission. Large Language Models (LLMs) can evaluate semantic sensitivity and contextual risks, enabling selective covert transmission of sensitive semantic information.  Conclusions  Research on intelligent semantic covert communication shows the advantages of coordinated semantic perception and physical-layer covert mechanisms. AI improves semantic extraction efficiency and strengthens adaptation to dynamic and complex environments. By integrating semantic understanding with covert transmission strategies, intelligent semantic covert communication supports both efficiency and security for ubiquitous 6G services.  Prospects   Future research on intelligent semantic covert communication should address several key challenges, including AI-enabled detection, unified semantic metrics, lightweight model design, multimodal semantic alignment, system interpretability, and semantic hallucination. Active threat detection and adaptive defense strategies are needed to counter AI-driven surveillance. Causal reasoning in Large Multimodal Models (LMMs) can help mitigate semantic hallucination and improve data transmission reliability. Advances in model compression and cloud-edge collaboration are also needed to deploy high-complexity AI models on resource-limited terminals. With the rapid development of AI, intelligent semantic covert communication is expected to provide core support for intelligent connectivity of everything and help build more secure, efficient, and reliable 6G networks.
An Inverse-Hybrid-Modeling Digital Twin System for Natural Gas Energy Metrology
LIU Bin, ZHONG Lu, FENG Quanyuan, CHEN Yihong
Available online  , doi: 10.11999/JEIT260289
Abstract:
  Objective  Global natural gas consumption continues to increase at an average annual rate of 3.2%. A 0.1% reduction in energy measurement error can reduce trade disputes by approximately $750 million per year. Traditional studies mainly use indirect methods for energy measurement. Among these methods, chromatographic analysis and acoustic velocity correlation are the most widely used, but both have clear application limits. Chromatographic analysis has a low interference error, but it shows delayed dynamic response at high flow rates and limited dynamic calibration capability. It also has poor adaptability to multi-gas-source switching, requires manual calibration, and has high operation and maintenance costs. The lack of interoperability standards for energy networks further increases the difficulty of system integration. Acoustic velocity correlation provides a low-latency dynamic response for flow measurement, but it has a high interference error. This error may increase when the content of a single component changes, such as when the hydrogen content increases from 5% to 10%. The method may even fail under complex operating conditions, such as multi-gas-source mixing and dynamic pressure fluctuations. To address these issues, new mechanism-modeling-oriented methods have been developed. The two most representative directions are mechanism-modeling-driven methods and hybrid-modeling methods. Both methods combine multi-source data fusion with virtual-physical interaction to establish mechanism models that link flow rate, other parameters, and energy. These methods provide a new approach for accurate energy measurement, but new challenges remain. Mechanism-modeling-driven methods are usually based on static flow modeling using Computational Fluid Dynamics (CFD). However, their dynamic parameter updates are slow, with delays of more than 30 s. They also have difficulty adapting to real-time operating-condition changes, rely on large labeled datasets, and have limited interpretability. Hybrid-modeling methods still face unresolved problems in collaborative optimization across multiple modules. In addition, existing studies lack support from industrial-grade verification platforms. These limits restrict their ability to solve the dynamic response delay, parameter identification difficulty, excessive physical simplification, and weak interference resistance of traditional natural gas energy metrology methods under complex conditions. Based on recent progress in mechanism-modeling-driven and hybrid-modeling methods, this study proposes an inverse-hybrid-modeling-driven digital twin system. The system introduces a Variational AutoEncoder (VAE)-based operating-condition feature extraction algorithm and a Dynamic Bayesian Network (DBN)-based parameter calibration mechanism. It also uses a Variational Expectation-Maximization (VEM) algorithm for offline calibration. The proposed system aims to improve the accuracy, adaptability, and interference resistance of natural gas energy metrology under complex operating conditions.  Methods   A natural gas energy metrology digital twin system based on inverse hybrid modeling is proposed. The system is built on a three-tier “algorithm-system-scenario” architecture. It integrates calorific value, flow, and energy mechanism models with multi-source real-time data streams. The VAE is used for unsupervised mining of operating-condition features. A parameter self-correction loop is then constructed by combining the DBN with VEM-based system calibration. Industrial-grade devices, including ultrasonic flowmeters and gas chromatographs, are integrated to ensure real-time data transmission and closed-loop control. The system covers key operating conditions, including dynamic pressure fluctuations, hydrogen-blended gas mixtures, and multi-gas-source switching. This design ensures strong adaptability between the model and practical applications. The system was continuously verified for 25 weeks on a full-scale industrial-grade experimental platform. The results show an operational delay of ≤3.8 s, data transmission jitter of ≤0.5 s, average daily energy consumption per device of ≤1.2 kW·h, Mean Time Between Failures (MTBF) of ≥4 100 h, energy measurement error of ≤0.25%, calorific value error of ≤0.12%, and flow indication error of ≤0.2%. The system also meets security requirements through industrial Ethernet encryption and hierarchical access control. It provides engineering support for intelligent pipeline-network optimization and standardized integration.  Results and Discussions  First, a multi-level hybrid modeling framework is established. Modular hybrid modeling is achieved through the algorithm-system-scenario three-tier architecture. Numerical methods combined with data are more flexible than purely analytical models and can represent complex multiphysics systems with fewer lumped physical parameters. These parameters may change during energy measurement under mechanical, energy, and hydrodynamic effects. The VAE and DBN are used to deeply integrate mechanism models with real-time data. This reduces the parameter synchronization delay to 3.8 s and supports fluid-acoustic co-simulation and rapid response under complex operating conditions, such as hydrogen-blended natural gas. Second, an integrated algorithm for inverse hybrid modeling and system calibration is proposed. By incorporating the VAE, DBN, and VEM algorithm, the inverse hybrid modeling algorithm forms a self-supervised, adaptive intelligent system with an internal closed-loop operation. The VAE encoder compresses high-dimensional operating-condition data into low-dimensional feature vectors. This enables unsupervised feature extraction without large labeled datasets. Based on the learned internal data distribution, the VAE can also generate perturbed data similar to the input data. These data are used to simulate abnormal operating conditions and verify interference resistance. The DBN constructs a continuous “prior-evidence-posterior” iterative cycle to support system self-correction and adaptive response to operating-condition changes. The VEM algorithm compensates for systematic errors that are difficult for the DBN to capture, thereby overcoming the limits of traditional static models.  Conclusions  This study describes and validates a hybrid digital twin system that combines experimental data-driven methods with physical models. The system successfully simulates the physical characteristics of natural gas energy metrology. A full-scale test platform was constructed, and the main system parameters were validated using experimental measurement data and compared with industry benchmarks. Each independent module in the algorithm-system-scenario three-tier hybrid modeling architecture, including calorific value measurement, flow calculation, and energy conversion, was continuously verified for 25 weeks. The results confirm strong consistency between model predictions and actual measurements. On the natural gas energy metrology digital twin experimental platform, systematic validation was performed for three core functions: flow measurement under dynamic conditions, multi-component calorific value determination, and energy accumulation. The results show that the output of the digital twin model matches the physical device measurement data with an accuracy of more than 99.5%. Under complex operating conditions, such as pressure pulsations and hydrogen-blended gas mixtures, the system maintains the measurement error within 0.5%. This performance is better than that of traditional methods and meets the Class A accuracy requirements for natural gas measurement. By introducing a multi-tier hybrid modeling framework, this study addresses the parameter identification difficulty and excessive physical simplification of traditional natural gas energy metrology methods. The integration of the VAE, DBN, and VEM algorithm enables unsupervised feature extraction under complex operating conditions and adaptive calibration of model parameters. This reduces dependence on prior physical knowledge and large labeled datasets. The experimental results show that the proposed method maintains high precision and strong stability under complex scenarios, including pressure pulsations and hydrogen-blended gas mixtures, where traditional models have difficulty providing accurate descriptions.
Non-Terrestrial Network Architecture and Key Technologies for Civil Aviation
LIU Xiangnan, QIU Yu, HUANG Zhipeng, ZHANG Haijun
Available online  , doi: 10.11999/JEIT260348
Abstract:
  Significance   Civil aviation communication systems are entering a new stage of development driven by the rapid growth of global air transportation, the increasing demand for intelligent air traffic management, and the continuous expansion of in-flight connectivity services. Traditional civil aviation communication systems mainly rely on high frequency radio, high frequency radio, terrestrial air-to-ground links, and conventional satellite communication systems. These technologies have supported aircraft operation, air traffic control, airline operational communication, and low-rate data transmission for a long time. However, they still face limitations when applied to future civil aviation scenarios characterized by global coverage, high-speed mobility, low latency, high reliability, and service diversification. Particularly, terrestrial networks are difficult to deploy in transoceanic routes, polar regions, deserts, mountains, and remote airspace, while traditional geostationary satellite systems suffer from large propagation delay and limited capacity. Current systems cannot fully meet the requirements of continuous aircraft access, real-time flight monitoring, engine health data transmission, aviation safety communication, and passenger broadband services. Non-Terrestrial Networks (NTNs) provide a promising technical path for overcoming these limitations. By integrating GEOstationary satellites (GEO), Medium Earth Orbit satellites (MEO), low Earth orbit satellites (LEO), Very Low Earth Orbit satellites (VLEO), High-Altitude Platform Stations (HAPS), Unmanned Aerial vehicles (UAV), electric Vertical Take Off and Landing (eVTOL), and terrestrial infrastructures, NTN can construct a multi-layer air-space-ground integrated communication system. Such a system is able to provide continuous coverage, flexible deployment, resilient connectivity, and differentiated service support for civil aviation. NTN is becoming an important enabling technology for future civil aviation communication systems and for the digital and intelligent transformation of the aviation industry.  Progress   This paper reviews the development of NTN technologies for civil aviation and summarizes key research progress from three aspects: network architecture, access and mobility management, and resource management and scheduling. (1) We propose an aviation-oriented NTN networking framework composed of three layers: the satellite edge layer, the airborne core layer, and the terrestrial assistance layer. The satellite edge layer includes GEO, MEO, LEO, and VLEO satellites connected through inter-satellite links. GEO satellites are suitable for wide-area broadcasting and non-real-time services, MEO satellites can support navigation and intermediate-delay services, LEO satellites are suitable for low-latency and high-capacity broadband access, and VLEO satellites can further reduce propagation delay for future near-real-time aviation applications. The airborne core layer includes civil aircraft, HAPS, UAVs, and eVTOL platforms. HAPS can act as a regional relay, edge computing node, or software-defined control carrier, while UAVs and eVTOL platforms can provide flexible low-altitude coverage, emergency communication, and local access support. The terrestrial assistance layer consists of terrestrial base stations and gateway stations, which support air-to-ground communication and satellite-terrestrial interconnection. Civil aviation services can be divided into air traffic control and air traffic management services, airline operational control services, and airline passenger communication or in-flight entertainment services. Through network slicing, these heterogeneous services can be logically isolated and managed over a shared air-space-ground infrastructure. In congestion, rain attenuation, or shortened visibility-window scenarios, safety slices should be protected with the highest priority, while passenger service slices can be rate-limited, buffered, or degraded. (2) We analyze the characteristics of NR-NTN access and air-to-ground direct access in civil aviation. NR-NTN can provide continuous coverage for oceanic, polar, desert, and remote flight routes through satellites or HAPS, while air-to-ground direct access can provide low-latency and high-rate links in areas where terrestrial base stations can be deployed. However, aircraft differ significantly from ordinary terrestrial terminals because their flight trajectory, altitude, speed, and route are highly predictable. Therefore, the key issue in aviation NTN access is not only how to execute random access, but how to predict the access window, timing compensation, frequency offset, and target access node before the aircraft enters the coverage area. By using satellite ephemeris, Global Navigation Satellite System information, aircraft trajectory, and velocity parameters, civil aircraft can predict satellite visibility and pre-compute timing advance, scheduling offset, and Doppler compensation before initiating access. This transforms random access from a passive response process into a proactive and predictive access process, thereby improving access certainty and synchronization stability in highly dynamic aviation scenarios. For mobility management, a signaling interaction process for aircraft handover is designed. Based on trajectory prediction and satellite visibility prediction, the network can select a target satellite or gateway with longer residence time and better service capability. Before the aircraft reaches the handover boundary, the source and target network sides can complete context preparation, user-plane path preparation, radio resource reservation, and protocol data unit session update. When the handover condition is triggered, the aircraft performs random access to the target satellite or beam and then switches the user-plane path. This “prediction–preparation–fast handover” mechanism can reduce service interruption and maintain session continuity. For safety-critical traffic, priority and isolation policies should remain consistent and auditable throughout session preparation, handover execution, and path switching. (3)We discuss computing and caching resource management in civil aviation NTN. As onboard computing capability is limited and aviation applications generate increasing computing demands, NTN can provide mobile edge computing and caching services through LEO satellites, HAPS, UAVs, and inter-satellite cooperation. The paper introduces several computing offloading modes, including on-orbit satellite collaborative offloading, network-level integrated offloading, and cloud-edge-terminal hybrid offloading. These mechanisms can support tasks such as aviation monitoring, trajectory analysis, intelligent inference, and in-flight service optimization. In addition, caching mechanisms such as onboard satellite caching, inter-satellite cooperative caching, and named-data-networking-based content caching can improve content delivery efficiency and service continuity. Cache placement should consider content popularity, regional demand prediction, visibility windows, cache prefetching, and cooperative cache sharing among different satellite layers.  Conclusions   NTN can effectively complement traditional civil aviation communication systems by filling coverage gaps in remote and oceanic airspace, enhancing service continuity, and supporting differentiated aviation services. The proposed aviation-oriented NTN architecture integrates multi-orbit satellites, HAPS, UAVs, civil aircraft, and terrestrial infrastructures into a unified framework. The on-demand isolated slicing mechanism can provide differentiated protection for ATC/ATM, AOC, and APC/IFE services. Ephemeris-map-assisted access and predictive mobility management can improve access reliability and reduce handover interruption in high-speed aviation scenarios. Computing offloading and cooperative caching further enhance the ability of NTN to support intelligent and data-intensive aviation applications.  Prospects   Future civil aviation NTN should evolve toward deeper integration of low-altitude networks, space networks, and terrestrial networks. Cross-domain topology visualization, link-state sharing, policy distribution, and programmable logical networks are essential for improving controllability and scalability. In mobility management, integrated cross-domain handover mechanisms should be developed to cope with satellite beam switching, terrestrial cell handover, and air-to-air relay reconstruction. In resource management, communication, navigation, computing, and caching resources should be jointly scheduled and transformed according to aviation service requirements. With continuous advances in NTN architecture, network slicing, predictive access, mobility management, computing offloading, and caching, NTN is expected to provide more efficient, stable, and intelligent communication support for civil aviation and to promote the digital transformation of future air transportation systems.
Research on Secure and Covert Transmission for UAV-assisted Visible Light Communication Systems
WU Mengru, LIN Jiale, LU Weidang, LI Bo, GUO Lei
Available online  , doi: 10.11999/JEIT260239
Abstract:
  Objective  Unmanned Aerial Vehicles (UAVs) can serve as aerial base stations for Visible Light Communication (VLC) because of their mobility and on-demand coverage capabilities. However, air-ground communication links are exposed to open environments, which makes VLC vulnerable to data eavesdropping and malicious detection. To address this issue, this paper proposes a secure and covert transmission strategy for a UAV-assisted VLC system from the perspectives of Physical Layer Security (PLS) and Covert Communication. The proposed strategy jointly optimizes UAV transmit power and hovering altitude to maximize the system secrecy capacity. The optimization is subject to covert communication requirements, illumination requirements, and operational constraints on UAV transmit power and hovering altitude.  Methods  This paper investigates secure and covert communication in a UAV-assisted VLC system. A UAV-assisted VLC system model is first established. In this model, a mobile UAV equipped with a Light-Emitting Diode (LED) is used to establish a VLC link with a legitimate ground user in the presence of an eavesdropper (Eve) and a warden (Willie). An optimization problem is then formulated to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude. To solve this problem, a Two-Layer OPtimization (TLOP) algorithm is proposed. The transformed problem is decomposed into two subproblems: an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude design problem. A closed-form expression for the optimal transmit power is derived for the inner-layer problem. A Particle Swarm Optimization (PSO) algorithm is then developed to solve the outer-layer problem.  Results and Discussions  In the simulations, the proposed optimization scheme is compared with two baseline schemes. First, the convergence of the proposed TLOP algorithm is verified (Fig. 3). The results show that the algorithm converges rapidly within a limited number of iterations. Second, the optimal UAV hovering altitude with respect to the UAV horizontal coordinates is illustrated under the spatial distribution (Fig. 4). The results indicate that the optimal hovering altitude decreases as the UAV approaches the legitimate ground user. The secrecy capacity with respect to the UAV horizontal coordinates is then presented (Fig. 5). The secrecy capacity increases as the UAV approaches the legitimate ground user. This is because the legitimate VLC channel gain increases when the UAV is closer to the user. In contrast, when the UAV approaches Eve and Willie, the security and covertness constraints become stricter. The UAV is then forced to reduce its transmit power or increase its hovering altitude, which decreases the system secrecy capacity. Furthermore, the secrecy capacity of all schemes increases as ϵ increases (Fig. 6). This is because a larger ϵ relaxes the covertness requirement. The UAV can therefore adjust its hovering altitude and transmit power more flexibly to increase the system secrecy capacity. In addition, the secrecy capacity decreases as the number of symbols increases (Fig. 7). This occurs because more symbols provide Willie with more signal samples for detection, thereby improving Willie’s detection capability. Finally, the secrecy capacity of all schemes decreases as the uncertainty-region radius of illegal nodes increases (Fig. 8). This trend occurs because greater location uncertainty forces the UAV to address potential threats over a wider area. The UAV must therefore adopt a more conservative strategy under worst-case eavesdropping and detection conditions. Overall, the simulation results confirm that the proposed scheme improves the secrecy capacity of the UAV-assisted VLC system.  Conclusions  This paper investigates secure and covert communication in a UAV-assisted VLC system. The objective is to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude under covert communication, illumination, transmit power, and hovering altitude constraints. Because the formulated problem is highly non-convex, a PSO-based TLOP algorithm is designed to solve it. The proposed algorithm decomposes the problem into an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude optimization problem. Simulation results show that the proposed algorithm converges rapidly and improves the system secrecy capacity compared with the baseline schemes.
Robust Optimization of Low-altitude Communication and Computation Resources in Uncertain Environments
GONG Yucheng, LI Bin, WANG Xinyi, FEI Zesong
Available online  , doi: 10.11999/JEIT260090
Abstract:
  Objective  Low-altitude edge computing networks provide flexible computing services and extended coverage for user equipment. However, quality of service is often degraded by uncertainty in task data size and by Unmanned Aerial Vehicle (UAV) position jitter caused by environmental disturbances. Existing robust methods commonly rely on deterministic uncertainty sets, which tend to be conservative and cannot accurately describe the stochastic distribution of task demands. To address these challenges, a robust energy minimization framework is proposed for multi-UAV-assisted Mobile Edge Computing (MEC) networks. The objective is to minimize the weighted sum of system energy consumption. This is achieved by developing a joint optimization model that coordinates UAV flight trajectories, task splitting decisions, and computation and communication resource allocation. The model explicitly accounts for the dual uncertainties of task data size and UAV trajectory.  Methods  To handle the nonconvexity and strong coupling among optimization variables, the problem is first modeled as a Markov Decision Process (MDP). A comprehensive state space is defined to characterize real-time system dynamics, and a continuous action space is designed for trajectory control and resource management. A Distributionally Robust Optimization Soft Actor-Critic (DRO-SAC) algorithm is then developed to solve the MDP. In this framework, an ambiguity set based on the L1-norm distance is constructed to characterize the distributional uncertainty of the task demand distribution. A maximum-entropy reinforcement learning mechanism is used to learn an optimal policy under the worst-case distribution within the ambiguity set. In this way, UAV trajectories, task splitting, and computation and communication resource allocation are jointly optimized to improve system robustness under dynamic environmental fluctuations.  Results and Discussions  The performance of the proposed DRO-SAC algorithm is evaluated through simulations. DRO-SAC achieves faster convergence and higher rewards than Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO) algorithms (Fig. 3). For energy consumption, the proposed method consistently achieves higher efficiency under different user densities (Fig. 4). The robustness of the system against position errors is also verified, with energy fluctuations kept at a low level (Fig. 5). Dynamic trajectory adjustment further confirms that the proposed method can provide effective user coverage while reducing system energy consumption (Fig. 6).  Conclusions  A DRO-SAC-based joint optimization framework is proposed to address uncertainty in task data size and UAV position jitter in multi-UAV-assisted MEC networks. By constructing an ambiguity set for the task demand distribution and optimizing the worst-case expected objective, the proposed method mitigates the limitations of traditional deterministic models in dynamic environments. Weighted system energy consumption is minimized while latency and safety constraints are satisfied. Simulation results demonstrate that the proposed scheme achieves stable convergence and high energy efficiency, even when communication and computation resources are limited and environmental parameters fluctuate strongly.
A Lightweight True Random Number Generator Based on Chain-Coupled Oscillation Rings
ZHANG Yuan, YING Haixuan, GAO Kai, YE Jin, WANG Shuang, ZHANG Jiliang
Available online  , doi: 10.11999/JEIT260377
Abstract:
  Objective  With the rapid growth of the Internet of Things, 5G/6G, and satellite Internet, resource-constrained devices increasingly require high-quality random numbers for key generation, authentication, masking, and other security functions. Although pseudo-random number generators are efficient, their outputs may be predictable once the seed or internal state is compromised. True random number generators (TRNGs) offer a hardware root of trust by extracting entropy from physical randomness, but many existing designs rely on multiple entropy sources or complex post-processing, leading to increased area and power consumption. To address this issue, this paper proposes a lightweight TRNG based on chain-coupled oscillation rings for high-quality randomness with very low FPGA overhead.  Methods  Starting from the state evolution of a Galois oscillation ring (GARO), this work demonstrates that ideal matched-delay conditions can result in periodic and predictable oscillation. However, in practical circuits, delay mismatch, jitter, and process variation disturb the ideal evolution and can be exploited as entropy sources. On this basis, a compact delay-feedback XOR ring is proposed to enhance state uncertainty, introduce feedback competition, and improve randomness through inter-stage delay differences. In addition, a second-order oscillation ring is incorporated to eliminate the all-zero stop state and provide continuous excitation. Multiple rings are then chain-coupled, enabling adjacent rings to mutually interfere with one another and thereby generate stronger irregular oscillations. The proposed design is modeled in MATLAB and implemented on a Xilinx Artix-7 FPGA. Finally, we evaluate its performance by NIST SP 800-22, NIST SP 800-90B, bias, autocorrelation, and voltage-temperature robustness tests.  Results and Discussions  Simulation confirms that the proposed structure avoids stable periodic locking and produces sustained irregular oscillation. Experimental results show that the TRNG passes all NIST SP 800-22 tests and achieves an average minimum entropy of 0.9936 in NIST SP 800-90B test, outperforming conventional RO and GARO-based TRNGs under similar conditions. The measured bias is only 0.0228%, and the autocorrelation remains well below the threshold, indicating excellent statistical independence. The design also maintains high entropy over temperatures from 0 °C to 80 °C and supply voltages from 0.9 V to 1.1 V. Implemented on Artix-7, our proposed TRNG achieves 200 Mbps throughput using only 11 LUTs and 4 DFFs, with 0.108 W power consumption.  Conclusions  This paper presents a lightweight chain-coupled oscillation-ring TRNG that exploits delay mismatch, phase disturbance, and feedback competition to generate high-quality physical randomness. The theoretical analysis clarifies how practical nonidealities transform ideal periodic oscillation into irregular oscillation, providing a design basis for compact oscillator-based entropy sources. By combining delay-feedback XOR rings with chain-coupled mutual disturbance and continuous excitation, the proposed design enhances entropy while avoiding excessive hardware overhead and complex post-processing. FPGA implementation and statistical evaluations verify high entropy, low bias, and high randomness under voltage and temperature variations. Therefore, the proposed TRNG achieves high randomness quality and high throughput while effectively reducing hardware overhead, making it suitable for resource-constrained security applications such as IoT terminals, lightweight cryptographic modules, and embedded authentication systems.
A Noise Reduction Strategy via Coprime-Spacing Subarrays for Biodiversity Acoustic Indices
CHEN Lei, XU Zhiyong, ZHAO Zhao
Available online  , doi: 10.11999/JEIT260237
Abstract:
  Objective  As a popular tool for rapid biodiversity assessment, acoustic indices have attracted increasing attention in the field of soundscape ecology in recent years. Nevertheless, most commonly used acoustic indices are susceptible to background noise. Traditional single-channel noise reduction strategies, including spectral subtraction, high-pass filtering, and threshold detection, have been widely adopted as preprocessing approaches to optimize the calculation of acoustic indices. However, when dealing with anthropogenic interference that overlaps with biotic signals in both time and frequency domains, the denoising capability of single-channel methods degrades severely. Although spatio-temporal adaptive whitening filtering based on microphone arrays provides a feasible approach for suppressing directional interference, it suffers from a non-uniform two-dimensional spatio-temporal amplitude response and the self-cancellation of target signal in the unconstrained interference cancellation. These disadvantages lead to distortion in the time-frequency distribution of target signals, causing acoustic index calculations to deviate from the ground truth. Therefore, this study aims to propose a noise reduction strategy via coprime-spacing subarrays for biodiversity acoustic indices. This method effectively suppresses directional interference while maximally preserving the time-frequency distribution structure of biotic signals.  Methods  The noise reduction strategy based on microphone array spatio-temporal adaptive whitening filtering is proposed, incorporating the Frequency-dependent Acoustic Diversity Index (FADI), which is insensitive to fluctuations in the array's two-dimensional spatio-temporal amplitude response. A noise-robust acoustic index method, termed Adaptive Interference Cancellation–Frequency-dependent Acoustic Diversity Index (AIC-FADI), is subsequently developed. Specifically, a non-uniform linear array is first constructed using three microphones to form two dual-element subarrays with coprime spacing. This design fully exploits the high spatial resolution of wide-spacing arrays to narrow the null width in the direction of interference. Meanwhile, it avoids the physical implementation difficulties and mutual coupling effects associated with small-spacing array designs caused by the ultra-wideband characteristics of target signals. The spatio-temporal adaptive whitening filtering is then performed on each coprime-spacing subarray separately, adaptively forming two-dimensional nulls within the interference support region, thereby suppressing directional anthropogenic interference in analytical data before index calculation. Next, a frequency-dependent threshold scheme is utilized to obtain the binary spectrogram for each coprime-spacing subarray output, abating the influence from gain differences along the frequency axis for a certain direction. Afterwards, by leveraging the high spatial resolution of wide-spacing arrays and the interleaved characteristics of spatial aliasing null positions between the spatio-temporal frequency responses of the two subarrays with coprime spacing, a pointwise maximum fusion is applied to the above two binary spectrograms. This process reconstructs the binary time-frequency distribution structure of target signals outside the interference support region, leading to a single binary spectrogram where biological sound components are preserved to a great extent and anthropogenic interference is considerably suppressed. Ultimately, from this single binary spectrogram, the proportions of non-zero time-frequency bins within each frequency band are calculated and forwarded to the entropy function, resulting in the final AIC-FADI result.  Results and Discussions  The simulation result indicates that the proposed AIC-FADI maintains numerical robustness across an SINR range down to –15 dB (the yellow line in Fig. 5), substantially outperforming the classical ADI version based on single-channel noise reduction algorithm (FADI) and other ADI versions based on single-array interference suppression processing mentioned in this paper (AIC-FADI-s, AIC-FADI1, and AIC-FADI2). The real-world experiment confirms that the proposed spatio-temporal adaptive whitening filtering effectively suppresses wideband interference signals in complex scenarios, thereby improving the SINR of the analyzed recording. This enables some weaker biotic signals to exceed their corresponding frequency-dependent adaptive thresholds, greatly reducing missed detection of the target signal. In addition, by performing pointwise maximum fusion of the binary spectrograms from the two coprime-spacing subarray outputs, AIC-FADI further alleviates the extent of target signal missed detection (Fig. 8). Nevertheless, the real-world experiments also reveal that the interference suppression performance of AIC-FADI degrades for highly time-varying interference components.  Conclusions  This paper addresses the challenge of calculating acoustic indices reliably in complex soundscapes where directional anthropogenic interference overlaps with biotic signals in both time and frequency domains. A noise reduction strategy using coprime-spacing subarrays is proposed, and a new noise-robust acoustic index (AIC-FADI) is then developed. The method is evaluated through simulations and real-world recordings, and the results show that: (1) By applying spatio-temporal adaptive whitening filtering on each coprime-spacing subarray followed by pointwise maximum fusion, the proposed method achieves both wideband interference suppression capability and target information fidelity in complex soundscapes containing strong interference. (2) As a result, the proposed AIC-FADI maintains numerical robustness down to –15 dB SINR, substantially outperforming the classical FADI algorithm and other ADI versions based on single-array interference suppression methods. (3) The proposed method provides a feasible technical solution for extending the practical application scenarios and spatio-temporal coverage of biodiversity acoustic indices in human-dominated areas. However, this study only considers directional interference that is relatively stable or slowly time-varying. Hence, the interference suppression performance degrades for highly time-varying or uncorrelated noise components. These challenges should be addressed in future work through more advanced signal processing techniques to further improve the robustness of acoustic indices in highly complex acoustic environments.
A Survey of Quantum Covert Communication Integration Schemes and Application Scenarios
SUN Yiheng, XU Yongjun, ZHANG Haibo, HUANG Zishan
Available online  , doi: 10.11999/JEIT260282
Abstract:
  Significance   With the growing demand for network communication security, research and development in covert communication and quantum communication have continued to evolve. However, current covert communication suffers from inherent security vulnerabilities; the transmission reliability of quantum communication has been limited by information eavesdropping and harmful interference. Therefore, quantum covert communication has become a research hotspot, integrating the advantages of both covert and quantum communication while addressing their respective security limitations. To this end, this paper provides a comprehensive survey of quantum covert communication integration schemes and application scenarios, including the principles of covert communication and typical enabling techniques; protocols for quantum communication and important quantum techniques; and three types of quantum covert communication integration schemes summarized by different application scenarios. This paper contributes to the design of advanced secure communication networks while offering guidance for the development of future quantum covert communication systems.  Progress   This paper presents a comprehensive survey of recent advances in quantum covert communication integration schemes and application scenarios, with an in-depth discussion of the principles of covert communication and key enabling techniques, such as Fluid Antenna (FA), Reconfigurable Intelligent Surface (RIS), and Unmanned Aerial Vehicle (UAV). FA actively reshapes wireless channel characteristics, particularly the spatial correlation of multipath components, by dynamically adjusting the transmitter physical configuration, thereby reducing information leakage. In Non-Line-of-Sight (NLoS) scenarios, RIS can dynamically alter the direction of reflected transmission of the incident signal, not only enhancing the Channel State Information (CSI) quality of the covert signal but also reducing signal leakage. In flexible or temporary communication networks, UAVs can increase CSI uncertainty, preventing unauthorized users from establishing a stable monitoring model and thereby complicating eavesdropping. Then, key protocols and significant techniques of quantum communication are introduced, including BB84, B92, and E91 for Quantum Key Distribution (QKD), and BF02, Two-Step for Quantum Secure Direct Communication (QSDC). Additionally, the quantum repeaters and Quantum Random Number Generator (QRNG) are reviewed. Based on different application scenarios, quantum covert communication integration schemes can be categorized into enabling, covert, and symbiotic integration schemes, depending on the integration mechanisms. To be specific, the enabling integration scheme leverages the unconditional security of quantum communication to address the security vulnerabilities in covert communication, the covert integration scheme utilizes enabling techniques in covert communication to reduce the detection probability of quantum communication, and the symbiotic integration scheme combines both advantages of covert communication and quantum communication to achieve mutual empowerment and deep symbiosis. Finally, critical challenges are highlighted, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles in large-scale applications. Promising directions for future research are also identified, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development.  Prospects   Despite remarkable progress in preliminary applications and specific scenarios, research on quantum covert communication remains in its infancy. As quantum covert communication scenarios become increasingly diverse and complex, future studies should prioritize challenges that restrict further development and large-scale application of quantum covert communication. The stringent hardware precision requirements are the primary challenge, limiting reliable transmission distance and stability. Low resource allocation efficiency is another challenge, as the quantum covert communication system that generates quantum entanglement over lossy channels remains subject to the Square Root Law (SRL) constraints, while signal transmission exhibits burstiness and dynamics. Additionally, high deployment costs and the lack of standardization present significant hurdles. To address the challenges mentioned, future directions should include R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development to facilitate the development of high-performance, large-scale, and multi-scenario quantum covert communication.  Conclusions  This paper provides a comprehensive survey of quantum covert communication with particular emphasis on integration schemes and application scenarios. The fundamentals and typical enabling techniques of covert communication are first reviewed, highlighting its Low Probability of Detection (LPD) secure paradigm and unique channel characteristics. The typical protocols and important techniques of quantum communication are then examined, including QKD, QSDC, quantum repeaters, and QRNG. Three types of quantum covert communication integration schemes have been further classified by different integration mechanisms and corresponding application scenarios. Finally, several existing challenges are identified, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles to large-scale applications. Relevant research directions are also outlined, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development. These directions are expected to serve as a valuable reference for advancing and standardizing quantum covert communication in future secure networks.
Millimeter-Wave Air-to-Ground Channel Prediction Assisted by Visual Information of the Propagation Environment
CHENG Yuanxun, HU Qingsong, ZHANG Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260274
Abstract:
  Objective  Accurate prediction of air-to-ground (A2G) channel states is essential for adaptive transmission and resource optimization in unmanned aerial vehicle (UAV) communications. In urban millimeter-wave scenarios, however, A2G links are highly sensitive to blockage, reflection, scattering, and the rapidly changing geometric relationship among the transmitter, the receiver, and surrounding buildings. As a result, the channel exhibits strong spatial and temporal nonstationarity, and conventional pilot- or feedback-based acquisition methods may become ineffective because the obtained channel state information is easily outdated. Recent data-driven approaches have shown potential, but many of them rely heavily on historical channel observations or directly use raw images as network inputs, which may introduce redundant visual information and weaken physical interpretability. To address these limitations, this paper proposes a vision-assisted millimeter-wave A2G channel prediction method that extracts low-dimensional geometric features from the propagation environment instead of using raw visual data directly. The objective is to preserve the key structural information governing channel evolution while reducing irrelevant redundancy, thereby improving the prediction of channel.  Methods  A communication-and-sensing integrated dataset with strict spatial and temporal alignment is established for millimeter-wave UAV A2G channel prediction. On the sensing side, a high-fidelity three-dimensional urban scenario containing 23 buildings, roads, and intersections is constructed in Unreal Engine 4.27, where synchronized RGB and depth images are collected through AirSim using a multirotor UAV equipped with RGB and depth cameras. The UAV flies along 10 preset trajectories at a height of 55 m with a spatial sampling interval of 1 m, yielding 2160 valid visual samples (Fig. 1, Fig. 2). On the communication side, the same scene is reconstructed in Wireless InSite, and the transmitter-receiver positions are synchronously updated along the same trajectories to ensure frame-level alignment between visual and channel data (Fig. 3). To obtain compact and physically meaningful environmental representations, a cross-modal spatial feature extraction scheme is developed. Buildings are first detected from RGB images using YOLO-V8 (Fig. 4), and the detected regions are then registered with depth images to reconstruct three-dimensional point clouds. After Euclidean clustering and axis-aligned bounding-box fitting, key geometric attributes, including planar position, height, and volume, are extracted. These features are combined with the transmitter-receiver distance to form the spatial feature vector of each frame, and their relevance to path loss, received power, and RMS delay spread is evaluated through cosine-similarity-based correlation analysis (Fig. 6). Based on the extracted features, a hybrid Transformer-MLP network is designed for channel prediction (Fig. 5). Building features are first projected into a latent space, and a stacked Transformer encoder is employed to capture global interactions among buildings through masked multi-head self-attention. Masked average pooling is then used to aggregate building-level representations into a scene-level environmental descriptor, which is concatenated with the link distance feature and fed into a multilayer perceptron regressor to predict the three target channel parameters.  Results and Discussions  The results confirm the effectiveness of the proposed spatial feature representation. Correlation analysis shows that the extracted geometric features are consistently related to path loss, received power, and RMS delay spread under different aggregation strategies (Fig. 6), indicating that compact building descriptors can effectively characterize the propagation environment. Among them, building height exhibits the strongest correlation with all three channel parameters, highlighting its important role in blockage, attenuation, and multipath propagation in urban millimeter-wave A2G channels. In prediction experiments, the proposed method accurately tracks the variation trends of all three targets. It remains effective in deep-fading and sharp-fluctuation regions for path loss prediction (Fig. 7), achieves high consistency with the ground truth for RMS delay spread (Fig. 8), and follows rapid local fluctuations of received power with good fidelity (Fig. 9). In contrast, the benchmark model only captures the general trend and shows larger deviations in peaks, valleys, and abrupt-changing intervals. Residual analysis further demonstrates the superiority of the proposed method. Its errors are more concentrated around zero and fluctuate within narrower ranges than those of the benchmark model across all three tasks (Fig. 10). Quantitatively, both the mean absolute error and the root mean squared error are reduced (Fig. 11). In addition, the model maintains acceptable complexity, with about 5.5 M parameters and a single-frame inference delay of about 3.4 ms, indicating good potential for real-time deployment.  Conclusions  A vision-assisted millimeter-wave A2G channel prediction method for UAV communications is proposed. By constructing a strictly aligned communication-and-sensing dataset and extracting low-dimensional spatial features with clear physical meaning, the method establishes an effective mapping from environmental geometry to channel parameters. The proposed Transformer-MLP framework achieves accurate prediction of path loss, received power, and RMS delay spread, while offering better interpretability, robustness, and efficiency than the benchmark model.
Transfer Learning Aided CNN for Efficient Data Detection in ReRAM with Sneak-Path Interference
DAI Bin, WU Anni
Available online  , doi: 10.11999/JEIT260354
Abstract:
  Objective  Sneak path interference (SPI) in resistive random-access memory (ReRAM) introduces unpredictable inter-cell correlations, significantly increasing the complexity of signal detection. Traditional detection methods typically rely on assumptions about known channel noise states, resulting in limited generalization capability in practical applications. To address this issue, three data detection methods based on convolutional neural networks (CNNs) are proposed, which can effectively model and mitigate interference without relying on prior channel information: first, a method combining constrained coding with a multi-layer CNN, which uses constrained coding to determine the sneak path interference state and recover data; second, a dual-CNN framework that first employs a lightweight CNN for sneak path interference identification, followed by a multi-layer CNN for refined detection; third, an approach incorporating transfer learning, which maintains detection accuracy while reducing the required training sample size to one-thousandth of that of traditional methods. Simulation results demonstrate that the proposed method achieves superior bit error rate (BER) performance under unknown channel conditions, with a BER reduction of at least half relative to existing algorithms, approaching the theoretical performance limit. Moreover, the integration of transfer learning reduces the required training samples from \begin{document}$ {10}^{6} $\end{document} to \begin{document}$ 1000 $\end{document}, corresponding to a reduction of three orders of magnitude.  Methods  To address distinct challenges in sneak path interference detection, this paper proposes three methods sequentially:1. The integrated constrained coding aided convolutional neural network (CC-CNN) detection framework effectively addresses the complex inter-cell correlations introduced by sneak path interference. This approach first employs constrained coding to detect the presence of interference and subsequently utilizes a CNN to learn and capture the random correlations under the influence of interference, thereby achieving accurate signal recovery.2. The dual-CNN-based detection method resolves the code rate loss associated with traditional constrained coding. By directly leveraging a CNN to learn and identify sneak path interference patterns from raw data, this method eliminates the need for redundant coding or additional overhead. It ensures high-precision interference detection while preserving the overall code rate performance of the system.3. The transfer learning-based CNN (TL-CNN) detection method overcomes the dependence of high-performance CNNs on large-scale training datasets. By reusing knowledge from pre-trained models, this method enables rapid adaptation to ReRAM signal detection tasks. It significantly reduces the required number of training samples while maintaining high detection accuracy and resource efficiency, thereby enhancing the feasibility of the solution in practical scenarios.  Results and Discussions  Simulation results demonstrate that the performance of the three proposed methods consistently approaches the theoretical lower bound (Fig.6), outperforming baseline methods such as the Belief Propagation (BP) detector, Deep Neural Network (DNN) detector, and Elementary Signal Estimator (ESE) detector. The two-step network achieves performance comparable to that of the single-step network while successfully avoiding code rate loss. Notably, the transfer learning-aided CNN attains near-optimal BER with only 1000 target domain samples, and its performance stabilizes when the sample size exceeds 1000 (Fig.7), fully validating its data efficiency. The integration of SK modules enables the models to effectively capture SPI-induced spatial correlations, while the transfer learning strategy ensures the models’ robust performance under different noise conditions.  Conclusions  The crossbar array architecture of ReRAM is susceptible to sneak-path interference during storage operations, leading to reduced data reliability. To address this issue, this paper proposes three deep learning-based detection methods. Type-I integrates constrained coding with a CNN to achieve efficient and fast interference detection. Type-II adopts a two-stage processing approach: it first classifies interference patterns in the memory array and then performs detection specifically on affected units, thereby ensuring high detection accuracy while minimizing coding rate loss. Type-III introduces a transfer learning framework that leverages a pre-trained model from the source domain, significantly reducing the number of training samples required in the target domain and effectively lowering training overhead. Experimental results show that under different noise conditions, all three proposed methods achieve performance close to the theoretical lower bound, providing an effective solution for enhancing the reliability of ReRAM storage systems.
Physical-layer Security in Visible Light Communications: Fundamental Theories, Key Techniques, and Future Challenges
WANG Jinyuan, YAN Xinrun, LIN Zihan, LI Yuanyuan, LI Zheng, ZHANG Xin
Available online  , doi: 10.11999/JEIT260338
Abstract:
  Significance   Due to the broadcast nature of optical signals, information security represents a critical research direction in visible light communication (VLC). Conventional encryption techniques address network security issues at the upper layers of the protocol stack through access control, cryptographic protection, and end-to-end encryption. However, their security relies on the assumption that eavesdroppers possess limited computational capabilities, an assumption that currently faces significant challenges. In recent years, physical layer security (PLS) has emerged as a novel information security paradigm and has attracted considerable attention from researchers worldwide. PLS exploits the randomness, heterogeneity, and distinctiveness between the main channel and the eavesdropping channel to achieve secure information transmission at the physical layer. To date, extensive research achievements have been made regarding PLS techniques in conventional radio frequency wireless communications (RFWC). Nevertheless, due to substantial differences in frequency bands, transmitted signals, power representations, and channel characteristics, PLS research results from RFWC systems cannot be directly applied to VLC. Although scholars worldwide have conducted research on VLC PLS technology, the foundational theories, key techniques, and future challenges involved in VLC PLS still lack a systematic review. To bridge this gap, this paper presents a comprehensive survey of VLC PLS technology.  Progress   To evaluate and enhance system performance, a classic VLC PLS system model—comprising the received signal model, the input constraint model, and the channel gain model—is initially established. A comprehensive theoretical framework for performance evaluation is then developed, encompassing instantaneous performance metrics, statistical performance metrics, and asymptotic performance metrics. Specifically, to characterize instantaneous performance, existing works on instantaneous secrecy capacity and instantaneous secrecy rate across different scenarios are summarized. As statistical performance metrics, average secrecy capacity, average secrecy rate, secrecy outage probability, probability of strictly positive secrecy capacity, and interception probability are analyzed. To demonstrate asymptotic performance, secrecy diversity order and secrecy degrees of freedom are derived. Furthermore, to enhance the PLS performance, advanced technologies, including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity, are summarized.  Prospects   Despite existing research achievements, numerous challenges remain in VLC PLS. This paper identifies four critical challenges: (i) Accurate PLS performance limit: Deriving exact expression of secrecy capacity under VLC's unique physical constraints remains challenging. (ii) Incomplete evaluation framework: Some key metrics widely used in RFWC have not been investigated in VLC, and the construction of a comprehensive VLC PLS performance evaluation framework remains unresolved. (iii) Limitations of existing methods: Conventional PLS performance enhancement methods typically adopt a “modeling-optimization-verification” separated research paradigm, often falling into a vicious cycle of “inaccurate modeling-suboptimal solutions-limited performance gains”. Therefore, it is imperative to integrate novel technologies (such as deep learning, reinforcement learning, and digital twins) to construct a data-model dual-driven framework for VLC PLS performance enhancement. (iv) Hardware platform gap: The absence of dedicated hardware platforms featuring adversarial topologies and real-time processing capabilities significantly impedes the practical deployment of VLC PLS technologies. Therefore, addressing these challenges is essential for transitioning VLC PLS from theoretical advances to commercial applications.  Conclusions  The broadcast nature of optical signals renders VLC systems vulnerable to eavesdropping attacks. This paper presents a comprehensive survey of PLS in VLC, covering system models, performance metrics (instantaneous, statistical, and asymptotic), and key performance enhancement technologies including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity. Despite significant progress, challenges remain in establishing accurate performance bounds, complete evaluation frameworks, novel enhancement techniques, and practical hardware implementations. By exploiting channel disparities at the physical layer without relying on complex encryption, PLS represents a paradigm shift in security assurance, paving the way for next-generation secure and reliable VLC networks.
Research on Energy Efficiency Optimization of Rotatable Hybrid Intelligent Reflecting Surface Communication
ZHANG Guangchi, GUO Xuan, WANG Luyao, CUI Miao, FU Hao
Available online  , doi: 10.11999/JEIT260119
Abstract:
  Objective  With the evolution of 6G communication networks, reconfigurable intelligent surfaces (RIS) have emerged as a pivotal technology for reshaping wireless environments and enhancing spectral efficiency. However, conventional fixed RIS architectures face two critical challenges in practical deployment: the “angle mismatch” loss, where the effective aperture significantly diminishes when users are located at large angles from the RIS normal, and the “energy consumption bottleneck,” caused by the high cumulative power consumption of radio frequency (RF) circuits and static control elements in large-scale arrays. Existing research often treats mechanical rotation and element switching in isolation, lacking a unified framework to balance the trade-off between mechanical/circuit energy consumption and communication gain. To address these limitations, this paper investigates a rotatable and switchable hybrid RIS (H-RIS) assisted downlink communication system. The primary objective is to maximize the system’s energy efficiency (EE) by jointly optimizing the base station transmit power, subarray activation states, physical rotation angles, and electronic phase shifts. This approach aims to introduce mechanical rotation degrees of freedom to compensate for path loss and employ dynamic switching mechanisms to reduce redundant power consumption, thereby achieving sustainable green communication.  Methods  A joint optimization framework is established for the H-RIS aided single-user multiple-input single-output (MISO) system. The system model explicitly accounts for the dynamic power consumption induced by mechanical rotation and the static power consumption of active subarrays. The resulting optimization problem is formulated as a non-convex mixed-Integer non-linear programming (MINLP) problem, involving coupled binary variables (activation status) and continuous variables (power, angles, phases). To solve this challenging problem, a block coordinate descent (BCD)-based alternating optimization (AO) algorithm is proposed to decouple the variables into three sub-problems.Firstly, to tackle the exponential complexity caused by binary switching variables, a channel contribution-based ranking strategy is developed. By performing eigenvalue decomposition on the cascaded channel correlation matrix, the priority of each subarray is quantified, reducing the search space from exponential to linear.Secondly, for the power allocation sub-problem, the non-convex fractional objective function is transformed into a parametric subtractive form using the Dinkelbach algorithm, which is then solved via the interior-point method.Thirdly, for the physical rotation and electronic phase optimization, the problem is decomposed into single-variable sub-problems. A Golden Section Search algorithm is employed to iteratively find the optimal rotation angle and phase shift for each subarray within bounded constraints, ensuring the monotonic convergence of the objective function.  Results and Discussions  Extensive simulations are conducted to evaluate the performance of the proposed H-RIS scheme compared with benchmark schemes, including “Only-Rotation” (always on), “Only-Switching” (fixed angle), and “Conventional” (fixed and always on).The simulation results regarding the maximum transmit power Pmax(Fig. 2 and Fig. 3) demonstrate that the proposed method achieves the highest energy efficiency across the entire power range. Specifically, in the low power regime, the proposed algorithm intelligently turns off redundant subarrays where the rate gain cannot offset the circuit power cost, thereby significantly outperforming the “Only-Rotation” scheme which suffers from high static power consumption.The impact of user distance is also analyzed (Fig. 4 and Fig. 5). Results indicate that the proposed scheme maintains high spectral efficiency comparable to the “Only-Rotation” scheme by dynamically adjusting the rotation angles to align with the Line-of-Sight (LoS) path, effectively compensating for the angle mismatch loss observed in the “Only-Switching” and “Conventional” schemes.Furthermore, the activation pattern of the subarray varies in a “U” shape with distance (Table 1), which allows for flexible adjustment of array size and orientation according to user-RIS geometry.  Conclusions  This paper proposes an energy-efficient transmission scheme for H-RIS aided communication systems by integrating mechanical rotation and dynamic switching capabilities. A low-complexity BCD-based algorithm is developed to jointly optimize the transceiver design. The results confirm that introducing mechanical rotation significantly mitigates the angle mismatch loss, while the proposed channel contribution-based switching strategy effectively eliminates redundant energy consumption. The proposed H-RIS architecture offers a superior trade-off between spectral efficiency and energy efficiency compared to traditional fixed RIS architectures, providing a viable solution for future green 6G networks.
A Radio Frequency Fingerprint Open-set Identification MethodCombining Multi-scale Wavelet Front-end and Hyperspherical Metric Learning
TIAN Xinyu, LI Zirui, ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260214
Abstract:
  Objective  Open-set Radio Frequency Fingerprint (RFF) identification under low Signal-to-Noise Ratio (SNR) conditions is challenging because fingerprint features are easily masked by noise, multipath effects induce nonlinear distortions, and existing methods struggle with feature extraction and unknown device detection. This study proposes a deep learning framework that integrates a multi-scale wavelet front-end with hyperspherical metric learning to achieve robust open-set RFF identification.  Methods  The proposed method, MS-RANet, comprises three key components. First, a multi-scale wavelet front-end based on one-dimensional stationary wavelet transform performs full-resolution, multi-scale decomposition of I/Q signals, preserving discriminative fingerprint information while suppressing noise. Second, a multi-scale residual attention network incorporates deep residual learning, global self-attention, and Bidirectional LSTM (BiLSTM) to enhance sensitivity to subtle fingerprint features and capture long-range temporal dependencies. Third, hyperspherical metric learning constrains the feature space onto a unit hypersphere, optimizing angular margins to produce compact intra-class and separable inter-class feature distributions. Unknown devices are subsequently detected using cosine similarity.  Results and Discussions  Experiments on a high-fidelity IEEE 802.11 simulation dataset demonstrate the effectiveness of MS-RANet. The method achieves an average classification accuracy of 65.34% across SNR levels from –5 dB to 20 dB, and an Area Under the Curve (AUC) of 0.81 at –5 dB SNR, outperforming DNN, GRU, CNN-LSTM, ResNet50, and DRSN-CA. Confusion matrices and Receiver Operating Characteristic (ROC) curves confirm robustness under extreme channel conditions. t-SNE visualization shows well-separated, compact clusters for known devices, while unknown samples are effectively isolated from known class regions. Ablation studies verify the contributions of the multi-scale wavelet front-end, global attention, BiLSTM, and hyperspherical metric learning modules.  Conclusions  This study presents a robust open-set RFF identification method combining a multi-scale wavelet front-end with hyperspherical metric learning. The framework exhibits strong noise resilience, enhanced feature discrimination, and reliable detection of unknown devices under low-SNR and multipath fading conditions. Future work will focus on reducing computational complexity, improving inference speed, evaluating generalization across diverse scenarios and protocols, and integrating the method with complementary physical-layer security mechanisms for collaborative authentication.
Semantic Relation-enhanced Adaptive Graph Representation Learning for Next POI Recommendation
WANG Zhuolu, XU Shenghua, WANG Yong, JIANG Shunshun
Available online  , doi: 10.11999/JEIT251357
Abstract:
  Objective  In recent years, next Point Of Interest (POI) recommendation has played an increasingly important role in Location-Based Social Networks (LBSNs). However, existing Graph Representation Learning (GRL)-based recommendation methods have struggled to balance node distributions across different domains (i.e., node types) effectively and have often overlooked feature differences among heterogeneous relations. Thus, complex semantic dependencies in contextual information cannot be fully captured when users’ temporal preference patterns are modeled.  Methods  To address these issues, a next POI recommendation method based on Semantic Relation-enhanced adaptive Graph Representation Learning (SR-GRL) is proposed. A heterogeneous transition graph is constructed to integrate three entity types, namely POIs, POI categories, and regions, and their complex interrelationships. An adaptive balanced random walk sampling strategy is designed to balance node distributions across different domains dynamically and to reduce information redundancy. A type-aware attention mechanism is then used to learn semantic associations among nodes through relation-specific transformation matrices, so that feature differences across node types can be identified effectively. The obtained disentangled POI representations are then used for spatiotemporal encoding of user check-in sequences, and a self-attention mechanism is applied to aggregate users, temporal preference features. Finally, next POI recommendation is generated through a Softmax function.  Results and Discussions  Experiments on the Foursquare datasets from Tokyo and New York and the Sina Weibo dataset from Shanghai show that, compared with state-of-the-art baselines, the SR-GRL method achieves Recall@10 improvements of 2.22%\begin{document}$ \sim $\end{document}24.16%, F1@10 improvements of 1.16%\begin{document}$ \sim $\end{document}10.48%, and NDCG@10 improvements of 3.01%\begin{document}$ \sim $\end{document}17.37%, indicating better recommendation performance.  Conclusions  Overall, the SR-GRL approach can balance the distributions of different node types dynamically and strengthen the modeling of complex semantic dependencies in heterogeneous contextual information.
Communication, Computation, and Caching Resource Collaboration for Heterogeneous Artificial Intelligence Generated Content Service Provisioning
WU Mengru, GAO Yu, ZHAO Bo, XU Bo, SUN Hao, GUO Lei
Available online  , doi: 10.11999/JEIT251300
Abstract:
  Objective  In the Artificial Intelligence of Things (AIoT), Edge Servers (ESs) provide intelligent content generation services to AIoT devices by utilizing cached Artificial Intelligence Generated Content (AIGC) models. However, the limited computing resources and caching capacity of ESs make it difficult to support the large-scale caching demands of heterogeneous AIGC services. To address this issue, a communication, computation, and caching resource collaboration scheme is proposed based on a combined cloud-edge and edge-edge collaborative framework. The scheme considers three representative AIGC services: lightweight AIGC services, computation-intensive AIGC services, and preprocessing-based AIGC services. The objective is to minimize the total AIGC service latency through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions.  Methods  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. First, an AIGC service-oriented AIoT system model is established to incorporate both cloud-edge and edge-edge collaboration. An optimization problem is then formulated to minimize the total latency of AIGC services through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions. Because the formulated problem is non-convex, an Alternating Optimization (AO) algorithm is proposed. The original problem is decomposed into three subproblems. These subproblems are solved using the Successive Convex Approximation (SCA) method, Karush-Kuhn-Tucker (KKT) conditions, and an improved Harris Hawks Optimization (HHO) algorithm.  Results and Discussions  Simulation experiments compare the proposed joint optimization scheme with three baseline methods: Particle Swarm Optimization (PSO), fixed resource allocation, and random offloading and caching. First, the convergence of the proposed AO algorithm is verified (Fig. 2). The results show that the algorithm converges rapidly within a limited number of iterations across different subproblems. Second, increasing transmission bandwidth significantly reduces the total AIGC service latency (Fig. 3). This occurs because each device obtains more bandwidth resources for task transmission, and the ES can allocate more bandwidth to deliver generated content in the downlink. Furthermore, the total AIGC service latency decreases as the ES storage capacity increases for all schemes (Fig. 4). Greater storage capacity enables the ES to store more AIGC models, which reduces the transmission delay between the ES and the cloud server. Moreover, when the required floating-point operations per bit increase, the total AIGC service latency rises significantly across all schemes (Fig. 5). Finally, the total AIGC service latency decreases as the maximum transmit power of the Base Station (BS) increases (Fig. 6). This occurs because higher BS transmit power improves the downlink signal-to-noise ratio, which increases the downlink transmission rate and reduces overall service latency. The proposed scheme demonstrates better performance than the baseline schemes, particularly under high computational demand.  Conclusions  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. The objective is to minimize total AIGC service latency through joint optimization of the transmit power of AIoT devices and BSs, computing resource allocation, AIGC model deployment, and service offloading decisions under computation and caching resource constraints. Because the formulated problem is a mixed-integer nonlinear programming problem, an efficient AO algorithm is developed. The original optimization problem is decomposed into three subproblems, which are solved using the SCA algorithm, KKT conditions, and the HHO algorithm, respectively. Simulation results show that the proposed algorithm reduces the total AIGC service latency compared with the baseline schemes.
Performance Analysis and Rapid Prediction of Long-range Underwater Acoustic Communications in Uncertain Deep-sea Environments
CHEN Xiangmei, TAI Yupeng, WANG Haibin, HU Chenghao, WANG Jun, WANG Diya
Available online  , doi: 10.11999/JEIT251244
Abstract:
  Objective  In complex and dynamically changing deep-sea environments, the performance of underwater acoustic communications shows substantial variability. Feedback-based channel estimation and parameter adaptation are impractical in long-range scenarios because platform constraints prevent reliable feedback channels and the slow propagation of sound introduces significant delay. In typical long-range systems, environmental dynamics are often ignored and communication parameters are selected heuristically, which frequently leads to mismatches with actual channel conditions and causes communication failures or reduced efficiency. Predictive methods able to assess performance in advance and support feed-forward parameter adjustment are therefore required. This study proposes a deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications under uncertain environmental conditions to enable efficient and reliable parameter–channel matching without feedback.  Methods  A feed-forward method for underwater acoustic communication performance analysis and rapid prediction is developed using deep-learning-based sound-field uncertainty estimation. A neural network is first used to estimate probability distributions of Transmission Loss (TL PDFs) at the receiver under dynamic environments. TL PDFs are then mapped to probability distributions of the Signal-to-Noise Ratio (SNR PDFs), enabling communication performance evaluation without real-time feedback. Statistical channel capacity and outage capacity are analyzed to characterize the theoretical upper limits of achievable rates in dynamic conditions. Finally, by integrating the SNR distribution with the bit-error-rate characteristics of a representative deep-sea single-carrier communication system under the corresponding channel, a rate–reliability prediction model is constructed. This model estimates the probability of reliable communication at different data rates and serves as a practical tool for forecasting link performance in highly dynamic and feedback-limited underwater acoustic environments.  Results and Discussions  The method is validated using simulation data and sea trial data. The TL PDFs predicted by the deep learning model show strong consistency with the traditional Monte Carlo (MC) method across multiple receiver locations (Fig. 6). Under identical computational settings, deep-learning-based TL PDF prediction reduces computation time by 2\begin{document}$ \sim $\end{document}3 orders of magnitude compared with the MC method. The chained mapping from TL PDFs to SNR PDFs and then to channel capacity metrics accurately represents the probabilistic features of communication performance under uncertain conditions (Fig. 7 and Fig. 8). The rate–reliability curves derived from the deep-learning-based TL PDFs are highly consistent with MC-based results. In the high sound-intensity region, prediction errors for reliable communication probabilities across data rates range from 0.1% to 3%, and in the low sound-intensity region errors are approximately 0.3% to 5% (Fig. 12). Sea trial results further indicate that predicted rate–reliability performance agrees well with measured data. In the convergence zone, deviations between predicted and measured reliability probabilities at each rate range from 0.9% to 4%, and in the shadow zone from 1% to 9% (Fig. 18). Under a 90% reliability requirement, the maximum achievable rates predicted by the method match the measurements in both the convergence and shadow zones, demonstrating accuracy and practical applicability in complex channel environments.  Conclusions  A deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications in uncertain deep-sea environments is developed and validated. The framework builds a chained mapping from environmental parameters to TL PDFs, SNR PDFs, and communication performance metrics, enabling quantitative capacity assessment under dynamic ocean conditions. Predictive “rate–reliability’’ profiles are obtained by integrating probabilistic propagation characteristics with the performance of a representative deep-sea single-carrier system under the corresponding channel, providing guidance for parameter selection without feedback. Sea trial results confirm strong agreement between predicted and measured performance. The proposed approach offers a technical pathway for feed-forward performance analysis and dynamic adaptation in long-range deep-sea communication systems, and can be extended to other communication scenarios in dynamic ocean environments.
Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning in Wireless Networks
LIU Jingyuan, MA Ke, XU Runchen, CHANG Zheng
Available online  , doi: 10.11999/JEIT251221
Abstract:
  Objective  Multimodal Federated Learning (MFL) uses complementary information from multiple modalities, yet in wireless edge networks it is restricted by limited energy and frequent missing modalities because many clients store only images or only reports. This study presents Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning (CREEMFL), which applies selective completion and joint communication–computation optimization to reduce training energy under latency and wireless constraints.  Methods  CREEMFL completes part of the incomplete samples by querying a public multimodal subset, and processes the remaining samples through zero padding. Each selected user downloads the global model, performs image-to-text or text-to-image retrieval, conducts local multimodal training, and uploads model updates for aggregation. An energy–delay model couples local computation and wireless communication and treats the required number of global rounds as a function of retrieval ratios. Based on this model, an energy minimization problem is formulated and solved using a two-layer algorithm with an outer search over retrieval ratios and an inner optimization of transmission time, Central Processing Unit (CPU) frequency, and transmit power.  Results and Discussions  Simulations on a single-cell wireless MFL system show that increasing the ratio of completing text from images improves test accuracy and reduces total energy. In contrast, a large ratio of completing images from text provides limited accuracy gain but increases energy consumption (Fig. 3, Fig. 4). Compared with four representative baselines, CREEMFL achieves shorter completion time and lower total energy across a wide range of maximum average transmit powers (Fig. 5, Fig. 6). For CREEMFL, increased system bandwidth further reduces completion time and energy consumption (Fig. 7, Fig. 8). Under different user modality compositions, CREEMFL also attains higher test accuracy than local training, zero padding, and cross-modal retrieval without energy optimization (Fig. 9).  Conclusions  CREEMFL integrates selective cross-modal retrieval and joint communication–computation optimization for energy-efficient MFL. By treating retrieval ratios as variables and modeling their effect on global convergence rounds, it captures the coupling between per-round costs and global training progress. Simulations verify that CREEMFL reduces training completion time and total energy while preserving classification accuracy in resource-constrained wireless edge networks.
Breakthrough in Solving NP-Complete Problems Using Electronic Probe Computers
XU Jin, YU Le, YANG Huihui, JI Siyuan, ZHANG Yu, YANG Anqi, LI Quanyou, LI Haisheng, ZHU Enqiang, SHI Xiaolong, WU Pu, SHAO Zehui, LENG Huang, LIU Xiaoqing
Available online  , doi: 10.11999/JEIT250352
Abstract:
This study presents a breakthrough in addressing NP-complete problems using a newly developed Electronic Probe Computer (EPC60). The system employs a hybrid serial–parallel computational model and performs large-scale parallel operations through seven probe operators. In benchmark tests on 3-coloring problems in graphs with 2,000 vertices, EPC60 achieves 100% accuracy, outperforming the mainstream solver Gurobi, which succeeds in only 6% of cases. Computation time is reduced from 15 days to 54 seconds. The system demonstrates high scalability and offers a general-purpose solution for complex optimization problems in areas such as supply chain management, finance, and telecommunications.  Objective   NP-complete problems pose a fundamental challenge in computer science. As problem size increases, the required computational effort grows exponentially, making it infeasible for traditional electronic computers to provide timely solutions. Alternative computational models have been proposed, with biological approaches—particularly DNA computing—demonstrating notable theoretical advances. However, DNA computing systems continue to face major limitations in practical implementation.  Methods  Computational Model: EPC is based on a non-Turing computational model in which data are multidimensional and processed in parallel. Its database comprises four types of graphs, and the probe library includes seven operators, each designed for specific graph operations. By executing parallel probe operations, EPC efficiently addresses NP-complete problems.Structural Features:EPC consists of four subsystems: a conversion system, input system, computation system, and output system. The conversion system transforms the target problem into a graph coloring problem; the input system allocates tasks to the computation system; the computation system performs parallel operations via probe computation cards; and the output system maps the solution back to the original problem format.EPC60 features a three-tier hierarchical hardware architecture comprising a control layer, optical routing layer, and probe computation layer. The control layer manages data conversion, format transformation, and task scheduling. The optical routing layer supports high-throughput data transmission, while the probe computation layer conducts large-scale parallel operations using probe computation cards.  Results and Discussions  EPC60 successfully solved 100 instances of the 3-coloring problem for graphs with 2,000 vertices, achieving a 100% success rate. In comparison, the mainstream solver Gurobi succeeded in only 6% of cases. Additionally, EPC60 rapidly solved two 3-coloring problems for graphs with 1,500 and 2,000 vertices, which Gurobi failed to resolve after 15 days of continuous computation on a high-performance workstation.Using an open-source dataset, we identified 1,000 3-colorable graphs with 1,000 vertices and 100 3-colorable graphs with 2,000 vertices. These correspond to theoretical complexities of O(1.3289n) for both cases. The test results are summarized in Table 1.Currently, EPC60 can directly solve 3-coloring problems for graphs with up to n vertices, with theoretical complexity of at least O(1.3289n).On April 15, 2023, a scientific and technological achievement appraisal meeting organized by the Chinese Institute of Electronics was held at Beijing Technology and Business University. A panel of ten senior experts conducted a comprehensive technical evaluation and Q&A session. The committee reached the following unanimous conclusions:1. The probe computer represents an original breakthrough in computational models.2. The system architecture design demonstrates significant innovation.3. The technical complexity reaches internationally leading levels.4. It provides a novel approach to solving NP-complete problems.Experts at the appraisal meeting stated, “This is a major breakthrough in computational science achieved by our country, with not only theoretical value but also broad application prospects.” In cybersecurity, EPC60 has also demonstrated remarkable potential. Supported by the National Key R&D Program of China (2019YFA0706400), Professor Xu Jin’s team developed an automated binary vulnerability mining system based on a function call graph model. Evaluation of the system using the Modbus Slave software showed over 95% vulnerability coverage, far exceeding the 75 vulnerabilities detected by conventional depth-first search algorithms. The system also discovered a previously unknown flaw, the “Unauthorized Access Vulnerability in Changyuan Shenrui PRS-7910 Data Gateway” (CNVD-2020-31406), highlighting EPC60’s efficacy in cybersecurity applications.The high efficiency of EPC60 derives from its unique computational model and hardware architecture. Given that all NP-complete problems can be polynomially reduced to one another, EPC60 provides a general-purpose solution framework. It is therefore expected to be applicable in a wide range of domains, including supply chain management, financial services, telecommunications, energy, and manufacturing.  Conclusions   The successful development of EPC offers a novel approach to solving NP-complete problems. As technological capabilities continue to evolve, EPC is expected to demonstrate strong computational performance across a broader range of application domains. Its distinctive computational model and hardware architecture also provide important insights for the design of next-generation computing systems.
Personalized Federated Learning Method Based on Collation Game and Knowledge Distillation
SUN Yanhua, SHI Yahui, LI Meng, YANG Ruizhe, SI Pengbo
Available online  , doi: 10.11999/JEIT221203
Abstract:
To overcome the limitation of the Federated Learning (FL) when the data and model of each client are all heterogenous and improve the accuracy, a personalized Federated learning algorithm with Collation game and Knowledge distillation (pFedCK) is proposed. Firstly, each client uploads its soft-predict on public dataset and download the most correlative of the k soft-predict. Then, this method apply the shapley value from collation game to measure the multi-wise influences among clients and quantify their marginal contribution to others on personalized learning performance. Lastly, each client identify it’s optimal coalition and then distill the knowledge to local model and train on private dataset. The results show that compared with the state-of-the-art algorithm, this approach can achieve superior personalized accuracy and can improve by about 10%.
The Range-angle Estimation of Target Based on Time-invariant and Spot Beam Optimization
Wei CHU, Yunqing LIU, Wenyug LIU, Xiaolong LI
Available online  , doi: 10.11999/JEIT210265
Abstract:
The application of Frequency Diverse Array and Multiple Input Multiple Output (FDA-MIMO) radar to achieve range-angle estimation of target has attracted more and more attention. The FDA can simultaneously obtain the degree of freedom of transmitting beam pattern in angle and range. However, its performance is degraded due to the periodicity and time-varying of the beam pattern. Therefore, an improved Estimating Signal Parameter via Rotational Invariance Techniques (ESPRIT) algorithm to estimate the target’s parameters based on a new waveform synthesis model of the Time Modulation and Range Compensation FDA-MIMO (TMRC-FDA-MIMO) radar is proposed. Finally, the proposed method is compared with identical frequency increment FDA-MIMO radar system, logarithmically increased frequency offset FDA-MIMO radar system and MUltiple SIgnal Classification (MUSIC) algorithm through the Cramer Rao lower bound and root mean square error of range and angle estimation, and the excellent performance of the proposed method is verified.
Satellite Navigation
Research on GRI Combination Design of eLORAN System
LIU Shiyao, ZHANG Shougang, HUA Yu
Available online  , doi: 10.11999/JEIT201066
Abstract:
To solve the problem of Group Repetition Interval (GRI) selection in the construction of the enhanced LORAN (eLORAN) system supplementary transmission station, a screening algorithm based on cross interference rate is proposed mainly from the mathematical point of view. Firstly, this method considers the requirement of second information, and on this basis, conducts a first screening by comparing the mutual Cross Rate Interference (CRI) with the adjacent Loran-C stations in the neighboring countries. Secondly, a second screening is conducted through permutation and pairwise comparison. Finally, the optimal GRI combination scheme is given by considering the requirements of data rate and system specification. Then, in view of the high-precision timing requirements for the new eLORAN system, an optimized selection is made in multiple optimal combinations. The analysis results show that the average interference rate of the optimal combination scheme obtained by this algorithm is comparable to that between the current navigation chains and can take into account the timing requirements, which can provide referential suggestions and theoretical basis for the construction of high-precision ground-based timing system.