SPIRE
Rethinking IRSTD: Single-Point Supervision Guided Encoder-Only Framework Is Enough for Infrared Small Target Detection
重新思考红外小目标检测:单点监督引导的纯编码器框架足以实现红外小目标检测
College of Electronic Science and Technology
National University of Defense Technology · Wei An Research Group
TL;DR: SPIRE turns centroid points into physically grounded probabilistic targets, regresses them with a lightweight high-resolution encoder, and localizes targets by peak extraction.
Overview
Infrared small target detection is commonly framed as dense mask reconstruction, with a decoder progressively restoring spatial resolution. SPIRE revisits that choice: when the operational goal is reliable centroid localization, supervision, representation, and output can all be centered on a probabilistic response instead of a reconstructed contour.
The method follows three steps: construct a structured response target from each centroid annotation, regress the response map on a fixed high-resolution branch, and recover target coordinates through local-maximum detection and sub-pixel refinement. The paper figure above opens at full resolution.
Construct a compact, center-enhanced regression target from each annotated centroid.
Regress the response on a fixed high-resolution branch without a decoder or skip connections.
Read target centroids directly through local maxima and sub-pixel refinement.
How PRPS Constructs the Regression Target
PRPS uses the normalized local radiometric profile Ck to modulate an unchanged Gaussian point-response prior Gk. Panel I uses an enlarged three-dimensional UAV-like surface to make this weighting visually explicit; during training, Ck is computed from the original infrared image patch.
A magnified three-dimensional surface of local target radiometry—not a mask or measured network response.
Ck
The unchanged unit-peak isotropic Gaussian is centered at the refined radiometric peak.
Gk
The weighted response retains local radiometric asymmetry while concentrating near the center.
Hk = Norm(Gk ⊙ Ck)
Schematic visualization of the paper-defined construction. In training, Ck comes from the min-max normalized local image patch; σ=2 and r=6 follow the paper. No target contour is reconstructed.
View figure provenanceQualitative Results Visualization
Explore verified, precomputed SPIRE outputs on three selected samples from the official SIRST-UAVB test split. All results use the released checkpoint and the same fixed evaluation settings.
Fixed settings: 640 × 640 · threshold 0.35 · δ = 5 px
Unseen real scenes
Real-World Video Demonstrations
Both videos use the SPIRE checkpoint trained on the SIRST-UAVB (2400:600) train/test split. Neither real-scene sequence was included in training. The successful detections provide qualitative evidence of cross-scene generalization beyond the benchmark; these clips are not part of the benchmark evaluation.
A small target is localized amid dense texture and a strong linear background feature.
A bright airborne target is localized against a high-contrast terrain background.
Red marker: SPIRE-predicted target location
Reported Results
Tables 1 and 2 reproduce all cross-method comparisons reported in the final paper. Table 1 follows the unified centroid protocol at δ = 5 pixels. Table 2 evaluates matching-threshold robustness on SIRST4 and reports throughput under the same input and hardware setting. In Table 1, bold and underlined values mark the best and second-best entries computed from the reported numbers.
Table 1 · Main benchmark
Comparison with representative methods
Pre, Rec (Pd), and F1 are reported in %. Fa is reported in units of 10−8. FLOPs and Params are measured at 640 × 640 input resolution.
| Method | Venue | SIRST-UAVB (2400:600) | SIRST4 (2285:1067) | FLOPs (G) ↓ | Params (M) ↓ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Pre ↑ | Rec (Pd) ↑ | F1 ↑ | Fa ↓ | Pre ↑ | Rec (Pd) ↑ | F1 ↑ | Fa ↓ | ||||
| ACM [11] | WACV '21 | 87.01 | 71.16 | 78.29 | 34.04 | 90.17 | 70.38 | 79.05 | 44.08 | 2.51 | 0.40 |
| ALCNet [12] | TGRS '21 | 95.06 | 81.11 | 87.53 | 12.72 | 91.45 | 72.41 | 80.82 | 38.90 | 2.36 | 0.43 |
| ISTDU-Net [16] | GRSL '22 | 79.28 | 78.08 | 78.67 | 61.54 | 93.28 | 72.03 | 81.29 | 29.83 | 49.65 | 2.75 |
| RDIAN [24] | TGRS '23 | 52.68 | 77.91 | 62.86 | 211.08 | 90.64 | 80.83 | 85.45 | 47.98 | 23.24 | 0.22 |
| DNANet [17] | TIP '23 | 94.48 | 89.54 | 91.95 | 15.77 | 93.99 | 81.20 | 87.13 | 29.82 | 89.13 | 4.70 |
| SCTransNet [37] | TGRS '24 | 98.27 | 95.95 | 97.09 | 5.09 | 81.20 | 86.39 | 83.72 | 114.98 | 63.22 | 11.19 |
| MSHNet [21] | CVPR '24 | 71.90 | 88.02 | 79.09 | 104.27 | 90.68 | 92.86 | 91.47 | 60.51 | 38.16 | 4.07 |
| SDSNet [38] | TGRS '25 | 97.60 | 96.29 | 96.94 | 7.12 | 87.35 | 88.27 | 87.81 | 73.48 | 42.42 | 2.49 |
| L2SKNet [28] | TGRS '25 | 95.83 | 89.04 | 92.31 | 11.70 | 92.89 | 91.42 | 92.16 | 40.20 | 43.09 | 0.90 |
| SPIRE (Ours) | — | 99.82 | 94.44 | 97.05 | 1.02 | 95.00 | 94.21 | 94.60 | 28.53 | 7.68 | 0.29 |
Table 2 · Robustness and speed
Multi-threshold comparison on SIRST4
Rec (Pd) and F1 are reported in %. Fa is reported in units of 10−8. Results at δ = 5, 8, and 10 are identical; FPS follows the paper's shared input and hardware setting.
| Method | δ = 3 px | δ = 5, 8, 10 px | FPS ↑ | ||||
|---|---|---|---|---|---|---|---|
| Rec (Pd) ↑ | F1 ↑ | Fa ↓ | Rec (Pd) ↑ | F1 ↑ | Fa ↓ | ||
| DNANet [17] | 80.52 | 86.41 | 33.72 | 81.20 | 87.13 | 29.82 | 80.86 |
| SCTransNet [37] | 80.15 | 81.18 | 99.42 | 86.39 | 83.72 | 114.98 | 67.09 |
| SDSNet [38] | 86.39 | 87.21 | 67.43 | 88.27 | 87.81 | 73.48 | 46.97 |
| SPIRE (Ours) | 92.93 | 94.74 | 18.61 | 94.21 | 94.60 | 28.53 | 261.2 |
All values are transcribed from the final paper. FLOPs and Params use 640 × 640 input resolution; Table 2 throughput is paper-reported and is not re-measured in the browser.
Citation
If this work supports your research, please consider citing the paper. See the Paper and Code.
@inproceedings{ni2026spire,
title={Rethinking IRSTD: Single-Point Supervision Guided
Encoder-Only Framework Is Enough for Infrared
Small Target Detection},
author={Ni, Rixiang and Chen, Jun and Li, Boyang and
Li, Yonghao and He, Wujiao and Wang, Yuji and
Ren, Feiyu and Yuan, Haoyang and An, Wei},
booktitle={European Conference on Computer Vision},
year={2026}
}