SKANN: a raw-waveform selective-kernel acoustic neural network
Conventional pipelines analyse spectrograms computed with a single, hand-set analysis window, yet no single window can resolve both slow machinery tonals and millisecond transients. SKANN replaces that fixed front end with a bank of dense learned 1-D filters at several kernel lengths in parallel. Long kernels capture the slow temporal structure of vessel-noise envelopes and the tonal band below 2 kHz; short kernels capture sharp transients such as clicks and snapping shrimp. Channel-wise selective-kernel attention learns which scales to trust for each input segment, and a compact 2-D convolutional backbone distils the fused feature map into a 512-dimensional embedding. No frequency band is hand-assigned; the network learns where the information lives.
Two variants are in use. The classification encoder shown in the capability overview (HybridSKEncoderV3) has eight branches spanning roughly 0.5 ms to 64 ms and a calibrated seven-class head; it is the classifier the PS12 pipeline runs on. The identity encoder described in the 2026 preprint has four kernel lengths from about 16 ms to 1 s and is trained with an additive angular-margin (ArcFace) objective, producing an L2-normalised hull embedding that is compared by cosine similarity.
An Indian provisional patent application (202611107132, filed 6 September 2026, applicant Oravont Systems LLP) covers aspects of the method. See Research.
Condition invariance as a property of the data
The ocean, the range and the recording chain all colour a signature. SKANN treats invariance to those conditions as a property of the training data rather than of the loss: an augmentation regime randomises recording-chain colouration, ambient noise and multipath while provably preserving the narrowband line structure that carries identity. Transforms that would move absolute line frequencies — the shaft and blade lines a stealth designer would recognise at once — are deliberately excluded, because they would erase the very evidence the model is meant to key on.
Physics-grounded synthesis where labels run out
Real underwater recordings are voluminous, unlabelled and expensive to annotate. Oravont's synthetic underwater acoustic dataset fills the gap for pretraining: 12,000 five-second clips at 16 kHz across four vessel classes — tanker, cargo ship, fishing vessel and small craft — and ambient ocean noise, generated from first-principle models with full-factorial parameter coverage. Sea noise follows piecewise Knudsen spectra by sea state; ship noise combines shaft-rate and blade-pass harmonics, generator and equipment lines, structural resonances, broadband flow noise and a physical cavitation-burst model. Every quantity is kept in physical units and verified against classical references, and the vessel classes use non-overlapping shaft-rate ranges — a decision forced by an earlier failure in downstream representation learning.
Used to pretrain a self-supervised encoder (Barlow Twins), the dataset produced an embedding space that separated its classes without a single label: a silhouette score of 0.97 and 100% k-NN accuracy across five classes, as reported in the dataset article. The dataset is released under CC BY 4.0.
GitHub: Underwater-Acoustic-Synthetic-Dataset · Read the article · Tools & data
A protocol that closes the easy routes
Underwater acoustic target recognition has largely settled on closed-set classification by vessel type, which does not answer whether a monitoring system has heard a particular hull before. The 2026 preprint formalises that question as open-set, cross-passage re-identification on public hydrophone data and specifies an evaluation protocol that removes the two easiest routes to a high score: hull-disjoint splits keyed to MMSI/IMO, galleries and queries drawn from disjoint passages of each hull, source-pure galleries, and an audio-adjudicated transit-deduplication gate.
On a 40-hull public gallery (96 queries against 98 passage candidates), cross-passage rank-1 is 0.25 for the SKANN embedding and 0.26 for a fully automated narrowband-tonal comparator — far below what closed-set numbers suggest, and statistically indistinguishable at the top of the ranking. The embedding orders the remainder of the list more reliably (AUC 0.82 against 0.76), and score fusion of the two reaches rank-1 0.35, the only contrast that attains nominal significance. Two further findings delimit what public data can support: ShipsEar cannot separate hull identity from recording channel under an identity protocol, and cross-network fine-tuning lifts performance on vessels seen during fine-tuning but is a null result on unseen ones.
Read together, these numbers support analyst triage over a ranked shortlist. They do not support treating either system as an identifier of record. The checkpoint, validation embeddings, transit map and per-query outputs are released under CC BY 4.0 so the numbers can be checked: doi:10.5281/zenodo.22160138.
Propagation-robust embeddings
A research collaboration in formation with IIT Delhi, in association with a DRDO/NSTL programme, extends re-identification toward propagation robustness by combining deep metric learning with physics-based acoustic-propagation modelling: the same hull, heard through different ranges, depths and sea states, should map to the same fingerprint.
Exploratory; nothing about the collaboration is finalised.
The other two tracks
Where the models go to work
Seven-class classification under PS12, open-set re-identification under DISC5, and DEMON running alongside as the transparent second opinion.
Open the track → 03 · StealthThe physics behind the training data
The acoustic-modelling know-how in the synthetic dataset and the augmentation regime — how machinery, propellers and hulls make noise — comes from stealth design.
Open the track →


