Oravont Systems LLP

Track 02 · Deep learning

The engine that reads acoustic signatures

Deep learning directly on raw hydrophone waveforms: learned multi-scale filterbanks instead of hand-set spectrograms, metric-learned embeddings instead of fixed class lists, physics-grounded synthetic data where labelled recordings run out, and evaluation protocols designed to be hard to flatter.

Preprint: arXiv:2609.07399
Architecture

SKANN: a raw-waveform selective-kernel acoustic neural network

Conventional pipelines analyse spectrograms computed with a single, hand-set analysis window, yet no single window can resolve both slow machinery tonals and millisecond transients. SKANN replaces that fixed front end with a bank of dense learned 1-D filters at several kernel lengths in parallel. Long kernels capture the slow temporal structure of vessel-noise envelopes and the tonal band below 2 kHz; short kernels capture sharp transients such as clicks and snapping shrimp. Channel-wise selective-kernel attention learns which scales to trust for each input segment, and a compact 2-D convolutional backbone distils the fused feature map into a 512-dimensional embedding. No frequency band is hand-assigned; the network learns where the information lives.

Two variants are in use. The classification encoder shown in the capability overview (HybridSKEncoderV3) has eight branches spanning roughly 0.5 ms to 64 ms and a calibrated seven-class head; it is the classifier the PS12 pipeline runs on. The identity encoder described in the 2026 preprint has four kernel lengths from about 16 ms to 1 s and is trained with an additive angular-margin (ArcFace) objective, producing an L2-normalised hull embedding that is compared by cosine similarity.

An Indian provisional patent application (202611107132, filed 6 September 2026, applicant Oravont Systems LLP) covers aspects of the method. See Research.

SKANN encoder: a raw 5-second waveform enters a four-branch selective-kernel filterbank spanning 16 ms to 1 s, fused by attention, through a five-stage convolutional backbone to a 512-dimensional hull fingerprint compared by cosine similarity The identity encoder of the 2026 preprint: a four-scale learned filterbank on the raw waveform (kernels of about 16 ms to 1 s), fused by selective-kernel attention, into a 512-dimensional hull embedding.
SKANN classification variant: eight parallel 1-D convolution branches with kernels from 0.5 ms to 64 ms, selective-kernel attention, a five-stage 2-D backbone and a seven-class head The classification variant (HybridSKEncoderV3): eight parallel kernels from 0.5 ms to 64 ms, selective-kernel attention across branches, a five-stage 2-D backbone and a calibrated seven-class head.
Invariance

Condition invariance as a property of the data

The ocean, the range and the recording chain all colour a signature. SKANN treats invariance to those conditions as a property of the training data rather than of the loss: an augmentation regime randomises recording-chain colouration, ambient noise and multipath while provably preserving the narrowband line structure that carries identity. Transforms that would move absolute line frequencies — the shaft and blade lines a stealth designer would recognise at once — are deliberately excluded, because they would erase the very evidence the model is meant to key on.

Training data

Physics-grounded synthesis where labels run out

Real underwater recordings are voluminous, unlabelled and expensive to annotate. Oravont's synthetic underwater acoustic dataset fills the gap for pretraining: 12,000 five-second clips at 16 kHz across four vessel classes — tanker, cargo ship, fishing vessel and small craft — and ambient ocean noise, generated from first-principle models with full-factorial parameter coverage. Sea noise follows piecewise Knudsen spectra by sea state; ship noise combines shaft-rate and blade-pass harmonics, generator and equipment lines, structural resonances, broadband flow noise and a physical cavitation-burst model. Every quantity is kept in physical units and verified against classical references, and the vessel classes use non-overlapping shaft-rate ranges — a decision forced by an earlier failure in downstream representation learning.

Used to pretrain a self-supervised encoder (Barlow Twins), the dataset produced an embedding space that separated its classes without a single label: a silhouette score of 0.97 and 100% k-NN accuracy across five classes, as reported in the dataset article. The dataset is released under CC BY 4.0.

GitHub: Underwater-Acoustic-Synthetic-Dataset · Read the article · Tools & data

Synthetic Underwater Acoustic Dataset: physics-grounded, 12,000 clips, self-supervised learning — Oravont Systems LLP 12,000 physics-grounded clips, four vessel classes plus ambient noise, released under CC BY 4.0.
Evaluation

A protocol that closes the easy routes

Underwater acoustic target recognition has largely settled on closed-set classification by vessel type, which does not answer whether a monitoring system has heard a particular hull before. The 2026 preprint formalises that question as open-set, cross-passage re-identification on public hydrophone data and specifies an evaluation protocol that removes the two easiest routes to a high score: hull-disjoint splits keyed to MMSI/IMO, galleries and queries drawn from disjoint passages of each hull, source-pure galleries, and an audio-adjudicated transit-deduplication gate.

On a 40-hull public gallery (96 queries against 98 passage candidates), cross-passage rank-1 is 0.25 for the SKANN embedding and 0.26 for a fully automated narrowband-tonal comparator — far below what closed-set numbers suggest, and statistically indistinguishable at the top of the ranking. The embedding orders the remainder of the list more reliably (AUC 0.82 against 0.76), and score fusion of the two reaches rank-1 0.35, the only contrast that attains nominal significance. Two further findings delimit what public data can support: ShipsEar cannot separate hull identity from recording channel under an identity protocol, and cross-network fine-tuning lifts performance on vessels seen during fine-tuning but is a null result on unseen ones.

Read together, these numbers support analyst triage over a ranked shortlist. They do not support treating either system as an identifier of record. The checkpoint, validation embeddings, transit map and per-query outputs are released under CC BY 4.0 so the numbers can be checked: doi:10.5281/zenodo.22160138.

Exploratory

Propagation-robust embeddings

A research collaboration in formation with IIT Delhi, in association with a DRDO/NSTL programme, extends re-identification toward propagation robustness by combining deep metric learning with physics-based acoustic-propagation modelling: the same hull, heard through different ranges, depths and sea states, should map to the same fingerprint.

Exploratory; nothing about the collaboration is finalised.

Concept illustration: propagation-robust acoustic vessel re-identification Concept illustration: the same hull heard through different propagation conditions should map to the same fingerprint.
How this connects

The other two tracks