CRAW: Audio Samples

David Chernin & Ethan Fetaya, Bar-Ilan University. All clips are unmodified (no attacks applied) — for perceptual quality comparison.

Component Ablation

The same clip watermarked at each stage of CRAW's design, so you can hear the fidelity recovery directly against the pristine original: Q-Former pooling and inference-time masking recover most of the perceptual quality given up when the robust distortion layer (DL) is added, at a small further cost from the final error-correcting code (+ECC, full CRAW). Clips are randomly sampled from the held-out test set.

Clip Original DL +QFormer +Mask +ECC Full CRAW
Clip 1
0:00
0:00
0:00
0:00
0:00
Clip 2
0:00
0:00
0:00
0:00
0:00
Clip 3
0:00
0:00
0:00
0:00
0:00
Clip 4
0:00
0:00
0:00
0:00
0:00
Clip 5
0:00
0:00
0:00
0:00
0:00
Clip 6
0:00
0:00
0:00
0:00
0:00
Clip 7
0:00
0:00
0:00
0:00
0:00
Clip 8
0:00
0:00
0:00
0:00
0:00
Clip 9
0:00
0:00
0:00
0:00
0:00
Clip 10
0:00
0:00
0:00
0:00
0:00
Clip 11
0:00
0:00
0:00
0:00
0:00
Clip 12
0:00
0:00
0:00
0:00
0:00
Clip 13
0:00
0:00
0:00
0:00
0:00
Clip 14
0:00
0:00
0:00
0:00
0:00
Clip 15
0:00
0:00
0:00
0:00
0:00
Clip 16
0:00
0:00
0:00
0:00
0:00
Clip 17
0:00
0:00
0:00
0:00
0:00
Clip 18
0:00
0:00
0:00
0:00
0:00
Clip 19
0:00
0:00
0:00
0:00
0:00
Clip 20
0:00
0:00
0:00
0:00
0:00

vs. Baselines

CRAW (full model) against the external baselines: WavMark, AudioSeal, TimbreWM, AWARE, and WMCodec, plus the original unwatermarked clip for reference. Clips are randomly sampled from the held-out test set.

Clip Original CRAW Ours WavMark AudioSeal TimbreWM AWARE WMCodec
Clip 1
0:00
0:00
0:00
0:00
0:00
0:00
0:00
Clip 2
0:00
0:00
0:00
0:00
0:00
0:00
0:00
Clip 3
0:00
0:00
0:00
0:00
0:00
0:00
0:00
Clip 4
0:00
0:00
0:00
0:00
0:00
0:00
0:00
Clip 5
0:00
0:00
0:00
0:00
0:00
0:00
0:00
Clip 6
0:00
0:00
0:00
0:00
0:00
0:00
0:00