Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<title>Abstract</title> <p>Voice-cloning attacks increasingly target the telephone channel, where a codec has compressed the audio and a live call must be judged from however much has arrived. We wrap AASIST, reused without retraining, to consume audio as a growing prefix, emit a per-duration calibrated confidence, and commit only when confident, deferring otherwise. All audio is ASVspoof 2019 LA passed through four real wideband codecs — Opus, Speex-WB, EVS, AMR-WB — with packet loss, not real telephone calls; the detector is frozen, so what we report is a decision layer and a channel characterization, not an accuracy result. The channel dominates: the same detector's equal-error rate (EER) spans 0.8% to 29% across codecs and bitrates, so a robustness claim means little without naming the channel. Recast as a sequential reject option on a matched risk–coverage curve, the streaming layer on a harsh channel (Speex-WB q4 + 20% loss) decides ~53% of calls at 2.9% committed-call error, against ~13% for the full-utterance detector at the same coverage and ~24% if it must decide every call. This is selection, not accuracy: a max-confidence selector credits ~83% of the apparent advantage to short early prefixes repeated to fill the model's fixed input window (tile-padding), and restricting it to prefixes ≥1.5 s collapses that share to ~13% — a collapse that replicates on two further harsh channels and a weaker raw-waveform backbone (RawNet2). The gain is bought by abstaining on ~47% of calls and rests on calibration that does not survive a channel mismatch. Calibration transfer is the actionable result: a clean-audio calibrator yields 3–3.5× the matched error on harsh low-bitrate channels, and asymmetrically — overconfidence commits wrongly, over-caution only defers — so calibration should be fit on a representative harsh condition. The channel-independent gain is latency, roughly halved (~1.6 s forced-inclusive vs. the 3.26 s grand mean).</p>