Recover a guessed stream's channels + bit depth by correlation too (step b)
Extends the two-path correlation from rate-only to the full layout, removing the "channels/bit-depth assumed = device" limitation. correlate_format tries each candidate de-interleaving (float32 / int16; mono..7.1) of the hook capture, runs the rate correlation per layout, and keeps whichever aligns with the loopback; a wrong de-interleaving is noise and won't. The catch: the hook can't know a guessed stream's real frame size, so its verify tap pads each render buffer to the device block -- which over-reads stale staging bytes for a stream with fewer channels/bits, scrambling the audio. So the tap is now self-describing: it prefixes each buffer with its frame count ([count][count*device_block bytes]), and the host strips the padding per candidate layout (take the real count*real_block of each chunk) before de-interleaving. - audio_correlate.hpp: ChunkedCapture + chunk-aware correlate_format + candidate layouts; absolute-margin confidence gate (the true layout scores ~1.0, a truly ambiguous alternative within ~0.001 -- 2ch@R == 1ch@2R for identical channels -- is correctly left unconfident). - audio_hook.cpp: chunked verify tap (free-space-checked so framing can't tear). - audio_format_verifier: parse chunks; recover_layout path. AudioMirror now corrects the full format. - audio_correlation_test: layout recovery from padded chunks (stereo float, 16-bit PCM, 5.1, mono). audio_verify_test gains scenario (b): 2ch on a multichannel endpoint with distinct per-channel content (new env-gated ToneSource mode) -> recovers ch=2/32-bit float end-to-end. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -410,8 +410,19 @@ HRESULT STDMETHODCALLTYPE hk_ReleaseBuffer(IAudioRenderClient* self, UINT32 num_
|
||||
const std::uint32_t block = g_streams[i].block_align.load(std::memory_order_relaxed);
|
||||
if (block != 0)
|
||||
{
|
||||
audio_ring_push(*vring, t_gb_data, readable_bytes(t_gb_data, num_frames * block),
|
||||
num_frames);
|
||||
// Self-describing chunk: [u32 frame-count][num_frames*block bytes]. The host can't
|
||||
// know the real frame size of a guessed stream, so it recovers the layout by
|
||||
// trying candidate de-interleavings -- but it needs the frame count to strip the
|
||||
// per-buffer padding (the guessed/device block over-reads a stream with fewer
|
||||
// channels/bits). Push both parts only if both fit and the payload is fully
|
||||
// readable, so a full ring or a short buffer can never tear the framing.
|
||||
const std::uint32_t want = num_frames * block;
|
||||
if (readable_bytes(t_gb_data, want) == want &&
|
||||
audio_ring_free_space(*vring) >= static_cast<std::uint32_t>(sizeof(num_frames)) + want)
|
||||
{
|
||||
audio_ring_push(*vring, &num_frames, sizeof(num_frames), 0);
|
||||
audio_ring_push(*vring, t_gb_data, want, num_frames);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user