pureboot: the activation window gets a behavioral gate, and honest per-poll constants under it

The window's per-poll cycle counts were hand-counted for a uint32_t
countdown, but every default window fits uint24_t, whose decrement chain
is one sbci shorter — so deployed loaders ran 9/10ths of their stated
seconds (a 328P's 8 s was 7.2 s on the wire). No golden-asm pin can hold
this: the loops compile in consumer context. pbwindow.py measures the
behavior instead: it installs a real application beside the loader
through the host tool's own plan_flash (surgery included), starts the
simulator with the line idle, and reads the cycle of the first transmit
— the application's banner, so that cycle is the window. Held at plus or
minus 2 percent per chip (pureboot.window), red at -10.0 percent against
the old constants, green with poll_cycles now counted for the narrow
countdown (hardware 9, software 7; window_polls() solves narrow-first
and adds the wide loop's cycle where the count forces uint32_t — a count
narrow only at the wide cost stays wide, so the choice cannot
oscillate). The autobaud window is its poll budget at the measured ten
cycles a poll, gated the same way (pureboot.window.autobaud), and the
README carries that arithmetic now. No version bump: timing-window
precision is not meaningful behavior, v7 stays.

The gate flushed out two runner gaps. The software bridge accepted any
falling edge as a start bit, so the device's own TX-init glitch decoded
as a stray byte; it re-samples mid-bit now and abandons a false start,
as silicon does. And after avr_reset, the idle-line re-raise was
silently dropped: ioport pin irqs are IRQ_FLAG_FILTERED and the irq's
cached value survives the reset the port latch does not, so the device
read the line stuck low, calibrate() measured reset-to-first-edge as one
wrapping pulse, and the first knock after a reset could boot the
application instead of locking — the intermittent autobaud failure.
bridge_reset forces a real transition (0 then 1, no cycles between).

The README's Autobaud column now carries each chip's worst
configuration — autobaud with OSCCAL baked, on a USART's own pins where
the chip has one (tinies: autobaud + OSCCAL) — the numbers the existing
pureboot_autobaud_osccal[_on_usart0] matrix points already gate;
sizes.py checks the column against exactly those targets. Tool sizes
and window prose updated with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 16:05:42 +02:00
parent c8ac61779e
commit 8e7cc86fb3
7 changed files with 297 additions and 50 deletions

View File

@@ -171,9 +171,11 @@ template <avr::hertz_t C, avr::baud_t B>
struct hardware_link {
using uart = avr::uart::usart<usart_unit, C, {.baud = B, .max_baud_error = 2.5_pct}>;
// The compiled idle poll: lds UCSR0A (2), sbrc skipping the exit (2),
// sbiw + sbci + sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 10;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// lds UCSR0A (2), sbrc skipping the exit (2), sbiw + sbci + brne (5).
// A uint32_t countdown pays one more sbci — window_polls() adds it where
// the count forces the wide type. Held by the pureboot.window gate.
static constexpr std::uint8_t poll_cycles = 9;
static void init()
{
@@ -206,9 +208,11 @@ struct software_link {
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, B>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, B>;
// The compiled idle poll: sbis skipping the exit (2), sbiw + sbci +
// sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 8;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// sbis skipping the exit (2), sbiw + sbci + brne (5). A uint32_t
// countdown pays one more sbci — window_polls() adds it where the count
// forces the wide type. Held by the pureboot.window gate.
static constexpr std::uint8_t poll_cycles = 7;
static void init()
{
@@ -317,17 +321,32 @@ void await_host()
}
}
#else
// The window as one 32-bit countdown, divided by the backend's counted
// poll-loop cycles. Whole seconds is all it promises.
// The window as one countdown, divided by the backend's counted poll-loop
// cycles. Whole seconds is all it promises. The per-poll cost depends on the
// countdown's own width (a uint32_t decrement chain is one sbci longer), and
// the width depends on the poll count — solved narrow-first: a count that
// fits 24 bits at the narrow cost keeps the narrow loop, anything else takes
// the wide loop at its own cost. A count fitting 24 bits only at the wide
// cost stays wide, so the choice cannot oscillate on the boundary.
consteval std::uint32_t polls_at(std::uint32_t per_poll)
{
return timeout_seconds * (dev::clock.hz / per_poll);
}
consteval bool narrow_window()
{
return polls_at(link::poll_cycles) <= 0xffffff;
}
consteval std::uint32_t window_polls()
{
return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles);
return polls_at(narrow_window() ? link::poll_cycles : link::poll_cycles + 1u);
}
// The countdown in the narrowest type that holds it: a fourth byte would
// cost a wider decrement chain at every poll for range most windows never
// use (the autobaud budget makes the same choice).
using window_t = std::conditional_t<window_polls() <= 0xffffff, avr::uint24_t, std::uint32_t>;
using window_t = std::conditional_t<narrow_window(), avr::uint24_t, std::uint32_t>;
bool pending_before_deadline()
{