pureboot: the activation window gets a behavioral gate, and honest per-poll constants under it

The window's per-poll cycle counts were hand-counted for a uint32_t
countdown, but every default window fits uint24_t, whose decrement chain
is one sbci shorter — so deployed loaders ran 9/10ths of their stated
seconds (a 328P's 8 s was 7.2 s on the wire). No golden-asm pin can hold
this: the loops compile in consumer context. pbwindow.py measures the
behavior instead: it installs a real application beside the loader
through the host tool's own plan_flash (surgery included), starts the
simulator with the line idle, and reads the cycle of the first transmit
— the application's banner, so that cycle is the window. Held at plus or
minus 2 percent per chip (pureboot.window), red at -10.0 percent against
the old constants, green with poll_cycles now counted for the narrow
countdown (hardware 9, software 7; window_polls() solves narrow-first
and adds the wide loop's cycle where the count forces uint32_t — a count
narrow only at the wide cost stays wide, so the choice cannot
oscillate). The autobaud window is its poll budget at the measured ten
cycles a poll, gated the same way (pureboot.window.autobaud), and the
README carries that arithmetic now. No version bump: timing-window
precision is not meaningful behavior, v7 stays.

The gate flushed out two runner gaps. The software bridge accepted any
falling edge as a start bit, so the device's own TX-init glitch decoded
as a stray byte; it re-samples mid-bit now and abandons a false start,
as silicon does. And after avr_reset, the idle-line re-raise was
silently dropped: ioport pin irqs are IRQ_FLAG_FILTERED and the irq's
cached value survives the reset the port latch does not, so the device
read the line stuck low, calibrate() measured reset-to-first-edge as one
wrapping pulse, and the first knock after a reset could boot the
application instead of locking — the intermittent autobaud failure.
bridge_reset forces a real transition (0 then 1, no cycles between).

The README's Autobaud column now carries each chip's worst
configuration — autobaud with OSCCAL baked, on a USART's own pins where
the chip has one (tinies: autobaud + OSCCAL) — the numbers the existing
pureboot_autobaud_osccal[_on_usart0] matrix points already gate;
sizes.py checks the column against exactly those targets. Tool sizes
and window prose updated with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 16:05:42 +02:00
parent c8ac61779e
commit 8e7cc86fb3
7 changed files with 297 additions and 50 deletions

View File

@@ -19,42 +19,42 @@ come out byte-identical linked at a different base.
## Chips
Sizes are the default configuration: the hardware USART0 at 115200 8N1 on a
16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at 57600 8N1 on
the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above). Every axis moves
per build — see *Configuration*. The autobaud column is the clock-free build,
which is the largest the space produces and the tightest fit in the matrix;
it carries the calibration machinery and no clock at all.
The Stock column is the default configuration: the hardware USART0 at 115200
8N1 on a 16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at
57600 8N1 on the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above).
Every axis moves per build — see *Configuration*. The Autobaud column is the
worst configuration the space produces for the chip: the clock-free build —
it alone carries the calibration machinery — with the `OSCCAL` trim baked
and, where the chip has a USART, the link deployed on that USART's own pins,
which the loader then has to release (*Pin ownership*). On default pins
without the trim the same loaders run 1030 B smaller.
| Chip | Flash | Loader at | Link | Stock | Autobaud |
|---|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 384 B | 452 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 388 B | 442 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 388 B | 442 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 388 B | 442 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 358 B | 470 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 360 B | 474 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 360 B | 474 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 378 B | 438 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 388 B | 448 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 390 B | 454 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 390 B | 454 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 390 B | 454 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 390 B | 454 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 384 B | 448 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 410 B | 474 B |
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 384 B | 456 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 388 B | 446 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 388 B | 446 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 388 B | 446 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 358 B | 476 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 360 B | 480 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 360 B | 480 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 378 B | 450 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 388 B | 460 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 390 B | 464 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 390 B | 464 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 390 B | 464 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 390 B | 464 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 384 B | 458 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 410 B | 484 B |
† No hardware boot section: the host patches the reset vector, and the budget
is 510 bytes, since the slot's last word is the trampoline.
The tightest fit in the whole space is the 1284s' autobaud build deployed on a
USART's own pins with the `OSCCAL` trim baked, 484 of its 512 — they alone
carry the far-flash machinery (ELPM reads, RAMPZ page commands), autobaud
alone carries the calibration loop, a bit-banged link on a USART's pins alone
has to release it (below), and the trim adds its one register write. Without
the trim that build is 478; on the default pins, 474. The flash bank riding
in a transfer's selector byte keeps even those chips' addressing the same
16-bit form every other chip uses, which is why they are no longer the
The tightest fit in the whole space is therefore the 1284s' 484 of their
512: they alone carry the far-flash machinery (ELPM reads, RAMPZ page
commands) on top of everything the column already stacks. The flash bank
riding in a transfer's selector byte keeps even those chips' addressing the
same 16-bit form every other chip uses, which is why they are no longer the
outlier they were.
The software UART enables the RX pull-up; TX idles high. All multi-byte wire
@@ -100,7 +100,10 @@ where a fixed-baud software build has to be rebuilt per clock and still drifts
out of tolerance. The cost is that it is software-serial only (a hardware USART
needs its divisor programmed) and that activation counts poll iterations rather
than seconds, since there is no clock to convert them against
(`PUREBOOT_AUTOBAUD_POLLS`, default 4,000,000).
(`PUREBOOT_AUTOBAUD_POLLS`, default 4,000,000). The wait spends ten cycles a
poll (measured, and held by the `pureboot.window.autobaud` gate), so the
default window is 40 M cycles: 5 s at 8 MHz, about 4.2 s at 9.6 MHz, 40 s at
1 MHz.
**Pick the rate by cycles a bit, and leave the oscillator room.** What the
calibration can measure is bounded by how many clock cycles one bit lasts, so a

View File

@@ -171,9 +171,11 @@ template <avr::hertz_t C, avr::baud_t B>
struct hardware_link {
using uart = avr::uart::usart<usart_unit, C, {.baud = B, .max_baud_error = 2.5_pct}>;
// The compiled idle poll: lds UCSR0A (2), sbrc skipping the exit (2),
// sbiw + sbci + sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 10;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// lds UCSR0A (2), sbrc skipping the exit (2), sbiw + sbci + brne (5).
// A uint32_t countdown pays one more sbci — window_polls() adds it where
// the count forces the wide type. Held by the pureboot.window gate.
static constexpr std::uint8_t poll_cycles = 9;
static void init()
{
@@ -206,9 +208,11 @@ struct software_link {
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, B>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, B>;
// The compiled idle poll: sbis skipping the exit (2), sbiw + sbci +
// sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 8;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// sbis skipping the exit (2), sbiw + sbci + brne (5). A uint32_t
// countdown pays one more sbci — window_polls() adds it where the count
// forces the wide type. Held by the pureboot.window gate.
static constexpr std::uint8_t poll_cycles = 7;
static void init()
{
@@ -317,17 +321,32 @@ void await_host()
}
}
#else
// The window as one 32-bit countdown, divided by the backend's counted
// poll-loop cycles. Whole seconds is all it promises.
// The window as one countdown, divided by the backend's counted poll-loop
// cycles. Whole seconds is all it promises. The per-poll cost depends on the
// countdown's own width (a uint32_t decrement chain is one sbci longer), and
// the width depends on the poll count — solved narrow-first: a count that
// fits 24 bits at the narrow cost keeps the narrow loop, anything else takes
// the wide loop at its own cost. A count fitting 24 bits only at the wide
// cost stays wide, so the choice cannot oscillate on the boundary.
consteval std::uint32_t polls_at(std::uint32_t per_poll)
{
return timeout_seconds * (dev::clock.hz / per_poll);
}
consteval bool narrow_window()
{
return polls_at(link::poll_cycles) <= 0xffffff;
}
consteval std::uint32_t window_polls()
{
return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles);
return polls_at(narrow_window() ? link::poll_cycles : link::poll_cycles + 1u);
}
// The countdown in the narrowest type that holds it: a fourth byte would
// cost a wider decrement chain at every poll for range most windows never
// use (the autobaud budget makes the same choice).
using window_t = std::conditional_t<window_polls() <= 0xffffff, avr::uint24_t, std::uint32_t>;
using window_t = std::conditional_t<narrow_window(), avr::uint24_t, std::uint32_t>;
bool pending_before_deadline()
{