HALF_DUPLEX deploys a shared line per backend. The hardware USART takes
the library's .half_duplex turn-around — RXD and TXD tied off-chip, each
reply byte held to transmit-complete before the line can be released
(m8 404 B, m328P 440, 1284P 460; the window poll runs through the
outlined release-line call at 18 or 22 cycles a poll, measured off the
built loops and held per chip by pureboot.window.halfduplex). The
software and autobaud links fold onto the RX pin — RX == TX spells the
same — and cost nothing: the frame's direction wrap is what the dropped
second-pin init paid, and the worst image in the space is unchanged at
the 1284s' 502 of 512, now with its one-wire twin proven equal across
the exhaustive matrix. The host gains --one-wire, the echo discard a
shared line requires: the adapter's echo is matched byte for byte and a
reply interleaving a blind write — a loader already in session
re-prompts inside the knock — is held for the reader. The device runner
models the shared line by direction (drives only while the firmware's
DDR reads input, decodes only while the firmware owns it, supplies the
host-side echo), extends the USART pin-ownership model to RXEN's hold
on RXD, and starts the pty USART from the datasheet's zeroed UCSR#B:
simavr's TXEN-set reset plus its clear-UDRE-on-TXEN-drop otherwise
wedges the first transmitter after a receiver-only program, which the
half-duplex window gate caught as a banner that never came. v7 is
tagged at its era's last commit; v8 changes nothing on the wire.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
libavr's calibrate() now bounds its measurement loop, starts the pulse
on an observed edge, and re-arms a rejected pulse on the remaining
budget instead of one-strike booting the application. The autobaud
images pay +16..20 B — every slot still fits, the worst now the 1284s'
502 of 512 — and the stock images are byte-identical, kept so by
fitting the loader's flag set per backend: -fno-ivopts stays on the
fixed-baud bodies it shrinks and comes off the autobaud body, where it
duplicated the calibration countdown into a 9-cycle loop against the
contracted seven.
One deployed constant moved and its gate caught it: the calibrate
wait's budget poll re-laid from ten cycles to nine (the exit branches
land where block layout puts them), so pureboot.window.autobaud
measured -10 % until AUTOBAUD_POLL_CYCLES and the README's derived
seconds were re-measured — the default autobaud window is 36 M cycles,
4.5 s at 8 MHz. Full gate green on all 37 chips; the README's autobaud
column carries each chip's rebuilt worst configuration, machine-checked
against the built trees.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The hand-over drains move to the explicit drain_unbounded() — both link
adapters drain only after their own write, so the frame is in flight by
construction and the bounded default's countdown would be dead bytes;
the images stay byte-identical. window_polls() states its arithmetic
through dev::cycles_for with the whole window converted before the
per-poll division — one truncation instead of one per second, same
instructions, only the countdown's immediate moves. Every size in the
matrix is unchanged; the full gate is green on all 37 chips.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
sizes.py merged every owned tree's rows and let the last one win, so a
stale reflect tree — last built before the window constants moved —
reported the atmega8's old stock size over the fresh build and failed
the README check with yesterday's number. Generated and reflect must
answer with the same bytes (the identity invariant), so the same target
measuring two sizes is a stale tree or an identity breach; collect()
refuses now, naming both trees. The stale reflect trees are removed —
the reflect sweep rebuilds them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
usart_of<0> is an incomplete type on the USART-less chips, and a static
member initializer with only non-dependent operands is checked when the
template is parsed, not when it is instantiated — so the address probe
broke every tiny build without hardware_link ever being named. The
lookup moves into a member function template taking the link's own baud
parameter, the dependence carrier that defers it to instantiation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The window gate's first full sweep caught it on the atmega8: 6.22 s
measured against 8 declared, the exact 7/9 of a poll modeled as an
extended-I/O lds + skip on a chip whose UCSRA sits in bit-addressable
I/O and compiles to a 2-cycle skip. poll_cycles now follows the status
register's home (7 below 0x40, 9 above). At 16 MHz over 7 cycles the
poll count no longer fits uint24_t, so the classic megas' stock windows
take the wide countdown — 8.000 s measured on all three, +4 B of stock
image (m8 362, m16/m32 364), README stock rows updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The window's per-poll cycle counts were hand-counted for a uint32_t
countdown, but every default window fits uint24_t, whose decrement chain
is one sbci shorter — so deployed loaders ran 9/10ths of their stated
seconds (a 328P's 8 s was 7.2 s on the wire). No golden-asm pin can hold
this: the loops compile in consumer context. pbwindow.py measures the
behavior instead: it installs a real application beside the loader
through the host tool's own plan_flash (surgery included), starts the
simulator with the line idle, and reads the cycle of the first transmit
— the application's banner, so that cycle is the window. Held at plus or
minus 2 percent per chip (pureboot.window), red at -10.0 percent against
the old constants, green with poll_cycles now counted for the narrow
countdown (hardware 9, software 7; window_polls() solves narrow-first
and adds the wide loop's cycle where the count forces uint32_t — a count
narrow only at the wide cost stays wide, so the choice cannot
oscillate). The autobaud window is its poll budget at the measured ten
cycles a poll, gated the same way (pureboot.window.autobaud), and the
README carries that arithmetic now. No version bump: timing-window
precision is not meaningful behavior, v7 stays.
The gate flushed out two runner gaps. The software bridge accepted any
falling edge as a start bit, so the device's own TX-init glitch decoded
as a stray byte; it re-samples mid-bit now and abandons a false start,
as silicon does. And after avr_reset, the idle-line re-raise was
silently dropped: ioport pin irqs are IRQ_FLAG_FILTERED and the irq's
cached value survives the reset the port latch does not, so the device
read the line stuck low, calibrate() measured reset-to-first-edge as one
wrapping pulse, and the first knock after a reset could boot the
application instead of locking — the intermittent autobaud failure.
bridge_reset forces a real transition (0 then 1, no cycles between).
The README's Autobaud column now carries each chip's worst
configuration — autobaud with OSCCAL baked, on a USART's own pins where
the chip has one (tinies: autobaud + OSCCAL) — the numbers the existing
pureboot_autobaud_osccal[_on_usart0] matrix points already gate;
sizes.py checks the column against exactly those targets. Tool sizes
and window prose updated with it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
--stay leaves the loader's final prompt in the USB pipeline; a fresh
invocation on a board that resets when its port opens then flushes too
early, trusts the stale prompt, and spends the new activation window on
a 2-second identity read against a device that never heard its knock —
collecting the application's banner as an unknown signature. Three
host-side moves, no device bytes: the line is drained until quiet
(bounded, 250 ms) before the port's first knock — once per port, since a
mid-session re-knock faces no foreign bytes and its own window is
already burning; the identity read_exact drops 2.0 to 0.5 s, dozens of
times the worst real answer, so any false prompt match leaves room for
the retry that already works; and the tool version drifts to 8. The
StaleDTRPort fixture models the whole moment — stale prompt in transit,
reset holding the device off the line, a finite window, the banner —
red against the old tool in exactly the field shape (unknown signature
from banner bytes), green now; LoaderPort answers its prompt to the
knock rather than to a read count, which the drain exposed as a
call-order coupling.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The de-string-2 pass upstream: every peripheral block behind generated
instance traits, int-keyed, the string layer gone. The port's share of it
is two spellings — the char usart_digit that existed to be pasted into
register names becomes the int unit the usart template now takes, and the
tsb tiers' one reg<"UBRR0"> is the flat hw::ubrr0 — plus the pbapp
harness probing has_usart<0>() instead of instance-name strings. Nine
loader codegen families rebuilt green through their full workflows (size
matrix and simulator protocol suites included); every image holds its
recorded size.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The in-file version window now carries the v7 line its own comment
claimed to hold; the GPIOR note counts words, not instructions; the
USART-release cost and citation match the silicon (two bytes on the
classics, §20.6.3); and the 512-byte claim reads as the slot bound it
is. The libavr pin advances over the review pass — images byte-identical.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Byte-identical images — the traits resolve the same database indices the
retired string forms did; the tightest image is compared outright.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four cuts, none touching what the loader can do. The entry stub stops
re-doing the reset logic's own SP write where the datasheet guarantees
RAMEND (stack::hardware — the classic megas keep theirs). The autobaud
unit moves into GPIOR2:GPIOR1 wherever the chip has the pair: one-word
accesses, no RAM object, and the host's measured-clock peek follows it
by version and geometry. 'J' rides the unified decode, carrying a
selector it ignores so its address is the same two reads as every other
command — the tool sends the bare form to older residents. run_app stops
insisting on a body of its own. The fleet lands at 358–410 B stock and
438–474 B autobaud; the tightest image in the space — the 1284s'
autobaud on a USART's own pins with the OSCCAL trim — drops from 510 to
484 of its 512. Every chip's suite is green on the wire that changed,
and the README's table is machine-checked against the built images.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The three things this repo took from under libavr's counter are now over
it, and the local copies fold away. The USART release on a software
link's pins is the library's init contract (its guard here becomes a
deletion, byte-identical images held by the gate); the WDRF routing test
is power::peek_reset_cause().watchdog instead of a hand lookup of the
flag's register; the tsb tiers' baud arithmetic is the public solver.
libavr pin advances over those three additions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The fixed-baud window counted down a uint32 where almost every window fits
24 bits; the countdown now takes avr::uint24_t when the poll budget allows
(the autobaud budget's own choice), uint32 past 16.7M polls — four bytes
off every fixed-baud image on every chip, the full suites green on the
changed window. The README size table is refreshed — its autobaud column
had also gone stale by the no-assembly pass's measurement-loop win, which
nothing gated: sizes.py check-readme now runs as the gate's final stage,
where every tree is freshly built and the table can actually be held.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pureboot_device and the tsb device, C until now, rewritten in C++23 with
every modeled behavior intact — the PGERS Z-mask and m48-discard ioctl
wraps, the GPIO bridge's timing and pacing, the tiny NVM's write-once
buffer, pin ownership, and the PB_PTY/TSB_PTY lines the harnesses parse.
The one linkage fact worth a comment: simavr's parts headers (uart_pty.h)
carry no C++ guards where its core headers do, so those includes sit in an
extern "C" block. Warning-clean at -Wall -Wextra on the build line; the
full protocol suites on all four sim-driven chips prove the conversion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A rig hiccup mid-walk records the scan check as failed and lets the suite
continue, matching its siblings' envelope; a nonexistent --port path
reports as an error instead of a traceback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
--scan's walk was unwalkable on POSIX: probe rates have no termios
B-constant, so the first off-nominal probe raised out of the loop. The port
speaks termios2 BOTHER now (red-proven on a pty at 9984 Bd), the probe's
open lives inside the walk's error handling, an fd no longer leaks on an
unmakeable rate, and the swallowed unknown-signature reply is named at
timeout instead of reported as silence. CMakePresets.json's generator emits
the submodule toolchain path it had drifted from — a hand edit on a
generated file, exactly the class rule 10 exists for — and presets.generated
gates the pair from here on (the marker CMake rejects at the
presets root stayed out; the check is the guard). The over-slot image guard
the tsb runner gained reaches the pureboot runner too; the GPIO bridge's
delivery comment states the hardware truth (RXC at the stop bit's sampling
point); the hardware suite gains the scan check — the one place the rate
physics is real; and the libavr pin advances over both audit rounds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Configuration gains the OSCCAL row; Deployment says what the build cannot
see (±10 % factory trim against a frame's ~±4 %, and silence that reads as
wiring); the update section names an OSCCAL bake as a link change in effect,
declared with --staged-baud; the host-tool section documents --scan and the
measured clock --info adds on an autobaud session; the version map gains 6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The RC-oscillator answer's host half. --scan probes ±10 % around the built
rate in 2 % steps, nearest first, one activation window (one reset) per
probe: a fixed-baud loader whose oscillator drifted answers at the ratio,
and the report gives the session workaround (--baud), the offset, the
OSCCAL direction at ~1 %/step, and the autobaud way out. The walk and the
advice are logic-tested (test_scan.py, red-proven on the trim direction) —
a pty carries bytes at any rate, so the wire cannot arbitrate them.
On an autobaud session --info now reads the measured bit period from
ram_start — the geometry table gains that column — and undoes the unit's
encoding ((cycles − 8) / 4, floored: libavr's spin granule and per-bit
overhead), so the printed clock is the true one within a granule; --clock
turns it into a stated drift. The autobaud end-to-end asserts the figure
inside exactly that envelope at both clock points.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The RC-oscillator answer's device half (dev/tasks.md in libavr): OSCCAL joins
pureboot_add_loader() as one optional byte, written at the top of run() before
the WDRF bail so the watchdog hand-over inherits the corrected clock too.
Orthogonal to the backend — an autobaud build may carry it purely for the
application. No value, no code: the stock image differs from v5 in exactly
the version's two bytes (the stamp and the 'b' immediate).
Measured: +6 B where OSCCAL takes sts (328P, 404→410), +4 B in low I/O
(t85, 402→406); the tightest image in the space (1284 autobaud on USART
pins, 504) carries the sts form at 510 of 512. New gates: the OSCCAL size
points on every chip, the wire-observed trim byte on both addressing
classes (test/pbosccal.py, red-green), and the autobaud unit pinned to
ram_start (test/check_unit.cmake, red-green) — the address --info's
measured-clock read is about to rely on.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The pin the whole history now encodes becomes the primary route: the
submodule default replaces FetchContent and the unpinned forge fallback,
LIBAVR_ROOT stays as the tandem-development override, the presets already
take the toolchain file from the submodule, and the Studio projects anchor
their include path there — correct by construction. The version tags and
the one-command historical build are documented beside the version map.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
master carries bootloader.atsln, so main does too. Two projects, because a
.cppproj is one binary at one flag set and this build has hundreds: the stock
328P pureboot loader (USART0 at 115200 on a 16 MHz crystal), and the tsb_asm
tier that occupies the same 512-byte section master's own tsb project targeted.
Both come out byte-identical to the Ninja build — 404 B and 510 B of .text —
in both configurations.
Debug keeps -Os and adds only -gdwarf-4. A loader's section is a correctness
bound, and -Og builds this source to 590 B: the link at 0x7e00 accepts that
without a diagnostic, 78 bytes past flash end, where rcall/rjmp wrap modulo
flash size and the image dies just after activation. Debug info costs no flash,
so the optimisation level stays where correctness needs it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A fourth tier answering one question: what does the full TinySafeBoot
feature set cost in C++ under pureboot's rules — no assembly, no
register variables, every pureboot lesson applied. 638 bytes, protocol
suite green: 198 below the idiomatic tier, 126 above the 512 B section,
and above the tiers that pay with the banned mechanisms (526 global
registers, 510 with two asm routines). The gap decomposes into the rent
policy-clean C++ pays for state held across calls — push/pop and
argument threading a global-register protocol avoids — and both
control-flow merges tried measured larger than the split cases they
replaced, while the data merge (one send loop over both memories) paid.
The tiers stay; this one keeps the floor an artifact instead of a claim.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A boot-linked image larger than its slot cannot execute on hardware, and
the naive copy smashed the heap beyond avr->flash — after which the
simulation misbehaved in ways that pointed everywhere but at the size:
phantom byte losses on the UART, garbage in SPMCSR, all downstream of
the overrun. The size gate had said it plainly; now the runner does too.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Writing 0x60-0x61 on an ATtiny13A reliably garbled the link, which looked like a
loader defect. It is the loader's own footing: .noinit lands at exactly 0x60,
size 2, and on an autobaud build that is unit_ — the measured bit period, and the
whole of its static RAM. Overwrite it and the next reply is timed against
garbage, so the symptom is a mangled prompt byte and no error, because nothing
went wrong except the rate both ends had agreed on.
Identical in kind to poking the stack at the top of SRAM, and cleared by a reset.
test/pbautobaud.py already steered its RAM round-trip clear of the bottom of SRAM
for this reason; only the README had not said it. Both regions are named there
now, beside the note that --poke does reach OSCCAL but that a session survives
only a step or two of moving the clock under itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The size matrix proves every image fits its slot and says nothing about the
numbers the README prints. Those drift silently: caller_page() took eight bytes
off every build at once, so all fifteen rows went stale together and no test
noticed, because nothing was over budget. sizes.py check-readme compares the
table against the built images; sizes.py max reports the largest image per chip
and anything over its slot.
It is a check rather than a generator, so the table stays prose someone can
write. No chip geometry lives here either: the image/budget pairs come out of
each build's own CTestTestfile.cmake, which is what the gate checks, so a chip
added or a budget changed needs no edit. Only trees a preset still owns are
read — a stale directory answers with a size that was true once.
16243 images across 37 chips today, none over budget, tightest tsb_asm at 510
of 512.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ATtiny13A run left autobaud's low-clock lock looking unreliable: 1/5 at
19200 on a 1.2 MHz RC part. The simulator does not reproduce it. At an exact
clock the calibration is solid down to ~36 cycles a bit and fails outright by
~31 — a sharp edge, not a fraying one — where the real part was already 1 in 5
by ~59. So the effect is the oscillator's own jitter and not backend logic, and
the two floors are different quantities about a factor of two apart.
Both are worth having. pureboot.autobaud gates a tight-bit point, since its two
existing clock points both sat near 100 cycles a bit and would not notice the
floor moving. The README carries the other half: both floors side by side, the
per-clock envelope measured on silicon, and the reason budgeting the logic's ~36
on an RC part is wrong.
It also carries the trap that produced the confusion. On a patched-vector chip an
erased application region walks back up into the loader, so every expired window
opens another and the host's retries eventually catch the pulse — 5/5 where the
same part with an application resident gives 1/5. Measure with an application in
place, or the fixture flatters the backend.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--update-loader works by entering copies of the *new* image and letting them
rewrite the resident. Those copies speak the link they were built for, but the
host went on knocking with the session's baud and backend — the resident's. Where
the image changed either, the staging copy was installed and then never answered:
resident untouched, and on a 1 KiB tiny the staging slot is the whole application
region, so the application was already gone.
The wire cannot be probed for it. 512 bytes of position-independent code carry no
header saying what rate they were built for, so the operator declares it:
--staged-baud and --staged-autobaud, applied from the jump into the staging copy
onward. Retuning goes through the open port — SetCommState or tcsetattr on the
live handle, never a reopen — because a DTR pulse would reset the copy being
talked to. Undeclared against a changed link it still cannot work, but the error
now names that as the cause instead of reporting the bare activation timeout that
sent the operator looking at wiring.
The README's idempotence claim needed the same qualification: from step 2 a
re-run must reach the new image, and after step 3 word 0 points at the staging
copy, so on a patched-vector part the resident's link reaches nothing at all.
Found on an ATtiny13A, where two controls differing only in the activation window
updated cleanly and so isolated the link as the variable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check.sh proves the protocol under simavr on every chip; it cannot prove a
board. Two things live only on silicon — an RC oscillator that is not on its
nominal, and a reset edge that has to come from somewhere — and until now the
scripts that reached them were per-session scratch on the machine holding the
programmer, which is where the ATtiny13A run's findings nearly stayed.
pbrig.py is the primitives, knowing nothing per-board: every deployment fact is
a flag or a PUREBOOT_* variable. Two rig facts are encoded in it because neither
is guessable and each cost a session to learn: an ISP access *is* the reset edge
where the adapter's DTR is unwired, so a session begins with an ISP touch and
knocks immediately after; and avrdude splits -U on colons, so a Windows drive
letter breaks the spec and every file goes as a bare name with avrdude run in
its own directory. Its `rate` subcommand is the one that turns "the loader is
silent, so the wiring must be wrong" into a number, by sweeping the host rate
against a fixed cycles-per-bit transmitter — PUREBOOT_HEARTBEAT makes the
existing fixture into one, software link only, since the hardware-link idle owes
the self-update tests its command loop.
pbhw.py takes every bound from the info block the loader reports, so one run
covers a 1 KiB tiny and a 128 KiB mega alike. Both are exercised on an ATtiny13A:
backup verified against a known-good capture, the clock measured at 9.048 MHz
against a 9.6 MHz nominal, and the suite 11/11.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pins move the image for exactly one reason — a bit-banged link on a USART's own
pins has to release that USART — and the matrix said outright that they were no
axis, so the tightest configuration in the space was one nothing built. Not
subtly, either: the 1284's slot ends at flash end, so that build does not merely
exceed the size test's limit, it fails to link. pureboot_{sw,autobaud}_on_usart
{0,1} are gate points in both matrix modes now, and the exhaustive sweep carries
the pins across its whole cross product. The hand-measured table is the gate's
output: 506 B of 512 for the 1284 autobaud on USART0's pins, 504 on USART1's.
pureboot.mute drives the defect itself — an application hands over with USART0
still enabled and the loader on those pins must still answer. Reaching that
needed the runner to know an enabled USART owns its TxD, which simavr does not
model at all: it wires a USART through IRQs and never takes the pin from the
port. It also brings UCSRnB up with TXEN already set where silicon clears the
register, so the runner restores the reset value for the USART it models — the
mute must come from the application, not from power-on. The fixture stays
silent, since nothing is listening on the USART it brings up.
test_handshake.py, written where no gate could run it, is pureboot.handshake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It was written next to the loader source; the harness lives at the repo root.
Not registered with ctest yet — it belongs beside pureboot.planner, which is
the other test of the host tool's pure logic.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The write guard's anchor was costing a materialised pointer and a byte swap to
use one byte of it. The libavr primitive answers it in a single load, which is
eight bytes off every build — and what lets the USART release fit the tightest
configuration in the space: the 1284 autobaud on USART-shared pins was 514 of
its 512 and is now 506, with the default pinning down from 510 to 502.
Verified on silicon: the guard still refuses an erase aimed at the slot it runs
from, and still permits one in the application region.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A software or autobaud link on a USART's own pins (PD0/PD1 on the mega328P, so
the Uno's USB bridge reaches it) was mute after an application handed over with
that USART still enabled: its TXEN keeps the USART owning the TX pin, so the
bit-banged transmitter cannot drive it — the loader locked and obeyed commands
but never answered. The link's init now clears the UCSRnB of the USART whose
TXD is its TX pin. Guarded with if constexpr on that pin match, so a link on
non-USART pins emits nothing: +4 bytes on a USART-pin build (494 of 512 for the
mega328P autobaud), zero on the default pb0/pb1 matrix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The drain fix changes the tool's activation behaviour; mark it. The loader
version window is unchanged — the wire protocol did not move.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The post-prompt settle loop in _handshake had no deadline, so a target that
never falls quiet — a board stuck in a reset loop, whose UART-reset garbage
carries a stray prompt byte — spun the tool forever. Bound it by the handshake
deadline; a real loader still settles on its first quiet read. Regression:
test/test_handshake.py (flood terminates, valid loader still connects).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
R/r/w/F collapse into G and g over a selector byte naming the space — flash,
EEPROM, data, fuse, SPM — with the flash bank in its high nibble. Four command
bodies, four transfer loops and four argument decodes become one of each, and
W joins the same decode instead of keeping an address form of its own. The
loader shrinks while gaining everything below: on the 1284P the stock build
goes 480 -> 432 B and the software one 496 -> 450.
What the freed space buys:
- Data space. On AVR one pointer spans SRAM, the register file and the whole
I/O space, so G over space 2 reads all three. pureboot keeps zero static
RAM and pushes no register, so at loader entry an application's SRAM is
still what the application left there — this is a post-mortem, not just a
poke hole. As its own command it needed a dispatch arm and a loop; as one
more space it is a single ld/st.
- Host-issued SPM. W fills the page buffer and stops; erase, write and RWW
re-enable are writes to space 4, which reach the same fused store-and-SPM
pair through the transfer's own address and data. Any SPM operation, lock
bits included, is now reachable and the loader carries no page-commit logic.
The four-cycle SPMCSR-to-SPM window is why that primitive stays fused: no
host can hit it across a serial link, and that — not the byte count — is
the floor on how low-level a bootloader's primitives can go.
- Byte addresses everywhere. The bank in the selector retires the
word-addressed wire the >64 KiB parts needed, so the 1284s stop being the
outlier.
SERIAL autobaud is a third backend on the same loader, over libavr's
software_autobaud: no clock, no baud, one binary per chip for every F_CPU and
every rate. Activation counts poll iterations rather than seconds and bounds
every wait, so a stray pulse cannot hold an unattended device.
b answers with the version and signature only; the host derives geometry from
the signature, which is what an autobaud build requires anyway. An update image
is a bare slot with no device to ask, so every image carries a six-byte stamp —
the same bytes b answers with, and the source of both — that the loader never
reads from flash and the host refuses to install a mismatch against. The
running-slot write guard moved onto the SPM commit, which covers erase and
write both where guarding W covered neither directly.
The position-independence lint now proves the property instead of a proxy for
it: the image must come out byte-identical linked at a different base.
-fno-move-loop-invariants left the tuned flag set — it was fitted to a command
loop carrying four transfer bodies and costs bytes now that it carries one.
Verified: the exhaustive matrix on all 37 chips (every clock x every baud x
every backend, non-standard rates included, plus the autobaud build) —
8174 size checks, no failures, tightest fit the 1284s' autobaud at 510 of 512.
Behavioral suites green on every chip class: t13a 10/10, t85 11/11, m8 13/13,
m16a 13/13, m48pa 13/13, 328P 23/23, 644A 17/17, 1284P 17/17. Data-space
round trip through --peek/--poke and the autobaud handshake are both red-green
proven.
The two prototype sources and their findings file go; the README carries the
protocol and dev/done.md in libavr carries the reasoning.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hardware testing found that a lone calibration pulse wedged the autobaud loader:
run() budgeted only the start-edge wait in measure(), and the rx() that read the
knock behind it was unbudgeted, so one stray low pulse held an unattended device
in the loader and the application never ran. Bound the whole activation — an
expired knock budget returns a byte that cannot be the knock, so control falls
back into the budgeted measure() and an idle line boots the app there.
That fix costs ~22 B, which neither version under review could absorb: the pure
one goes 508 -> 530 on the 1284P and the register one 512 -> 534, both over a
512 B slot. Their margin was never spare capacity, it was the space the missing
fix should have occupied. So the choice between them is moot; both are kept for
the record and no longer built.
pureboot_autobaud_uni.cpp replaces them at 464 B. It is pureboot 5: one read
command and one write command over named spaces (G/g, sel8, addr16, n8) instead
of four per-memory bodies, which collapses four transfer loops into one. The
selector's high nibble carries flash's bank, so the shared cursor stays 16 bits
and no command speaks word addresses. Three things fall out of the freed space:
RAM read/write — the missing feature, and with it arbitrary I/O access, since
AVR maps peripherals into the data space; host-issued SPM, so W's hardcoded
erase/write/RWW tail becomes three writes to a space and any SPM operation is
reachable; and W on the same selector-and-address decode as everything else.
Strictly pure throughout: no inline asm, no global register variable, and no
GPIOR either — the unit lives in a .noinit static, so the loader claims no chip
resource and the chips without GPIOR stop being a special case.
pureboot.py speaks both generations, keyed on the version, so the fixed-baud
path is untouched; --peek/--poke reach the new data space. pbautobaud.py adds a
RAM round-trip and a regression for the hang: a lone pulse must still let the
app boot. All 37 chips plus the 12-preset reflect spot set build and size-test
green, 444-466 B, worst case 46 B under budget. Sim suites 100%: 1284P 17/17,
328P 23/23. Only real-hardware acceptance remains (pureboot/autobaud.md).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
pureboot.py --autobaud sends the 0xC0 calibration pulse and a single knock at
the host's chosen baud, reads the slimmed info block, and derives the full
geometry from the signature (AUTOBAUD_GEOMETRY, a table over every pureboot
chip). Everything downstream — flash, EEPROM, fuses, hand-over, verify — is the
fixed-baud path unchanged; the dropped write guard is host-transparent.
test/pbautobaud.py drives each variant over the GPIO⇄pty software-UART bridge
through the calibration handshake and a flash + EEPROM + fuse round-trip
cross-checked against the simulator's ground-truth memory, then repeats at
double the F_CPU with the same binary — the clock-agnostic property autobaud
exists for. Wired as pureboot.autobaud_pure/reg on the near-flash 328P and the
word-addressed 1284P. A wrong measured unit fails the flash/verify, so the test
also pins the codegen-coupled calibration constant against a toolchain bump.
Both variants green in sim on both chips at two clocks each; the fixed-baud
suite is unaffected. Only real-hardware acceptance on an RC part remains
(pureboot/autobaud.md).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one
clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments
no longer need a per-clock build.
Two source files, differing only in the write-guard/purity tradeoff:
pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch
registers, running-slot write guard dropped, 508 B on the 1284) stays strictly
pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard
kept, 512 B) keeps every feature at the cost of that single GRV. Both fit
512/510 on all 37 chips and share two licensed simplifications: a slimmed info
block (version + signature; the host derives geometry from the chip database)
and a single-byte activation knock.
pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that
set the target, and the compiler-knob path to it. Size-tested on every chip via
pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the
host calibration handshake, and real-hardware acceptance remain.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A sticky WDRF diverts every reset past the activation window (deliberate, so
an app can reboot instantly, at the cost of a possible lockout); an EEPROM
address past E2END wraps onto low EEPROM (the host bounds it, not the loader).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
clang-format bin-packs braced lists to the column limit, collapsing the
'b' reply's byte layout into dense rows. A minimal clang-format-off span
keeps each wire byte on its own line, where the layout is legible against
the protocol. Whitespace only; image byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The unaligned-W and U2X-hand-over fixes change the loader's observable
on-wire behavior, and the --stay reconnect fix changes the host tool, so
both move: loader version 3 -> 4, tool VERSION 2 -> 3. The protocol and info
block are unchanged, so OLDEST_LOADER stays 1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A prompt byte alone does not: one left over from a previous session can
still be in the pipeline while the port opening resets the device into a
fresh window, where the bare command that follows is discarded. Each
attempt is now the whole handshake, retried until the block comes back.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The in-page bits of a W address are dropped so the fill always walks from
the page base; the wire contract is one page of data for any address
inside it, on both the byte- and the word-addressed path. +2 B.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every plausible oscillator against every rate it reaches against every
backend, on one chip per size-bearing class, under --full only. The baud
ladder becomes a reachability predicate the enumeration filters on, so an
unreachable point drops out instead of aborting the configure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>