The RC-oscillator answer's host half. --scan probes ±10 % around the built
rate in 2 % steps, nearest first, one activation window (one reset) per
probe: a fixed-baud loader whose oscillator drifted answers at the ratio,
and the report gives the session workaround (--baud), the offset, the
OSCCAL direction at ~1 %/step, and the autobaud way out. The walk and the
advice are logic-tested (test_scan.py, red-proven on the trim direction) —
a pty carries bytes at any rate, so the wire cannot arbitrate them.
On an autobaud session --info now reads the measured bit period from
ram_start — the geometry table gains that column — and undoes the unit's
encoding ((cycles − 8) / 4, floored: libavr's spin granule and per-bit
overhead), so the printed clock is the true one within a granule; --clock
turns it into a stated drift. The autobaud end-to-end asserts the figure
inside exactly that envelope at both clock points.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The RC-oscillator answer's device half (dev/tasks.md in libavr): OSCCAL joins
pureboot_add_loader() as one optional byte, written at the top of run() before
the WDRF bail so the watchdog hand-over inherits the corrected clock too.
Orthogonal to the backend — an autobaud build may carry it purely for the
application. No value, no code: the stock image differs from v5 in exactly
the version's two bytes (the stamp and the 'b' immediate).
Measured: +6 B where OSCCAL takes sts (328P, 404→410), +4 B in low I/O
(t85, 402→406); the tightest image in the space (1284 autobaud on USART
pins, 504) carries the sts form at 510 of 512. New gates: the OSCCAL size
points on every chip, the wire-observed trim byte on both addressing
classes (test/pbosccal.py, red-green), and the autobaud unit pinned to
ram_start (test/check_unit.cmake, red-green) — the address --info's
measured-clock read is about to rely on.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The pin the whole history now encodes becomes the primary route: the
submodule default replaces FetchContent and the unpinned forge fallback,
LIBAVR_ROOT stays as the tandem-development override, the presets already
take the toolchain file from the submodule, and the Studio projects anchor
their include path there — correct by construction. The version tags and
the one-command historical build are documented beside the version map.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
master carries bootloader.atsln, so main does too. Two projects, because a
.cppproj is one binary at one flag set and this build has hundreds: the stock
328P pureboot loader (USART0 at 115200 on a 16 MHz crystal), and the tsb_asm
tier that occupies the same 512-byte section master's own tsb project targeted.
Both come out byte-identical to the Ninja build — 404 B and 510 B of .text —
in both configurations.
Debug keeps -Os and adds only -gdwarf-4. A loader's section is a correctness
bound, and -Og builds this source to 590 B: the link at 0x7e00 accepts that
without a diagnostic, 78 bytes past flash end, where rcall/rjmp wrap modulo
flash size and the image dies just after activation. Debug info costs no flash,
so the optimisation level stays where correctness needs it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A fourth tier answering one question: what does the full TinySafeBoot
feature set cost in C++ under pureboot's rules — no assembly, no
register variables, every pureboot lesson applied. 638 bytes, protocol
suite green: 198 below the idiomatic tier, 126 above the 512 B section,
and above the tiers that pay with the banned mechanisms (526 global
registers, 510 with two asm routines). The gap decomposes into the rent
policy-clean C++ pays for state held across calls — push/pop and
argument threading a global-register protocol avoids — and both
control-flow merges tried measured larger than the split cases they
replaced, while the data merge (one send loop over both memories) paid.
The tiers stay; this one keeps the floor an artifact instead of a claim.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A boot-linked image larger than its slot cannot execute on hardware, and
the naive copy smashed the heap beyond avr->flash — after which the
simulation misbehaved in ways that pointed everywhere but at the size:
phantom byte losses on the UART, garbage in SPMCSR, all downstream of
the overrun. The size gate had said it plainly; now the runner does too.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Writing 0x60-0x61 on an ATtiny13A reliably garbled the link, which looked like a
loader defect. It is the loader's own footing: .noinit lands at exactly 0x60,
size 2, and on an autobaud build that is unit_ — the measured bit period, and the
whole of its static RAM. Overwrite it and the next reply is timed against
garbage, so the symptom is a mangled prompt byte and no error, because nothing
went wrong except the rate both ends had agreed on.
Identical in kind to poking the stack at the top of SRAM, and cleared by a reset.
test/pbautobaud.py already steered its RAM round-trip clear of the bottom of SRAM
for this reason; only the README had not said it. Both regions are named there
now, beside the note that --poke does reach OSCCAL but that a session survives
only a step or two of moving the clock under itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The size matrix proves every image fits its slot and says nothing about the
numbers the README prints. Those drift silently: caller_page() took eight bytes
off every build at once, so all fifteen rows went stale together and no test
noticed, because nothing was over budget. sizes.py check-readme compares the
table against the built images; sizes.py max reports the largest image per chip
and anything over its slot.
It is a check rather than a generator, so the table stays prose someone can
write. No chip geometry lives here either: the image/budget pairs come out of
each build's own CTestTestfile.cmake, which is what the gate checks, so a chip
added or a budget changed needs no edit. Only trees a preset still owns are
read — a stale directory answers with a size that was true once.
16243 images across 37 chips today, none over budget, tightest tsb_asm at 510
of 512.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ATtiny13A run left autobaud's low-clock lock looking unreliable: 1/5 at
19200 on a 1.2 MHz RC part. The simulator does not reproduce it. At an exact
clock the calibration is solid down to ~36 cycles a bit and fails outright by
~31 — a sharp edge, not a fraying one — where the real part was already 1 in 5
by ~59. So the effect is the oscillator's own jitter and not backend logic, and
the two floors are different quantities about a factor of two apart.
Both are worth having. pureboot.autobaud gates a tight-bit point, since its two
existing clock points both sat near 100 cycles a bit and would not notice the
floor moving. The README carries the other half: both floors side by side, the
per-clock envelope measured on silicon, and the reason budgeting the logic's ~36
on an RC part is wrong.
It also carries the trap that produced the confusion. On a patched-vector chip an
erased application region walks back up into the loader, so every expired window
opens another and the host's retries eventually catch the pulse — 5/5 where the
same part with an application resident gives 1/5. Measure with an application in
place, or the fixture flatters the backend.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--update-loader works by entering copies of the *new* image and letting them
rewrite the resident. Those copies speak the link they were built for, but the
host went on knocking with the session's baud and backend — the resident's. Where
the image changed either, the staging copy was installed and then never answered:
resident untouched, and on a 1 KiB tiny the staging slot is the whole application
region, so the application was already gone.
The wire cannot be probed for it. 512 bytes of position-independent code carry no
header saying what rate they were built for, so the operator declares it:
--staged-baud and --staged-autobaud, applied from the jump into the staging copy
onward. Retuning goes through the open port — SetCommState or tcsetattr on the
live handle, never a reopen — because a DTR pulse would reset the copy being
talked to. Undeclared against a changed link it still cannot work, but the error
now names that as the cause instead of reporting the bare activation timeout that
sent the operator looking at wiring.
The README's idempotence claim needed the same qualification: from step 2 a
re-run must reach the new image, and after step 3 word 0 points at the staging
copy, so on a patched-vector part the resident's link reaches nothing at all.
Found on an ATtiny13A, where two controls differing only in the activation window
updated cleanly and so isolated the link as the variable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check.sh proves the protocol under simavr on every chip; it cannot prove a
board. Two things live only on silicon — an RC oscillator that is not on its
nominal, and a reset edge that has to come from somewhere — and until now the
scripts that reached them were per-session scratch on the machine holding the
programmer, which is where the ATtiny13A run's findings nearly stayed.
pbrig.py is the primitives, knowing nothing per-board: every deployment fact is
a flag or a PUREBOOT_* variable. Two rig facts are encoded in it because neither
is guessable and each cost a session to learn: an ISP access *is* the reset edge
where the adapter's DTR is unwired, so a session begins with an ISP touch and
knocks immediately after; and avrdude splits -U on colons, so a Windows drive
letter breaks the spec and every file goes as a bare name with avrdude run in
its own directory. Its `rate` subcommand is the one that turns "the loader is
silent, so the wiring must be wrong" into a number, by sweeping the host rate
against a fixed cycles-per-bit transmitter — PUREBOOT_HEARTBEAT makes the
existing fixture into one, software link only, since the hardware-link idle owes
the self-update tests its command loop.
pbhw.py takes every bound from the info block the loader reports, so one run
covers a 1 KiB tiny and a 128 KiB mega alike. Both are exercised on an ATtiny13A:
backup verified against a known-good capture, the clock measured at 9.048 MHz
against a 9.6 MHz nominal, and the suite 11/11.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pins move the image for exactly one reason — a bit-banged link on a USART's own
pins has to release that USART — and the matrix said outright that they were no
axis, so the tightest configuration in the space was one nothing built. Not
subtly, either: the 1284's slot ends at flash end, so that build does not merely
exceed the size test's limit, it fails to link. pureboot_{sw,autobaud}_on_usart
{0,1} are gate points in both matrix modes now, and the exhaustive sweep carries
the pins across its whole cross product. The hand-measured table is the gate's
output: 506 B of 512 for the 1284 autobaud on USART0's pins, 504 on USART1's.
pureboot.mute drives the defect itself — an application hands over with USART0
still enabled and the loader on those pins must still answer. Reaching that
needed the runner to know an enabled USART owns its TxD, which simavr does not
model at all: it wires a USART through IRQs and never takes the pin from the
port. It also brings UCSRnB up with TXEN already set where silicon clears the
register, so the runner restores the reset value for the USART it models — the
mute must come from the application, not from power-on. The fixture stays
silent, since nothing is listening on the USART it brings up.
test_handshake.py, written where no gate could run it, is pureboot.handshake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It was written next to the loader source; the harness lives at the repo root.
Not registered with ctest yet — it belongs beside pureboot.planner, which is
the other test of the host tool's pure logic.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The write guard's anchor was costing a materialised pointer and a byte swap to
use one byte of it. The libavr primitive answers it in a single load, which is
eight bytes off every build — and what lets the USART release fit the tightest
configuration in the space: the 1284 autobaud on USART-shared pins was 514 of
its 512 and is now 506, with the default pinning down from 510 to 502.
Verified on silicon: the guard still refuses an erase aimed at the slot it runs
from, and still permits one in the application region.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A software or autobaud link on a USART's own pins (PD0/PD1 on the mega328P, so
the Uno's USB bridge reaches it) was mute after an application handed over with
that USART still enabled: its TXEN keeps the USART owning the TX pin, so the
bit-banged transmitter cannot drive it — the loader locked and obeyed commands
but never answered. The link's init now clears the UCSRnB of the USART whose
TXD is its TX pin. Guarded with if constexpr on that pin match, so a link on
non-USART pins emits nothing: +4 bytes on a USART-pin build (494 of 512 for the
mega328P autobaud), zero on the default pb0/pb1 matrix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The drain fix changes the tool's activation behaviour; mark it. The loader
version window is unchanged — the wire protocol did not move.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The post-prompt settle loop in _handshake had no deadline, so a target that
never falls quiet — a board stuck in a reset loop, whose UART-reset garbage
carries a stray prompt byte — spun the tool forever. Bound it by the handshake
deadline; a real loader still settles on its first quiet read. Regression:
test/test_handshake.py (flood terminates, valid loader still connects).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
R/r/w/F collapse into G and g over a selector byte naming the space — flash,
EEPROM, data, fuse, SPM — with the flash bank in its high nibble. Four command
bodies, four transfer loops and four argument decodes become one of each, and
W joins the same decode instead of keeping an address form of its own. The
loader shrinks while gaining everything below: on the 1284P the stock build
goes 480 -> 432 B and the software one 496 -> 450.
What the freed space buys:
- Data space. On AVR one pointer spans SRAM, the register file and the whole
I/O space, so G over space 2 reads all three. pureboot keeps zero static
RAM and pushes no register, so at loader entry an application's SRAM is
still what the application left there — this is a post-mortem, not just a
poke hole. As its own command it needed a dispatch arm and a loop; as one
more space it is a single ld/st.
- Host-issued SPM. W fills the page buffer and stops; erase, write and RWW
re-enable are writes to space 4, which reach the same fused store-and-SPM
pair through the transfer's own address and data. Any SPM operation, lock
bits included, is now reachable and the loader carries no page-commit logic.
The four-cycle SPMCSR-to-SPM window is why that primitive stays fused: no
host can hit it across a serial link, and that — not the byte count — is
the floor on how low-level a bootloader's primitives can go.
- Byte addresses everywhere. The bank in the selector retires the
word-addressed wire the >64 KiB parts needed, so the 1284s stop being the
outlier.
SERIAL autobaud is a third backend on the same loader, over libavr's
software_autobaud: no clock, no baud, one binary per chip for every F_CPU and
every rate. Activation counts poll iterations rather than seconds and bounds
every wait, so a stray pulse cannot hold an unattended device.
b answers with the version and signature only; the host derives geometry from
the signature, which is what an autobaud build requires anyway. An update image
is a bare slot with no device to ask, so every image carries a six-byte stamp —
the same bytes b answers with, and the source of both — that the loader never
reads from flash and the host refuses to install a mismatch against. The
running-slot write guard moved onto the SPM commit, which covers erase and
write both where guarding W covered neither directly.
The position-independence lint now proves the property instead of a proxy for
it: the image must come out byte-identical linked at a different base.
-fno-move-loop-invariants left the tuned flag set — it was fitted to a command
loop carrying four transfer bodies and costs bytes now that it carries one.
Verified: the exhaustive matrix on all 37 chips (every clock x every baud x
every backend, non-standard rates included, plus the autobaud build) —
8174 size checks, no failures, tightest fit the 1284s' autobaud at 510 of 512.
Behavioral suites green on every chip class: t13a 10/10, t85 11/11, m8 13/13,
m16a 13/13, m48pa 13/13, 328P 23/23, 644A 17/17, 1284P 17/17. Data-space
round trip through --peek/--poke and the autobaud handshake are both red-green
proven.
The two prototype sources and their findings file go; the README carries the
protocol and dev/done.md in libavr carries the reasoning.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hardware testing found that a lone calibration pulse wedged the autobaud loader:
run() budgeted only the start-edge wait in measure(), and the rx() that read the
knock behind it was unbudgeted, so one stray low pulse held an unattended device
in the loader and the application never ran. Bound the whole activation — an
expired knock budget returns a byte that cannot be the knock, so control falls
back into the budgeted measure() and an idle line boots the app there.
That fix costs ~22 B, which neither version under review could absorb: the pure
one goes 508 -> 530 on the 1284P and the register one 512 -> 534, both over a
512 B slot. Their margin was never spare capacity, it was the space the missing
fix should have occupied. So the choice between them is moot; both are kept for
the record and no longer built.
pureboot_autobaud_uni.cpp replaces them at 464 B. It is pureboot 5: one read
command and one write command over named spaces (G/g, sel8, addr16, n8) instead
of four per-memory bodies, which collapses four transfer loops into one. The
selector's high nibble carries flash's bank, so the shared cursor stays 16 bits
and no command speaks word addresses. Three things fall out of the freed space:
RAM read/write — the missing feature, and with it arbitrary I/O access, since
AVR maps peripherals into the data space; host-issued SPM, so W's hardcoded
erase/write/RWW tail becomes three writes to a space and any SPM operation is
reachable; and W on the same selector-and-address decode as everything else.
Strictly pure throughout: no inline asm, no global register variable, and no
GPIOR either — the unit lives in a .noinit static, so the loader claims no chip
resource and the chips without GPIOR stop being a special case.
pureboot.py speaks both generations, keyed on the version, so the fixed-baud
path is untouched; --peek/--poke reach the new data space. pbautobaud.py adds a
RAM round-trip and a regression for the hang: a lone pulse must still let the
app boot. All 37 chips plus the 12-preset reflect spot set build and size-test
green, 444-466 B, worst case 46 B under budget. Sim suites 100%: 1284P 17/17,
328P 23/23. Only real-hardware acceptance remains (pureboot/autobaud.md).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
pureboot.py --autobaud sends the 0xC0 calibration pulse and a single knock at
the host's chosen baud, reads the slimmed info block, and derives the full
geometry from the signature (AUTOBAUD_GEOMETRY, a table over every pureboot
chip). Everything downstream — flash, EEPROM, fuses, hand-over, verify — is the
fixed-baud path unchanged; the dropped write guard is host-transparent.
test/pbautobaud.py drives each variant over the GPIO⇄pty software-UART bridge
through the calibration handshake and a flash + EEPROM + fuse round-trip
cross-checked against the simulator's ground-truth memory, then repeats at
double the F_CPU with the same binary — the clock-agnostic property autobaud
exists for. Wired as pureboot.autobaud_pure/reg on the near-flash 328P and the
word-addressed 1284P. A wrong measured unit fails the flash/verify, so the test
also pins the codegen-coupled calibration constant against a toolchain bump.
Both variants green in sim on both chips at two clocks each; the fixed-baud
suite is unaffected. Only real-hardware acceptance on an RC part remains
(pureboot/autobaud.md).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one
clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments
no longer need a per-clock build.
Two source files, differing only in the write-guard/purity tradeoff:
pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch
registers, running-slot write guard dropped, 508 B on the 1284) stays strictly
pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard
kept, 512 B) keeps every feature at the cost of that single GRV. Both fit
512/510 on all 37 chips and share two licensed simplifications: a slimmed info
block (version + signature; the host derives geometry from the chip database)
and a single-byte activation knock.
pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that
set the target, and the compiler-knob path to it. Size-tested on every chip via
pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the
host calibration handshake, and real-hardware acceptance remain.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A sticky WDRF diverts every reset past the activation window (deliberate, so
an app can reboot instantly, at the cost of a possible lockout); an EEPROM
address past E2END wraps onto low EEPROM (the host bounds it, not the loader).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
clang-format bin-packs braced lists to the column limit, collapsing the
'b' reply's byte layout into dense rows. A minimal clang-format-off span
keeps each wire byte on its own line, where the layout is legible against
the protocol. Whitespace only; image byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The unaligned-W and U2X-hand-over fixes change the loader's observable
on-wire behavior, and the --stay reconnect fix changes the host tool, so
both move: loader version 3 -> 4, tool VERSION 2 -> 3. The protocol and info
block are unchanged, so OLDEST_LOADER stays 1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A prompt byte alone does not: one left over from a previous session can
still be in the pipeline while the port opening resets the device into a
fresh window, where the bare command that follows is discarded. Each
attempt is now the whole handshake, retried until the block comes back.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The in-page bits of a W address are dropped so the fill always walks from
the page base; the wire contract is one page of data for any address
inside it, on both the byte- and the word-addressed path. +2 B.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every plausible oscillator against every rate it reaches against every
backend, on one chip per size-bearing class, under --full only. The baud
ladder becomes a reachability predicate the enumeration filters on, so an
unreachable point drops out instead of aborting the configure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The word-addressed 1284s were the one family deploying in a 1 KiB slot,
because the far-flash machinery (ELPM reads, RAMPZ page commands, a
word-addressed wire) did not fit 512 B. It does now: 478 B stock, 494 B in
the heaviest configuration the build can produce. They take the 644s'
geometry, where the smallest boot section holds the resident slot and its
staging slot together. The loader's version goes to 3; the host tool did
not change, so its own version stays 2 and only the window it speaks
widens.
Most of the saving is one restructure. The info block and a flash read are
the same act, so giving all four streamed commands one address-and-count
path leaves exactly one call site for the flash streamer: it inlines into
the never-returning command loop and its 24-bit cursor stops being saved
and restored around every transmit. Around it, the ack byte moved out of
line, the wire's byte pair is bit_cast into the word it already is, the
fuse loop ends on its count, the info block's in-slot offset is taken as
the one-byte relocation it is, and -fno-expensive-optimizations gives way
to -fno-move-loop-invariants -fno-tree-ter. Every chip shrank 14-18 B.
The size matrix grew the axes it was missing: the USART1 instance across
the whole clock ladder, and the shape a slow baud gives a software UART —
past 255 delay iterations libavr takes the 16-bit delay loop, which the
ladder default never selects and which was 4 B over the 1284's slot the
first time it was built.
The protocol fixture stopped deriving the loader entry from the flash
size; on the 1284s it had been jumping a slot low and reaching the loader
only because erased flash walked it up.
Docs and comments were consolidated across the port in the same pass: the
README carries a per-chip size table instead of prose, and prose that
restated the code is gone — 190 lines, no behaviour with it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The info block's third byte was a protocol version that never moved in the
loader's lifetime. It is the pureboot version now, and this change is
version 2: the one number that says what a deployed loader is. Every loader
already in the field answers 1.
The protocol keeps no number of its own — a pureboot version implies it, and
the host tool is what holds that map. pureboot.py states the loader-version
window it speaks (OLDEST_LOADER/NEWEST_LOADER; a version that changes the
protocol becomes the new floor there), so a loader newer than the tool is
refused by name rather than decoded on the assumption nothing moved, while an
older one is read, identified and installed like any other. The tool carries
its own version, free to drift from the loader's: --version prints it and the
window, --info leads with the device's, --update-loader names the version it
installs.
Tests: the planner unit pins the window — every version in it decodes, one
above it is refused, an older loader's image is still found — and the live
suite pins the built loader against the tool beside it, so a bump that reaches
only one of them fails. The image is byte-identical to the previous build but
for that byte.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The headline numbers described the stock deployments but claimed "every
configured variant of them", which the matrix contradicts: choosing the
software UART where the chip has a USART costs 8-46 B, so the megas reach
460-462 rather than 452, and the 1284s' software-serial build is 546 B —
inside their 1 KiB boot sector, but not inside 512.
Also names the actual tightest chip. The 1284 looks like it at 506, but it
deploys in 1 KiB with 478 B spare; against its own budget the ATmega328P
has the least room, 50 B.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The temporary page buffer is write-once per word, so a page filled over
one an earlier writer left dirty programs the stale words. The same
datasheet clause carries the cure: the buffer auto-erases after a page
write (§26.2.1; §19.2 on the tinies), so the corruption clears itself by
happening, and rewriting the page programs correctly.
The loader therefore clears the buffer nowhere. The tinies' CTPB and the
m48s' RWWSRE discard are gone; the boot-sectioned megas keep only the
trailing RWWSRE they need anyway to re-enable the RWW section for
read-back, which discards the buffer as a side effect and keeps them off
the path entirely. 434 B on the tiny13s, 438-442 on the tiny25/45/85,
430 on the m48s; the megas are unchanged, the 1284s still 506.
The host takes over the guarantee: a flash page that reads back wrong is
rewritten up to RETRIES times before the run stops. Both read-back paths
repair — verify_pages for programming, and write_differing, which is the
loader-update path where a page left wrong is a half-written loader slot.
That one is not hypothetical: deleting the discard made attiny85
pureboot.rehome fail deterministically there, the only flow still
assuming the old contract.
Protocol-visible, so README's W command says it: one W may program the
wrong bytes after a refused page, or after an application that
self-programmed entered without a reset, and a host that programs without
reading back cannot trust it.
Tests: pureboot.dirty drives the case the loader declines to guard — the
fixture application dirties every buffer word and jumps in with no reset
(hardware forbids that on a boot-sectioned mega, but simavr dispatches SPM
from anywhere, which is what makes it constructible) — and asserts a bare
verify sees the corruption, the repairing verify fixes it in one rewrite,
and it stays fixed. pbreloc asserts the same shape after a refusal.
test_planner covers the bound against a fake device: one bad write
repaired in a single rewrite, a page that never comes good stopping after
exactly three.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
tools/make_presets.py emits the uniform pipeline the hand-grown file had
drifted from — generated configure/build/test presets and workflows for
all 37 chips, reflect configure/build for libavr's 12-chip spot set —
and tools/check.sh runs every chip's workflow (--full adds the reflect
spot) as the port's gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
--info prints the decoded info block field by field and --fuses each
byte on its own line plus the BOOTSZ/BOOTRST meaning on boot-sectioned
megas. Transfers that take wire time draw a transient progress bar on
stderr when it is a tty — logs, pipes and the tests see only the
summary lines. -v/--verbose narrates decisions: knock counts, the
programming plan, update state handling and per-phase page counts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Clock, baud, serial backend (hardware USART 0/1 or the software UART on
any pins) and the activation window all resolve through one CMake
function, pureboot_add_loader() in pureboot/CMakeLists.txt — the unit a
downstream project consumes. The default baud is the fastest standard
rate within 2.5 % (the same best-divisor search libavr's solver runs),
gated on software builds by the polled receiver's 100-cycles-a-bit
floor; every explicit pick is re-checked by the compile's static asserts.
The size matrix builds each axis that can move the image — backend x
clock ladder x USART instance, per chip — against the slot budget, and
two nondefault deployments run the whole protocol suite live: the 328P
on its shipped 1 MHz fuses over software serial on TX=PB1/RX=PB5
(pureboot.custom), and the 644A over USART1 (pureboot.usart1). The sim
runner takes -l to bridge any link, paces a fully quiet bridge toward
real time (a free-running 8 M-cycle window loses the reset-race knock),
and the fixture application speaks the deployment it is built for.
The loader itself shed bytes on the way: the return-address high byte
spelled through byteswap (the double swap folds to the one-byte pick),
the info-block address composed instead of bit_cast, and libavr's new
polled-UART helpers replacing the port's uart::detail reaches. Every
combination fits: 458-506 B across the megas' whole matrix, 470-484 B
on the tinies, 556-562 B in the 1284s' 1 KiB slot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The update flow's first step wrote the staging content over whatever the
staging slot held; with a loader running right there (programmed by hand
onto erased flash), that write met the copy's own running-slot guard on
the composed through-word and the tool stopped at its verify — although
the copy is exactly an installed staging copy, able to stream the new
resident like any other. The install is now skipped when the slot holds a
complete loader: its info block where every image carries it, matching
the device's byte for byte, and the slot unchanged since the update began
(the state file's snapshot) — so a resumed half-written install still
differs from its snapshot and takes the install path, which completes it.
pbrehome gains the staging-slot position (an older build at stage
streaming a newer resident in); the README's wrong "cannot re-home from
the staging slot" claim is corrected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The README's deployment section now says what each build artifact is for:
the .hex is the programmer artifact (self-addressed into the top slot),
the .bin the self-update image — bare slot bytes a programmer would put
at address 0, where a boot-sectioned mega cannot even heal itself (SPM
only runs from the boot section) but a patched-vector chip runs the
position-independent copy and re-homes a build through the ordinary
--update-loader flow: the staging install and the word-0 redirect both
execute outside page 0's slot, so the running-slot guard never blocks it.
pbrehome.py is the acceptance test (misplaced at 0, guard intact,
re-home, app flash over the stale copy, banner); the staging slot is the
one position that cannot re-home itself, documented. The preflight's
wrong-chip refusal and loader_image's handling of padded images (peeled
to the slot content by the embedded base) are documented and the padded
case pinned in the planner.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The chip table becomes family blocks covering all 37 targets. The m48s
are a new deployment class: no boot section, so the tiny profile spoken
over the hardware USART — host-patched reset vector, trampoline
hand-over, a 510-byte budget (474 B built), no fuse preflight — while
their RWWSRE store stays the buffer discard (Atmel-8271 §26.2); the
device keys the patch flag and the CPU-halt waits on the curated
boot-section capability and the discard on the RWWSRE bit itself. The
644s' 64 KiB is exactly the 16-bit byte space: plain LPM, byte wire
addresses, 498 B in a 512-byte slot — and their 1 KiB minimum boot
section holds the resident and staging slots together, so self-update
needs no fuse step (the update test's slot pick now keys word-flash on
base >= 64 KiB; base + slot merely touching the boundary stays
byte-addressed). The 1284 joins the 1284P's word-addressed 1 KiB slot at
558 B. BOOT_FUSE gains every boot-sectioned family's ladder and fuse
byte; the planner exercises them all. The sim scaffolding keys
patch-vector-ness instead of the atmega name prefix, the fixture app
picks its clock by family (the tiny25/45/13 builds surfaced the 16 MHz
fallthrough as garbled banners), and the runner's wrapped flash ioctl
performs the m48 discard simavr's no-RWW cores turn into a stray buffer
fill. Sizes across the fleet: 466-504 B megas, 474 B m48s, 498 B 644s,
488-502 B tinies, 558 B 1284s — every chip passing
size/pi/planner/protocol/reloc/update.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The deployment section claimed the 1284P has no standalone profile, runs
BOOTSZ = 512 words always, and self-updates with no fuse change, reset
landing at 0x1f800 — internally contradictory (a 512-word section starts
at 0x1fc00, and the section holding both 1 KiB slots is 1024 words) and
contradicted by update_preflight, which refuses a self-update unless the
boot section covers two slots. The text described a 512-byte-slot
geometry this chip's loader cannot have. In truth the 328P profile table
maps onto the 1284P doubled: standalone = 512 words (the smallest
section is exactly the 1 KiB slot, reset at the loader base), self-update
= 1024 words with the loader-first reset walking the staging slot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The far machinery (ELPM reads, RAMPZ page commands, wire-word math)
costs ~46 B over the m328P's 504, and the tsb-calibrated C++-to-asm
gap says no implementation of this feature set reaches 512 on this
chip — a boundary its hardware does not have anyway: the 1284P's
smallest boot sector is 1 KiB. The slot therefore becomes
per-geometry (512 B, or 1 KiB past 64 KiB), which the host derives
from the word-addressing flag; slot arithmetic unifies (the index is
the wire high byte with its low bit dropped in either unit), the
update preflight demands a two-slot boot section in the chip's own
terms, and pbapp's hand-back jumps to the real slot base. libavr's
far primitives split their RAMPZ/Z asm operands (a page never
crosses 64 KiB, so callers keep a byte and a 16-bit cursor — the
32-bit address folds away; flash_load_far's byte form becomes the
out-RAMPZ+elpm pair avr-libc's pgm_read_byte_far rebuilds per call),
and the host splits reads at 64 KiB boundaries. All ten chips pass
the full suite — the 1284P at 558 B including protocol, relocation,
and the power-fail self-update — with pureboot byte-identical across
generated and reflect modes everywhere, and the original three
chips' images unchanged to the byte (488/502/504).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Device: boot-section detection probes SPMCR beside SPMCSR, the link
picks any hardware USART through the instance-aware lookups (URSEL
chips included), WDRF reads MCUSR-or-MCUCSR, and the >64 KiB shape
lands — word-addressed wire flash (info flag bit 1, page byte 0 means
256, base as a word address), far reads through flash_load_far, a
single 32-bit byte-cursor page walk (the 256-byte page wraps its low
byte exactly), and slot arithmetic in words (the return address
already is one). Host: addresses stay bytes internally and scale at
the wire, the boot-fuse decode becomes a per-signature table (byte
index + BOOTSZ ladder — the m168A's lives in EXTENDED), and the
planner tests pin every chip's ladder plus the word-addressed info
decode. Tests: the device runner serves every mega over the USART pty,
pbapp banners over the right link, the update rehearsal synthesizes
its assumed fuses from the tool's own table, and the PI lint tracks
the renamed info symbol. All six classic-mega/168A targets pass the
full suite (size, PI, planner, protocol, reloc, self-update) at
466–504 B; the 1284P builds await a libavr far-path slimming to make
its 512.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
libavr now exposes avr::hw::db.signature (compile-time, from the ATDF), so the
info block drops its per-chip hardcoded signature() for the db constant. The
loaders are byte-identical across modes with the correct signature, sizes
unchanged (488/502/504).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The host tool takes either form — load_image() parses Intel HEX by extension
and treats anything else as raw bytes — but the build emitted only the HEX, so
the raw path had no artifact behind it. The reloc and update tests each shell
out to objcopy at runtime to produce one for themselves.
add_hex_output becomes add_image_outputs and emits both forms. The .bin is
byte-identical to the plain `objcopy -O binary` those tests generate (-R .eeprom
strips nothing the loaders carry), and decodes equal to the HEX payload — 504 B
at 0x7e00 either way for pureboot. Sizes come out at the flash sizes exactly
(504/510/836/526), so nothing stretches to the .data load address.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The build emits an Intel HEX beside every loader image, but --update-loader
could not consume one: load_image() anchors every image at address zero, and
a loader HEX links at its base, so it decoded to a 32760-byte blob carrying
504 bytes of loader at the end. staging_content() then refused it as "loader
image is 32760 B, the slot holds 512" - an error naming neither the cause nor
the raw .bin the tool wanted instead.
Drop the blank below the base in the update path. The base comes from the
image's own info block rather than the device's, so an image built for
another target survives the slice intact and the preflight still reports it
as another target rather than failing to find an info block at all.
Verified on an ATmega328P: the full self-update flow driven straight from
pureboot_timeout-5s.hex, resident slot byte-for-byte against the image
afterwards, application preserved; both refusal paths unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The host tool was standard-library-only but POSIX-only with it: termios and
select() bound the port layer, and importing termios failed outright on
Windows, so the module could not even load there.
Split Port into PosixPort (unchanged) and a WindowsPort over the Win32 serial
API through ctypes, picked by os.name; every call site keeps the Port name.
kernel32 only, so the standard-library constraint holds.
Windows has no select() for a COM handle, so the read deadlines move into the
driver as COMMTIMEOUTS, re-armed per read: read_available() ends on a gap
longer than a USB-serial latency timer coalesces (16 ms on FTDI parts),
read_exact() on the count or its deadline. Opening asserts DTR and RTS as a
POSIX open does, so a board wiring DTR to reset still pulses it. A failed
configuration closes the handle before raising - a COM handle is exclusive,
and the leak met the next open as "Access is denied". Win32 takes any integer
baud and a driver may accept one its hardware cannot produce (an FT232R
reports back a baud of 3 and keeps the old divisor), so obvious nonsense is
refused where termios' table would have.
Tested against an ATmega328P on COM6: info, fuses, both memories programmed
and verified, session reconnect, hand-over, the loader self-update, and the
write guard on its own slot. test_planner runs on Windows now as well.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
avrdude programs Intel-HEX, not ELF, and the build produced only ELFs — so
flashing a loader to a real chip meant running objcopy by hand. add_hex_output()
hangs a POST_BUILD objcopy on each loader image: the three tsb tiers through
add_tsb_variant, pureboot, and the re-timed pureboot9. .eeprom is dropped, being
its own avrdude update.
It uses the toolchain file's CMAKE_OBJCOPY rather than a hardcoded path, so
every chip preset emits hex, not just the mega. pbapp keeps its ELF alone: the
update test converts it to a raw binary itself, and it is not a flashing target.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
check_pi.py asserts the two link-time facts position independence rests on
(no absolute jmp/call; the info block within the image's first 256 bytes);
the size gates drop to 510 on the tinies for the trampoline word.
New per-chip tests beside the reworked protocol test: the planner units
(programming orders and their recovery properties, the surgery, staging
composition, boot-fuse decode, and the update preflight's error/warning
matrix over synthetic fuse bytes), the relocated-copy sweep (the identical
image installed one slot lower serves the full command set — the PI
acceptance test, and the one that caught the temporary-buffer trap), and
the self-update end-to-end: --update-loader to a re-timed build
(pureboot9, byte-different by PUREBOOT_TIMEOUT alone), then every
power-fail phase killed mid-write, restarted from the runner's flash dump,
and completed by a re-run with the application intact throughout. The mega
rounds run the BOOTRST-unprogrammed profile: the fixture application's 'L'
jump is the application-owned loader entry that profile relies on.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The image now runs from any 512-byte slot with every command intact:
control flow stays PC-relative, the write guard keys on the running slot
(the return-address anchor, computed once), the info block is addressed
from that same anchor as a byte pair (no absolute 16-bit address in the
image), and the application jump is an indirect call through a noipa-
laundered pointer to the absolute entry. 'J' — jump to a wire word
address, the one transfer primitive — replaces 'G': the host knows the
application entry from the info block, and moving between loader copies
needs arbitrary targets. The activation window is a compile-time 8 s
(PUREBOOT_TIMEOUT overrides), counted as a single calibrated poll loop.
A refused page no longer poisons the write-once temporary buffer (a real
silicon trap: the next write would program the drained data): every page
write discards the buffer first — CTPB on the tinies, on the mega the same
RWWSRE store that re-enables RWW after programming. The tinies' post-op
busy-waits go with it: their CPU halts through page erase and write.
488 / 502 / 504 B on t13a / t85 / mega — under the tinies' 510-byte budget,
whose last slot word is the host-managed trampoline: the resident's holds
the application entry, a staging copy's the jump through which an abandoned
update still times out into a loader.
The host tool updates the loader with itself: --update-loader installs the
identical image one slot below the resident, jumps into it, lets it rewrite
the resident, and restores the staging region from a state file — each
phase idempotent off the flash state, resumable after any interruption
(t13a: the staging slot carries the reset vector, written last in and
first out; t85: word 0 redirected around the resident rewrite; mega:
fuse-matrix preflight with a hard BOOTSZ gate and --assume-fuses for
simulators). Application flashing recovers by reset from any interruption:
patched page 0 and trampoline first, erase descending, and a walk-region
refusal behind --force on BOOTRST-below-loader megas.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cancel the GPIO bridge's cycle timers with the state they drive: avr_reset
drops the TX latch, whose falling edge starts a spurious decode before
bridge_reset runs, and the stale sampler then interleaves with the loader's
first real answer through the shared shift state — the first post-reset
replies came back corrupted and the knock retries burned the activation
window into the application.
Wrap the mega's registered flash ioctl to re-dispatch page erases with Z
masked to the page boundary: simavr's PGERS handler erases spm_pagesize
bytes from Z & ~1 (its PGWRT path masks correctly), wiping the neighbouring
page when Z sits past the page start, which hardware permits (§26.8.1).
Model the write-once temporary buffer in the tiny NVM module — silicon
refuses a second load per word until the buffer clears, and a last-write-
wins model masks real firmware bugs.
Optional arguments select the reset vector (the mega's fuse profiles) and a
raw flash image to resume from (power-fail tests re-enter a dumped state).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pureboot.py: reject an empty image file with a clear error instead of
an IndexError deep in the vector-surgery planner; tighten the erase
docstring (order is irrelevant there — every target byte is the same
value, unlike a real flash where page 0 must go last).
pureboot_device.c: the GPIO bridge's bit_cycles used plain truncating
division where the firmware computes its own bit period with
round-to-nearest (uart.hpp: (Clock.hz + Baud.bd/2)/Baud.bd) — one
cycle off per bit on both tinies, harmless in practice but needless
drift against a firmware built to a different constant. Matched
exactly. Also clear the queued-bytes/decode-in-progress bridge state
on the test-only reset signal, so a future reset-mid-transfer scenario
can't feed a freshly reset chip bytes queued for its previous life.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>