49 Commits

Author SHA1 Message Date
a4da885e36 pureboot: the unified autobaud loader, and the hang that settled the decision
Hardware testing found that a lone calibration pulse wedged the autobaud loader:
run() budgeted only the start-edge wait in measure(), and the rx() that read the
knock behind it was unbudgeted, so one stray low pulse held an unattended device
in the loader and the application never ran. Bound the whole activation — an
expired knock budget returns a byte that cannot be the knock, so control falls
back into the budgeted measure() and an idle line boots the app there.

That fix costs ~22 B, which neither version under review could absorb: the pure
one goes 508 -> 530 on the 1284P and the register one 512 -> 534, both over a
512 B slot. Their margin was never spare capacity, it was the space the missing
fix should have occupied. So the choice between them is moot; both are kept for
the record and no longer built.

pureboot_autobaud_uni.cpp replaces them at 464 B. It is pureboot 5: one read
command and one write command over named spaces (G/g, sel8, addr16, n8) instead
of four per-memory bodies, which collapses four transfer loops into one. The
selector's high nibble carries flash's bank, so the shared cursor stays 16 bits
and no command speaks word addresses. Three things fall out of the freed space:
RAM read/write — the missing feature, and with it arbitrary I/O access, since
AVR maps peripherals into the data space; host-issued SPM, so W's hardcoded
erase/write/RWW tail becomes three writes to a space and any SPM operation is
reachable; and W on the same selector-and-address decode as everything else.

Strictly pure throughout: no inline asm, no global register variable, and no
GPIOR either — the unit lives in a .noinit static, so the loader claims no chip
resource and the chips without GPIOR stop being a special case.

pureboot.py speaks both generations, keyed on the version, so the fixed-baud
path is untouched; --peek/--poke reach the new data space. pbautobaud.py adds a
RAM round-trip and a regression for the hang: a lone pulse must still let the
app boot. All 37 chips plus the 12-preset reflect spot set build and size-test
green, 444-466 B, worst case 46 B under budget. Sim suites 100%: 1284P 17/17,
328P 23/23. Only real-hardware acceptance remains (pureboot/autobaud.md).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 00:48:54 +02:00
335e494a31 pureboot: autobaud host support and simavr end-to-end for both variants
pureboot.py --autobaud sends the 0xC0 calibration pulse and a single knock at
the host's chosen baud, reads the slimmed info block, and derives the full
geometry from the signature (AUTOBAUD_GEOMETRY, a table over every pureboot
chip). Everything downstream — flash, EEPROM, fuses, hand-over, verify — is the
fixed-baud path unchanged; the dropped write guard is host-transparent.

test/pbautobaud.py drives each variant over the GPIO⇄pty software-UART bridge
through the calibration handshake and a flash + EEPROM + fuse round-trip
cross-checked against the simulator's ground-truth memory, then repeats at
double the F_CPU with the same binary — the clock-agnostic property autobaud
exists for. Wired as pureboot.autobaud_pure/reg on the near-flash 328P and the
word-addressed 1284P. A wrong measured unit fails the flash/verify, so the test
also pins the codegen-coupled calibration constant against a toolchain bump.

Both variants green in sim on both chips at two clocks each; the fixed-baud
suite is unaffected. Only real-hardware acceptance on an RC part remains
(pureboot/autobaud.md).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 22:45:22 +02:00
799709efcf pureboot: autobaud variant, two versions for review
Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one
clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments
no longer need a per-clock build.

Two source files, differing only in the write-guard/purity tradeoff:
pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch
registers, running-slot write guard dropped, 508 B on the 1284) stays strictly
pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard
kept, 512 B) keeps every feature at the cost of that single GRV. Both fit
512/510 on all 37 chips and share two licensed simplifications: a slimmed info
block (version + signature; the host derives geometry from the chip database)
and a single-byte activation knock.

pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that
set the target, and the compiler-knob path to it. Size-tested on every chip via
pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the
host calibration handshake, and real-hardware acceptance remain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 22:15:38 +02:00
7ae80087b3 docs: the watchdog-lockout and EEPROM-wrap gotchas
A sticky WDRF diverts every reset past the activation window (deliberate, so
an app can reboot instantly, at the cost of a possible lockout); an EEPROM
address past E2END wraps onto low EEPROM (the host bounds it, not the loader).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 02:01:33 +02:00
b477f53ca5 pureboot: the info block reads as a table, one wire byte per line
clang-format bin-packs braced lists to the column limit, collapsing the
'b' reply's byte layout into dense rows. A minimal clang-format-off span
keeps each wire byte on its own line, where the layout is legible against
the protocol. Whitespace only; image byte-identical.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 01:19:09 +02:00
84d3f679c2 style: clang-format the W-fix line
Layout only, byte-identical output.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 01:13:51 +02:00
b5020a1e20 pureboot 4: the loader carries its fixes' identity
The unaligned-W and U2X-hand-over fixes change the loader's observable
on-wire behavior, and the --stay reconnect fix changes the host tool, so
both move: loader version 3 -> 4, tool VERSION 2 -> 3. The protocol and info
block are unchanged, so OLDEST_LOADER stays 1.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 01:09:38 +02:00
1dbf0089d6 pureboot: the info block is what proves a knock landed
A prompt byte alone does not: one left over from a previous session can
still be in the pipeline while the port opening resets the device into a
fresh window, where the bare command that follows is discarded. Each
attempt is now the whole handshake, retried until the block comes back.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 23:49:37 +02:00
39dfe40dbf pureboot: W addresses a page, not a word in it
The in-page bits of a W address are dropped so the fill always walks from
the page base; the wire contract is one page of data for any address
inside it, on both the byte- and the word-addressed path. +2 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 23:43:41 +02:00
8470ce3da0 test: the exhaustive clock x baud x backend size matrix
Every plausible oscillator against every rate it reaches against every
backend, on one chip per size-bearing class, under --full only. The baud
ladder becomes a reachability predicate the enumeration filters on, so an
unreachable point drops out instead of aborting the configure.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 23:32:28 +02:00
d634741015 pureboot 3: a 512-byte slot on every chip, the 1284s included
The word-addressed 1284s were the one family deploying in a 1 KiB slot,
because the far-flash machinery (ELPM reads, RAMPZ page commands, a
word-addressed wire) did not fit 512 B. It does now: 478 B stock, 494 B in
the heaviest configuration the build can produce. They take the 644s'
geometry, where the smallest boot section holds the resident slot and its
staging slot together. The loader's version goes to 3; the host tool did
not change, so its own version stays 2 and only the window it speaks
widens.

Most of the saving is one restructure. The info block and a flash read are
the same act, so giving all four streamed commands one address-and-count
path leaves exactly one call site for the flash streamer: it inlines into
the never-returning command loop and its 24-bit cursor stops being saved
and restored around every transmit. Around it, the ack byte moved out of
line, the wire's byte pair is bit_cast into the word it already is, the
fuse loop ends on its count, the info block's in-slot offset is taken as
the one-byte relocation it is, and -fno-expensive-optimizations gives way
to -fno-move-loop-invariants -fno-tree-ter. Every chip shrank 14-18 B.

The size matrix grew the axes it was missing: the USART1 instance across
the whole clock ladder, and the shape a slow baud gives a software UART —
past 255 delay iterations libavr takes the 16-bit delay loop, which the
ladder default never selects and which was 4 B over the 1284's slot the
first time it was built.

The protocol fixture stopped deriving the loader entry from the flash
size; on the 1284s it had been jumping a slot low and reaching the loader
only because erased flash walked it up.

Docs and comments were consolidated across the port in the same pass: the
README carries a per-chip size table instead of prose, and prose that
restated the code is gone — 190 lines, no behaviour with it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-22 20:49:23 +02:00
ab842c8d99 pureboot: a version number for the loader, not for the protocol
The info block's third byte was a protocol version that never moved in the
loader's lifetime. It is the pureboot version now, and this change is
version 2: the one number that says what a deployed loader is. Every loader
already in the field answers 1.

The protocol keeps no number of its own — a pureboot version implies it, and
the host tool is what holds that map. pureboot.py states the loader-version
window it speaks (OLDEST_LOADER/NEWEST_LOADER; a version that changes the
protocol becomes the new floor there), so a loader newer than the tool is
refused by name rather than decoded on the assumption nothing moved, while an
older one is read, identified and installed like any other. The tool carries
its own version, free to drift from the loader's: --version prints it and the
window, --info leads with the device's, --update-loader names the version it
installs.

Tests: the planner unit pins the window — every version in it decodes, one
above it is refused, an older loader's image is still found — and the live
suite pins the built loader against the tool beside it, so a bump that reaches
only one of them fails. The image is byte-identical to the previous build but
for that byte.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-22 17:20:34 +02:00
a58b59c60f docs: correct the size figures for the configured variants
The headline numbers described the stock deployments but claimed "every
configured variant of them", which the matrix contradicts: choosing the
software UART where the chip has a USART costs 8-46 B, so the megas reach
460-462 rather than 452, and the 1284s' software-serial build is 546 B —
inside their 1 KiB boot sector, but not inside 512.

Also names the actual tightest chip. The 1284 looks like it at 506, but it
deploys in 1 KiB with 478 B spare; against its own budget the ATmega328P
has the least room, 50 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 17:08:24 +02:00
ceacb61ba1 pureboot: no SPM buffer discard, the host repairs instead
The temporary page buffer is write-once per word, so a page filled over
one an earlier writer left dirty programs the stale words. The same
datasheet clause carries the cure: the buffer auto-erases after a page
write (§26.2.1; §19.2 on the tinies), so the corruption clears itself by
happening, and rewriting the page programs correctly.

The loader therefore clears the buffer nowhere. The tinies' CTPB and the
m48s' RWWSRE discard are gone; the boot-sectioned megas keep only the
trailing RWWSRE they need anyway to re-enable the RWW section for
read-back, which discards the buffer as a side effect and keeps them off
the path entirely. 434 B on the tiny13s, 438-442 on the tiny25/45/85,
430 on the m48s; the megas are unchanged, the 1284s still 506.

The host takes over the guarantee: a flash page that reads back wrong is
rewritten up to RETRIES times before the run stops. Both read-back paths
repair — verify_pages for programming, and write_differing, which is the
loader-update path where a page left wrong is a half-written loader slot.
That one is not hypothetical: deleting the discard made attiny85
pureboot.rehome fail deterministically there, the only flow still
assuming the old contract.

Protocol-visible, so README's W command says it: one W may program the
wrong bytes after a refused page, or after an application that
self-programmed entered without a reset, and a host that programs without
reading back cannot trust it.

Tests: pureboot.dirty drives the case the loader declines to guard — the
fixture application dirties every buffer word and jumps in with no reset
(hardware forbids that on a boot-sectioned mega, but simavr dispatches SPM
from anywhere, which is what makes it constructible) — and asserts a bare
verify sees the corruption, the repairing verify fixes it in one rewrite,
and it stays fixed. pbreloc asserts the same shape after a refusal.
test_planner covers the bound against a fake device: one bad write
repaired in a single rewrite, a page that never comes good stopping after
exactly three.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 16:29:27 +02:00
dd7490a3df build: generated presets and a check entry point
tools/make_presets.py emits the uniform pipeline the hand-grown file had
drifted from — generated configure/build/test presets and workflows for
all 37 chips, reflect configure/build for libavr's 12-chip spot set —
and tools/check.sh runs every chip's workflow (--full adds the reflect
spot) as the port's gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
bc24b63e65 host: one fact per line, progress bars, --verbose
--info prints the decoded info block field by field and --fuses each
byte on its own line plus the BOOTSZ/BOOTRST meaning on boot-sectioned
megas. Transfers that take wire time draw a transient progress bar on
stderr when it is a tty — logs, pipes and the tests see only the
summary lines. -v/--verbose narrates decisions: knock counts, the
programming plan, update state handling and per-phase page counts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
332b244ac3 pureboot: every deployment axis is a build parameter
Clock, baud, serial backend (hardware USART 0/1 or the software UART on
any pins) and the activation window all resolve through one CMake
function, pureboot_add_loader() in pureboot/CMakeLists.txt — the unit a
downstream project consumes. The default baud is the fastest standard
rate within 2.5 % (the same best-divisor search libavr's solver runs),
gated on software builds by the polled receiver's 100-cycles-a-bit
floor; every explicit pick is re-checked by the compile's static asserts.

The size matrix builds each axis that can move the image — backend x
clock ladder x USART instance, per chip — against the slot budget, and
two nondefault deployments run the whole protocol suite live: the 328P
on its shipped 1 MHz fuses over software serial on TX=PB1/RX=PB5
(pureboot.custom), and the 644A over USART1 (pureboot.usart1). The sim
runner takes -l to bridge any link, paces a fully quiet bridge toward
real time (a free-running 8 M-cycle window loses the reset-race knock),
and the fixture application speaks the deployment it is built for.

The loader itself shed bytes on the way: the return-address high byte
spelled through byteswap (the double swap folds to the one-byte pick),
the info-block address composed instead of bit_cast, and libavr's new
polled-UART helpers replacing the port's uart::detail reaches. Every
combination fits: 458-506 B across the megas' whole matrix, 470-484 B
on the tinies, 556-562 B in the 1284s' 1 KiB slot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
61b687120c pureboot: a loader in the staging slot is the staging copy — leave it there
The update flow's first step wrote the staging content over whatever the
staging slot held; with a loader running right there (programmed by hand
onto erased flash), that write met the copy's own running-slot guard on
the composed through-word and the tool stopped at its verify — although
the copy is exactly an installed staging copy, able to stream the new
resident like any other. The install is now skipped when the slot holds a
complete loader: its info block where every image carries it, matching
the device's byte for byte, and the slot unchanged since the update began
(the state file's snapshot) — so a resumed half-written install still
differs from its snapshot and takes the install path, which completes it.
pbrehome gains the staging-slot position (an older build at stage
streaming a newer resident in); the README's wrong "cannot re-home from
the staging slot" claim is corrected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 21:55:15 +02:00
ef58d5d363 pureboot: artifact roles, wrong-chip refusal, and misplaced-loader re-homing
The README's deployment section now says what each build artifact is for:
the .hex is the programmer artifact (self-addressed into the top slot),
the .bin the self-update image — bare slot bytes a programmer would put
at address 0, where a boot-sectioned mega cannot even heal itself (SPM
only runs from the boot section) but a patched-vector chip runs the
position-independent copy and re-homes a build through the ordinary
--update-loader flow: the staging install and the word-0 redirect both
execute outside page 0's slot, so the running-slot guard never blocks it.
pbrehome.py is the acceptance test (misplaced at 0, guard intact,
re-home, app flash over the stale copy, banner); the staging slot is the
one position that cannot re-home itself, documented. The preflight's
wrong-chip refusal and loader_image's handling of padded images (peeled
to the slot content by the embedded base) are documented and the padded
case pinned in the planner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 21:18:10 +02:00
03bb1f56cf pureboot: every libavr chip — 37 loaders, the m48 class, the 644 geometry
The chip table becomes family blocks covering all 37 targets. The m48s
are a new deployment class: no boot section, so the tiny profile spoken
over the hardware USART — host-patched reset vector, trampoline
hand-over, a 510-byte budget (474 B built), no fuse preflight — while
their RWWSRE store stays the buffer discard (Atmel-8271 §26.2); the
device keys the patch flag and the CPU-halt waits on the curated
boot-section capability and the discard on the RWWSRE bit itself. The
644s' 64 KiB is exactly the 16-bit byte space: plain LPM, byte wire
addresses, 498 B in a 512-byte slot — and their 1 KiB minimum boot
section holds the resident and staging slots together, so self-update
needs no fuse step (the update test's slot pick now keys word-flash on
base >= 64 KiB; base + slot merely touching the boundary stays
byte-addressed). The 1284 joins the 1284P's word-addressed 1 KiB slot at
558 B. BOOT_FUSE gains every boot-sectioned family's ladder and fuse
byte; the planner exercises them all. The sim scaffolding keys
patch-vector-ness instead of the atmega name prefix, the fixture app
picks its clock by family (the tiny25/45/13 builds surfaced the 16 MHz
fallthrough as garbled banners), and the runner's wrapped flash ioctl
performs the m48 discard simavr's no-RWW cores turn into a stray buffer
fill. Sizes across the fleet: 466-504 B megas, 474 B m48s, 498 B 644s,
488-502 B tinies, 558 B 1284s — every chip passing
size/pi/planner/protocol/reloc/update.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 18:53:04 +02:00
6a5e313530 pureboot: correct the 1284P deployment profiles in the README
The deployment section claimed the 1284P has no standalone profile, runs
BOOTSZ = 512 words always, and self-updates with no fuse change, reset
landing at 0x1f800 — internally contradictory (a 512-word section starts
at 0x1fc00, and the section holding both 1 KiB slots is 1024 words) and
contradicted by update_preflight, which refuses a self-update unless the
boot section covers two slots. The text described a 512-byte-slot
geometry this chip's loader cannot have. In truth the 328P profile table
maps onto the 1284P doubled: standalone = 512 words (the smallest
section is exactly the 1 KiB slot, reset at the loader base), self-update
= 1024 words with the loader-first reset walking the staging slot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 16:48:46 +02:00
a28fccf475 pureboot: the 1284P rides a 1 KiB slot — its own boot-sector minimum
The far machinery (ELPM reads, RAMPZ page commands, wire-word math)
costs ~46 B over the m328P's 504, and the tsb-calibrated C++-to-asm
gap says no implementation of this feature set reaches 512 on this
chip — a boundary its hardware does not have anyway: the 1284P's
smallest boot sector is 1 KiB. The slot therefore becomes
per-geometry (512 B, or 1 KiB past 64 KiB), which the host derives
from the word-addressing flag; slot arithmetic unifies (the index is
the wire high byte with its low bit dropped in either unit), the
update preflight demands a two-slot boot section in the chip's own
terms, and pbapp's hand-back jumps to the real slot base. libavr's
far primitives split their RAMPZ/Z asm operands (a page never
crosses 64 KiB, so callers keep a byte and a 16-bit cursor — the
32-bit address folds away; flash_load_far's byte form becomes the
out-RAMPZ+elpm pair avr-libc's pgm_read_byte_far rebuilds per call),
and the host splits reads at 64 KiB boundaries. All ten chips pass
the full suite — the 1284P at 558 B including protocol, relocation,
and the power-fail self-update — with pureboot byte-identical across
generated and reflect modes everywhere, and the original three
chips' images unchanged to the byte (488/502/504).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:54:40 +02:00
92503fbb2d pureboot: the classic megas and the word-addressed 1284P groundwork
Device: boot-section detection probes SPMCR beside SPMCSR, the link
picks any hardware USART through the instance-aware lookups (URSEL
chips included), WDRF reads MCUSR-or-MCUCSR, and the >64 KiB shape
lands — word-addressed wire flash (info flag bit 1, page byte 0 means
256, base as a word address), far reads through flash_load_far, a
single 32-bit byte-cursor page walk (the 256-byte page wraps its low
byte exactly), and slot arithmetic in words (the return address
already is one). Host: addresses stay bytes internally and scale at
the wire, the boot-fuse decode becomes a per-signature table (byte
index + BOOTSZ ladder — the m168A's lives in EXTENDED), and the
planner tests pin every chip's ladder plus the word-addressed info
decode. Tests: the device runner serves every mega over the USART pty,
pbapp banners over the right link, the update rehearsal synthesizes
its assumed fuses from the tool's own table, and the PI lint tracks
the renamed info symbol. All six classic-mega/168A targets pass the
full suite (size, PI, planner, protocol, reloc, self-update) at
466–504 B; the 1284P builds await a libavr far-path slimming to make
its 512.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 12:00:19 +02:00
8ba023ed6b pureboot: take the device signature from the chip database
libavr now exposes avr::hw::db.signature (compile-time, from the ATDF), so the
info block drops its per-chip hardcoded signature() for the db constant. The
loaders are byte-identical across modes with the correct signature, sizes
unchanged (488/502/504).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:38:20 +02:00
443fd39d31 build: emit a raw .bin beside the HEX for every loader image
The host tool takes either form — load_image() parses Intel HEX by extension
and treats anything else as raw bytes — but the build emitted only the HEX, so
the raw path had no artifact behind it. The reloc and update tests each shell
out to objcopy at runtime to produce one for themselves.

add_hex_output becomes add_image_outputs and emits both forms. The .bin is
byte-identical to the plain `objcopy -O binary` those tests generate (-R .eeprom
strips nothing the loaders carry), and decodes equal to the HEX payload — 504 B
at 0x7e00 either way for pureboot. Sizes come out at the flash sizes exactly
(504/510/836/526), so nothing stretches to the .data load address.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 23:17:59 +02:00
08a0d77109 pureboot: take the emitted loader HEX for --update-loader
The build emits an Intel HEX beside every loader image, but --update-loader
could not consume one: load_image() anchors every image at address zero, and
a loader HEX links at its base, so it decoded to a 32760-byte blob carrying
504 bytes of loader at the end. staging_content() then refused it as "loader
image is 32760 B, the slot holds 512" - an error naming neither the cause nor
the raw .bin the tool wanted instead.

Drop the blank below the base in the update path. The base comes from the
image's own info block rather than the device's, so an image built for
another target survives the slice intact and the preflight still reports it
as another target rather than failing to find an info block at all.

Verified on an ATmega328P: the full self-update flow driven straight from
pureboot_timeout-5s.hex, resident slot byte-for-byte against the image
afterwards, application preserved; both refusal paths unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 22:46:50 +02:00
260ab2e4e3 pureboot: drive the serial port on Windows too
The host tool was standard-library-only but POSIX-only with it: termios and
select() bound the port layer, and importing termios failed outright on
Windows, so the module could not even load there.

Split Port into PosixPort (unchanged) and a WindowsPort over the Win32 serial
API through ctypes, picked by os.name; every call site keeps the Port name.
kernel32 only, so the standard-library constraint holds.

Windows has no select() for a COM handle, so the read deadlines move into the
driver as COMMTIMEOUTS, re-armed per read: read_available() ends on a gap
longer than a USB-serial latency timer coalesces (16 ms on FTDI parts),
read_exact() on the count or its deadline. Opening asserts DTR and RTS as a
POSIX open does, so a board wiring DTR to reset still pulses it. A failed
configuration closes the handle before raising - a COM handle is exclusive,
and the leak met the next open as "Access is denied". Win32 takes any integer
baud and a driver may accept one its hardware cannot produce (an FT232R
reports back a baud of 3 and keeps the old divisor), so obvious nonsense is
refused where termios' table would have.

Tested against an ATmega328P on COM6: info, fuses, both memories programmed
and verified, session reconnect, hand-over, the loader self-update, and the
write guard on its own slot. test_planner runs on Windows now as well.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 22:46:14 +02:00
e4aaf5bc62 build: emit an Intel-HEX beside every loader image
avrdude programs Intel-HEX, not ELF, and the build produced only ELFs — so
flashing a loader to a real chip meant running objcopy by hand. add_hex_output()
hangs a POST_BUILD objcopy on each loader image: the three tsb tiers through
add_tsb_variant, pureboot, and the re-timed pureboot9. .eeprom is dropped, being
its own avrdude update.

It uses the toolchain file's CMAKE_OBJCOPY rather than a hardcoded path, so
every chip preset emits hex, not just the mega. pbapp keeps its ELF alone: the
update test converts it to a raw binary itself, and it is not a flashing target.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 21:57:26 +02:00
653602058a pureboot: PI lint, tightened gates, and the self-update test suite
check_pi.py asserts the two link-time facts position independence rests on
(no absolute jmp/call; the info block within the image's first 256 bytes);
the size gates drop to 510 on the tinies for the trampoline word.

New per-chip tests beside the reworked protocol test: the planner units
(programming orders and their recovery properties, the surgery, staging
composition, boot-fuse decode, and the update preflight's error/warning
matrix over synthetic fuse bytes), the relocated-copy sweep (the identical
image installed one slot lower serves the full command set — the PI
acceptance test, and the one that caught the temporary-buffer trap), and
the self-update end-to-end: --update-loader to a re-timed build
(pureboot9, byte-different by PUREBOOT_TIMEOUT alone), then every
power-fail phase killed mid-write, restarted from the runner's flash dump,
and completed by a re-run with the application intact throughout. The mega
rounds run the BOOTRST-unprogrammed profile: the fixture application's 'L'
jump is the application-owned loader entry that profile relies on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:57 +02:00
57d1c009cc pureboot: one position-independent binary — its own staging loader
The image now runs from any 512-byte slot with every command intact:
control flow stays PC-relative, the write guard keys on the running slot
(the return-address anchor, computed once), the info block is addressed
from that same anchor as a byte pair (no absolute 16-bit address in the
image), and the application jump is an indirect call through a noipa-
laundered pointer to the absolute entry. 'J' — jump to a wire word
address, the one transfer primitive — replaces 'G': the host knows the
application entry from the info block, and moving between loader copies
needs arbitrary targets. The activation window is a compile-time 8 s
(PUREBOOT_TIMEOUT overrides), counted as a single calibrated poll loop.

A refused page no longer poisons the write-once temporary buffer (a real
silicon trap: the next write would program the drained data): every page
write discards the buffer first — CTPB on the tinies, on the mega the same
RWWSRE store that re-enables RWW after programming. The tinies' post-op
busy-waits go with it: their CPU halts through page erase and write.

488 / 502 / 504 B on t13a / t85 / mega — under the tinies' 510-byte budget,
whose last slot word is the host-managed trampoline: the resident's holds
the application entry, a staging copy's the jump through which an abandoned
update still times out into a loader.

The host tool updates the loader with itself: --update-loader installs the
identical image one slot below the resident, jumps into it, lets it rewrite
the resident, and restores the staging region from a state file — each
phase idempotent off the flash state, resumable after any interruption
(t13a: the staging slot carries the reset vector, written last in and
first out; t85: word 0 redirected around the resident rewrite; mega:
fuse-matrix preflight with a hard BOOTSZ gate and --assume-fuses for
simulators). Application flashing recovers by reset from any interruption:
patched page 0 and trampoline first, erase descending, and a walk-region
refusal behind --force on BOOTRST-below-loader megas.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:43 +02:00
06bb06d994 pureboot: harden the sim device runner
Cancel the GPIO bridge's cycle timers with the state they drive: avr_reset
drops the TX latch, whose falling edge starts a spurious decode before
bridge_reset runs, and the stale sampler then interleaves with the loader's
first real answer through the shared shift state — the first post-reset
replies came back corrupted and the knock retries burned the activation
window into the application.

Wrap the mega's registered flash ioctl to re-dispatch page erases with Z
masked to the page boundary: simavr's PGERS handler erases spm_pagesize
bytes from Z & ~1 (its PGWRT path masks correctly), wiping the neighbouring
page when Z sits past the page start, which hardware permits (§26.8.1).
Model the write-once temporary buffer in the tiny NVM module — silicon
refuses a second load per word until the buffer clears, and a last-write-
wins model masks real firmware bugs.

Optional arguments select the reset vector (the mega's fuse profiles) and a
raw flash image to resume from (power-fail tests re-enter a dumped state).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:10 +02:00
9a9a7ef0e8 pureboot: review-pass fixes to the host tool and device runner
pureboot.py: reject an empty image file with a clear error instead of
an IndexError deep in the vector-surgery planner; tighten the erase
docstring (order is irrelevant there — every target byte is the same
value, unlike a real flash where page 0 must go last).

pureboot_device.c: the GPIO bridge's bit_cycles used plain truncating
division where the firmware computes its own bit period with
round-to-nearest (uart.hpp: (Clock.hz + Baud.bd/2)/Baud.bd) — one
cycle off per bit on both tinies, harmless in practice but needless
drift against a firmware built to a different constant. Matched
exactly. Also clear the queued-bytes/decode-in-progress bridge state
on the test-only reset signal, so a future reset-mid-transfer scenario
can't feed a freshly reset chip bytes queued for its previous life.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:45:19 +02:00
4acf358dda pureboot: gitignore python bytecode cache
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:43:56 +02:00
11c10f986d pureboot: stop tracking the python bytecode cache
A stray __pycache__/*.pyc from a local test run got swept into the
previous commit's git add. Untracked and gitignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:43:36 +02:00
f587f0a26e pureboot: host tool and end-to-end protocol tests, all three chips
pureboot.py (Python stdlib only): images as raw binary or Intel HEX,
flash and EEPROM programming with read-back verify, erase composites,
fuse and info readout, activation-timeout configuration, and the
tinies' reset-vector surgery — the trampoline word below the loader,
page 0 written last.

The test spawns a simavr device (pureboot_device.c) — the mega's USART
as a pty; on the tinies a cycle-timed GPIO<->pty bridge for the polled
software UART plus the NVM module simavr's tiny cores lack (their SPM
opcode ioctls into a void and silently does nothing) — and drives it
with the real tool: knock from reset (erased-flash walk on the tinies),
program and verify both memories, timeout write, session reconnect, an
external reset through the patched vector, hand-over, and the fixture
application's banner. Results are cross-checked against ground-truth
memory dumps and an independent decode of the surgery's rjmp words,
red-verified against a sabotaged encoder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 05:54:15 +02:00
27c768a102 pureboot: the device — one pure C++ source, 512 bytes, every chip
No inline assembly, no global register variables; libavr does the
datasheet work. The device speaks primitives — flash read/page-program,
EEPROM read/write, fuse read, info block, EEPROM-resident activation
timeout, hand-over — and verify, erase, reset-vector surgery, and
timeout configuration live in the host tool. 490 B on the ATtiny13A,
510 B on the ATtiny85, 484 B on the ATmega328P, each linked into the
top 512 bytes of flash; per-chip size tests gate all three.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 05:33:28 +02:00
d989b5b8cd tsb: third size pass — restructure to the oracle's shape
The second pass concluded the 168 B tricks->asm gap was per-call ABI
cost. Most of it was structure. Rebuilt around the oracle's own shape —
argless noinline primitives over a whole-loader call-saved register
protocol (g_addr in Y, count r16, window r7, direction latch r6), a
top-down erase_below whose loop tests against zero and hands callers
g_addr = 0 for free, bounded rx everywhere (a silent host unwinds to
the app from any state, as the oracle does), and a named tsb_app entry
that --pmem-wrap-around=32k relaxes to the wrapped rjmp:

  tsb_asm    510 B in the 512 B section (oracle: 500), C++ except rx
             and the page-store loop — the two routines whose remaining
             cost is the calling convention itself (~30 asm lines, was
             ~280)
  tsb_tricks 526 B, no assembly at all (was 666)
  tsb_pure   836 B, still one readable function per command (was 842)

Every g_* update placement works around a GCC 16.1 wrong-code bug
(stores into global register variables deleted when only callees read
them — repro and rules in libavr dev/lessons.md). Also fixes two
latent hardware bugs all earlier tiers carried, masked by simavr's
zeroed register file: the crt-less entries never established
__zero_reg__ = 0, and the direction latch was read before written —
power-on registers are undefined.

All tiers full oracle feature parity, protocol tests green in both
libavr modes, .text byte-identical across modes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:00:27 +02:00
314e422b19 tsb: beat the first-pass size floors (tricks 666, pure 842)
tricks 778->666: always_inline every single-call handler into the
[[noreturn]] reset entry (which pays no prologue, so their push/pop of
call-saved registers vanishes), walk the page pointer in Y (adiw, base
recovered as g_addr-page) instead of recomputing Z=base+offset, bring
the UART up in the two registers that are not already at their reset
value, and seed the activation counter as __uint24.

pure 896->842: TU-local internal linkage (proper hygiene, and it lets
the compiler inline the one-call handlers), a byte-wide activation
count, __uint24 timeout. Still one readable function per command.

asm unchanged at 498: its C++-expressible parts are already C++; the
core stays asm (the 666 B all-tricks tier is 168 B over — per-call ABI
tax, not a feature). All three cross-mode byte-identical, protocol green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 20:15:49 +02:00
3e0717ef06 tsb: drive each tier to its size floor
asm 502->498 B (below the oracle's 500): the stack bring-up moves to plain C++,
and a register is reserved for the config-page high byte instead of reloading it
at each app-flash-boundary compare. tricks 808->778 B: shared erase/rww helpers
plus the libavr half-duplex W1C fix. pure 950->896 B and no SRAM: streams
rx->SPM/EEPROM instead of staging a 128 B page buffer. All three keep full oracle
feature parity and stay byte-identical across modes; protocol tests (round-trip +
password + emergency erase) green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 18:47:38 +02:00
812be910c1 tsb: document the three tiers at full parity in the build file
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:50:18 +02:00
34529c8031 tsb: protocol test covers the password gate and emergency erase
Each scenario group now runs on its own freshly-reset device: the round-trip
on a blank config page, plus a password-config device that must be sent the
password after the knock to activate, and an emergency-erase device where a
0-byte + two confirms wipes flash, EEPROM and the config page (verified by
reading all three back as 0xff). All three tiers pass every group.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:49:20 +02:00
06bd56af50 tsb: pure and tricks tiers reach full oracle feature parity
Both tiers gain the features the asm tier already carries — one-wire
half-duplex (via libavr's new .half_duplex), the config-page activation
timeout, and emergency erase (password \0 + double-confirm wipes flash,
EEPROM and the config page) — on top of the watchdog bail, password gate and
config/flash/EEPROM read-write they already had. pure stays idiomatic
(flash_table info block, one function per command) at 950 B; tricks keeps its
compiler trickery (call-saved global-register page walk, unified runtime-flag
paths pinned noinline/noclone, streaming stores, arithmetic command decode)
at 808 B. Both byte-identical across generated and reflect modes; the size
gradient across the three tiers is now 502 / 808 / 950 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:47:21 +02:00
75bb84fa3e tsb: asm tier reaches full oracle feature parity at 502 B
Rewrite the inline-asm tier so it matches the hand-written fixed-baud oracle's
feature set inside the 512 B boot section: watchdog-reset bail, one-wire
half-duplex (RXEN/TXEN toggled per direction, TX turnaround guard),
config-page activation timeout, the password gate (wrong byte hangs draining
the UART), emergency erase (password \0 + double-confirm wipes flash, EEPROM
and the config page), and config/flash/EEPROM read-write. Every geometry,
baud and info-block constant comes from libavr consteval; only the dense
control flow is hand-written. 502 B, byte-identical across generated and
reflect modes.

Test harness: seed the config page from TSB_CONFIG so the password and
emergency-erase paths are exercisable, and clear simavr's AVR_UART_FLAG_POLL_
SLEEP — a host-CPU-saving usleep(1)-per-idle-poll hack that models no hardware
and paces a one-wire loader (which releases TX between bytes) in real time,
distorting protocol timing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:23:16 +02:00
778b8a0100 tsb: vendor the fixed-baud assembly oracle as the size/feature bar
The Seed Robotics native-UART fixed-baud TinySafeBoot (GPLv3), reference
only — not built. Assembles to 500 B with the full feature set, proving
≤512 B and full feature parity are simultaneously reachable. Also drops the
stale empty stk500v2/ leftover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:29:58 +02:00
48d11fe8e3 tsb: use the named register surface
Direct register access now reads through the named surface
(hw::mcusr::wdrf.test(), hw::ucsr0b::write(...)) instead of the string form,
matching how libavr itself is written. Zero-overhead: pure 740 B, tricks 658 B,
asm 508 B unchanged, all byte-identical across modes, protocol green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:55:01 +02:00
bb4784ae0c tsb: refactor the pure tier onto libavr sugar
The showcase tier now leans on the helpers it fed back instead of reaching under
them: the info block is an avr::flash_table (no raw [[gnu::progmem]]), a page is
filled with spm::fill(addr, span) (no hand-packed lo|hi<<8 loop), and the
WDT-reset bail reads field<"MCUSR","WDRF">::test() (no read() & {}(1).value).

Zero-overhead throughout: .text stays 740 B, byte-identical across generated and
reflect modes, protocol test green. The info block streams through the existing
address-based send_flash rather than a range-for over the flash_table — the
range-for is a distinct loop that cannot share the loader's one flash streamer,
so it would add 14 B for no functional gain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:33:53 +02:00
d07586652d tsb: slim the port branch to the libavr reimplementation
main carried the whole pre-libavr tree beside the port: the other-bootloader
directories (blink, stk500v2), the Atmel Studio solution/project, and — dead in
the tsb dir itself — four submodule links to the superseded io/flash/uart/type
libraries the libavr sources never include. None are build inputs; CMake drives
the three variants through FetchContent. master keeps the full legacy tree
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:12:56 +02:00
3a2f2a23c2 tsb: drop the local -O3 strip, now handled by the libavr toolchain
The -O3 leak is fixed upstream (cmake/release-os.cmake via CMAKE_PROJECT_INCLUDE),
so the port no longer needs its own string(REPLACE); a Release build is -Os
through the toolchain file. Verified: all three variants build at their sizes
(508/658/740) and pass the size + protocol ctest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 10:52:13 +02:00
4fa058653e tsb: reimplement TinySafeBoot on libavr in three size tiers
The native-UART fixed-baud TinySafeBoot protocol, ported onto libavr as a
crt-free boot-section loader, in three variants that trade clarity for size:

  tsb_pure   740 B  idiomatic C++: SRAM page buffer, separate flash/EEPROM
                    leaves, shared framing; the polled `unused` guard posture.
  tsb_tricks 658 B  unified runtime-flag paths (noinline/noclone), call-saved
                    global-register page walk — attributes only, no asm.
  tsb_asm    508 B  streaming store + hand-rolled UART/SPM/EEPROM/erase loops;
                    fits the 512 B boot section (BOOTSZ=11). Trims the optional
                    password gate and WDT-reset bail — unreachable in C++ with
                    both (hand-asm is ~15 % denser). Tiers 1-2 keep them and
                    live in the 1 KB section they fit.

All three are .text byte-identical across libavr's generated and reflect modes.
The CMake build strips the leaked -O3 (a Release build is silently -O3, not the
-Os this loader is measured against) and gates each variant's size against its
section. A simavr harness (test/device.c + test/tsbtest.py) drives the real wire
protocol over a pty and flashes the device; the size and protocol tests run in
ctest. Verified byte-for-byte against the reference tsbloader_adv (C#/mono):
activate, read info, flash write + verify.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 05:00:51 +02:00
26 changed files with 1985 additions and 2684 deletions

3
.gitmodules vendored Normal file
View File

@@ -0,0 +1,3 @@
[submodule "libavr"]
path = libavr
url = ../libavr.git

View File

@@ -8,6 +8,9 @@ include(FetchContent)
if(NOT LIBAVR_ROOT AND DEFINED ENV{LIBAVR_ROOT})
set(LIBAVR_ROOT $ENV{LIBAVR_ROOT})
endif()
if(NOT LIBAVR_ROOT)
set(LIBAVR_ROOT ${CMAKE_CURRENT_SOURCE_DIR}/libavr)
endif()
if(LIBAVR_ROOT)
FetchContent_Declare(libavr SOURCE_DIR ${LIBAVR_ROOT})
else()
@@ -94,10 +97,6 @@ endfunction()
# tsb_pure — pure idiomatic libavr, one function per command, TU-local
# (internal linkage), streaming (no SRAM page buffer): 836 B in
# the 1 KB section.
# tsb_policy — the policy floor: pureboot's rules (no asm, no register
# variables) with every pureboot lesson applied. 638 B in the
# 1 KB section — the measured evidence that the 512 B fit is a
# property of the mechanisms philosophy #5 bans.
#
# add_tsb_variant(<name> <boot-section-bytes>)
function(add_tsb_variant name bytes)
@@ -125,13 +124,8 @@ endfunction()
# chips build pureboot alone.
if(LIBAVR_MCU STREQUAL "atmega328p")
add_tsb_variant(tsb_asm 512)
add_tsb_variant(tsb_policy 1024)
add_tsb_variant(tsb_pure 1024)
add_tsb_variant(tsb_tricks 1024)
# The policy tier's floor is measured with the loop flags pureboot's size
# work found (a loader's loop bodies all contain calls); the other tiers
# keep the flag set their recorded floors were measured with — none.
target_compile_options(tsb_policy PRIVATE -fno-move-loop-invariants -fno-tree-ter)
endif()
# pureboot — the pure-constraint port (see pureboot/README.md): one source,
@@ -159,17 +153,10 @@ if(PROJECT_IS_TOP_LEVEL)
if(Python3_FOUND)
add_test(NAME pureboot.pi
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/check_pi.py
${CMAKE_OBJDUMP} ${CMAKE_OBJCOPY} ${CMAKE_CXX_COMPILER} ${LIBAVR_MCU}
$<TARGET_FILE:pureboot>
${CMAKE_BINARY_DIR}/CMakeFiles/pureboot.dir/pureboot/pureboot.cpp.obj
${PUREBOOT_BASE_HEX})
${CMAKE_OBJDUMP} ${CMAKE_NM} $<TARGET_FILE:pureboot> ${PUREBOOT_BASE_HEX})
add_test(NAME pureboot.planner
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_planner.py
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py)
add_test(NAME pureboot.handshake
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_handshake.py)
add_test(NAME pureboot.updatelink
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_update_link.py)
endif()
# The protocol test flashes this fixture through the loader with the real
@@ -248,12 +235,12 @@ if(PROJECT_IS_TOP_LEVEL)
# The size matrix: every configuration axis that could move the image
# size — the serial backend (different code), the USART instance
# (different registers), the clock (different constants), the baud
# through the shapes its bit timing takes, and the pins through the one
# thing they decide (whether a bit-banged link has to release the USART
# that owns them) — each combination must still fit the chip's slot
# budget. The timeout is a constant and adds no axis. The stock build is
# one point of this matrix and already has its test.
# (different registers), the clock (different constants), and the baud
# through the shapes its bit timing takes — each combination must still
# fit the chip's slot budget. Pins are size-neutral (port and bit are
# immediate operands) and the timeout is a constant, so neither adds an
# axis. The stock build is one point of this matrix and already has its
# test.
function(pureboot_size_variant name)
pureboot_add_loader(${name} ${ARGN})
add_test(NAME ${name}.size
@@ -261,35 +248,37 @@ if(PROJECT_IS_TOP_LEVEL)
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
# The autobaud loader: one clock-agnostic image per chip, so it has no
# clock x baud axis of its own — the matrix below sweeps those for the
# fixed-baud builds, and this one binary has to serve all of them at run
# time. Size-tested against the same per-chip budget as every other variant.
pureboot_add_loader(pureboot_autobaud SERIAL autobaud)
add_test(NAME pureboot_autobaud.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:pureboot_autobaud>
# The autobaud loader (pureboot/autobaud.md): one clock-agnostic image per
# chip, no clock × baud axis, size-tested against the same per-chip budget on
# every chip.
#
# Only the unified version is built. Its two predecessors —
# pureboot_autobaud_pure.cpp at 508 B and pureboot_autobaud_reg.cpp at 512 —
# had 4 B and 0 B of margin on the 1284P, and the fix for the activation hang
# costs ~22, which puts them at 530 and 534. Neither can ship, so the choice
# the branch existed to offer is settled by measurement rather than taste.
# The sources stay for the record; autobaud.md carries the numbers.
function(pureboot_autobaud_variant name source)
pureboot_add_autobaud(${name} ${source})
add_test(NAME ${name}.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:${name}>
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
pureboot_autobaud_variant(pureboot_autobaud_uni pureboot_autobaud_uni.cpp)
# One point of the exhaustive matrix, named from its resolved parameters
# so the enumeration cannot collide with itself. `pins` is empty for the
# default pair, or the index of the USART whose own pins a bit-banged
# link sits on. Unreachable rates drop out here rather than aborting the
# configure.
function(pureboot_matrix_point hz baud link pins)
set(_name pbm_${hz}_${baud}_${link})
# so the enumeration cannot collide with itself. Unreachable rates drop
# out here rather than aborting the configure.
function(pureboot_matrix_point hz baud link)
if(link STREQUAL "software")
pureboot_baud_feasible(${hz} ${baud} 1 _ok)
set(_args SERIAL software)
if(NOT pins STREQUAL "")
list(APPEND _args RX ${PUREBOOT_USART${pins}_RX} TX ${PUREBOOT_USART${pins}_TX})
set(_name ${_name}_on${pins})
endif()
else()
pureboot_baud_feasible(${hz} ${baud} 0 _ok)
set(_args USART ${link})
endif()
if(_ok)
pureboot_size_variant(${_name} CLOCK ${hz} BAUD ${baud} ${_args})
pureboot_size_variant(pbm_${hz}_${baud}_${link} CLOCK ${hz} BAUD ${baud} ${_args})
endif()
endfunction()
@@ -316,25 +305,24 @@ if(PROJECT_IS_TOP_LEVEL)
# sites), the largest image the space produces and a shape the ladder
# default — always the *fastest* rate a clock reaches — never picks.
#
# Every chip runs the full cross product: the size-bearing classes (flash
# addressing, hand-over shape, page size, USART inventory) are what make
# the image differ, and a chip outside them is expected to match its class
# — but "expected" is what a matrix is for, and the whole sweep is cheap
# enough to run rather than reason about. PUREBOOT_FULL_MATRIX is what
# selects it; the compact matrix below is the per-commit default.
# Bounded to one chip per size-bearing class: flash addressing (the
# word-addressed 1284), hand-over shape (the patched vector on the tinies
# and m48s), page size, and USART inventory. Everything else in the image
# is chip-independent code, so a further chip buys builds and no
# coverage; every chip outside the set carries the compact matrix.
get_property(_full_bauds GLOBAL PROPERTY PUREBOOT_BAUD_LADDER)
list(APPEND _full_bauds 16000 4800 2400 1200)
if(DEFINED ENV{PUREBOOT_FULL_MATRIX})
set(_matrix_spot attiny13a attiny85 atmega48pa atmega8a atmega168pa
atmega328p atmega164a atmega644a atmega1284p)
if(DEFINED ENV{PUREBOOT_FULL_MATRIX} AND LIBAVR_MCU IN_LIST _matrix_spot)
foreach(_matrix_hz IN LISTS _full_clocks)
foreach(_matrix_baud IN LISTS _full_bauds)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software)
if(PUREBOOT_HAS_USART)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 0)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0 "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0)
endif()
if(PUREBOOT_HAS_USART1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1 "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1)
endif()
endforeach()
endforeach()
@@ -353,42 +341,11 @@ if(PROJECT_IS_TOP_LEVEL)
endforeach()
list(GET _matrix_clocks -1 _matrix_top_hz)
pureboot_size_variant(pureboot_sw_wide CLOCK ${_matrix_top_hz} BAUD 9600 SERIAL software)
# The pin axis at the widest software image — the slowest ladder rate
# against the fastest clock, whose bit spin needs the 16-bit delay
# loop — with the USART release on top of it. The exhaustive sweep
# above carries the same axis across its whole cross product.
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_sw_wide_on_usart0 CLOCK ${_matrix_top_hz} BAUD 9600
SERIAL software RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_sw_wide_on_usart1 CLOCK ${_matrix_top_hz} BAUD 9600
SERIAL software RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
endif()
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_usart1 USART 1)
endif()
# The pin axis at its fixed points, in both matrix modes. The autobaud
# loader carries no clock and no baud, so the sweep has nothing to vary
# for it — yet it is the tightest image in the space, and on a USART's
# own pins it pays the release too: that combination is the one that
# overflowed the 1284's slot. The software build on those pins is the
# same deployment the mute test drives.
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_sw_on_usart0 SERIAL software
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
pureboot_size_variant(pureboot_autobaud_on_usart0 SERIAL autobaud
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_sw_on_usart1 SERIAL software
RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
pureboot_size_variant(pureboot_autobaud_on_usart1 SERIAL autobaud
RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
endif()
# One configured deployment end to end — a real board's shape rather
# than the stock assumption: the ATmega328P on its shipped 1 MHz fuses,
# the software UART on hand-picked pins (TX = PB1, RX = PB5), the ladder
@@ -417,33 +374,6 @@ if(PROJECT_IS_TOP_LEVEL)
set_tests_properties(pureboot.custom PROPERTIES TIMEOUT 180)
endif()
# Hand-over with the USART that owns the loader's pins left enabled — the
# state an application reaches by jumping in without a reset, and the one
# that made a bit-banged loader on PD0/PD1 (where the Uno's USB bridge
# lands) receive and obey while answering nothing. Run where it was found
# on silicon; the runner supplies the pin ownership simavr has no model
# for, which is what lets this fail when the release is gone.
if(LIBAVR_MCU STREQUAL "atmega328p" AND DEFINED PB_DEVICE)
get_target_property(_mute_hz pureboot_sw_on_usart0 PUREBOOT_HZ)
get_target_property(_mute_baud pureboot_sw_on_usart0 PUREBOOT_BAUD)
get_target_property(_mute_link pureboot_sw_on_usart0 PUREBOOT_LINK)
add_executable(pbapp_handover test/pbapp.cpp)
target_link_libraries(pbapp_handover PRIVATE libavr)
target_compile_definitions(pbapp_handover PRIVATE PUREBOOT_CLOCK_HZ=${_mute_hz}
PUREBOOT_BAUD=${_mute_baud} PUREBOOT_HANDOVER)
add_custom_command(TARGET pbapp_handover POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_handover> $<TARGET_FILE:pbapp_handover>.bin)
add_test(NAME pureboot.mute
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbmute.py
${PB_DEVICE} $<TARGET_FILE:pureboot_sw_on_usart0> ${PUREBOOT_SIM_MCU} ${_mute_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_mute_baud}
$<TARGET_FILE:pbapp_handover>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbmute-work ${_mute_link})
set_tests_properties(pureboot.mute PROPERTIES TIMEOUT 180)
endif()
# The second USART, driven for real on one chip: instance selection is
# compile-checked everywhere, but only a live session proves the loader
# initialized and polls the USART it claims to. The fixture application
@@ -482,12 +412,14 @@ if(PROJECT_IS_TOP_LEVEL)
add_custom_command(TARGET pbapp_autobaud POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_autobaud> $<TARGET_FILE:pbapp_autobaud>.bin)
add_test(NAME pureboot.autobaud
foreach(_variant uni)
add_test(NAME pureboot.autobaud_${_variant}
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbautobaud.py
${PB_DEVICE} $<TARGET_FILE:pureboot_autobaud> ${PUREBOOT_SIM_MCU}
${PB_DEVICE} $<TARGET_FILE:pureboot_autobaud_${_variant}> ${PUREBOOT_SIM_MCU}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} $<TARGET_FILE:pbapp_autobaud>.bin
1000000 9600 ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbautobaud-work)
set_tests_properties(pureboot.autobaud PROPERTIES TIMEOUT 240)
${CMAKE_BINARY_DIR}/pbautobaud-${_variant}-work)
set_tests_properties(pureboot.autobaud_${_variant} PROPERTIES TIMEOUT 240)
endforeach()
endif()
endif()

View File

@@ -6,7 +6,7 @@
"hidden": true,
"generator": "Ninja",
"binaryDir": "${sourceDir}/build/${presetName}",
"toolchainFile": "$env{LIBAVR_ROOT}/cmake/avr-toolchain.cmake",
"toolchainFile": "${sourceDir}/libavr/cmake/avr-toolchain.cmake",
"cacheVariables": {
"CMAKE_BUILD_TYPE": "Release",
"CMAKE_EXPORT_COMPILE_COMMANDS": "ON",

1
libavr Submodule

Submodule libavr added at edc77ca43f

View File

@@ -116,17 +116,6 @@ else()
math(EXPR _pb_limit "${_pb_slot} - 2")
endif()
# The pins each USART owns. A bit-banged link deployed on them has to release
# that USART before it can drive the line, and those instructions are the one
# way the choice of pins moves the image — so a size matrix needs them as an
# axis even though pins are otherwise immediate operands. Uniform across every
# mega libavr covers: USART0 (the classics' un-numbered USART included) on
# PD0/PD1, USART1 on PD2/PD3.
set(_pb_usart0_rx pd0)
set(_pb_usart0_tx pd1)
set(_pb_usart1_rx pd2)
set(_pb_usart1_tx pd3)
# simavr names its cores after the base dies; the A revisions run on them
# (the 644PA on the 644P core).
set(_pb_sim_mcu ${LIBAVR_MCU})
@@ -144,8 +133,6 @@ set_property(GLOBAL PROPERTY PUREBOOT_WRAP "${_pb_wrap}")
set_property(GLOBAL PROPERTY PUREBOOT_DEFAULT_HZ ${_pb_hz})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART ${_pb_has_usart})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART1 ${_pb_has_usart1})
set_property(GLOBAL PROPERTY PUREBOOT_USART0_TX ${_pb_usart0_tx})
set_property(GLOBAL PROPERTY PUREBOOT_USART1_TX ${_pb_usart1_tx})
# The port's own build (tests, the size matrix) reads the geometry from the
# parent scope; a downstream consumer gets the same variables for free.
@@ -158,10 +145,6 @@ set(PUREBOOT_DEFAULT_HZ ${_pb_hz} PARENT_SCOPE)
set(PUREBOOT_HAS_USART ${_pb_has_usart} PARENT_SCOPE)
set(PUREBOOT_HAS_USART1 ${_pb_has_usart1} PARENT_SCOPE)
set(PUREBOOT_SIM_MCU ${_pb_sim_mcu} PARENT_SCOPE)
set(PUREBOOT_USART0_RX ${_pb_usart0_rx} PARENT_SCOPE)
set(PUREBOOT_USART0_TX ${_pb_usart0_tx} PARENT_SCOPE)
set(PUREBOOT_USART1_RX ${_pb_usart1_rx} PARENT_SCOPE)
set(PUREBOOT_USART1_TX ${_pb_usart1_tx} PARENT_SCOPE)
# The rates a default may pick, fastest first.
set_property(GLOBAL PROPERTY PUREBOOT_BAUD_LADDER 115200 57600 38400 19200 9600)
@@ -211,20 +194,13 @@ function(pureboot_default_baud clock software outvar)
endfunction()
# pureboot_add_loader(<name> [CLOCK <hz>] [BAUD <bd>]
# [SERIAL auto|hardware|software|autobaud] [USART <n>]
# [SERIAL auto|hardware|software] [USART <n>]
# [RX <pin>] [TX <pin>] [TIMEOUT <s>])
#
# The loader target plus its flashable images (<name>.hex for a programmer,
# <name>.bin for --update-loader). The resolved deployment is stamped on the
# target as PUREBOOT_HZ / PUREBOOT_BAUD / PUREBOOT_LINK (the link spelled
# usart0, usart1, or sw:<RX>,<TX> with a trailing @<n> where those pins are a
# USART's own) — what a test harness speaks to it with.
#
# SERIAL autobaud measures the host's bit timing at run time, so the image
# carries no clock and no baud: CLOCK and BAUD are not build parameters there,
# and one binary per chip serves every F_CPU and every rate. The stamped
# PUREBOOT_HZ/PUREBOOT_BAUD then record what a harness should *drive* it at,
# not what it was built for.
# usart0, usart1 or sw:<RX>,<TX>) — what a test harness speaks to it with.
function(pureboot_add_loader name)
cmake_parse_arguments(PB "" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT" "" ${ARGN})
if(PB_UNPARSED_ARGUMENTS)
@@ -246,8 +222,8 @@ function(pureboot_add_loader name)
if(NOT PB_SERIAL)
set(PB_SERIAL auto)
endif()
if(DEFINED PB_USART AND NOT PB_SERIAL MATCHES "^(auto|hardware)$")
message(FATAL_ERROR "pureboot_add_loader(${name}): USART ${PB_USART} contradicts SERIAL ${PB_SERIAL}")
if(DEFINED PB_USART AND PB_SERIAL STREQUAL "software")
message(FATAL_ERROR "pureboot_add_loader(${name}): USART ${PB_USART} contradicts SERIAL software")
endif()
if(DEFINED PB_USART)
set(PB_SERIAL hardware)
@@ -276,7 +252,7 @@ function(pureboot_add_loader name)
set(PB_SERIAL software)
endif()
endif()
if(PB_SERIAL MATCHES "^(software|autobaud)$")
if(PB_SERIAL STREQUAL "software")
if(NOT PB_RX)
set(PB_RX pb0)
endif()
@@ -288,26 +264,12 @@ function(pureboot_add_loader name)
message(FATAL_ERROR "pureboot_add_loader(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
if(PB_SERIAL STREQUAL "autobaud")
set(_serial_defines PUREBOOT_AUTOBAUD PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
else()
set(_serial_defines PUREBOOT_SOFT_SERIAL PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
endif()
# sw:<RX>,<TX> as port letter and bit, upcased — with @<n> where
# the TX pin is a USART's own TXD, since a harness driving that
# link has to know the USART owns the pin until the loader
# releases it.
# sw:<RX>,<TX> as port letter and bit, upcased.
string(SUBSTRING ${PB_RX} 1 2 _rx_pin)
string(SUBSTRING ${PB_TX} 1 2 _tx_pin)
string(TOUPPER "sw:${_rx_pin},${_tx_pin}" _link)
string(REPLACE "SW" "sw" _link ${_link})
get_property(_tx0 GLOBAL PROPERTY PUREBOOT_USART0_TX)
get_property(_tx1 GLOBAL PROPERTY PUREBOOT_USART1_TX)
if(_usart AND PB_TX STREQUAL _tx0)
set(_link "${_link}@0")
elseif(_usart1 AND PB_TX STREQUAL _tx1)
set(_link "${_link}@1")
endif()
endif()
endif()
if(NOT PB_BAUD)
@@ -318,29 +280,23 @@ function(pureboot_add_loader name)
endif()
endif()
if(PB_SERIAL STREQUAL "autobaud")
# No clock and no baud reach the image; the window is a poll budget.
set(_defines ${_serial_defines})
else()
set(_defines PUREBOOT_CLOCK_HZ=${PB_CLOCK} PUREBOOT_BAUD=${PB_BAUD} PUREBOOT_TIMEOUT=${PB_TIMEOUT}
${_serial_defines})
endif()
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/pureboot.cpp)
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE ${_defines})
# Codegen shaping for the loader TU only. At -Os GCC otherwise rewrites the
# byte-stream loops' counters into end-pointer forms that cost registers
# (-fno-ivopts, -fno-split-wide-types), leaves register pressure on the
# table with the default allocator (-fira-algorithm=priority), and keeps
# expression temporaries in registers (-fno-tree-ter) — but every loop body
# here contains a call, so a register held across it costs more than the
# load-immediate it saves. The set is fitted to the loader's body and has to
# be re-measured when that body changes: -fno-move-loop-invariants belonged
# here while the command loop carried four transfer bodies and costs bytes
# now that it carries one.
# Codegen shaping for the loader TU only, worth 1436 B depending on the
# chip. At -Os GCC otherwise rewrites the byte-stream loops' counters into
# end-pointer forms that cost registers (-fno-ivopts,
# -fno-split-wide-types), leaves register pressure on the table with the
# default allocator (-fira-algorithm=priority), and keeps loop-invariant
# immediates and expression temporaries in registers
# (-fno-move-loop-invariants, -fno-tree-ter) — but every loop body here
# contains a call, so a register held across it costs more than the
# load-immediate it saves.
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-tree-ter -fno-split-wide-types)
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
@@ -355,3 +311,37 @@ function(pureboot_add_loader name)
PUREBOOT_LINK ${_link})
endfunction()
# pureboot_add_autobaud(<name> <source> [RX <pin>] [TX <pin>])
#
# An autobaud software-serial loader from <source> (pureboot_autobaud_*.cpp).
# Autobaud measures the host's bit timing at runtime, so the image carries no
# clock and no baud — one binary per chip runs at any F_CPU. Same per-chip
# geometry, link and codegen flags as pureboot_add_loader(); only the clock and
# baud axes fall away. Two source files are under review (autobaud.md):
# pureboot_autobaud_pure.cpp and pureboot_autobaud_reg.cpp.
function(pureboot_add_autobaud name source)
cmake_parse_arguments(PB "" "RX;TX" "" ${ARGN})
if(NOT PB_RX)
set(PB_RX pb0)
endif()
if(NOT PB_TX)
set(PB_TX pb1)
endif()
foreach(_pin ${PB_RX} ${PB_TX})
if(NOT _pin MATCHES "^p[a-h][0-7]$")
message(FATAL_ERROR "pureboot_add_autobaud(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
get_property(_base_hex GLOBAL PROPERTY PUREBOOT_BASE_HEX)
get_property(_app GLOBAL PROPERTY PUREBOOT_APP)
get_property(_wrap GLOBAL PROPERTY PUREBOOT_WRAP)
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/${source})
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
endfunction()

View File

@@ -8,52 +8,46 @@ erase, reset-vector surgery, updating the loader itself — lives in the host
tool (`pureboot.py`).
The image is **position-independent**: control flow is PC-relative, the
transfer paths take wire addresses, the write guard protects the slot the code
is *running* in (from the runtime return address), nothing else is
flash-resident to address at all, and the application jump is an indirect call
read/write paths take wire addresses, the write guard protects the slot the
code is *running* in (from the runtime return address), the info block is
addressed from that same anchor, and the application jump is an indirect call
to an absolute entry. The identical binary therefore runs from any slot with
every command intact, which makes pureboot **its own staging loader**: the host
installs the same binary one slot below the resident, jumps into it, and lets
it rewrite the resident. The lint holds it to that literally — the image must
come out byte-identical linked at a different base.
every command intact, which makes pureboot **its own staging loader**: the
host installs the same binary one slot below the resident, jumps into it, and
lets it rewrite the resident.
## Chips
Sizes are the default configuration: the hardware USART0 at 115200 8N1 on a
16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at 57600 8N1 on
the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above). Every axis moves
per build — see *Configuration*. The autobaud column is the clock-free build,
which is the largest the space produces and the tightest fit in the matrix;
it carries the calibration machinery and no clock at all.
per build — see *Configuration*; the largest image any of them produces is a
software UART at a slow baud, which on the 1284s is 494 B, the tightest fit in
the whole matrix at 18 B spare.
| Chip | Flash | Loader at | Link | Stock | Autobaud |
|---|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 394 B | 464 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 398 B | 468 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 402 B | 472 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 402 B | 472 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 364 B | 478 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 366 B | 482 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 366 B | 482 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 392 B | 468 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 402 B | 478 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 404 B | 482 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 404 B | 482 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 404 B | 482 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 404 B | 482 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 398 B | 476 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 424 B | 502 B |
| Chip | Flash | Loader at | Link | Size |
|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 416 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 420 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 424 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 424 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 396 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 400 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 400 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 414 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 434 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 432 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 478 B |
† No hardware boot section: the host patches the reset vector, and the budget
is 510 bytes, since the slot's last word is the trampoline.
The tightest fit in the whole space is the 1284s' autobaud build deployed on a
USART's own pins, 506 of its 512 — they alone carry the far-flash machinery
(ELPM reads, RAMPZ page commands), autobaud alone carries the calibration loop,
and a bit-banged link on a USART's pins alone has to release it (below). The
same build on the default pins is 502. The flash bank riding in a transfer's
selector byte keeps even those chips' addressing the same 16-bit form every
other chip uses, which is why they are no longer the outlier they were.
The 1284s are the heaviest because they alone carry the far-flash machinery —
ELPM reads, RAMPZ page commands, a word-addressed wire.
The software UART enables the RX pull-up; TX idles high. All multi-byte wire
quantities are little-endian.
@@ -68,7 +62,7 @@ repo's build and by a downstream project alike:
|---|---|---|
| `CLOCK <hz>` | the clock the board runs | 16 MHz megas, 8 MHz t25/45/85, 9.6 MHz t13s |
| `BAUD <bd>` | the wire rate | the ladder below |
| `SERIAL auto\|hardware\|software\|autobaud` | the link backend | `auto`: the hardware USART where the chip has one |
| `SERIAL auto\|hardware\|software` | the link backend | `auto`: the hardware USART where the chip has one |
| `USART <n>` | the USART instance (x4 megas carry two) | 0 |
| `RX <pin>`, `TX <pin>` | software-UART pins | `pb0`, `pb1` |
| `TIMEOUT <s>` | the activation window | 8 |
@@ -80,45 +74,6 @@ receiver's 100-cycles-a-bit floor. Whatever is picked or overridden is
re-checked in the compile: an infeasible combination, or a USART the chip does
not have, fails with a named static assert.
Putting a bit-banged link on a USART's own pins is a supported deployment, and
the usual one where a board's USB bridge is wired to RXD/TXD: the link's `init`
clears that USART's `UCSRnB` first, because while its `TXEN` is set the USART —
not the port register — owns the TX pin, and a loader entered from an
application that left it enabled would receive and obey while answering nothing
(§20.2). It costs four bytes, and only on those pins.
`SERIAL autobaud` takes neither: the loader **measures** the host's bit timing
at run time, so `CLOCK` and `BAUD` are not build parameters there and one
binary per chip serves every clock and every rate. It is for the deployments
whose clock is not known at build time and does not hold still — the internal
RC oscillator, ±10 % from the factory and moving with supply and temperature —
where a fixed-baud software build has to be rebuilt per clock and still drifts
out of tolerance. The cost is that it is software-serial only (a hardware USART
needs its divisor programmed) and that activation counts poll iterations rather
than seconds, since there is no clock to convert them against
(`PUREBOOT_AUTOBAUD_POLLS`, default 4,000,000).
**Pick the rate by cycles a bit, and leave the oscillator room.** What the
calibration can measure is bounded by how many clock cycles one bit lasts, so a
rate is only ever sensible relative to the clock. Two different floors matter:
| | cycles a bit |
|---|---|
| the logic's floor — exact clock, simulated | solid to ~36, fails outright by ~31 (`pureboot.autobaud` gates a point here) |
| a factory-trimmed internal RC, measured on an ATtiny13A | reliable at ~118; already locking 1 attempt in 5 by ~59 |
The gap is the oscillator's own jitter, and no exact-clock simulation shows it.
So on an RC part, **budget about 100 cycles a bit** — the same order as the
fixed-baud software receiver's floor — rather than the logic's ~36. Measured
envelope on that ATtiny13A, with an application resident: 9.6 and 4.8 MHz reach
115200, 1.2 MHz reaches 9600, 600 kHz reaches 4800, 128 kHz reaches 2400.
One trap in testing this: on a patched-vector chip an **erased** application
region walks straight back up into the loader, so every expired window opens
another one and the host's retries eventually catch the pulse. That reads as far
more reliable than the same part with an application resident, which gets one
window per reset. Measure with an application in place.
A downstream project brings its usual libavr setup (the `libavr` target, the
chip via the `LIBAVR_MCU` toolchain preset), consumes this directory, and
states its deployment — an ATmega328P on its shipped 1 MHz fuses with the
@@ -158,20 +113,9 @@ window; any other byte is discarded and awaited again, so line noise can delay
the loader but never lock it. A window expiring on an idle line boots the
application.
An autobaud build opens differently, because it has to learn the rate before it
can read a byte at all: the host sends the **calibration byte 0xC0** — a start
bit plus six zero data bits form one low pulse of seven bit-times — and the
loader times that pulse into its bit period. A single `p` then activates; the
pulse has already proven a host is present, which the two-byte knock exists to
establish elsewhere. Both waits are bounded, so a stray low pulse with no host
behind it costs one window and then boots the application rather than holding
the loader.
The window is a compile-time constant (`TIMEOUT`, 8 s by default), so the whole
EEPROM belongs to the application — pureboot keeps no state of its own.
Re-timing a deployed loader is a self-update with a re-timed build. An autobaud
build counts poll iterations instead (`PUREBOOT_AUTOBAUD_POLLS`), there being
no clock to turn into seconds.
Re-timing a deployed loader is a self-update with a re-timed build.
## Session
@@ -181,116 +125,75 @@ write and sends the prompt `+` (0x2b), which is therefore also the previous
command's completion ack. A session is: await `+`, send a command, read its
reply, repeat.
Addresses are **byte addresses within a 64 KiB bank**, and the bank rides in
the command's selector byte, so no command has to speak word addresses. `J` is
the exception: it takes a word address, because that is what the hardware's own
jump takes. EEPROM and data-space addresses and all counts are bytes.
On chips whose flash exceeds 64 KiB (the 1284s — info-block flag bit 1) the
`R`/`W` flash addresses are **word** addresses; everywhere else they are byte
addresses (the 644s' 64 KiB is exactly the 16-bit byte space). EEPROM
addresses and all counts are bytes.
The loader trusts the host to keep addresses in range: it does not bound them
against the chip. **Gotcha:** a write (or read) that runs past `E2END` wraps
EEAR is only as wide as the array, so an address past the end truncates onto
low EEPROM and the write silently overwrites it. Keeping transfers within the
real sizes is the host's job (the shipped tool does); the flash budget is
better spent on features than on re-checking a bound the host already holds.
against the info block. **Gotcha:** a `w` (or `r`) that runs past `E2END` wraps
EEAR is only as wide as the array, so an address past the end truncates onto
low EEPROM and the write silently overwrites it. Keeping writes within the
advertised sizes is the host's job (the shipped tool does); the flash budget
is better spent on features than on re-checking a bound the host already holds.
| Cmd | Arguments | Reply |
|---|---|---|
| `b` | — | 4 bytes: the pureboot version, then the three signature bytes |
| `G` | sel8, addr16, n8 | n bytes from the selected space (n = 0 means 256) |
| `g` | sel8, addr16, n8, then n data bytes | `+` per byte, sent once its write has begun |
| `W` | sel8, addr16, then one page of data | — (completion = next prompt) |
| `b` | — | the 12-byte info block |
| `R` | addr16, n8 | n flash bytes (n = 0 means 256) |
| `W` | addr16 (any address in the page), then one page of data | — (completion = next prompt) |
| `r` | addr16, n8 | n EEPROM bytes (n = 0 means 256) |
| `w` | addr16, n8, then n data bytes | `+` per byte, sent once its write has begun |
| `F` | — | 4 bytes: low fuse, lock, extended fuse, high fuse |
| `J` | word address (16-bit) | `+`, then execution continues there |
| other | — | ignored; the loop re-prompts (send a junk byte, await `+`, to resync) |
`G` and `g` are one letter in two cases, which is the whole command set for
every memory: the **selector** byte's low nibble names the space and its high
nibble carries the flash bank.
| Space | | |
|---|---|---|
| 0 | flash | read-only here; it is written through `W` and the SPM space |
| 1 | EEPROM | |
| 2 | data | SRAM — and with it the register file and every I/O register, which share the data address space on AVR |
| 3 | fuse and lock | index 0..3 in the hardware's own Z order: low, lock, extended, high |
| 4 | SPM | write-only: the byte goes to SPMCSR and fires the instruction at the address |
The data space is worth more than it looks. pureboot keeps **zero static RAM**
and pushes no register, so at loader entry an application's SRAM is still
whatever the application left there, bar the handful of bytes of return-address
stack — which makes `G` over space 2 a post-mortem of a running application,
not just a poke hole. The same address space carries the register file and the
I/O registers, so peripheral state is readable too; reading some of those has
side effects (reading UDR clears its flags), which is the host's business to
know.
Programming a page is therefore `W` to fill the buffer, then a `g` to the SPM
space for the erase, another for the write, and on a boot-sectioned chip a
third to re-enable the RWW section — `0x03`, `0x05` and `0x11`, the SPMCSR
encodings every part pureboot targets shares. The loader carries no page-commit
logic of its own, and the same primitive reaches every other SPM operation,
lock bits included.
The SPM store and the SPM instruction must issue within four cycles of each
other (§26.2), which no host can hit across a serial link — so this one
primitive is *fused* rather than being a poke of SPMCSR followed by a poke of
something else. That four-cycle window is the floor on how low-level a
bootloader's primitives can go; it is not a byte-count decision.
An SPM command aimed at the 512-byte slot the loader is **running in** is
dropped, so a broken host cannot brick the running copy, while a staged copy
one slot lower may rewrite the resident — which is what a self-update is.
`W` streams exactly one SPM page (size from the info block) into the buffer,
then erases and programs — except pages inside the 512-byte slot
the loader is *running* in, which are drained and left alone, so a broken host
cannot brick the running copy and a staged copy may rewrite the resident.
The loader never clears the SPM buffer before a fill, so **one `W` may program
the wrong bytes, and the host is what fixes it**. The buffer is write-once per
word until cleared, and two things leave words in it: a refused page, and —
where SPM runs from anywhere, the tinies and the m48s — an application that
self-programmed before entering. The next page write takes those stale words
and clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on
the tinies), so repeating it programs correctly. The host therefore verifies
every page it writes and rewrites what comes back wrong (three retries, then it
self-programmed before entering. The next `W` takes those stale words and
clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on the
tinies), so repeating it programs correctly. The host therefore verifies every
page it writes and rewrites what comes back wrong (three retries, then it
stops).
`g` is host-paced: send the next byte only after the previous byte's `+`. Fuse
*writing* does not exist — SPM reaches flash and boot lock bits only.
`w` is host-paced: send the next byte only after the previous byte's `+`. `F`
returns the bytes in the hardware's Z order; on a chip without an extended
fuse byte that slot carries no meaning. Fuse *writing* does not exist — SPM
reaches flash and boot lock bits only.
`J` is the one control-transfer primitive: it runs the application (word 0 or
the trampoline word, both derived from the chip) and moves between loader
the trampoline word, both known from the info block) and moves between loader
copies during a self-update. A jump to a slot's base re-enters that copy's own
startup, which must then be knocked afresh.
`b` answers with the loader's identity — its version and the chip's signature —
and nothing else. Everything else the host needs (page size, loader base,
EEPROM size, whether the reset vector must be patched, how many flash banks)
follows from the signature, and the host holds that table; the loader derived
the same facts from its own chip database at build time, so nothing is guessed,
it is simply not sent twice.
The info block (`b`):
An update image, though, is a bare 512-byte slot with no device to ask, and
installing one built for another chip bricks the target. Every loader image
therefore carries a six-byte **stamp**`'P'`, `'B'`, the version, the three
signature bytes — which the loader itself never reads and the host tool refuses
to install a mismatch against.
| Offset | Content |
|---|---|
| 02 | `'P'`, `'B'`, pureboot version (3) |
| 35 | device signature |
| 6 | SPM page size in bytes (0 means 256) |
| 78 | loader base — application flash ends here (a word address when bit 1 is set) |
| 910 | EEPROM size |
| 11 | bit 0: host must patch the reset vector (no hardware boot section); bit 1: flash wire addresses are word addresses |
## Version
`b`'s first byte is the **pureboot version** — the loader's one identity
number, and the only way to tell what a deployed loader is. Nothing else is
numbered: the wire protocol has no version, a pureboot version implies it, and
the host tool holds that map. The tool states the window of loader versions it
speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a version that
changes the protocol becomes the new floor there. A loader newer than the tool
is refused by name rather than decoded on the assumption that nothing moved.
Two generations exist. **1 through 4** speak one session — a 12-byte info block
from `b`, and a command per memory (`R`/`W` flash, `r`/`w` EEPROM, `F` fuses).
**5** replaced those with the single `G`/`g` pair over selector-named spaces
above; the shipped tool speaks both, choosing on the version it reads, so a
deployed pureboot 4 stays drivable and self-updatable to 5.
Collapsing four command bodies into one transfer loop is what paid for the
version: the data space, the host-issued SPM operations and the fuses now share
the loop, the cursor and the argument decode that `R`/`r`/`w` each carried a
copy of. The loader shrank while gaining all three.
The info block's third byte is the **pureboot version** — the loader's one
identity number, and the only way to tell what a deployed loader is. Nothing
else is numbered: the wire protocol has no version, a pureboot version implies
it, and the host tool holds that map. The tool states the window of loader
versions it speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a
version that changes the protocol becomes the new floor there. None has so
far: 1 through 3 speak the identical session. A loader newer than the tool is
refused by name rather than decoded on the assumption that nothing moved.
The tool carries its own version, free to drift; `--version` prints it and the
window.
@@ -349,27 +252,11 @@ with any pureboot build — a re-timed window, a newer version — using the
loader itself as its own staging loader. The image is the loader's own 512
bytes as a raw binary, or the Intel HEX the build emits beside it.
One thing the image cannot tell the host: **which link it speaks.** The update
works by entering copies of the *new* image (steps 3 and 4 below), so a build
made for another baud or another backend answers on that one and not on the
session's — and 512 bytes of position-independent code carry no header to read
it from. Where the new image's link differs, name it:
```sh
# a 57600 fixed-baud resident, replaced by an autobaud build
pureboot.py --port … --baud 57600 --update-loader ab.bin --staged-autobaud
# …or by a 38400 build of the same backend
pureboot.py --port … --baud 57600 --update-loader sw38400.bin --staged-baud 38400
```
The host retunes on the open port, so no DTR pulse resets the copy it is talking
to. Omit them against a changed link and the update stops after installing the
staging copy, saying so and naming this as the cause.
The preflight refuses an image built for another chip: the stamp every pureboot
binary carries must resolve to the device's own geometry, and the error names
both. Die revisions share their base signature and geometry, so their images
are interchangeable — as the silicon is.
The preflight refuses an image built for another chip: the info block embedded
in every pureboot binary (signature, page size, loader base, EEPROM size,
flags) must match the device's own, and the error names both. Die revisions
share their base signature and geometry, so their images are interchangeable —
as the silicon is.
1. The staging slot `[base512, base)` is saved to a host-side state file (on
the 1 KB tiny13s that is the whole application, vectors included).
@@ -386,15 +273,10 @@ are interchangeable — as the silicon is.
content, and the state file is discarded.
Every phase is idempotent and keyed off the actual flash state, so re-running
the same command after any interruption resumes and completes — with one
qualification, which is the link again: from step 2 on, the copy the re-run has
to reach is the *new* image, so a resumed run needs the same `--staged-*` as the
first one. On a patched-vector part step 3 also re-aims word 0 at the staging
copy, so after that point a reset reaches the new image's link and **only** that
one; a re-run on the resident's link finds nothing at all. The state file carries
the only bytes not recoverable from the device; losing it mid-update still
completes the update, and the staging region comes back by reflashing the
application. A boot-sectioned mega needs its fuses for the preflight — read
the same command after any interruption resumes and completes. The state file
carries the only bytes not recoverable from the device; losing it mid-update
still completes the update, and the staging region comes back by reflashing
the application. A boot-sectioned mega needs its fuses for the preflight — read
from the device, or supplied with `--assume-fuses` where reading is impossible
(simulators).
@@ -411,8 +293,8 @@ to reset gets its reset pulse and opens the activation window by itself.
--info --fuses --flash app.hex
Operations run in a fixed order within one session: info, fuses, loader
update, flash (erase / program / read / verify), EEPROM (the same), then
`--peek`/`--poke` — then the loader hands over to the application. `--stay` keeps the session alive
update, flash (erase / program / read / verify), EEPROM (the same) then the
loader hands over to the application. `--stay` keeps the session alive
instead, and a later invocation reconnects into it. `--flash` and `--eeprom`
verify by read-back unless `--no-verify`, and a flash page that reads back
wrong is rewritten up to three times before the run stops (see `W` above).
@@ -420,33 +302,9 @@ wrong is rewritten up to three times before the run stops (see `W` above).
extension. `--force` overrides the refusable safety checks — today, flashing
application data into a mega's reset walk region.
`--autobaud` opens with the calibration pulse instead of the plain knock, for a
loader built `SERIAL autobaud`; the rest of the session is identical, at
whatever `--baud` the host chose.
`--peek ADDR[:N]` and `--poke ADDR:HEX` reach the data space (pureboot 5) —
SRAM, and through the same address space the register file and every I/O
register. Reading an I/O register can have side effects (reading UDR clears its
flags), which is the caller's business to know.
Reads are safe anywhere; **two small regions cannot be written without ending the
session,** because they are what the loader is standing on:
- the **top of SRAM**, where its stack lives — a handful of bytes below RAMEND;
- on an **autobaud** build, the **two bytes at RAMSTART**: the measured bit
period, in `.noinit`, which is the whole of that loader's static RAM. Overwrite
it and its next reply is timed against garbage. On an ATtiny13A that is
`0x60..0x61`, and the symptom is a mangled prompt byte rather than any error —
the loader is fine, it simply is no longer speaking the agreed rate.
Both are self-inflicted rather than defects, and a reset clears them. Note also
that `--poke` can write OSCCAL, which does take effect — but a session can only
survive a step or two of it before the clock walks the link out of the rate
autobaud locked to, and OSCCAL reverts on reset regardless.
Readouts come one fact per line: `--info` prints the device's version and
signature and the geometry that follows from them, `--fuses` each fuse byte
plus, on a boot-sectioned mega, its decoded meaning. Transfers that take wire time draw a transient progress bar on stderr
Readouts come one fact per line: `--info` decodes the info block field by
field, `--fuses` each fuse byte plus, on a boot-sectioned mega, its decoded
meaning. Transfers that take wire time draw a transient progress bar on stderr
when it is a tty. `-v`/`--verbose` adds the decisions as they happen: knock
counts, the programming plan, update state handling and per-phase page counts.
@@ -463,31 +321,17 @@ Per chip preset, `ctest` runs:
the fastest clock — where a software UART's per-bit spin outgrows its
one-register delay loop and takes the 16-bit one. That is the largest image
the configuration space produces, and a shape the ladder default (always the
*fastest* rate a clock reaches) never picks. Pins are an axis for one reason
only, and it is enough: a bit-banged link on a USART's own pins has to
release that USART, so `pureboot_{sw,autobaud}_on_usart{0,1}` build there
too. The timeout is a constant and is no axis;
- `pureboot_autobaud.size` — the clock-free build, which has no clock or baud
axis of its own: one binary per chip has to serve every point the matrix
below sweeps;
- `pbm_*.size` — with `PUREBOOT_FULL_MATRIX=1`, the exhaustive cross product
replacing that compact matrix, on **every** chip: every plausible oscillator
(the internal ones, the CKDIV8 floor, the plain and the UART crystals) ×
every rate reachable from it × every backend, unreachable combinations
dropping out rather than aborting the configure. Thousands of points per
chip, and cheap enough to run rather than reason about;
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, no
flash-resident section but `.text`, and the image byte-identical when linked
at a different base — which is position independence itself rather than a
proxy for it;
- `pureboot.handshake` — the host tool's activation must not hang on a target
that never falls quiet: the drain after a prompt is bounded by the handshake
deadline, and a well-behaved loader still connects;
- `pureboot.updatelink` — an update whose image changes the baud or the backend
must follow the staging copy onto *its* link, since that copy is the new image;
and where nothing was declared, the failure must name the link rather than
report a bare activation timeout, because by then the staging slot is written
and on a 1 KiB tiny that was the application;
*fastest* rate a clock reaches) never picks. Pins are immediate operands and
the timeout is a constant: neither is an axis;
- `pbm_*.size` — under `--full`, the exhaustive cross product replacing that
compact matrix: every plausible oscillator (the internal ones, the CKDIV8
floor, the plain and the UART crystals) × every rate reachable from it ×
every backend, unreachable combinations dropping out rather than aborting
the configure. Bounded to one chip per size-bearing class — flash
addressing, hand-over shape, page size, USART inventory — since everything
else in the image is chip-independent code;
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, the
info block within the image's first 256 bytes;
- `pureboot.planner` — the host tool's pure logic: programming orders and their
recovery properties, the surgery, the staging composition, the boot-fuse
decode, the update preflight over synthetic fuse bytes, and the repairing
@@ -510,12 +354,6 @@ Per chip preset, `ctest` runs:
- `pureboot.usart1` (644A) — the same suite over the second hardware USART:
instance selection is compile-checked everywhere, but only a live session
proves the loader polls the USART it claims;
- `pureboot.mute` (328P) — a software link on USART0's own pins, entered from an
application that handed over with that USART still enabled: the loader must
still answer, which it does only because it releases it. The pin ownership is
the runner's, not simavr's — simavr wires a USART through IRQs and never takes
the pin from the port, so without that model the state under test could not
arise at all;
- `pureboot.dirty` (328P) — entering the loader from a running application over
an SPM buffer it deliberately dirtied, the case the loader declines to guard:
a bare verify must see the corruption and the repairing verify must fix it in
@@ -523,53 +361,7 @@ Per chip preset, `ctest` runs:
anywhere, which is what makes the path constructible;
- `pureboot.update` — the full `--update-loader` flow, then every power-fail
phase: the device is killed mid-write, restarted from its flash dump, and a
re-run must complete the update with the application intact;
- `pureboot.autobaud` (328P, 1284P) — the clock-free build over the GPIO⇄pty
bridge: the calibration handshake, a flash + EEPROM + fuse round trip against
the simulator's own memory, a data-space round trip, the hand-over — then the
same binary again at double the clock, which is the property the backend
exists for. A lone calibration pulse with no knock behind it must still let
the application boot, so no wait in activation can be unbounded.
re-run must complete the update with the application intact.
`size`, `pi`, `planner` and `handshake` are host logic and run anywhere; the
`size`, `pi` and `planner` are host logic and run anywhere; the
simulator-driven targets need simavr and a pty, so they are POSIX-only.
## Hardware
The suite above proves the protocol on every chip; it cannot prove a *board*.
Two things live only on silicon: an RC oscillator that is not on its nominal, and
a reset edge that has to come from somewhere. `tools/pbrig.py` and
`tools/pbhw.py` cover that, and know nothing per-board — every deployment fact
is a flag or a `PUREBOOT_*` environment variable.
```sh
export PUREBOOT_PROGRAMMER=atmelice_isp PUREBOOT_PART=t13 PUREBOOT_PORT=COM6
tools/pbrig.py backup rig-backup/ # verified, before anything is written
tools/pbhw.py --autobaud --loader build/ab.bin --app build/pbapp.hex --marker APP
```
`pbrig.py` is the primitives — `signature`, `reset`, `flash`, `fuses`, `backup`,
`rate` — and the module `pbhw.py` builds on. Two rig facts are encoded in it
because neither is guessable: an **ISP access is the reset edge** (the part runs
the moment the programmer releases it, which is the only edge available when the
adapter's DTR is not wired to reset, so a session begins with an ISP touch and
knocks immediately after), and **avrdude splits `-U` on colons**, so a Windows
path's drive letter breaks the spec and every file is passed as a bare name with
avrdude run in its own directory.
`pbrig.py rate` is the one that turns "the loader is silent, so the wiring must
be wrong" into a number. Against a fixture built with `PUREBOOT_HEARTBEAT` — a
*fixed* cycles-per-bit transmitter — it sweeps the host rate, and the band where
the marker still decodes brackets the part's true bit rate; with the clock the
image was built for, that is the clock the part is really running at. No
instrument beyond the adapter already attached. An ATtiny13A measured this way
came out at 9.072 MHz against its 9.6 MHz nominal, 5.5 % — inside the
datasheet's ±10 % and outside what an 8N1 frame survives, which is the whole
case for the autobaud backend on such a part.
`pbhw.py` takes its bounds from the info block the loader reports, so one run
covers a 1 KiB tiny and a 128 KiB mega alike: identity, the EEPROM round trip
and erase, an application flashed and verified and then *seen running*, the
application region read and erased, the loader slot proven intact across that
erase by an independent ISP read, and an oversized image refused. It overwrites
the application flash and EEPROM, which is why `backup` comes first.

251
pureboot/autobaud.md Normal file
View File

@@ -0,0 +1,251 @@
# pureboot autobaud — findings, and how the version question settled itself
The autobaud loader measures the host's bit timing **at runtime** from a
calibration pulse, so the image carries no clock: one clock-agnostic binary per
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
for the software-serial deployments — the RC-oscillator parts (the tinies,
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
each needs a per-clock build. Autobaud erases that axis. That is the win —
deployment, not bytes.
This document records how the fit was established, why the two versions that
were under review are both dead, and what the loader looks like now.
## The decision, settled by measurement
Two autobaud loaders were built for review, differing in one tradeoff — the
running-slot write guard against strict purity. `pureboot_autobaud_pure.cpp`
landed at **508 B** on the 1284P (4 B spare) and `pureboot_autobaud_reg.cpp` at
**512** (zero spare).
Then hardware testing found a defect that neither could absorb.
**A single spurious calibration pulse wedged the loader.** `run()` budgeted only
the start-edge wait inside `measure()`; the `rx()` that read the knock behind it
was unbudgeted and blocked forever. One stray low pulse on an unattended device
— EMI, or a host that opens the port and never knocks — held the loader in its
activation loop and the application never ran. On a field device that is a hang,
not a hiccup, and it is exactly the deployment autobaud is for.
The fix is to bound the whole activation: an expired knock budget returns a byte
that cannot be the knock, so control falls back into the budgeted `measure()`,
and a line that stays idle boots the application there. It costs about 22 bytes.
| 1284P, with the activation fix | size | 512 B budget |
|---|---|---|
| `pureboot_autobaud_pure.cpp` | 530 | **over by 18** |
| `pureboot_autobaud_reg.cpp` | 534 | **over by 22** |
| `pureboot_autobaud_uni.cpp` | **464** | **48 B spare** |
Both candidates were unshippable, and the margin they were competing over was
never real — it was the space the missing fix should have occupied. So the
choice is not between them. It is the third loader below, which fits with room
to spare *and* carries features neither had. The two sources stay in the tree
for the record; only the unified one is built.
## The unified loader
The insight that paid was the one that had already paid once: **merging command
bodies removes cost that moving them around only redistributes.** Folding `R`,
`r` and `w` into a single address-and-count path had been worth 14 B earlier.
Pushed further — one read command and one write command over *named spaces*
it is worth far more, because four transfer loops collapse into one.
`pureboot_autobaud_uni.cpp` is pureboot 5. It is strictly pure: no inline
assembly, no global register variable, and **no GPIOR either** — the measured
unit lives in a plain static, so the loader claims no chip resource an
application might want, and the GPIOR-versus-static question disappears along
with the chips that have no GPIOR.
### The protocol
| command | arguments | |
|---|---|---|
| `b` | — | version, then the three signature bytes |
| `J` | addr16 | ack, then jump (word address) |
| `W` | sel8, addr16, page bytes | fill the flash page buffer |
| `G` | sel8, addr16, n8 | read n bytes (0 means 256) |
| `g` | sel8, addr16, n8, then n bytes | write, each byte acked |
`sel` is `space | bank << 4`. The low nibble names the space; the high nibble is
flash's third address byte, so every transfer speaks a **byte** address inside a
64 KiB bank and no command has to carry word addresses. The host must not span a
bank boundary in one transfer — it already chunks by page, so nothing it does
comes close.
| space | | |
|---|---|---|
| 0 | flash | `lpm`/`elpm` |
| 1 | EEPROM | |
| 2 | data | SRAM — and with it the register file and every I/O register, which share the data address space on AVR |
| 3 | fuse and lock | |
| 4 | SPM | write-only: the data byte goes to SPMCSR and fires the instruction at the selected address |
Three things follow from that table that the loader never had:
- **RAM read and write**, the missing feature. In a per-command design it would
have cost a fresh dispatch arm and a fresh loop, ~30 B, on a loader with 4 B
spare. As one more space on a shared loop it is a single `ld`/`st`. It also
hands the host arbitrary **I/O register access** for free, because AVR maps
the peripherals into the same address space.
- **Host-driven SPM.** `W` used to end with a hardcoded erase, write and RWW
re-enable — 42 B. Those are now three writes to the SPM space, reusing the
store path's own address, data byte and ack. The host pays three extra
round-trips per page (18 wire bytes against 256 of data) and gains the ability
to issue *any* SPM operation, lock bits included.
- **`W` on the same footing as everything else.** It takes the same selector and
the same byte address instead of a word address of its own, which made flash
addressing uniform across the protocol *and* was 20 B cheaper than keeping its
private convention.
Read and write are `G` and `g` — the same letter, one bit apart — so the
transfer loop picks its direction with a one-word skip rather than a compare.
### Why the SPM space has to be one primitive
It is tempting to go further and expose a generic "poke this I/O register", from
which the host could drive SPM itself. The hardware forbids it: `out SPMCSR, x`
and the `spm` that follows must issue within four cycles, and EEPROM's
EEMPE→EEPE window is the same shape. A host cannot hit a four-cycle window
across a serial link. **Atomicity is the floor, and the atomic unit must be
resident.** That is the real limit on how low-level a bootloader's primitives
can go — not the byte count.
## Where the bytes went
The starting point was the 508 B pure build, disassembled and attributed:
| phase | bytes |
|---|---|
| `link::rx` + `link::tx` (bit-banged UART) | 102 |
| autobaud measure + knock | 62 |
| reset vector, pin init, WDRF check, jump, ack, EEPROM wait | 50 |
| command loop head + dispatch tree | 42 |
| `W` program flash page | 98 |
| `R` / `r` / `w` / `F` bodies, plus the shared address decode | 122 |
| `b` info, `J` jump | 32 |
**208 B — 41% — is physical layer and activation**, which no protocol change can
touch. The command bodies were the entire addressable surface, and they were
four copies of one idea.
The route from a first attempt to the final loader, all on the 1284P:
| step | size |
|---|---|
| unified `G`/`P`, `load`/`store` outlined, `__uint24` cursor | 600 |
| …`load`/`store` inlined; unit in `.noinit`, not `.bss` | 534 |
| …16-bit cursor with the bank in the selector; direction as a command bit | 510 |
| …erase/write/RWW moved out to the SPM space | 484 |
| …`W` sharing the selector-and-address decode | **464** |
Three of those steps are worth keeping as lessons:
- **Outlining `load`/`store` cost more than the four bodies they replaced.** As
functions they were 110 B against the 106 B of inline bodies — the AVR ABI's
argument marshalling plus prologue ate the entire saving. Inlined into the one
shared loop they cost only their own instructions. The call-site lesson cuts
both ways: *merging* call sites pays, *creating* one does not.
- **A `.bss` static drags in `__do_clear_bss`** — 18 B of startup code to zero a
variable that is always measured before it is read. `.noinit` is correct here
and free.
- **A three-byte cursor taxes every space.** Widening the shared cursor so flash
could reach past 64 KiB put an extra increment on EEPROM and RAM reads that
never need it. Moving the bank into the selector byte kept the cursor at
sixteen bits and cost nothing on the wire.
### What did not work
- **Encoding the space in the command byte** (so dispatch becomes masking rather
than a compare tree) cannot carry the bank's four bits alongside a space. The
cheap half of the idea survived as the direction bit; the rest lost to the
selector byte, which is also more extensible.
- **A generic primitive interpreter** — a loader with no logic at all, driven
entirely by the host — is not reachable on AVR. Harvard architecture means the
program counter cannot fetch from data space, so the classic "upload a flash
algorithm into RAM and jump to it" bootstrap is impossible, and on every
boot-sectioned part SPM only takes effect from the boot section anyway. What
remains is a fixed primitive set: still a protocol, still logic, only at a
different granularity. debugWIRE reaches that design point only because its
interpreter is *in silicon*; it costs the loader nothing because it is not in
the loader.
- Below 512 B the saved bytes are largely unspendable on the boot-sectioned
chips: the 328P's smallest boot section is exactly 512 B, and the 1284P's is
1024 B, of which pureboot already occupies only the top half. The margin
matters as headroom for correctness fixes — as this defect showed — not as
flash returned to the application. On the patch-vector parts, which have no
boot section, it *is* returned: on the ATtiny13 the loader is 43% of a 1 KiB
part, and every byte is real.
## The codegen coupling, still load-bearing
`count >> 2` is exact only because the calibration pulse's bit-count (7, from
the 0xC0 byte) equals the poll loop's cycles per iteration (7 — `sbis` 1,
`rjmp` 2, `adiw` 2, `rjmp` 2). The loop shape survived every restructuring here,
verified in the disassembly, but a toolchain bump that reshapes it would break
the lock silently. `test/pbautobaud.py` is what pins it: a wrong unit fails the
flash verify.
## Sizes — every chip
Budget 510 B on the patch-vector parts, 512 elsewhere. The 1284P is no longer
the tight one: the bank nibble made far flash *cheaper* than the near-flash
arithmetic it replaced.
| size | chips | budget | spare |
|---|---|---|---|
| 444 | ATtiny13, 13A | 510 | 66 |
| 448 | ATmega48, 48A, 48P, 48PA; ATtiny25 | 510 | 62 |
| 452 | ATtiny45, 85 | 510 | 58 |
| 460 | ATmega644, 644A, 644P, 644PA | 512 | 52 |
| 464 | **ATmega1284, 1284P**; ATmega8, 8A, 88, 88A, 88P, 88PA | 512 | 48 |
| 466 | ATmega16, 16A, 32, 32A; 164A/P/PA, 168/A/P/PA, 324A/P/PA, 328, 328P | 512 | 46 |
All 37 chips build and size-test green, plus the 12-preset reflect spot set
(guidance rule 4 — the reflect matrix is never run in full), which matches its
generated counterpart byte for byte on every chip in the set. Worst case across
the whole set is **466 B, 46 under budget**.
## Host tool and simulation
- **`pureboot.py`** speaks both generations. `Info.version >= 5` selects the
unified path; everything below it keeps the four-command protocol, so the
fixed-baud loader is untouched. `--autobaud` sends the 0xC0 pulse and one
knock, then derives full geometry from the signature. New: `--peek ADDR[:N]`
and `--poke ADDR:HEX` reach the data space.
- **`test/pbautobaud.py`** drives the loader over the GPIO⇄pty software-UART
bridge through the calibration handshake, a flash + EEPROM + fuse round-trip
cross-checked against the simulator's own memory, a RAM read/write round-trip,
and a hand-over to the fixture application — then repeats at double the F_CPU
with the same binary, which is the clock-agnostic property autobaud exists
for. Run on the near-flash 328P and the word-addressed 1284P.
- It also **pins the activation hang**: the test sends a lone calibration pulse
with no knock behind it and requires the application to boot. Against the
unfixed loader that assertion never returns.
## What remains
- **Real-hardware acceptance.** A cycle-exact simulator cannot produce what
autobaud exists for: a real RC oscillator at ±10% with drift and jitter.
simavr proves the arithmetic and the fit at exact clocks; only silicon proves
the feature. Drive an internal-oscillator ATtiny at a fixed host baud and
confirm lock plus a full flash and verify.
- **A generic `spm::command()` in libavr.** The SPM space issues a runtime
command through `spm::detail::page_command` where RAMPZ exists, and falls back
to a dispatch over the known operations where it does not — the one
preprocessor branch in the file. A two-line library addition would make it
uniform and save a few bytes on the 36 non-RAMPZ chips, none of which are
tight.
- **Retire or revive the two dead variants.** They are kept only as the record
of the measurement; nothing builds them.
## Files
- `pureboot_autobaud_uni.cpp` — the loader. pureboot 5.
- `pureboot_autobaud_pure.cpp`, `pureboot_autobaud_reg.cpp` — superseded, not
built; 530 and 534 B on the 1284P once the activation hang is fixed.
- `pureboot.py``--autobaud`, the unified transfer path, `--peek`/`--poke`.
- `test/pbautobaud.py` — the end-to-end sim test and the hang regression.
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
probe at 506 B, off-tree and gitignored; a size reference only. The unified
loader is 42 B under it, with features the probe never had.

View File

@@ -26,17 +26,14 @@ constexpr std::uint8_t ack = '+';
// Deployment parameters come from the build (pureboot_add_loader()). The
// signature is not one of them: the chip database is the only universal
// source — a tiny13A cannot read its own signature row from code. An autobaud
// build carries no clock and no baud at all; it measures both.
#if !defined(PUREBOOT_AUTOBAUD) && (!defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD))
// source — a tiny13A cannot read its own signature row from code.
#if !defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD)
#error \
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud — create loader targets with pureboot_add_loader(), or PUREBOOT_AUTOBAUD for a clock-free one (README.md)"
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud — create loader targets with pureboot_add_loader() (README.md)"
#endif
#if !defined(PUREBOOT_AUTOBAUD)
using dev = avr::device<{.clock = avr::hertz_t{PUREBOOT_CLOCK_HZ}}>;
constexpr avr::baud_t wire_baud{PUREBOOT_BAUD};
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
@@ -50,110 +47,60 @@ consteval std::int16_t wdrf_field()
// runs from anywhere (Atmel-8271 §26) — keep the application's relocated
// reset vector in the word under the slot.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = avr::hw::curated::has_boot_section();
// Past 64 KiB one bank of flash does not cover the chip, so a transfer's
// selector byte carries the bank and the wire address stays a byte address
// within it. 'J' is the exception: it is a word address everywhere, because
// that is what the hardware's own jump takes.
constexpr bool banked_flash = spm::flash_bytes > 65536;
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one). A slot is 256 of
// those — one value of a wire address's high byte, where 512 bytes span two.
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint16_t wire_base =
word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base);
// A compile-time window, so the whole EEPROM belongs to the application;
// re-timing a deployed loader is a self-update with a re-timed build. An
// autobaud build has no clock to convert seconds against and counts polls.
// re-timing a deployed loader is a self-update with a re-timed build.
#if !defined(PUREBOOT_TIMEOUT)
#define PUREBOOT_TIMEOUT 8
#endif
constexpr std::uint8_t timeout_seconds = PUREBOOT_TIMEOUT;
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr avr::uint24_t autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The loader's one identity number. The protocol carries none of its own —
// a version implies it, and the host tool holds that map (README.md).
constexpr std::uint8_t version = 5;
constexpr std::uint8_t version = 4;
// The image's identity stamp, for the host tool rather than for the wire: an
// update image is a bare 512-byte slot, and without this nothing in it says
// which chip it was built for. The tool refuses to install an image whose
// stamp does not match the device — flashing a foreign loader bricks the
// target, and the loader itself cannot check what has already replaced it.
//
// Never read from flash by the loader — 'b' answers out of this array, but at
// constant indices, so those fold to immediates and no runtime address of it
// is ever formed. `used` keeps the compiler from dropping the copy the host
// needs and `retain` keeps --gc-sections from collecting it.
// The 'b' reply, byte for byte (layout: README.md). Flash-resident because
// no crt copies a .data image — and flash_table's storage carries the word
// alignment 'b' needs to halve the address on the large chips.
// One wire byte per line: this is the reply's layout, not a list.
// clang-format off
[[gnu::used, gnu::retain, gnu::section(".text.stamp")]]
inline constexpr std::uint8_t identity_stamp[]{
'P', 'B', // the magic the host scans an image for
version, // and from here on, exactly what 'b' answers
inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{
'P',
'B',
version,
avr::hw::db.signature[0],
avr::hw::db.signature[1],
avr::hw::db.signature[2],
};
static_cast<std::uint8_t>(page), // 0 means 256
wire_base & 0xff,
wire_base >> 8,
avr::hw::db.mem.eeprom_size & 0xff,
avr::hw::db.mem.eeprom_size >> 8,
static_cast<std::uint8_t>((boot_section ? 0 : 1) | (word_flash ? 2 : 0)), // patch-vector, word-addressed
}>
info_data;
// clang-format on
// Where the identity proper starts: past the magic the host scans for.
constexpr std::uint8_t stamp_identity = 2;
// The address spaces a transfer can name, in a selector byte's low nibble.
// Flash is 0 so it is the cheapest to select.
//
// spm_ops is the one that is not memory: a write there hands its byte to
// SPMCSR and fires the instruction at the transfer's address, which is how
// page erase, page write and RWW re-enable reach the wire without the loader
// carrying a command for each. The hardware's four-cycle store-to-SPM window
// is why this is one fused primitive and not a poke of SPMCSR — no host can
// hit that window across a serial link.
enum : std::uint8_t { sp_flash = 0, sp_eeprom = 1, sp_data = 2, sp_fuse = 3, sp_spm = 4 };
// A selector's high nibble is the flash bank — the address bits above the
// 16-bit wire address, RAMPZ on the chips that have one. Keeping it here
// rather than widening the wire address is what lets one 16-bit cursor serve
// every space: a 24-bit cursor would pay its extra byte on EEPROM and data
// reads that can never need it.
[[gnu::always_inline]] inline std::uint8_t space_of(std::uint8_t selector)
{
return selector & 0x0f;
}
[[gnu::always_inline]] inline std::uint8_t bank_of(std::uint8_t selector)
{
return static_cast<std::uint8_t>(selector >> 4);
}
// The slot a flash address falls in, as one byte. A slot is half as many words
// as bytes, so the word address's high byte is exactly this index — which is
// what lets the write guard compare a single byte, and what the running copy's
// own return address yields for free.
constexpr std::uint8_t slot_shift = std::countr_zero(slot_bytes);
constexpr std::uint8_t bank_shift = 16 - slot_shift;
[[gnu::always_inline]] inline std::uint8_t slot_of([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
{
const auto within = static_cast<std::uint8_t>(at >> slot_shift);
if constexpr (banked_flash)
return static_cast<std::uint8_t>((bank << bank_shift) | within);
else
return within;
}
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL /
// PUREBOOT_AUTOBAUD, defaulting to the chip's USART0 where it has one. The
// software receiver is the polled one: the vector table belongs to the
// application. Templates on the clock, so only the selected backend
// instantiates. pending() is the cheap line test the activation window polls;
// drain() holds until the last frame is off the wire, so a hand-over cannot
// let the target's re-init clip the ack.
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL,
// defaulting to the chip's USART0 where it has one. The software receiver is
// the polled one: the vector table belongs to the application. Templates on
// the clock, so only the selected backend instantiates. pending() is the
// cheap line test the activation window polls; drain() holds until the last
// frame is off the wire, so a hand-over cannot let the target's re-init clip
// the ack.
#if defined(PUREBOOT_SOFT_SERIAL) && defined(PUREBOOT_USART)
#error "PUREBOOT_SOFT_SERIAL and PUREBOOT_USART select opposing serial backends"
#endif
#if defined(PUREBOOT_AUTOBAUD) && defined(PUREBOOT_USART)
#error "PUREBOOT_AUTOBAUD measures a software link; it cannot drive a hardware USART"
#endif
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
@@ -166,30 +113,9 @@ constexpr char usart_digit = '0' + PUREBOOT_USART;
constexpr char usart_digit = '0';
#endif
// Release a hardware USART the application may have left enabled onto a
// bit-banged link's pins. A software transmitter drives its TX pin through the
// port register, but while that USART's TXEN is set the USART owns the pin and
// the port write does nothing — the loader would receive and obey yet never
// answer. Writing UCSRnB zero hands the pin back to the port. Guarded on the
// pin actually being a USART's TXD, so a link on non-USART pins emits nothing.
template <char Inst, avr::io::pin Tx>
[[gnu::always_inline]] inline void release_usart_on()
{
if constexpr (avr::uart::has_usart<Inst>())
if constexpr (avr::uart::detail::usart_pin<Inst>("TXD") == Tx)
avr::hw::reg_impl<avr::uart::detail::ureg<Inst, "UCSR#B">()>::write(0);
}
template <avr::io::pin Tx>
[[gnu::always_inline]] inline void release_usarts_on()
{
release_usart_on<'0', Tx>();
release_usart_on<'1', Tx>();
}
template <avr::hertz_t C, avr::baud_t B>
template <avr::hertz_t C>
struct hardware_link {
using uart = avr::uart::usart<usart_digit, C, {.baud = B, .max_baud_error = 2.5_pct}>;
using uart = avr::uart::usart<usart_digit, C, {.baud = wire_baud, .max_baud_error = 2.5_pct}>;
// The compiled idle poll: lds UCSR0A (2), sbrc skipping the exit (2),
// sbiw + sbci + sbci + brne (6).
@@ -221,10 +147,10 @@ struct hardware_link {
}
};
template <avr::hertz_t C, avr::baud_t B>
template <avr::hertz_t C>
struct software_link {
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, B>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, B>;
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, wire_baud>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, wire_baud>;
// The compiled idle poll: sbis skipping the exit (2), sbiw + sbci +
// sbci + brne (6).
@@ -233,7 +159,6 @@ struct software_link {
static void init()
{
avr::init<rx_t, tx_t>();
release_usarts_on<avr::PUREBOOT_TX>();
}
static bool pending()
@@ -257,45 +182,14 @@ struct software_link {
}
};
// The clock-free link: the bit period is measured from the host's calibration
// pulse instead of derived from a clock, so one image serves every F_CPU and
// every rate. Activation differs in kind from the other two — there is no
// clock to time a window against — so this backend brings its own, below.
struct autobaud_link {
using uart = avr::uart::software_autobaud<avr::PUREBOOT_RX, avr::PUREBOOT_TX>;
static void init()
{
avr::init<uart>();
release_usarts_on<avr::PUREBOOT_TX>();
}
static std::uint8_t rx()
{
return uart::template read<off>();
}
static void tx(std::uint8_t byte)
{
uart::template write<off>(byte);
}
static void drain()
{
uart::drain();
}
};
#if defined(PUREBOOT_AUTOBAUD)
using link = autobaud_link;
#elif defined(PUREBOOT_USART)
#if defined(PUREBOOT_USART)
static_assert(avr::uart::has_usart<usart_digit>(), "PUREBOOT_USART selects a hardware USART this chip does not have");
using link = hardware_link<dev::clock, wire_baud>;
using link = hardware_link<dev::clock>;
#elif defined(PUREBOOT_SOFT_SERIAL)
using link = software_link<dev::clock, wire_baud>;
using link = software_link<dev::clock>;
#else
using link = std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock, wire_baud>,
software_link<dev::clock, wire_baud>>;
using link =
std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock>, software_link<dev::clock>>;
#endif
// The application's entry, pinned by the linker (--defsym): word 0 on a
@@ -315,27 +209,6 @@ extern "C" [[noreturn]] void pureboot_app();
jump(pureboot_app);
}
// Activation: a bounded wait for the host, then the knock. Both forms boot the
// application when the window closes on an idle line, and both bound *every*
// wait — a knock awaited without a deadline would let one stray edge hold an
// unattended device in the loader forever.
#if defined(PUREBOOT_AUTOBAUD)
// The window is a fixed poll budget: with no clock, whole seconds cannot be
// timed. A uint24_t holds it — a fourth byte would cost two words at every
// countdown step for range never used.
void await_host()
{
for (;;) {
if (!link::uart::calibrate(autobaud_budget))
run_app();
// The calibration pulse has already proven a host is there, so one
// byte activates. A knock that never arrives falls back to calibrate(),
// whose own budget then boots the application.
if (link::uart::template read<off>(autobaud_budget) == 'p')
return;
}
}
#else
// The window as one 32-bit countdown, divided by the backend's counted
// poll-loop cycles. Whole seconds is all it promises.
consteval std::uint32_t window_polls()
@@ -362,14 +235,6 @@ std::uint8_t rx_deadline()
return link::rx();
}
void await_host()
{
// 'p' then 'b', each under a fresh window; anything else is line noise.
while (rx_deadline() != 'p' || rx_deadline() != 'b') {
}
}
#endif
// Inlined: read across a call, the first byte strands in a call-saved
// register the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
@@ -387,92 +252,132 @@ void await_host()
return std::bit_cast<std::uint16_t>(pair);
}
// Out of line: several sites send it, and a call is shorter than a
// load-immediate at each.
// Counts arrive in the wire's 8-bit form: 0 means 256. Both streamers fold
// into the one command that reads flash, which is what lets the far one's
// 24-bit cursor sit in the command loop's own call-saved registers.
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
// The 24-bit cursor as the machine holds it — the RAMPZ byte and a 16-bit Z,
// carried apart; the reassembled address folds away inside the far load.
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
// Carrying the wrap is smaller than the flat 32-bit cursor GCC
// builds without it.
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
// Out of line: three sites send it, and a call is shorter than three
// load-immediates.
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
// A wire address and its selector's bank as the flash address they name.
[[gnu::always_inline]] inline spm::flash_address_t flash_address([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
if constexpr (banked_flash)
return (static_cast<spm::flash_address_t>(bank) << 16) | at;
else
return at;
do
link::tx(ee::read(address++));
while (--count);
}
// One byte out of any space. Every accessor shares the transfer's cursor, its
// loop and its call site, so a space costs only its own instruction rather
// than a body, a loop and a dispatch arm of its own.
[[gnu::always_inline]] inline std::uint8_t load(std::uint8_t space, [[maybe_unused]] std::uint8_t bank,
std::uint16_t at)
{
if (space == sp_eeprom)
return ee::read(at);
if (space == sp_data)
return *reinterpret_cast<volatile std::uint8_t *>(at);
if (space == sp_fuse)
return spm::read_fuse<off>(static_cast<spm::fuse>(at));
if constexpr (banked_flash)
return avr::flash_load_far<std::uint8_t>(flash_address(bank, at));
else
return avr::flash_load(reinterpret_cast<const std::uint8_t *>(at));
}
// One byte into a writable space. Flash is not one of them — it arrives a
// page at a time through 'W' and is committed through sp_spm — and the fuses
// are not writable at all: SPM reaches flash and boot lock bits only.
[[gnu::always_inline]] inline void store(std::uint8_t space, std::uint8_t bank, std::uint16_t at, std::uint8_t value,
std::uint8_t slot_high)
{
if (space == sp_data) {
*reinterpret_cast<volatile std::uint8_t *>(at) = value;
return;
}
if (space == sp_spm) {
// The running-slot write guard. An SPM command aimed at the slot this
// code executes from is dropped, so a broken host cannot brick the
// running loader — while a copy one slot lower may still rewrite the
// resident one, which is what a self-update is. Guarding the commit
// rather than the page fill covers erase and write both, and leaves a
// refused page's words in the buffer: harmless, since the next page
// write auto-erases it (§26.2.1).
if (slot_of(bank, at) != slot_high)
spm::command<off>(value, flash_address(bank, at));
// Only a boot-sectioned mega runs on while its RWW section programs;
// everywhere else the CPU halts through erase and write, so the wait
// is already over by the time it returns.
if constexpr (boot_section)
spm::wait();
return;
}
// Host-paced: the ack goes out once the write has begun, so the next byte
// arrives while it completes and nothing is missed without a buffer.
ee::write<off>(at, value);
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, and only that: the erase and the write that
// commit it are host-issued sp_spm stores, which reach the same fused
// store-and-SPM pair through the transfer path's own address and data.
// One page into the SPM buffer, then erase and program — except the slot
// this code is running in (`slot_high`, from run()), which is drained and
// left alone. A broken host therefore cannot brick the running loader, and a
// copy one slot lower may rewrite the resident one.
//
// Nothing discards the buffer first: it is write-once per word (§26.2.1), so
// filling over a refused page or an application's leavings programs stale
// words — but a page write auto-erases it (§26.2.1; §19.2 on the tinies), so
// that write clears the condition and the host's read-back rewrites the page.
void fill_page(std::uint8_t bank, std::uint16_t at)
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
{
// The address names a page, so its in-page bits are dropped and the walk
// starts at the page base; the low byte of the cursor is the whole in-page
// offset, since a page is aligned and never crosses a bank.
std::uint16_t z = at & ~static_cast<std::uint16_t>(page - 1);
// starts at the page base — one induction either way: a byte-addressed
// wire address walks the page itself (the offset bits wrap back to zero),
// while a word one becomes a byte cursor once. The slot index is the wire
// address's high byte — on byte-addressed chips the byte address's, with
// the low bit dropped, since a slot is two of those.
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
// A page is aligned, so it never crosses 64 KiB: RAMPZ is a per-page
// constant and the 16-bit Z's low byte is the whole in-page offset.
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(flash_address(bank, z), word_of({low, high}));
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z) & (page - 1));
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
}
if (page_high != slot_high) {
// Only a boot-sectioned mega runs on while its RWW section programs;
// everywhere else the CPU halts through erase and write.
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
// Programming leaves the RWW section disabled; reads need it back on. The
// same store discards the buffer (§26.2.2), so a boot-sectioned mega never
// meets the stale-word case above.
if constexpr (boot_section)
spm::rww_enable<off>();
}
// The four fuse and lock bytes in the hardware's own Z order: low, lock,
// extended, high.
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
@@ -484,13 +389,19 @@ void fill_page(std::uint8_t bank, std::uint16_t at)
link::init();
// The slot this copy runs in, which the write guard follows: the return
// address is a word address and a slot is half as many words as bytes, so
// its high byte is the slot index outright. No absolute address is ever
// formed, so the image stays position-independent.
const auto slot_high = avr::startup::caller_page();
// The high byte of the slot this copy runs at, which the write guard and
// the info block both follow: the return address is a word address, so its
// high byte is the 256-word slot index, doubled back into byte terms where
// the wire counts bytes. Taken as byteswap's low byte — the builtin already
// swaps the two stacked bytes, and the double swap folds away, where `>> 8`
// would leave the swap materialized.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
await_host();
// 'p' then 'b', each under a fresh window; anything else is line noise.
while (rx_deadline() != 'p' || rx_deadline() != 'b') {
}
for (;;) {
// No prompt while an EEPROM write runs: it blocks SPM and fuse reads
@@ -505,43 +416,47 @@ void fill_page(std::uint8_t bank, std::uint16_t at)
link::drain();
jump(target);
}
case 'b': // identity: the version, then the three signature bytes
// Straight out of the stamp, so the wire and the image can never
// disagree about what this loader is. The indices are constant and
// the array is constexpr, so these are immediates, not flash reads:
// nothing here needs the stamp's runtime address.
for (std::uint8_t at = stamp_identity; at != sizeof identity_stamp; ++at)
link::tx(identity_stamp[at]);
break;
case 'W': // fill one flash page buffer: sel8, addr16, then page bytes
case 'G': // read: sel8, addr16, n8 (0 = 256)
case 'g': { // write: sel8, addr16, n8, then n bytes, each acked
// One decode, one cursor and one loop for every space and both
// directions: a command per memory would carry a copy of all three
// each. 'W' joins the same decode rather than keeping an address
// form of its own, so flash addressing is uniform across every
// command that names it.
const std::uint8_t selector = link::rx();
const std::uint8_t space = space_of(selector);
const std::uint8_t bank = bank_of(selector);
std::uint16_t at = rx16();
if (command == 'W') {
fill_page(bank, at);
case 'b': // info block, read relative to the running slot
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for all four: 'b' is a flash read
// whose arguments the loader already knows, so it joins the
// wire-argument three rather than streaming from a call site of its
// own. That leaves one flash streamer in the image, and lets its
// cursor live in this never-returning loop's own call-saved
// registers instead of being saved and restored around a call.
std::uint16_t address;
std::uint8_t count;
if (command == 'b') {
// The block sits in the image's first 256 bytes (check_pi.py
// asserts it) and slots are 512-aligned, so the low byte of its
// link address is its offset in any slot — halved where wire
// units are words. The high byte is runtime data, so no
// absolute address is ever materialized.
const auto link_byte =
static_cast<std::uint8_t>(reinterpret_cast<std::uint16_t>(info_data.storage.data()));
const std::uint8_t low = word_flash ? static_cast<std::uint8_t>(link_byte >> 1) : link_byte;
address = static_cast<std::uint16_t>(low | (slot_high << 8));
count = static_cast<std::uint8_t>(info_data.size());
} else {
address = rx16();
count = link::rx();
}
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
std::uint8_t count = link::rx();
do {
// Read and write are one letter apart in case, so the direction
// is a single bit and the loop picks it with a one-word skip.
if (command & 0x20) {
store(space, bank, at, link::rx(), slot_high);
tx_ack();
} else
link::tx(load(space, bank, at));
++at;
} while (--count);
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
}
default: // unknown bytes are ignored; the loop re-acks
break;
}

View File

@@ -24,7 +24,7 @@ else:
import termios
PROMPT = b"+"
VERSION = 5 # this tool's own version — free to drift from a loader's
VERSION = 3 # this tool's own version — free to drift from a loader's
# The loader versions this tool speaks. A pureboot version implies its wire
# protocol, which carries no number of its own, so this window is where that
# map lives: every version so far speaks the same protocol, and one that
@@ -34,38 +34,36 @@ NEWEST_LOADER = 5
SLOT = 512 # the loader slot, on every chip
RETRIES = 3 # rewrites of a page that reads back wrong, before the run stops
# pureboot 5 replaced the four per-memory commands with one pair: 'G' reads and
# 'g' writes, each taking a selector byte, a 16-bit address and a count, over
# the spaces below. The loader carries one transfer loop instead of four bodies
# which is what buys the data space and the host-issued SPM operations.
# pureboot 5 replaced the four per-memory commands with one unified pair: 'G'
# reads and 'g' writes, both taking a selector byte, a 16-bit address and a
# count, over the spaces below. The loader carries one transfer loop instead of
# four bodies, which is what paid for RAM access and for the calibration fix.
UNIFIED_LOADER = 5
SP_FLASH, SP_EEPROM, SP_RAM, SP_FUSE, SP_SPM = 0, 1, 2, 3, 4
# A selector's high nibble is the flash bank — the address bits above the 16-bit
# wire address — so a transfer names a byte address within one 64 KiB bank and
# no command has to speak word addresses. No single transfer may cross a bank
# boundary; the host chunks to keep that true.
# A selector's high nibble is flash's third address byte, so a transfer names a
# byte address within a 64 KiB bank and never has to speak word addresses. No
# single transfer may cross a bank boundary — the host chunks to keep that true.
def selector(space, address):
return space | ((address >> 16) << 4)
# The SPM operations pureboot 5 leaves to the host: a write to SP_SPM hands its
# byte to SPMCSR and fires the instruction at the selected flash address. Every
# part pureboot targets agrees on these encodings.
# data byte to SPMCSR and fires the instruction at the selected flash address.
# Every part pureboot targets agrees on these encodings.
SPM_ERASE, SPM_WRITE, SPM_RWWSRE = 0x03, 0x05, 0x11
# Calibration byte for an autobaud loader: 0xC0 is a start bit plus six zero
# Calibration byte for the autobaud loader: 0xC0 is a start bit plus six zero
# data bits — one low pulse of seven bit-times, which the loader times into its
# per-bit unit. Sent at whatever baud the host chose; the loader locks to it.
CALIBRATE = 0xC0
# pureboot 5 answers 'b' with its version and the chip signature; the host
# derives the rest of the geometry from the signature rather than reading a
# table off the device. flash, page, eeprom, patch-vector per distinct
# signature, over every chip pureboot targets (the loader computes the same
# from its chip database at build time). Die revisions that share a signature
# share this row, as they share the silicon.
CHIP_GEOMETRY = {
# The autobaud loader slims its info block to the version and signature; the host
# derives the rest of the geometry from the signature. flash, page, eeprom,
# patch-vector per distinct signature, over every chip pureboot targets (the
# loader computes the same from its chip database at build time). Die revisions
# that share a signature share this row, as they share the silicon.
AUTOBAUD_GEOMETRY = {
# signature : (flash, page, eeprom, patch_vector)
(0x1E, 0x90, 0x07): (1024, 32, 64, True), # ATtiny13/13A
(0x1E, 0x91, 0x08): (2048, 32, 128, True), # ATtiny25
@@ -146,13 +144,6 @@ class Progress:
class PosixPort:
"""A raw serial port with deadline-based reads, over termios."""
@staticmethod
def _speed(baud):
try:
return getattr(termios, f"B{baud}")
except AttributeError:
raise Error(f"unsupported baud rate {baud}") from None
def __init__(self, path, baud):
self.fd = os.open(path, os.O_RDWR | os.O_NOCTTY)
attrs = termios.tcgetattr(self.fd)
@@ -160,20 +151,14 @@ class PosixPort:
attrs[1] = 0 # oflag
attrs[2] = termios.CREAD | termios.CLOCAL | termios.CS8 # cflag
attrs[3] = 0 # lflag
attrs[4] = attrs[5] = self._speed(baud)
try:
speed = getattr(termios, f"B{baud}")
except AttributeError:
raise Error(f"unsupported baud rate {baud}") from None
attrs[4] = attrs[5] = speed
attrs[6][termios.VMIN] = 0
attrs[6][termios.VTIME] = 0
termios.tcsetattr(self.fd, termios.TCSANOW, attrs)
self.baud = baud
def set_baud(self, baud):
"""Retune the port without closing it — the fd stays open, so no DTR
pulse and no reset. That matters: the only caller is mid-session with a
loader copy that a reset would throw away."""
attrs = termios.tcgetattr(self.fd)
attrs[4] = attrs[5] = self._speed(baud)
termios.tcsetattr(self.fd, termios.TCSANOW, attrs)
self.baud = baud
def close(self):
os.close(self.fd)
@@ -302,7 +287,6 @@ if os.name == "nt":
# timeout would otherwise stay at the driver's default — which
# may be "wait forever" — until the first read.
self._deadline(_GAP_MS, 1000)
self.baud = baud
except Error:
# An open port outlives the exception otherwise, and a COM
# handle is exclusive: the next attempt would meet its own
@@ -310,21 +294,6 @@ if os.name == "nt":
self.close()
raise
def set_baud(self, baud):
"""Retune the port on its live handle — SetCommState only, so the
handle is never reopened and DTR never drops. That matters: the only
caller is mid-session with a loader copy a reset would throw away."""
if baud < 50:
raise Error(f"unsupported baud rate {baud}")
dcb = _DCB()
dcb.DCBlength = ctypes.sizeof(_DCB)
if not _k32.GetCommState(self.handle, ctypes.byref(dcb)):
_fail("cannot read the port state")
dcb.BaudRate = baud
if not _k32.SetCommState(self.handle, ctypes.byref(dcb)):
_fail(f"cannot retune the port to {baud} baud")
self.baud = baud
def close(self):
_k32.CloseHandle(self.handle)
@@ -390,26 +359,18 @@ class Info:
"""The 12-byte info block."""
@classmethod
def from_identity(cls, raw):
"""pureboot 5's reply: the version and the chip signature. The rest of
the geometry is looked up from the signature the loader derived the
same facts from its chip database at build time, so nothing is guessed,
it is simply not sent. Reconstructs a block in the older layout, so
every derived attribute below is shared with the loaders that do send
one.
The base is where application flash ends, which is a property of the
chip and not of the copy answering: a loader staged one slot lower
reports the same geometry the resident one does, exactly as the loaders
that send a block do. Which slot a copy runs in matters only to its own
write guard, which is the loader's business."""
def from_slim(cls, raw):
"""The autobaud loader's slimmed reply version and signature only —
with the rest of the geometry looked up from the signature (the loader
derived it from the same chip facts at build time). Reconstructs the full
block so every derived attribute matches the fixed-baud path exactly."""
if len(raw) != 4:
raise Error(f"bad identity reply: {raw.hex()}")
raise Error(f"bad slim info block: {raw.hex()}")
version, signature = raw[0], tuple(raw[1:4])
geometry = CHIP_GEOMETRY.get(signature)
geometry = AUTOBAUD_GEOMETRY.get(signature)
if geometry is None:
sig = " ".join(f"{b:02x}" for b in signature)
raise Error(f"unknown signature {sig} — this tool has no geometry for it")
raise Error(f"unknown signature {sig} — this tool has no autobaud geometry for it")
flash, page, eeprom, patch = geometry
base = flash - SLOT
word_flash = flash > 0x10000
@@ -432,11 +393,8 @@ class Info:
self.signature = raw[3:6]
self.page = raw[6] or 256 # the wire count convention: 0 means 256
self.patch_vector = bool(raw[11] & 1)
# Bit 1: the flash runs past what one 16-bit address covers. Through
# pureboot 4 that made flash addresses words on the wire; pureboot 5
# keeps them bytes and carries the bank in the selector instead. Every
# address in this tool stays a byte address either way and converts at
# the wire.
# Bit 1: flash addresses are words on the wire. Every address here
# stays a byte address and converts at the wire.
self.word_flash = bool(raw[11] & 2)
scale = 2 if self.word_flash else 1
self.base = (raw[7] | (raw[8] << 8)) * scale
@@ -466,7 +424,7 @@ class Info:
f"version pureboot {self.version}",
f"signature {' '.join(f'{b:02x}' for b in self.signature)}",
f"flash {self.flash_size} B, {self.page} B pages"
+ (", past one 16-bit bank" if self.word_flash else ""),
+ (", word-addressed wire" if self.word_flash else ""),
f"application 0x0000..{self.base - 1:#06x} ({self.base} B)",
f"loader {self.base:#06x} ({SLOT} B slot)",
f"staging {self.stage:#06x}",
@@ -483,78 +441,70 @@ class Loader:
def __init__(self, port):
self.port = port
self.info = None
# Set once a session is established over an autobaud link, so a
# re-entry after 'J' repeats the handshake that worked.
self.autobaud = False
# The link this session is speaking. It moves when the host follows a
# staging copy built for another one (enter_copy).
self.baud = getattr(port, "baud", None)
self._link_declared = False
def _read_identity(self):
"""The 'b' reply, in either of the two layouts a loader may send.
pureboot 5 answers with its version and the signature; older loaders
answer with a 12-byte block. The version byte cannot be mistaken for
the older block's 'P', so four bytes are enough to tell them apart."""
head = self.port.read_exact(4, 2.0)
if head[0:2] == b"PB":
return Info(head + self.port.read_exact(8, 2.0))
return Info.from_identity(head)
def _handshake(self, wait, knock, what):
"""One activation, retried until the loader answers or the window
closes. The identity reply is what proves the loader is listening — a
prompt byte alone does not, since one left over from a previous session
can still be in the pipeline while the port opening resets the device
into a fresh window, where a command without its knock is discarded.
Each attempt is therefore the whole handshake. This also converges into
an already-live session: the knock bytes are ignored there and the
drain absorbs whatever they produced."""
def connect(self, wait):
"""Knock until the info block comes back. The block is what proves the
loader is listening — a prompt byte alone does not, since one left over
from a previous session can still be in the pipeline while the port
opening resets the device into a fresh activation window, where a
command without its knock is discarded. Each attempt is therefore the
whole handshake, retried until it produces the block or the window
closes. Also converges into a live session: the knock bytes are ignored
there and the drain absorbs whatever they produced."""
deadline = time.monotonic() + wait
knocks = 0
while True:
self.port.flush_input()
self.port.write(knock)
self.port.write(b"pb")
knocks += 1
if PROMPT in self.port.read_available(0.4):
# Settle: absorb a real loader's trailing bytes before asking
# for the identity. Bounded by the deadline so a target that
# never falls quiet — a board stuck in a reset loop, whose
# garbage carries a stray prompt — cannot spin here forever.
while self.port.read_available(0.3):
if time.monotonic() > deadline:
break
pass
self.port.write(b"b")
try:
# A version the tool cannot speak is the loader's own
# answer, not a failed knock: Info reports it rather than
# sending the tool round the loop again.
self.info = self._read_identity()
except Error as failed:
if "pureboot" in str(failed):
raise
self.info = None
if self.info is not None:
block = self.port.read_exact(12, 2.0)
except Error:
block = b""
# A version the tool cannot speak is the loader's own answer,
# not a failed knock: Info reports it rather than retrying.
if block[0:2] == b"PB":
self.info = Info(block)
self._expect_prompt()
verbose(f"loader answered {what} {knocks}; identity read")
verbose(f"loader answered knock {knocks}; info block read")
return self.info
if time.monotonic() > deadline:
raise Error("no answer — reset the device within its activation window")
def connect(self, wait):
"""Knock 'p' then 'b' and read the identity."""
return self._handshake(wait, b"pb", "knock")
def connect_autobaud(self, wait):
"""The autobaud handshake. In place of the p+b knock the host sends the
calibration pulse — one seven-bit-time low pulse at the host's chosen
baud, which the loader times into its per-bit unit then a single 'p'
the loader decodes at the rate it just measured. A lost pulse, or a
knock landing while the loader is mid-frame, simply fails to answer and
leaves the measurement loop waiting for the next pulse, so the retry in
_handshake covers it."""
self.autobaud = True
return self._handshake(wait, bytes((CALIBRATE, ord("p"))), "calibration")
"""The autobaud handshake. Instead of the p+b knock, the host sends the
0xC0 calibration pulse — a single seven-bit-time low pulse at the host's
chosen baud which the loader times into its per-bit unit, then a single
'p' knock the loader decodes at the rate it just measured. As with
connect(), each attempt is the whole handshake, retried until the slim
info block comes back or the window closes: a lost pulse or a knock that
lands while the loader is mid-frame simply fails to answer, and the
loader's measurement loop is back waiting for the next pulse."""
deadline = time.monotonic() + wait
knocks = 0
while True:
self.port.flush_input()
self.port.write(bytes((CALIBRATE, ord("p"))))
knocks += 1
if PROMPT in self.port.read_available(0.4):
while self.port.read_available(0.3):
pass
self.port.write(b"b")
try:
block = self.port.read_exact(4, 2.0)
except Error:
block = b""
if len(block) == 4:
self.info = Info.from_slim(block)
self._expect_prompt()
verbose(f"loader locked on knock {knocks}; slim info read")
return self.info
if time.monotonic() > deadline:
raise Error("no answer — reset the device within its activation window")
def _expect_prompt(self, timeout=2.0):
byte = self.port.read_exact(1, timeout)
@@ -697,41 +647,11 @@ class Loader:
self.port.write(bytes((ord("J"), word_address & 0xFF, word_address >> 8)))
self._expect_prompt()
def enter_copy(self, byte_address, wait, link=None):
def enter_copy(self, byte_address, wait):
"""Jump into the loader copy at `byte_address` and knock it — a slot
base is that copy's entry stub, so it can only land there.
`link` is that copy's own `(baud, autobaud)`, for when it is not this
session's. A staging copy *is* the new image, so it speaks the rate and
backend it was built for; the host has to be told which, because 512
bytes of position-independent code carry no header to read it from.
Retuning goes through the open port, so no DTR pulse resets the copy that
is now running — and the session keeps the new link afterwards, since
every later jump lands in the same image.
"""
baud, autobaud = link if link is not None else (self.baud, self.autobaud)
if link is not None:
self._link_declared = True
base is that copy's entry stub, so it can only land there."""
self.jump(byte_address // 2)
if baud is not None and baud != self.baud:
self.port.set_baud(baud)
self.baud = baud
self.autobaud = autobaud
try:
return self.connect_autobaud(wait) if autobaud else self.connect(wait)
except Error as unheard:
if self._link_declared:
raise
# The bare activation timeout sends the operator to look at wiring,
# while on a patched-vector part the application region is already
# gone. Name the one cause that fits: the copy answers on its own
# link, not the resident's.
raise Error(
f"the copy at {byte_address:#06x} did not answer on this session's "
f"link ({baud} Bd, {'autobaud' if autobaud else 'fixed baud'}). An "
f"image built for another baud or backend speaks that one instead — "
f"say which with --staged-baud / --staged-autobaud"
) from unheard
return self.connect(wait)
def run_application(self):
self.jump(self.info.app_entry_word)
@@ -898,26 +818,12 @@ def mega_boot(info, fuse_bytes):
def image_info(image):
"""What a pureboot binary says about itself, or None.
An update image is a bare slot: nothing about it names the chip it was
built for, and installing a foreign one bricks the target — so every
loader carries a stamp for this. Through pureboot 4 the stamp is the
12-byte info block the device also serves; pureboot 5 serves its identity
from immediates and carries a 6-byte stamp (magic, version, signature)
that only this exists for, from which the geometry is looked up exactly as
it is for a live device.
Searched once per known version, so the magic stays three selective bytes
rather than two that code could carry by chance."""
"""The info block embedded in a pureboot binary, or None. Searched once
per known version, so the magic stays three selective bytes rather than
two that code could carry by chance."""
for version in range(OLDEST_LOADER, NEWEST_LOADER + 1):
at = image.find(b"PB" + bytes((version,)))
if at < 0:
continue
if version >= UNIFIED_LOADER:
if at <= len(image) - 6:
return Info.from_identity(image[at + 2 : at + 6])
elif at <= len(image) - 12:
if 0 <= at <= len(image) - 12:
return Info(image[at : at + 12])
return None
@@ -1079,16 +985,10 @@ def patch_word0(loader, page0, target_base):
return bytes(patched)
def op_update_loader(loader, wait, path, state_path, fuse_bytes, staged_link=None):
def op_update_loader(loader, wait, path, state_path, fuse_bytes):
"""Replace the resident loader with `path`, using the loader as its own
staging loader. Every phase is idempotent and keyed off the flash state,
so a re-run resumes; the state file carries what the staging slot held.
`staged_link` is the new image's own `(baud, autobaud)` where it differs from
this session's — the copies the host enters *are* that image, so they answer
on its link and not the resident's. Note what this does to the idempotence
above: once the staging copy is installed, the resumable state is only
reachable on the new link, so a re-run has to name it too."""
so a re-run resumes; the state file carries what the staging slot held."""
info = loader.info
image = loader_image(path)
for warning in update_preflight(image, info, fuse_bytes):
@@ -1113,10 +1013,7 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes, staged_link=Non
# it and matching byte for byte, and the slot unchanged since this update
# began, so a half-written install takes the path below instead.
current = loader.read_flash(info.stage, SLOT)
# The whole slot is searched: a loader's stamp sits wherever its image put
# it, which is the end of the code on pureboot 5 and the front of it
# before that.
staged_loader = image_info(current)
staged_loader = image_info(current[:268])
if staged_loader is not None and staged_loader.raw == info.raw and current == state.staging:
print("staging slot already holds a loader — left in place")
else:
@@ -1133,7 +1030,7 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes, staged_link=Non
# routes through the resident, word 0 is re-aimed at the staging copy for
# the rewrite, so a power loss mid-rewrite still resets into a loader.
verbose(f"entering the staging copy at {info.stage:#06x}")
loader.enter_copy(info.stage, wait, link=staged_link)
loader.enter_copy(info.stage, wait)
redirect = info.patch_vector and info.stage != 0
if redirect:
verbose("word 0 re-aimed at the staging copy for the rewrite")
@@ -1370,15 +1267,6 @@ def main():
parser.add_argument("--fuses", action="store_true", help="read the fuse and lock bytes")
parser.add_argument("--update-loader", metavar="FILE", help="replace the loader with this pureboot binary")
parser.add_argument("--state", metavar="FILE", help="update state file (default: FILE.pbstate)")
# The update enters the staging copy, which is the new image and so speaks
# the link *it* was built for. Nothing in the image says which, so where it
# differs from this session's these name it and the host follows.
parser.add_argument("--staged-baud", metavar="BD", type=int,
help="the baud the --update-loader image was built for, where it "
"differs from --baud")
parser.add_argument("--staged-autobaud", action=argparse.BooleanOptionalAction, default=None,
help="whether that image is an autobaud build, where it differs "
"from --autobaud")
parser.add_argument("--assume-fuses", metavar="HEX8", help="fuse bytes low,lock,ext,high as 8 hex digits "
"(overrides reading them — e.g. under a simulator that cannot)")
parser.add_argument("--erase-flash", action="store_true", help="0xff over the application flash")
@@ -1428,14 +1316,7 @@ def main():
fuse_bytes = read
if args.update_loader:
state = args.state or args.update_loader + ".pbstate"
staged_link = None
if args.staged_baud is not None or args.staged_autobaud is not None:
staged_link = (
args.staged_baud if args.staged_baud is not None else args.baud,
args.staged_autobaud if args.staged_autobaud is not None else args.autobaud,
)
op_update_loader(loader, args.wait, args.update_loader, state, fuse_bytes,
staged_link)
op_update_loader(loader, args.wait, args.update_loader, state, fuse_bytes)
if args.flash:
op_flash(loader, args.flash, args.erase_flash, not args.no_verify, fuse_bytes, args.force)
elif args.erase_flash:

View File

@@ -0,0 +1,396 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 530 B on the 1284P against a 512 B slot. Its 4 B of margin was
// never spare capacity; it was the space the missing fix should have occupied.
// pureboot_autobaud_uni.cpp replaces it at 464 B with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE PURE VERSION — no inline assembly, no global register variables,
// exactly the constraints the fixed-baud loader keeps.
//
// Two source files exist for review (autobaud.md):
// this one, pure, and pureboot_autobaud_reg.cpp, which keeps the running-slot
// write guard at the cost of one global register variable. They differ only in
// where the measured unit lives and whether the guard is present.
//
// What this version trades to fit 512 B in pure C++ (the 1284 at 508), each
// licensed by "the host guarantees safety" (README.md) and the owner's approval
// to simplify the info block:
// - the measured per-bit unit lives in the two general-purpose I/O scratch
// registers (GPIOR) where the chip has them, in a static otherwise —
// reached through libavr's named register surface, so no asm and no global
// register variable; the loader stays pure;
// - the info block is slimmed to the version and the signature — the chip's
// identity — from which the host derives page size, loader base, EEPROM
// size and the addressing flags via its own chip database;
// - no running-slot write guard: the host never programs the loader's own
// slot, and a broken host bricking the target is the host's bug;
// - a single-byte activation knock: the calibration pulse already proves a
// host is present.
//
// Position independence is kept and is in fact total here: control flow is
// PC-relative, the wire carries addresses, and with the slimmed info block and
// no write guard nothing anchors on the runtime address at all.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 4;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations). It lives
// in the two adjacent general-purpose I/O scratch registers (GPIOR1:GPIOR2)
// where the chip has them — in/out reach them in one word where a static's
// lds/sts take two, and there is no .bss to clear — and in a plain static
// otherwise (the t13, m8 and m16/32 have no GPIOR). Both are pure: the named
// register surface, no inline asm, no global register variable.
constexpr bool have_gpior = hw::db.reg_index("GPIOR1") >= 0 && hw::db.reg_index("GPIOR2") >= 0;
std::uint16_t unit_backing;
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
if constexpr (Gpior)
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
else
return unit_backing;
}
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
if constexpr (Gpior) {
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
} else
unit_backing = u;
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program. No running-slot write
// guard: the host guarantees it never targets the loader's own slot (the pure
// version's one dropped safety net, licensed — README.md).
void program_flash(std::uint16_t wire_address)
{
spm::flash_address_t address;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
}
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop — its cursor then lives in the loop's own call-saved
// registers instead of being saved and restored around a call
// (the call-site-count lesson, autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16());
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -0,0 +1,371 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 534 B on the 1284P against a 512 B slot, so the global register
// variable it broke purity for buys nothing. pureboot_autobaud_uni.cpp replaces
// it at 464 B, strictly pure and with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE REGISTER VERSION — keeps the running-slot write guard, at the cost
// of one global register variable (r4) holding the measured unit. That variable
// is pureboot's single, deliberate break from its no-global-register-variable
// rule, present only in this variant; everything else stays pure C++.
//
// Two source files exist for review (autobaud.md):
// pureboot_autobaud_pure.cpp, fully pure but dropping the write guard, and this
// one. They differ only in where the measured unit lives (a call-saved register
// here, GPIOR/RAM there) and whether the guard is present.
//
// The register buys ~32 B over a RAM home — an outlined rx/tx reads it with one
// move where a static costs an lds — and that is what lets the write guard stay
// while the image still fits 512 B (the 1284 at 512, exactly). The unit is
// written once through a noinline setter so the store lands immediately before a
// ret: GCC otherwise deletes a global-register store whose only readers are
// callees (autobaud.md, upstream bug 6).
//
// Simplifications shared with the pure version, each licensed: a slimmed info
// block (version + signature; the host derives geometry from its chip database)
// and a single-byte activation knock (the calibration pulse already proves a
// host). Position independence is kept: control flow is PC-relative and the
// write guard anchors on the runtime return address, as the fixed-baud loader.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations) in a
// call-saved register that the serial callees read directly, so no wire path
// threads it and rx/tx reach it with a move, not a load. Written only through
// set_unit() below.
register std::uint16_t g_unit asm("r4");
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock; only the software-UART pins are a parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint8_t version = 4;
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The one store into g_unit, isolated so it lands right before the ret: a
// global-register store whose only later readers are callees is dropped
// otherwise (autobaud.md, upstream bug 6).
[[gnu::noinline]] void set_unit(std::uint16_t v)
{
g_unit = v;
}
// The autobaud software link: bit-banged with cycle-counted delays, but the
// per-bit delay is g_unit, measured from the host's calibration pulse.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into g_unit. The host sends 0xC0 — a start bit
// plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so count >> 2 is the bit
// period in _delay_loop_2's four-cycle iterations. 0 means the budget
// expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
set_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
_delay_loop_2(static_cast<std::uint16_t>(g_unit + (g_unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(g_unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
out_t::clear(); // start bit
_delay_loop_2(g_unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(g_unit);
}
out_t::set(); // stop bit
_delay_loop_2(g_unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program — except the slot this
// code is running in (`slot_high`, from run()), which is drained and left
// alone. A broken host therefore cannot brick the running loader, and a copy one
// slot lower may rewrite the resident one.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
{
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
}
if (page_high != slot_high) {
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// The high byte of the slot this copy runs at, which the write guard
// follows: the return address is a word address, so its high byte is the
// 256-word slot index, doubled back into byte terms on a byte-addressed
// chip. Taken as byteswap's low byte — the builtin already swaps the two
// stacked bytes, and the double swap folds away.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
// Measure the calibration pulse into g_unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop (autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -0,0 +1,411 @@
// pureboot autobaud, the unified-primitive version — VARIANT A, an explicit
// space byte. One read command and one write command carry a space selector, so
// flash, EEPROM, RAM and the fuses share a single cursor, a single transfer loop
// and a single argument decode instead of one command body each.
//
// Strictly pure: no inline assembly, no global register variables, and no GPIOR
// either — the measured unit lives in a plain static, so the loader claims no
// chip resource an application might want. (PUREBOOT_UNIT_GPIOR=1 puts it back
// in the I/O scratch registers, kept only as a measurement axis.)
//
// Against pureboot_autobaud_pure.cpp this version:
// - adds RAM read and write, which the loader has never had. Because AVR maps
// the register file and the whole I/O space into the data address space,
// that one space also gives the host arbitrary peripheral access for free;
// - collapses 'R' (read flash), 'r' (read EEPROM), 'w' (write EEPROM) and 'F'
// (fuses) — four bodies, four loops — into 'G' and 'P' over four spaces;
// - fixes the activation hang: a lone calibration pulse used to leave the
// loader blocked forever in the knock's rx(), so a stray edge on an
// unattended device wedged it in the loader and the application never ran.
//
// Position independence is kept and is total: control flow is PC-relative, the
// wire carries addresses, and nothing anchors on the runtime address.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The unit's home. A plain static by default — pureboot claims no GPIOR, so the
// application keeps both scratch registers. The GPIOR spelling is retained
// behind a macro purely so the two can be measured against each other.
#if !defined(PUREBOOT_UNIT_GPIOR)
#define PUREBOOT_UNIT_GPIOR 0
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 5;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The two SPM commands the loader still has to recognise by value, on the chips
// where it cannot issue a runtime one generically. Taken from the chip's own
// definitions rather than spelled 3 and 5 — though every part pureboot targets
// agrees on those, which is what lets the host send the raw SPMCSR byte.
constexpr std::uint8_t spm_erase = __BOOT_PAGE_ERASE;
constexpr std::uint8_t spm_write = __BOOT_PAGE_WRITE;
// The spaces a transfer can name. Flash is 0 so it is the cheap default.
//
// sp_spm is the one that is not memory: a write there hands its data byte to
// SPMCSR and fires the instruction at the given flash address, so page erase,
// page write and RWW re-enable become host-issued commands instead of a
// hardcoded tail inside 'W'. The store side already owns an address, a data
// byte and an ack, so the whole sequence costs only the fused out/spm pair. It
// also lets the host reach every other SPM operation — lock bits included —
// which the loader previously had no way to expose.
enum : std::uint8_t { sp_flash = 0, sp_eeprom = 1, sp_ram = 2, sp_fuse = 3, sp_spm = 4 };
// A transfer's selector byte is `space | bank << 4`: the low nibble names the
// space, the high nibble carries flash's third address byte (RAMPZ) on the
// chips that have one. Putting the bank here rather than widening the address
// keeps the shared cursor sixteen bits for every space — a three-byte cursor
// costs its extra increment on EEPROM and RAM reads too, which never need it.
// The host must not span a bank boundary in one transfer; it already chunks by
// page, so nothing it does today comes close.
// .noinit, not .bss: the unit is always measured before it is read, so it needs
// no zeroing — and a zeroed .bss would drag in __do_clear_bss, 18 bytes of
// startup code for a variable that is written before its first use.
[[gnu::section(".noinit")]] std::uint16_t unit_backing;
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
#if PUREBOOT_UNIT_GPIOR
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
#else
return unit_backing;
#endif
}
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
#if PUREBOOT_UNIT_GPIOR
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
#else
unit_backing = u;
#endif
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
//
// The shape of the counting loop is load-bearing, not incidental: the >> 2
// is exact only while the pulse's bit-count equals the loop's cycles per
// iteration. Both are 7 here (sbis 1 + rjmp 2 + adiw 2 + rjmp 2). Reshaping
// this loop silently changes the lock; test/pbautobaud.py is what pins it.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. That is the whole hang fix: no wait during activation is unbounded,
// so a stray calibration pulse can no longer wedge the loader.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
// One byte out of any space. The four accessors share the cursor, the loop and
// the call site — the whole point of the unified commands — so each costs only
// its own instruction rather than a body, a loop and a dispatch arm.
[[gnu::always_inline]] inline std::uint8_t load(std::uint8_t space, std::uint8_t bank, std::uint16_t at)
{
if (space == sp_eeprom)
return ee::read(at);
if (space == sp_ram)
return *reinterpret_cast<volatile std::uint8_t *>(at);
if (space == sp_fuse)
return spm::read_fuse<off>(static_cast<spm::fuse>(at));
if constexpr (word_flash)
return avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(bank) << 16) | at);
else
return avr::flash_load(reinterpret_cast<const std::uint8_t *>(at));
}
// One byte into a writable space. Flash is not one of them — it arrives a page
// at a time through 'W' — and the fuses are not writable at all.
[[gnu::always_inline]] inline void store(std::uint8_t space, [[maybe_unused]] std::uint8_t bank, std::uint16_t at,
std::uint8_t value)
{
if (space == sp_ram) {
*reinterpret_cast<volatile std::uint8_t *>(at) = value;
return;
}
if (space == sp_spm) {
// The fused store-and-fire: SPMCSR takes the data byte and the SPM
// issues against Z in the same unscheduled pair the hardware's
// four-cycle window demands, which is exactly why this is one
// primitive and not a poke of SPMCSR followed by a poke of anything
// else. A host cannot hit that window across a serial link.
// The preprocessor rather than `if constexpr` only because
// spm::detail::page_command does not exist at all where there is no
// RAMPZ, and a discarded constexpr branch outside a template is still
// name-checked. A generic spm::command() in libavr would let the
// runtime command through on every chip and retire the dispatch below.
#if defined(RAMPZ)
spm::detail::page_command(value, (static_cast<spm::flash_address_t>(bank) << 16) | at);
#else
if (value == spm_erase)
spm::erase_page<off>(at);
else if (value == spm_write)
spm::write_page<off>(at);
else if constexpr (boot_section)
spm::rww_enable<off>();
#endif
if constexpr (boot_section)
spm::wait();
return;
}
ee::write<off>(at, value);
}
// One page into the SPM buffer, and only that: the erase, the write and the RWW
// re-enable that used to follow are now three host-issued writes to sp_spm,
// which reach the same fused out/spm pair through the store path's own address
// and data. No running-slot write guard: the host guarantees it never targets
// the loader's own slot (licensed — README.md).
void program_flash([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
{
std::uint16_t z = at & ~static_cast<std::uint16_t>(page - 1);
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
#if defined(RAMPZ)
spm::fill<off>((static_cast<spm::flash_address_t>(bank) << 16) | z, word_of({low, high}));
#else
spm::fill<off>(z, word_of({low, high}));
#endif
z += 2;
} while (static_cast<std::uint8_t>(z) & (page - 1));
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock —
// both under the poll budget. An expired budget (no host) boots the
// application; anything but 'p', including the knock timing out, re-measures
// and so returns to the budgeted wait that boots it.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'W': // program one flash page: sel8, addr16, then page bytes
case 'G': // read: sel8, addr16, n8 (0 = 256)
case 'g': { // write: sel8, addr16, n8, then n bytes each acked
// The unified transfer. One decode, one cursor, one loop for every
// space and both directions — the four command bodies this replaces
// each carried their own copy of all three. Read and write are the
// same letter in the two cases, so the direction is bit 5 of the
// command and the loop tests it with a one-word skip. 'W' joins the
// same selector-and-address decode rather than keeping a word
// address of its own, which makes flash addressing uniform across
// every command that names it and costs nothing to share.
const std::uint8_t sel = link::rx();
const std::uint8_t space = sel & 0x0f;
const std::uint8_t bank = static_cast<std::uint8_t>(sel >> 4);
std::uint16_t at = rx16();
if (command == 'W') {
program_flash(bank, at);
break;
}
std::uint8_t count = link::rx();
do {
if (command & 0x20) {
store(space, bank, at, link::rx());
tx_ack();
} else
link::tx(load(space, bank, at));
++at;
} while (--count);
break;
}
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,75 +1,47 @@
#!/usr/bin/env python3
"""Position-independence lint: the property that lets the identical image run
from any slot, asserted from the built ELF and its object.
"""Position-independence lint: the two link-time facts that let the identical
image run from any slot, asserted from the built ELF.
1. No absolute jmp/call — -mrelax normally guarantees it, but a branch that
grows out of relaxation range would break it silently.
2. Nothing flash-resident to address: the image is .text alone, so there is
no table whose runtime address has to be reconstructed.
3. The image is byte-identical when linked at a different base. This is
position independence itself rather than a proxy for it — an absolute
address anywhere in the image would move with the link and show up as a
differing byte.
2. The info block within the image's first 256 bytes: 'b' rebuilds its
address as (running slot high byte : link address low byte).
Usage: check_pi.py <objdump> <objcopy> <cxx> <mcu> <elf> <object> <text_start_hex>
Usage: check_pi.py <objdump> <nm> <elf> <text_start_hex>
"""
import os
import re
import subprocess
import sys
import tempfile
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
objdump, objcopy, cxx, mcu, elf, obj, text_start = sys.argv[1:]
objdump, nm, elf, text_start = sys.argv[1:]
text_start = int(text_start, 0)
listing = subprocess.run([objdump, "-d", elf], capture_output=True, text=True, check=True).stdout
absolute = [line for line in listing.splitlines() if re.search(r"\t(jmp|call)\t", line)]
absolute = [
line
for line in listing.splitlines()
if re.search(r"\t(jmp|call)\t", line)
]
if absolute:
fail("absolute control flow in the image:\n" + "\n".join(absolute))
print("FAIL: absolute control flow in the image:")
print("\n".join(absolute))
sys.exit(1)
# Allocated flash beyond .text would be data the running copy has to find.
# Only ALLOC sections reach the device at all; .comment and the debug
# sections ride along in the ELF container and are never flashed. objdump
# prints each section's flags on the line following its header.
headers = subprocess.run([objdump, "-h", elf], capture_output=True, text=True, check=True).stdout.splitlines()
for index, line in enumerate(headers):
fields = line.split()
if len(fields) < 6 or not fields[0].isdigit():
continue
name, size = fields[1], int(fields[2], 16)
flags = headers[index + 1] if index + 1 < len(headers) else ""
if "ALLOC" not in flags or not size:
continue
if name not in (".text", ".noinit", ".bss"):
fail(f"flash-resident section {name} ({size} bytes): the image must be .text alone")
symbols = subprocess.run([nm, "-C", elf], capture_output=True, text=True, check=True).stdout
info = [line for line in symbols.splitlines() if "flash_table" in line and "::storage" in line]
if len(info) != 1:
print(f"FAIL: expected one info-block storage symbol, found {len(info)}")
sys.exit(1)
address = int(info[0].split()[0], 16)
offset = address - text_start
if not 0 <= offset < 256:
print(f"FAIL: info block at image offset {offset:#x}, must sit in the first 256 bytes")
sys.exit(1)
# Relink at a different base and compare the bytes.
with tempfile.TemporaryDirectory() as work:
elsewhere = text_start - 0x200 if text_start >= 0x200 else text_start + 0x200
images = []
for base, tag in ((text_start, "here"), (elsewhere, "there")):
relinked = os.path.join(work, f"{tag}.elf")
binary = os.path.join(work, f"{tag}.bin")
subprocess.run(
[cxx, f"-mmcu={mcu}", "-nostartfiles", f"-Wl,--section-start=.text={base:#x}",
"-Wl,--defsym=pureboot_app=0", "-mrelax", obj, "-o", relinked],
check=True, capture_output=True)
subprocess.run([objcopy, "-O", "binary", relinked, binary], check=True)
images.append(open(binary, "rb").read())
if images[0] != images[1]:
differing = [i for i, (a, b) in enumerate(zip(*images)) if a != b]
fail(f"the image changes when linked at {elsewhere:#x} instead of {text_start:#x}: "
f"{len(differing)} byte(s) differ, first at offset {differing[0]:#x}")
print(f"PI lint: control flow PC-relative, .text only, identical linked at {text_start:#x} and {elsewhere:#x}")
print(f"PI lint: control flow PC-relative, info block at offset {offset:#x}")
if __name__ == "__main__":

View File

@@ -65,15 +65,6 @@ int main(int argc, char *argv[])
fprintf(stderr, "device: cannot read %s\n", argv[1]);
return 1;
}
// An image that runs past flash end cannot execute on hardware, and a
// naive copy of it would smash the heap beyond avr->flash — after which
// the simulation misbehaves in ways that point everywhere but here.
// Refuse it loudly instead.
if (boot_base + fw.flashsize > avr->flashend + 1) {
fprintf(stderr, "device: %u B at 0x%x runs past flash end 0x%x — image does not fit its slot\n",
(unsigned)fw.flashsize, boot_base, avr->flashend);
return 1;
}
memcpy(avr->flash + boot_base, fw.flash, fw.flashsize);
avr->pc = boot_base;
avr->codeend = avr->flashend;

View File

@@ -11,10 +11,6 @@
// reset reaches those loaders through the patched vector (or the runner
// models BOOTRST), so the application owes them nothing.
//
// PUREBOOT_HANDOVER drops the listening and jumps straight in, leaving the
// USART enabled behind it — the hand-over state a loader bit-banging on that
// USART's own pins has to survive.
//
// The fixture speaks the deployment its loader was built for: the same
// PUREBOOT_* defines configure it, and without them it assumes the stock
// deployment (the crystal/RC clock table below, the chip's natural link).
@@ -68,29 +64,16 @@ struct link {
{
tx_t::write(static_cast<std::uint8_t>(c));
}
// The loader sits in the top slot — 512 bytes on every chip. The jump
// takes a word address, which is what makes the >64 KiB chips' entry
// reachable through a 16-bit pointer at all.
static void enter_loader()
{
constexpr std::uint32_t slot = 512;
reinterpret_cast<void (*)()>(static_cast<std::uint16_t>((avr::hw::db.mem.flash_size - slot) / 2))();
}
[[noreturn]] static void idle()
{
#if defined(PUREBOOT_HANDOVER)
// Hand back at once, with this USART still enabled — the state that
// leaves a bit-banged loader on its pins mute unless the loader
// releases it. Unconditional because there is no command wire to
// wait on: that loader's link is the pins, not this peripheral.
enter_loader();
__builtin_unreachable();
#else
// 'L' hands back to the loader in the top slot — 512 bytes on every
// chip. The jump takes a word address, which is what makes the
// >64 KiB chips' entry reachable through a 16-bit pointer at all.
constexpr std::uint32_t slot = 512;
for (;;) {
auto command = tx_t::read_blocking();
if (command == 'L')
enter_loader();
reinterpret_cast<void (*)()>(static_cast<std::uint16_t>((avr::hw::db.mem.flash_size - slot) / 2))();
// 'D' leaves every word of the SPM page buffer dirty, so that a
// following 'L' enters the loader with the buffer it never clears.
if (command == 'D') {
@@ -99,7 +82,6 @@ struct link {
tx('D');
}
}
#endif
}
};
@@ -117,27 +99,8 @@ struct link<C, false> {
}
[[noreturn]] static void idle()
{
#if defined(PUREBOOT_HEARTBEAT)
// Repeat the banner forever, which turns the fixture into a fixed
// cycles-per-bit transmitter: `tools/pbrig.py rate` sweeps the host rate
// against it to find the part's true bit rate, and from that the clock
// its RC oscillator is really running at. Only the *bit* timing carries
// the measurement — the delay merely spaces the lines out, so its own
// error does not matter. Software link only: the hardware-link idle owes
// the self-update tests a command loop, and a crystal deployment has
// nothing to measure.
while (true) {
tx('A');
tx('P');
tx('P');
tx('\r');
tx('\n');
dev::delay<50_ms>();
}
#else
while (true) {
}
#endif
}
};
@@ -146,13 +109,8 @@ struct link<C, false> {
int main()
{
avr::init<typename link<dev::clock>::tx_t>();
#if !defined(PUREBOOT_HANDOVER)
link<dev::clock>::tx('A');
link<dev::clock>::tx('P');
link<dev::clock>::tx('P');
#endif
// The hand-over fixture stays silent: nothing is listening on the USART it
// brings up — the loader it hands to speaks those pins directly — so its
// banner would be a write into a peer that does not exist.
link<dev::clock>::idle();
}

View File

@@ -141,45 +141,13 @@ def main():
print(f" {label}: locked at {hz} Hz / {baud} Bd, flash+EEPROM verified"
+ (", hand-over ok" if hand_over else ""))
def must_lock(hz, baud, label):
"""The calibration alone, at a tight bit period. Nothing is programmed —
the question is only whether the loader can still measure the pulse."""
dump = os.path.join(workdir, f"flash_{label}.bin")
device = pbsim.Device(device_bin, elf, mcu, str(hz), base_hex, page, baud, dump,
link="sw:B0,B1")
try:
port = pb.Port(device.pty, baud)
try:
live = pb.Loader(port).connect_autobaud(15)
if live.version != pb.NEWEST_LOADER:
fail(f"{label}: loader reports pureboot {live.version}")
finally:
port.close()
finally:
device.stop()
print(f" {label}: locked at {hz} Hz / {baud} Bd ({hz / baud:.0f} cycles a bit)")
# The app fixture is built for one clock; the hand-over banners there. A
# second point at double that clock, same loader binary, proves the lock is
# measured, not baked in — the whole point of autobaud. (Doubling keeps the
# bit period healthy; halving would drop it below the software UART's floor.)
round_trip(app_hz, app_baud, "clock-a", hand_over=True)
round_trip(app_hz * 2, app_baud, "clock-b", hand_over=False)
# Both points above sit near 100 cycles a bit, which is comfortable. The
# calibration's real floor is far tighter, and it is worth a gate: measured
# here, the lock is solid down to ~36 cycles a bit and fails outright by ~31
# — a sharp edge, not a fraying one. This pins the tightest standard rate the
# fixture's clock reaches, so a change that raises the floor is caught.
#
# It does *not* bound what a real deployment can use. On silicon the
# oscillator's own jitter costs roughly a factor of two: an ATtiny13A on its
# factory RC trim was reliable at ~118 cycles a bit and already locking only
# 1 attempt in 5 by ~59, which no exact-clock simulation can show. The
# deployable envelope is a README matter; this is the logic's floor.
must_lock(app_hz, app_baud * 2, "tight-bit")
print("pbautobaud: calibration lock and flash/EEPROM/fuse round-trip pass at both clocks, "
"and the tight bit period still locks")
print("pbautobaud: calibration lock and flash/EEPROM/fuse round-trip pass at both clocks")
if __name__ == "__main__":

View File

@@ -1,73 +0,0 @@
#!/usr/bin/env python3
"""Hand-over with a USART left enabled on the loader's own pins.
A software or autobaud link deployed on a USART's TxD is mute if an
application hands over with that USART still enabled: TXEN keeps the USART
owning the pin, so the bit-banged transmitter's port writes go nowhere and the
loader receives and obeys while answering nothing. The link's init releases it.
The state is reached the way silicon reaches it — an application that sets up
its USART and jumps in with no reset between, so nothing clears UCSRnB for it.
The pin ownership itself is modelled by the device runner: simavr wires a
USART through IRQs alone and never takes the pin from the port, so without
that the mute could not happen here at all (test/pureboot_device.c).
Usage: pbmute.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <app_bin> <tool_py> <workdir> <link>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, app_bin, tool, workdir, link = sys.argv[1:]
page, baud = int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
if "@" not in link:
fail(f"the link {link} names no owning USART — nothing would be under test")
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "dump.bin")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, link=link)
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
loader.connect(25)
resident = loader.info.version
# Install the fixture and let it take over. It brings up the USART
# that owns these pins and jumps straight back in.
pb.op_flash(loader, app_bin, erase=False, verify=True)
loader.run_application()
# The loader is running again with that USART enabled behind it. Only
# the release makes it audible; without it the connect times out.
loader = pb.Loader(port)
try:
loader.connect(25)
except pb.Error as error:
fail(f"the loader never answered after the hand-over — the USART still owns its TX pin ({error})")
if loader.info.version != resident:
fail(f"identity changed across the hand-over: {resident} then {loader.info.version}")
# Answering is not enough: it has to still be a working loader.
pb.verify_pages(loader, pb.plan_flash(open(app_bin, "rb").read(), loader.info))
port.close()
finally:
device.stop()
print("pbmute: a loader on a USART's own pins answers after a hand-over that left it enabled")
if __name__ == "__main__":
main()

View File

@@ -55,12 +55,6 @@ def main():
# the page byte is the wire's 0-means-256.
mega = mcu.startswith("atmega")
patch = not mega or mcu.startswith("atmega48")
# Where SRAM begins: the x8 and x4 megas push it past their extended I/O
# space, everything else starts right after the plain I/O registers. The
# loader keeps no statics and its stack sits at RAMEND, so the first SRAM
# byte is free for the data-space probe below.
classic = mcu in ("atmega8", "atmega8a", "atmega16", "atmega16a", "atmega32", "atmega32a")
ram_base = 0x0100 if mega and not classic else 0x0060
word_flash = base + pb.SLOT > 0x10000
wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0)
@@ -79,20 +73,12 @@ def main():
if needed not in out:
fail(f"session 1 output lacks {needed!r}")
# Session 2: reconnect into the live session, verify, dump, exercise
# the data space; hand over is deferred — the pty must be reopened for
# the APP banner first.
probe = "c0ffee"
# Session 2: reconnect into the live session, verify, dump, hand over
# is deferred — the pty must be reopened for the APP banner first.
out = pbsim.run_tool(tool, device.pty, baud, "--verify-flash", app_bin, "--verify-eeprom", ee_path,
"--read-flash", read_flash, "--read-eeprom", read_eeprom,
"--poke", f"{ram_base:#x}:{probe}", "--peek", f"{ram_base:#x}:3", "--stay")
"--read-flash", read_flash, "--read-eeprom", read_eeprom, "--stay")
if out.count("verify:") != 2:
fail("session 2 did not verify both memories")
# What went into SRAM must come back out of it: the data space is one
# more selector on the same transfer as flash and EEPROM, so a wrong
# selector decode would show up here and nowhere else.
if probe not in out.replace(" ", ""):
fail(f"data-space round trip at {ram_base:#x} did not read back {probe}\n{out}")
eeprom_back = open(read_eeprom, "rb").read()
if eeprom_back[: len(ee_image)] != ee_image:
@@ -124,16 +110,12 @@ def main():
f"{pb.OLDEST_LOADER}..{pb.NEWEST_LOADER}")
# A W addressed inside a page rather than at its base must still
# consume exactly one page and prompt. The loader's own slot is the
# target — the guard refuses to commit it — and the payload is
# erased-state bytes, so the probe can disturb neither the image nor
# the page buffer it leaves behind. Hand-built rather than through
# write_page(), which would follow the fill with its erase and
# write; the point here is that the fill alone consumes exactly one
# page whatever the address's low bits say.
wire = base + 1
port.write(bytes((ord("W"), pb.selector(pb.SP_FLASH, wire), wire & 0xFF, (wire >> 8) & 0xFF))
+ b"\xff" * page)
# consume exactly one page and prompt. The loader's own slot is
# the target — it is drained and never programmed — and the
# payload is erased-state bytes, so the probe can disturb neither
# the image nor the page buffer it leaves behind.
wire = wire_base + 1
port.write(bytes((ord("W"), wire & 0xFF, wire >> 8)) + b"\xff" * page)
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("unaligned W did not return to the prompt")

View File

@@ -10,9 +10,7 @@
//
// The link follows the chip's natural default (USART0 on the megas, the
// software UART on PB0/PB1 elsewhere) unless -l overrides it: `-l usart1`
// for the second instance, `-l sw:B5,B1` for a software build's RX,TX pins,
// and `-l sw:D0,D1@0` where those pins are a USART's own — see the pin
// ownership the bridge models below.
// for the second instance, `-l sw:B5,B1` for a software build's RX,TX pins.
//
// simavr's tiny cores decode the SPM opcode but attach no NVM module — SPM
// is a silent no-op (the mega's boot section has one, avr_flash). The
@@ -48,7 +46,6 @@ static int link_software;
static char uart_digit = '0';
static char sw_rx_port = 'B', sw_tx_port = 'B';
static int sw_rx_bit = 0, sw_tx_bit = 1;
static char sw_tx_owner = 0; // the USART whose TXD the software link sits on
static const char *dump_path;
static uint32_t reset_pc;
static volatile sig_atomic_t reset_requested;
@@ -64,13 +61,9 @@ static int parse_link(const char *spec)
link_software = 1;
if (spec[2] == '\0')
return 0;
char owner = 0;
int fields = sscanf(spec + 2, ":%c%d,%c%d@%c", &sw_rx_port, &sw_rx_bit, &sw_tx_port, &sw_tx_bit, &owner);
if (fields == 4 || fields == 5) {
sw_tx_owner = owner;
if (sscanf(spec + 2, ":%c%d,%c%d", &sw_rx_port, &sw_rx_bit, &sw_tx_port, &sw_tx_bit) == 4)
return 0;
}
}
return -1;
}
@@ -189,66 +182,19 @@ static avr_cycle_count_t tx_sample(avr_t *mcu, avr_cycle_count_t when, void *par
{
(void)mcu;
(void)param;
if (tx_bit < 8) {
tx_shift = (uint8_t)((tx_shift >> 1) | (tx_level ? 0x80 : 0));
if (++tx_bit < 8)
return when + bit_cycles;
/* The byte is not delivered until its stop bit has passed. A real
* receiver cannot answer sooner, and a host that did would put its
* start bit on the wire while the device is still driving the stop
* bit — which the device, transmitting, is not watching for. */
return when + bit_cycles;
}
if (write(pty_master, &tx_shift, 1) != 1)
fprintf(stderr, "device: pty write lost a byte\n");
tx_active = 0;
return 0;
}
// A USART owns its TxD pin whenever its transmitter is enabled, and the port
// register cannot drive it (§20.2 / Atmel-8271 §19.2) — which is why a
// bit-banged link deployed on those pins is mute until it clears UCSRnB.
// simavr wires a USART entirely through IRQs and never touches the port pin
// model, so the ownership does not exist there and the mute cannot happen:
// supply it, or the very state this models is untestable. The link spec's
// trailing @n names the USART; without one the pins are nobody's.
static avr_uart_t *tx_owner;
static int tx_pin_taken(void)
{
return tx_owner && avr_regbit_get(avr, tx_owner->txen);
}
// simavr leaves TXEN set in UCSRnB out of reset, where silicon clears the
// whole register (§20.11.3) — which would hand the pin to a USART no code has
// enabled, making a freshly reset chip mute for reasons hardware does not
// have. Reset it the way the datasheet does, so the ownership starts from
// nobody's and only an application that really enables the USART takes it.
static void reset_tx_owner(void)
{
if (tx_owner)
avr_regbit_clear(avr, tx_owner->txen);
}
static void find_tx_owner(void)
{
for (avr_io_t *io = avr->io_port; io; io = io->next)
if (io->kind && strcmp(io->kind, "uart") == 0 && ((avr_uart_t *)io)->name == sw_tx_owner) {
tx_owner = (avr_uart_t *)io;
reset_tx_owner();
return;
}
fprintf(stderr, "device: no USART%c to own the software link's TX pin\n", sw_tx_owner);
}
static void tx_hook(avr_irq_t *irq, uint32_t value, void *param)
{
(void)irq;
(void)param;
if (tx_pin_taken()) { // the USART holds the line; the port write goes nowhere
tx_level = 1;
return;
}
int level = value & 1;
if (!tx_active && tx_level == 1 && level == 0) { // start edge
tx_active = 1;
@@ -371,8 +317,7 @@ int main(int argc, char *argv[])
fprintf(stderr,
"usage: %s [-l link] <pureboot.elf> <mcu> <hz> <base_hex> <page> <baud> <flash_dump>"
" [reset_hex] [resume_flash]\n"
" -l link: usart0 | usart1 | sw[:B0,B1[@0]] (RX,TX, then the USART owning\n"
" them); default: the chip's own\n"
" -l link: usart0 | usart1 | sw[:B0,B1] (RX,TX); default: the chip's own\n"
" reset_hex: reset vector (default: base with a boot section, else 0)\n"
" resume_flash: raw full-flash image loaded instead of the ELF — a prior\n"
" run's dump, for power-fail resume tests\n",
@@ -458,8 +403,6 @@ int main(int argc, char *argv[])
printf("PB_PTY %s\n", uart_pty.pty.slavename);
} else {
bit_cycles = (avr->frequency + baud / 2) / baud; // matches uart.hpp's own rounding exactly
if (sw_tx_owner)
find_tx_owner();
rx_pin = avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_rx_port), (unsigned)sw_rx_bit);
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_tx_port), (unsigned)sw_tx_bit), tx_hook,
NULL);
@@ -497,7 +440,6 @@ int main(int argc, char *argv[])
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
} else {
bridge_reset();
reset_tx_owner();
}
}
if (link_software && ++since_poll >= 2000) {

View File

@@ -1,105 +0,0 @@
#!/usr/bin/env python3
"""Host-tool activation handshake: it must not hang on a flooding target.
`_handshake` drains the line after it sees a prompt, to absorb a real loader's
trailing bytes before it asks for the identity. That drain must be bounded: a
target that never falls quiet — a board stuck in a reset loop presents exactly
this, ~60 reboots/s of UART-reset garbage in which a stray 0x2b reads as a
prompt — otherwise spins the tool forever. Regression for that hang, plus a
control that a well-behaved loader still connects.
Stdlib only, no device: host-tool logic, so it runs on every chip's preset
beside pureboot.planner.
"""
import importlib.util
import pathlib
import threading
import time
PB = pathlib.Path(__file__).resolve().parents[1] / "pureboot" / "pureboot.py"
_spec = importlib.util.spec_from_file_location("pureboot", PB)
pb = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(pb)
P = F = 0
def check(name, ok):
global P, F
P, F = P + (1 if ok else 0), F + (0 if ok else 1)
print(f" [{'PASS' if ok else 'FAIL'}] {name}")
class FloodPort:
"""A line that never falls quiet: read_available always returns bytes, and
they contain a prompt. No identity ever completes."""
def flush_input(self):
pass
def write(self, data):
pass
def read_available(self, wait):
time.sleep(0.01) # a real read waits; keep the busy loop off a core
return b"+\x00\xff"
def read_exact(self, count, timeout):
raise pb.Error("no identity")
class LoaderPort:
"""A well-behaved pureboot 5: one prompt to the knock, then quiet, then the
slim identity (version 5 + m328p signature) and a closing prompt."""
def __init__(self):
self.reads = self.exacts = 0
def flush_input(self):
pass
def write(self, data):
pass
def read_available(self, wait):
self.reads += 1
return b"+" if self.reads == 1 else b"" # prompt once, then settle quiet
def read_exact(self, count, timeout):
self.exacts += 1
return b"\x05\x1e\x95\x0f" if self.exacts == 1 else b"+" # identity, then prompt
def terminates(port, wait, budget):
"""Run connect_autobaud in a thread; True if it returns/raises within
`budget` seconds rather than hanging."""
done = threading.Event()
def run():
try:
pb.Loader(port).connect_autobaud(wait)
except Exception:
pass
finally:
done.set()
threading.Thread(target=run, daemon=True).start()
return done.wait(budget)
def main():
# the hang: a flooding target must not spin the drain forever. With wait=0.5
# the whole handshake has to give up well inside a few seconds.
check("flooding target: handshake terminates, drain is bounded",
terminates(FloodPort(), wait=0.5, budget=4.0))
# the control: a real loader still connects and reads identity.
info = pb.Loader(LoaderPort()).connect_autobaud(2.0)
check("well-behaved loader still connects (version 5)", info.version == 5)
print(f"\n {P} passed, {F} failed")
return 1 if F else 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -30,14 +30,8 @@ def info_of(pb, base, page, patch, flash, signature=(0x1E, 0x93, 0x0B), word_fla
scale = 2 if word_flash else 1
wire_base = base // scale
flags = (1 if patch else 0) | (2 if word_flash else 0)
# The EEPROM size comes from the signature, as it must: pureboot 5 derives
# the whole geometry from the signature rather than sending it, so a
# synthetic block that disagreed with its own signature would describe a
# chip that cannot exist.
eeprom = pb.CHIP_GEOMETRY[signature][2]
raw = bytes((0x50, 0x42, pb.NEWEST_LOADER if version is None else version,
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8,
eeprom & 0xFF, eeprom >> 8, flags))
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8, 0, 2, flags))
info = pb.Info(raw)
if info.flash_size != flash:
fail(f"info_of({base:#x}) decodes to {info.flash_size:#x} of flash, not {flash:#x}")
@@ -168,16 +162,11 @@ def main():
fail("mega staging content should be the bare image")
expect_error("mega staging size", lambda: pb.staging_content(image + b"!", mega), "512")
# The image stamp: found in a synthetic binary, absent in noise. pureboot
# 5 stamps the magic, its version and the signature, and the geometry is
# looked up from there — so what comes back must equal what a live device
# of the same chip reports.
stamp = bytes((0x50, 0x42, pb.NEWEST_LOADER)) + bytes(tiny.signature)
binary = bytes((0xAA,)) * 10 + stamp + bytes((0xBB,)) * 10
# The embedded info block: found in a synthetic binary, absent in noise.
binary = bytes((0xAA,)) * 10 + tiny.raw + bytes((0xBB,)) * 10
found = pb.image_info(binary)
if found is None or found.raw != tiny.raw:
fail(f"image_info misreads the v{pb.NEWEST_LOADER} stamp: "
f"{found.raw.hex() if found else None} != {tiny.raw.hex()}")
fail("image_info misses the embedded block")
if pb.image_info(bytes((0xAA,)) * 40) is not None:
fail("image_info invents a block")
# An older loader's image stays readable, so a deployed build can be

View File

@@ -1,167 +0,0 @@
#!/usr/bin/env python3
"""Self-update across a link change: the host must follow the staging copy.
`--update-loader` installs the new image in the staging slot and then *enters
it* to have it rewrite the resident. That copy is the new image, so it speaks the
new image's baud and backend — but the host was talking to the *resident*. Where
the two differ, the host kept knocking at the old rate in the old mode, the
staging copy never answered, and the update stranded: staging installed, resident
untouched, and on a 1 KiB tiny the application region (which *is* the staging
slot there) already gone.
The wire cannot be probed for this — 512 bytes of position-independent code carry
no header saying what rate they were built for — so the operator declares it, and
a mismatch with nothing declared has to say so instead of reporting a bare
timeout.
Stdlib only, no device: host-tool logic, so it runs on every chip's preset beside
pureboot.planner.
"""
import importlib.util
import pathlib
PB = pathlib.Path(__file__).resolve().parents[1] / "pureboot" / "pureboot.py"
_spec = importlib.util.spec_from_file_location("pureboot", PB)
pb = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(pb)
IDENTITY = b"\x05\x1e\x95\x0f" # pureboot 5 + m328p signature
P = F = 0
def check(name, ok, detail=""):
global P, F
P, F = P + (1 if ok else 0), F + (0 if ok else 1)
print(f" [{'PASS' if ok else 'FAIL'}] {name}" + (f"{detail}" if detail else ""))
class TwoLinkPort:
"""A board whose resident and staging copy answer on different links.
Only the rate currently set decides who can be heard, which is the physical
truth: a loader's bit timing is a cycle count, so a copy built for another
rate is unreadable until the host retunes. The knock bytes carry the mode, so
a backend mismatch is caught the same way.
"""
def __init__(self, resident=(57600, False), staged=(38400, False)):
self.resident, self.staged = resident, staged
self.baud = resident[0]
self.entered = False # a 'J' has handed control to the staging copy
self.switches = [] # every retune the host asked for
self.pending = bytearray() # what the device has queued to send
# --- the part under test needs this to exist at all
def set_baud(self, baud):
self.baud = baud
self.switches.append(baud)
def flush_input(self):
self.pending.clear()
def _audible(self, knock=None):
baud, autobaud = self.staged if self.entered else self.resident
if self.baud != baud:
return False
if knock is None:
return True
return knock == (bytes((pb.CALIBRATE, ord("p"))) if autobaud else b"pb")
def write(self, data):
data = bytes(data)
if data[:1] == b"J" and len(data) == 3:
# The resident acks the jump, then control moves to the copy.
if self._audible():
self.pending += pb.PROMPT
self.entered = True
elif data in (b"pb", bytes((pb.CALIBRATE, ord("p")))):
if self._audible(data):
self.pending += pb.PROMPT
elif data == b"b":
if self._audible():
self.pending += IDENTITY + pb.PROMPT
def read_available(self, wait):
out, self.pending = bytes(self.pending), bytearray()
return out
def read_exact(self, count, timeout):
if len(self.pending) < count:
raise pb.Error(f"timeout: got {len(self.pending)} of {count} bytes")
out, self.pending = bytes(self.pending[:count]), self.pending[count:]
return out
def connected(port):
"""A Loader already in session with the resident."""
loader = pb.Loader(port)
loader.connect(2.0)
return loader
def main():
# The control first: where the staged image keeps the resident's link, the
# flow works and needs no retune. This is the case that always passed, and
# it is what made the bug look like "self-update is broken" rather than
# "self-update cannot change the link".
port = TwoLinkPort(resident=(57600, False), staged=(57600, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0)
check("same link: staging copy entered", True)
except pb.Error as error:
check("same link: staging copy entered", False, str(error))
# A baud change, declared. The host must retune before knocking.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(38400, False))
check("baud change declared: entered after retuning", 38400 in port.switches,
f"switches={port.switches}")
except (pb.Error, TypeError) as error:
check("baud change declared: entered after retuning", False, repr(error))
# A backend change, declared: the knock itself has to become the calibration
# pulse, or an autobaud staging copy never hears a thing.
port = TwoLinkPort(resident=(57600, False), staged=(57600, True))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(57600, True))
check("backend change declared: entered as autobaud", True)
except (pb.Error, TypeError) as error:
check("backend change declared: entered as autobaud", False, repr(error))
# Nothing declared against a changed link: it still cannot work, but the
# error has to name the cause. A bare "no answer" sent the operator looking
# at the wiring while the application region sat erased.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 0.3)
check("undeclared mismatch: reported", False, "unexpectedly succeeded")
except pb.Error as error:
text = str(error).lower()
check("undeclared mismatch: error names the link, not just a timeout",
"link" in text or "baud" in text or "backend" in text, str(error))
except TypeError as error:
check("undeclared mismatch: error names the link, not just a timeout",
False, repr(error))
# The resident's own link must be restored for the caller: a declared
# staging link is for the copy, and the tool talks to the new resident after.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(38400, False))
check("session records the link it is now speaking", loader.baud == 38400,
f"loader.baud={getattr(loader, 'baud', None)}")
except (pb.Error, TypeError, AttributeError) as error:
check("session records the link it is now speaking", False, repr(error))
print(f"\n {P} passed, {F} failed")
return 1 if F else 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -1,210 +0,0 @@
#!/usr/bin/env python3
"""Hardware acceptance suite for a pureboot deployment.
`tools/check.sh` proves the protocol under simavr on every chip. This proves one
*board*: that the loader actually installed on it answers, that the memories
round-trip over the real link, that the application it flashes runs afterwards,
and that the refusals which keep a 512-byte slot alive still fire. Run it once
when a board is brought up, and again whenever the deployment moves — a new
clock, a new backend, new pins.
Every check derives its bounds from the info block the loader itself reports, so
nothing here is per-chip: the same run covers a 1 KiB tiny whose application
region is 510 usable bytes and a 128 KiB mega whose flash needs a bank in the
selector.
**This overwrites the board's application flash and EEPROM.** Capture them first
with `pbrig.py backup`, which verifies what it captured.
tools/pbhw.py --programmer atmelice_isp --part t13 --port COM6 \
--autobaud --loader build/ab.bin --app build/pbapp.hex \
--marker APP
"""
from __future__ import annotations
import argparse
import pathlib
import sys
import tempfile
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
import pbrig # noqa: E402
class Suite:
def __init__(self, rig: pbrig.Rig, work: pathlib.Path):
self.rig = rig
self.work = work
self.results: list[tuple[str, bool, str]] = []
def check(self, name: str, ok: bool, detail: str = "") -> bool:
self.results.append((name, ok, detail))
print(f" {'PASS' if ok else 'FAIL'} {name}" + (f" {detail}" if detail else ""))
return ok
@staticmethod
def _brief(text: str, limit: int = 78) -> str:
return " | ".join(l.strip() for l in text.splitlines() if l.strip())[:limit]
# ----------------------------------------------------------------- checks
def identity(self) -> object | None:
"""The info block, which every later check takes its bounds from."""
module = pbrig.load_pureboot(self.rig.d.pureboot)
self.rig.reset()
port = module.Port(self.rig.d.port, self.rig.d.baud)
try:
loader = module.Loader(port)
if self.rig.d.autobaud:
loader.connect_autobaud(self.rig.d.wait)
else:
loader.connect(self.rig.d.wait)
info = loader.info
self.check("identity read", True, info.describe())
return info
except Exception as error: # noqa: BLE001 — a dead link is a result
self.check("identity read", False, str(error)[:70])
return None
finally:
try:
port.close()
except Exception:
pass
def eeprom(self, info) -> None:
size = info.eeprom_size
if not size:
print(" skip EEPROM (this part has none)")
return
# A pattern no erase or partial write could produce by accident.
pattern = bytes((i * 7 + 3) & 0xFF for i in range(size))
image = self.work / "ee.bin"
image.write_bytes(pattern)
rc, out = self.rig.pureboot("--eeprom", str(image), "--verify-eeprom", str(image))
self.check(f"EEPROM write + verify ({size} B)", rc == 0, self._brief(out))
back = self.work / "ee-back.bin"
rc, out = self.rig.pureboot("--read-eeprom", str(back))
got = back.read_bytes() if back.exists() else b""
self.check("EEPROM reads back what was written", got == pattern, f"{len(got)} B")
self.rig.pureboot("--erase-eeprom")
erased = self.work / "ee-erased.bin"
self.rig.pureboot("--read-eeprom", str(erased))
got = erased.read_bytes() if erased.exists() else b""
self.check("EEPROM erase leaves 0xff", got == b"\xff" * size, f"{len(got)} B")
def application(self, info, app: pathlib.Path, marker: str) -> None:
rc, out = self.rig.pureboot("--flash", str(app), "--verify-flash", str(app))
self.check(f"application flash + verify ({app.name})", rc == 0, self._brief(out))
if marker:
# The tool hands over as it ends its session, so the application is
# already running; opening the port does not reset a board whose DTR
# is unwired, so this simply listens.
data = self.rig.capture(seconds=2.5)
seen = marker.encode() in data
sample = "".join(chr(b) if 32 <= b < 127 else "." for b in data[:40])
self.check(f"application runs (emits {marker!r})", seen, f"|{sample}|")
back = self.work / "app-back.bin"
rc, out = self.rig.pureboot("--read-flash", str(back))
got = back.read_bytes() if back.exists() else b""
self.check("application flash reads back", rc == 0 and len(got) == info.base,
f"{len(got)} B of {info.base}")
def erase_and_guard(self, info, loader_image: pathlib.Path | None) -> None:
rc, out = self.rig.pureboot("--erase-flash")
self.check("application region erases", rc == 0, self._brief(out))
# The slot must be untouched by an application erase, which only an
# independent read can show — so this one goes over ISP, not the link.
whole = self.work / "whole.bin"
if not self.rig.read_memory("flash", whole, "r"):
self.check("loader slot survives the erase", False, "ISP read failed")
return
image = whole.read_bytes()
image += b"\xff" * (info.flash_size - len(image))
# Erased application flash, up to the trampoline word the host composes
# on a patched-vector part.
limit = info.base - 2 if info.patch_vector else info.base
self.check("erased application region is 0xff",
set(image[0:limit]) <= {0xFF}, f"0x0000..{limit:#06x}")
if loader_image and loader_image.exists():
want = loader_image.read_bytes()
got = image[info.base:info.base + len(want)]
self.check("loader slot survives the erase", got == want,
f"{len(want)} B at {info.base:#06x}")
else:
print(" skip loader slot comparison (pass --loader <image.bin>)")
def refusals(self, info) -> None:
# One word too many: a patched-vector part spends the slot's last word
# on the trampoline, so its application stops two bytes short.
limit = info.base - 2 if info.patch_vector else info.base
oversized = self.work / "oversized.bin"
oversized.write_bytes(bytes(limit + 2))
rc, out = self.rig.pureboot("--flash", str(oversized))
self.check(f"image over {limit} B refused", rc != 0, self._brief(out))
# ------------------------------------------------------------------- run
def run(self, app: pathlib.Path | None, loader_image: pathlib.Path | None,
marker: str) -> int:
print("identity")
info = self.identity()
if info is None:
print("\nthe loader never answered; nothing below can be trusted")
return 1
print("\nEEPROM")
self.eeprom(info)
if app:
print("\napplication")
self.application(info, app, marker)
else:
print("\nskip application checks (pass --app <image.hex>)")
print("\nerase and the write guard")
self.erase_and_guard(info, loader_image)
print("\nrefusals")
self.refusals(info)
passed = sum(1 for _, ok, _ in self.results if ok)
print(f"\n{passed}/{len(self.results)} passed")
return 0 if passed == len(self.results) else 1
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="hardware acceptance suite for one pureboot deployment",
epilog="overwrites the board's application flash and EEPROM — back them up first")
pbrig.Deployment.add_arguments(parser)
parser.add_argument("--app", type=pathlib.Path,
help="application image to flash (test/pbapp.cpp built for this deployment)")
parser.add_argument("--loader", type=pathlib.Path,
help="the resident loader's .bin, to prove the slot survives an erase")
parser.add_argument("--marker", default="",
help="text the application emits when it runs, e.g. APP")
args = parser.parse_args(argv)
rig = pbrig.Rig(pbrig.Deployment.from_args(args))
print(f"rig: {args.part} on {args.programmer}, link {args.port} at {args.baud} Bd"
f"{' (autobaud)' if args.autobaud else ''}")
print("this overwrites the application flash and EEPROM\n")
with tempfile.TemporaryDirectory(prefix="pbhw-") as temporary:
return Suite(rig, pathlib.Path(temporary)).run(args.app, args.loader, args.marker)
if __name__ == "__main__":
try:
sys.exit(main())
except pbrig.Error as error:
print(f"error: {error}", file=sys.stderr)
sys.exit(2)

View File

@@ -1,427 +0,0 @@
#!/usr/bin/env python3
"""Hardware rig driver for pureboot: an ISP programmer beside a serial link.
The simulated suites (`test/pb*.py`) prove the protocol; this drives the same
loader on real silicon, where the things a cycle-exact simulator cannot model
live — an RC oscillator off its nominal, a reset edge that has to come from
somewhere, a serial bridge with its own idea of what a baud is.
Nothing here knows a port name, a part or a programmer. Every deployment fact
arrives from the command line or the environment, so the same script serves any
board: see `Deployment`. As a module it is the reset/flash/talk primitives that
`pbhw.py` builds its acceptance suite from; as a command it is the handful of
one-shot operations worth having on a rig — most importantly `backup`, which is
the only thing standing between a fuse experiment and an unrecoverable part.
Two rig facts are encoded here because they are not guessable and cost a
session each to learn:
* **An ISP access resets the part**, and it runs again the moment the programmer
releases it. That is the only reset edge available when the serial adapter's
DTR is not wired to reset — so a loader session begins with an ISP touch and
knocks immediately after, which is what `Rig.pureboot()` does.
* **avrdude splits `-U memory:op:file:format` on colons**, so a Windows path's
drive letter breaks the spec. Every file argument is therefore passed as a
bare filename with avrdude run in that file's own directory.
"""
from __future__ import annotations
import argparse
import dataclasses
import importlib.util
import os
import pathlib
import subprocess
import sys
import time
HERE = pathlib.Path(__file__).resolve().parent
DEFAULT_PUREBOOT = HERE.parent / "pureboot" / "pureboot.py"
# Memories worth capturing before an experiment, and the format each is read in.
# Fuses and lock are per-part: a part without an extended fuse simply fails that
# one read, which `backup` reports and steps over rather than aborting on.
BACKUP_MEMORIES = (
("flash", "i", "hex"),
("flash", "r", "bin"),
("eeprom", "i", "hex"),
("eeprom", "r", "bin"),
("lfuse", "h", "hex"),
("hfuse", "h", "hex"),
("efuse", "h", "hex"),
("lock", "h", "hex"),
("calibration", "h", "hex"),
)
class Error(Exception):
pass
def bitclock_for(hz: int) -> str:
"""A safe ISP bitclock for a part *currently running* at `hz`.
SCK must stay under a quarter of the target clock, so the bitclock follows
the clock in force — not the one about to be fused in. Halving that ceiling
again costs nothing on a link that moves a few hundred bytes and buys margin
against an oscillator that is already known to be off its nominal.
"""
ceiling = hz // 8
for candidate in (1000, 4000, 8000, 32000, 125000, 400000):
if candidate <= ceiling:
best = candidate
else:
break
else:
best = 400000
if ceiling < 1000:
raise Error(f"a part at {hz} Hz is too slow to reach over ISP safely")
return f"{best // 1000}kHz"
@dataclasses.dataclass
class Deployment:
"""Everything about one board. No default names a real device."""
port: str = "" # serial device the loader speaks on
baud: int = 57600 # host rate; for autobaud, the rate to drive
autobaud: bool = False # send the calibration pulse instead of p+b
programmer: str = "" # avrdude -c
part: str = "" # avrdude -p
avrdude: str = "avrdude"
bitclock: str = "125kHz" # see bitclock_for()
pureboot: pathlib.Path = DEFAULT_PUREBOOT
wait: int = 12 # seconds the host keeps knocking
@classmethod
def from_env(cls) -> "Deployment":
"""Environment defaults, so a rig's facts live in one place per machine."""
return cls(
port=os.environ.get("PUREBOOT_PORT", ""),
baud=int(os.environ.get("PUREBOOT_BAUD", "57600")),
autobaud=os.environ.get("PUREBOOT_AUTOBAUD", "") not in ("", "0"),
programmer=os.environ.get("PUREBOOT_PROGRAMMER", ""),
part=os.environ.get("PUREBOOT_PART", ""),
avrdude=os.environ.get("AVRDUDE", "avrdude"),
bitclock=os.environ.get("PUREBOOT_BITCLOCK", "125kHz"),
pureboot=pathlib.Path(os.environ.get("PUREBOOT_TOOL", str(DEFAULT_PUREBOOT))),
)
@staticmethod
def add_arguments(parser: argparse.ArgumentParser) -> None:
"""Deployment flags, shared by this tool and pbhw.py."""
env = Deployment.from_env()
parser.add_argument("--port", default=env.port, help="serial device the loader speaks on")
parser.add_argument("--baud", type=int, default=env.baud,
help="host rate (for autobaud, the rate to drive)")
parser.add_argument("--autobaud", action="store_true", default=env.autobaud,
help="send the calibration pulse instead of the p+b knock")
parser.add_argument("--programmer", default=env.programmer, help="avrdude -c, e.g. atmelice_isp")
parser.add_argument("--part", default=env.part, help="avrdude -p, e.g. t13 or m328p")
parser.add_argument("--avrdude", default=env.avrdude, help="path to avrdude")
parser.add_argument("--bitclock", default=env.bitclock, help="ISP bitclock, e.g. 125kHz or 8kHz")
parser.add_argument("--pureboot", type=pathlib.Path, default=env.pureboot,
help="path to pureboot.py")
parser.add_argument("--wait", type=int, default=env.wait, help="seconds to keep knocking")
@classmethod
def from_args(cls, args: argparse.Namespace) -> "Deployment":
return cls(port=args.port, baud=args.baud, autobaud=args.autobaud,
programmer=args.programmer, part=args.part, avrdude=args.avrdude,
bitclock=args.bitclock, pureboot=args.pureboot, wait=args.wait)
def load_pureboot(path: pathlib.Path = DEFAULT_PUREBOOT):
"""The host tool as a module — its Port and Loader, not a subprocess.
Used where a subprocess cannot express what is needed: a poke followed by a
peek in the *same* session, or a raw read at an arbitrary baud.
"""
spec = importlib.util.spec_from_file_location("pureboot", path)
if spec is None or spec.loader is None:
raise Error(f"cannot load the host tool from {path}")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
class Rig:
"""One board: its programmer on one side, its serial link on the other."""
def __init__(self, deployment: Deployment):
self.d = deployment
if not deployment.programmer or not deployment.part:
raise Error("a rig needs --programmer and --part")
# ------------------------------------------------------------- programmer
def avrdude(self, *args: str, cwd: pathlib.Path | None = None,
bitclock: str | None = None, timeout: int = 300) -> subprocess.CompletedProcess:
command = [self.d.avrdude, "-c", self.d.programmer, "-p", self.d.part,
"-B", bitclock or self.d.bitclock, *args]
return subprocess.run(command, capture_output=True, text=True,
cwd=None if cwd is None else str(cwd), timeout=timeout)
@staticmethod
def _ok(result: subprocess.CompletedProcess) -> bool:
return result.returncode == 0
def reset(self, bitclock: str | None = None) -> None:
"""An ISP access, which resets the part; it runs when avrdude exits."""
self.avrdude("-U", "signature:r:-:h", bitclock=bitclock)
def signature(self, bitclock: str | None = None) -> str:
result = self.avrdude("-U", "signature:r:-:h", bitclock=bitclock)
for line in reversed(result.stdout.splitlines()):
if line.strip().startswith("0x"):
return line.strip()
raise Error(f"no signature read: {(result.stderr or result.stdout).strip()[:200]}")
def read_memory(self, memory: str, destination: pathlib.Path, fmt: str = "r",
bitclock: str | None = None) -> bool:
"""Read `memory` into `destination`, whose directory avrdude runs in."""
destination = pathlib.Path(destination).resolve()
destination.parent.mkdir(parents=True, exist_ok=True)
result = self.avrdude("-U", f"{memory}:r:{destination.name}:{fmt}",
cwd=destination.parent, bitclock=bitclock)
# A memory the part does not have (a tiny's extended fuse) leaves avrdude
# happy and the file empty. An empty capture is a miss, not a backup.
return self._ok(result) and destination.exists() and destination.stat().st_size > 0
def write_memory(self, memory: str, source: pathlib.Path, fmt: str = "i",
erase: bool = False, bitclock: str | None = None) -> bool:
source = pathlib.Path(source).resolve()
args = ["-U", f"{memory}:w:{source.name}:{fmt}"]
if erase:
args.insert(0, "-e")
result = self.avrdude(*args, cwd=source.parent, bitclock=bitclock)
return "verified" in (result.stdout + result.stderr)
def flash_hex(self, image: pathlib.Path, erase: bool = True,
bitclock: str | None = None) -> bool:
return self.write_memory("flash", image, "i", erase=erase, bitclock=bitclock)
def read_fuses(self, bitclock: str | None = None) -> dict[str, str]:
out: dict[str, str] = {}
for fuse in ("lfuse", "hfuse", "efuse", "lock"):
result = self.avrdude("-U", f"{fuse}:r:-:h", bitclock=bitclock)
values = [l.strip() for l in result.stdout.splitlines() if l.strip().startswith("0x")]
if values:
out[fuse] = values[-1]
return out
def write_fuses(self, bitclock: str | None = None, **fuses: str) -> bool:
"""Write named fuses. A fuse change moves the clock the *next* access is
timed against, so pass a bitclock safe for both sides of the change."""
args: list[str] = []
for name, value in fuses.items():
args += ["-U", f"{name}:w:{value}:m"]
if not args:
return True
result = self.avrdude(*args, bitclock=bitclock)
text = result.stdout + result.stderr
return "verified" in text or "written" in text
# ------------------------------------------------------------ backup
def backup(self, directory: pathlib.Path, prefix: str = "") -> dict[str, bool]:
"""Capture every memory worth keeping, then prove it by a second read.
A backup nobody verified is a guess. Each memory is read twice and the
two reads compared; a mismatch is reported rather than quietly stored.
"""
directory = pathlib.Path(directory).resolve()
directory.mkdir(parents=True, exist_ok=True)
stem = prefix or self.d.part
status: dict[str, bool] = {}
for memory, fmt, extension in BACKUP_MEMORIES:
name = f"{stem}-{memory}.{extension}"
if not self.read_memory(memory, directory / name, fmt):
status[f"{memory}.{extension}"] = False
continue
if extension == "bin": # only the raw form is worth comparing byte-wise
again = directory / f".{name}.again"
self.read_memory(memory, again, fmt)
same = again.exists() and again.read_bytes() == (directory / name).read_bytes()
again.unlink(missing_ok=True)
status[f"{memory}.{extension}"] = same
else:
status[f"{memory}.{extension}"] = True
return status
# ------------------------------------------------------------ serial link
def pureboot(self, *args: str, reset_first: bool = True, baud: int | None = None,
autobaud: bool | None = None, timeout: int = 300,
bitclock: str | None = None) -> tuple[int, str]:
"""Reset, then knock immediately — see the module docstring.
Returns the host tool's exit status and its combined output, so a caller
can assert on what it printed as well as on whether it succeeded.
"""
if reset_first:
self.reset(bitclock=bitclock)
command = [sys.executable, str(self.d.pureboot), "--port", self.d.port,
"--baud", str(self.d.baud if baud is None else baud),
"--wait", str(self.d.wait)]
if self.d.autobaud if autobaud is None else autobaud:
command.append("--autobaud")
command += [str(a) for a in args]
try:
result = subprocess.run(command, capture_output=True, text=True, timeout=timeout)
except subprocess.TimeoutExpired as expired:
return 99, f"TIMEOUT after {timeout}s\n{expired.stdout or ''}{expired.stderr or ''}"
return result.returncode, (result.stdout or "") + (result.stderr or "")
def capture(self, seconds: float = 2.0, baud: int | None = None) -> bytes:
"""Listen to whatever the board is saying, at an arbitrary rate.
Opening the port does not reset a board whose DTR is unwired, so this can
sample a running application repeatedly without disturbing it — which is
what makes the rate sweep below possible.
"""
module = load_pureboot(self.d.pureboot)
port = module.Port(self.d.port, self.d.baud if baud is None else baud)
try:
data = b""
deadline = time.monotonic() + seconds
while time.monotonic() < deadline:
chunk = port.read_available(0.2)
if chunk:
data += chunk
return data
finally:
try:
port.close()
except Exception:
pass
def measure_rate(rig: Rig, marker: bytes, built_baud: int, nominal_hz: int | None = None,
span_percent: float = 12.0, step_percent: float = 0.5,
seconds: float = 0.75) -> dict:
"""Find a transmitting board's true bit rate, using only the serial port.
The board must be emitting something recognisable at a *fixed* cycles-per-bit
— `test/pbapp.cpp` built with PUREBOOT_HEARTBEAT does. Since its bit timing is
a cycle count, its wire rate scales with its actual clock, so the host rates
at which `marker` still decodes bracket that rate; the centre of the band is
the answer, and with the clock the image was built for it gives the real one.
This is the measurement that turns "the loader is silent, so the wiring must
be wrong" into a number, and it needs no instrument beyond the adapter
already attached.
"""
steps = int(span_percent / step_percent)
clean: list[int] = []
samples: list[tuple[int, int, bool]] = []
for index in range(-steps, steps + 1):
baud = int(round(built_baud * (1 + index * step_percent / 100.0)))
if baud <= 0:
continue
data = rig.capture(seconds=seconds, baud=baud)
hit = marker in data
samples.append((baud, len(data), hit))
if hit:
clean.append(baud)
result: dict = {"samples": samples, "clean": clean, "built_baud": built_baud}
if clean:
low, high = min(clean), max(clean)
centre = (low + high) / 2.0
result |= {"low": low, "high": high, "centre": centre,
"half_width_percent": (high - low) / 2.0 / centre * 100.0,
"error_percent": (centre / built_baud - 1.0) * 100.0}
if nominal_hz:
result["measured_hz"] = nominal_hz * centre / built_baud
return result
# ------------------------------------------------------------------- command
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="pureboot hardware rig: ISP reset/flash beside the serial link")
Deployment.add_arguments(parser)
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser("signature", help="read the part signature over ISP")
sub.add_parser("reset", help="reset the part (an ISP access) and let it run")
sub.add_parser("fuses", help="read the fuse and lock bytes")
p = sub.add_parser("flash", help="program a hex image over ISP")
p.add_argument("image", type=pathlib.Path)
p.add_argument("--no-erase", action="store_true", help="do not chip-erase first")
p = sub.add_parser("backup", help="capture and verify every memory")
p.add_argument("directory", type=pathlib.Path)
p.add_argument("--prefix", default="", help="filename stem (default: the part name)")
p = sub.add_parser("rate", help="measure the board's true bit rate and clock")
p.add_argument("--marker", default="APP", help="text the board emits (default: APP)")
p.add_argument("--built-baud", type=int, required=True,
help="the baud the running image was built for")
p.add_argument("--nominal-hz", type=int, default=0,
help="the clock the image was built for, to report the real one")
p.add_argument("--span", type=float, default=12.0, help="sweep +-this many percent")
p.add_argument("--step", type=float, default=0.5, help="sweep step in percent")
p.add_argument("--verbose", action="store_true", help="print every step")
p = sub.add_parser("bitclock", help="a safe ISP bitclock for a clock in force")
p.add_argument("hz", type=int)
args = parser.parse_args(argv)
if args.command == "bitclock":
print(bitclock_for(args.hz))
return 0
rig = Rig(Deployment.from_args(args))
if args.command == "signature":
print(rig.signature())
elif args.command == "reset":
rig.reset()
print("reset")
elif args.command == "fuses":
for name, value in rig.read_fuses().items():
print(f"{name:<6} {value}")
elif args.command == "flash":
ok = rig.flash_hex(args.image, erase=not args.no_erase)
print(f"{args.image.name}: {'verified' if ok else 'FAILED'}")
return 0 if ok else 1
elif args.command == "backup":
status = rig.backup(args.directory, args.prefix)
for name, ok in status.items():
print(f" {'ok ' if ok else 'FAIL'} {name}")
missing = [n for n, ok in status.items() if not ok]
# Fuses a part does not have are expected misses, not failures.
fatal = [n for n in missing if not n.startswith(("efuse", "calibration"))]
print(f"\n{len(status) - len(missing)}/{len(status)} captured into {args.directory}")
return 1 if fatal else 0
elif args.command == "rate":
result = measure_rate(rig, args.marker.encode(), args.built_baud,
args.nominal_hz or None, args.span, args.step)
if args.verbose:
for baud, size, hit in result["samples"]:
print(f" {baud:7d} Bd {size:5d} B {'MARKER' if hit else ''}")
if not result["clean"]:
print(f"no capture contained {args.marker!r} at any rate — is the board "
f"transmitting, and on the pin this port is wired to?")
return 1
print(f"clean band {result['low']}..{result['high']} Bd")
print(f"centre {result['centre']:.0f} Bd "
f"(+-{result['half_width_percent']:.1f} %)")
print(f"vs built {result['built_baud']} Bd ({result['error_percent']:+.1f} %)")
if "measured_hz" in result:
print(f"true clock {result['measured_hz'] / 1e6:.3f} MHz")
return 0
if __name__ == "__main__":
try:
sys.exit(main())
except Error as error:
print(f"error: {error}", file=sys.stderr)
sys.exit(2)

View File

@@ -1,160 +0,0 @@
#!/usr/bin/env python3
"""What the loader images actually measure, and whether the README still agrees.
The size matrix asserts every image fits its slot; it says nothing about the
numbers the README prints, and those drift. Every row of that table was eight
bytes stale once `startup::caller_page()` landed — common code, so every build
moved at once and no test noticed, because none of them was over budget.
Two questions, both answered from built trees:
sizes.py max the largest image per chip, and anything over budget
sizes.py check-readme the README's per-chip table against what is built
Nothing here knows a chip's geometry. The (image, budget) pairs come from each
build's own `CTestTestfile.cmake` — the same values the gate checks — so the
slot rules stay where they belong, in `pureboot/CMakeLists.txt`, and a chip
added or a budget changed needs no edit here. Only trees a configure preset
still owns are read: a stale directory keeps its last build, and a loader built
before a slot changed will happily report a size that was true once
(`tools/prune-build-trees.sh` in libavr removes them).
Sizes come from `avr-size`, and a target is only as current as its last build —
run the gate first if you want the table checked against today's source.
"""
from __future__ import annotations
import argparse
import pathlib
import re
import shutil
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parents[1]
# add_test(<name>.size ... -DELF=<path> ... -DLIMIT=<n> ...) — the gate's own
# pairing of an image with the budget it must fit.
# ctest writes the name as a bracket argument ([=[name.size]=]) and quotes the
# rest, so the name starts after the bracket and the path ends at the quote.
SIZE_TEST = re.compile(r'add_test\(\s*\[=\[(?P<name>[^\]]+?)\.size\]=\][^\n]*?'
r'-DELF=(?P<elf>[^"\s]+)[^\n]*?-DLIMIT=(?P<limit>\d+)')
def avr_size() -> str:
for env in (ROOT / "../../toolchain").resolve().glob("avr-gcc-*/bin/avr-size"):
if env.is_file():
return str(env)
found = shutil.which("avr-size")
if not found:
sys.exit("no avr-size found (build the toolchain, or put it on PATH)")
return found
def preset_dirs() -> list[pathlib.Path]:
"""Build trees a configure preset still owns, newest-listed first."""
listing = subprocess.run(["cmake", "--list-presets"], cwd=ROOT, capture_output=True, text=True)
names = re.findall(r'^\s*"(.+)"$', listing.stdout, re.MULTILINE)
if not names:
sys.exit("cmake --list-presets returned nothing — run from a configured checkout")
return [d for d in (ROOT / "build" / n for n in names) if (d / "CTestTestfile.cmake").is_file()]
def measure(paths: list[str], tool: str) -> dict[str, int]:
""".text per ELF, in one avr-size call per batch."""
sizes: dict[str, int] = {}
for start in range(0, len(paths), 400):
batch = [p for p in paths[start:start + 400] if pathlib.Path(p).is_file()]
if not batch:
continue
out = subprocess.run([tool, *batch], capture_output=True, text=True).stdout
for line in out.splitlines()[1:]:
fields = line.split()
if len(fields) >= 6 and fields[0].isdigit():
sizes[fields[5]] = int(fields[0])
return sizes
def collect() -> dict[str, list[tuple[str, int, int]]]:
"""chip -> [(target, text, limit)], from every owned build tree."""
tool = avr_size()
found: dict[str, list[tuple[str, str, int]]] = {}
for tree in preset_dirs():
chip = tree.name.split("-")[0]
for match in SIZE_TEST.finditer((tree / "CTestTestfile.cmake").read_text()):
found.setdefault(chip, []).append((match["name"], match["elf"], int(match["limit"])))
sizes = measure([elf for rows in found.values() for _, elf, _ in rows], tool)
measured = {
chip: sorted(((name, sizes[elf], limit) for name, elf, limit in rows if elf in sizes),
key=lambda row: -row[1])
for chip, rows in sorted(found.items())
}
# A configured-but-unbuilt preset registers its tests with no images behind
# them; it is not a chip with nothing to say, it is a chip not built yet.
return {chip: rows for chip, rows in measured.items() if rows}
def cmd_max(args) -> int:
measured = collect()
if not measured:
sys.exit("nothing built — configure and build a preset first")
over = []
print(f"{'chip':<13} {'largest image':<34} {'.text':>6} {'budget':>7} headroom")
for chip, rows in measured.items():
name, text, limit = rows[0]
flag = "OVER" if text > limit else f"{limit - text:>5} B"
print(f"{chip:<13} {name:<34} {text:>6} {limit:>7} {flag}")
over += [(chip, n, t, l) for n, t, l in rows if t > l]
total = sum(len(rows) for rows in measured.values())
print(f"\n{total} images across {len(measured)} chips")
if over:
print("\nOVER BUDGET:")
for chip, name, text, limit in over:
print(f" {chip} {name}: {text} > {limit}")
return 1
tightest = min(((chip, n, t, l) for chip, rows in measured.items() for n, t, l in rows),
key=lambda row: row[3] - row[2])
chip, name, text, limit = tightest
print(f"tightest fit: {chip} {name}{text} of {limit}, {limit - text} B spare")
return 0
def cmd_check_readme(args) -> int:
"""The README's per-chip table, against the stock and autobaud builds."""
readme = (ROOT / "pureboot" / "README.md").read_text()
measured = collect()
rows = re.findall(r"^\|\s*(AT\w+[^|]*?)\s*\|[^|]*\|[^|]*\|[^|]*\|\s*(\d+) B\s*\|\s*(\d+) B\s*\|$",
readme, re.MULTILINE)
if not rows:
sys.exit("no size table found in pureboot/README.md")
bad = skipped = 0
for chips, stock_doc, auto_doc in rows:
# "ATmega48, 48A, 48P, 48PA †" — the first name is the family's base.
chip = re.sub(r"[^a-z0-9]", "", chips.split(",")[0].strip().lower())
built = {name: text for name, text, _ in measured.get(chip, [])}
for target, documented in (("pureboot", stock_doc), ("pureboot_autobaud", auto_doc)):
if target not in built:
skipped += 1
continue
if built[target] != int(documented):
print(f" {chip:<12} {target:<18} README says {documented} B, built is {built[target]} B")
bad += 1
if bad:
print(f"\n{bad} row(s) stale — update pureboot/README.md")
return 1
print(f"README size table matches every built image ({len(rows)} rows"
+ (f", {skipped} not built" if skipped else "") + ")")
return 0
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
subs = parser.add_subparsers(dest="cmd", required=True)
subs.add_parser("max", help="largest image per chip, and anything over budget")
subs.add_parser("check-readme", help="the README's size table against what is built")
args = parser.parse_args()
return {"max": cmd_max, "check-readme": cmd_check_readme}[args.cmd](args)
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -1,302 +0,0 @@
// TinySafeBoot on libavr — the policy floor: pureboot's rules, measured.
//
// The full TinySafeBoot feature set — watchdog bail, one-wire half-duplex,
// config-page activation timeout, password gate, emergency erase, and
// config/flash/EEPROM read-write — under philosophy #5 exactly as pureboot
// obeys it: no assembly, no register variables; code, attributes, and flags
// only. Every lesson pureboot's development produced is applied — the
// library's half-duplex serial and startup entry, lean bring-up from reset
// state, one merged send loop over both memories, oracle-shaped loop bounds,
// locals threaded through noinline primitives, pureboot's codegen flags —
// and the result is 638 bytes: 198 below the idiomatic tier, and 126 above
// the 512 B boot section the tricks/asm tiers reach with the banned
// mechanisms (526/510). This tier exists to keep that number an artifact
// rather than a claim: the gap to 512 is the rent of policy-clean C++ —
// helpers that hold a cursor across rx()/tx() pay push/pop and argument
// threading where a global-register protocol pays nothing, and both
// control-flow merges tried (a parametrized paged session, a merged store
// loop) measured larger than the split cases they replaced. TSB's wire fixes
// the per-command loop shapes on the device, so pureboot 5's one-transfer-
// loop collapse has no purchase here.
//
// The wire protocol is strict request/response, which is what makes the
// shared line safe: the device drives it only between a received command and
// its reply, and releases it (the library's half-duplex choreography)
// whenever it waits.
#include <libavr/libavr.hpp>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
using dev = avr::device<{.clock = 16_MHz}>;
// One-wire: RX and TX share the line, exactly as the native-UART TSB expects.
using serial_t = dev::uart0<{.baud = 115200_Bd, .max_baud_error = 3_pct, .half_duplex = true}>;
inline constexpr serial_t serial{};
namespace tsb {
namespace {
// The loader is purely polled — it never enables interrupts — so every SPM and
// EEPROM lock folds to nothing under this posture.
constexpr auto off = avr::irq::guard_policy::unused;
// The handshake bytes, identical across every TSB host.
constexpr std::uint8_t confirm = '!';
constexpr std::uint8_t request = '?';
constexpr std::uint8_t knock = '@';
// Boot geometry for the 1 KB boot section (BOOTSZ=10); the page size and the
// flash/EEPROM extents are the chip database's to know. app_end is the config
// page (TSB's LASTPAGE), one page below the boot section.
constexpr std::uint16_t page = spm::page_bytes;
constexpr std::uint16_t boot_bytes = 1024;
constexpr std::uint16_t app_end = spm::flash_bytes - boot_bytes - page;
constexpr std::uint16_t eeprom_end = avr::hw::db.mem.eeprom_size - 1;
// Lockout-proof floor for the activation window (the oracle's F_CPU/1MHz).
constexpr std::uint8_t act_min = 16;
// Post-activation window: the host gets seconds, not milliseconds, mid-session.
constexpr std::uint8_t comm_window = 200;
// Firmware version stamp: YY*512 + MM*32 + DD, the encoding the host decodes.
constexpr std::uint16_t build_date = 26 * 512 + 7 * 32 + 27;
// The 16-byte device-info block, streamed out on activation.
// clang-format off
[[gnu::progmem]] constexpr std::uint8_t info[16] = {
'T', 'S', 'B',
build_date & 0xFF, build_date >> 8,
0xF3, // status: native-UART fixed-baud lineage
avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2],
page / 2, // page size in words
(app_end / 2) & 0xFF, (app_end / 2) >> 8, // app-flash boundary, words
eeprom_end & 0xFF, eeprom_end >> 8,
0xAA, 0xAA, // ATmega processor-type marker (bytes 14 == 15)
};
// clang-format on
// The receive window, pre-floored where it is set. In .noinit: there is no
// crt to clear a .bss image, and run() stores it before the first receive.
[[gnu::section(".noinit")]] std::uint8_t window;
const std::uint8_t *flash_ptr(std::uint16_t addr)
{
return reinterpret_cast<const std::uint8_t *>(addr);
}
// Bounded byte receive: poll under nested countdowns, 0 on silence. The 0
// then falls through every compare — not a knock, not a confirm, not a
// command — so a silent host unwinds the loader to the application from
// anywhere, and a mid-session cable pull cannot wedge it. The line release on
// a direction change is the serial backend's.
[[gnu::noinline]] std::uint8_t rx()
{
std::uint16_t outer = static_cast<std::uint16_t>(window) << 8;
do {
std::uint8_t fine = 0;
do {
if (auto byte = serial.read())
return *byte;
} while (--fine);
} while (--outer);
return 0;
}
// One-wire transmit: the backend takes the line with a turn-around guard and
// holds it until the whole frame is out.
[[gnu::noinline]] void tx(std::uint8_t byte)
{
serial.write(byte);
}
// '?', then hand back the host's reply for the callers' one-byte compare.
[[gnu::noinline]] std::uint8_t rcnf()
{
tx(request);
return rx();
}
// The one send loop: the info block, the config page, application flash and
// EEPROM pages all stream through here.
[[gnu::noinline]] void send_block(bool eep, std::uint16_t at, std::uint8_t count)
{
do {
tx(eep ? ee::read(at) : avr::flash_load(flash_ptr(at)));
++at;
} while (--count);
}
// One EEPROM byte in — shared by the emergency wipe and the 'E' stream.
[[gnu::noinline]] void eeput(std::uint16_t at, std::uint8_t value)
{
ee::write<off>(at, value);
}
// Wait out a running SPM op, then re-open the RWW section — after every page
// op and before handing over, as the oracle does.
[[gnu::noinline]] void settle()
{
spm::wait();
spm::rww_enable<off>();
}
// One host page straight into the erased flash page at `at` — through the SPM
// word buffer (low byte then high), no SRAM staging — then committed. `at`
// names a page base, so the cursor's low byte reaching the boundary ends the
// walk.
[[gnu::noinline]] void store_flash_page(std::uint16_t at)
{
do {
std::uint8_t low = rx();
std::uint8_t high = rx();
spm::fill<off>(at, std::bit_cast<std::uint16_t>(std::array{low, high}));
at += 2;
} while (static_cast<std::uint8_t>(at) & (page - 1));
spm::write_page<off>(at - page);
settle();
}
extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --defsym=tsb_app=0
[[noreturn]] void appjump()
{
settle();
tsb_app();
}
// Step one page down and erase it — the erase shared by the whole-app walk,
// the config rewrite and the emergency wipe; hands the stepped address back.
[[gnu::noinline]] std::uint16_t erase_below(std::uint16_t at)
{
at -= page;
spm::erase_page<off>(at);
settle();
return at;
}
// Erase the whole application, top-down like the oracle: the loop bound is a
// compare with zero, and the returned 0 is the address every caller wants
// next.
[[gnu::noinline]] std::uint16_t erase_application()
{
std::uint16_t at = app_end;
do {
at = erase_below(at);
} while (at != 0);
return at;
}
[[noreturn]] void run()
{
// A watchdog reset hands straight back to the application, as the
// reference loader does, rather than re-entering the bootloader.
if (avr::hw::mcusr::wdrf.test())
appjump();
// Lean bring-up from reset state: UCSR0C already reads 8N1, UBRR0H reads
// 0, and the half-duplex write()/read() raise TXEN0/RXEN0 on first use —
// only the divisor low byte and U2X0 need a store. The solver still does
// the datasheet work; the asserts pin the reset-state assumptions.
{
constexpr auto sol = avr::uart::detail::solve_baud(dev::clock, 115200_Bd);
static_assert(sol.u2x && sol.ubrr < 256, "lean bring-up writes UBRR0L only, with U2X0");
avr::hw::reg<"UBRR0">::write(static_cast<std::uint8_t>(sol.ubrr));
avr::hw::ucsr0a::write(avr::hw::ucsr0a::u2x0(1));
}
// Activation: 3×'@', each inside the config page's timeout window
// (floored so a corrupt page cannot lock the loader out); anything else —
// including silence — hands over.
window = avr::flash_load(flash_ptr(app_end + 2)) | act_min;
for (std::uint8_t k = 3; k; --k)
if (rx() != knock)
appjump();
window = comm_window;
// Password gate (config page from app_end+3, 0xff-terminated; a blank
// page is no password). A wrong byte blanks the comparison and drains the
// line forever, so a wrong password can never fall through; a 0 requests
// emergency erase behind two confirms. On pass the info block goes out;
// the emergency path skips it and drops into the command loop.
std::uint16_t at = app_end + 3;
std::uint8_t mask = 0xff;
for (;;) {
std::uint8_t expected = avr::flash_load(flash_ptr(at)) & mask;
++at;
if (expected == 0xff) {
send_block(false, reinterpret_cast<std::uint16_t>(&info[0]), sizeof info);
break;
}
std::uint8_t got = rx();
if (got == 0) {
if (mask == 0)
continue;
if (rcnf() != confirm || rcnf() != confirm)
appjump();
std::uint16_t a = erase_application();
do {
eeput(a, 0xff);
} while (++a <= eeprom_end);
erase_below(app_end + page);
break;
}
if (got != expected)
mask = 0;
}
for (;;) {
tx(confirm); // Mainloop ready
const std::uint8_t command = rx();
switch (command) {
case 'f': // read application flash, one page per host '!'
for (std::uint16_t a = 0; a < app_end; a += page) {
if (rx() != confirm)
break;
send_block(false, a, page);
}
break;
case 'e': // read EEPROM, one page per host '!', until the host stops
for (std::uint16_t a = 0;; a += page) {
if (rx() != confirm)
break;
send_block(true, a, page);
}
break;
case 'F': { // erase the application, then take pages behind '?'
std::uint16_t a = erase_application();
for (; rcnf() == confirm; a += page)
store_flash_page(a);
break;
}
case 'E': // take EEPROM pages behind '?', each write host-paced
for (std::uint16_t a = 0; rcnf() == confirm;) {
std::uint8_t count = page;
do {
eeput(a, rx());
++a;
} while (--count);
}
break;
case 'c': // read the config page
read_config:
send_block(false, app_end, page);
break;
case 'C': // replace the config page, then echo it back to verify
if (rcnf() != confirm)
break;
store_flash_page(erase_below(app_end + page));
goto read_config;
default: // 'q' or any other byte runs the application
appjump();
}
}
}
} // namespace
} // namespace tsb
// Reset lands at the boot section base (BOOTRST): the entry stub in .vectors
// is laid first and does the one line of crt a crt-less image needs.
template struct avr::startup::entry<tsb::run>;