37 Commits

Author SHA1 Message Date
a58b59c60f docs: correct the size figures for the configured variants
The headline numbers described the stock deployments but claimed "every
configured variant of them", which the matrix contradicts: choosing the
software UART where the chip has a USART costs 8-46 B, so the megas reach
460-462 rather than 452, and the 1284s' software-serial build is 546 B —
inside their 1 KiB boot sector, but not inside 512.

Also names the actual tightest chip. The 1284 looks like it at 506, but it
deploys in 1 KiB with 478 B spare; against its own budget the ATmega328P
has the least room, 50 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 17:08:24 +02:00
ceacb61ba1 pureboot: no SPM buffer discard, the host repairs instead
The temporary page buffer is write-once per word, so a page filled over
one an earlier writer left dirty programs the stale words. The same
datasheet clause carries the cure: the buffer auto-erases after a page
write (§26.2.1; §19.2 on the tinies), so the corruption clears itself by
happening, and rewriting the page programs correctly.

The loader therefore clears the buffer nowhere. The tinies' CTPB and the
m48s' RWWSRE discard are gone; the boot-sectioned megas keep only the
trailing RWWSRE they need anyway to re-enable the RWW section for
read-back, which discards the buffer as a side effect and keeps them off
the path entirely. 434 B on the tiny13s, 438-442 on the tiny25/45/85,
430 on the m48s; the megas are unchanged, the 1284s still 506.

The host takes over the guarantee: a flash page that reads back wrong is
rewritten up to RETRIES times before the run stops. Both read-back paths
repair — verify_pages for programming, and write_differing, which is the
loader-update path where a page left wrong is a half-written loader slot.
That one is not hypothetical: deleting the discard made attiny85
pureboot.rehome fail deterministically there, the only flow still
assuming the old contract.

Protocol-visible, so README's W command says it: one W may program the
wrong bytes after a refused page, or after an application that
self-programmed entered without a reset, and a host that programs without
reading back cannot trust it.

Tests: pureboot.dirty drives the case the loader declines to guard — the
fixture application dirties every buffer word and jumps in with no reset
(hardware forbids that on a boot-sectioned mega, but simavr dispatches SPM
from anywhere, which is what makes it constructible) — and asserts a bare
verify sees the corruption, the repairing verify fixes it in one rewrite,
and it stays fixed. pbreloc asserts the same shape after a refusal.
test_planner covers the bound against a fake device: one bad write
repaired in a single rewrite, a page that never comes good stopping after
exactly three.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 16:29:27 +02:00
dd7490a3df build: generated presets and a check entry point
tools/make_presets.py emits the uniform pipeline the hand-grown file had
drifted from — generated configure/build/test presets and workflows for
all 37 chips, reflect configure/build for libavr's 12-chip spot set —
and tools/check.sh runs every chip's workflow (--full adds the reflect
spot) as the port's gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
bc24b63e65 host: one fact per line, progress bars, --verbose
--info prints the decoded info block field by field and --fuses each
byte on its own line plus the BOOTSZ/BOOTRST meaning on boot-sectioned
megas. Transfers that take wire time draw a transient progress bar on
stderr when it is a tty — logs, pipes and the tests see only the
summary lines. -v/--verbose narrates decisions: knock counts, the
programming plan, update state handling and per-phase page counts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
332b244ac3 pureboot: every deployment axis is a build parameter
Clock, baud, serial backend (hardware USART 0/1 or the software UART on
any pins) and the activation window all resolve through one CMake
function, pureboot_add_loader() in pureboot/CMakeLists.txt — the unit a
downstream project consumes. The default baud is the fastest standard
rate within 2.5 % (the same best-divisor search libavr's solver runs),
gated on software builds by the polled receiver's 100-cycles-a-bit
floor; every explicit pick is re-checked by the compile's static asserts.

The size matrix builds each axis that can move the image — backend x
clock ladder x USART instance, per chip — against the slot budget, and
two nondefault deployments run the whole protocol suite live: the 328P
on its shipped 1 MHz fuses over software serial on TX=PB1/RX=PB5
(pureboot.custom), and the 644A over USART1 (pureboot.usart1). The sim
runner takes -l to bridge any link, paces a fully quiet bridge toward
real time (a free-running 8 M-cycle window loses the reset-race knock),
and the fixture application speaks the deployment it is built for.

The loader itself shed bytes on the way: the return-address high byte
spelled through byteswap (the double swap folds to the one-byte pick),
the info-block address composed instead of bit_cast, and libavr's new
polled-UART helpers replacing the port's uart::detail reaches. Every
combination fits: 458-506 B across the megas' whole matrix, 470-484 B
on the tinies, 556-562 B in the 1284s' 1 KiB slot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 23:42:57 +02:00
61b687120c pureboot: a loader in the staging slot is the staging copy — leave it there
The update flow's first step wrote the staging content over whatever the
staging slot held; with a loader running right there (programmed by hand
onto erased flash), that write met the copy's own running-slot guard on
the composed through-word and the tool stopped at its verify — although
the copy is exactly an installed staging copy, able to stream the new
resident like any other. The install is now skipped when the slot holds a
complete loader: its info block where every image carries it, matching
the device's byte for byte, and the slot unchanged since the update began
(the state file's snapshot) — so a resumed half-written install still
differs from its snapshot and takes the install path, which completes it.
pbrehome gains the staging-slot position (an older build at stage
streaming a newer resident in); the README's wrong "cannot re-home from
the staging slot" claim is corrected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 21:55:15 +02:00
ef58d5d363 pureboot: artifact roles, wrong-chip refusal, and misplaced-loader re-homing
The README's deployment section now says what each build artifact is for:
the .hex is the programmer artifact (self-addressed into the top slot),
the .bin the self-update image — bare slot bytes a programmer would put
at address 0, where a boot-sectioned mega cannot even heal itself (SPM
only runs from the boot section) but a patched-vector chip runs the
position-independent copy and re-homes a build through the ordinary
--update-loader flow: the staging install and the word-0 redirect both
execute outside page 0's slot, so the running-slot guard never blocks it.
pbrehome.py is the acceptance test (misplaced at 0, guard intact,
re-home, app flash over the stale copy, banner); the staging slot is the
one position that cannot re-home itself, documented. The preflight's
wrong-chip refusal and loader_image's handling of padded images (peeled
to the slot content by the embedded base) are documented and the padded
case pinned in the planner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 21:18:10 +02:00
03bb1f56cf pureboot: every libavr chip — 37 loaders, the m48 class, the 644 geometry
The chip table becomes family blocks covering all 37 targets. The m48s
are a new deployment class: no boot section, so the tiny profile spoken
over the hardware USART — host-patched reset vector, trampoline
hand-over, a 510-byte budget (474 B built), no fuse preflight — while
their RWWSRE store stays the buffer discard (Atmel-8271 §26.2); the
device keys the patch flag and the CPU-halt waits on the curated
boot-section capability and the discard on the RWWSRE bit itself. The
644s' 64 KiB is exactly the 16-bit byte space: plain LPM, byte wire
addresses, 498 B in a 512-byte slot — and their 1 KiB minimum boot
section holds the resident and staging slots together, so self-update
needs no fuse step (the update test's slot pick now keys word-flash on
base >= 64 KiB; base + slot merely touching the boundary stays
byte-addressed). The 1284 joins the 1284P's word-addressed 1 KiB slot at
558 B. BOOT_FUSE gains every boot-sectioned family's ladder and fuse
byte; the planner exercises them all. The sim scaffolding keys
patch-vector-ness instead of the atmega name prefix, the fixture app
picks its clock by family (the tiny25/45/13 builds surfaced the 16 MHz
fallthrough as garbled banners), and the runner's wrapped flash ioctl
performs the m48 discard simavr's no-RWW cores turn into a stray buffer
fill. Sizes across the fleet: 466-504 B megas, 474 B m48s, 498 B 644s,
488-502 B tinies, 558 B 1284s — every chip passing
size/pi/planner/protocol/reloc/update.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 18:53:04 +02:00
6a5e313530 pureboot: correct the 1284P deployment profiles in the README
The deployment section claimed the 1284P has no standalone profile, runs
BOOTSZ = 512 words always, and self-updates with no fuse change, reset
landing at 0x1f800 — internally contradictory (a 512-word section starts
at 0x1fc00, and the section holding both 1 KiB slots is 1024 words) and
contradicted by update_preflight, which refuses a self-update unless the
boot section covers two slots. The text described a 512-byte-slot
geometry this chip's loader cannot have. In truth the 328P profile table
maps onto the 1284P doubled: standalone = 512 words (the smallest
section is exactly the 1 KiB slot, reset at the loader base), self-update
= 1024 words with the loader-first reset walking the staging slot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 16:48:46 +02:00
a28fccf475 pureboot: the 1284P rides a 1 KiB slot — its own boot-sector minimum
The far machinery (ELPM reads, RAMPZ page commands, wire-word math)
costs ~46 B over the m328P's 504, and the tsb-calibrated C++-to-asm
gap says no implementation of this feature set reaches 512 on this
chip — a boundary its hardware does not have anyway: the 1284P's
smallest boot sector is 1 KiB. The slot therefore becomes
per-geometry (512 B, or 1 KiB past 64 KiB), which the host derives
from the word-addressing flag; slot arithmetic unifies (the index is
the wire high byte with its low bit dropped in either unit), the
update preflight demands a two-slot boot section in the chip's own
terms, and pbapp's hand-back jumps to the real slot base. libavr's
far primitives split their RAMPZ/Z asm operands (a page never
crosses 64 KiB, so callers keep a byte and a 16-bit cursor — the
32-bit address folds away; flash_load_far's byte form becomes the
out-RAMPZ+elpm pair avr-libc's pgm_read_byte_far rebuilds per call),
and the host splits reads at 64 KiB boundaries. All ten chips pass
the full suite — the 1284P at 558 B including protocol, relocation,
and the power-fail self-update — with pureboot byte-identical across
generated and reflect modes everywhere, and the original three
chips' images unchanged to the byte (488/502/504).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:54:40 +02:00
92503fbb2d pureboot: the classic megas and the word-addressed 1284P groundwork
Device: boot-section detection probes SPMCR beside SPMCSR, the link
picks any hardware USART through the instance-aware lookups (URSEL
chips included), WDRF reads MCUSR-or-MCUCSR, and the >64 KiB shape
lands — word-addressed wire flash (info flag bit 1, page byte 0 means
256, base as a word address), far reads through flash_load_far, a
single 32-bit byte-cursor page walk (the 256-byte page wraps its low
byte exactly), and slot arithmetic in words (the return address
already is one). Host: addresses stay bytes internally and scale at
the wire, the boot-fuse decode becomes a per-signature table (byte
index + BOOTSZ ladder — the m168A's lives in EXTENDED), and the
planner tests pin every chip's ladder plus the word-addressed info
decode. Tests: the device runner serves every mega over the USART pty,
pbapp banners over the right link, the update rehearsal synthesizes
its assumed fuses from the tool's own table, and the PI lint tracks
the renamed info symbol. All six classic-mega/168A targets pass the
full suite (size, PI, planner, protocol, reloc, self-update) at
466–504 B; the 1284P builds await a libavr far-path slimming to make
its 512.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 12:00:19 +02:00
8ba023ed6b pureboot: take the device signature from the chip database
libavr now exposes avr::hw::db.signature (compile-time, from the ATDF), so the
info block drops its per-chip hardcoded signature() for the db constant. The
loaders are byte-identical across modes with the correct signature, sizes
unchanged (488/502/504).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:38:20 +02:00
443fd39d31 build: emit a raw .bin beside the HEX for every loader image
The host tool takes either form — load_image() parses Intel HEX by extension
and treats anything else as raw bytes — but the build emitted only the HEX, so
the raw path had no artifact behind it. The reloc and update tests each shell
out to objcopy at runtime to produce one for themselves.

add_hex_output becomes add_image_outputs and emits both forms. The .bin is
byte-identical to the plain `objcopy -O binary` those tests generate (-R .eeprom
strips nothing the loaders carry), and decodes equal to the HEX payload — 504 B
at 0x7e00 either way for pureboot. Sizes come out at the flash sizes exactly
(504/510/836/526), so nothing stretches to the .data load address.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 23:17:59 +02:00
08a0d77109 pureboot: take the emitted loader HEX for --update-loader
The build emits an Intel HEX beside every loader image, but --update-loader
could not consume one: load_image() anchors every image at address zero, and
a loader HEX links at its base, so it decoded to a 32760-byte blob carrying
504 bytes of loader at the end. staging_content() then refused it as "loader
image is 32760 B, the slot holds 512" - an error naming neither the cause nor
the raw .bin the tool wanted instead.

Drop the blank below the base in the update path. The base comes from the
image's own info block rather than the device's, so an image built for
another target survives the slice intact and the preflight still reports it
as another target rather than failing to find an info block at all.

Verified on an ATmega328P: the full self-update flow driven straight from
pureboot_timeout-5s.hex, resident slot byte-for-byte against the image
afterwards, application preserved; both refusal paths unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 22:46:50 +02:00
260ab2e4e3 pureboot: drive the serial port on Windows too
The host tool was standard-library-only but POSIX-only with it: termios and
select() bound the port layer, and importing termios failed outright on
Windows, so the module could not even load there.

Split Port into PosixPort (unchanged) and a WindowsPort over the Win32 serial
API through ctypes, picked by os.name; every call site keeps the Port name.
kernel32 only, so the standard-library constraint holds.

Windows has no select() for a COM handle, so the read deadlines move into the
driver as COMMTIMEOUTS, re-armed per read: read_available() ends on a gap
longer than a USB-serial latency timer coalesces (16 ms on FTDI parts),
read_exact() on the count or its deadline. Opening asserts DTR and RTS as a
POSIX open does, so a board wiring DTR to reset still pulses it. A failed
configuration closes the handle before raising - a COM handle is exclusive,
and the leak met the next open as "Access is denied". Win32 takes any integer
baud and a driver may accept one its hardware cannot produce (an FT232R
reports back a baud of 3 and keeps the old divisor), so obvious nonsense is
refused where termios' table would have.

Tested against an ATmega328P on COM6: info, fuses, both memories programmed
and verified, session reconnect, hand-over, the loader self-update, and the
write guard on its own slot. test_planner runs on Windows now as well.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 22:46:14 +02:00
e4aaf5bc62 build: emit an Intel-HEX beside every loader image
avrdude programs Intel-HEX, not ELF, and the build produced only ELFs — so
flashing a loader to a real chip meant running objcopy by hand. add_hex_output()
hangs a POST_BUILD objcopy on each loader image: the three tsb tiers through
add_tsb_variant, pureboot, and the re-timed pureboot9. .eeprom is dropped, being
its own avrdude update.

It uses the toolchain file's CMAKE_OBJCOPY rather than a hardcoded path, so
every chip preset emits hex, not just the mega. pbapp keeps its ELF alone: the
update test converts it to a raw binary itself, and it is not a flashing target.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 21:57:26 +02:00
653602058a pureboot: PI lint, tightened gates, and the self-update test suite
check_pi.py asserts the two link-time facts position independence rests on
(no absolute jmp/call; the info block within the image's first 256 bytes);
the size gates drop to 510 on the tinies for the trampoline word.

New per-chip tests beside the reworked protocol test: the planner units
(programming orders and their recovery properties, the surgery, staging
composition, boot-fuse decode, and the update preflight's error/warning
matrix over synthetic fuse bytes), the relocated-copy sweep (the identical
image installed one slot lower serves the full command set — the PI
acceptance test, and the one that caught the temporary-buffer trap), and
the self-update end-to-end: --update-loader to a re-timed build
(pureboot9, byte-different by PUREBOOT_TIMEOUT alone), then every
power-fail phase killed mid-write, restarted from the runner's flash dump,
and completed by a re-run with the application intact throughout. The mega
rounds run the BOOTRST-unprogrammed profile: the fixture application's 'L'
jump is the application-owned loader entry that profile relies on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:57 +02:00
57d1c009cc pureboot: one position-independent binary — its own staging loader
The image now runs from any 512-byte slot with every command intact:
control flow stays PC-relative, the write guard keys on the running slot
(the return-address anchor, computed once), the info block is addressed
from that same anchor as a byte pair (no absolute 16-bit address in the
image), and the application jump is an indirect call through a noipa-
laundered pointer to the absolute entry. 'J' — jump to a wire word
address, the one transfer primitive — replaces 'G': the host knows the
application entry from the info block, and moving between loader copies
needs arbitrary targets. The activation window is a compile-time 8 s
(PUREBOOT_TIMEOUT overrides), counted as a single calibrated poll loop.

A refused page no longer poisons the write-once temporary buffer (a real
silicon trap: the next write would program the drained data): every page
write discards the buffer first — CTPB on the tinies, on the mega the same
RWWSRE store that re-enables RWW after programming. The tinies' post-op
busy-waits go with it: their CPU halts through page erase and write.

488 / 502 / 504 B on t13a / t85 / mega — under the tinies' 510-byte budget,
whose last slot word is the host-managed trampoline: the resident's holds
the application entry, a staging copy's the jump through which an abandoned
update still times out into a loader.

The host tool updates the loader with itself: --update-loader installs the
identical image one slot below the resident, jumps into it, lets it rewrite
the resident, and restores the staging region from a state file — each
phase idempotent off the flash state, resumable after any interruption
(t13a: the staging slot carries the reset vector, written last in and
first out; t85: word 0 redirected around the resident rewrite; mega:
fuse-matrix preflight with a hard BOOTSZ gate and --assume-fuses for
simulators). Application flashing recovers by reset from any interruption:
patched page 0 and trampoline first, erase descending, and a walk-region
refusal behind --force on BOOTRST-below-loader megas.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:43 +02:00
06bb06d994 pureboot: harden the sim device runner
Cancel the GPIO bridge's cycle timers with the state they drive: avr_reset
drops the TX latch, whose falling edge starts a spurious decode before
bridge_reset runs, and the stale sampler then interleaves with the loader's
first real answer through the shared shift state — the first post-reset
replies came back corrupted and the knock retries burned the activation
window into the application.

Wrap the mega's registered flash ioctl to re-dispatch page erases with Z
masked to the page boundary: simavr's PGERS handler erases spm_pagesize
bytes from Z & ~1 (its PGWRT path masks correctly), wiping the neighbouring
page when Z sits past the page start, which hardware permits (§26.8.1).
Model the write-once temporary buffer in the tiny NVM module — silicon
refuses a second load per word until the buffer clears, and a last-write-
wins model masks real firmware bugs.

Optional arguments select the reset vector (the mega's fuse profiles) and a
raw flash image to resume from (power-fail tests re-enter a dumped state).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:37:10 +02:00
9a9a7ef0e8 pureboot: review-pass fixes to the host tool and device runner
pureboot.py: reject an empty image file with a clear error instead of
an IndexError deep in the vector-surgery planner; tighten the erase
docstring (order is irrelevant there — every target byte is the same
value, unlike a real flash where page 0 must go last).

pureboot_device.c: the GPIO bridge's bit_cycles used plain truncating
division where the firmware computes its own bit period with
round-to-nearest (uart.hpp: (Clock.hz + Baud.bd/2)/Baud.bd) — one
cycle off per bit on both tinies, harmless in practice but needless
drift against a firmware built to a different constant. Matched
exactly. Also clear the queued-bytes/decode-in-progress bridge state
on the test-only reset signal, so a future reset-mid-transfer scenario
can't feed a freshly reset chip bytes queued for its previous life.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:45:19 +02:00
4acf358dda pureboot: gitignore python bytecode cache
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:43:56 +02:00
11c10f986d pureboot: stop tracking the python bytecode cache
A stray __pycache__/*.pyc from a local test run got swept into the
previous commit's git add. Untracked and gitignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 10:43:36 +02:00
f587f0a26e pureboot: host tool and end-to-end protocol tests, all three chips
pureboot.py (Python stdlib only): images as raw binary or Intel HEX,
flash and EEPROM programming with read-back verify, erase composites,
fuse and info readout, activation-timeout configuration, and the
tinies' reset-vector surgery — the trampoline word below the loader,
page 0 written last.

The test spawns a simavr device (pureboot_device.c) — the mega's USART
as a pty; on the tinies a cycle-timed GPIO<->pty bridge for the polled
software UART plus the NVM module simavr's tiny cores lack (their SPM
opcode ioctls into a void and silently does nothing) — and drives it
with the real tool: knock from reset (erased-flash walk on the tinies),
program and verify both memories, timeout write, session reconnect, an
external reset through the patched vector, hand-over, and the fixture
application's banner. Results are cross-checked against ground-truth
memory dumps and an independent decode of the surgery's rjmp words,
red-verified against a sabotaged encoder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 05:54:15 +02:00
27c768a102 pureboot: the device — one pure C++ source, 512 bytes, every chip
No inline assembly, no global register variables; libavr does the
datasheet work. The device speaks primitives — flash read/page-program,
EEPROM read/write, fuse read, info block, EEPROM-resident activation
timeout, hand-over — and verify, erase, reset-vector surgery, and
timeout configuration live in the host tool. 490 B on the ATtiny13A,
510 B on the ATtiny85, 484 B on the ATmega328P, each linked into the
top 512 bytes of flash; per-chip size tests gate all three.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 05:33:28 +02:00
d989b5b8cd tsb: third size pass — restructure to the oracle's shape
The second pass concluded the 168 B tricks->asm gap was per-call ABI
cost. Most of it was structure. Rebuilt around the oracle's own shape —
argless noinline primitives over a whole-loader call-saved register
protocol (g_addr in Y, count r16, window r7, direction latch r6), a
top-down erase_below whose loop tests against zero and hands callers
g_addr = 0 for free, bounded rx everywhere (a silent host unwinds to
the app from any state, as the oracle does), and a named tsb_app entry
that --pmem-wrap-around=32k relaxes to the wrapped rjmp:

  tsb_asm    510 B in the 512 B section (oracle: 500), C++ except rx
             and the page-store loop — the two routines whose remaining
             cost is the calling convention itself (~30 asm lines, was
             ~280)
  tsb_tricks 526 B, no assembly at all (was 666)
  tsb_pure   836 B, still one readable function per command (was 842)

Every g_* update placement works around a GCC 16.1 wrong-code bug
(stores into global register variables deleted when only callees read
them — repro and rules in libavr dev/lessons.md). Also fixes two
latent hardware bugs all earlier tiers carried, masked by simavr's
zeroed register file: the crt-less entries never established
__zero_reg__ = 0, and the direction latch was read before written —
power-on registers are undefined.

All tiers full oracle feature parity, protocol tests green in both
libavr modes, .text byte-identical across modes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:00:27 +02:00
314e422b19 tsb: beat the first-pass size floors (tricks 666, pure 842)
tricks 778->666: always_inline every single-call handler into the
[[noreturn]] reset entry (which pays no prologue, so their push/pop of
call-saved registers vanishes), walk the page pointer in Y (adiw, base
recovered as g_addr-page) instead of recomputing Z=base+offset, bring
the UART up in the two registers that are not already at their reset
value, and seed the activation counter as __uint24.

pure 896->842: TU-local internal linkage (proper hygiene, and it lets
the compiler inline the one-call handlers), a byte-wide activation
count, __uint24 timeout. Still one readable function per command.

asm unchanged at 498: its C++-expressible parts are already C++; the
core stays asm (the 666 B all-tricks tier is 168 B over — per-call ABI
tax, not a feature). All three cross-mode byte-identical, protocol green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 20:15:49 +02:00
3e0717ef06 tsb: drive each tier to its size floor
asm 502->498 B (below the oracle's 500): the stack bring-up moves to plain C++,
and a register is reserved for the config-page high byte instead of reloading it
at each app-flash-boundary compare. tricks 808->778 B: shared erase/rww helpers
plus the libavr half-duplex W1C fix. pure 950->896 B and no SRAM: streams
rx->SPM/EEPROM instead of staging a 128 B page buffer. All three keep full oracle
feature parity and stay byte-identical across modes; protocol tests (round-trip +
password + emergency erase) green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 18:47:38 +02:00
812be910c1 tsb: document the three tiers at full parity in the build file
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:50:18 +02:00
34529c8031 tsb: protocol test covers the password gate and emergency erase
Each scenario group now runs on its own freshly-reset device: the round-trip
on a blank config page, plus a password-config device that must be sent the
password after the knock to activate, and an emergency-erase device where a
0-byte + two confirms wipes flash, EEPROM and the config page (verified by
reading all three back as 0xff). All three tiers pass every group.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:49:20 +02:00
06bd56af50 tsb: pure and tricks tiers reach full oracle feature parity
Both tiers gain the features the asm tier already carries — one-wire
half-duplex (via libavr's new .half_duplex), the config-page activation
timeout, and emergency erase (password \0 + double-confirm wipes flash,
EEPROM and the config page) — on top of the watchdog bail, password gate and
config/flash/EEPROM read-write they already had. pure stays idiomatic
(flash_table info block, one function per command) at 950 B; tricks keeps its
compiler trickery (call-saved global-register page walk, unified runtime-flag
paths pinned noinline/noclone, streaming stores, arithmetic command decode)
at 808 B. Both byte-identical across generated and reflect modes; the size
gradient across the three tiers is now 502 / 808 / 950 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:47:21 +02:00
75bb84fa3e tsb: asm tier reaches full oracle feature parity at 502 B
Rewrite the inline-asm tier so it matches the hand-written fixed-baud oracle's
feature set inside the 512 B boot section: watchdog-reset bail, one-wire
half-duplex (RXEN/TXEN toggled per direction, TX turnaround guard),
config-page activation timeout, the password gate (wrong byte hangs draining
the UART), emergency erase (password \0 + double-confirm wipes flash, EEPROM
and the config page), and config/flash/EEPROM read-write. Every geometry,
baud and info-block constant comes from libavr consteval; only the dense
control flow is hand-written. 502 B, byte-identical across generated and
reflect modes.

Test harness: seed the config page from TSB_CONFIG so the password and
emergency-erase paths are exercisable, and clear simavr's AVR_UART_FLAG_POLL_
SLEEP — a host-CPU-saving usleep(1)-per-idle-poll hack that models no hardware
and paces a one-wire loader (which releases TX between bytes) in real time,
distorting protocol timing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 16:23:16 +02:00
778b8a0100 tsb: vendor the fixed-baud assembly oracle as the size/feature bar
The Seed Robotics native-UART fixed-baud TinySafeBoot (GPLv3), reference
only — not built. Assembles to 500 B with the full feature set, proving
≤512 B and full feature parity are simultaneously reachable. Also drops the
stale empty stk500v2/ leftover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:29:58 +02:00
48d11fe8e3 tsb: use the named register surface
Direct register access now reads through the named surface
(hw::mcusr::wdrf.test(), hw::ucsr0b::write(...)) instead of the string form,
matching how libavr itself is written. Zero-overhead: pure 740 B, tricks 658 B,
asm 508 B unchanged, all byte-identical across modes, protocol green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:55:01 +02:00
bb4784ae0c tsb: refactor the pure tier onto libavr sugar
The showcase tier now leans on the helpers it fed back instead of reaching under
them: the info block is an avr::flash_table (no raw [[gnu::progmem]]), a page is
filled with spm::fill(addr, span) (no hand-packed lo|hi<<8 loop), and the
WDT-reset bail reads field<"MCUSR","WDRF">::test() (no read() & {}(1).value).

Zero-overhead throughout: .text stays 740 B, byte-identical across generated and
reflect modes, protocol test green. The info block streams through the existing
address-based send_flash rather than a range-for over the flash_table — the
range-for is a distinct loop that cannot share the loader's one flash streamer,
so it would add 14 B for no functional gain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:33:53 +02:00
d07586652d tsb: slim the port branch to the libavr reimplementation
main carried the whole pre-libavr tree beside the port: the other-bootloader
directories (blink, stk500v2), the Atmel Studio solution/project, and — dead in
the tsb dir itself — four submodule links to the superseded io/flash/uart/type
libraries the libavr sources never include. None are build inputs; CMake drives
the three variants through FetchContent. master keeps the full legacy tree
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:12:56 +02:00
3a2f2a23c2 tsb: drop the local -O3 strip, now handled by the libavr toolchain
The -O3 leak is fixed upstream (cmake/release-os.cmake via CMAKE_PROJECT_INCLUDE),
so the port no longer needs its own string(REPLACE); a Release build is -Os
through the toolchain file. Verified: all three variants build at their sizes
(508/658/740) and pass the size + protocol ctest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 10:52:13 +02:00
4fa058653e tsb: reimplement TinySafeBoot on libavr in three size tiers
The native-UART fixed-baud TinySafeBoot protocol, ported onto libavr as a
crt-free boot-section loader, in three variants that trade clarity for size:

  tsb_pure   740 B  idiomatic C++: SRAM page buffer, separate flash/EEPROM
                    leaves, shared framing; the polled `unused` guard posture.
  tsb_tricks 658 B  unified runtime-flag paths (noinline/noclone), call-saved
                    global-register page walk — attributes only, no asm.
  tsb_asm    508 B  streaming store + hand-rolled UART/SPM/EEPROM/erase loops;
                    fits the 512 B boot section (BOOTSZ=11). Trims the optional
                    password gate and WDT-reset bail — unreachable in C++ with
                    both (hand-asm is ~15 % denser). Tiers 1-2 keep them and
                    live in the 1 KB section they fit.

All three are .text byte-identical across libavr's generated and reflect modes.
The CMake build strips the leaked -O3 (a Release build is silently -O3, not the
-Os this loader is measured against) and gates each variant's size against its
section. A simavr harness (test/device.c + test/tsbtest.py) drives the real wire
protocol over a pty and flashes the device; the size and protocol tests run in
ctest. Verified byte-for-byte against the reference tsbloader_adv (C#/mono):
activate, read info, flash write + verify.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 05:00:51 +02:00
16 changed files with 748 additions and 656 deletions

3
.gitmodules vendored Normal file
View File

@@ -0,0 +1,3 @@
[submodule "libavr"]
path = libavr
url = ../libavr.git

View File

@@ -8,6 +8,9 @@ include(FetchContent)
if(NOT LIBAVR_ROOT AND DEFINED ENV{LIBAVR_ROOT}) if(NOT LIBAVR_ROOT AND DEFINED ENV{LIBAVR_ROOT})
set(LIBAVR_ROOT $ENV{LIBAVR_ROOT}) set(LIBAVR_ROOT $ENV{LIBAVR_ROOT})
endif() endif()
if(NOT LIBAVR_ROOT)
set(LIBAVR_ROOT ${CMAKE_CURRENT_SOURCE_DIR}/libavr)
endif()
if(LIBAVR_ROOT) if(LIBAVR_ROOT)
FetchContent_Declare(libavr SOURCE_DIR ${LIBAVR_ROOT}) FetchContent_Declare(libavr SOURCE_DIR ${LIBAVR_ROOT})
else() else()
@@ -231,13 +234,12 @@ if(PROJECT_IS_TOP_LEVEL)
endif() endif()
# The size matrix: every configuration axis that could move the image # The size matrix: every configuration axis that could move the image
# size — the serial backend (different code), the USART instance # size — the serial backend (different code), the clock and its ladder
# (different registers), the clock (different constants), and the baud # baud (different constants and divisor shapes), the USART instance
# through the two shapes its bit timing takes — each combination must # (different register class) — each combination must still fit the
# still fit the chip's slot budget. Pins are size-neutral (port and bit # chip's slot budget. Pins are size-neutral (port and bit are immediate
# are immediate operands) and the timeout is a constant, so neither adds # operands) and the timeout is a constant, so neither adds an axis. The
# an axis. The stock build is one point of this matrix and already has # stock build is one point of this matrix and already has its test.
# its test.
function(pureboot_size_variant name) function(pureboot_size_variant name)
pureboot_add_loader(${name} ${ARGN}) pureboot_add_loader(${name} ${ARGN})
add_test(NAME ${name}.size add_test(NAME ${name}.size
@@ -261,26 +263,11 @@ if(PROJECT_IS_TOP_LEVEL)
if(PUREBOOT_HAS_USART AND NOT _matrix_hz EQUAL _pb_stock_hz) if(PUREBOOT_HAS_USART AND NOT _matrix_hz EQUAL _pb_stock_hz)
pureboot_size_variant(pureboot_hw_${_matrix_khz}k CLOCK ${_matrix_hz} SERIAL hardware) pureboot_size_variant(pureboot_hw_${_matrix_khz}k CLOCK ${_matrix_hz} SERIAL hardware)
endif() endif()
if(PUREBOOT_HAS_USART1 AND NOT _matrix_hz EQUAL _pb_stock_hz)
pureboot_size_variant(pureboot_usart1_${_matrix_khz}k CLOCK ${_matrix_hz} USART 1)
endif()
endforeach() endforeach()
if(PUREBOOT_HAS_USART1) if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_usart1 USART 1) pureboot_size_variant(pureboot_usart1 USART 1)
endif() endif()
# The baud axis, whose one size-bearing shape the ladder never picks: a
# software UART spins out each bit with _delay_loop_1 while the count
# fits a byte and with the 16-bit _delay_loop_2 beyond it, two words more
# setup at every one of its five sites — the largest image the
# configuration space produces. The ladder default takes the *fastest*
# rate a clock reaches, which always lands in the byte, so the wide form
# needs the slowest ladder rate against the fastest clock to appear. The
# hardware USART has no such shape: its baud is a divisor constant, and
# the ladder's U2X solutions are already its larger form.
list(GET _matrix_clocks -1 _matrix_top_hz)
pureboot_size_variant(pureboot_sw_wide CLOCK ${_matrix_top_hz} BAUD 9600 SERIAL software)
# One configured deployment end to end — a real board's shape rather # One configured deployment end to end — a real board's shape rather
# than the stock assumption: the ATmega328P on its shipped 1 MHz fuses, # than the stock assumption: the ATmega328P on its shipped 1 MHz fuses,
# the software UART on hand-picked pins (TX = PB1, RX = PB5), the ladder # the software UART on hand-picked pins (TX = PB1, RX = PB5), the ladder

View File

@@ -6,7 +6,7 @@
"hidden": true, "hidden": true,
"generator": "Ninja", "generator": "Ninja",
"binaryDir": "${sourceDir}/build/${presetName}", "binaryDir": "${sourceDir}/build/${presetName}",
"toolchainFile": "$env{LIBAVR_ROOT}/cmake/avr-toolchain.cmake", "toolchainFile": "${sourceDir}/libavr/cmake/avr-toolchain.cmake",
"cacheVariables": { "cacheVariables": {
"CMAKE_BUILD_TYPE": "Release", "CMAKE_BUILD_TYPE": "Release",
"CMAKE_EXPORT_COMPILE_COMMANDS": "ON", "CMAKE_EXPORT_COMPILE_COMMANDS": "ON",

1
libavr Submodule

Submodule libavr added at e81dad0131

View File

@@ -1,16 +1,27 @@
# pureboot as a consumable CMake unit: the per-chip geometry, the default baud # pureboot as a consumable CMake unit: the per-chip geometry, the default
# ladder, and pureboot_add_loader() — the one way a loader target is created. # baud ladder, and pureboot_add_loader() — the one way a loader target is
# A downstream project brings its usual libavr setup (the `libavr` target and # created, both by this port's own build and by a downstream project. A
# downstream project brings its usual libavr setup (the `libavr` target and
# the LIBAVR_MCU toolchain preset), adds this directory, and states its # the LIBAVR_MCU toolchain preset), adds this directory, and states its
# deployment; every argument is optional (README.md): # deployment:
# #
# add_subdirectory(bootloader/pureboot) # add_subdirectory(bootloader/pureboot)
# pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5) # pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5)
#
# Every argument is optional — CLOCK defaults to the family assumption
# below, BAUD to the fastest standard rate the clock reaches within 2.5 %
# (the ladder), SERIAL to the chip's hardware USART where it has one
# (`hardware`/`software` force a backend, USART 1 picks the second
# instance), RX/TX to pb0/pb1 for the software UART, TIMEOUT to 8 s.
# Infeasible picks fail the build by name: libavr's baud-error and
# software-UART cycle-floor static asserts re-check whatever is passed.
# Per-family geometry, deployment defaults, and the linker wrap the PC modulo # Per-family geometry: flash/page/EEPROM sizes and the linker wrap the PC
# needs. The slot is 512 bytes on every chip. The USART flags mirror the # modulo needs, the loader slot (each chip's smallest boot sector — 1 KiB on
# hardware inventory the loader's own static asserts check — the plain 644 is # the word-addressed 1284s), and the deployment defaults (crystal assumption
# the x4 family's one single-USART die (Atmel-2593). # on the megas, calibrated RC on the tinies). The USART flags mirror the
# hardware inventory the loader's own static asserts check (the plain 644 is
# the x4 family's one single-USART die, Atmel-2593).
set(_pb_has_usart 1) set(_pb_has_usart 1)
set(_pb_has_usart1 0) set(_pb_has_usart1 0)
if(LIBAVR_MCU MATCHES "^attiny13a?$") if(LIBAVR_MCU MATCHES "^attiny13a?$")
@@ -80,9 +91,10 @@ elseif(LIBAVR_MCU MATCHES "^atmega324(a|p|pa)$")
set(_pb_eeprom 1024) set(_pb_eeprom 1024)
set(_pb_has_usart1 1) set(_pb_has_usart1 1)
elseif(LIBAVR_MCU MATCHES "^atmega644(a|p|pa)?$") elseif(LIBAVR_MCU MATCHES "^atmega644(a|p|pa)?$")
# 64 KiB is exactly the 16-bit byte space, so plain LPM still reaches # 64 KiB is exactly the 16-bit byte space: plain LPM reaches everything,
# everything and the wire stays byte-addressed. The plain 644 is the # and the smallest boot section (1 KiB) holds the loader and its staging
# family's one single-USART die. # slot together (see README.md). The plain 644 is the family's one
# single-USART die.
set(_pb_flash 65536) set(_pb_flash 65536)
set(_pb_wrap -Wl,--pmem-wrap-around=64k) set(_pb_wrap -Wl,--pmem-wrap-around=64k)
set(_pb_page 256) set(_pb_page 256)
@@ -92,25 +104,33 @@ elseif(LIBAVR_MCU MATCHES "^atmega644(a|p|pa)?$")
set(_pb_has_usart1 1) set(_pb_has_usart1 1)
endif() endif()
elseif(LIBAVR_MCU MATCHES "^atmega1284p?$") elseif(LIBAVR_MCU MATCHES "^atmega1284p?$")
# 128 KiB: wire addresses are words, reads go through ELPM, and the PC's # 128 KiB: wire flash addresses are word addresses, reads go through
# modulo wrap exceeds what --pmem-wrap-around models. # ELPM, and the PC's modulo wrap exceeds what --pmem-wrap-around models.
# The slot is 1 KiB — this chip's own smallest boot sector; the far
# machinery cannot fit 512 B (see README.md).
set(_pb_flash 131072) set(_pb_flash 131072)
set(_pb_wrap "") set(_pb_wrap "")
set(_pb_page 256) set(_pb_page 256)
set(_pb_hz 16000000) set(_pb_hz 16000000)
set(_pb_eeprom 4096) set(_pb_eeprom 4096)
set(_pb_slot 1024)
set(_pb_limit 1024)
set(_pb_has_usart1 1) set(_pb_has_usart1 1)
else() else()
message(FATAL_ERROR "pureboot: no geometry for ${LIBAVR_MCU}") message(FATAL_ERROR "pureboot: no geometry for ${LIBAVR_MCU}")
endif() endif()
set(_pb_slot 512) if(NOT DEFINED _pb_slot)
set(_pb_slot 512)
endif()
math(EXPR _pb_base "${_pb_flash} - ${_pb_slot}") math(EXPR _pb_base "${_pb_flash} - ${_pb_slot}")
math(EXPR _pb_base_hex "${_pb_base}" OUTPUT_FORMAT HEXADECIMAL) math(EXPR _pb_base_hex "${_pb_base}" OUTPUT_FORMAT HEXADECIMAL)
# Patched-vector chips hand over through the trampoline word below the slot, # Patched-vector chips hand over through the trampoline word below the slot,
# which is also the slot's own last word — their budget is slot 2. # which is also the slot's own last word — their budget is slot 2.
if(LIBAVR_MCU MATCHES "^atmega" AND NOT LIBAVR_MCU MATCHES "^atmega48") if(LIBAVR_MCU MATCHES "^atmega" AND NOT LIBAVR_MCU MATCHES "^atmega48")
set(_pb_app 0) set(_pb_app 0)
set(_pb_limit ${_pb_slot}) if(NOT DEFINED _pb_limit)
set(_pb_limit ${_pb_slot})
endif()
else() else()
math(EXPR _pb_app "${_pb_base} - 2") math(EXPR _pb_app "${_pb_base} - 2")
math(EXPR _pb_limit "${_pb_slot} - 2") math(EXPR _pb_limit "${_pb_slot} - 2")
@@ -146,11 +166,12 @@ set(PUREBOOT_HAS_USART ${_pb_has_usart} PARENT_SCOPE)
set(PUREBOOT_HAS_USART1 ${_pb_has_usart1} PARENT_SCOPE) set(PUREBOOT_HAS_USART1 ${_pb_has_usart1} PARENT_SCOPE)
set(PUREBOOT_SIM_MCU ${_pb_sim_mcu} PARENT_SCOPE) set(PUREBOOT_SIM_MCU ${_pb_sim_mcu} PARENT_SCOPE)
# The fastest standard rate the clock reaches within 2.5 %, by the same # The fastest standard rate the clock reaches within 2.5 % the same
# best-of-U2X-and-plain divisor search libavr's solve_baud runs, so a default # best-of-U2X-and-plain divisor search libavr's solve_baud runs, so a
# never trips the compile-time error it is checked against. A software build # default never trips the compile-time error it is checked against. A
# also needs the polled receiver's 100-cycles-a-bit floor: at low clocks the # software build additionally requires the polled receiver's 100-cycles-a-bit
# U2X divisor reaches rates the bit-banged sampler cannot. # floor (its own static assert): at low clocks the U2X divisor still reaches
# rates the bit-banged sampler cannot, so the backend gates the ladder.
function(pureboot_default_baud clock software outvar) function(pureboot_default_baud clock software outvar)
foreach(baud 115200 57600 38400 19200 9600) foreach(baud 115200 57600 38400 19200 9600)
math(EXPR _cycles "${clock} / ${baud}") math(EXPR _cycles "${clock} / ${baud}")
@@ -182,10 +203,11 @@ endfunction()
# [SERIAL auto|hardware|software] [USART <n>] # [SERIAL auto|hardware|software] [USART <n>]
# [RX <pin>] [TX <pin>] [TIMEOUT <s>]) # [RX <pin>] [TX <pin>] [TIMEOUT <s>])
# #
# The loader target plus its flashable images (<name>.hex for a programmer, # Creates the loader target plus its flashable images (<name>.hex for a
# <name>.bin for --update-loader). The resolved deployment is stamped on the # programmer, <name>.bin for --update-loader) and stamps the resolved
# target as PUREBOOT_HZ / PUREBOOT_BAUD / PUREBOOT_LINK (the link spelled # deployment on the target: the PUREBOOT_HZ, PUREBOOT_BAUD and PUREBOOT_LINK
# usart0, usart1 or sw:<RX>,<TX>) — what a test harness speaks to it with. # properties (the link as usart0/usart1/sw:<RX>,<TX> — what a test harness
# needs to speak to the build).
function(pureboot_add_loader name) function(pureboot_add_loader name)
cmake_parse_arguments(PB "" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT" "" ${ARGN}) cmake_parse_arguments(PB "" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT" "" ${ARGN})
if(PB_UNPARSED_ARGUMENTS) if(PB_UNPARSED_ARGUMENTS)
@@ -250,7 +272,8 @@ function(pureboot_add_loader name)
endif() endif()
endforeach() endforeach()
set(_serial_defines PUREBOOT_SOFT_SERIAL PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX}) set(_serial_defines PUREBOOT_SOFT_SERIAL PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
# sw:<RX>,<TX> as port letter and bit, upcased. # The link spec a test harness drives a GPIO bridge with: sw:<RX>,<TX>
# as the port letter and bit, the loader's own pin naming upcased.
string(SUBSTRING ${PB_RX} 1 2 _rx_pin) string(SUBSTRING ${PB_RX} 1 2 _rx_pin)
string(SUBSTRING ${PB_TX} 1 2 _tx_pin) string(SUBSTRING ${PB_TX} 1 2 _tx_pin)
string(TOUPPER "sw:${_rx_pin},${_tx_pin}" _link) string(TOUPPER "sw:${_rx_pin},${_tx_pin}" _link)
@@ -271,22 +294,21 @@ function(pureboot_add_loader name)
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/pureboot.cpp) add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/pureboot.cpp)
target_link_libraries(${name} PRIVATE libavr) target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE ${_defines}) target_compile_definitions(${name} PRIVATE ${_defines})
# Codegen shaping for the loader TU only, worth 1436 B depending on the # Codegen shaping for the loader TU only, worth ~40 B on every chip and
# chip. At -Os GCC otherwise rewrites the byte-stream loops' counters into # what carries the far-flash 1284 build under 512. At -Os GCC otherwise
# end-pointer forms that cost registers (-fno-ivopts, # rewrites the byte-stream loops' counters into end-pointer forms that
# -fno-split-wide-types), leaves register pressure on the table with the # cost registers (-fno-ivopts, -fno-split-wide-types), leaves register
# default allocator (-fira-algorithm=priority), and keeps loop-invariant # pressure on the table with the default allocator
# immediates and expression temporaries in registers # (-fira-algorithm=priority), and spends bytes on rewrites a
# (-fno-move-loop-invariants, -fno-tree-ter) — but every loop body here # straight-line loader gains nothing from.
# contains a call, so a register held across it costs more than the
# load-immediate it saves.
target_compile_options(${name} PRIVATE target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types) -fno-ivopts -fira-algorithm=priority -fno-expensive-optimizations -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex} target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap}) -Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>) add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
# The ELF is a container, never flashed: .hex for a programmer, .bin (the # The ELF is a container (symbols, section headers), never flashed; the
# slot's bare bytes) for --update-loader. # flashable forms sit beside it: .hex for a programmer, .bin (the slot's
# bare bytes) for the host tool's raw path and --update-loader.
add_custom_command(TARGET ${name} POST_BUILD add_custom_command(TARGET ${name} POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O ihex -R .eeprom COMMAND ${CMAKE_OBJCOPY} -O ihex -R .eeprom
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.hex $<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.hex

View File

@@ -2,61 +2,46 @@
A serial bootloader on [libavr](https://git.blackmark.me/avr/libavr), pure by A serial bootloader on [libavr](https://git.blackmark.me/avr/libavr), pure by
constraint: one C++ source, no inline assembly, no global register variables constraint: one C++ source, no inline assembly, no global register variables
(attributes and compiler flags allowed), **512 bytes on every chip libavr (attributes and compiler flags allowed), built for **every chip libavr
targets — all 37**. The device speaks primitives; every composite — verify, targets — all 37 — in 512 bytes each**: 434 B on the tiny13s, 438442 B on
erase, reset-vector surgery, updating the loader itself — lives in the host the tiny25/45/85, 412452 B across the megas, and 506 B on the
tool (`pureboot.py`). ATmega1284/1284P, whose far-flash machinery (ELPM reads, RAMPZ page commands,
word-addressed wire) is the heaviest. Those are the stock deployments;
choosing the software UART where the chip has a USART costs 846 B more (a
bit-bang against a peripheral), which every chip still absorbs inside its
slot — on the 1284s that means their 1 KiB boot sector, where the
software-serial image lands at 546 B. Bringing the 1284's default build
under 512 at all is what the loop-placement attributes on the byte streamers
(`pureboot.cpp`) and the codegen flags on the loader TU (`CMakeLists.txt`)
are for; measured against each chip's own budget the tightest is the
ATmega328P, 50 B spare. Clock, baud, serial backend and
pins are per-build configuration (below); the size matrix in the test suite
holds every combination inside its slot. The device speaks primitives; every
composite — verify, erase, reset-vector surgery, updating the loader itself —
lives in the host tool (`pureboot.py`).
The 1284s still *deploy* in a 1 KiB slot, their smallest boot sector being
512 words; at 506 B the image would also fit the 644's
two-512-byte-slots-per-boot-sector geometry.
The image is **position-independent**: control flow is PC-relative, the The image is **position-independent**: control flow is PC-relative, the
read/write paths take wire addresses, the write guard protects the slot the read/write paths take wire addresses, the write guard protects the slot the
code is *running* in (from the runtime return address), the info block is code is *running* in (from the runtime return address), the info block is
addressed from that same anchor, and the application jump is an indirect call addressed from that same anchor, and the application jump is an indirect
to an absolute entry. The identical binary therefore runs from any slot with call to an absolute entry. The identical binary therefore runs from any
every command intact, which makes pureboot **its own staging loader**: the slot with every command intact which makes pureboot **its own staging
host installs the same binary one slot below the resident, jumps into it, and loader**: the host installs the same binary one slot below the resident,
lets it rewrite the resident. jumps into it, and lets it rewrite the resident. The slot is 512 bytes
(1 KiB on the word-addressed large chips, matching their boot-sector
## Chips minimum); on the tinies the budget is 510, not 512: a slot's last word
belongs to the host-managed trampoline (below).
Sizes are the default configuration: the hardware USART0 at 115200 8N1 on a
16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at 57600 8N1 on
the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above). Every axis moves
per build — see *Configuration*; the largest image any of them produces is a
software UART at a slow baud, which on the 1284s is 494 B, the tightest fit in
the whole matrix at 18 B spare.
| Chip | Flash | Loader at | Link | Size |
|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 416 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 420 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 424 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 424 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 396 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 400 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 400 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 414 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 434 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 432 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 478 B |
† No hardware boot section: the host patches the reset vector, and the budget
is 510 bytes, since the slot's last word is the trampoline.
The 1284s are the heaviest because they alone carry the far-flash machinery —
ELPM reads, RAMPZ page commands, a word-addressed wire.
The software UART enables the RX pull-up; TX idles high. All multi-byte wire
quantities are little-endian.
## Configuration ## Configuration
Every deployment axis is a build parameter of `pureboot_add_loader()` (in Every deployment axis is a build parameter, resolved by the CMake function
`pureboot/CMakeLists.txt`) — the one way a loader target is created, by this `pureboot_add_loader()` (in `pureboot/CMakeLists.txt`) — the one way a
repo's build and by a downstream project alike: loader target is created, by this repo's own build and by a downstream
project alike:
| Argument | Meaning | Default | | Argument | Meaning | Default |
|---|---|---| |---|---|---|
@@ -70,14 +55,15 @@ repo's build and by a downstream project alike:
The default baud is the fastest of 115200/57600/38400/19200/9600 the clock The default baud is the fastest of 115200/57600/38400/19200/9600 the clock
reaches within 2.5 % — the same U2X-included divisor search libavr's baud reaches within 2.5 % — the same U2X-included divisor search libavr's baud
solver runs — and on a software build additionally within the polled solver runs — and on a software build additionally within the polled
receiver's 100-cycles-a-bit floor. Whatever is picked or overridden is receiver's 100-cycles-a-bit floor. 16 MHz lands 115200, 8 MHz 57600,
re-checked in the compile: an infeasible combination, or a USART the chip does 1 MHz 9600. Whatever is picked or overridden is re-checked in the compile:
not have, fails with a named static assert. an infeasible clock/baud/backend combination, or a USART the chip does not
have, fails with a named static assert.
A downstream project brings its usual libavr setup (the `libavr` target, the A downstream project brings its usual libavr setup (the `libavr` target,
chip via the `LIBAVR_MCU` toolchain preset), consumes this directory, and the chip via the `LIBAVR_MCU` toolchain preset), consumes this directory,
states its deployment — an ATmega328P on its shipped 1 MHz fuses with the and states its deployment — for example an ATmega328P on its shipped
software UART on hand-picked pins, say: 1 MHz fuses with the software UART on hand-picked pins:
```cmake ```cmake
FetchContent_Declare(bootloader GIT_REPOSITORY git@git.blackmark.me:avr/bootloader.git GIT_TAG main) FetchContent_Declare(bootloader GIT_REPOSITORY git@git.blackmark.me:avr/bootloader.git GIT_TAG main)
@@ -88,40 +74,57 @@ pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5)
``` ```
The function emits the ELF plus `myboot.hex` (the programmer artifact) and The function emits the ELF plus `myboot.hex` (the programmer artifact) and
`myboot.bin` (the self-update image), prints the size, and stamps the resolved `myboot.bin` (the self-update image), prints the size, and stamps the
deployment on the target as the `PUREBOOT_HZ`, `PUREBOOT_BAUD` and resolved deployment on the target as the `PUREBOOT_HZ`, `PUREBOOT_BAUD`
`PUREBOOT_LINK` properties — what a flashing script or test harness needs to and `PUREBOOT_LINK` properties — what a flashing script or test harness
speak to the build. This exact deployment runs the full protocol suite in CI needs to speak to the build. This exact example deployment runs the full
(`pureboot.custom`). protocol suite in CI (`pureboot.custom`).
## Link
The stock builds assume the family's natural deployment; any axis moves
per build (above).
| Chip | Serial | Baud | Clock assumed |
|---|---|---|---|
| every ATmega | the hardware USART (USART0), RXD/TXD per pinout | 115200 8N1 | 16 MHz crystal |
| ATtiny25/45/85 | software UART, RX = PB0, TX = PB1 | 57600 8N1 | 8 MHz internal RC |
| ATtiny13/13A | software UART, RX = PB0, TX = PB1 | 57600 8N1 | 9.6 MHz internal RC |
The software-UART RX pin has its pull-up enabled; TX idles high. All
multi-byte quantities on the wire are little-endian.
## Activation ## Activation
Reset enters the loader (BOOTRST on the boot-sectioned megas, the patched Reset enters the loader (BOOTRST on the boot-sectioned megas; the patched
reset vector elsewhere) — except a watchdog reset, which hands straight to the reset vector on the tinies and the boot-section-less m48s) — except a
application, since the application owns its watchdog and must clear WDRF watchdog reset, which hands straight to the application (the application
itself. owns its watchdog; it must clear WDRF itself, which also releases the
WDRF-forced WDE).
The host then knocks `p` then `b`, each awaited byte under a fresh activation The host then has one activation window per awaited byte to knock: `p` then
window; any other byte is discarded and awaited again, so line noise can delay `b`. Each awaited byte gets a fresh window; any other byte is discarded and
the loader but never lock it. A window expiring on an idle line boots the awaited again (line noise cannot lock the loader, only delay it). A window
application. expiring with an idle line boots the application.
The window is a compile-time constant (`TIMEOUT`, 8 s by default), so the whole The window length is a compile-time constant 8 s by default, another
EEPROM belongs to the application — pureboot keeps no state of its own. value via `pureboot_add_loader(... TIMEOUT <s>)` (the stock target keeps
Re-timing a deployed loader is a self-update with a re-timed build. the `PUREBOOT_TIMEOUT` cache variable) — so the whole EEPROM belongs to
the application; pureboot never uses it for its own state. Re-timing a
deployed loader is a self-update with a re-timed build (below).
## Session ## Session
After the knock the loader stays in its command loop until `J` jumps away or After the knock the loader stays in its command loop until `J` jumps away or
the chip resets. Before reading each command it waits for any pending EEPROM the chip resets. Before reading each command it waits for any pending EEPROM
write and sends the prompt `+` (0x2b), which is therefore also the previous write to finish and sends the prompt `+` (0x2b) — the prompt is therefore
command's completion ack. A session is: await `+`, send a command, read its also the completion ack of the previous command. A session is: await `+`,
reply, repeat. send a command, read its reply, repeat.
On chips whose flash exceeds 64 KiB (the 1284s — info-block flag bit 1) the On chips whose flash exceeds 64 KiB (the 1284s — info-block flag bit 1) the
`R`/`W` flash addresses are **word** addresses; everywhere else they are byte `R`/`W` flash addresses are **word** addresses; everywhere else they are byte
addresses (the 644s' 64 KiB is exactly the 16-bit byte space). EEPROM addresses (the 644s' 64 KiB is exactly the 16-bit byte space and stays
addresses and all counts are bytes. byte-addressed). EEPROM addresses are always bytes, counts always bytes.
| Cmd | Arguments | Reply | | Cmd | Arguments | Reply |
|---|---|---| |---|---|---|
@@ -134,75 +137,86 @@ addresses and all counts are bytes.
| `J` | word address (16-bit) | `+`, then execution continues there | | `J` | word address (16-bit) | `+`, then execution continues there |
| other | — | ignored; the loop re-prompts (send a junk byte, await `+`, to resync) | | other | — | ignored; the loop re-prompts (send a junk byte, await `+`, to resync) |
`W` streams exactly one page-aligned SPM page (size from the info block) into `W` streams exactly one SPM page (size from the info block) into the buffer,
the buffer, then erases and programs — except pages inside the 512-byte slot then erases and programs; the address must be page-aligned. Pages inside the
the loader is *running* in, which are drained and left alone, so a broken host 512-byte slot the loader is *running* in are drained but never programmed — a
cannot brick the running copy and a staged copy may rewrite the resident. broken host cannot brick the running copy, and a staged copy may rewrite the
resident slot.
The loader never clears the SPM buffer before a fill, so **one `W` may program The loader never clears the SPM buffer before a fill, so **one `W` may
the wrong bytes, and the host is what fixes it**. The buffer is write-once per program the wrong bytes, and the host is what fixes it**. The buffer is
word until cleared, and two things leave words in it: a refused page, and — write-once per word until cleared, and two things leave words in it: a
where SPM runs from anywhere, the tinies and the m48s — an application that refused page (drained, never programmed) and — where SPM runs from anywhere,
self-programmed before entering. The next `W` takes those stale words and the tinies and the m48s — an application that self-programmed before
clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on the entering. The next `W` takes those stale words, and clears them: a page write
tinies), so repeating it programs correctly. The host therefore verifies every auto-erases the buffer (§26.2.1; §19.2 on the tinies), so repeating it
page it writes and rewrites what comes back wrong (three retries, then it programs correctly. The host therefore verifies every page it writes and
stops). rewrites what comes back wrong (three retries, then it stops); a host that
programs without reading back cannot trust the first `W` after either event.
`w` is host-paced: send the next byte only after the previous byte's `+`. `F` `w` is host-paced: send the next byte only after the previous
returns the bytes in the hardware's Z order; on a chip without an extended byte's `+`. `F` returns the bytes in the hardware's Z order; on a chip
fuse byte that slot carries no meaning. Fuse *writing* does not exist — SPM without an extended fuse byte (the ATtiny13A) that slot carries no meaning.
reaches flash and boot lock bits only. Fuse *writing* does not exist: SPM reaches flash (and, on the mega, lock
bits) only — fuse bytes are external-programming territory by hardware.
`J` is the one control-transfer primitive: it runs the application (word 0 or `J` is the one control-transfer primitive: the host uses it to run the
the trampoline word, both known from the info block) and moves between loader application (word 0 on the mega, the trampoline word on the tinies — both
copies during a self-update. A jump to a slot's base re-enters that copy's own known from the info block) and to move between loader copies during a
startup, which must then be knocked afresh. self-update. A jump to a loader slot's base re-enters that copy's own
startup; it must then be knocked afresh.
The info block (`b`): The info block (`b`):
| Offset | Content | | Offset | Content |
|---|---| |---|---|
| 02 | `'P'`, `'B'`, pureboot version (3) | | 02 | `'P'`, `'B'`, protocol version (1) |
| 35 | device signature | | 35 | device signature |
| 6 | SPM page size in bytes (0 means 256) | | 6 | SPM page size in bytes (0 means 256) |
| 78 | loader base — application flash ends here (a word address when bit 1 is set) | | 78 | loader base — application flash ends here (a word address when bit 1 is set) |
| 910 | EEPROM size | | 910 | EEPROM size |
| 11 | bit 0: host must patch the reset vector (no hardware boot section); bit 1: flash wire addresses are word addresses | | 11 | bit 0: host must patch the reset vector (no hardware boot section); bit 1: flash wire addresses are word addresses |
## Version Composites are the host's job: verify = read back and compare, erase =
write `0xff` (per page for flash, per byte for EEPROM).
The info block's third byte is the **pureboot version** — the loader's one
identity number, and the only way to tell what a deployed loader is. Nothing
else is numbered: the wire protocol has no version, a pureboot version implies
it, and the host tool holds that map. The tool states the window of loader
versions it speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a
version that changes the protocol becomes the new floor there. None has so
far: 1 through 3 speak the identical session. A loader newer than the tool is
refused by name rather than decoded on the assumption that nothing moved.
The tool carries its own version, free to drift; `--version` prints it and the
window.
## Deployment ## Deployment
The build leaves three artifacts per chip. The ELF is a container for the The build leaves three artifacts per chip. The ELF is a container for the
tests and objcopy, never flashed. The **.hex is the programmer artifact**: it tests and objcopy never flashed. The **.hex is the programmer artifact**:
carries its own addresses and lands the loader in its top slot, touching it carries its own addresses and lands the loader in its top slot,
nothing else. The **.bin is the self-update image** — the slot's bare bytes. touching nothing else. The **.bin is the self-update image** — the slot's
bare bytes with no addressing, which a programmer would put at address 0.
On a boot-sectioned mega a copy at 0 is dead weight (SPM only executes
from the boot section, so it cannot even heal itself — reflash the .hex);
on the patched-vector chips it *runs* (the image is position-independent
and reset enters word 0), reports its canonical geometry, and the ordinary
`--update-loader` flow re-homes a build into the top slot from any
position — the staging install and the word-0 redirect execute from
copies outside page 0's slot, and a copy sitting in the staging slot
itself is recognized as the installed staging copy and left in place (it
streams the new resident like any staged copy, so an older build installs
a newer one). `pureboot.rehome` is the acceptance test for both
positions. Flashing the application afterwards overwrites the stale copy,
vector surgery included.
**Boot-sectioned megas**: program the loader at `flash 512` with an external **Boot-sectioned megas**: program the loader at `flash slot` with an
programmer. Every such mega has a BOOTSZ step whose boot section is exactly external programmer. Every such mega has a BOOTSZ step whose boot section
the 512-byte slot — the second-smallest step on the 8 KiB and 16 KiB chips, is exactly the loader slot — 512 B, the second-smallest step on the 8 KiB
the smallest on the 32 KiB ones — so the ATmega328P profiles below apply to and 16 KiB chips (m8, m88, m16, m168, m164), the smallest on the 32 KiB
every one of them with its own addresses; the per-chip BOOTSZ ladders live in ones (m32, m328, m324); on the 1284s that step is the smallest, 512 words,
the host tool (`BOOT_FUSE`). which is why their slot is 1 KiB — so the ATmega328P profiles below apply
to every one of them with its own addresses and slot size; the per-chip
BOOTSZ ladders live in the host tool (`BOOT_FUSE`). The 1284s' numbers:
standalone = BOOTSZ 512 words (reset at the loader base 0x1fc00);
self-update = 1024 words, covering both 1 KiB slots, the loader-first
reset landing at 0x1f800 — the staging slot, walked across when erased.
The **644s and 1284s** are the geometry's sweet spot: their smallest boot The **644s** are the geometry's sweet spot: their smallest boot section
section (512 words = 1 KiB) is exactly *two* slots, so the resident and its (512 words = 1 KiB) is exactly *two* 512-byte slots, so the resident and
staging slot both live inside the minimum section. Self-update needs no fuse its staging slot both live inside the minimum section — self-update needs
step up, and the standalone profile does not exist reset lands one erased no fuse step up, and the standalone profile does not exist (reset lands at
slot below the loader (0xfc00 / 0x1fc00) and walks up into it. 0xfc00, one erased slot below the loader: the loader-first walk built in).
ATmega328P profiles (addresses for its 32 KiB): ATmega328P profiles (addresses for its 32 KiB):
@@ -212,135 +226,154 @@ ATmega328P profiles (addresses for its 32 KiB):
| 512 words (1 KB) | unprogrammed | *Self-update, app-first*: reset always boots the application, which owns all 31.5 KB and must offer its own jump to 0x7e00 to reach the loader (a virgin chip reaches it by reset across erased flash). Updates are power-fail-safe except mid-rewrite of the resident slot itself (no reset path leads to the staging copy then). | | 512 words (1 KB) | unprogrammed | *Self-update, app-first*: reset always boots the application, which owns all 31.5 KB and must offer its own jump to 0x7e00 to reach the loader (a virgin chip reaches it by reset across erased flash). Updates are power-fail-safe except mid-rewrite of the resident slot itself (no reset path leads to the staging copy then). |
| 512 words (1 KB) | programmed | *Self-update, loader-first*: reset lands at 0x7c00 — the staging slot, normally erased, so execution walks up into the loader; during an update it is the staging copy itself, so a mid-rewrite power loss recovers by reset. The loss windows move to the staging install/retire page writes instead (page-write scale). The host keeps `[0x7c00, 0x7e00)` clear of application data (`--force` overrides). | | 512 words (1 KB) | programmed | *Self-update, loader-first*: reset lands at 0x7c00 — the staging slot, normally erased, so execution walks up into the loader; during an update it is the staging copy itself, so a mid-rewrite power loss recovers by reset. The loss windows move to the staging install/retire page writes instead (page-write scale). The host keeps `[0x7c00, 0x7e00)` clear of application data (`--force` overrides). |
Applications are flashed unmodified here — word 0 stays the application's own Applications are flashed unmodified — word 0 stays the application's own
reset vector, and the hand-over jumps to 0. reset vector, and the hand-over jumps to 0.
**Patched-vector chips — the tinies and the m48s** (no boot section; the m48s' **Patched-vector chips — the tinies and the m48s** (no boot section; the
SPM runs from the entire flash, Atmel-8271 §26): program the loader at m48s' SPM runs from the entire flash, Atmel-8271 §26): program the loader
`flash 512`; erased flash below it walks up into the loader, so a virgin at `flash 512`; erased flash below it walks up into the loader, so a
chip activates. Flashing an application then takes reset-vector surgery: word virgin chip activates. When flashing an application the host performs
0 becomes an `rjmp` to the loader base, and the application's own entry is reset-vector surgery: word 0 is rewritten to `rjmp` to the loader base, and
re-encoded as a trampoline `rjmp` in the word just below the loader the application's own entry is re-encoded as a trampoline `rjmp` in the
(`base 2`, where the hand-over jumps). Every other vector stays the word just below the loader (`base 2`, where the hand-over jumps). Every
application's. The patched page 0 and the trampoline page are written *first* other vector stays the application's. The patched page 0 and the trampoline
and an erase runs top-down, so from the first write on an interruption still page are written *first*, so from the first write on an interrupted flash
resets into the loader. still resets into the loader; an erase runs top-down for the same reason.
The m48s speak this profile over their hardware USART — no fuse preflight,
A .bin programmed at address 0 by mistake is dead weight on a boot-sectioned BOOTRST does not exist there.
mega (SPM only executes from the boot section — reflash the .hex), but *runs*
on a patched-vector chip, and the ordinary `--update-loader` flow re-homes it
into the top slot from there (`pureboot.rehome`).
## Updating the loader ## Updating the loader
`pureboot.py --update-loader new_pureboot.bin` replaces the resident loader `pureboot.py --update-loader new_pureboot.bin` replaces the resident loader
with any pureboot build — a re-timed window, a newer version — using the with any pureboot build — a re-timed window, a newer protocol — using the
loader itself as its own staging loader. The image is the loader's own 512 loader itself as its own staging loader. The image is the loader's own 512
bytes as a raw binary, or the Intel HEX the build emits beside it. bytes as a raw binary, or the Intel HEX the build emits beside it, which
links the loader at its base inside an otherwise blank flash image:
The preflight refuses an image built for another chip: the info block embedded The preflight refuses an image built for another chip: the info block
in every pureboot binary (signature, page size, loader base, EEPROM size, embedded in every pureboot binary (signature, page size, loader base,
flags) must match the device's own, and the error names both. Die revisions EEPROM size, flags) must match the device's own, and the error names both.
share their base signature and geometry, so their images are interchangeable — Die revisions share their base signature and geometry, so their images are
as the silicon is. interchangeable — as the silicon is. `loader_image()` also accepts a
padded image (a raw .bin padded from 0, or a whole-flash read-back with
the loader resident) and peels it to the slot content by the embedded base.
1. The staging slot `[base512, base)` is saved to a host-side state file (on 1. The staging slot `[baseslot, base)` is saved to a host-side state file
the 1 KB tiny13s that is the whole application, vectors included). (on the 1 KB tiny13s that is the whole application, vectors included).
2. The resident installs the update image there. On the patched-vector chips 2. The resident installs the identical update image there. On the
the host composes the slot's last word as a jump to the resident base, so patched-vector chips the host composes the slot's last word — the same
even an abandoned staging copy times out into a loader. A loader already address as the resident's trampoline — as a jump to the resident base,
sitting whole in the staging slot is left as the staging copy instead — so even an abandoned staging copy times out into a loader, never into
rewriting it would only meet its own running-slot guard. garbage. A loader already sitting whole in the staging slot (its info
3. `J` enters the staging copy, which rewrites the resident slot. Where a block in place, the slot unchanged since the update began) is left as
patched reset vector routes through the resident, the host first re-aims the staging copy instead — rewriting it would only meet its own
word 0 at the staging copy, so a power loss mid-rewrite still resets into a running-slot guard.
loader; on the tiny13s the staging slot carries the reset vector itself. 3. `J` enters the staging copy, which rewrites the resident slot. On the
patched-vector chips whose staging slot sits away from page 0 the host
first re-aims word 0 at the staging copy, so a power loss mid-rewrite
still resets into a loader; on the tiny13s the staging slot carries the
reset vector itself.
4. `J` enters the new resident, which restores the staging slot's saved 4. `J` enters the new resident, which restores the staging slot's saved
content, and the state file is discarded. content (word 0 and the trampoline with it) and the state file is
discarded.
Every phase is idempotent and keyed off the actual flash state, so re-running Every phase is idempotent and keyed off the actual flash state: re-running
the same command after any interruption resumes and completes. The state file the same command after any interruption resumes and completes. The state
carries the only bytes not recoverable from the device; losing it mid-update file carries the only bytes not recoverable from the device; if it is lost
still completes the update, and the staging region comes back by reflashing mid-update the update still completes, and the staging region is restored by
the application. A boot-sectioned mega needs its fuses for the preflight — read reflashing the application. A boot-sectioned mega needs its fuses for the
from the device, or supplied with `--assume-fuses` where reading is impossible preflight (BOOTSZ gate, profile notes) — read from the device, or supplied
(simulators). with `--assume-fuses` where reading is impossible (simulators); the
patched-vector chips need none.
## Host tool ## Host tool
`pureboot.py` — Python 3, standard library only. The port layer is the one `pureboot.py` — Python 3, standard library only. The port layer is the one
platform-specific part: termios drives any tty on POSIX (a USB adapter as well platform-specific part: termios drives any tty on POSIX (a USB adapter as
as a simavr pty), the Win32 serial API through `ctypes` drives a COM port on well as a simavr pty), the Win32 serial API through `ctypes` drives a COM
Windows (`--port COM6`; the `\\.\` form for two-digit ports is supplied by the port on Windows (`--port COM6`; the `\\.\` form for two-digit ports is
tool). Opening the port asserts DTR and RTS on both, so a board that wires DTR supplied by the tool). Opening the port asserts DTR and RTS on both, so a
to reset gets its reset pulse and opens the activation window by itself. board that wires DTR to reset gets its reset pulse and opens the activation
window by itself.
pureboot.py --port /dev/ttyUSB0 --baud 57600 \ pureboot.py --port /dev/ttyUSB0 --baud 57600 \
--info --fuses --flash app.hex --info --fuses --flash app.hex
Operations run in a fixed order within one session: info, fuses, loader Operations run in a fixed order within one session: info, fuses, loader
update, flash (erase / program / read / verify), EEPROM (the same) — then the update, flash (erase / program / read / verify), EEPROM (erase / program /
loader hands over to the application. `--stay` keeps the session alive read / verify) — then the loader hands over to the application; `--stay`
instead, and a later invocation reconnects into it. `--flash` and `--eeprom` keeps the session alive instead, and a later invocation reconnects into it
verify by read-back unless `--no-verify`, and a flash page that reads back (the knock converges there too). `--flash` and `--eeprom` verify by
wrong is rewritten up to three times before the run stops (see `W` above). read-back unless `--no-verify`, and a flash page that reads back wrong is
`--verify-flash` only reports. Images are raw binary, or Intel HEX by rewritten up to three times before the run stops — the loader leaves one
extension. `--force` overrides the refusable safety checks — today, flashing recoverable way for a page to land wrong (see `W` above), and rewriting is
application data into a mega's reset walk region. what clears it. `--verify-flash` only reports. Images are raw binary, or
Intel HEX by extension. `--force` overrides the refusable safety checks (today: flashing
application data into a mega's reset walk region).
Readouts come one fact per line: `--info` decodes the info block field by Readouts come one fact per line: `--info` prints the decoded info block
field, `--fuses` each fuse byte plus, on a boot-sectioned mega, its decoded field by field, `--fuses` each fuse byte on its own line — plus, on a
meaning. Transfers that take wire time draw a transient progress bar on stderr boot-sectioned mega, the decoded meaning (where the BOOTSZ section starts,
when it is a tty. `-v`/`--verbose` adds the decisions as they happen: knock what BOOTRST does to reset). Transfers that take wire time — programming,
counts, the programming plan, update state handling and per-phase page counts. reading, erasing, verifying, the update phases — draw a transient progress
bar on stderr when it is a tty; logs and pipes see only the summary lines.
`-v`/`--verbose` adds the decisions as they happen: knock counts, the
programming plan (vector-surgery targets, skipped blank pages), update
state handling and per-phase page counts.
## Tests ## Tests
`tools/check.sh` runs every chip's workflow (`--full` adds the reflect-mode `tools/check.sh` runs every chip's workflow (`tools/check.sh --full` adds
builds of libavr's spot set; `tools/make_presets.py` regenerates the presets). the reflect-mode builds of libavr's spot set; `tools/make_presets.py`
Per chip preset, `ctest` runs: regenerates the presets). Per chip preset, `ctest` runs:
- `pureboot.size` — the 510-byte (patched-vector) / 512-byte budget; - `pureboot.size` — the 510-byte (tinies) / 512-byte (mega) budget;
- `pureboot_*.size` — the size matrix: the serial backends × the clock ladder - `pureboot_*.size` — the size matrix: the serial backends × the clock
(1/8/16 MHz; the t13s' own RC menu), the USART1 instance across that same ladder (1/8/16 MHz; the t13s' own RC menu), plus the USART1 build on the
ladder on the x4 chips, and `pureboot_sw_wide`, the slowest ladder rate at x4 chips — every configuration axis that could move the image, each
the fastest clock — where a software UART's per-bit spin outgrows its variant against the same slot budget (pins are immediate operands and the
one-register delay loop and takes the 16-bit one. That is the largest image timeout is a constant: size-neutral);
the configuration space produces, and a shape the ladder default (always the - `pureboot.custom` (328P) — the configured-deployment acceptance test: the
*fastest* rate a clock reaches) never picks. Pins are immediate operands and 1 MHz software-serial TX=PB1/RX=PB5 build from the configuration example
the timeout is a constant: neither is an axis; drives the full protocol suite through the runner's GPIO bridge, fixture
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, the application included;
info block within the image's first 256 bytes; - `pureboot.usart1` (644A) — the same protocol suite over the second
- `pureboot.planner` — the host tool's pure logic: programming orders and their hardware USART: instance selection is compile-checked everywhere, but
recovery properties, the surgery, the staging composition, the boot-fuse only a live session proves the loader polls the USART it claims;
decode, the update preflight over synthetic fuse bytes, and the repairing - `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`
verify against a fake device; in the image, the info block within its first 256 bytes;
- `pureboot.planner` — the host tool's pure logic: programming orders and
their recovery properties, the surgery, the staging composition, the
boot-fuse decode, the update preflight's error/warning matrix over
synthetic fuse bytes, and the repairing verify against a fake device — one
bad write repaired in a single rewrite, a page that never comes good
stopping after exactly three;
- `pureboot.protocol` — end to end against a simavr device - `pureboot.protocol` — end to end against a simavr device
(`test/pureboot_device.c`: a hardware USART as a pty, or a cycle-timed (`test/pureboot_device.c` a hardware USART as a pty, or a cycle-timed
GPIO⇄pty bridge for a software-UART build, plus the SPM/NVM module simavr's GPIO⇄pty bridge for a software-UART build, selected with `-l` to match
tiny cores lack) driven by the real host tool through knock-from-reset, the loader's link; plus the SPM/NVM module simavr's tiny cores lack)
program + verify of both memories, session reconnect, an external reset driven by the real host tool through
through the patched vector, and the hand-over to a fixture application whose knock-from-reset, program + verify of both memories, session reconnect, an
banner proves the launch — cross-checked against the simulator's external reset through the patched vector, and the hand-over to a fixture
ground-truth memory dumps and an independent decode of the surgery; application whose banner proves the launch — cross-checked against the
- `pureboot.reloc` — the identical image one slot below the resident serves the simulator's ground-truth memory dumps and an independent decode of the
complete command set from there; surgery's rjmp words;
- `pureboot.rehome` (t85) — a loader programmed at address 0 or in the staging - `pureboot.reloc` — the identical image installed one slot below the
slot re-homes into the top slot through the ordinary update flow; resident serves the complete command set from there (the
- `pureboot.custom` (328P) — the configuration example's 1 MHz software-serial position-independence acceptance test);
build driving the full protocol suite, proving the plumbing produces a - `pureboot.dirty` (328P) — entering the loader from a running application
working loader and not just one that fits; with no reset between, over an SPM page buffer the fixture deliberately
- `pureboot.usart1` (644A) — the same suite over the second hardware USART: dirtied: the case the loader declines to guard against. A bare verify must
instance selection is compile-checked everywhere, but only a live session see the corruption, the repairing verify must fix it in one rewrite, and a
proves the loader polls the USART it claims; plain verify afterwards must pass. On the boot-sectioned megas hardware
- `pureboot.dirty` (328P) — entering the loader from a running application over forbids the state outright (SPM runs only from the boot section, and reset
an SPM buffer it deliberately dirtied, the case the loader declines to guard: erases the buffer), but simavr dispatches SPM from anywhere — which is what
a bare verify must see the corruption and the repairing verify must fix it in makes the path constructible at all;
one rewrite. Hardware forbids the state here, but simavr dispatches SPM from - `pureboot.update` — the full `--update-loader` flow to a re-timed build,
anywhere, which is what makes the path constructible; then every power-fail phase: the device is killed mid-write, restarted
- `pureboot.update` — the full `--update-loader` flow, then every power-fail from its flash dump, and a re-run must complete the update with the
phase: the device is killed mid-write, restarted from its flash dump, and a application intact throughout.
re-run must complete the update with the application intact.
`size`, `pi` and `planner` are host logic and run anywhere; the `size`, `pi`, and `planner` are host logic and run anywhere; the
simulator-driven targets need simavr and a pty, so they are POSIX-only. simulator-driven targets need simavr and a pty, so they are POSIX-only
on Windows the tool is exercised against real hardware.

View File

@@ -1,14 +1,28 @@
// pureboot — a serial bootloader on libavr: one C++ source, no inline // pureboot — a serial bootloader on libavr, pure by constraint: one C++
// assembly, no global register variables, 512 bytes on every chip libavr // source with no inline assembly and no global register variables, built for
// targets. The device speaks primitives; every composite (verify, erase, // every chip libavr targets, 512 bytes on each. The device speaks primitives
// reset-vector surgery, self-update) lives in the host tool. Protocol, // — read/program flash, read/write EEPROM, fuse bytes, an info block, a jump
// deployment and configuration: README.md next to this file. // — and everything composite (verify, erase, reset-vector surgery, updating
// the loader itself) lives in the host tool. Protocol reference: README.md
// next to this file.
// //
// The image is position-independent — PC-relative control flow, wire // The image is position-independent: control flow is PC-relative, the write
// addresses in, the write guard and the info block both anchored on the // and read paths take wire addresses, the write guard refuses the 512-byte
// runtime return address — so the identical binary runs from any slot. That // slot the code is *running* in (taken from the runtime return address), the
// is what makes a copy one slot below able to rewrite the resident one, and // info block is read relative to that same anchor, and the application jump
// every change here has to keep it (test/check_pi.py). // is an indirect call to an absolute entry. The identical binary therefore
// runs from any 512-byte slot with every command intact: flashed one slot
// below the resident loader it becomes the staging loader that rewrites the
// resident — how pureboot updates itself, host-driven, with no other
// firmware involved.
//
// Entry: reset lands in avr::startup::entry below (BOOTRST on the
// boot-sectioned megas; the patched reset vector — or erased flash walking
// up into the loader — on the tinies and the boot-section-less m48s). A
// watchdog reset hands straight to the application. Otherwise the
// host has one activation window per awaited knock byte ("pb"); an idle line
// boots the application. A session then stays in the command loop until 'J'
// jumps away or the chip resets.
#include <libavr/libavr.hpp> #include <libavr/libavr.hpp>
@@ -19,14 +33,17 @@ namespace ee = avr::eeprom;
namespace pureboot { namespace pureboot {
namespace { namespace {
// Purely polled: every interrupt guard folds to nothing. // Purely polled interrupts stay off, every guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused; constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+'; constexpr std::uint8_t ack = '+';
// Deployment parameters come from the build (pureboot_add_loader()). The // Per-deployment personality, passed in by the build pureboot_add_loader()
// signature is not one of them: the chip database is the only universal // (the CMake function next to this file) resolves the defaults: the clock the
// source — a tiny13A cannot read its own signature row from code. // board actually runs, the wire baud, the serial backend and its pins. The
// device signature needs no configuring — it comes from the chip database
// (avr::hw::db.signature), the only universal source, since the tiny13A
// cannot even read its signature row from code.
#if !defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD) #if !defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD)
#error \ #error \
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud — create loader targets with pureboot_add_loader() (README.md)" "PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud — create loader targets with pureboot_add_loader() (README.md)"
@@ -42,51 +59,70 @@ consteval std::int16_t wdrf_field()
return avr::hw::db.field_index(reg, "WDRF"); return avr::hw::db.field_index(reg, "WDRF");
} }
// The loader owns the top 512 bytes; a staging copy goes in the slot below. // Geometry: the resident loader owns the top slot of flash — 512 bytes,
// Chips without a hardware boot section — the tinies and the m48s, whose SPM // except on the >64 KiB chips whose own smallest boot sector is 1 KiB (the
// runs from anywhere (Atmel-8271 §26) — keep the application's relocated // 1284s): there the slot is 1 KiB, matching the hardware boundary the
// reset vector in the word under the slot. // 512-byte figure comes from everywhere else. The word below the slot is
constexpr std::uint16_t slot_bytes = 512; // the trampoline (the application's relocated reset vector) on chips
// without a hardware boot section — the tinies and the m48s, whose SPM
// runs from anywhere (Atmel-8271 §26). A boot section also means the CPU
// runs on while the RWW section programs; everywhere else it halts through
// the operation.
constexpr std::uint16_t slot_bytes = spm::flash_bytes > 65536 ? 1024 : 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes; constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes; constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = avr::hw::curated::has_boot_section(); constexpr bool boot_section = avr::hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash // Past 64 KiB a byte address no longer fits the wire's 16 bits, so on the
// addresses there are word addresses ('J' always was one). A slot is 256 of // large chips every flash address on the wire — and all slot arithmetic —
// those — one value of a wire address's high byte, where 512 bytes span two. // is a word address instead ('J' always was one). A slot spans the same
// wire-high-byte pair in either unit (512 B = 2 x 256 bytes, 1 KiB =
// 2 x 256 words), so the slot index is the high byte with its low bit
// dropped everywhere.
constexpr bool word_flash = spm::flash_bytes > 65536; constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint16_t wire_base = constexpr std::uint16_t wire_base =
word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base); word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base);
constexpr std::uint16_t wire_page_mask = word_flash ? (page / 2 - 1) : (page - 1);
// A compile-time window, so the whole EEPROM belongs to the application; // The activation window, in seconds, is a compile-time constant (the build
// re-timing a deployed loader is a self-update with a re-timed build. // may override it): the whole EEPROM belongs to the application, and
// re-timing the loader is a bootloader self-update with a re-timed binary.
#if !defined(PUREBOOT_TIMEOUT) #if !defined(PUREBOOT_TIMEOUT)
#define PUREBOOT_TIMEOUT 8 #define PUREBOOT_TIMEOUT 8
#endif #endif
constexpr std::uint8_t timeout_seconds = PUREBOOT_TIMEOUT; constexpr std::uint8_t timeout_seconds = PUREBOOT_TIMEOUT;
// The loader's one identity number. The protocol carries none of its own — // The 12-byte info block the host reads with the 'b' command, flash-resident
// a version implies it, and the host tool holds that map (README.md). // through flash_table (there is no crt to copy a .data image, and its storage
constexpr std::uint8_t version = 3; // carries the word alignment 'b' needs to halve the address on the large
// chips). The page byte is the wire count convention: 0 means 256.
// The 'b' reply, byte for byte (layout: README.md). Flash-resident because
// no crt copies a .data image — and flash_table's storage carries the word
// alignment 'b' needs to halve the address on the large chips.
inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{ inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{
'P', 'B', version, avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2], 'P',
static_cast<std::uint8_t>(page), // 0 means 256 'B',
wire_base & 0xff, wire_base >> 8, avr::hw::db.mem.eeprom_size & 0xff, avr::hw::db.mem.eeprom_size >> 8, 1, // magic, protocol version
static_cast<std::uint8_t>((boot_section ? 0 : 1) | (word_flash ? 2 : 0)), // patch-vector, word-addressed avr::hw::db.signature[0],
avr::hw::db.signature[1],
avr::hw::db.signature[2],
static_cast<std::uint8_t>(page),
wire_base & 0xff,
wire_base >> 8, // app flash ends here; resident loader base (a word address on large chips)
avr::hw::db.mem.eeprom_size & 0xff,
avr::hw::db.mem.eeprom_size >> 8,
// bit 0: host must patch the reset vector (no hardware boot section);
// bit 1: flash wire addresses are word addresses
static_cast<std::uint8_t>((boot_section ? 0 : 1) | (word_flash ? 2 : 0)),
}> }>
info_data; info_data;
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL, // The serial link. PUREBOOT_USART forces a hardware USART instance,
// defaulting to the chip's USART0 where it has one. The software receiver is // PUREBOOT_SOFT_SERIAL the polled software UART (no vector — the table
// the polled one: the vector table belongs to the application. Templates on // belongs to the application) on PUREBOOT_RX/PUREBOOT_TX; with neither, the
// the clock, so only the selected backend instantiates. pending() is the // chip's first USART where it has one and the software UART elsewhere. Both
// cheap line test the activation window polls; drain() holds until the last // are class templates on the clock so only the selected backend is ever
// frame is off the wire, so a hand-over cannot let the target's re-init clip // instantiated. pending() is the cheap line test the activation window
// the ack. // polls; rx() then picks the byte up; drain() holds until the last
// transmitted frame is fully on the wire (the jump hand-over must not let
// the target's re-init clip the ack).
#if defined(PUREBOOT_SOFT_SERIAL) && defined(PUREBOOT_USART) #if defined(PUREBOOT_SOFT_SERIAL) && defined(PUREBOOT_USART)
#error "PUREBOOT_SOFT_SERIAL and PUREBOOT_USART select opposing serial backends" #error "PUREBOOT_SOFT_SERIAL and PUREBOOT_USART select opposing serial backends"
#endif #endif
@@ -181,10 +217,12 @@ using link =
std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock>, software_link<dev::clock>>; std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock>, software_link<dev::clock>>;
#endif #endif
// The application's entry, pinned by the linker (--defsym): word 0 on a // The application's entry, an absolute address the linker pins (--defsym in
// boot-sectioned mega, the trampoline at base 2 elsewhere. Reaching it must // CMakeLists.txt): 0x0000 on the mega (word 0 stays the application's own
// not depend on where this copy runs, so the jump goes through a pointer, and // vector — BOOTRST re-vectors a reset into the loader in hardware) and the
// [[gnu::noipa]] keeps the constant from folding back into a relative call. // trampoline word at base - 2 on the tinies. Reaching it must not depend on
// where this copy runs, so the jump goes through a pointer: [[gnu::noipa]]
// keeps the constant from folding back into a PC-relative call.
extern "C" [[noreturn]] void pureboot_app(); extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)()) [[gnu::noipa, noreturn]] void jump(void (*target)())
@@ -198,8 +236,10 @@ extern "C" [[noreturn]] void pureboot_app();
jump(pureboot_app); jump(pureboot_app);
} }
// The window as one 32-bit countdown, divided by the backend's counted // One activation window is a single 32-bit poll countdown. The divisor is
// poll-loop cycles. Whole seconds is all it promises. // the backend's counted poll-loop cycles (its own comment reads them off the
// compiled loop); whole-second precision is all the window promises, so the
// nearest cycle count is plenty.
consteval std::uint32_t window_polls() consteval std::uint32_t window_polls()
{ {
return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles); return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles);
@@ -215,8 +255,8 @@ bool pending_before_deadline()
return false; return false;
} }
// A knock byte under the deadline: an idle window means no host, so the // A knock byte under the activation deadline: an idle line means no host is
// application runs. // there, and the application runs.
std::uint8_t rx_deadline() std::uint8_t rx_deadline()
{ {
if (!pending_before_deadline()) if (!pending_before_deadline())
@@ -224,26 +264,22 @@ std::uint8_t rx_deadline()
return link::rx(); return link::rx();
} }
// Inlined: read across a call, the first byte strands in a call-saved // Inlined into its call sites: reading two bytes across a call otherwise
// register the caller has to push and pop. // strands the first in a call-saved register the caller must push/pop; folded
// into the (noreturn) command loop that cost disappears.
[[gnu::always_inline]] inline std::uint16_t rx16() [[gnu::always_inline]] inline std::uint16_t rx16()
{ {
std::uint16_t low = link::rx(); std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8)); return static_cast<std::uint16_t>(low | (link::rx() << 8));
} }
// The wire's byte pair as the word it is — AVR is little-endian too, so the // The streamers take the count in the wire's 8-bit form: 0 means 256.
// cast is the identity a shift-and-or spelling makes the compiler rediscover. //
// Callers read into named variables first: the wire order is a sequence of // Two functions, because they want opposite placement and placement is an
// reads, not an argument order. // attribute: the byte-addressed loop is small enough to inline into both
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair) // callers, the word-addressed one stays out of line but flattened — a call to
{ // the transmit inside it would strand the 24-bit cursor in callee-saved
return std::bit_cast<std::uint16_t>(pair); // registers. `word_flash` picks at the call site.
}
// Counts arrive in the wire's 8-bit form: 0 means 256. Both streamers fold
// into the one command that reads flash, which is what lets the far one's
// 24-bit cursor sit in the command loop's own call-saved registers.
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count) [[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{ {
do do
@@ -251,16 +287,17 @@ std::uint8_t rx_deadline()
while (--count); while (--count);
} }
// The 24-bit cursor as the machine holds it the RAMPZ byte and a 16-bit Z, // The 24-bit cursor as the machine holds it: the RAMPZ byte and a 16-bit Z,
// carried apart; the reassembled address folds away inside the far load. // carried explicitly (the reassembled 32-bit address folds away inside the
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count) // inlined far load).
[[maybe_unused, gnu::flatten, gnu::noinline]] void send_flash_far(std::uint16_t address, std::uint8_t count)
{ {
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15); std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1); std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do { do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z)); link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
// Carrying the wrap is smaller than the flat 32-bit cursor GCC // The protocol never reads across 64 KiB, but carrying the wrap is
// builds without it. // smaller than the flat 32-bit cursor GCC builds without it.
if (++z == 0) if (++z == 0)
++rampz; ++rampz;
} while (--count); } while (--count);
@@ -274,13 +311,6 @@ std::uint8_t rx_deadline()
send_flash_near(address, count); send_flash_near(address, count);
} }
// Out of line: three sites send it, and a call is shorter than three
// load-immediates.
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count) void send_eeprom(std::uint16_t address, std::uint8_t count)
{ {
do do
@@ -288,62 +318,74 @@ void send_eeprom(std::uint16_t address, std::uint8_t count)
while (--count); while (--count);
} }
// Host-paced: the ack goes out once the write has begun, so the next byte // EEPROM write, host-paced: each ack goes out once the byte's write has
// arrives while it completes and nothing is missed without a buffer. // begun, so the next byte arrives while it completes and the following
// write's own ready-wait sees an idle line. Nothing is ever missed, on
// either serial backend, without a buffer.
void store_eeprom(std::uint16_t address, std::uint8_t count) void store_eeprom(std::uint16_t address, std::uint8_t count)
{ {
do { do {
ee::write<off>(address++, link::rx()); ee::write<off>(address++, link::rx());
tx_ack(); link::tx(ack);
} while (--count); } while (--count);
} }
// One page into the SPM buffer, then erase and program — except the slot // One flash page: stream the bytes into the SPM buffer as little-endian
// this code is running in (`slot_high`, from run()), which is drained and // words, then erase and program — except the 512-byte slot this code runs
// left alone. A broken host therefore cannot brick the running loader, and a // in, which is drained but never programmed, so a copy can never erase
// copy one slot lower may rewrite the resident one. // itself. `slot_high` is the high byte of that running slot's base (run()
// // derives it); a broken host thus cannot brick the running loader, and a
// Nothing discards the buffer first: it is write-once per word (§26.2.1), so // copy flashed one slot lower may rewrite the slot above it — how pureboot
// filling over a refused page or an application's leavings programs stale // updates itself.
// words — but a page write auto-erases it (§26.2.1; §19.2 on the tinies), so
// that write clears the condition and the host's read-back rewrites the page.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high) void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
{ {
// One induction either way: a byte-addressed wire address walks the page // No discard before the fill: the buffer is write-once per word
// itself (aligned, so the offset bits wrap to zero), while a word one // (§26.2.1), so filling over one a refused page or an application left
// becomes a byte cursor once. The slot index is the wire address's high // dirty programs stale words — but a page write auto-erases the buffer
// byte — on byte-addressed chips the byte address's, with the low bit // (§26.2.1; §19.2 on the tinies), so that write clears the condition and
// dropped, since a slot is two of those. // the host's read-back rewrites the page.
// One induction either way. On the byte-addressed chips the wire address
// itself walks the page (aligned, so the offset bits wrap to zero); on
// the word-addressed large chips the wire word address becomes a 32-bit
// byte cursor once, and their 256-byte page makes its low byte the whole
// in-page offset. The slot index is one high byte of the wire address —
// two values on byte-addressed chips (the & ~1), bits 16:9 re-packed on
// the large ones.
spm::flash_address_t address; spm::flash_address_t address;
std::uint8_t page_high; std::uint8_t page_high;
if constexpr (word_flash) { if constexpr (word_flash) {
// A page is aligned, so it never crosses 64 KiB: RAMPZ is a per-page // Pages are aligned, so one page never crosses a 64 KiB boundary:
// constant and the 16-bit Z's low byte is the whole in-page offset. // RAMPZ is a per-page constant and the fill cursor is a 16-bit Z
// whose low byte is the whole in-page offset (256-byte pages). The
// slot index is simply the wire word address's high byte.
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15); const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1); const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1);
std::uint16_t z = z0; std::uint16_t z = z0;
do { do {
std::uint8_t low = link::rx(); std::uint8_t low = link::rx();
std::uint8_t high = link::rx(); std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high})); spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z,
static_cast<std::uint16_t>(low | (high << 8)));
z += 2; z += 2;
} while (static_cast<std::uint8_t>(z)); } while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0; address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8); page_high = static_cast<std::uint8_t>(wire_address >> 8) & 0xfe;
} else { } else {
address = static_cast<spm::flash_address_t>(wire_address); address = static_cast<spm::flash_address_t>(wire_address);
do { do {
std::uint8_t low = link::rx(); std::uint8_t low = link::rx();
std::uint8_t high = link::rx(); std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high})); spm::fill<off>(address, static_cast<std::uint16_t>(low | (high << 8)));
address += 2; address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1)); } while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe; page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
} }
if (page_high != slot_high) { if (page_high != slot_high) {
// Only a boot-sectioned mega runs on while its RWW section programs; // The tinies and the m48s halt the CPU through the erase and the
// everywhere else the CPU halts through erase and write. // write, so only the boot-sectioned megas — running on while their
// RWW section programs — wait.
spm::erase_page<off>(address); spm::erase_page<off>(address);
if constexpr (boot_section) if constexpr (boot_section)
spm::wait(); spm::wait();
@@ -351,92 +393,89 @@ void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
if constexpr (boot_section) if constexpr (boot_section)
spm::wait(); spm::wait();
} }
// Programming leaves the RWW section disabled; reads need it back on. The // The megas program with their RWW section disabled; reads need it back
// same store discards the buffer (§26.2.2), so a boot-sectioned mega never // on. The same store discards the buffer (§26.2.2), so they never meet
// meets the stale-word case above. // the stale-word case above. boot_section implies an RWW section.
if constexpr (boot_section) if constexpr (boot_section)
spm::rww_enable<off>(); spm::rww_enable<off>();
} }
// The four fuse and lock bytes in the hardware's own Z order: low, lock, // The four fuse/lock bytes in the hardware's own Z order: low, lock,
// extended, high. // extended, high. Writing fuses is not a thing self-programming can do on
// AVR — SPM reaches flash (and boot lock bits) only.
void send_fuses() void send_fuses()
{ {
std::uint8_t which = 0; std::uint8_t which = 0;
do do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which))); link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4); while (++which & 3);
} }
[[noreturn]] void run() [[noreturn]] void run()
{ {
// A watchdog reset belongs to the application, whose watchdog stays forced // A watchdog reset belongs to the application (whose watchdog stays
// on until it clears WDRF — no activation window in its way. // forced on until it clears WDRF) — no activation window in its way.
// The flag register is MCUSR, or the classic megas' MCUCSR.
if (avr::hw::field_impl<wdrf_field()>::test()) if (avr::hw::field_impl<wdrf_field()>::test())
run_app(); run_app();
link::init(); link::init();
// The high byte of the slot this copy runs at, which the write guard and // The high byte of the 512-byte-aligned base this copy runs at: the
// the info block both follow: the return address is a word address, so its // return address is a word address, whose high byte is the 256-word slot
// high byte is the 256-word slot index, doubled back into byte terms where // index — on byte-addressed chips doubled back into byte terms.
// the wire counts bytes. Taken as byteswap's low byte — the builtin already // program_flash refuses this one slot and the info block is addressed
// swaps the two stacked bytes, and the double swap folds away, where `>> 8` // from it, so both follow wherever the code was flashed. The high byte is
// would leave the swap materialized. // spelled as byteswap's low byte: the builtin's value is itself built by
// swapping the two stacked bytes, and the double swap folds to the single
// byte pick a hand assembler writes — `>> 8` leaves the swap materialized.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0)); const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words)); const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1); const std::uint8_t slot_high = word_flash ? ra_high & 0xfe : static_cast<std::uint8_t>(ra_high << 1);
// 'p' then 'b', each under a fresh window; anything else is line noise. // The knock: 'p' then 'b', each under a fresh window; any other byte is
// line noise and waits again. Falling out of a window runs the app.
while (rx_deadline() != 'p' || rx_deadline() != 'b') { while (rx_deadline() != 'p' || rx_deadline() != 'b') {
} }
for (;;) { for (;;) {
// No prompt while an EEPROM write runs: it blocks SPM and fuse reads // No prompt while an EEPROM write runs: a pending write blocks SPM
// (§26.2.1), and the prompt is the previous command's completion ack. // and fuse reads (§26.2.1), and the ack tells the host all is done.
ee::wait(); ee::wait();
tx_ack(); link::tx(ack);
const std::uint8_t command = link::rx(); const std::uint8_t command = link::rx();
switch (command) { switch (command) {
case 'b': { // info block, read relative to the running slot
// The block sits in the image's first 256 bytes (the build lint
// asserts it), and slots are 512-aligned — so the low byte of its
// link address (in wire units: bytes, or words on the large
// chips) is its offset in any slot, and the high byte of its
// runtime address is the running slot's. Composed from the two
// bytes — the high half is runtime data, so no absolute address
// is ever materialized.
const auto link_low = reinterpret_cast<std::uint16_t>(info_data.storage.data());
const std::uint8_t low =
word_flash ? static_cast<std::uint8_t>(link_low >> 1) : static_cast<std::uint8_t>(link_low);
send_flash(static_cast<std::uint16_t>(low | (slot_high << 8)), static_cast<std::uint8_t>(info_data.size()));
break;
}
case 'J': { // jump to a wire word address: hand-over and staging transfer case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16()); auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack(); link::tx(ack);
link::drain(); link::drain();
jump(target); jump(target);
} }
case 'b': // info block, read relative to the running slot
case 'R': // read flash: addr16, n8 (0 = 256) case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8 case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for all four: 'b' is a flash read std::uint16_t address = rx16();
// whose arguments the loader already knows, so it joins the std::uint8_t count = link::rx();
// wire-argument three rather than streaming from a call site of its if (command == 'R')
// own. That leaves one flash streamer in the image, and lets its
// cursor live in this never-returning loop's own call-saved
// registers instead of being saved and restored around a call.
std::uint16_t address;
std::uint8_t count;
if (command == 'b') {
// The block sits in the image's first 256 bytes (check_pi.py
// asserts it) and slots are 512-aligned, so the low byte of its
// link address is its offset in any slot — halved where wire
// units are words. The high byte is runtime data, so no
// absolute address is ever materialized.
const auto link_byte =
static_cast<std::uint8_t>(reinterpret_cast<std::uint16_t>(info_data.storage.data()));
const std::uint8_t low = word_flash ? static_cast<std::uint8_t>(link_byte >> 1) : link_byte;
address = static_cast<std::uint16_t>(low | (slot_high << 8));
count = static_cast<std::uint8_t>(info_data.size());
} else {
address = rx16();
count = link::rx();
}
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count); send_flash(address, count);
else if (command == 'r')
send_eeprom(address, count);
else
store_eeprom(address, count);
break; break;
} }
case 'W': // program one flash page: addr16, page bytes case 'W': // program one flash page: addr16, page bytes

View File

@@ -1,13 +1,24 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""pureboot host tool — the smart half of the protocol (README.md). """pureboot host tool — the smart half of the pureboot protocol (README.md).
The device exposes primitives; everything composite is here: HEX/raw images, The device exposes primitives; this tool composes them: image loading (raw
programming with repairing read-back verification, the reset-vector surgery binary or Intel HEX), flash programming with read-back verification, erase as
the boot-section-less chips need, and the self-update that stages the loader writing 0xff, EEPROM programming, fuse and info readout, the hand-over jump,
one slot lower and lets it rewrite the resident. and — on chips without a hardware boot section — the reset-vector surgery
that re-homes the application's entry through the trampoline word below the
loader. Page 0 and the trampoline are written first, so every interruption
point of a flash leaves the chip reset-recoverable into the loader.
Standard library only. The port is termios on POSIX and the Win32 serial API It also updates the loader itself (--update-loader): pureboot's image is
through ctypes on Windows, so any tty or COM port works. position-independent, so the tool installs the identical binary one 512-byte
slot below the resident loader, jumps into that staging copy, lets it rewrite
the resident slot, and restores what the staging slot held — resumable at
every phase from the flash state plus a host-side state file carrying the
saved bytes.
Python standard library only; the serial port is driven with termios on POSIX
and the Win32 serial API (through ctypes) on Windows, so any tty or COM port
works — a USB adapter as well as a simavr pty.
""" """
import argparse import argparse
@@ -24,20 +35,16 @@ else:
import termios import termios
PROMPT = b"+" PROMPT = b"+"
VERSION = 2 # this tool's own version — free to drift from a loader's PROTOCOL_VERSION = 1
# The loader versions this tool speaks. A pureboot version implies its wire SLOT = 512 # the loader slot on byte-addressed chips; word-addressed ones (>64 KiB) use 1 KiB — their own smallest boot sector
# protocol, which carries no number of its own, so this window is where that
# map lives: every version so far speaks the same protocol, and one that
# changes it becomes the new floor here.
OLDEST_LOADER = 1
NEWEST_LOADER = 3
SLOT = 512 # the loader slot, on every chip
RETRIES = 3 # rewrites of a page that reads back wrong, before the run stops RETRIES = 3 # rewrites of a page that reads back wrong, before the run stops
VERBOSE = False VERBOSE = False
def verbose(message): def verbose(message):
"""Detail printed only under --verbose: decisions and derived facts, not
per-byte chatter — the progress bar carries the bulk transfers."""
if VERBOSE: if VERBOSE:
print(f" {message}") print(f" {message}")
@@ -47,9 +54,11 @@ class Error(Exception):
class Progress: class Progress:
"""A transient bar on stderr, drawn only for a tty and erased when done — """A transient in-place bar on stderr for the operations that take wire
logs and pipes see only the summary line each operation prints. No label time. Drawn only when stderr is a tty — logs, pipes and the test harness
or a zero total disables it, so callers can pass one unconditionally.""" see nothing — and erased once done; the summary line each operation
prints afterwards is the persistent record. A zero total (or no label)
disables it, so callers can pass one through unconditionally."""
def __init__(self, label, total, unit="pages"): def __init__(self, label, total, unit="pages"):
self.label, self.total, self.unit = label, total, unit self.label, self.total, self.unit = label, total, unit
@@ -304,26 +313,24 @@ class Info:
def __init__(self, raw): def __init__(self, raw):
if len(raw) != 12 or raw[0:2] != b"PB": if len(raw) != 12 or raw[0:2] != b"PB":
raise Error(f"bad info block: {raw.hex()}") raise Error(f"bad info block: {raw.hex()}")
self.version = raw[2] if raw[2] != PROTOCOL_VERSION:
if not OLDEST_LOADER <= self.version <= NEWEST_LOADER: raise Error(f"protocol version {raw[2]}, tool speaks {PROTOCOL_VERSION}")
raise Error(
f"pureboot {self.version}: this tool (version {VERSION}) speaks pureboot "
f"{OLDEST_LOADER}..{NEWEST_LOADER} — a newer loader needs a newer tool"
)
self.raw = bytes(raw) self.raw = bytes(raw)
self.signature = raw[3:6] self.signature = raw[3:6]
self.page = raw[6] or 256 # the wire count convention: 0 means 256 self.page = raw[6] or 256 # the wire count convention: 0 means 256
self.patch_vector = bool(raw[11] & 1) self.patch_vector = bool(raw[11] & 1)
# Bit 1: flash addresses are words on the wire. Every address here # Large chips speak word addresses for flash (bit 1); the host keeps
# stays a byte address and converts at the wire. # every address in bytes and converts at the wire.
self.word_flash = bool(raw[11] & 2) self.word_flash = bool(raw[11] & 2)
scale = 2 if self.word_flash else 1 scale = 2 if self.word_flash else 1
self.base = (raw[7] | (raw[8] << 8)) * scale self.base = (raw[7] | (raw[8] << 8)) * scale
self.eeprom_size = raw[9] | (raw[10] << 8) self.eeprom_size = raw[9] | (raw[10] << 8)
self.flash_size = self.base + SLOT self.slot = 1024 if self.word_flash else SLOT
self.stage = self.base - SLOT # where a staging copy of the loader goes self.flash_size = self.base + self.slot
# The hand-over target as 'J' takes it: the trampoline below the self.stage = self.base - self.slot # where a staging copy of the loader goes
# loader, or word 0 where BOOTRST re-vectors reset in hardware. # The hand-over target, as the word address 'J' takes: the trampoline
# below the loader (tinies), or word 0 (mega — the application's own
# reset vector; BOOTRST re-vectors a reset into the loader instead).
self.app_entry_word = (self.base - 2) // 2 if self.patch_vector else 0 self.app_entry_word = (self.base - 2) // 2 if self.patch_vector else 0
def describe(self): def describe(self):
@@ -336,18 +343,17 @@ class Info:
) )
def lines(self): def lines(self):
"""One fact per line — what --info prints.""" """The info block as one fact per line — what --info prints."""
if self.patch_vector: if self.patch_vector:
hand_over = f"host-patched reset vector, trampoline at {self.base - 2:#06x}" hand_over = f"host-patched reset vector, trampoline at {self.base - 2:#06x}"
else: else:
hand_over = "hardware boot section, jump to word 0" hand_over = "hardware boot section, jump to word 0"
return ( return (
f"version pureboot {self.version}",
f"signature {' '.join(f'{b:02x}' for b in self.signature)}", f"signature {' '.join(f'{b:02x}' for b in self.signature)}",
f"flash {self.flash_size} B, {self.page} B pages" f"flash {self.flash_size} B, {self.page} B pages"
+ (", word-addressed wire" if self.word_flash else ""), + (", word-addressed wire" if self.word_flash else ""),
f"application 0x0000..{self.base - 1:#06x} ({self.base} B)", f"application 0x0000..{self.base - 1:#06x} ({self.base} B)",
f"loader {self.base:#06x} ({SLOT} B slot)", f"loader {self.base:#06x} ({self.slot} B slot)",
f"staging {self.stage:#06x}", f"staging {self.stage:#06x}",
f"EEPROM {self.eeprom_size} B", f"EEPROM {self.eeprom_size} B",
f"hand-over {hand_over}", f"hand-over {hand_over}",
@@ -355,18 +361,19 @@ class Info:
class Loader: class Loader:
"""A session. Between commands the loader has prompted and awaits a """A pureboot session. Between commands the loader has prompted `+` and
command byte; every method restores that, except jump() — after which the awaits a command byte; every method restores that invariant — except
target must be knocked afresh.""" jump(), after which the target must be knocked afresh."""
def __init__(self, port): def __init__(self, port):
self.port = port self.port = port
self.info = None self.info = None
def connect(self, wait): def connect(self, wait):
"""Knock until the window answers, then read the info block. Also """Knock until the activation window answers, then read the info
converges into a live session: the knock bytes are ignored there and block. Also converges when the loader already sits in its command
the drain absorbs whatever they produced.""" loop: the knock bytes are ignored-or-executed there, and the drain
absorbs whatever they produced."""
self.port.flush_input() self.port.flush_input()
deadline = time.monotonic() + wait deadline = time.monotonic() + wait
knocks = 0 knocks = 0
@@ -452,13 +459,16 @@ class Loader:
return self._command(b"F", 4, 2.0) return self._command(b"F", 4, 2.0)
def jump(self, word_address): def jump(self, word_address):
"""The device acks, then execution continues at the word address.""" """'J': the device acks, then execution continues at the word
address — a loader slot's base (whose copy must then be knocked
afresh) or the application entry."""
self.port.write(bytes((ord("J"), word_address & 0xFF, word_address >> 8))) self.port.write(bytes((ord("J"), word_address & 0xFF, word_address >> 8)))
self._expect_prompt() self._expect_prompt()
def enter_copy(self, byte_address, wait): def enter_copy(self, byte_address, wait):
"""Jump into the loader copy at `byte_address` and knock it — a slot """Jump into the loader copy at `byte_address` and knock it. Ending
base is that copy's entry stub, so it can only land there.""" up in the copy addressed is guaranteed by construction: a jump to a
slot base lands in that slot's entry stub."""
self.jump(byte_address // 2) self.jump(byte_address // 2)
return self.connect(wait) return self.connect(wait)
@@ -554,14 +564,15 @@ def plan_flash(image, info):
def covered(pages, info, skip_blank): def covered(pages, info, skip_blank):
"""Pages in programming order, optionally dropping all-0xff ones (sound """Pages in programming order; optionally dropping all-0xff pages (sound
only over erased flash, and never a load-bearing page). only over erased flash) — never a load-bearing one.
A patched vector puts page 0 first and the trampoline page second, so from With a patched vector (tinies), the patched page 0 goes first and the
the first write on a reset lands in the loader and its fall-through on the trampoline page second: from the first write on, a reset lands in the
application entry — every interruption point recoverable. A hardware boot loader and the loader's own fall-through lands on the application entry,
section re-vectors reset regardless; page 0 goes last there, which so every interruption point of the flash is recoverable. With a hardware
maximizes what an interrupted image retains.""" boot section a reset re-vectors to the loader regardless; ascending
order, page 0 last, maximizes what an interrupted image retains."""
trampoline_page = info.base - info.page if info.patch_vector else None trampoline_page = info.base - info.page if info.patch_vector else None
first = [0, trampoline_page] if info.patch_vector else [] first = [0, trampoline_page] if info.patch_vector else []
rest = [a for a in sorted(pages) if a not in first] rest = [a for a in sorted(pages) if a not in first]
@@ -627,21 +638,18 @@ def mega_boot(info, fuse_bytes):
def image_info(image): def image_info(image):
"""The info block embedded in a pureboot binary, or None. Searched once """The info block embedded in a pureboot binary, or None."""
per known version, so the magic stays three selective bytes rather than at = image.find(b"PB" + bytes((PROTOCOL_VERSION,)))
two that code could carry by chance.""" return Info(image[at : at + 12]) if 0 <= at <= len(image) - 12 else None
for version in range(OLDEST_LOADER, NEWEST_LOADER + 1):
at = image.find(b"PB" + bytes((version,)))
if 0 <= at <= len(image) - 12:
return Info(image[at : at + 12])
return None
def loader_image(path): def loader_image(path):
"""An update image as the slot's own content: a raw binary already is, """A loader update image, as the slot's own content. A raw binary is that
while a HEX carries the blank below the loader's base, which is peeled off already; an Intel HEX links the loader at its base inside an otherwise
here. The base comes from the image's own block, not the device's, so a blank flash image, and load_image() anchors every image at zero, so the
foreign image survives intact for the preflight to reject by name.""" blank below the base is dropped here. The base comes from the image's own
info block rather than the device's, so an image built for somewhere else
survives intact and the preflight can say so."""
image = load_image(path) image = load_image(path)
embedded = image_info(image) embedded = image_info(image)
if embedded and len(image) > embedded.base: if embedded and len(image) > embedded.base:
@@ -650,17 +658,18 @@ def loader_image(path):
def staging_content(image, info): def staging_content(image, info):
"""The staging slot's content: the image, padding, and — where the """The 512-byte staging-slot content: the image, padding, and — on
hand-over jumps through the word below the resident — that word, which for chips whose hand-over jumps through the word below the resident loader —
a staging copy is its own last one. Composed as an rjmp to the resident, that word, which for a staging copy is the slot's own last word: an rjmp
so an abandoned staging copy still falls through into a loader.""" to the resident base. The staging copy's fall-through and 'J'-free exit
budget = SLOT - 2 if info.patch_vector else SLOT both land in a loader instead of garbage."""
if len(image) > budget: slot = info.slot
raise Error(f"loader image is {len(image)} B, the slot holds {budget}") if len(image) > (slot - 2 if info.patch_vector else slot):
content = bytearray(image) + bytearray([0xFF] * (SLOT - len(image))) raise Error(f"loader image is {len(image)} B, the slot holds {slot - 2 if info.patch_vector else slot}")
content = bytearray(image) + bytearray([0xFF] * (slot - len(image)))
if info.patch_vector: if info.patch_vector:
through = rjmp_to((info.base - 2) // 2, info.base // 2, info.flash_size // 2) through = rjmp_to((info.base - 2) // 2, info.base // 2, info.flash_size // 2)
content[SLOT - 2], content[SLOT - 1] = through & 0xFF, through >> 8 content[slot - 2], content[slot - 1] = through & 0xFF, through >> 8
return bytes(content) return bytes(content)
@@ -668,10 +677,7 @@ def update_preflight(image, info, fuse_bytes):
"""Errors and warnings before any flash is touched. Returns warnings.""" """Errors and warnings before any flash is touched. Returns warnings."""
embedded = image_info(image) embedded = image_info(image)
if embedded is None: if embedded is None:
raise Error( raise Error("no pureboot info block in the update image — not a pureboot binary?")
"no pureboot info block in the update image — not a pureboot binary, "
f"or a version this tool ({VERSION}) does not know"
)
if embedded.raw[3:] != info.raw[3:]: if embedded.raw[3:] != info.raw[3:]:
raise Error( raise Error(
f"update image is for another target: it declares " f"update image is for another target: it declares "
@@ -686,7 +692,7 @@ def update_preflight(image, info, fuse_bytes):
raise Error( raise Error(
f"cannot self-update: the staging slot {info.stage:#06x} lies below the " f"cannot self-update: the staging slot {info.stage:#06x} lies below the "
f"boot section ({bls_start:#06x}) where SPM is disabled " f"boot section ({bls_start:#06x}) where SPM is disabled "
f"— a boot section of at least two slots ({2 * SLOT} B, BOOTSZ) is " f"— a boot section of at least two slots ({2 * info.slot} B, BOOTSZ) is "
f"required, and only an external programmer can change fuses" f"required, and only an external programmer can change fuses"
) )
if not bootrst: if not bootrst:
@@ -728,7 +734,7 @@ class UpdateState:
self.data = { self.data = {
"signature": info.signature.hex(), "signature": info.signature.hex(),
"base": info.base, "base": info.base,
"staging": loader.read_flash(info.stage, SLOT).hex(), "staging": loader.read_flash(info.stage, info.slot).hex(),
"page0": loader.read_flash(0, info.page).hex() if info.patch_vector else "", "page0": loader.read_flash(0, info.page).hex() if info.patch_vector else "",
} }
with open(self.path, "w") as f: with open(self.path, "w") as f:
@@ -747,8 +753,9 @@ class UpdateState:
def write_differing(loader, base, content, order=None, label=None): def write_differing(loader, base, content, order=None, label=None):
"""Program the pages of `content` at `base` that differ from flash, so a """Program the pages of `content` at `base` that differ from flash
resumed phase redoes only what an interruption left.""" idempotent, so a resumed phase redoes only what an interruption left.
A label puts the compare-and-program loop on the progress bar."""
page = loader.info.page page = loader.info.page
offsets = list(order) if order is not None else list(range(0, len(content), page)) offsets = list(order) if order is not None else list(range(0, len(content), page))
written = 0 written = 0
@@ -761,8 +768,9 @@ def write_differing(loader, base, content, order=None, label=None):
bar.step() bar.step()
if label: if label:
verbose(f"{label}: {written} of {len(offsets)} pages differed") verbose(f"{label}: {written} of {len(offsets)} pages differed")
# The same bounded repair as verify_pages: here a page left wrong is a # Page-wise read-back with the same bounded repair as verify_pages: this
# half-written loader slot. # is the loader-update path, where a page left wrong is a half-written
# loader slot.
for retry in range(RETRIES + 1): for retry in range(RETRIES + 1):
bad = [ bad = [
offset offset
@@ -784,8 +792,8 @@ def write_differing(loader, base, content, order=None, label=None):
def patch_word0(loader, page0, target_base): def patch_word0(loader, page0, target_base):
"""Re-aim word 0 at `target_base` — the resume insurance around """Rewrite page 0 with its word 0 re-aimed at `target_base` — the
rewriting a loader slot the reset path goes through.""" resume insurance around rewriting a loader slot the reset path uses."""
info = loader.info info = loader.info
patched = bytearray(page0) patched = bytearray(page0)
word = rjmp_to(0, target_base // 2, info.flash_size // 2) word = rjmp_to(0, target_base // 2, info.flash_size // 2)
@@ -795,17 +803,16 @@ def patch_word0(loader, page0, target_base):
def op_update_loader(loader, wait, path, state_path, fuse_bytes): def op_update_loader(loader, wait, path, state_path, fuse_bytes):
"""Replace the resident loader with `path`, using the loader as its own """Replace the resident loader with `path`, using the loader itself as
staging loader. Every phase is idempotent and keyed off the flash state, its own staging loader. Every phase is idempotent and keyed off the
so a re-run resumes; the state file carries what the staging slot held.""" actual flash state, so a re-run after any interruption resumes; the
state file carries the bytes the staging slot held."""
info = loader.info info = loader.info
image = loader_image(path) image = loader_image(path)
for warning in update_preflight(image, info, fuse_bytes): for warning in update_preflight(image, info, fuse_bytes):
print(f"note: {warning}") print(f"note: {warning}")
update = image_info(image) # the preflight proved it is there
verbose(f"installing pureboot {update.version} over pureboot {info.version}")
staged = staging_content(image, info) staged = staging_content(image, info)
resident = bytes(image) + bytes([0xFF] * (SLOT - len(image))) resident = bytes(image) + bytes([0xFF] * (info.slot - len(image)))
page = info.page page = info.page
state = UpdateState(state_path) state = UpdateState(state_path)
@@ -815,29 +822,38 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes):
verbose(f"saving the staging slot to {state_path}") verbose(f"saving the staging slot to {state_path}")
state.load_or_save(loader) state.load_or_save(loader)
# A loader already sitting whole in the staging slot IS the staging copy: # Install the staging copy — unless a loader already sits whole in the
# rewriting it would only meet its own running-slot guard. Any pureboot # staging slot (a build programmed there by hand): that copy IS the
# with the device's info block serves, since a staged copy only streams # installed staging copy, and rewriting it would only trip its own
# pages. "Whole" needs both checks — the block where every image carries # running-slot guard on the composed through-word. Any pureboot with
# it and matching byte for byte, and the slot unchanged since this update # the device's own info block serves — the staged copy just streams
# began, so a half-written install takes the path below instead. # pages, so an older build installs a newer resident all the same. Two
current = loader.read_flash(info.stage, SLOT) # checks make "already a loader" mean a *complete* one: the block must
# sit where every image carries it (within the slot's first 256 bytes
# — the build's position lint), matching the device's block byte for
# byte, and the slot must be unchanged since this update began (the
# state file's snapshot) — a resumed, half-written install differs
# from its snapshot and takes the install path below, which completes
# it page by page.
current = loader.read_flash(info.stage, info.slot)
staged_loader = image_info(current[:268]) staged_loader = image_info(current[:268])
if staged_loader is not None and staged_loader.raw == info.raw and current == state.staging: if staged_loader is not None and staged_loader.raw == info.raw and current == state.staging:
print("staging slot already holds a loader — left in place") print("staging slot already holds a loader — left in place")
else: else:
# Where the staging slot starts at address 0 (the 1 KB tiny13s) its # On a chip whose staging slot starts at address 0 (the 1 KB
# first page carries the reset vector, so it goes last: until then a # tiny13s), its first page carries the reset vector: written last,
# reset still reaches the old resident. # so any earlier interruption still resets into the old resident,
order = list(range(0, SLOT, page)) # and from then on resets enter the staging copy.
order = list(range(0, info.slot, page))
if info.stage == 0: if info.stage == 0:
order = order[1:] + [0] order = order[1:] + [0]
if write_differing(loader, info.stage, staged, order, label="staging copy"): if write_differing(loader, info.stage, staged, order, label="staging copy"):
print(f"staging copy installed at {info.stage:#06x}") print(f"staging copy installed at {info.stage:#06x}")
# Enter it and let it rewrite the resident. Where a patched reset vector # Enter it and let it rewrite the resident slot. Where a patched reset
# routes through the resident, word 0 is re-aimed at the staging copy for # vector routes through the resident (a tiny with the staging slot away
# the rewrite, so a power loss mid-rewrite still resets into a loader. # from page 0), word 0 is re-aimed at the staging copy around the
# rewrite, so a power failure mid-rewrite still resets into a loader.
verbose(f"entering the staging copy at {info.stage:#06x}") verbose(f"entering the staging copy at {info.stage:#06x}")
loader.enter_copy(info.stage, wait) loader.enter_copy(info.stage, wait)
redirect = info.patch_vector and info.stage != 0 redirect = info.patch_vector and info.stage != 0
@@ -855,19 +871,20 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes):
if redirect: if redirect:
verbose("word 0 restored") verbose("word 0 restored")
write_differing(loader, 0, state.page0) write_differing(loader, 0, state.page0)
order = list(range(0, SLOT, page)) order = list(range(0, info.slot, page))
if info.stage == 0: if info.stage == 0:
order = [0] + order[1:] order = [0] + order[1:]
write_differing(loader, info.stage, state.staging, order, label="staging restore") write_differing(loader, info.stage, state.staging, order, label="staging restore")
state.discard() state.discard()
print(f"loader updated: pureboot {update.version}, {len(image)} B at {info.base:#06x}, staging region restored") print(f"loader updated: {len(image)} B at {info.base:#06x}, staging region restored")
def check_walk_region(pages, info, fuse_bytes, force): def check_walk_region(pages, info, fuse_bytes, force):
"""BOOTRST programmed below the loader means reset reaches it only by """With BOOTRST programmed but targeting below the loader, reset reaches
walking across erased flash; application data in that span would divert the loader only by walking across erased flash from the boot-section
reset into itself. Needs the fuses (--fuses or --assume-fuses).""" start; application data in that span would divert reset into itself.
Only checkable when the fuses are known (--fuses or --assume-fuses)."""
if info.patch_vector or fuse_bytes is None: if info.patch_vector or fuse_bytes is None:
return return
bootrst, bls_start = mega_boot(info, fuse_bytes) bootrst, bls_start = mega_boot(info, fuse_bytes)
@@ -886,9 +903,10 @@ def check_walk_region(pages, info, fuse_bytes, force):
def op_erase_flash(loader): def op_erase_flash(loader):
"""0xff over the application area, descending where the reset vector is """0xff over the whole application area. Descending on a patched-vector
patched: page 0 goes last, so an interrupted erase still resets into the chip: page 0 — the patched reset vector — goes last, so an interrupted
loader and once it is gone, the erased walk reaches it anyway.""" erase still resets into the loader, and once it is gone the whole area
is erased and the reset walk reaches the loader anyway."""
blank = bytes([0xFF] * loader.info.page) blank = bytes([0xFF] * loader.info.page)
addresses = range(0, loader.info.base, loader.info.page) addresses = range(0, loader.info.base, loader.info.page)
with Progress("erase", len(addresses)) as bar: with Progress("erase", len(addresses)) as bar:
@@ -924,10 +942,11 @@ def op_flash(loader, path, erase, verify, fuse_bytes=None, force=False):
def verify_pages(loader, pages, repair=False): def verify_pages(loader, pages, repair=False):
"""Read every page back and compare. With `repair`, a mismatch is """Read every page back and compare. With `repair`, a mismatched page is
rewritten and re-read up to RETRIES times first: a page filled over a rewritten and re-read, up to RETRIES times before it is raised: a page
dirty SPM buffer takes stale words, and the write that took them cleared filled over a dirty SPM buffer takes stale words, and the write that took
the buffer, so one rewrite settles it. Anything still wrong is not that.""" them cleared the buffer, so one rewrite settles it. Anything still wrong
after three is not that, and stops the run."""
repaired = 0 repaired = 0
with Progress("verify", len(pages)) as bar: with Progress("verify", len(pages)) as bar:
for address in sorted(pages): for address in sorted(pages):
@@ -1033,8 +1052,6 @@ def main():
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
description="pureboot host tool", epilog="operations run in the order listed above" description="pureboot host tool", epilog="operations run in the order listed above"
) )
parser.add_argument("--version", action="version", version=f"%(prog)s {VERSION} "
f"(speaks pureboot {OLDEST_LOADER}..{NEWEST_LOADER})")
parser.add_argument("--port", required=True, help="serial device: COM6, /dev/ttyUSB0, or a simavr pty") parser.add_argument("--port", required=True, help="serial device: COM6, /dev/ttyUSB0, or a simavr pty")
parser.add_argument("--baud", type=int, default=115200, help="115200 mega, 57600 tinies") parser.add_argument("--baud", type=int, default=115200, help="115200 mega, 57600 tinies")
parser.add_argument("--wait", type=float, default=30.0, help="seconds to keep knocking") parser.add_argument("--wait", type=float, default=30.0, help="seconds to keep knocking")

View File

@@ -1,11 +1,17 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Position-independence lint: the two link-time facts that let the identical """Position-independence lint for the pureboot image.
image run from any slot, asserted from the built ELF.
1. No absolute jmp/call — -mrelax normally guarantees it, but a branch that The self-staging design lets the identical binary run from any 512-byte
grows out of relaxation range would break it silently. slot, which holds only if nothing in the image addresses itself absolutely.
2. The info block within the image's first 256 bytes: 'b' rebuilds its Two link-time facts guarantee it, both asserted here from the built ELF:
address as (running slot high byte : link address low byte).
1. No absolute jmp/call opcodes — all control flow is PC-relative
(rjmp/rcall/ijmp/icall). -mrelax normally guarantees this; a code
change that grows a branch out of relaxation range would break it
silently.
2. The info block sits within the image's first 256 bytes: the 'b'
command rebuilds its address as (running slot high byte : low byte of
the link address), which needs the offset to fit that low byte.
Usage: check_pi.py <objdump> <nm> <elf> <text_start_hex> Usage: check_pi.py <objdump> <nm> <elf> <text_start_hex>
""" """

View File

@@ -66,10 +66,9 @@ struct link {
} }
[[noreturn]] static void idle() [[noreturn]] static void idle()
{ {
// 'L' hands back to the loader in the top slot — 512 bytes on every // 'L' hands back to the loader at the top slot — 512 bytes, or the
// chip. The jump takes a word address, which is what makes the // 1 KiB the >64 KiB chips use.
// >64 KiB chips' entry reachable through a 16-bit pointer at all. constexpr std::uint32_t slot = avr::hw::db.mem.flash_size > 65536 ? 1024 : 512;
constexpr std::uint32_t slot = 512;
for (;;) { for (;;) {
auto command = tx_t::read_blocking(); auto command = tx_t::read_blocking();
if (command == 'L') if (command == 'L')

View File

@@ -1,13 +1,15 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Dirty-page-buffer acceptance test: with no discard in the loader, a page """Dirty-page-buffer acceptance test: the loader carries no buffer discard,
filled over words an earlier writer left takes those instead. The whole so a page filled over words an earlier writer left behind programs those
contract is asserted — a bare verify sees the corruption, the repairing instead. This asserts the whole contract — the corruption is real and a bare
verify fixes it in one rewrite, and it stays fixed. verify sees it, the repairing verify fixes it in one rewrite (the write that
took the stale words auto-erased the buffer), and it stays fixed.
The state is reached the one way the loader cannot prevent: an application The state is reached the way the loader cannot prevent: an application
dirties the buffer and jumps in with no reset between. Boot-sectioned megas dirties the buffer and jumps in with no reset between. Real boot-sectioned
forbid that outright (SPM runs only from the boot section, Atmel-8271 §26.2), megas forbid that outright SPM executes only from the boot section
but simavr dispatches SPM from anywhere, which is what makes it constructible. (Atmel-8271 §26.2) — but simavr dispatches SPM from anywhere, which is what
makes the path constructible at all.
Usage: pbdirty.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page> Usage: pbdirty.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <app_bin> <tool_py> <workdir> <baud> <app_bin> <tool_py> <workdir>

View File

@@ -1,13 +1,19 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Re-homing acceptance test: an image programmed somewhere other than its """Re-homing acceptance test: a pureboot image programmed somewhere other
canonical slot must still be a working loader, and the ordinary than its canonical top slot must still be a working loader
position-independent, guarding its accidental slot — and the ordinary
--update-loader flow must put a build into the top slot from there. --update-loader flow must put a build into the top slot from there.
Two positions. Address 0, a raw .bin handed to a programmer: the staging Two positions are exercised. Address 0 (a raw .bin handed to a programmer,
install and the word-0 redirect run from copies outside page 0's slot, so the which defaults to offset 0): the staging install and the word-0 redirect
running-slot guard never blocks them. And the staging slot itself, where a both run from copies whose slots are not page 0's, so the running-slot
loader already sitting there IS the staging copy — recognized by its embedded guard never blocks the flow. The staging slot itself: a loader already
block and left in place, then streaming the new resident like any staged copy. sitting there IS the installed staging copy — the tool recognizes it by
its embedded info block and leaves it in place instead of tripping the
copy's own guard on the composed through-word — and that (older) copy
streams the new resident like any staged copy. In both cases flashing an
application through the healed resident overwrites the stale copy, vector
surgery included, and the banner proves the launch.
Usage: pbrehome.py <device_bin> <pureboot_elf> <update_bin> <mcu> <hz> Usage: pbrehome.py <device_bin> <pureboot_elf> <update_bin> <mcu> <hz>
<base_hex> <page> <baud> <app_bin> <tool_py> <workdir> <base_hex> <page> <baud> <app_bin> <tool_py> <workdir>
@@ -84,7 +90,7 @@ def main():
# The staging slot: erased flash with the loader sitting exactly where # The staging slot: erased flash with the loader sitting exactly where
# a staging copy would — the tool must leave it in place and let it # a staging copy would — the tool must leave it in place and let it
# stream the (different) update build into the resident slot. # stream the (different) update build into the resident slot.
stage = base - pb.SLOT stage = base - 512
rehome_from(pbsim, pb, device_bin, elf, hex(stage), hex(stage), update_bin, base, page, baud, app_bin, workdir, rehome_from(pbsim, pb, device_bin, elf, hex(stage), hex(stage), update_bin, base, page, baud, app_bin, workdir,
mcu, hz) mcu, hz)
print("re-home from the staging slot: converged") print("re-home from the staging slot: converged")

View File

@@ -1,9 +1,11 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Position-independence acceptance test: the identical binary, flashed one """Position-independence acceptance test: the identical pureboot binary,
slot below the resident, must serve the complete command set from there. The flashed one slot below the resident loader, must serve the complete command
info block must come back byte-identical, the write guard must refuse the set from there. The resident installs it (through-word composed by the host
staged copy's own slot and permit the resident's, and the staged copy must be layer), 'J' transfers control, and every command is exercised against the
able to rewrite the resident verbatim. staged copy — the info block must come back byte-identical, the write guard
must protect the staged copy's own slot and permit the resident's, and the
staged copy must be able to rewrite the resident slot verbatim.
Usage: pbreloc.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page> Usage: pbreloc.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <tool_py> <workdir> <baud> <tool_py> <workdir>
@@ -81,7 +83,7 @@ def main():
# Restore the resident image through the staged copy, then 'J' back # Restore the resident image through the staged copy, then 'J' back
# into it and prove it lives. # into it and prove it lives.
resident = image + b"\xff" * (pb.SLOT - len(image)) resident = image + b"\xff" * (info.slot - len(image))
pb.write_differing(loader, base, resident) pb.write_differing(loader, base, resident)
back_info = loader.enter_copy(base, 25) back_info = loader.enter_copy(base, 25)
if back_info.raw != resident_info: if back_info.raw != resident_info:

View File

@@ -1,13 +1,15 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""End-to-end protocol test: drive the simavr device with the real host tool """End-to-end pureboot protocol test: spawn the simavr device, then drive it
over its pty through flash, EEPROM, fuse and hand-over scenarios, and with the real host tool (pureboot.py, as a subprocess over the device's pty)
cross-check the tool's view against the simulator's ground-truth dumps. through flash + EEPROM + fuse + hand-over scenarios, and cross-check
the tool's view against the simulator's ground-truth memory dumps.
Usage: pbtest.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page> Usage: pbtest.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <eeprom_size> <app_bin> <tool_py> <workdir> [link] <baud> <eeprom_size> <app_bin> <tool_py> <workdir> [link]
The optional link is the runner's -l spec (usart1, sw:B5,B1, ...), for a The optional link is the runner's -l spec (usart1, sw:B5,B1, ...) for a
loader built off the chip's natural serial default. loader built off the chip's natural serial default.
Exits 0 if every scenario passes.
""" """
import os import os
@@ -55,11 +57,11 @@ def main():
# the page byte is the wire's 0-means-256. # the page byte is the wire's 0-means-256.
mega = mcu.startswith("atmega") mega = mcu.startswith("atmega")
patch = not mega or mcu.startswith("atmega48") patch = not mega or mcu.startswith("atmega48")
word_flash = base + pb.SLOT > 0x10000 word_flash = base + 512 > 0x10000
wire_base = base // 2 if word_flash else base wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0) flags = (1 if patch else 0) | (2 if word_flash else 0)
info = pb.Info( info = pb.Info(
bytes([ord("P"), ord("B"), pb.NEWEST_LOADER, 0, 0, 0, page & 0xFF]) bytes([ord("P"), ord("B"), 1, 0, 0, 0, page & 0xFF])
+ bytes([wire_base & 0xFF, wire_base >> 8, eeprom_size & 0xFF, eeprom_size >> 8]) + bytes([wire_base & 0xFF, wire_base >> 8, eeprom_size & 0xFF, eeprom_size >> 8])
+ bytes([flags]) + bytes([flags])
) )
@@ -69,7 +71,7 @@ def main():
# Session 1: knock from reset, identify, program everything, stay. # Session 1: knock from reset, identify, program everything, stay.
out = pbsim.run_tool(tool, device.pty, baud, "--info", "--fuses", "--flash", app_bin, out = pbsim.run_tool(tool, device.pty, baud, "--info", "--fuses", "--flash", app_bin,
"--eeprom", ee_path, "--stay") "--eeprom", ee_path, "--stay")
for needed in ("version", "signature", "fuses", "verify:", "stays"): for needed in ("signature", "fuses", "verify:", "stays"):
if needed not in out: if needed not in out:
fail(f"session 1 output lacks {needed!r}") fail(f"session 1 output lacks {needed!r}")
@@ -99,12 +101,7 @@ def main():
port = pb.Port(device.pty, baud) port = pb.Port(device.pty, baud)
try: try:
loader = pb.Loader(port) loader = pb.Loader(port)
live = loader.connect(15) loader.connect(15)
# The loader built from this tree and the tool beside it must
# agree on where the version numbering stands: a bump the tool
# was never told about is a loader it would refuse to speak to.
if live.version != pb.NEWEST_LOADER:
fail(f"loader reports pureboot {live.version}, the tool's newest is {pb.NEWEST_LOADER}")
loader.run_application() loader.run_application()
banner = port.read_exact(3, 5.0) banner = port.read_exact(3, 5.0)
if banner != b"APP": if banner != b"APP":
@@ -125,7 +122,7 @@ def main():
# loader, the trampoline on the application's own entry (patched-vector # loader, the trampoline on the application's own entry (patched-vector
# chips only — a boot-sectioned mega's word 0 stays the application's). # chips only — a boot-sectioned mega's word 0 stays the application's).
if patch: if patch:
flash_words = (base + pb.SLOT) // 2 flash_words = (base + 512) // 2
app = open(app_bin, "rb").read() app = open(app_bin, "rb").read()
word0 = flash_true[0] | (flash_true[1] << 8) word0 = flash_true[0] | (flash_true[1] << 8)
if rjmp_decode(word0, 0, flash_words) != base // 2: if rjmp_decode(word0, 0, flash_words) != base // 2:

View File

@@ -1,12 +1,17 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Self-update end-to-end: an application is flashed, the loader replaces """Self-update end-to-end: an application is flashed, then the loader
itself with a re-timed build, and every power-fail phase is rehearsed by replaces itself with a re-timed build through the host tool's
killing the device mid-write, restarting it from its flash dump, and letting --update-loader — and the power-fail phases of that update are rehearsed by
a re-run complete the update. killing the simulated device mid-write, restarting it from its flash dump,
and letting a re-run complete the update.
The boot-sectioned megas run the BOOTRST-unprogrammed profile reset boots The boot-sectioned megas run the BOOTRST-unprogrammed profile (reset boots
the application, whose 'L' is the application-owned loader entry — with the application; the fixture application's 'L' jump is the application-owned
--assume-fuses standing in for the fuse read simavr cannot model. loader entry), with --assume-fuses standing in for the fuse read simavr
cannot model. The patched-vector chips — the tinies and the m48s — reset
into a loader at every phase by construction: the t13a because its staging
slot carries the reset vector itself, the others through the word-0 redirect
the tool plants around the resident rewrite.
Usage: pbupdate.py <device_bin> <pureboot_elf> <update_elf> <mcu> <hz> Usage: pbupdate.py <device_bin> <pureboot_elf> <update_elf> <mcu> <hz>
<base_hex> <page> <baud> <app_bin> <tool_py> <workdir> <base_hex> <page> <baud> <app_bin> <tool_py> <workdir>
@@ -45,7 +50,7 @@ def assumed_fuses(pb, image):
image's embedded signature.""" image's embedded signature."""
info = pb.image_info(image) info = pb.image_info(image)
which, ladder = pb.BOOT_FUSE[bytes(info.signature[1:3])] which, ladder = pb.BOOT_FUSE[bytes(info.signature[1:3])]
bits = min((b for b in ladder if ladder[b] * 2 >= 2 * pb.SLOT), key=lambda b: ladder[b]) bits = min((b for b in ladder if ladder[b] * 2 >= 2 * info.slot), key=lambda b: ladder[b])
fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF)) fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF))
fuses[which] = 0xF8 | (bits << 1) | 1 fuses[which] = 0xF8 | (bits << 1) | 1
return bytes(fuses) return bytes(fuses)
@@ -88,13 +93,16 @@ def main():
# The m48s are megas without a boot section: patched vector, no fuse # The m48s are megas without a boot section: patched vector, no fuse
# preflight, and the same reset-to-0 the tinies get. # preflight, and the same reset-to-0 the tinies get.
patch = not mega or mcu.startswith("atmega48") patch = not mega or mcu.startswith("atmega48")
# Word-addressed (>64 KiB) chips use the 1 KiB slot; their loader base
# itself sits beyond the 16-bit byte space — the 644's base + slot only
# touches the 64 KiB boundary and stays byte-addressed.
slot = 1024 if base >= 0x10000 and mega else 512
reset_hex = "0" if mega else None # the boot-sectioned mega runs BOOTRST-unprogrammed here reset_hex = "0" if mega else None # the boot-sectioned mega runs BOOTRST-unprogrammed here
sys.path.insert(0, os.path.dirname(os.path.abspath(tool))) sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim import pbsim
import pureboot as pb import pureboot as pb
slot = pb.SLOT
os.makedirs(workdir, exist_ok=True) os.makedirs(workdir, exist_ok=True)
objcopy = os.environ.get("PB_OBJCOPY", "avr-objcopy") objcopy = os.environ.get("PB_OBJCOPY", "avr-objcopy")
images = {} images = {}

View File

@@ -1,8 +1,9 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Host-tool unit tests — the planning and policy logic, no simulator: """Host-tool unit tests — the pure planning and policy logic, no simulator:
programming orders and their recovery properties, the reset-vector surgery, the flash-programming orders and their recovery properties, the reset-vector
the staging composition, the boot-fuse decode, and the update preflight over surgery, the staging-slot composition, the mega boot-fuse decode, and the
fuse combinations simavr cannot model. update preflight's error/warning matrix (fuse combinations simavr cannot
model reach it here as synthetic bytes).
Usage: test_planner.py <tool_py> Usage: test_planner.py <tool_py>
""" """
@@ -26,15 +27,14 @@ def expect_error(what, fn, *needles):
fail(f"{what}: no error raised") fail(f"{what}: no error raised")
def info_of(pb, base, page, patch, flash, signature=(0x1E, 0x93, 0x0B), word_flash=False, version=None): def info_of(pb, base, page, patch, flash, signature=(0x1E, 0x93, 0x0B), word_flash=False):
scale = 2 if word_flash else 1 scale = 2 if word_flash else 1
wire_base = base // scale wire_base = base // scale
flags = (1 if patch else 0) | (2 if word_flash else 0) flags = (1 if patch else 0) | (2 if word_flash else 0)
raw = bytes((0x50, 0x42, pb.NEWEST_LOADER if version is None else version, raw = bytes((0x50, 0x42, 1, *signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8,
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8, 0, 2, flags)) 0, 2, flags))
info = pb.Info(raw) info = pb.Info(raw)
if info.flash_size != flash: assert info.flash_size == flash
fail(f"info_of({base:#x}) decodes to {info.flash_size:#x} of flash, not {flash:#x}")
return info return info
@@ -55,21 +55,6 @@ def main():
tiny = info_of(pb, 0x1E00, 64, True, 0x2000) tiny = info_of(pb, 0x1E00, 64, True, 0x2000)
mega = info_of(pb, 0x7E00, 128, False, 0x8000, signature=(0x1E, 0x95, 0x0F)) mega = info_of(pb, 0x7E00, 128, False, 0x8000, signature=(0x1E, 0x95, 0x0F))
# Versioning: the block's third byte is the loader's version, and the tool
# speaks a window of them. Every version in the window decodes, so an older
# deployed loader stays usable; one above the window is refused by name,
# since which version changed the protocol is knowledge only the tool
# holds, and it holds none about a version it has never heard of.
for version in range(pb.OLDEST_LOADER, pb.NEWEST_LOADER + 1):
if info_of(pb, 0x1E00, 64, True, 0x2000, version=version).version != version:
fail(f"pureboot {version} does not decode")
expect_error(
"unknown loader version",
lambda: info_of(pb, 0x1E00, 64, True, 0x2000, version=pb.NEWEST_LOADER + 1),
f"pureboot {pb.NEWEST_LOADER + 1}",
"newer tool",
)
# mega_boot: BOOTSZ words and the BOOTRST sense per chip — the fuse byte # mega_boot: BOOTSZ words and the BOOTRST sense per chip — the fuse byte
# index (EXTENDED on the x8 line except the m328s' HIGH, HIGH elsewhere) # index (EXTENDED on the x8 line except the m328s' HIGH, HIGH elsewhere)
# and the per-family ladders (Atmel-2486/2466/2503/2545/8271/DS40002065/ # and the per-family ladders (Atmel-2486/2466/2503/2545/8271/DS40002065/
@@ -94,7 +79,9 @@ def main():
((0x1E, 0x97, 0x05), 0x20000, 3, {0b11: 0x1FC00, 0b10: 0x1F800, 0b01: 0x1F000, 0b00: 0x1E000}), # 1284P ((0x1E, 0x97, 0x05), 0x20000, 3, {0b11: 0x1FC00, 0b10: 0x1F800, 0b01: 0x1F000, 0b00: 0x1E000}), # 1284P
) )
for signature, flash, which, ladder in cases: for signature, flash, which, ladder in cases:
chip = info_of(pb, flash - pb.SLOT, 128 if flash < 0x20000 else 0, False, flash, # Word-addressed chips carry the 1 KiB slot (their smallest boot sector).
slot = 1024 if flash > 0x10000 else 512
chip = info_of(pb, flash - slot, 128 if flash < 0x20000 else 0, False, flash,
signature=signature, word_flash=flash > 0x10000) signature=signature, word_flash=flash > 0x10000)
for bits, start in ladder.items(): for bits, start in ladder.items():
fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF)) fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF))
@@ -107,12 +94,10 @@ def main():
if prog or at != start: if prog or at != start:
fail(f"mega_boot {signature[1]:02x}{signature[2]:02b} unprogrammed: {prog} {at:#07x}") fail(f"mega_boot {signature[1]:02x}{signature[2]:02b} unprogrammed: {prog} {at:#07x}")
# Word-addressed info decode: the 1284P's base and page ride the wire # Word-addressed info decode: the 1284P's base/page ride the wire scaled,
# scaled — a 17-bit base halved into the block's two bytes, a 256-byte page # and its slot is 1 KiB.
# spelled 0 — and its slot is the same 512 bytes as everywhere else, so its big = info_of(pb, 0x1FC00, 0, False, 0x20000, signature=(0x1E, 0x97, 0x05), word_flash=True)
# staging slot lands inside the 1 KiB minimum boot section. if big.page != 256 or big.base != 0x1FC00 or big.stage != 0x1F800 or big.slot != 1024:
big = info_of(pb, 0x1FE00, 0, False, 0x20000, signature=(0x1E, 0x97, 0x05), word_flash=True)
if big.page != 256 or big.base != 0x1FE00 or big.stage != 0x1FC00:
fail(f"word-addressed info decode: page {big.page}, base {big.base:#x}, stage {big.stage:#x}") fail(f"word-addressed info decode: page {big.page}, base {big.base:#x}, stage {big.stage:#x}")
# Surgery: word 0 lands on the loader, the trampoline on the original # Surgery: word 0 lands on the loader, the trampoline on the original
@@ -169,12 +154,6 @@ def main():
fail("image_info misses the embedded block") fail("image_info misses the embedded block")
if pb.image_info(bytes((0xAA,)) * 40) is not None: if pb.image_info(bytes((0xAA,)) * 40) is not None:
fail("image_info invents a block") fail("image_info invents a block")
# An older loader's image stays readable, so a deployed build can be
# identified and installed like any other.
old = info_of(pb, 0x1E00, 64, True, 0x2000, version=pb.OLDEST_LOADER)
found_old = pb.image_info(bytes((0xAA,)) * 10 + old.raw)
if found_old is None or found_old.version != pb.OLDEST_LOADER:
fail("image_info misses an older loader's block")
# loader_image must peel a padded image down to the slot content: a raw # loader_image must peel a padded image down to the slot content: a raw
# .bin padded from address 0 (or a whole-flash read-back with the loader # .bin padded from address 0 (or a whole-flash read-back with the loader
@@ -215,15 +194,6 @@ def main():
if pb.update_preflight(bytes((0xAA,)) * 8 + tiny.raw, tiny, None) != []: if pb.update_preflight(bytes((0xAA,)) * 8 + tiny.raw, tiny, None) != []:
fail("tiny preflight should pass without fuses") fail("tiny preflight should pass without fuses")
# The 1284s' smallest boot section (512 words) is exactly the resident
# slot plus its staging slot, so self-update is possible at the minimum
# BOOTSZ — no fuse step up, the 644's geometry. That holds only while a
# slot is 512 B: at 1 KiB the staging slot would fall outside the section
# and the preflight would refuse.
notes = pb.update_preflight(bytes((0xAA,)) * 8 + big.raw, big, fuses(0xFE))
if not any("staging slot" in n for n in notes):
fail(f"1284 minimum-BOOTSZ notes: {notes}")
# The walk-region refusal: BOOTRST aimed below the loader plus app data # The walk-region refusal: BOOTRST aimed below the loader plus app data
# in the walk span errors without --force; erased spans and unprogrammed # in the walk span errors without --force; erased spans and unprogrammed
# BOOTRST pass. # BOOTRST pass.