56 Commits
v4 ... main

Author SHA1 Message Date
cbb3c58300 test: the format and ASCII rules stop being a habit here too
libavr's guidance binds this repo, and its gate checked every chip's codegen
without ever checking whether the sources it compiled were clang-format clean
or ASCII. `libavr_format_test()` does both now, over this tree alone - the
oracle's assembly needs no exclusion, being in neither glob, which is the right
answer for a vendored reference whose text is the artifact.

The sizes this repo prints were already gated: `sizes.py check-readme` is the
shape the rest of the fleet has now copied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 23:20:36 +02:00
c4b7769f84 build: the libavr pin advances to the sweep's own record
Documentation only - the guideline sweep's condensed entry, the three measured
facts about class-type constants it produced, and the port filings it left
open. No header, tool or generated input moves, so every image is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 17:02:53 +02:00
735ffab7dc fix: the four tiers stop describing features they do not have, and three gates start failing
The reading pass over this repo found the tiers disagreeing with themselves,
and every fix here was measured.

**The turn-around guard is real code.** `tsb_asm` and `tsb_tricks` wrote
`for (std::uint8_t guard = 46; guard; --guard) ;` between taking the one-wire
line and the first UDR0 store, under a comment naming it a turn-around guard.
It has no side effect, so GCC deleted it - `sts UCSR0B` went straight to
`sts UDR0` - while the hand-written oracle spends six bytes on that wait and
libavr's own half-duplex spends them through `delay::cycles`. Two of four
tiers described a feature they did not have, which made the size gradient a
comparison between different loaders. `avr::delay::cycles<one bit time>()`
bottoms out in asm and cannot be deleted.

**The entry belongs to the library, and hand-rolling it was expensive.** Three
tiers wrote their own naked `.vectors` stub with `asm volatile("clr
__zero_reg__")` - which design.md fences to libavr and never a port, and which
`tsb_tricks` denied having in its own title line. `avr::startup::entry` also
keeps the body `noinline` for a stated reason: avr-ld must not shrink a
`.vectors` section, so a loader inlined into one forfeits call relaxation
everywhere. `tsb_pure` came out **836 -> 734** bytes for that alone.
`stack::hardware` - the reset value this part guarantees, with the write kept
where a part does not - saved another four, which is what let `tsb_asm` afford
the guard it had been four bytes short of. It fills its 512-byte section
exactly now, with the whole feature set.

**`tsb_pure` had no receive timeout.** Its `rx()` was `read_blocking()`, so a
silent host wedged the password gate and the command loop forever - the one
fix the oracle's own header lists by name, and one the other three tiers
implement. It is bounded now, and 0-on-silence falls through every compare as
theirs does.

Three gates could pass without proving anything. `sizes.py check-readme`
reported a match when every row's lookup missed; `check_size.cmake` used
`CMAKE_MATCH_1` without checking the match succeeded, which is the guard its
sibling `check_unit.cmake` has and it is the size gate; `check_pi.py` raised
IndexError instead of reporting a position-independence break that changed the
image's length. And `check.sh` spelled the 37-chip list a second time beside
make_presets.py, where a chip added to one and missed in the other is a
silently unbuilt chip - it reads the presets now, and produces the same 37 and
12.

tsbtest.py gains the scenario nothing covered: a wrong password byte must
neither activate the loader nor reach the emergency erase behind it. Red-green
on a tier with the refusal removed.

Smaller, all measured or checked: the signature is `hw::db.signature` in every
tier as the page size and EEPROM end beside it already were; `act_min` derives
from the clock; pureboot.py's `rjmp` helpers refuse a part past rjmp's
4096-word reach rather than silently folding an offset (unreachable today, the
ATtiny85 sits exactly on it); the host tool calls space 2 `data` as the wire
and the loader do; `.clangd` strips the fifth GCC-only flag the build passes;
pbrig's bitclock guard reads its own ladder; pbreloc's unexplained retry is
gone, the write being reliable on five runs without it; and the four tier
sizes live in oracle/README.md's table instead of four file headers and a
CMake comment.

`--poke` before `--peek` turned out to be right - pbtest.py round-trips a poke
through the peek behind it - so the parser order and README say so now.

Every chip green, the README size table matching every image.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 16:41:50 +02:00
be78f38f3f build: the libavr pin advances past the audit sweep, and two gates start meaning something
The pin crosses libavr's phase-6 close and the guideline sweep behind it.
All 37 chips green, 43 tests each, the README size table matching every
built image, and all 13602 flash images byte-identical to the previous pin.

The bump broke one gate and exposed another as ornamental.

`check_unit.cmake` matched the autobaud loader's measured unit by the symbol
`unit_E`; libavr's rule-46 sweep renamed the member to `m_unit`, which the
mangling spells `6m_unitE`. On the RAM-home chips the check went red and said
so. On the GPIOR chips it went green - the branch that asserts the unit is
*not* in RAM passes on an empty match, and an empty match is what a stale
regex returns for every image. Both branches mean something again.

`tools/check.sh` ran the 37-chip loop under `set -e`, so the first red chip
ended the gate and the 36 behind it were never built - a stale size canary on
attiny13 would have been an alibi for every loader after it. It accumulates
now and fails at the end naming every red preset, which is the shape libavr's
own check.sh carries and the reason it carries it.

The port's own sweep, verified by byte identity: the four TSB tiers' 16-byte
info block is `std::to_array` rather than an extent written beside the
sixteen elements the compiler can count, the three-member serial and loader
configs break one member per line, the turn-around loops are braced, and the
test fixture's config pair is a deduced `std::array` (rules 36, 40, 34). Two
comments stop narrating how the code came to be and one stops citing a repro
at a path it left two phases ago (rules 12, 13).

pureboot's identity stamp stays the raw array rule 36 bans, and now says why:
its reads must fold to immediates because the bytes are in program memory and
a formed address is dereferenced as data space. As a `std::array` the read
loop stopped unrolling and emitted exactly that - measured at +8 B and a
wrong answer on the wire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 15:05:07 +02:00
d3d8d07905 editor: clangd and cmake work from a committed vscode workspace
The three files libavr's consumers carry, in the vendored shape: this repo
rides as a submodule (wateralarm), so .clangd names no database -- the
checkout that is opened as a folder names build/atmega328p-generated in
.vscode/settings.json -- and holds the stand-ins clang needs for GCC's AVR
dialect plus the removals for the codegen flags the loader TUs carry and
clang has no spelling for (-fira-algorithm, -fno-split-wide-types,
-fno-tree-ter, -fno-ivopts). The libavr pin advances to the editor-audit
fixes. The preset builds green from the pin and clangd reports zero errors
on pureboot.cpp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:49:01 +02:00
b4243a89a3 docs: the one-wire section's sizes catch up with v9
The v9 update moved the size table and the tightest-fit paragraph and
missed this section: the autobaud + OSCCAL twins measure 480 now, not
502, and the hardware half-duplex trio is 392/428/440. The claims
around the numbers were true all along - one-wire still measures the
same as two-wire, and the +42..50 delta over stock still holds exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:35:55 +02:00
e4390d2ba8 build: the libavr pin advances past phase 6, at byte parity everywhere
The pin crosses libavr's phase 6 - the renamed system surface, the named
serial configs, the receiver-tolerance table, the paged SPM receipts -
and every loader image comes out size-identical: the full matrix on six
representative chips (the exhaustive cross product on three of them),
the stock and autobaud columns untouched, the four tsb tiers back on
their recorded floors at 510/526/638/836.

Byte parity was not free, and the two libavr defects it surfaced were
fixed there rather than absorbed here. The EEPROM write procedure's
step 2 - the SPMEN spin - had landed unconditionally and cost every
build six bytes for a wait a polled loader can never take; it is scoped
now, and the loaders state the datasheet's own omission clause
(spm_interlock::omitted, DS40002061B 8.6.3). The blocking page
erase/write grew an internal wait the tiers' settle() already provides,
so the tiers issue the command form and pureboot keeps its host-driven
sp_spm path.

What the port states rather than inherits: the stock 115200 at 16 MHz
sits +2.1 % past the receiver-tolerance table libavr now holds rates
to, so the hardware links say .allow_baud_error = true - the same
2.5 % envelope pureboot_baud_feasible() has always enforced, proven on
silicon across the fleet. rx_ready() reads readable() now.

Alongside the pin: rule 33's ASCII sweep over every source (docs keep
their typography), rule 34's InsertBraces in .clang-format with the
tree reformatted, std::array over the simavr runners' raw buffers, and
the stale Studio size in ide/README.md replaced by the claim its
check-flags gate actually holds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:43:44 +02:00
0cb83ff36f build: the libavr pin advances past the consumer-report fixes
timer::engine gains stop()/start() and a runtime TOP, adc gains
disable()/enable(), and libavr_programming_targets() stops leaving .fuse bytes
in the flash HEX. Every one of them is additive, and this port adopts none of
them yet: its 687 built images come out byte-identical across the pin change,
which is what the advance is here to keep true.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:18:55 +02:00
8aa1721709 README: which slot the 510 bytes belong to, and the way in without a reset
The Chips footnote said "the slot's last word" against a table whose subject is
the loader, where it means the *lower* slot's — the trampoline holding the
application's relocated reset vector, and a staging copy's own last word during
a self-update. The resident has all 512 bytes of the slot it runs in; 510 is
what an image must fit so a copy staged one slot down leaves that word alone.

And a section on entering from a running application, for the boards whose
adapter does not drive reset and which therefore have no edge to open a window
with. Deciding when to jump stays the application's business — a console
command, a held pin, an idle timeout — so there is no knock detector and no
header here, only the mechanics: the base as a --defsym symbol, and the three
things that must be true first (interrupts off, WDRF clear, and any peripheral
holding the link released, since a loader entered by a jump inherits the
application's registers rather than reset values).

Not the noipa indirect call run_app() uses, which is the obvious thing to copy
and the wrong one: that is a position-independence measure belonging to a loader
that runs the same image from either slot. An application is linked at a fixed
base, so a plain call to the symbol comes out `call 0x7e00` in four bytes where
the laundered form spends two ldi's and a helper call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:12:34 +02:00
2655d058f4 pbhw: an absent probe is not a destroyed loader
The slot-survives-erase check reads back over ISP by design — an independent
reader is the only witness worth having about a loader that has just been
asked to erase around itself. With the one probe on another board it printed
"ISP read failed" as a red, which is the same word a destroyed loader would
get, and it is permanently red on the two deployments with no ISP header at
all.

It names its witness now and falls back to the link when there is no
programmer, saying that the loader is then reporting on its own slot — weaker
for exactly the reason it is worth having, since a destroyed loader could not
answer at all. An absent instrument is a fact about the bench and a wrong byte
is a verdict on the subject; a check that prints them identically stops being
read.

Both paths exercised on hardware: ISP on the Uno, the link on the ATtiny13A.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 20:25:47 +02:00
e9c3897d4f build: the libavr pin advances to the v9 era
Built and tested against it in a clean checkout of this port, through its own
submodule rather than a working-tree override, so the pin is what was proved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 18:48:31 +02:00
546b1589a3 pureboot v9: seal every command, and stop guarding what the seal covers
'W' handed the loader a whole page with no ack inside it and sp_spm handed any
wire byte to SPMCSR, so a dropped byte re-aligned the stream and page data
arrived where commands belong. That is how a page-address byte became
BLBSET|SELFPRGEN on the tempmon board and programmed its lock bits.

The first answer was to refuse that one command. It was the wrong shape twice
over: it forbade a lock-bit write the owner may want, and it left every other
command decided by bytes nobody checked. v9 checks them instead. One header for
every command — opcode, selector, address, count, seal — folded and compared
before the command is decoded, and *answered* before any payload moves: '+'
accepts, 0xd4 (the ack inverted) refuses and nothing happened. An ack cannot do
this job; it reports a command that has already run.

It is smaller than v8 everywhere: 1284P 506→480, m8 498→480, 328P 484→468,
t13A 474→460. The seal costs 14 bytes; bit opcodes in place of the letters pay
for it twice over, since a letter costs a compare and a branch where a bit costs
a skip. Both guards go — the lock-bit refusal because the seal covers it, the
running-slot write guard because what it defended against was a wire fault
naming an address and a wire fault can no longer name one. That one is a real
trade: a host bug aimed at the running slot now lands. It buys a resident copy
that can write its own slot, which is the only self-update route on a chip whose
boot section *is* the slot.

Two things the tests caught, both introduced here. Removing the invalid-opcode
arm made every byte a command, so the knock stopped being harmless against a
loader already in session and ate the five bytes behind it — identify moves to
bit 5, which both 'p' and 'b' carry, so the knock is inert again and version
discovery still works before the version is known. And the SPM value rides the
count field because a data byte would arrive after the seal was checked.

pbselfwrite and pbglitch are the new gates, both red-green: the same erase of
the running page refused unsealed and performed sealed, and every header byte
damaged after sealing refused where the identical damage before sealing is
obeyed. Both judge by the simulator's flash, not the loader's opinion of it.
pbreloc and pbrehome lose their write-guard probes, which is what those two
gates replace. Defeating the seal in the loader turns seven tests red.

37 of 37 chips green with the exhaustive size matrix; README protocol section
and every size row rewritten. pbhw gains an adversarial --seal-rounds sweep for
the bench.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 18:38:41 +02:00
eb213e1025 one-wire: a lost echo and a dead line are not the same report
The blind-write path said "lost to the device's ack" for any missing echo, and
the count is what distinguishes two different faults. Some bytes lost is the
device's ack winning the line against the host's series resistor — ordinary,
and what the knock retry absorbs. *Every* byte lost is nothing coming back at
all, which means the line is not free: a pin held low, a wedge, or an RX that
is not on it.

Found pointing the wrong way on purpose-built hardware. This rig's LED demo
ends by driving every port pin low, and one of them is the shared link — so a
knock into a finished demo got no echo whatsoever and was told the device had
acked, when nothing had answered and nothing could. Same retry either way, but
blaming an ack that never happened sends the reader to the protocol when the
answer is a pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:32:18 +02:00
3e4bfbaf48 pbhw: the marker check assumed a board whose port-open is not a reset
Its comment said "opening the port does not reset a board whose DTR is
unwired, so this simply listens" — true of the tiny it was written against,
false of an Arduino, and this is the generic harness. Where DTR is wired to
reset, that open resets the part and the activation window comes first, so a
fixture emitting its banner once says it on the far side of a wait the suite
cannot know the length of: the window is a compile-time constant and nothing
on the wire reports it. The suite read the silence as an application that
never ran, on a board where it demonstrably had.

So --marker-wait, defaulting to the 2.5 s that was hardcoded, and a failure
that names the window as the candidate rather than leaving the next person to
suspect the loader. The other half is the fixture: PUREBOOT_HEARTBEAT makes
the observation independent of when the listener arrives, which is what the
rig's own builds now pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 17:12:08 +02:00
579ca81b27 one-wire: the knock's lost byte is the wiring, and the diagnosis was unreachable
Measured on an ATtiny13A with the link folded onto PB3 and the FTDI's TX
reaching it through 1 k: a knock aimed at a loader already in session loses
its second byte every time, 8 runs of 8, never intermittently. The first byte
draws a prompt while the second is still going out and the device's push-pull
ack wins the line against the resistor, so that byte is destroyed rather than
delayed — which is what the README predicted and the sim bridge cannot show,
since it arbitrates the line by queueing.

The recovery for it existed and could not run. Two defects:

OneWirePort.write read its echo with read_exact, whose contract is to raise, so
the "one-wire echo missing — is the adapter's RX tied to the line?" message was
unreachable on any line that simply fell quiet, and a bare "timeout: got 0 of 1
bytes" surfaced in its place. The one message the class exists to produce could
never be produced. The read is speculative and is now read_available.

And any raise from write aborted _handshake before the retry loop that exists
to absorb exactly this, whose docstring already claimed it "converges into an
already-live session" — true on a pty, impossible on real wiring. The knock is
now the one write marked blind: a missing echo there is a property of the
shared line, counted and reported under -v rather than raised. Every other
write is ack-paced and cannot collide, so a missing echo there still means an
RX that is not on the line, and still raises.

Both gates green on Windows (31/31 m328p, 16/16 t13a); on hardware the
reconnect now converges on the first knock, the surviving prompt being all the
handshake needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 16:29:17 +02:00
5d520a1ff9 pbhw: --one-wire never reached the suite's own sessions
The flag was plumbed through pbrig.Deployment to the host-tool subprocess
calls and nowhere else, so identity() and scan() opened a raw port and drove
a shared line as though it were two wires. On real one-wire hardware the
adapter's echo answers the knock before the device does, so the suite would
have died at its very first check — "the loader never answered; nothing below
can be trusted" — for the one deployment the flag exists to test, and every
result after it is gated on that check passing.

Both now open through pbrig.Rig.open_port(), which applies the deployment's
link mode. The gap underneath was that only the subprocess path could reach
those facts at all; anything driving the protocol in-process had to restate
them, and did not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:25:46 +02:00
f71d76a815 pureboot v8: one-wire on every backend
HALF_DUPLEX deploys a shared line per backend. The hardware USART takes
the library's .half_duplex turn-around — RXD and TXD tied off-chip, each
reply byte held to transmit-complete before the line can be released
(m8 404 B, m328P 440, 1284P 460; the window poll runs through the
outlined release-line call at 18 or 22 cycles a poll, measured off the
built loops and held per chip by pureboot.window.halfduplex). The
software and autobaud links fold onto the RX pin — RX == TX spells the
same — and cost nothing: the frame's direction wrap is what the dropped
second-pin init paid, and the worst image in the space is unchanged at
the 1284s' 502 of 512, now with its one-wire twin proven equal across
the exhaustive matrix. The host gains --one-wire, the echo discard a
shared line requires: the adapter's echo is matched byte for byte and a
reply interleaving a blind write — a loader already in session
re-prompts inside the knock — is held for the reader. The device runner
models the shared line by direction (drives only while the firmware's
DDR reads input, decodes only while the firmware owns it, supplies the
host-side echo), extends the USART pin-ownership model to RXEN's hold
on RXD, and starts the pty USART from the datasheet's zeroed UCSR#B:
simavr's TXEN-set reset plus its clear-UDRE-on-TXEN-drop otherwise
wedges the first transmitter after a receiver-only program, which the
half-duplex window gate caught as a banner that never came. v7 is
tagged at its era's last commit; v8 changes nothing on the wire.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 02:32:11 +02:00
47419400f6 build: the pin advances over the one-wire serial feature
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 02:31:55 +02:00
bad8b6b43e build: the pin advances over the bounded calibration
libavr's calibrate() now bounds its measurement loop, starts the pulse
on an observed edge, and re-arms a rejected pulse on the remaining
budget instead of one-strike booting the application. The autobaud
images pay +16..20 B — every slot still fits, the worst now the 1284s'
502 of 512 — and the stock images are byte-identical, kept so by
fitting the loader's flag set per backend: -fno-ivopts stays on the
fixed-baud bodies it shrinks and comes off the autobaud body, where it
duplicated the calibration countdown into a 9-cycle loop against the
contracted seven.

One deployed constant moved and its gate caught it: the calibrate
wait's budget poll re-laid from ten cycles to nine (the exit branches
land where block layout puts them), so pureboot.window.autobaud
measured -10 % until AUTOBAUD_POLL_CYCLES and the README's derived
seconds were re-measured — the default autobaud window is 36 M cycles,
4.5 s at 8 MHz. Full gate green on all 37 chips; the README's autobaud
column carries each chip's rebuilt worst configuration, machine-checked
against the built trees.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 22:49:55 +02:00
5d1b4497d4 build: the libavr pin advances over the drain and delay contracts
The hand-over drains move to the explicit drain_unbounded() — both link
adapters drain only after their own write, so the frame is in flight by
construction and the bounded default's countdown would be dead bytes;
the images stay byte-identical. window_polls() states its arithmetic
through dev::cycles_for with the whole window converted before the
per-poll division — one truncation instead of one per second, same
instructions, only the countdown's immediate moves. Every size in the
matrix is unchanged; the full gate is green on all 37 chips.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 17:31:22 +02:00
a743ea64a3 tool: a size collision across build trees is an error, not a coin toss
sizes.py merged every owned tree's rows and let the last one win, so a
stale reflect tree — last built before the window constants moved —
reported the atmega8's old stock size over the fresh build and failed
the README check with yesterday's number. Generated and reflect must
answer with the same bytes (the identity invariant), so the same target
measuring two sizes is a stale tree or an identity breach; collect()
refuses now, naming both trees. The stale reflect trees are removed —
the reflect sweep rebuilds them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:32:35 +02:00
aa66cfccff pureboot: the poll-cost lookup rides the baud parameter
usart_of<0> is an incomplete type on the USART-less chips, and a static
member initializer with only non-dependent operands is checked when the
template is parsed, not when it is instantiated — so the address probe
broke every tiny build without hardware_link ever being named. The
lookup moves into a member function template taking the link's own baud
parameter, the dependence carrier that defers it to instantiation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:13:44 +02:00
459d463283 pureboot: the classic megas' idle poll is a bit-skip — 7 cycles, and their stock window goes wide
The window gate's first full sweep caught it on the atmega8: 6.22 s
measured against 8 declared, the exact 7/9 of a poll modeled as an
extended-I/O lds + skip on a chip whose UCSRA sits in bit-addressable
I/O and compiles to a 2-cycle skip. poll_cycles now follows the status
register's home (7 below 0x40, 9 above). At 16 MHz over 7 cycles the
poll count no longer fits uint24_t, so the classic megas' stock windows
take the wide countdown — 8.000 s measured on all three, +4 B of stock
image (m8 362, m16/m32 364), README stock rows updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:11:34 +02:00
8e7cc86fb3 pureboot: the activation window gets a behavioral gate, and honest per-poll constants under it
The window's per-poll cycle counts were hand-counted for a uint32_t
countdown, but every default window fits uint24_t, whose decrement chain
is one sbci shorter — so deployed loaders ran 9/10ths of their stated
seconds (a 328P's 8 s was 7.2 s on the wire). No golden-asm pin can hold
this: the loops compile in consumer context. pbwindow.py measures the
behavior instead: it installs a real application beside the loader
through the host tool's own plan_flash (surgery included), starts the
simulator with the line idle, and reads the cycle of the first transmit
— the application's banner, so that cycle is the window. Held at plus or
minus 2 percent per chip (pureboot.window), red at -10.0 percent against
the old constants, green with poll_cycles now counted for the narrow
countdown (hardware 9, software 7; window_polls() solves narrow-first
and adds the wide loop's cycle where the count forces uint32_t — a count
narrow only at the wide cost stays wide, so the choice cannot
oscillate). The autobaud window is its poll budget at the measured ten
cycles a poll, gated the same way (pureboot.window.autobaud), and the
README carries that arithmetic now. No version bump: timing-window
precision is not meaningful behavior, v7 stays.

The gate flushed out two runner gaps. The software bridge accepted any
falling edge as a start bit, so the device's own TX-init glitch decoded
as a stray byte; it re-samples mid-bit now and abandons a false start,
as silicon does. And after avr_reset, the idle-line re-raise was
silently dropped: ioport pin irqs are IRQ_FLAG_FILTERED and the irq's
cached value survives the reset the port latch does not, so the device
read the line stuck low, calibrate() measured reset-to-first-edge as one
wrapping pulse, and the first knock after a reset could boot the
application instead of locking — the intermittent autobaud failure.
bridge_reset forces a real transition (0 then 1, no cycles between).

The README's Autobaud column now carries each chip's worst
configuration — autobaud with OSCCAL baked, on a USART's own pins where
the chip has one (tinies: autobaud + OSCCAL) — the numbers the existing
pureboot_autobaud_osccal[_on_usart0] matrix points already gate;
sizes.py checks the column against exactly those targets. Tool sizes
and window prose updated with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:05:42 +02:00
c8ac61779e tool: survive our own leftovers — drain the fresh port, shorten the identity read
--stay leaves the loader's final prompt in the USB pipeline; a fresh
invocation on a board that resets when its port opens then flushes too
early, trusts the stale prompt, and spends the new activation window on
a 2-second identity read against a device that never heard its knock —
collecting the application's banner as an unknown signature. Three
host-side moves, no device bytes: the line is drained until quiet
(bounded, 250 ms) before the port's first knock — once per port, since a
mid-session re-knock faces no foreign bytes and its own window is
already burning; the identity read_exact drops 2.0 to 0.5 s, dozens of
times the worst real answer, so any false prompt match leaves room for
the retry that already works; and the tool version drifts to 8. The
StaleDTRPort fixture models the whole moment — stale prompt in transit,
reset holding the device off the line, a finite window, the banner —
red against the old tool in exactly the field shape (unknown signature
from banner bytes), green now; LoaderPort answers its prompt to the
knock rather than to a read count, which the drain exposed as a
call-order coupling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:05:16 +02:00
4362886c39 build: the libavr pin advances over the trait projection
The de-string-2 pass upstream: every peripheral block behind generated
instance traits, int-keyed, the string layer gone. The port's share of it
is two spellings — the char usart_digit that existed to be pasted into
register names becomes the int unit the usart template now takes, and the
tsb tiers' one reg<"UBRR0"> is the flat hw::ubrr0 — plus the pbapp
harness probing has_usart<0>() instead of instance-name strings. Nine
loader codegen families rebuilt green through their full workflows (size
matrix and simulator protocol suites included); every image holds its
recorded size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 20:08:10 +02:00
0513d07e87 build: the libavr pin advances to current main
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:24:10 +02:00
d4ab28aa17 build: the libavr pin advances to current main
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 06:48:37 +02:00
b60f182105 review: the port's findings — the version map speaks v7, costs told true
The in-file version window now carries the v7 line its own comment
claimed to hold; the GPIOR note counts words, not instructions; the
USART-release cost and citation match the silicon (two bytes on the
classics, §20.6.3); and the 512-byte claim reads as the slot bound it
is. The libavr pin advances over the review pass — images byte-identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 20:05:42 +02:00
5bc35a9733 build: the libavr pin advances over the instance traits
Byte-identical images — the traits resolve the same database indices the
retired string forms did; the tightest image is compared outright.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 17:18:07 +02:00
8f6319c068 pureboot 7: the same features in fewer words on every chip
Four cuts, none touching what the loader can do. The entry stub stops
re-doing the reset logic's own SP write where the datasheet guarantees
RAMEND (stack::hardware — the classic megas keep theirs). The autobaud
unit moves into GPIOR2:GPIOR1 wherever the chip has the pair: one-word
accesses, no RAM object, and the host's measured-clock peek follows it
by version and geometry. 'J' rides the unified decode, carrying a
selector it ignores so its address is the same two reads as every other
command — the tool sends the bare form to older residents. run_app stops
insisting on a body of its own. The fleet lands at 358–410 B stock and
438–474 B autobaud; the tightest image in the space — the 1284s'
autobaud on a USART's own pins with the OSCCAL trim — drops from 510 to
484 of its 512. Every chip's suite is green on the wire that changed,
and the README's table is machine-checked against the built images.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 17:02:38 +02:00
a3ea099105 pureboot: the detail reaches become the library's own API
The three things this repo took from under libavr's counter are now over
it, and the local copies fold away. The USART release on a software
link's pins is the library's init contract (its guard here becomes a
deletion, byte-identical images held by the gate); the WDRF routing test
is power::peek_reset_cause().watchdog instead of a hand lookup of the
flag's register; the tsb tiers' baud arithmetic is the public solver.
libavr pin advances over those three additions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 16:12:35 +02:00
c535c4756c pureboot: the activation countdown in the narrowest type that holds it
The fixed-baud window counted down a uint32 where almost every window fits
24 bits; the countdown now takes avr::uint24_t when the poll budget allows
(the autobaud budget's own choice), uint32 past 16.7M polls — four bytes
off every fixed-baud image on every chip, the full suites green on the
changed window. The README size table is refreshed — its autobaud column
had also gone stale by the no-assembly pass's measurement-loop win, which
nothing gated: sizes.py check-readme now runs as the gate's final stage,
where every tree is freshly built and the table can actually be held.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 14:47:29 +02:00
321ff8a4ee test: the device runners speak the C++ the rest of the repo does
pureboot_device and the tsb device, C until now, rewritten in C++23 with
every modeled behavior intact — the PGERS Z-mask and m48-discard ioctl
wraps, the GPIO bridge's timing and pacing, the tiny NVM's write-once
buffer, pin ownership, and the PB_PTY/TSB_PTY lines the harnesses parse.
The one linkage fact worth a comment: simavr's parts headers (uart_pty.h)
carry no C++ guards where its core headers do, so those includes sit in an
extern "C" block. Warning-clean at -Wall -Wextra on the build line; the
full protocol suites on all four sim-driven chips prove the conversion.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 13:59:03 +02:00
531ae6c8dc build: the libavr pin advances over the audit rounds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:52:05 +02:00
b3f41caf6e audit: round three on the port — the hardware scan check fails soft
A rig hiccup mid-walk records the scan check as failed and lets the suite
continue, matching its siblings' envelope; a nonexistent --port path
reports as an error instead of a traceback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:31:32 +02:00
7716e1e291 audit: the port's pass — the scan that could not walk, and the drift a generator ends
--scan's walk was unwalkable on POSIX: probe rates have no termios
B-constant, so the first off-nominal probe raised out of the loop. The port
speaks termios2 BOTHER now (red-proven on a pty at 9984 Bd), the probe's
open lives inside the walk's error handling, an fd no longer leaks on an
unmakeable rate, and the swallowed unknown-signature reply is named at
timeout instead of reported as silence. CMakePresets.json's generator emits
the submodule toolchain path it had drifted from — a hand edit on a
generated file, exactly the class rule 10 exists for — and presets.generated
gates the pair from here on (the  marker CMake rejects at the
presets root stayed out; the check is the guard). The over-slot image guard
the tsb runner gained reaches the pureboot runner too; the GPIO bridge's
delivery comment states the hardware truth (RXC at the stop bit's sampling
point); the hardware suite gains the scan check — the one place the rate
physics is real; and the libavr pin advances over both audit rounds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:12:17 +02:00
1b18f10f4e docs: the OSCCAL axis, --scan, and the RC-oscillator deployment risk
Configuration gains the OSCCAL row; Deployment says what the build cannot
see (±10 % factory trim against a frame's ~±4 %, and silence that reads as
wiring); the update section names an OSCCAL bake as a link change in effect,
declared with --staged-baud; the host-tool section documents --scan and the
measured clock --info adds on an autobaud session; the version map gains 6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 01:03:15 +02:00
52c4cdab32 pureboot.py: --scan walks a silent loader's rate; --info decodes the clock
The RC-oscillator answer's host half. --scan probes ±10 % around the built
rate in 2 % steps, nearest first, one activation window (one reset) per
probe: a fixed-baud loader whose oscillator drifted answers at the ratio,
and the report gives the session workaround (--baud), the offset, the
OSCCAL direction at ~1 %/step, and the autobaud way out. The walk and the
advice are logic-tested (test_scan.py, red-proven on the trim direction) —
a pty carries bytes at any rate, so the wire cannot arbitrate them.

On an autobaud session --info now reads the measured bit period from
ram_start — the geometry table gains that column — and undoes the unit's
encoding ((cycles − 8) / 4, floored: libavr's spin granule and per-bit
overhead), so the printed clock is the true one within a granule; --clock
turns it into a stated drift. The autobaud end-to-end asserts the figure
inside exactly that envelope at both clock points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 01:01:35 +02:00
69f089e53a pureboot 6: a build-time OSCCAL trim, applied ahead of every reset path
The RC-oscillator answer's device half (dev/tasks.md in libavr): OSCCAL joins
pureboot_add_loader() as one optional byte, written at the top of run() before
the WDRF bail so the watchdog hand-over inherits the corrected clock too.
Orthogonal to the backend — an autobaud build may carry it purely for the
application. No value, no code: the stock image differs from v5 in exactly
the version's two bytes (the stamp and the 'b' immediate).

Measured: +6 B where OSCCAL takes sts (328P, 404→410), +4 B in low I/O
(t85, 402→406); the tightest image in the space (1284 autobaud on USART
pins, 504) carries the sts form at 510 of 512. New gates: the OSCCAL size
points on every chip, the wire-observed trim byte on both addressing
classes (test/pbosccal.py, red-green), and the autobaud unit pinned to
ram_start (test/check_unit.cmake, red-green) — the address --info's
measured-clock read is about to rely on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 00:53:14 +02:00
fe0d9f8790 build: the submodule is how libavr arrives; the era already carries it
The pin the whole history now encodes becomes the primary route: the
submodule default replaces FetchContent and the unpinned forge fallback,
LIBAVR_ROOT stays as the tandem-development override, the presets already
take the toolchain file from the submodule, and the Studio projects anchor
their include path there — correct by construction. The version tags and
the one-command historical build are documented beside the version map.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 00:28:49 +02:00
3ce817ea03 ide: the Atmel Studio solution master has, on the libavr port
master carries bootloader.atsln, so main does too. Two projects, because a
.cppproj is one binary at one flag set and this build has hundreds: the stock
328P pureboot loader (USART0 at 115200 on a 16 MHz crystal), and the tsb_asm
tier that occupies the same 512-byte section master's own tsb project targeted.
Both come out byte-identical to the Ninja build — 404 B and 510 B of .text —
in both configurations.

Debug keeps -Os and adds only -gdwarf-4. A loader's section is a correctness
bound, and -Og builds this source to 590 B: the link at 0x7e00 accepts that
without a diagnostic, 78 bytes past flash end, where rcall/rjmp wrap modulo
flash size and the image dies just after activation. Debug info costs no flash,
so the optimisation level stays where correctness needs it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 22:15:28 +02:00
af0dd15a77 tsb: the policy floor, measured and kept
A fourth tier answering one question: what does the full TinySafeBoot
feature set cost in C++ under pureboot's rules — no assembly, no
register variables, every pureboot lesson applied. 638 bytes, protocol
suite green: 198 below the idiomatic tier, 126 above the 512 B section,
and above the tiers that pay with the banned mechanisms (526 global
registers, 510 with two asm routines). The gap decomposes into the rent
policy-clean C++ pays for state held across calls — push/pop and
argument threading a global-register protocol avoids — and both
control-flow merges tried measured larger than the split cases they
replaced, while the data merge (one send loop over both memories) paid.
The tiers stay; this one keeps the floor an artifact instead of a claim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:57:13 +02:00
a54075e526 test: the device runner refuses an image that runs past flash end
A boot-linked image larger than its slot cannot execute on hardware, and
the naive copy smashed the heap beyond avr->flash — after which the
simulation misbehaved in ways that pointed everywhere but at the size:
phantom byte losses on the UART, garbage in SPMCSR, all downstream of
the overrun. The size gate had said it plainly; now the runner does too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:57:13 +02:00
aec430c2e1 docs: name the two data-space regions a poke cannot survive
Writing 0x60-0x61 on an ATtiny13A reliably garbled the link, which looked like a
loader defect. It is the loader's own footing: .noinit lands at exactly 0x60,
size 2, and on an autobaud build that is unit_ — the measured bit period, and the
whole of its static RAM. Overwrite it and the next reply is timed against
garbage, so the symptom is a mangled prompt byte and no error, because nothing
went wrong except the rate both ends had agreed on.

Identical in kind to poking the stack at the top of SRAM, and cleared by a reset.
test/pbautobaud.py already steered its RAM round-trip clear of the bottom of SRAM
for this reason; only the README had not said it. Both regions are named there
now, beside the note that --poke does reach OSCCAL but that a session survives
only a step or two of moving the clock under itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:48:09 +02:00
9a8a5b0082 tools: measure the images, and hold the README to what is built
The size matrix proves every image fits its slot and says nothing about the
numbers the README prints. Those drift silently: caller_page() took eight bytes
off every build at once, so all fifteen rows went stale together and no test
noticed, because nothing was over budget. sizes.py check-readme compares the
table against the built images; sizes.py max reports the largest image per chip
and anything over its slot.

It is a check rather than a generator, so the table stays prose someone can
write. No chip geometry lives here either: the image/budget pairs come out of
each build's own CTestTestfile.cmake, which is what the gate checks, so a chip
added or a budget changed needs no edit. Only trees a preset still owns are
read — a stale directory answers with a size that was true once.

16243 images across 37 chips today, none over budget, tightest tsb_asm at 510
of 512.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:45:20 +02:00
7702c6b700 test: gate autobaud's logic floor, and say what an RC oscillator costs
The ATtiny13A run left autobaud's low-clock lock looking unreliable: 1/5 at
19200 on a 1.2 MHz RC part. The simulator does not reproduce it. At an exact
clock the calibration is solid down to ~36 cycles a bit and fails outright by
~31 — a sharp edge, not a fraying one — where the real part was already 1 in 5
by ~59. So the effect is the oscillator's own jitter and not backend logic, and
the two floors are different quantities about a factor of two apart.

Both are worth having. pureboot.autobaud gates a tight-bit point, since its two
existing clock points both sat near 100 cycles a bit and would not notice the
floor moving. The README carries the other half: both floors side by side, the
per-clock envelope measured on silicon, and the reason budgeting the logic's ~36
on an RC part is wrong.

It also carries the trap that produced the confusion. On a patched-vector chip an
erased application region walks back up into the loader, so every expired window
opens another and the host's retries eventually catch the pulse — 5/5 where the
same part with an application resident gives 1/5. Measure with an application in
place, or the fixture flatters the backend.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:44:21 +02:00
77dd45aeca pureboot.py: an update follows the staging copy onto its own link
--update-loader works by entering copies of the *new* image and letting them
rewrite the resident. Those copies speak the link they were built for, but the
host went on knocking with the session's baud and backend — the resident's. Where
the image changed either, the staging copy was installed and then never answered:
resident untouched, and on a 1 KiB tiny the staging slot is the whole application
region, so the application was already gone.

The wire cannot be probed for it. 512 bytes of position-independent code carry no
header saying what rate they were built for, so the operator declares it:
--staged-baud and --staged-autobaud, applied from the jump into the staging copy
onward. Retuning goes through the open port — SetCommState or tcsetattr on the
live handle, never a reopen — because a DTR pulse would reset the copy being
talked to. Undeclared against a changed link it still cannot work, but the error
now names that as the cause instead of reporting the bare activation timeout that
sent the operator looking at wiring.

The README's idempotence claim needed the same qualification: from step 2 a
re-run must reach the new image, and after step 3 word 0 points at the staging
copy, so on a patched-vector part the resident's link reaches nothing at all.

Found on an ATtiny13A, where two controls differing only in the activation window
updated cleanly and so isolated the link as the variable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:31:27 +02:00
433bec3e58 tools: a hardware harness, so a board can be proven and not just a protocol
check.sh proves the protocol under simavr on every chip; it cannot prove a
board. Two things live only on silicon — an RC oscillator that is not on its
nominal, and a reset edge that has to come from somewhere — and until now the
scripts that reached them were per-session scratch on the machine holding the
programmer, which is where the ATtiny13A run's findings nearly stayed.

pbrig.py is the primitives, knowing nothing per-board: every deployment fact is
a flag or a PUREBOOT_* variable. Two rig facts are encoded in it because neither
is guessable and each cost a session to learn: an ISP access *is* the reset edge
where the adapter's DTR is unwired, so a session begins with an ISP touch and
knocks immediately after; and avrdude splits -U on colons, so a Windows drive
letter breaks the spec and every file goes as a bare name with avrdude run in
its own directory. Its `rate` subcommand is the one that turns "the loader is
silent, so the wiring must be wrong" into a number, by sweeping the host rate
against a fixed cycles-per-bit transmitter — PUREBOOT_HEARTBEAT makes the
existing fixture into one, software link only, since the hardware-link idle owes
the self-update tests its command loop.

pbhw.py takes every bound from the info block the loader reports, so one run
covers a 1 KiB tiny and a 128 KiB mega alike. Both are exercised on an ATtiny13A:
backup verified against a known-good capture, the clock measured at 9.048 MHz
against a 9.6 MHz nominal, and the suite 11/11.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:06:15 +02:00
e89000f73e pureboot: the pin axis, and the mute it was hiding
Pins move the image for exactly one reason — a bit-banged link on a USART's own
pins has to release that USART — and the matrix said outright that they were no
axis, so the tightest configuration in the space was one nothing built. Not
subtly, either: the 1284's slot ends at flash end, so that build does not merely
exceed the size test's limit, it fails to link. pureboot_{sw,autobaud}_on_usart
{0,1} are gate points in both matrix modes now, and the exhaustive sweep carries
the pins across its whole cross product. The hand-measured table is the gate's
output: 506 B of 512 for the 1284 autobaud on USART0's pins, 504 on USART1's.

pureboot.mute drives the defect itself — an application hands over with USART0
still enabled and the loader on those pins must still answer. Reaching that
needed the runner to know an enabled USART owns its TxD, which simavr does not
model at all: it wires a USART through IRQs and never takes the pin from the
port. It also brings UCSRnB up with TXEN already set where silicon clears the
register, so the runner restores the reset value for the USART it models — the
mute must come from the application, not from power-on. The fixture stays
silent, since nothing is listening on the USART it brings up.

test_handshake.py, written where no gate could run it, is pureboot.handshake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 22:08:56 +02:00
b0737f7cc0 test: move the handshake regression in beside the rest
It was written next to the loader source; the harness lives at the repo root.
Not registered with ctest yet — it belongs beside pureboot.planner, which is
the other test of the host tool's pure logic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 18:22:19 +02:00
07f93caba8 pureboot: take the running slot from avr::startup::caller_page
The write guard's anchor was costing a materialised pointer and a byte swap to
use one byte of it. The libavr primitive answers it in a single load, which is
eight bytes off every build — and what lets the USART release fit the tightest
configuration in the space: the 1284 autobaud on USART-shared pins was 514 of
its 512 and is now 506, with the default pinning down from 510 to 502.

Verified on silicon: the guard still refuses an erase aimed at the slot it runs
from, and still permits one in the application region.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 17:57:47 +02:00
45f10f843a pureboot: release a USART left enabled on the software link's pins
A software or autobaud link on a USART's own pins (PD0/PD1 on the mega328P, so
the Uno's USB bridge reaches it) was mute after an application handed over with
that USART still enabled: its TXEN keeps the USART owning the TX pin, so the
bit-banged transmitter cannot drive it — the loader locked and obeyed commands
but never answered. The link's init now clears the UCSRnB of the USART whose
TXD is its TX pin. Guarded with if constexpr on that pin match, so a link on
non-USART pins emits nothing: +4 bytes on a USART-pin build (494 of 512 for the
mega328P autobaud), zero on the default pb0/pb1 matrix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 17:22:48 +02:00
34b47048ca pureboot.py: bump the tool version to 5
The drain fix changes the tool's activation behaviour; mark it. The loader
version window is unchanged — the wire protocol did not move.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 17:11:03 +02:00
392035923f pureboot.py: bound the activation drain against a flooding target
The post-prompt settle loop in _handshake had no deadline, so a target that
never falls quiet — a board stuck in a reset loop, whose UART-reset garbage
carries a stray prompt byte — spun the tool forever. Bound it by the handshake
deadline; a real loader still settles on its first quiet read. Regression:
test/test_handshake.py (flood terminates, valid loader still connects).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 17:10:40 +02:00
f98ed406b8 pureboot 5: one command pair for every memory, and a clock-free backend
R/r/w/F collapse into G and g over a selector byte naming the space — flash,
EEPROM, data, fuse, SPM — with the flash bank in its high nibble. Four command
bodies, four transfer loops and four argument decodes become one of each, and
W joins the same decode instead of keeping an address form of its own. The
loader shrinks while gaining everything below: on the 1284P the stock build
goes 480 -> 432 B and the software one 496 -> 450.

What the freed space buys:

  - Data space. On AVR one pointer spans SRAM, the register file and the whole
    I/O space, so G over space 2 reads all three. pureboot keeps zero static
    RAM and pushes no register, so at loader entry an application's SRAM is
    still what the application left there — this is a post-mortem, not just a
    poke hole. As its own command it needed a dispatch arm and a loop; as one
    more space it is a single ld/st.
  - Host-issued SPM. W fills the page buffer and stops; erase, write and RWW
    re-enable are writes to space 4, which reach the same fused store-and-SPM
    pair through the transfer's own address and data. Any SPM operation, lock
    bits included, is now reachable and the loader carries no page-commit logic.
    The four-cycle SPMCSR-to-SPM window is why that primitive stays fused: no
    host can hit it across a serial link, and that — not the byte count — is
    the floor on how low-level a bootloader's primitives can go.
  - Byte addresses everywhere. The bank in the selector retires the
    word-addressed wire the >64 KiB parts needed, so the 1284s stop being the
    outlier.

SERIAL autobaud is a third backend on the same loader, over libavr's
software_autobaud: no clock, no baud, one binary per chip for every F_CPU and
every rate. Activation counts poll iterations rather than seconds and bounds
every wait, so a stray pulse cannot hold an unattended device.

b answers with the version and signature only; the host derives geometry from
the signature, which is what an autobaud build requires anyway. An update image
is a bare slot with no device to ask, so every image carries a six-byte stamp —
the same bytes b answers with, and the source of both — that the loader never
reads from flash and the host refuses to install a mismatch against. The
running-slot write guard moved onto the SPM commit, which covers erase and
write both where guarding W covered neither directly.

The position-independence lint now proves the property instead of a proxy for
it: the image must come out byte-identical linked at a different base.
-fno-move-loop-invariants left the tuned flag set — it was fitted to a command
loop carrying four transfer bodies and costs bytes now that it carries one.

Verified: the exhaustive matrix on all 37 chips (every clock x every baud x
every backend, non-standard rates included, plus the autobaud build) —
8174 size checks, no failures, tightest fit the 1284s' autobaud at 510 of 512.
Behavioral suites green on every chip class: t13a 10/10, t85 11/11, m8 13/13,
m16a 13/13, m48pa 13/13, 328P 23/23, 644A 17/17, 1284P 17/17. Data-space
round trip through --peek/--poke and the autobaud handshake are both red-green
proven.

The two prototype sources and their findings file go; the README carries the
protocol and dev/done.md in libavr carries the reasoning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 15:50:48 +02:00
52 changed files with 6608 additions and 3042 deletions

View File

@@ -7,8 +7,9 @@ TabWidth: 4
UseTab: ForIndentation
AlignEscapedNewlines: DontAlign
AllowShortFunctionsOnASingleLine: Empty
AlwaysBreakTemplateDeclarations: true
BreakTemplateDeclarations: Yes
BreakBeforeBraces: Custom
BraceWrapping:
AfterFunction: true
InsertBraces: true
...

38
.clangd Normal file
View File

@@ -0,0 +1,38 @@
# Editor accommodations for the second frontend. No compilation database is
# named here: this repo rides as a submodule in its consumers, and this file
# travels with it — a consumer's own database then covers these sources, with
# that project's loader flags. The checkout that is opened as a folder names
# its build tree in .vscode/settings.json instead.
CompileFlags:
Add:
# clang has no 24-bit integer and GCC's are keywords, not macros, so the
# editor needs a stand-in for avr::uint24_t. The next width up is the only
# one available — clang rejects _BitInt(24) on this target.
- -D__uint24=unsigned long
- -D__int24=long
# clangd forwards the driver's system includes but not its own header
# directory, so <stdint.h> resolves to avr-libc's, which still gates the
# limit and constant macros on the C++98 opt-in.
- -D__STDC_LIMIT_MACROS
- -D__STDC_CONSTANT_MACROS
# isr::emit spells a vector number into [[gnu::signal(N)]], which clang
# rejects rather than ignores — enough of them in one TU to reach the
# default limit of 19 inside the headers and truncate the parse.
- -ferror-limit=0
Remove:
# Codegen shaping the loader TUs carry and clang has no spelling for.
- -fira-algorithm=*
- -fno-split-wide-types
- -fno-tree-ter
- -fno-ivopts
- -fno-move-loop-invariants
# The build promotes warnings for the compiler that has to be right about
# them; in the editor the flag paints a second frontend's opinions in the
# colour reserved for things that do not compile.
- -Werror
Diagnostics:
Suppress:
# clang's AVR `signal` attribute takes no arguments and it knows none of
# progmem, naked or OS_main. A misspelling is what the build is for.
- attribute_wrong_number_arguments
- unknown-attributes

6
.vscode/extensions.json vendored Normal file
View File

@@ -0,0 +1,6 @@
{
"recommendations": [
"llvm-vs-code-extensions.vscode-clangd",
"ms-vscode.cmake-tools"
]
}

37
.vscode/settings.json vendored Normal file
View File

@@ -0,0 +1,37 @@
{
// clangd is the language server; the cpptools engine would parse every file
// a second time and disagree, since nothing tells it about a cross
// compiler.
"C_Cpp.intelliSenseEngine": "disabled",
// --query-driver lets clangd ask the cross compiler for its own system
// includes and target. The database is named here rather than in .clangd
// because that file travels with the driver into a consumer's submodule,
// where a build tree of this repo's own need not exist.
"clangd.arguments": [
"--compile-commands-dir=${workspaceFolder}/build/atmega328p-generated",
"--query-driver=**avr-g++*",
"--header-insertion=never"
],
// The presets are the build interface; the toolchain file inside the
// libavr submodule is the one place the compiler is chosen. The submodule
// carries no local/toolchain for it to discover, so the prefix is named
// here, for the window that opens this folder.
"cmake.useCMakePresets": "always",
"cmake.environment": {
"LIBAVR_TOOLCHAIN": "D:/dev/libavr/local/toolchain/avr-gcc-16.1.0-mingw"
},
"cmake.configureOnOpen": true,
"cmake.options.statusBarVisibility": "compact",
"files.watcherExclude": {
"**/build/**": true,
"**/libavr/**": true
},
"files.associations": {
".clangd": "yaml",
".clang-format": "yaml"
}
}

View File

@@ -2,64 +2,73 @@ cmake_minimum_required(VERSION 3.28)
project(tsb_libavr LANGUAGES CXX)
# libavr from a local checkout (LIBAVR_ROOT) or the forge; the toolchain file
# comes from the same checkout via CMakePresets.json.
include(FetchContent)
# libavr rides as the pinned submodule; LIBAVR_ROOT (cache or environment)
# overrides it for tandem development against a working tree. The toolchain
# file comes from the submodule via CMakePresets.json either way.
if(NOT LIBAVR_ROOT AND DEFINED ENV{LIBAVR_ROOT})
set(LIBAVR_ROOT $ENV{LIBAVR_ROOT})
endif()
if(NOT LIBAVR_ROOT)
set(LIBAVR_ROOT ${CMAKE_CURRENT_SOURCE_DIR}/libavr)
endif()
if(LIBAVR_ROOT)
FetchContent_Declare(libavr SOURCE_DIR ${LIBAVR_ROOT})
else()
FetchContent_Declare(libavr GIT_REPOSITORY git@git.blackmark.me:avr/libavr.git GIT_TAG main)
if(NOT EXISTS ${LIBAVR_ROOT}/CMakeLists.txt)
message(FATAL_ERROR "libavr not found at ${LIBAVR_ROOT} - run: git submodule update --init libavr")
endif()
FetchContent_MakeAvailable(libavr)
add_subdirectory(${LIBAVR_ROOT} libavr-build)
include(${LIBAVR_ROOT}/cmake/checks.cmake)
if(PROJECT_IS_TOP_LEVEL)
add_compile_options(-Werror) # warnings are errors for the port's own code
enable_testing()
# Rules 11 and 33 over this repo's own sources. The oracle's assembly needs
# no exclusion: it is neither formatted nor ASCII-checked, being in neither
# glob, which is the right answer for a vendored reference whose text is
# the artifact.
libavr_format_test()
# The behavioral tests drive the real wire protocols over a simavr pty
# (as the host tools do) and actually flash the device. The runners are
# host programs built at configure time against libsimavr; if they or
# Python are missing, only the size tests run.
find_program(_host_cc NAMES cc gcc)
# host programs built at configure time against libsimavr (C++23 - what
# the distribution's compiler speaks in full); if they or Python are
# missing, only the size tests run.
find_program(_host_cxx NAMES c++ g++)
find_package(Python3 COMPONENTS Interpreter)
if(_host_cc AND Python3_FOUND)
if(_host_cxx AND Python3_FOUND)
set(PB_DEVICE ${CMAKE_BINARY_DIR}/pureboot_device)
execute_process(
COMMAND ${_host_cc} -O2 -I/usr/include/simavr -I/usr/include/simavr/parts
-o ${PB_DEVICE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pureboot_device.c
COMMAND ${_host_cxx} -std=c++23 -Wall -Wextra -O2
-I/usr/include/simavr -I/usr/include/simavr/parts
-o ${PB_DEVICE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pureboot_device.cpp
-lsimavr -lsimavrparts -lelf -lutil
RESULT_VARIABLE _pbdev_res ERROR_VARIABLE _pbdev_err)
if(NOT _pbdev_res EQUAL 0)
message(STATUS "pureboot_device not built (${_pbdev_err}) protocol tests skipped")
message(STATUS "pureboot_device not built (${_pbdev_err}) - protocol tests skipped")
unset(PB_DEVICE)
endif()
if(LIBAVR_MCU STREQUAL "atmega328p")
set(TSB_DEVICE ${CMAKE_BINARY_DIR}/tsb_device)
execute_process(
COMMAND ${_host_cc} -O2 -I/usr/include/simavr -I/usr/include/simavr/parts
-o ${TSB_DEVICE} ${CMAKE_CURRENT_SOURCE_DIR}/test/device.c
COMMAND ${_host_cxx} -std=c++23 -Wall -Wextra -O2
-I/usr/include/simavr -I/usr/include/simavr/parts
-o ${TSB_DEVICE} ${CMAKE_CURRENT_SOURCE_DIR}/test/device.cpp
-lsimavr -lsimavrparts -lelf
RESULT_VARIABLE _dev_res ERROR_VARIABLE _dev_err)
if(NOT _dev_res EQUAL 0)
message(STATUS "tsb_device not built (${_dev_err}) protocol tests skipped")
message(STATUS "tsb_device not built (${_dev_err}) - protocol tests skipped")
unset(TSB_DEVICE)
endif()
endif()
endif()
endif()
# The ELF is only a container (symbols, section headers) and is never flashed
# The ELF is only a container (symbols, section headers) and is never flashed -
# and the host tool's load_image() dispatches on extension, so handing it one
# would silently program the header bytes. Every loader image therefore gets
# both flashable forms beside it at link time: .hex for avrdude, and .bin for
# the host tool's raw path (which is what the reloc and update tests convert to
# on the fly). .eeprom is dropped EEPROM content is its own update.
# on the fly). .eeprom is dropped - EEPROM content is its own update.
function(add_image_outputs name)
add_custom_command(TARGET ${name} POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O ihex -R .eeprom
@@ -68,35 +77,39 @@ function(add_image_outputs name)
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.bin)
endfunction()
# The TinySafeBoot protocol reimplemented on libavr in three variants that trade
# The TinySafeBoot protocol reimplemented on libavr in variants that trade
# clarity for size. Each links into the ATmega328P boot section (BOOTSZ selects
# its size; BOOTRST vectors a reset to its base) with -nostartfiles a polled
# loader has no use for the crt or the vector table. The naked entry sits in
# .vectors, laid first, and runs. The boot base is FLASHEND+1 minus the section
# size; the linker section-start and the source's boot_bytes agree. tsb_app is
# its size; BOOTRST vectors a reset to its base) with -nostartfiles - a polled
# loader has no use for the crt or the vector table. The entry sits in
# .vectors, laid first, and runs - avr::startup::entry on the policy tier,
# the experiment tiers' own naked stubs elsewhere, each documented in its
# source. The boot base is FLASHEND+1 minus the section size; the linker
# section-start and the source's boot_bytes agree. tsb_app is
# the application's reset vector, pinned to 0 here so the loaders jump to a
# named function; --pmem-wrap-around lets relaxation turn that absolute jump
# into the wrapped rjmp AVR's modulo-flash PC actually executes.
# All three implement the full oracle feature set (see oracle/README.md):
# All four implement the full oracle feature set (see oracle/README.md):
# watchdog bail, one-wire half-duplex, config-page activation timeout, password
# gate, emergency erase, config/flash/EEPROM read-write. They differ only in how,
# and the size gradient is the cost of that "how" — see dev/lessons.md.
# tsb_asm the tricks tier's C++ with exactly two routines in asm (the
# bounded rx and the page-store loop the two whose remaining
# cost is the C ABI itself): 510 B in the 512 B section the
# hand-written 500 B oracle occupies. Everything else, from
# bring-up to dispatch, is C++ on libavr.
# tsb_tricks — no asm at all: the whole-loader register allocation lives in
# and the size gradient is the cost of that "how".
# tsb_asm - the tricks tier's C++ with exactly two routines in asm: the
# bounded rx and the page-store loop, the two whose remaining
# cost is the C ABI itself. Everything else, bring-up to
# dispatch, is C++ on libavr.
# tsb_tricks - no asm at all: the whole-loader register allocation lives in
# global register variables (Y walks the page pointer), every
# helper is a tiny noinline primitive placed by the
# global-register store rules, pages stream straight to
# SPM/EEPROM, and the bring-up is the two reset-non-default
# registers only. 526 B in the 1 KB section (BOOTSZ=10) — 14
# over the oracle's section, from 168 over at this tier's first
# floor.
# tsb_pure — pure idiomatic libavr, one function per command, TU-local
# (internal linkage), streaming (no SRAM page buffer): 836 B in
# the 1 KB section.
# SPM/EEPROM.
# tsb_pure - pure idiomatic libavr, one function per command, TU-local
# (internal linkage), streaming (no SRAM page buffer).
# tsb_policy - the policy floor: pureboot's rules (no asm, no register
# variables) with every pureboot lesson applied, and the
# measured evidence that the 512 B fit is a property of the
# mechanisms philosophy #5 bans.
#
# What each measures is oracle/README.md's table, which is the one place the
# four numbers and the hand-written loader's own are compared.
#
# add_tsb_variant(<name> <boot-section-bytes>)
function(add_tsb_variant name bytes)
@@ -124,14 +137,19 @@ endfunction()
# chips build pureboot alone.
if(LIBAVR_MCU STREQUAL "atmega328p")
add_tsb_variant(tsb_asm 512)
add_tsb_variant(tsb_policy 1024)
add_tsb_variant(tsb_pure 1024)
add_tsb_variant(tsb_tricks 1024)
# The policy tier's floor is measured with the loop flags pureboot's size
# work found (a loader's loop bodies all contain calls); the other tiers
# keep the flag set their recorded floors were measured with - none.
target_compile_options(tsb_policy PRIVATE -fno-move-loop-invariants -fno-tree-ter)
endif()
# pureboot the pure-constraint port (see pureboot/README.md): one source,
# pureboot - the pure-constraint port (see pureboot/README.md): one source,
# no inline assembly, no global register variables, every libavr chip,
# fitting each chip's smallest boot sector. The geometry and the
# pureboot_add_loader() deployment function live in pureboot/CMakeLists.txt
# pureboot_add_loader() deployment function live in pureboot/CMakeLists.txt -
# the unit a downstream project consumes; everything below is this port's
# own build: the stock loaders, their tests, and the size matrix. The
# distinct binary dir keeps the `pureboot` target's output name free.
@@ -139,7 +157,7 @@ add_subdirectory(pureboot pureboot-cmake)
# The stock loader: the family-default deployment (crystal/RC clock, the
# chip's natural link, default pins). The activation window stays a cache
# variable re-timing a deployed loader is a self-update with a re-timed
# variable - re-timing a deployed loader is a self-update with a re-timed
# build. pureboot9 is that re-timed build, and what the update test installs.
set(PUREBOOT_TIMEOUT 8 CACHE STRING "pureboot activation window, seconds")
pureboot_add_loader(pureboot TIMEOUT ${PUREBOOT_TIMEOUT})
@@ -153,10 +171,25 @@ if(PROJECT_IS_TOP_LEVEL)
if(Python3_FOUND)
add_test(NAME pureboot.pi
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/check_pi.py
${CMAKE_OBJDUMP} ${CMAKE_NM} $<TARGET_FILE:pureboot> ${PUREBOOT_BASE_HEX})
${CMAKE_OBJDUMP} ${CMAKE_OBJCOPY} ${CMAKE_CXX_COMPILER} ${LIBAVR_MCU}
$<TARGET_FILE:pureboot>
${CMAKE_BINARY_DIR}/CMakeFiles/pureboot.dir/pureboot/pureboot.cpp.obj
${PUREBOOT_BASE_HEX})
add_test(NAME pureboot.planner
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_planner.py
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py)
add_test(NAME pureboot.scan
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_scan.py
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py)
# CMakePresets.json is generated; hand edits drift the moment the
# generator reruns, so the gate holds the pair together.
add_test(NAME presets.generated
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/tools/make_presets.py
--check)
add_test(NAME pureboot.handshake
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_handshake.py)
add_test(NAME pureboot.updatelink
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_update_link.py)
endif()
# The protocol test flashes this fixture through the loader with the real
@@ -175,6 +208,38 @@ if(PROJECT_IS_TOP_LEVEL)
${CMAKE_BINARY_DIR}/pbtest-work)
set_tests_properties(pureboot.protocol PROPERTIES TIMEOUT 180)
# The activation window as a measured duration: application installed,
# line idle, the first transmit is the application's banner - its
# cycle is the window the source declares, held to +/-2 % (one
# mis-counted cycle per poll is a 10 % shift).
add_test(NAME pureboot.window
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbwindow.py
--device ${PB_DEVICE} --loader $<TARGET_FILE:pureboot>
--mcu ${PUREBOOT_SIM_MCU} --hz ${_pb_stock_hz}
--base ${PUREBOOT_BASE_HEX} --page ${PUREBOOT_PAGE}
--baud ${_pb_stock_baud} --app $<TARGET_FILE:pbapp>.bin
--seconds ${PUREBOOT_TIMEOUT}
--tool ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
--workdir ${CMAKE_BINARY_DIR}/pbwindow-work)
set_tests_properties(pureboot.window PROPERTIES TIMEOUT 300)
# The half-duplex loader's window, same gate: its poll runs through
# readable()'s release-line test, whose outlined call re-shapes the
# whole loop - a per-class cycle count (poll_cost() in pureboot.cpp)
# that only the built image can prove, chip by chip.
if(PUREBOOT_HAS_USART)
add_test(NAME pureboot.window.halfduplex
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbwindow.py
--device ${PB_DEVICE} --loader $<TARGET_FILE:pureboot_hd>
--mcu ${PUREBOOT_SIM_MCU} --hz ${_pb_stock_hz}
--base ${PUREBOOT_BASE_HEX} --page ${PUREBOOT_PAGE}
--baud ${_pb_stock_baud} --app $<TARGET_FILE:pbapp>.bin
--seconds ${PUREBOOT_TIMEOUT}
--tool ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
--workdir ${CMAKE_BINARY_DIR}/pbwindow-hd-work)
set_tests_properties(pureboot.window.halfduplex PROPERTIES TIMEOUT 300)
endif()
# The position-independence acceptance test: the identical image,
# installed one slot lower, must serve the full command set.
add_test(NAME pureboot.reloc
@@ -186,8 +251,30 @@ if(PROJECT_IS_TOP_LEVEL)
set_tests_properties(pureboot.reloc PROPERTIES TIMEOUT 180
ENVIRONMENT "PB_OBJCOPY=${CMAKE_OBJCOPY}")
# The seal against the one command that proves it: erasing the page the
# loader runs from, refused unsealed and honoured sealed. Destroys the
# loader by design, so it gets a device of its own.
add_test(NAME pureboot.selfwrite
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbselfwrite.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud}
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbselfwrite-work)
set_tests_properties(pureboot.selfwrite PROPERTIES TIMEOUT 180)
# The seal against a link that damages bytes on purpose: every header
# field flipped after sealing must be refused, and the identical flip
# applied before sealing must be obeyed.
add_test(NAME pureboot.glitch
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbglitch.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud}
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbglitch-work)
set_tests_properties(pureboot.glitch PROPERTIES TIMEOUT 180)
# Entering the loader from a running application with no reset
# between, over a page buffer the application dirtied the case the
# between, over a page buffer the application dirtied - the case the
# loader declines to guard and the host repairs. Hardware forbids the
# state here (SPM runs only from the boot section); simavr does not,
# which is what makes it constructible.
@@ -218,7 +305,7 @@ if(PROJECT_IS_TOP_LEVEL)
endif()
# The self-update end-to-end: the re-timed build (same source, only
# the timeout differs a byte-different image) replaces the resident
# the timeout differs - a byte-different image) replaces the resident
# through --update-loader, with every power-fail phase rehearsed from
# the runner's flash dumps.
pureboot_add_loader(pureboot9 TIMEOUT 9)
@@ -234,13 +321,13 @@ if(PROJECT_IS_TOP_LEVEL)
endif()
# The size matrix: every configuration axis that could move the image
# size the serial backend (different code), the USART instance
# (different registers), the clock (different constants), and the baud
# through the shapes its bit timing takes — each combination must still
# fit the chip's slot budget. Pins are size-neutral (port and bit are
# immediate operands) and the timeout is a constant, so neither adds an
# axis. The stock build is one point of this matrix and already has its
# test.
# size - the serial backend (different code), the USART instance
# (different registers), the clock (different constants), the baud
# through the shapes its bit timing takes, and the pins through the one
# thing they decide (whether a bit-banged link has to release the USART
# that owns them) - each combination must still fit the chip's slot
# budget. The timeout is a constant and adds no axis. The stock build is
# one point of this matrix and already has its test.
function(pureboot_size_variant name)
pureboot_add_loader(${name} ${ARGN})
add_test(NAME ${name}.size
@@ -248,43 +335,64 @@ if(PROJECT_IS_TOP_LEVEL)
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
# The autobaud loader (pureboot/autobaud.md): one clock-agnostic image per
# chip, no clock × baud axis, size-tested against the same per-chip budget on
# every chip.
#
# Only the unified version is built. Its two predecessors —
# pureboot_autobaud_pure.cpp at 508 B and pureboot_autobaud_reg.cpp at 512 —
# had 4 B and 0 B of margin on the 1284P, and the fix for the activation hang
# costs ~22, which puts them at 530 and 534. Neither can ship, so the choice
# the branch existed to offer is settled by measurement rather than taste.
# The sources stay for the record; autobaud.md carries the numbers.
function(pureboot_autobaud_variant name source)
pureboot_add_autobaud(${name} ${source})
add_test(NAME ${name}.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:${name}>
# The autobaud loader: one clock-agnostic image per chip, so it has no
# clock x baud axis of its own - the matrix below sweeps those for the
# fixed-baud builds, and this one binary has to serve all of them at run
# time. Size-tested against the same per-chip budget as every other variant.
pureboot_add_loader(pureboot_autobaud SERIAL autobaud)
add_test(NAME pureboot_autobaud.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:pureboot_autobaud>
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
pureboot_autobaud_variant(pureboot_autobaud_uni pureboot_autobaud_uni.cpp)
# The measured unit's home is wire contract, not layout accident: the
# host reads the bit period from it (--info's measured clock). In the
# GPIOR home the image must carry no RAM copy at all; in the RAM home it
# is the loader's only RAM object, at the very start of SRAM.
add_test(NAME pureboot_autobaud.unit
COMMAND ${CMAKE_COMMAND} -DOBJDUMP=${CMAKE_OBJDUMP} -DELF=$<TARGET_FILE:pureboot_autobaud>
-DRAM_START=${PUREBOOT_RAM_START} -DGPIOR=${PUREBOOT_UNIT_GPIOR}
-P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_unit.cmake)
# One point of the exhaustive matrix, named from its resolved parameters
# so the enumeration cannot collide with itself. Unreachable rates drop
# out here rather than aborting the configure.
function(pureboot_matrix_point hz baud link)
# so the enumeration cannot collide with itself. `pins` is empty for the
# default pair, or the index of the USART whose own pins a bit-banged
# link sits on. Unreachable rates drop out here rather than aborting the
# configure.
# The optional trailing argument is the one-wire shape of the same link:
# ONE_WIRE folds a software point onto its RX pin (the default, or the
# named USART's RXD), HALF_DUPLEX is the hardware USART's turn-around.
function(pureboot_matrix_point hz baud link pins)
set(_name pbm_${hz}_${baud}_${link})
if(link STREQUAL "software")
pureboot_baud_feasible(${hz} ${baud} 1 _ok)
set(_args SERIAL software)
if(NOT pins STREQUAL "")
list(APPEND _args RX ${PUREBOOT_USART${pins}_RX} TX ${PUREBOOT_USART${pins}_TX})
set(_name ${_name}_on${pins})
endif()
if(ARGC GREATER 4 AND ARGV4 STREQUAL "ONE_WIRE")
if(NOT pins STREQUAL "")
set(_args SERIAL software RX ${PUREBOOT_USART${pins}_RX} TX ${PUREBOOT_USART${pins}_RX})
else()
list(APPEND _args RX pb0 TX pb0)
endif()
set(_name ${_name}_1w)
endif()
else()
pureboot_baud_feasible(${hz} ${baud} 0 _ok)
set(_args USART ${link})
if(ARGC GREATER 4 AND ARGV4 STREQUAL "HALF_DUPLEX")
list(APPEND _args HALF_DUPLEX)
set(_name ${_name}_hd)
endif()
endif()
if(_ok)
pureboot_size_variant(pbm_${hz}_${baud}_${link} CLOCK ${hz} BAUD ${baud} ${_args})
pureboot_size_variant(${_name} CLOCK ${hz} BAUD ${baud} ${_args})
endif()
endfunction()
# Clock points: the shipped-fuse floor (CKDIV8), the calibrated RC, and
# the crystal the stock build assumes (the tiny13's ladder is its own RC
# menu it has no crystal option).
# menu - it has no crystal option).
if(LIBAVR_MCU MATCHES "^attiny13")
set(_matrix_clocks 1200000 4800000 9600000)
set(_full_clocks 128000 600000 1200000 4800000 9600000)
@@ -295,34 +403,40 @@ if(PROJECT_IS_TOP_LEVEL)
endif()
# The exhaustive cross product: every clock a deployment plausibly runs
# the internal oscillators, the shipped CKDIV8 floor, the plain
# crystals and the UART crystals against every rate, against every
# - the internal oscillators, the shipped CKDIV8 floor, the plain
# crystals and the UART crystals - against every rate, against every
# backend. Beyond the ladder the list carries the slow rates a
# sub-megahertz oscillator is left with, which no ladder rate reaches
# (16000 Bd is the only rate the 128 kHz oscillator holds exactly); at
# the fast clocks those same rates also select the software UART's
# 16-bit _delay_loop_2 bit spin (two words more setup at each of its five
# sites), the largest image the space produces and a shape the ladder
# default always the *fastest* rate a clock reaches never picks.
# default - always the *fastest* rate a clock reaches - never picks.
#
# Bounded to one chip per size-bearing class: flash addressing (the
# word-addressed 1284), hand-over shape (the patched vector on the tinies
# and m48s), page size, and USART inventory. Everything else in the image
# is chip-independent code, so a further chip buys builds and no
# coverage; every chip outside the set carries the compact matrix.
# Every chip runs the full cross product: the size-bearing classes (flash
# addressing, hand-over shape, page size, USART inventory) are what make
# the image differ, and a chip outside them is expected to match its class
# - but "expected" is what a matrix is for, and the whole sweep is cheap
# enough to run rather than reason about. PUREBOOT_FULL_MATRIX is what
# selects it; the compact matrix below is the per-commit default.
get_property(_full_bauds GLOBAL PROPERTY PUREBOOT_BAUD_LADDER)
list(APPEND _full_bauds 16000 4800 2400 1200)
set(_matrix_spot attiny13a attiny85 atmega48pa atmega8a atmega168pa
atmega328p atmega164a atmega644a atmega1284p)
if(DEFINED ENV{PUREBOOT_FULL_MATRIX} AND LIBAVR_MCU IN_LIST _matrix_spot)
if(DEFINED ENV{PUREBOOT_FULL_MATRIX})
foreach(_matrix_hz IN LISTS _full_clocks)
foreach(_matrix_baud IN LISTS _full_bauds)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software "" ONE_WIRE)
if(PUREBOOT_HAS_USART)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 0)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 0 ONE_WIRE)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0 "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0 "" HALF_DUPLEX)
endif()
if(PUREBOOT_HAS_USART1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software 1 ONE_WIRE)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1 "")
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1 "" HALF_DUPLEX)
endif()
endforeach()
endforeach()
@@ -341,16 +455,103 @@ if(PROJECT_IS_TOP_LEVEL)
endforeach()
list(GET _matrix_clocks -1 _matrix_top_hz)
pureboot_size_variant(pureboot_sw_wide CLOCK ${_matrix_top_hz} BAUD 9600 SERIAL software)
# The pin axis at the widest software image - the slowest ladder rate
# against the fastest clock, whose bit spin needs the 16-bit delay
# loop - with the USART release on top of it. The exhaustive sweep
# above carries the same axis across its whole cross product.
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_sw_wide_on_usart0 CLOCK ${_matrix_top_hz} BAUD 9600
SERIAL software RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_sw_wide_on_usart1 CLOCK ${_matrix_top_hz} BAUD 9600
SERIAL software RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
endif()
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_usart1 USART 1)
endif()
# One configured deployment end to end — a real board's shape rather
# The pin axis at its fixed points, in both matrix modes. The autobaud
# loader carries no clock and no baud, so the sweep has nothing to vary
# for it - yet it is the tightest image in the space, and on a USART's
# own pins it pays the release too: that combination is the one that
# overflowed the 1284's slot. The software build on those pins is the
# same deployment the mute test drives.
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_sw_on_usart0 SERIAL software
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
pureboot_size_variant(pureboot_autobaud_on_usart0 SERIAL autobaud
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_sw_on_usart1 SERIAL software
RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
pureboot_size_variant(pureboot_autobaud_on_usart1 SERIAL autobaud
RX ${PUREBOOT_USART1_RX} TX ${PUREBOOT_USART1_TX})
endif()
# The OSCCAL axis at its fixed points: the stock shape, and the tightest
# image in the space with the trim on top - the axis adds one register
# write, and these points hold both of its addressing encodings to every
# chip's budget.
pureboot_size_variant(pureboot_osccal OSCCAL 0x9c)
pureboot_size_variant(pureboot_autobaud_osccal SERIAL autobaud OSCCAL 0x9c)
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_autobaud_osccal_on_usart0 SERIAL autobaud OSCCAL 0x9c
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_TX})
endif()
# The one-wire axis at its fixed points, in both matrix modes (the
# exhaustive sweep carries the same shapes across its cross product):
# the software link folded onto one pin, the tightest autobaud image
# likewise - on the default pin and on the USART's own RXD, whose
# release the now-driven shared pin needs where a receive-only link
# would not - and the hardware USART's half-duplex turn-around, stock
# and at the widest fixed-baud shape.
# The two spellings deliberately split across the two points: HALF_DUPLEX
# folds TX onto RX, RX == TX states the same thing directly.
pureboot_size_variant(pureboot_1w SERIAL software RX pb0 HALF_DUPLEX)
pureboot_size_variant(pureboot_1w_autobaud_osccal SERIAL autobaud OSCCAL 0x9c RX pb0 TX pb0)
if(PUREBOOT_HAS_USART)
pureboot_size_variant(pureboot_1w_on_usart0 SERIAL software
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_RX})
pureboot_size_variant(pureboot_1w_autobaud_osccal_on_usart0 SERIAL autobaud OSCCAL 0x9c
RX ${PUREBOOT_USART0_RX} TX ${PUREBOOT_USART0_RX})
pureboot_size_variant(pureboot_hd HALF_DUPLEX)
list(GET _matrix_clocks -1 _hd_top_hz)
pureboot_size_variant(pureboot_hd_wide CLOCK ${_hd_top_hz} BAUD 9600 HALF_DUPLEX)
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_usart1_hd USART 1 HALF_DUPLEX)
endif()
# The trim byte, observed through the wire from the first prompt - one
# chip per OSCCAL addressing class: extended I/O on the 328P (data 0x66,
# an sts - DS40002061B section 36), plain I/O on the 85 (data 0x51, an out -
# Atmel-2586 section 21).
if(LIBAVR_MCU MATCHES "^(atmega328p|attiny85)$" AND DEFINED PB_DEVICE)
if(LIBAVR_MCU STREQUAL "atmega328p")
set(_osccal_addr 0x66)
else()
set(_osccal_addr 0x51)
endif()
get_target_property(_osccal_hz pureboot_osccal PUREBOOT_HZ)
get_target_property(_osccal_baud pureboot_osccal PUREBOOT_BAUD)
add_test(NAME pureboot.osccal
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbosccal.py
${PB_DEVICE} $<TARGET_FILE:pureboot_osccal> ${PUREBOOT_SIM_MCU}
${_osccal_hz} ${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_osccal_baud}
${_osccal_addr} 0x9c ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbosccal-work)
set_tests_properties(pureboot.osccal PROPERTIES TIMEOUT 120)
endif()
# One configured deployment end to end - a real board's shape rather
# than the stock assumption: the ATmega328P on its shipped 1 MHz fuses,
# the software UART on hand-picked pins (TX = PB1, RX = PB5), the ladder
# baud (9600). The full protocol suite runs against it, fixture
# application included, over the runner's GPIO bridge proving the
# application included, over the runner's GPIO bridge - proving the
# configuration plumbing produces a working loader, not just one that
# fits.
if(LIBAVR_MCU STREQUAL "atmega328p" AND DEFINED PB_DEVICE)
@@ -374,6 +575,85 @@ if(PROJECT_IS_TOP_LEVEL)
set_tests_properties(pureboot.custom PROPERTIES TIMEOUT 180)
endif()
# Hand-over with the USART that owns the loader's pins left enabled - the
# state an application reaches by jumping in without a reset, and the one
# that made a bit-banged loader on PD0/PD1 (where the Uno's USB bridge
# lands) receive and obey while answering nothing. Run where it was found
# on silicon; the runner supplies the pin ownership simavr has no model
# for, which is what lets this fail when the release is gone.
if(LIBAVR_MCU STREQUAL "atmega328p" AND DEFINED PB_DEVICE)
get_target_property(_mute_hz pureboot_sw_on_usart0 PUREBOOT_HZ)
get_target_property(_mute_baud pureboot_sw_on_usart0 PUREBOOT_BAUD)
get_target_property(_mute_link pureboot_sw_on_usart0 PUREBOOT_LINK)
add_executable(pbapp_handover test/pbapp.cpp)
target_link_libraries(pbapp_handover PRIVATE libavr)
target_compile_definitions(pbapp_handover PRIVATE PUREBOOT_CLOCK_HZ=${_mute_hz}
PUREBOOT_BAUD=${_mute_baud} PUREBOOT_HANDOVER)
add_custom_command(TARGET pbapp_handover POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_handover> $<TARGET_FILE:pbapp_handover>.bin)
add_test(NAME pureboot.mute
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbmute.py
${PB_DEVICE} $<TARGET_FILE:pureboot_sw_on_usart0> ${PUREBOOT_SIM_MCU} ${_mute_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_mute_baud}
$<TARGET_FILE:pbapp_handover>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbmute-work ${_mute_link})
set_tests_properties(pureboot.mute PROPERTIES TIMEOUT 180)
# The same hand-over against the one-wire deployment on that USART's
# RXD: RXEN forces the shared pin's direction, so a loader that only
# released the transmit-side hold would read the wire and answer into
# a pin it cannot drive. The host runs with the --one-wire echo
# discard, which the bridge's shared-line model feeds for real.
get_target_property(_mute1w_link pureboot_1w_on_usart0 PUREBOOT_LINK)
add_test(NAME pureboot.mute.onewire
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbmute.py
${PB_DEVICE} $<TARGET_FILE:pureboot_1w_on_usart0> ${PUREBOOT_SIM_MCU} ${_mute_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_mute_baud}
$<TARGET_FILE:pbapp_handover>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbmute-1w-work ${_mute1w_link})
set_tests_properties(pureboot.mute.onewire PROPERTIES TIMEOUT 180)
# The full protocol suite over one shared pin: the loader folded onto
# PB0, the bridge following the pin's direction, the fixture
# bannering as a guest on the same line, and the host discarding its
# own echo throughout.
get_target_property(_1w_hz pureboot_1w PUREBOOT_HZ)
get_target_property(_1w_baud pureboot_1w PUREBOOT_BAUD)
get_target_property(_1w_link pureboot_1w PUREBOOT_LINK)
add_executable(pbapp_1w test/pbapp.cpp)
target_link_libraries(pbapp_1w PRIVATE libavr)
target_compile_definitions(pbapp_1w PRIVATE PUREBOOT_CLOCK_HZ=${_1w_hz}
PUREBOOT_BAUD=${_1w_baud} PUREBOOT_SOFT_SERIAL
PUREBOOT_RX=pb0 PUREBOOT_TX=pb0)
add_custom_command(TARGET pbapp_1w POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_1w> $<TARGET_FILE:pbapp_1w>.bin)
add_test(NAME pureboot.onewire
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbtest.py
${PB_DEVICE} $<TARGET_FILE:pureboot_1w> ${PUREBOOT_SIM_MCU} ${_1w_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_1w_baud} ${PUREBOOT_EEPROM}
$<TARGET_FILE:pbapp_1w>.bin ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pb1w-work ${_1w_link})
set_tests_properties(pureboot.onewire PROPERTIES TIMEOUT 180)
# The hardware USART's half-duplex turn-around, end to end: every
# reply byte runs drive-line, TXC-hold, release - against simavr's
# RXEN-gated receiver, which drops input to a disabled receiver the
# way silicon does. The pty is a two-wire transport, so the host
# needs no echo discard here; the off-chip tie itself is the
# hardware bench's item.
add_test(NAME pureboot.halfduplex
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbtest.py
${PB_DEVICE} $<TARGET_FILE:pureboot_hd> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud} ${PUREBOOT_EEPROM}
$<TARGET_FILE:pbapp>.bin ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbhd-work)
set_tests_properties(pureboot.halfduplex PROPERTIES TIMEOUT 180)
endif()
# The second USART, driven for real on one chip: instance selection is
# compile-checked everywhere, but only a live session proves the loader
# initialized and polls the USART it claims to. The fixture application
@@ -397,10 +677,10 @@ if(PROJECT_IS_TOP_LEVEL)
set_tests_properties(pureboot.usart1 PROPERTIES TIMEOUT 180)
endif()
# The autobaud variants driven end to end over the software-UART bridge (both
# under review — pureboot/autobaud.md): the host sends the 0xC0 calibration
# pulse, the loader times it, locks, and programs. Run on the near-flash 328P
# and the word-addressed 1284P the two flash-addressing classes and each
# The autobaud loader driven end to end over the software-UART bridge:
# the host sends the 0xC0 calibration pulse, the loader times it, locks,
# and programs. Run on the near-flash 328P
# and the word-addressed 1284P - the two flash-addressing classes - and each
# at two clocks with the one binary, which is the clock-agnostic property
# autobaud exists for (test/pbautobaud.py). The fixture application banners
# over the same software link at the first clock's rate.
@@ -412,14 +692,51 @@ if(PROJECT_IS_TOP_LEVEL)
add_custom_command(TARGET pbapp_autobaud POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_autobaud> $<TARGET_FILE:pbapp_autobaud>.bin)
foreach(_variant uni)
add_test(NAME pureboot.autobaud_${_variant}
add_test(NAME pureboot.autobaud
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbautobaud.py
${PB_DEVICE} $<TARGET_FILE:pureboot_autobaud_${_variant}> ${PUREBOOT_SIM_MCU}
${PB_DEVICE} $<TARGET_FILE:pureboot_autobaud> ${PUREBOOT_SIM_MCU}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} $<TARGET_FILE:pbapp_autobaud>.bin
1000000 9600 ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbautobaud-${_variant}-work)
set_tests_properties(pureboot.autobaud_${_variant} PROPERTIES TIMEOUT 240)
endforeach()
${CMAKE_BINARY_DIR}/pbautobaud-work)
set_tests_properties(pureboot.autobaud PROPERTIES TIMEOUT 240)
# The tightest deployment in the space, end to end: the autobaud
# loader folded onto the USART's own RXD with the OSCCAL trim baked
# - one-wire calibration, the receive-side release, and the host's
# echo discard, over the same two-clock sweep. One chip carries it;
# the shape is chip-independent.
if(LIBAVR_MCU STREQUAL "atmega328p")
get_target_property(_ab1w_link pureboot_1w_autobaud_osccal_on_usart0 PUREBOOT_LINK)
add_executable(pbapp_autobaud_1w test/pbapp.cpp)
target_link_libraries(pbapp_autobaud_1w PRIVATE libavr)
target_compile_definitions(pbapp_autobaud_1w PRIVATE PUREBOOT_CLOCK_HZ=1000000
PUREBOOT_BAUD=9600 PUREBOOT_SOFT_SERIAL
PUREBOOT_RX=pd0 PUREBOOT_TX=pd0)
add_custom_command(TARGET pbapp_autobaud_1w POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_autobaud_1w> $<TARGET_FILE:pbapp_autobaud_1w>.bin)
add_test(NAME pureboot.autobaud.onewire
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbautobaud.py
${PB_DEVICE} $<TARGET_FILE:pureboot_1w_autobaud_osccal_on_usart0>
${PUREBOOT_SIM_MCU} ${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE}
$<TARGET_FILE:pbapp_autobaud_1w>.bin
1000000 9600 ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbautobaud-1w-work ${_ab1w_link})
set_tests_properties(pureboot.autobaud.onewire PROPERTIES TIMEOUT 240)
endif()
# The autobaud window: the calibration poll budget, at the measured
# 10 cycles a poll (pbwindow.py pins the constant the README's
# seconds arithmetic uses; the budget itself is the clock-free knob).
add_test(NAME pureboot.window.autobaud
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbwindow.py
--device ${PB_DEVICE} --loader $<TARGET_FILE:pureboot_autobaud>
--mcu ${PUREBOOT_SIM_MCU} --hz 1000000
--base ${PUREBOOT_BASE_HEX} --page ${PUREBOOT_PAGE}
--baud 9600 --app $<TARGET_FILE:pbapp_autobaud>.bin
--autobaud-polls 4000000 --link sw
--tool ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
--workdir ${CMAKE_BINARY_DIR}/pbwindow-autobaud-work)
set_tests_properties(pureboot.window.autobaud PROPERTIES TIMEOUT 300)
endif()
endif()

78
ide/README.md Normal file
View File

@@ -0,0 +1,78 @@
# Atmel Studio
`master` carries `bootloader.atsln`, so this branch does too: `ide/bootloader.atsln`
builds the loaders from the same sources Ninja does, to a **byte-identical
`.text`** — the stock 328P pureboot deployment and the `tsb_asm` tier in its
512-byte section (`check-flags.py` below is what holds the flag sets equal, so
the sizes are Ninja's own). CMake remains the build system; the solution is
here so the port opens in Studio as its predecessor did.
## The two projects, and why two
pureboot is a chip × backend × clock × baud matrix — `pureboot_add_loader()`
resolves a deployment into compile definitions — and a `.cppproj` is one binary
at one set of flags, so a project can only ever be one point of it. `pureboot`
is that point: the stock 328P deployment, USART0 at 115200 on a 16 MHz crystal,
an 8-second activation window. `tsb_asm` is the TinySafeBoot tier that occupies
the same 512-byte section `master`'s `tsb` project targeted.
The other three tsb tiers (`tsb_pure`, `tsb_tricks`, `tsb_policy`) are not here.
They differ from `tsb_asm` in their source file, their section size, and — for
`tsb_policy` — two loop flags; nothing about that is a Studio concern, and what
they exist to demonstrate is a size gradient only the CMake size tests measure.
Adding one is a copy of `tsb_asm/tsb_asm.cppproj` in its own directory, with its
name, its GUID, its source path and its `--section-start` changed (`0x7c00` for
the 1 KiB tiers), plus four lines in the solution.
`avrdevice` is a project property, so each project gets its own directory:
Studio builds into `<project dir>/<Configuration>` whatever `OutputDirectory`
says, and two projects sharing a directory would share one object file.
## Debug keeps `-Os`
Both configurations compile at `-Os`; Debug adds only `-gdwarf-4`. The `.text`
is therefore identical in both, which is the point — a loader's section is a
**correctness** bound and not a budget. `-Og` builds this same source to 590 B,
and linking it at `--section-start=.text=0x7e00` on a 32 KiB part puts 78 bytes
past flash end **without a diagnostic**: `rcall`/`rjmp` targets there wrap
modulo flash size, so the image dies right after activation. A debug
configuration that silently produces that is worse than none, and DWARF costs no
flash, so the optimisation level stays where correctness needs it.
## What Studio needs from the machine
libavr from the **submodule**, found at
`$(MSBuildProjectDirectory)\..\..\libavr\include` — correct by construction, and
anchored to the project because a plain relative path resolves against the
generated makefile's directory (the configuration's output directory), not the
project's. There is no `LIBAVR_ROOT` escape hatch: a variable exported in a
shell is invisible to Studio launched from the Start menu, and the failure reads
as a missing `libavr/libavr.hpp` — which is what the submodule answers.
A GCC 16.1 toolchain registered as flavour `avr-g++-16.1.0`, nothing older
reaching `-std=c++26`.
## Generating and gating
One generated file is required before a project will load at all, and one command
per project checks the flags have not drifted (both from libavr's
`tools/atmelstudio/`):
```sh
for name in pureboot tsb_asm; do
python libavr/tools/atmelstudio/componentinfo.py \
"ide/$name/$name.componentinfo.xml" --device ATmega328P
python libavr/tools/atmelstudio/check-flags.py \
--solution ide/bootloader.atsln --project "$name" --target "$name" \
--compile-commands build/atmega328p-generated/compile_commands.json \
--log "build/as-$name.log"
done
```
`--project` because one reference describes one binary; `--target` because
`pureboot.cpp` is compiled by every point of the size matrix and the flags
differ per point, so the basename alone does not name a reference. Release is
what the gate compares — the presets define no debug build, and Debug differs
from Release only in `-gdwarf-4`.
Legacy (the yazoalfa-era submodules) stays on `master`.

28
ide/bootloader.atsln Normal file
View File

@@ -0,0 +1,28 @@
Microsoft Visual Studio Solution File, Format Version 12.00
# Atmel Studio Solution File, Format Version 11.00
VisualStudioVersion = 14.0.23107.0
MinimumVisualStudioVersion = 10.0.40219.1
Project("{E66E83B9-2572-4076-B26E-6BE79FF3018A}") = "pureboot", "pureboot\pureboot.cppproj", "{99067222-32D5-49E3-B4F8-5ABA0F7722B7}"
EndProject
Project("{E66E83B9-2572-4076-B26E-6BE79FF3018A}") = "tsb_asm", "tsb_asm\tsb_asm.cppproj", "{6618D3BE-7EB3-49A2-9113-F128E396FF06}"
EndProject
Global
GlobalSection(SolutionConfigurationPlatforms) = preSolution
Debug|AVR = Debug|AVR
Release|AVR = Release|AVR
EndGlobalSection
GlobalSection(ProjectConfigurationPlatforms) = postSolution
{99067222-32D5-49E3-B4F8-5ABA0F7722B7}.Debug|AVR.ActiveCfg = Debug|AVR
{99067222-32D5-49E3-B4F8-5ABA0F7722B7}.Debug|AVR.Build.0 = Debug|AVR
{99067222-32D5-49E3-B4F8-5ABA0F7722B7}.Release|AVR.ActiveCfg = Release|AVR
{99067222-32D5-49E3-B4F8-5ABA0F7722B7}.Release|AVR.Build.0 = Release|AVR
{6618D3BE-7EB3-49A2-9113-F128E396FF06}.Debug|AVR.ActiveCfg = Debug|AVR
{6618D3BE-7EB3-49A2-9113-F128E396FF06}.Debug|AVR.Build.0 = Debug|AVR
{6618D3BE-7EB3-49A2-9113-F128E396FF06}.Release|AVR.ActiveCfg = Release|AVR
{6618D3BE-7EB3-49A2-9113-F128E396FF06}.Release|AVR.Build.0 = Release|AVR
EndGlobalSection
GlobalSection(SolutionProperties) = preSolution
HideSolutionNode = FALSE
EndGlobalSection
EndGlobal

View File

@@ -0,0 +1,118 @@
<?xml version="1.0" encoding="utf-8"?>
<Project DefaultTargets="Build" xmlns="http://schemas.microsoft.com/developer/msbuild/2003" ToolsVersion="14.0">
<PropertyGroup>
<SchemaVersion>2.0</SchemaVersion>
<ProjectVersion>7.0</ProjectVersion>
<ToolchainName>com.Atmel.AVRGCC8.CPP</ToolchainName>
<ProjectGuid>99067222-32d5-49e3-b4f8-5aba0f7722b7</ProjectGuid>
<avrdevice>ATmega328P</avrdevice>
<avrdeviceseries>none</avrdeviceseries>
<OutputType>Executable</OutputType>
<Language>CPP</Language>
<OutputFileName>$(MSBuildProjectName)</OutputFileName>
<OutputFileExtension>.elf</OutputFileExtension>
<OutputDirectory>$(MSBuildProjectDirectory)\$(Configuration)</OutputDirectory>
<AssemblyName>pureboot</AssemblyName>
<Name>pureboot</Name>
<RootNamespace>pureboot</RootNamespace>
<ToolchainFlavour>avr-g++-16.1.0</ToolchainFlavour>
<KeepTimersRunning>true</KeepTimersRunning>
<OverrideVtor>false</OverrideVtor>
<CacheFlash>true</CacheFlash>
<ProgFlashFromRam>true</ProgFlashFromRam>
<RamSnippetAddress>0x20000000</RamSnippetAddress>
<UncachedRange />
<preserveEEPROM>true</preserveEEPROM>
<OverrideVtorValue>exception_table</OverrideVtorValue>
<BootSegment>2</BootSegment>
<ResetRule>0</ResetRule>
<eraseonlaunchrule>0</eraseonlaunchrule>
<EraseKey />
<AsfFrameworkConfig>
<framework-data xmlns="">
<options />
<configurations />
<files />
<documentation help="" />
<offline-documentation help="" />
<dependencies>
<content-extension eid="atmel.asf" uuidref="Atmel.ASF" version="3.52.0" />
</dependencies>
</framework-data>
</AsfFrameworkConfig>
</PropertyGroup>
<PropertyGroup Condition=" '$(Configuration)' == 'Release' ">
<ToolchainSettings>
<AvrGccCpp>
<avrgcc.common.Device>-mmcu=atmega328p</avrgcc.common.Device>
<avrgcc.common.outputfiles.hex>True</avrgcc.common.outputfiles.hex>
<avrgcc.common.outputfiles.lss>True</avrgcc.common.outputfiles.lss>
<avrgcc.common.outputfiles.eep>True</avrgcc.common.outputfiles.eep>
<avrgcc.common.outputfiles.srec>True</avrgcc.common.outputfiles.srec>
<avrgcc.common.outputfiles.usersignatures>False</avrgcc.common.outputfiles.usersignatures>
<avrgcccpp.compiler.symbols.DefSymbols>
<ListValues>
<Value>NDEBUG</Value>
<Value>PUREBOOT_CLOCK_HZ=16000000</Value>
<Value>PUREBOOT_BAUD=115200</Value>
<Value>PUREBOOT_TIMEOUT=8</Value>
</ListValues>
</avrgcccpp.compiler.symbols.DefSymbols>
<avrgcccpp.compiler.directories.IncludePaths>
<ListValues>
<Value>$(MSBuildProjectDirectory)\..\..\libavr\include</Value>
</ListValues>
</avrgcccpp.compiler.directories.IncludePaths>
<avrgcccpp.compiler.optimization.level>Optimize for size (-Os)</avrgcccpp.compiler.optimization.level>
<avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>
<avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>
<avrgcccpp.compiler.warnings.AllWarnings>True</avrgcccpp.compiler.warnings.AllWarnings>
<avrgcccpp.compiler.miscellaneous.OtherFlags>-std=c++26 -Wextra -Werror -mrelax -fno-exceptions -fno-rtti -fno-threadsafe-statics -fno-ivopts -fira-algorithm=priority -fno-tree-ter -fno-split-wide-types</avrgcccpp.compiler.miscellaneous.OtherFlags>
<avrgcccpp.linker.optimization.GarbageCollectUnusedSections>True</avrgcccpp.linker.optimization.GarbageCollectUnusedSections>
<avrgcccpp.linker.miscellaneous.LinkerFlags>-mrelax -nostartfiles -Wl,--section-start=.text=0x7e00 -Wl,--defsym=pureboot_app=0 -Wl,--pmem-wrap-around=32k</avrgcccpp.linker.miscellaneous.LinkerFlags>
</AvrGccCpp>
</ToolchainSettings>
</PropertyGroup>
<PropertyGroup Condition=" '$(Configuration)' == 'Debug' ">
<ToolchainSettings>
<AvrGccCpp>
<avrgcc.common.Device>-mmcu=atmega328p</avrgcc.common.Device>
<avrgcc.common.outputfiles.hex>True</avrgcc.common.outputfiles.hex>
<avrgcc.common.outputfiles.lss>True</avrgcc.common.outputfiles.lss>
<avrgcc.common.outputfiles.eep>True</avrgcc.common.outputfiles.eep>
<avrgcc.common.outputfiles.srec>True</avrgcc.common.outputfiles.srec>
<avrgcc.common.outputfiles.usersignatures>False</avrgcc.common.outputfiles.usersignatures>
<avrgcccpp.compiler.symbols.DefSymbols>
<ListValues>
<Value>DEBUG</Value>
<Value>PUREBOOT_CLOCK_HZ=16000000</Value>
<Value>PUREBOOT_BAUD=115200</Value>
<Value>PUREBOOT_TIMEOUT=8</Value>
</ListValues>
</avrgcccpp.compiler.symbols.DefSymbols>
<avrgcccpp.compiler.directories.IncludePaths>
<ListValues>
<Value>$(MSBuildProjectDirectory)\..\..\libavr\include</Value>
</ListValues>
</avrgcccpp.compiler.directories.IncludePaths>
<avrgcccpp.compiler.optimization.level>Optimize for size (-Os)</avrgcccpp.compiler.optimization.level>
<avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>
<avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>
<avrgcccpp.compiler.warnings.AllWarnings>True</avrgcccpp.compiler.warnings.AllWarnings>
<avrgcccpp.compiler.miscellaneous.OtherFlags>-std=c++26 -Wextra -Werror -mrelax -fno-exceptions -fno-rtti -fno-threadsafe-statics -fno-ivopts -fira-algorithm=priority -fno-tree-ter -fno-split-wide-types -gdwarf-4</avrgcccpp.compiler.miscellaneous.OtherFlags>
<avrgcccpp.linker.optimization.GarbageCollectUnusedSections>True</avrgcccpp.linker.optimization.GarbageCollectUnusedSections>
<avrgcccpp.linker.miscellaneous.LinkerFlags>-mrelax -nostartfiles -Wl,--section-start=.text=0x7e00 -Wl,--defsym=pureboot_app=0 -Wl,--pmem-wrap-around=32k</avrgcccpp.linker.miscellaneous.LinkerFlags>
</AvrGccCpp>
</ToolchainSettings>
</PropertyGroup>
<ItemGroup>
<Compile Include="..\..\pureboot\pureboot.cpp">
<SubType>compile</SubType>
<Link>pureboot\pureboot.cpp</Link>
</Compile>
</ItemGroup>
<ItemGroup>
<Folder Include="pureboot" />
</ItemGroup>
<Import Project="$(AVRSTUDIO_EXE_PATH)\Vs\Compiler.targets" />
</Project>

112
ide/tsb_asm/tsb_asm.cppproj Normal file
View File

@@ -0,0 +1,112 @@
<?xml version="1.0" encoding="utf-8"?>
<Project DefaultTargets="Build" xmlns="http://schemas.microsoft.com/developer/msbuild/2003" ToolsVersion="14.0">
<PropertyGroup>
<SchemaVersion>2.0</SchemaVersion>
<ProjectVersion>7.0</ProjectVersion>
<ToolchainName>com.Atmel.AVRGCC8.CPP</ToolchainName>
<ProjectGuid>6618d3be-7eb3-49a2-9113-f128e396ff06</ProjectGuid>
<avrdevice>ATmega328P</avrdevice>
<avrdeviceseries>none</avrdeviceseries>
<OutputType>Executable</OutputType>
<Language>CPP</Language>
<OutputFileName>$(MSBuildProjectName)</OutputFileName>
<OutputFileExtension>.elf</OutputFileExtension>
<OutputDirectory>$(MSBuildProjectDirectory)\$(Configuration)</OutputDirectory>
<AssemblyName>tsb_asm</AssemblyName>
<Name>tsb_asm</Name>
<RootNamespace>tsb_asm</RootNamespace>
<ToolchainFlavour>avr-g++-16.1.0</ToolchainFlavour>
<KeepTimersRunning>true</KeepTimersRunning>
<OverrideVtor>false</OverrideVtor>
<CacheFlash>true</CacheFlash>
<ProgFlashFromRam>true</ProgFlashFromRam>
<RamSnippetAddress>0x20000000</RamSnippetAddress>
<UncachedRange />
<preserveEEPROM>true</preserveEEPROM>
<OverrideVtorValue>exception_table</OverrideVtorValue>
<BootSegment>2</BootSegment>
<ResetRule>0</ResetRule>
<eraseonlaunchrule>0</eraseonlaunchrule>
<EraseKey />
<AsfFrameworkConfig>
<framework-data xmlns="">
<options />
<configurations />
<files />
<documentation help="" />
<offline-documentation help="" />
<dependencies>
<content-extension eid="atmel.asf" uuidref="Atmel.ASF" version="3.52.0" />
</dependencies>
</framework-data>
</AsfFrameworkConfig>
</PropertyGroup>
<PropertyGroup Condition=" '$(Configuration)' == 'Release' ">
<ToolchainSettings>
<AvrGccCpp>
<avrgcc.common.Device>-mmcu=atmega328p</avrgcc.common.Device>
<avrgcc.common.outputfiles.hex>True</avrgcc.common.outputfiles.hex>
<avrgcc.common.outputfiles.lss>True</avrgcc.common.outputfiles.lss>
<avrgcc.common.outputfiles.eep>True</avrgcc.common.outputfiles.eep>
<avrgcc.common.outputfiles.srec>True</avrgcc.common.outputfiles.srec>
<avrgcc.common.outputfiles.usersignatures>False</avrgcc.common.outputfiles.usersignatures>
<avrgcccpp.compiler.symbols.DefSymbols>
<ListValues>
<Value>NDEBUG</Value>
</ListValues>
</avrgcccpp.compiler.symbols.DefSymbols>
<avrgcccpp.compiler.directories.IncludePaths>
<ListValues>
<Value>$(MSBuildProjectDirectory)\..\..\libavr\include</Value>
</ListValues>
</avrgcccpp.compiler.directories.IncludePaths>
<avrgcccpp.compiler.optimization.level>Optimize for size (-Os)</avrgcccpp.compiler.optimization.level>
<avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>
<avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>
<avrgcccpp.compiler.warnings.AllWarnings>True</avrgcccpp.compiler.warnings.AllWarnings>
<avrgcccpp.compiler.miscellaneous.OtherFlags>-std=c++26 -Wextra -Werror -mrelax -fno-exceptions -fno-rtti -fno-threadsafe-statics</avrgcccpp.compiler.miscellaneous.OtherFlags>
<avrgcccpp.linker.optimization.GarbageCollectUnusedSections>True</avrgcccpp.linker.optimization.GarbageCollectUnusedSections>
<avrgcccpp.linker.miscellaneous.LinkerFlags>-mrelax -nostartfiles -Wl,--section-start=.text=0x7e00 -Wl,--defsym=tsb_app=0 -Wl,--pmem-wrap-around=32k</avrgcccpp.linker.miscellaneous.LinkerFlags>
</AvrGccCpp>
</ToolchainSettings>
</PropertyGroup>
<PropertyGroup Condition=" '$(Configuration)' == 'Debug' ">
<ToolchainSettings>
<AvrGccCpp>
<avrgcc.common.Device>-mmcu=atmega328p</avrgcc.common.Device>
<avrgcc.common.outputfiles.hex>True</avrgcc.common.outputfiles.hex>
<avrgcc.common.outputfiles.lss>True</avrgcc.common.outputfiles.lss>
<avrgcc.common.outputfiles.eep>True</avrgcc.common.outputfiles.eep>
<avrgcc.common.outputfiles.srec>True</avrgcc.common.outputfiles.srec>
<avrgcc.common.outputfiles.usersignatures>False</avrgcc.common.outputfiles.usersignatures>
<avrgcccpp.compiler.symbols.DefSymbols>
<ListValues>
<Value>DEBUG</Value>
</ListValues>
</avrgcccpp.compiler.symbols.DefSymbols>
<avrgcccpp.compiler.directories.IncludePaths>
<ListValues>
<Value>$(MSBuildProjectDirectory)\..\..\libavr\include</Value>
</ListValues>
</avrgcccpp.compiler.directories.IncludePaths>
<avrgcccpp.compiler.optimization.level>Optimize for size (-Os)</avrgcccpp.compiler.optimization.level>
<avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareFunctionsForGarbageCollection>
<avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>True</avrgcccpp.compiler.optimization.PrepareDataForGarbageCollection>
<avrgcccpp.compiler.warnings.AllWarnings>True</avrgcccpp.compiler.warnings.AllWarnings>
<avrgcccpp.compiler.miscellaneous.OtherFlags>-std=c++26 -Wextra -Werror -mrelax -fno-exceptions -fno-rtti -fno-threadsafe-statics -gdwarf-4</avrgcccpp.compiler.miscellaneous.OtherFlags>
<avrgcccpp.linker.optimization.GarbageCollectUnusedSections>True</avrgcccpp.linker.optimization.GarbageCollectUnusedSections>
<avrgcccpp.linker.miscellaneous.LinkerFlags>-mrelax -nostartfiles -Wl,--section-start=.text=0x7e00 -Wl,--defsym=tsb_app=0 -Wl,--pmem-wrap-around=32k</avrgcccpp.linker.miscellaneous.LinkerFlags>
</AvrGccCpp>
</ToolchainSettings>
</PropertyGroup>
<ItemGroup>
<Compile Include="..\..\tsb\tsb_asm.cpp">
<SubType>compile</SubType>
<Link>tsb\tsb_asm.cpp</Link>
</Compile>
</ItemGroup>
<ItemGroup>
<Folder Include="tsb" />
</ItemGroup>
<Import Project="$(AVRSTUDIO_EXE_PATH)\Vs\Compiler.targets" />
</Project>

2
libavr

Submodule libavr updated: edc77ca43f...1412b7173b

View File

@@ -38,11 +38,23 @@ avra -I /usr/share/avra tsb-fixedbaud.asm # after uncommenting .include "m328P
```
**500 bytes with every feature** — the proof that ≤512 B and full feature parity
are simultaneously reachable. The port's `tsb_asm` tier meets the same bar at
510 B in the same 512 B section, written in C++ on libavr except the two
routines whose remaining cost is the calling convention itself (the bounded rx
and the page-store loop); `tsb_tricks` needs no assembly at all at 526 B, and
`tsb_pure` stays fully idiomatic at 836 B, both in the 1 KB section.
are simultaneously reachable. The port's four tiers reach it from the other
side, and the gradient between them is the cost of the mechanisms each is
allowed:
| tier | bytes | section | what it is allowed |
|---|---|---|---|
| oracle | 500 | 512 B | hand-written assembly, the reference |
| `tsb_asm` | 512 | 512 B | C++ on libavr, two routines in asm |
| `tsb_tricks` | 528 | 1 KB | no asm; global register variables |
| `tsb_policy` | 630 | 1 KB | pureboot's rules: no asm, no register variables |
| `tsb_pure` | 776 | 1 KB | idiomatic libavr throughout |
The two routines `tsb_asm` keeps are the ones whose remaining cost is the
calling convention itself: the bounded rx and the page-store loop. It fills
its section exactly, with the same one-bit-time turn-around guard the oracle
spends six bytes on - every tier implements the whole feature set, which is
what makes the column a gradient rather than four different loaders.
The oracle targets 20 MHz / 33333 baud; the port targets 16 MHz / 115200 baud
(what the simavr protocol test drives). Baud and geometry differ, code size and

View File

@@ -1,5 +1,5 @@
# pureboot as a consumable CMake unit: the per-chip geometry, the default baud
# ladder, and pureboot_add_loader() the one way a loader target is created.
# ladder, and pureboot_add_loader() - the one way a loader target is created.
# A downstream project brings its usual libavr setup (the `libavr` target and
# the LIBAVR_MCU toolchain preset), adds this directory, and states its
# deployment; every argument is optional (README.md):
@@ -9,7 +9,7 @@
# Per-family geometry, deployment defaults, and the linker wrap the PC modulo
# needs. The slot is 512 bytes on every chip. The USART flags mirror the
# hardware inventory the loader's own static asserts check the plain 644 is
# hardware inventory the loader's own static asserts check - the plain 644 is
# the x4 family's one single-USART die (Atmel-2593).
set(_pb_has_usart 1)
set(_pb_has_usart1 0)
@@ -107,7 +107,7 @@ set(_pb_slot 512)
math(EXPR _pb_base "${_pb_flash} - ${_pb_slot}")
math(EXPR _pb_base_hex "${_pb_base}" OUTPUT_FORMAT HEXADECIMAL)
# Patched-vector chips hand over through the trampoline word below the slot,
# which is also the slot's own last word their budget is slot 2.
# which is also the slot's own last word - their budget is slot - 2.
if(LIBAVR_MCU MATCHES "^atmega" AND NOT LIBAVR_MCU MATCHES "^atmega48")
set(_pb_app 0)
set(_pb_limit ${_pb_slot})
@@ -116,6 +116,17 @@ else()
math(EXPR _pb_limit "${_pb_slot} - 2")
endif()
# The pins each USART owns. A bit-banged link deployed on them has to release
# that USART before it can drive the line, and those instructions are the one
# way the choice of pins moves the image - so a size matrix needs them as an
# axis even though pins are otherwise immediate operands. Uniform across every
# mega libavr covers: USART0 (the classics' un-numbered USART included) on
# PD0/PD1, USART1 on PD2/PD3.
set(_pb_usart0_rx pd0)
set(_pb_usart0_tx pd1)
set(_pb_usart1_rx pd2)
set(_pb_usart1_tx pd3)
# simavr names its cores after the base dies; the A revisions run on them
# (the 644PA on the 644P core).
set(_pb_sim_mcu ${LIBAVR_MCU})
@@ -125,6 +136,28 @@ elseif(LIBAVR_MCU STREQUAL "atmega644pa")
set(_pb_sim_mcu atmega644p)
endif()
# Where SRAM begins: the classic megas keep it right after the plain I/O
# registers, the x8/x4 generations push it past their extended I/O file, and
# the tinies match the classics. An autobaud loader keeps its measured unit
# in GPIOR2:GPIOR1 wherever the chip has the pair (data 0x32 on the
# t25/45/85, 0x4A from the x8 generation on) and as the first RAM object at
# SRAM start where it does not (the t13s and classic megas). The host reads
# whichever home applies (pureboot.py's geometry), and the unit-position
# test holds the image to the same split.
if(LIBAVR_MCU MATCHES "^atmega(8|16|32)a?$")
set(_pb_ram 0x60)
set(_pb_unit_gpior "")
elseif(LIBAVR_MCU MATCHES "^atmega")
set(_pb_ram 0x100)
set(_pb_unit_gpior 0x4A)
elseif(LIBAVR_MCU MATCHES "^attiny13")
set(_pb_ram 0x60)
set(_pb_unit_gpior "")
else()
set(_pb_ram 0x60)
set(_pb_unit_gpior 0x32)
endif()
# The function runs in its caller's scope, so everything it needs crosses
# scopes as global properties.
set_property(GLOBAL PROPERTY PUREBOOT_BASE_HEX ${_pb_base_hex})
@@ -133,6 +166,10 @@ set_property(GLOBAL PROPERTY PUREBOOT_WRAP "${_pb_wrap}")
set_property(GLOBAL PROPERTY PUREBOOT_DEFAULT_HZ ${_pb_hz})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART ${_pb_has_usart})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART1 ${_pb_has_usart1})
set_property(GLOBAL PROPERTY PUREBOOT_USART0_RX ${_pb_usart0_rx})
set_property(GLOBAL PROPERTY PUREBOOT_USART0_TX ${_pb_usart0_tx})
set_property(GLOBAL PROPERTY PUREBOOT_USART1_RX ${_pb_usart1_rx})
set_property(GLOBAL PROPERTY PUREBOOT_USART1_TX ${_pb_usart1_tx})
# The port's own build (tests, the size matrix) reads the geometry from the
# parent scope; a downstream consumer gets the same variables for free.
@@ -142,9 +179,15 @@ set(PUREBOOT_SLOT ${_pb_slot} PARENT_SCOPE)
set(PUREBOOT_LIMIT ${_pb_limit} PARENT_SCOPE)
set(PUREBOOT_EEPROM ${_pb_eeprom} PARENT_SCOPE)
set(PUREBOOT_DEFAULT_HZ ${_pb_hz} PARENT_SCOPE)
set(PUREBOOT_RAM_START ${_pb_ram} PARENT_SCOPE)
set(PUREBOOT_UNIT_GPIOR "${_pb_unit_gpior}" PARENT_SCOPE)
set(PUREBOOT_HAS_USART ${_pb_has_usart} PARENT_SCOPE)
set(PUREBOOT_HAS_USART1 ${_pb_has_usart1} PARENT_SCOPE)
set(PUREBOOT_SIM_MCU ${_pb_sim_mcu} PARENT_SCOPE)
set(PUREBOOT_USART0_RX ${_pb_usart0_rx} PARENT_SCOPE)
set(PUREBOOT_USART0_TX ${_pb_usart0_tx} PARENT_SCOPE)
set(PUREBOOT_USART1_RX ${_pb_usart1_rx} PARENT_SCOPE)
set(PUREBOOT_USART1_TX ${_pb_usart1_tx} PARENT_SCOPE)
# The rates a default may pick, fastest first.
set_property(GLOBAL PROPERTY PUREBOOT_BAUD_LADDER 115200 57600 38400 19200 9600)
@@ -190,19 +233,39 @@ function(pureboot_default_baud clock software outvar)
endif()
endforeach()
message(FATAL_ERROR "pureboot: no standard baud rate fits a ${clock} Hz clock within 2.5 % "
" pass BAUD <rate> to deploy a non-standard one")
" - pass BAUD <rate> to deploy a non-standard one")
endfunction()
# pureboot_add_loader(<name> [CLOCK <hz>] [BAUD <bd>]
# [SERIAL auto|hardware|software] [USART <n>]
# [RX <pin>] [TX <pin>] [TIMEOUT <s>])
# [SERIAL auto|hardware|software|autobaud] [USART <n>]
# [RX <pin>] [TX <pin>] [TIMEOUT <s>] [OSCCAL <byte>]
# [HALF_DUPLEX])
#
# The loader target plus its flashable images (<name>.hex for a programmer,
# <name>.bin for --update-loader). The resolved deployment is stamped on the
# target as PUREBOOT_HZ / PUREBOOT_BAUD / PUREBOOT_LINK (the link spelled
# usart0, usart1 or sw:<RX>,<TX>) — what a test harness speaks to it with.
# usart0, usart1, or sw:<RX>,<TX> with a trailing @<n> where those pins are a
# USART's own) - what a test harness speaks to it with.
#
# HALF_DUPLEX is the one-wire deployment, per backend: on the hardware USART
# it enables the library's .half_duplex turn-around (RXD and TXD tied
# together off-chip); on a software or autobaud link it puts both directions
# on the RX pin - the same thing RX == TX spells directly.
#
# SERIAL autobaud measures the host's bit timing at run time, so the image
# carries no clock and no baud: CLOCK and BAUD are not build parameters there,
# and one binary per chip serves every F_CPU and every rate. The stamped
# PUREBOOT_HZ/PUREBOOT_BAUD then record what a harness should *drive* it at,
# not what it was built for.
#
# OSCCAL bakes a measured oscillator trim into the loader (README.md: the
# RC-oscillator deployment answer): the byte is written at the top of run(),
# so every reset path - the watchdog hand-over included - runs on the
# corrected clock. Orthogonal to the backend: an autobaud build may carry it
# purely for the application's benefit, its own link being clock-free. No
# value, no code.
function(pureboot_add_loader name)
cmake_parse_arguments(PB "" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT" "" ${ARGN})
cmake_parse_arguments(PB "HALF_DUPLEX" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT;OSCCAL" "" ${ARGN})
if(PB_UNPARSED_ARGUMENTS)
message(FATAL_ERROR "pureboot_add_loader(${name}): unknown arguments ${PB_UNPARSED_ARGUMENTS}")
endif()
@@ -222,8 +285,8 @@ function(pureboot_add_loader name)
if(NOT PB_SERIAL)
set(PB_SERIAL auto)
endif()
if(DEFINED PB_USART AND PB_SERIAL STREQUAL "software")
message(FATAL_ERROR "pureboot_add_loader(${name}): USART ${PB_USART} contradicts SERIAL software")
if(DEFINED PB_USART AND NOT PB_SERIAL MATCHES "^(auto|hardware)$")
message(FATAL_ERROR "pureboot_add_loader(${name}): USART ${PB_USART} contradicts SERIAL ${PB_SERIAL}")
endif()
if(DEFINED PB_USART)
set(PB_SERIAL hardware)
@@ -239,23 +302,38 @@ function(pureboot_add_loader name)
message(FATAL_ERROR "pureboot_add_loader(${name}): ${LIBAVR_MCU} has no hardware USART")
endif()
set(_serial_defines PUREBOOT_USART=${PB_USART})
if(PB_HALF_DUPLEX)
list(APPEND _serial_defines PUREBOOT_HALF_DUPLEX)
endif()
set(_link usart${PB_USART})
else()
if(PB_SERIAL STREQUAL "auto")
if(_usart AND (PB_RX OR PB_TX))
message(WARNING "pureboot_add_loader(${name}): RX/TX apply to the software UART, "
"which auto does not pick on ${LIBAVR_MCU} SERIAL software to force it")
"which auto does not pick on ${LIBAVR_MCU} - SERIAL software to force it")
endif()
if(_usart)
set(_link usart0)
if(PB_HALF_DUPLEX)
set(_serial_defines PUREBOOT_HALF_DUPLEX)
endif()
else()
set(PB_SERIAL software)
endif()
endif()
if(PB_SERIAL STREQUAL "software")
if(PB_SERIAL MATCHES "^(software|autobaud)$")
if(NOT PB_RX)
set(PB_RX pb0)
endif()
if(PB_HALF_DUPLEX)
# One-wire: both directions on the RX pin. RX == TX spells
# the same deployment directly.
if(PB_TX AND NOT PB_TX STREQUAL PB_RX)
message(FATAL_ERROR "pureboot_add_loader(${name}): HALF_DUPLEX puts both "
"directions on RX (${PB_RX}); TX ${PB_TX} contradicts it")
endif()
set(PB_TX ${PB_RX})
endif()
if(NOT PB_TX)
set(PB_TX pb1)
endif()
@@ -264,12 +342,35 @@ function(pureboot_add_loader name)
message(FATAL_ERROR "pureboot_add_loader(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
if(PB_SERIAL STREQUAL "autobaud")
set(_serial_defines PUREBOOT_AUTOBAUD PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
else()
set(_serial_defines PUREBOOT_SOFT_SERIAL PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
# sw:<RX>,<TX> as port letter and bit, upcased.
endif()
# sw:<RX>,<TX> as port letter and bit, upcased - with @<n> where
# the TX pin is a USART's own TXD, since a harness driving that
# link has to know the USART owns the pin until the loader
# releases it.
string(SUBSTRING ${PB_RX} 1 2 _rx_pin)
string(SUBSTRING ${PB_TX} 1 2 _tx_pin)
string(TOUPPER "sw:${_rx_pin},${_tx_pin}" _link)
string(REPLACE "SW" "sw" _link ${_link})
get_property(_tx0 GLOBAL PROPERTY PUREBOOT_USART0_TX)
get_property(_tx1 GLOBAL PROPERTY PUREBOOT_USART1_TX)
get_property(_rx0 GLOBAL PROPERTY PUREBOOT_USART0_RX)
get_property(_rx1 GLOBAL PROPERTY PUREBOOT_USART1_RX)
if(_usart AND PB_TX STREQUAL _tx0)
set(_link "${_link}@0")
elseif(_usart1 AND PB_TX STREQUAL _tx1)
set(_link "${_link}@1")
elseif(PB_TX STREQUAL PB_RX AND _usart AND PB_RX STREQUAL _rx0)
# One-wire on a USART's RXD: RXEN forces that pin's direction
# (section 20.7.3), so the driven shared pin is held exactly like a
# TXD - the harness models the hold either way.
set(_link "${_link}@0")
elseif(PB_TX STREQUAL PB_RX AND _usart1 AND PB_RX STREQUAL _rx1)
set(_link "${_link}@1")
endif()
endif()
endif()
if(NOT PB_BAUD)
@@ -280,23 +381,45 @@ function(pureboot_add_loader name)
endif()
endif()
if(PB_SERIAL STREQUAL "autobaud")
# No clock and no baud reach the image; the window is a poll budget.
set(_defines ${_serial_defines})
else()
set(_defines PUREBOOT_CLOCK_HZ=${PB_CLOCK} PUREBOOT_BAUD=${PB_BAUD} PUREBOOT_TIMEOUT=${PB_TIMEOUT}
${_serial_defines})
endif()
if(DEFINED PB_OSCCAL)
math(EXPR _osccal "${PB_OSCCAL}" OUTPUT_FORMAT DECIMAL)
if(_osccal LESS 0 OR _osccal GREATER 255)
message(FATAL_ERROR "pureboot_add_loader(${name}): OSCCAL ${PB_OSCCAL} is not one byte")
endif()
list(APPEND _defines PUREBOOT_OSCCAL=${_osccal})
endif()
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/pureboot.cpp)
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE ${_defines})
# Codegen shaping for the loader TU only, worth 1436 B depending on the
# chip. At -Os GCC otherwise rewrites the byte-stream loops' counters into
# end-pointer forms that cost registers (-fno-ivopts,
# -fno-split-wide-types), leaves register pressure on the table with the
# default allocator (-fira-algorithm=priority), and keeps loop-invariant
# immediates and expression temporaries in registers
# (-fno-move-loop-invariants, -fno-tree-ter) — but every loop body here
# contains a call, so a register held across it costs more than the
# load-immediate it saves.
# Codegen shaping for the loader TU only. At -Os GCC otherwise rewrites the
# byte-stream loops' counters into end-pointer forms that cost registers
# (-fno-ivopts, -fno-split-wide-types), leaves register pressure on the
# table with the default allocator (-fira-algorithm=priority), and keeps
# expression temporaries in registers (-fno-tree-ter) - but every loop body
# here contains a call, so a register held across it costs more than the
# load-immediate it saves. The set is fitted to the loader's body and has to
# be re-measured when that body changes: -fno-move-loop-invariants belonged
# here while the command loop carried four transfer bodies and costs bytes
# now that it carries one, and -fno-ivopts is fitted per backend - an
# autobaud body needs ivopts to keep the calibration countdown a single
# induction variable (without it the counter is duplicated and the
# measurement loop runs 9 cycles instead of its contracted 7), while the
# fixed-baud bodies still measure smaller with it off.
if(PB_SERIAL STREQUAL "autobaud")
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
-fira-algorithm=priority -fno-tree-ter -fno-split-wide-types)
else()
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-tree-ter -fno-split-wide-types)
endif()
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
@@ -311,37 +434,3 @@ function(pureboot_add_loader name)
PUREBOOT_LINK ${_link})
endfunction()
# pureboot_add_autobaud(<name> <source> [RX <pin>] [TX <pin>])
#
# An autobaud software-serial loader from <source> (pureboot_autobaud_*.cpp).
# Autobaud measures the host's bit timing at runtime, so the image carries no
# clock and no baud — one binary per chip runs at any F_CPU. Same per-chip
# geometry, link and codegen flags as pureboot_add_loader(); only the clock and
# baud axes fall away. Two source files are under review (autobaud.md):
# pureboot_autobaud_pure.cpp and pureboot_autobaud_reg.cpp.
function(pureboot_add_autobaud name source)
cmake_parse_arguments(PB "" "RX;TX" "" ${ARGN})
if(NOT PB_RX)
set(PB_RX pb0)
endif()
if(NOT PB_TX)
set(PB_TX pb1)
endif()
foreach(_pin ${PB_RX} ${PB_TX})
if(NOT _pin MATCHES "^p[a-h][0-7]$")
message(FATAL_ERROR "pureboot_add_autobaud(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
get_property(_base_hex GLOBAL PROPERTY PUREBOOT_BASE_HEX)
get_property(_app GLOBAL PROPERTY PUREBOOT_APP)
get_property(_wrap GLOBAL PROPERTY PUREBOOT_WRAP)
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/${source})
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
endfunction()

View File

@@ -2,52 +2,67 @@
A serial bootloader on [libavr](https://git.blackmark.me/avr/libavr), pure by
constraint: one C++ source, no inline assembly, no global register variables
(attributes and compiler flags allowed), **512 bytes on every chip libavr
targets — all 37**. The device speaks primitives; every composite — verify,
(attributes and compiler flags allowed), **a 512-byte slot on every chip
libavr targets — all 37**. The device speaks primitives; every composite — verify,
erase, reset-vector surgery, updating the loader itself — lives in the host
tool (`pureboot.py`).
The image is **position-independent**: control flow is PC-relative, the
read/write paths take wire addresses, the write guard protects the slot the
code is *running* in (from the runtime return address), the info block is
addressed from that same anchor, and the application jump is an indirect call
to an absolute entry. The identical binary therefore runs from any slot with
every command intact, which makes pureboot **its own staging loader**: the
host installs the same binary one slot below the resident, jumps into it, and
lets it rewrite the resident.
transfer paths take wire addresses, nothing is flash-resident to address at
all, and the application jump is an indirect call to an absolute entry. It does
not know which slot it occupies and does not need to. The identical binary
therefore runs from any slot with
every command intact, which makes pureboot **its own staging loader**: the host
installs the same binary one slot below the resident, jumps into it, and lets
it rewrite the resident. The lint holds it to that literally — the image must
come out byte-identical linked at a different base.
## Chips
Sizes are the default configuration: the hardware USART0 at 115200 8N1 on a
16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at 57600 8N1 on
the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above). Every axis moves
per build — see *Configuration*; the largest image any of them produces is a
software UART at a slow baud, which on the 1284s is 494 B, the tightest fit in
the whole matrix at 18 B spare.
The Stock column is the default configuration: the hardware USART0 at 115200
8N1 on a 16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at
57600 8N1 on the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above).
Every axis moves per build — see *Configuration*. The Autobaud column is the
worst configuration the space produces for the chip: the clock-free build —
it alone carries the calibration machinery — with the `OSCCAL` trim baked
and, where the chip has a USART, the link deployed on that USART's own pins,
which the loader then has to release (*Pin ownership*). Folding the same
build onto a single pin (*One-wire*) measures identically on every chip, so
the column covers that twin too. On default pins without the trim the same
loaders run 410 B smaller.
| Chip | Flash | Loader at | Link | Size |
|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 416 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 420 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 424 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 424 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 396 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 400 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 400 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 414 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 434 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 432 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 478 B |
| Chip | Flash | Loader at | Link | Stock | Autobaud |
|---|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 372 B | 460 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 376 B | 452 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 376 B | 452 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 376 B | 452 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 350 B | 480 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 352 B | 484 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 352 B | 484 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 366 B | 454 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 376 B | 464 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 378 B | 468 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 378 B | 468 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 378 B | 468 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 378 B | 468 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 372 B | 462 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 390 B | 480 B |
† No hardware boot section: the host patches the reset vector, and the budget
is 510 bytes, since the slot's last word is the trampoline.
† No hardware boot section: the host patches the reset vector, and an image's
budget is 510 bytes. The last word of the **lower** slot belongs to the host —
it is the trampoline holding the application's relocated reset vector, and it
is a staging copy's own last word during a self-update — so an image must fit
below it. The resident loader has all 512 bytes of the slot it runs in; 510 is
what an image must fit so that a copy of it staged one slot down leaves that
word alone.
The 1284s are the heaviest because they alone carry the far-flash machinery —
ELPM reads, RAMPZ page commands, a word-addressed wire.
The tightest fit in the whole space is therefore the 1284s' 480 of their
512: they alone carry the far-flash machinery (ELPM reads, RAMPZ page
commands) on top of everything the column already stacks. The flash bank
riding in a transfer's selector byte keeps even those chips' addressing the
same 16-bit form every other chip uses, which is why they are no longer the
outlier they were.
The software UART enables the RX pull-up; TX idles high. All multi-byte wire
quantities are little-endian.
@@ -62,10 +77,12 @@ repo's build and by a downstream project alike:
|---|---|---|
| `CLOCK <hz>` | the clock the board runs | 16 MHz megas, 8 MHz t25/45/85, 9.6 MHz t13s |
| `BAUD <bd>` | the wire rate | the ladder below |
| `SERIAL auto\|hardware\|software` | the link backend | `auto`: the hardware USART where the chip has one |
| `SERIAL auto\|hardware\|software\|autobaud` | the link backend | `auto`: the hardware USART where the chip has one |
| `USART <n>` | the USART instance (x4 megas carry two) | 0 |
| `RX <pin>`, `TX <pin>` | software-UART pins | `pb0`, `pb1` |
| `TIMEOUT <s>` | the activation window | 8 |
| `OSCCAL <byte>` | a measured oscillator trim, applied before anything runs | none — no value, no code |
| `HALF_DUPLEX` | one-wire: both directions on one line (*One-wire* below) | off |
The default baud is the fastest of 115200/57600/38400/19200/9600 the clock
reaches within 2.5 % — the same U2X-included divisor search libavr's baud
@@ -74,15 +91,96 @@ receiver's 100-cycles-a-bit floor. Whatever is picked or overridden is
re-checked in the compile: an infeasible combination, or a USART the chip does
not have, fails with a named static assert.
Putting a bit-banged link on a USART's own pins is a supported deployment, and
the usual one where a board's USB bridge is wired to RXD/TXD: the link's `init`
clears that USART's `UCSRnB` first, because while its `TXEN` is set the USART —
not the port register — owns the TX pin, and a loader entered from an
application that left it enabled would receive and obey while answering nothing
(§20.6.3). It costs one store — four bytes on the extended-I/O chips, two on
the classic megas — and only on those pins.
`SERIAL autobaud` takes neither: the loader **measures** the host's bit timing
at run time, so `CLOCK` and `BAUD` are not build parameters there and one
binary per chip serves every clock and every rate. It is for the deployments
whose clock is not known at build time and does not hold still — the internal
RC oscillator, ±10 % from the factory and moving with supply and temperature —
where a fixed-baud software build has to be rebuilt per clock and still drifts
out of tolerance. The cost is that it is software-serial only (a hardware USART
needs its divisor programmed) and that activation counts poll iterations rather
than seconds, since there is no clock to convert them against
(`PUREBOOT_AUTOBAUD_POLLS`, default 4,000,000). The wait spends nine cycles a
poll (measured, and held by the `pureboot.window.autobaud` gate), so the
default window is 36 M cycles: 4.5 s at 8 MHz, 3.75 s at 9.6 MHz, 36 s at
1 MHz.
**Pick the rate by cycles a bit, and leave the oscillator room.** What the
calibration can measure is bounded by how many clock cycles one bit lasts, so a
rate is only ever sensible relative to the clock. Two different floors matter:
| | cycles a bit |
|---|---|
| the logic's floor — exact clock, simulated | solid to ~36, fails outright by ~31 (`pureboot.autobaud` gates a point here) |
| a factory-trimmed internal RC, measured on an ATtiny13A | reliable at ~118; already locking 1 attempt in 5 by ~59 |
The gap is the oscillator's own jitter, and no exact-clock simulation shows it.
So on an RC part, **budget about 100 cycles a bit** — the same order as the
fixed-baud software receiver's floor — rather than the logic's ~36. Measured
envelope on that ATtiny13A, with an application resident: 9.6 and 4.8 MHz reach
115200, 1.2 MHz reaches 9600, 600 kHz reaches 4800, 128 kHz reaches 2400.
One trap in testing this: on a patched-vector chip an **erased** application
region walks straight back up into the loader, so every expired window opens
another one and the host's retries eventually catch the pulse. That reads as far
more reliable than the same part with an application resident, which gets one
window per reset. Measure with an application in place.
## One-wire
`HALF_DUPLEX` puts both directions on one line — the deployment for a board
with a single spare pin, or a native-UART bootloader's shared-line wiring.
Each backend has its shape:
- **Software and autobaud links** fold onto the RX pin (`RX == TX` spells
the same deployment directly). The pin idles as the receiver's pull-up
input; each transmitted frame takes the pin's direction and hands it back
with the stop bit's level already on the pull-up, so neither flip makes
an edge. This costs nothing: the frame's direction wrap is exactly what
the dropped second-pin init paid, and the tightest image in the space —
the 1284s' autobaud + `OSCCAL` on their USART's RXD — measures the same
480 bytes one-wire as two-wire. On a USART's own pin the release applies
as ever, RXD included: `RXEN` forces that pin's direction (§20.7.3),
which a receive-only link could live with and a driven shared pin cannot.
- **The hardware USART** (`SERIAL hardware`/`auto` + `HALF_DUPLEX`) uses
libavr's `.half_duplex` turn-around — exactly one direction enabled at a
time, each written byte held to transmit-complete before the line can be
released — and needs RXD and TXD tied together off-chip. It costs
+42…50 B over the stock loader (m8 392, m328P 428, 1284P 440 — all far
inside the slot); the activation window is unchanged, its poll merely
runs through the release-line test (18 cycles a poll in bit-addressable
I/O, 22 in extended — measured, and held per chip by
`pureboot.window.halfduplex`).
Host wiring, for an FTDI-style adapter: **adapter TX through ~1 kΩ to the
line, adapter RX and the MCU pin directly on it.** The resistor lets the MCU
win the line while it answers; the price is that the adapter reads back every
byte it transmits. `pureboot.py --one-wire` consumes that echo byte for byte
— a missing echo is reported as the wiring fault it is, and a device reply
that lands between the echoes of the knock (a loader already in session
re-prompts mid-knock) is held for the reader. The knock is the protocol's
one blind multi-byte write, so on real wiring its second byte can be lost to
that collision outright; the tool's knock retries absorb it. Everything else
is ack-paced and cannot collide.
A downstream project brings its usual libavr setup (the `libavr` target, the
chip via the `LIBAVR_MCU` toolchain preset), consumes this directory, and
states its deployment — an ATmega328P on its shipped 1 MHz fuses with the
software UART on hand-picked pins, say:
software UART on hand-picked pins, say. A submodule pins the loader version
(the tags name them; this repo pins its own libavr the same way), where
FetchContent tracks whatever `main` is:
```cmake
FetchContent_Declare(bootloader GIT_REPOSITORY git@git.blackmark.me:avr/bootloader.git GIT_TAG main)
FetchContent_MakeAvailable(bootloader)
add_subdirectory(${bootloader_SOURCE_DIR}/pureboot pureboot)
# git submodule add <forge>/avr/bootloader.git bootloader — or FetchContent
add_subdirectory(bootloader/pureboot pureboot)
pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5)
```
@@ -113,87 +211,261 @@ window; any other byte is discarded and awaited again, so line noise can delay
the loader but never lock it. A window expiring on an idle line boots the
application.
An autobaud build opens differently, because it has to learn the rate before it
can read a byte at all: the host sends the **calibration byte 0xC0** — a start
bit plus six zero data bits form one low pulse of seven bit-times — and the
loader times that pulse into its bit period. A single `p` then activates; the
pulse has already proven a host is present, which the two-byte knock exists to
establish elsewhere. Both waits are bounded, so a stray low pulse with no host
behind it costs one window and then boots the application rather than holding
the loader.
The window is a compile-time constant (`TIMEOUT`, 8 s by default), so the whole
EEPROM belongs to the application — pureboot keeps no state of its own.
Re-timing a deployed loader is a self-update with a re-timed build.
Re-timing a deployed loader is a self-update with a re-timed build. An autobaud
build counts poll iterations instead (`PUREBOOT_AUTOBAUD_POLLS`), there being
no clock to turn into seconds.
### Entering from a running application
A window opens at a reset and nowhere else, which assumes a reset edge exists —
a button, or DTR wired to it. On a board with neither, the way in has to come
from the application, by jumping to the loader base.
Deciding *when* to jump is the application's business and deliberately not
pureboot's: a console command, a held pin, a magic byte, an idle timeout — each
board's answer differs, and none of them belongs in a 512-byte loader. pureboot
offers no knock detector and no consumable header for this. What follows is the
handful of non-obvious mechanics.
Give the linker the base as a symbol rather than casting a literal to a
function pointer, so the address is stated once and in the same units the
loader is linked at:
```cmake
target_link_options(app PRIVATE "LINKER:--defsym=pureboot_loader=0x7e00")
```
```cpp
extern "C" [[noreturn]] void pureboot_loader();
// ... then, with interrupts off and the watchdog disabled:
pureboot_loader();
```
That is the whole jump: an ordinary call to an absolute symbol, which the
linker resolves and `-mrelax` shortens where it can — `call 0x7e00` in four
bytes, or an `rjmp` on a part small enough for one.
It is worth saying what *not* to copy here, because pureboot's own hand-over
(`run_app()`) looks different: it launders its target through a
`[[gnu::noipa]]` indirect call. That is a position-independence measure, and it
belongs to the loader alone — the same image runs from either slot, so it must
never bake an absolute address. An application is linked at a fixed base and
has no such problem; using the indirect form costs two `ldi`s and a helper call
to reach the same place a plain call reaches in four bytes.
Three things must be true before the jump:
- **Interrupts off and the watchdog disabled.** The loader is polled and
vector-less; an ISR landing in it vectors into the application's table.
- **WDRF clear.** pureboot hands straight back to the application on a watchdog
reset (above), and it only *peeks* MCUSR — so a jump arriving with WDRF still
set opens no window at all.
- **Release any peripheral holding the link.** A loader entered by a jump
inherits the application's registers rather than reset values: with `TXEN0`
still set the USART owns TxD, and a bit-banging loader then receives
perfectly and answers into a pin it does not control. Clearing `UCSR0B`
before jumping is the whole fix, and the symptom without it is a loader that
is mute rather than deaf, which reads as a dead board.
## Session
After the knock the loader stays in its command loop until `J` jumps away or
the chip resets. Before reading each command it waits for any pending EEPROM
After the knock the loader stays in its command loop until a jump takes it away
or the chip resets. Before reading each command it waits for any pending EEPROM
write and sends the prompt `+` (0x2b), which is therefore also the previous
command's completion ack. A session is: await `+`, send a command, read its
reply, repeat.
verdict, then its reply, repeat.
On chips whose flash exceeds 64 KiB (the 1284s — info-block flag bit 1) the
`R`/`W` flash addresses are **word** addresses; everywhere else they are byte
addresses (the 644s' 64 KiB is exactly the 16-bit byte space). EEPROM
addresses and all counts are bytes.
There is no invalid opcode: every byte begins a command, and it is the seal
rather than a table of known letters that rejects noise. The knock is the one
thing that must survive being sent into a loader already in session, which is
why both its bytes — `p` and `b` — carry the identify bit: they answer the
identity and consume nothing else.
Addresses are **byte addresses within a 64 KiB bank**, and the bank rides in
the command's selector byte, so no command has to speak word addresses. The jump
is the exception: its address is a word address, because that is what the
hardware's own jump takes — it still carries a selector byte (reserved,
ignored) so its decode is the same three reads as every other command's.
EEPROM and data-space addresses and all counts are bytes.
The loader trusts the host to keep addresses in range: it does not bound them
against the info block. **Gotcha:** a `w` (or `r`) that runs past `E2END` wraps
EEAR is only as wide as the array, so an address past the end truncates onto
low EEPROM and the write silently overwrites it. Keeping writes within the
advertised sizes is the host's job (the shipped tool does); the flash budget
is better spent on features than on re-checking a bound the host already holds.
against the chip. **Gotcha:** a write (or read) that runs past `E2END` wraps
EEAR is only as wide as the array, so an address past the end truncates onto
low EEPROM and the write silently overwrites it. Keeping transfers within the
real sizes is the host's job (the shipped tool does); the flash budget is
better spent on features than on re-checking a bound the host already holds.
| Cmd | Arguments | Reply |
Every command but the identity has **one header shape** — six bytes, the last
of them a seal over the other five:
```text
op8 sel8 addr_lo8 addr_hi8 n8 seal8
seal = op ^ sel ^ addr_lo ^ addr_hi ^ n ^ 0x5A
```
The loader folds those fields and compares before it decodes what the command
is, then **answers the seal**: `+` accepts, `0xD4` refuses and nothing has
happened. The verdict comes before any payload, which is what keeps a refusal
local — a fill or a write burst sends its data only once the header is in, so
a rejected header never leaves the host pushing bytes into a loader that has
gone back to reading commands.
The opcode is bits, not a letter. A transfer is the absence of the other three,
and its direction is the low bit.
| Opcode | | Arguments after the header | Reply |
|---|---|---|---|
| 0x00 | read | none | verdict, then n bytes from the selected space (n = 0 means 256), then `+` |
| 0x01 | write | n data bytes | verdict, then `+` per byte once its write has begun, then `+` |
| 0x04 | fill | one page of data | verdict, then `+` when the page is in |
| 0x08 | jump | none | verdict, then execution continues at the word address |
| 0x20 | identify | *unsealed, one byte on its own* | 4 bytes: the version, then the three signature bytes, then `+` |
The **selector** byte's low nibble names the space and its high nibble carries
the flash bank.
| Space | | |
|---|---|---|
| `b` | — | the 12-byte info block |
| `R` | addr16, n8 | n flash bytes (n = 0 means 256) |
| `W` | addr16 (any address in the page), then one page of data | — (completion = next prompt) |
| `r` | addr16, n8 | n EEPROM bytes (n = 0 means 256) |
| `w` | addr16, n8, then n data bytes | `+` per byte, sent once its write has begun |
| `F` | — | 4 bytes: low fuse, lock, extended fuse, high fuse |
| `J` | word address (16-bit) | `+`, then execution continues there |
| other | — | ignored; the loop re-prompts (send a junk byte, await `+`, to resync) |
| 0 | flash | read-only here; it is written through the fill and the SPM command |
| 1 | EEPROM | |
| 2 | data | SRAM — and with it the register file and every I/O register, which share the data address space on AVR |
| 3 | fuse and lock | index 0..3 in the hardware's own Z order: low, lock, extended, high |
| 4 | SPM | write-only: the byte goes to SPMCSR and fires the instruction at the address |
`W` streams exactly one SPM page (size from the info block) into the buffer,
then erases and programs — except pages inside the 512-byte slot
the loader is *running* in, which are drained and left alone, so a broken host
cannot brick the running copy and a staged copy may rewrite the resident.
The data space is worth more than it looks. pureboot keeps **zero static RAM**
and pushes no register, so at loader entry an application's SRAM is still
whatever the application left there, bar the handful of bytes of return-address
stack — which makes a read over space 2 a post-mortem of a running application,
not just a poke hole. The same address space carries the register file and the
I/O registers, so peripheral state is readable too; reading some of those has
side effects (reading UDR clears its flags), which is the host's business to
know.
The loader never clears the SPM buffer before a fill, so **one `W` may program
Programming a page is therefore a fill to load the buffer, then an SPM command
for the erase, another for the write, and on a boot-sectioned chip a third to
re-enable the RWW section — `0x03`, `0x05` and `0x11`, the SPMCSR encodings
every part pureboot targets shares. The loader carries no page-commit logic of
its own, and the same primitive reaches every other SPM operation, lock bits
included.
An SPM command has **no data phase**: its SPMCSR byte rides the header's count
field, where the seal covers it. That is the whole reason the field is
overloaded — a byte arriving behind the header would arrive after the seal had
been checked, and the one command that cannot be taken back is exactly the one
that must not be decided by an unchecked byte. Setting the lock bits is
therefore an ordinary sealed command (`0x09`) rather than something the loader
refuses: deliberate is expressible, accidental is not reachable.
The SPM store and the SPM instruction must issue within four cycles of each
other (§26.2), which no host can hit across a serial link — so this one
primitive is *fused* rather than being a poke of SPMCSR followed by a poke of
something else. That four-cycle window is the floor on how low-level a
bootloader's primitives can go; it is not a byte-count decision.
Nothing refuses an address, the loader's own slot included. Through pureboot 8
a running-slot write was dropped; the seal replaced that guard, because what
the guard defended against was a wire fault naming an address, and a wire fault
can no longer name one. What it costs is that a host bug aimed at the running
slot now lands. What it buys is that a resident copy can write its own slot —
which is the only route a self-update has on a chip whose boot section *is* the
512-byte slot, where no staged copy can run SPM at all: the resident plants a
primitive in its own spare space and an application-side installer drives it.
A copy that erases the page it is executing from does not come back, so which
page matters; erasing any other page of its own slot it survives. That needs a
page the image does not reach into, which the stock builds have and the biggest
do not: a 378 B loader on a 128-byte-page mega leaves 384..511 entirely free,
while the 480 B autobaud build reaches into it and has none.
The loader never clears the SPM buffer before a fill, so **one fill may program
the wrong bytes, and the host is what fixes it**. The buffer is write-once per
word until cleared, and two things leave words in it: a refused page, and —
where SPM runs from anywhere, the tinies and the m48s — an application that
self-programmed before entering. The next `W` takes those stale words and
clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on the
tinies), so repeating it programs correctly. The host therefore verifies every
page it writes and rewrites what comes back wrong (three retries, then it
self-programmed before entering. The next page write takes those stale words
and clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on
the tinies), so repeating it programs correctly. The host therefore verifies
every page it writes and rewrites what comes back wrong (three retries, then it
stops).
`w` is host-paced: send the next byte only after the previous byte's `+`. `F`
returns the bytes in the hardware's Z order; on a chip without an extended
fuse byte that slot carries no meaning. Fuse *writing* does not exist — SPM
reaches flash and boot lock bits only.
A write is host-paced: send the next byte only after the previous byte's `+`. Fuse
*writing* does not exist — SPM reaches flash and boot lock bits only.
`J` is the one control-transfer primitive: it runs the application (word 0 or
the trampoline word, both known from the info block) and moves between loader
The jump is the one control-transfer primitive: it runs the application (word 0 or
the trampoline word, both derived from the chip) and moves between loader
copies during a self-update. A jump to a slot's base re-enters that copy's own
startup, which must then be knocked afresh.
The info block (`b`):
Identify answers with the loader's identity — its version and the chip's signature —
and nothing else. Everything else the host needs (page size, loader base,
EEPROM size, whether the reset vector must be patched, how many flash banks)
follows from the signature, and the host holds that table; the loader derived
the same facts from its own chip database at build time, so nothing is guessed,
it is simply not sent twice.
| Offset | Content |
|---|---|
| 02 | `'P'`, `'B'`, pureboot version (3) |
| 35 | device signature |
| 6 | SPM page size in bytes (0 means 256) |
| 78 | loader base — application flash ends here (a word address when bit 1 is set) |
| 910 | EEPROM size |
| 11 | bit 0: host must patch the reset vector (no hardware boot section); bit 1: flash wire addresses are word addresses |
An update image, though, is a bare 512-byte slot with no device to ask, and
installing one built for another chip bricks the target. Every loader image
therefore carries a six-byte **stamp**`'P'`, `'B'`, the version, the three
signature bytes — which the loader itself never reads and the host tool refuses
to install a mismatch against.
## Version
The info block's third byte is the **pureboot version** — the loader's one
identity number, and the only way to tell what a deployed loader is. Nothing
else is numbered: the wire protocol has no version, a pureboot version implies
it, and the host tool holds that map. The tool states the window of loader
versions it speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a
version that changes the protocol becomes the new floor there. None has so
far: 1 through 3 speak the identical session. A loader newer than the tool is
refused by name rather than decoded on the assumption that nothing moved.
The identity's first byte is the **pureboot version** — the loader's one identity
number, and the only way to tell what a deployed loader is. Nothing else is
numbered: the wire protocol has no version, a pureboot version implies it, and
the host tool holds that map. The tool states the window of loader versions it
speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a version that
changes the protocol becomes the new floor there. A loader newer than the tool
is refused by name rather than decoded on the assumption that nothing moved.
Two generations exist. **1 through 4** speak one session — a 12-byte info block
from `b`, and a command per memory (`R`/`W` flash, `r`/`w` EEPROM, `F` fuses).
**5** replaced those with the single `G`/`g` pair over selector-named spaces
above; the shipped tool speaks both, choosing on the version it reads, so a
deployed pureboot 4 stays drivable and self-updatable to 5. **6** changes
nothing on the wire: it marks the builds that may carry a baked `OSCCAL` trim
(Configuration), so a tool driving an update knows such images exist. **7**
moves `J` onto the unified decode — it gains the selector byte the table
shows, which older loaders do not read, so the tool sends each form to the
version that speaks it — and re-homes the autobaud unit into the GPIOR pair
on the chips that have one (Session: what must not be written), which is
where `--info`'s measured clock now reads it on those parts. **8** changes
nothing on the wire either: it marks the builds whose deployment may be
one-wire (*One-wire* above) — the hardware USART's half-duplex turn-around,
or a software link folded onto a single pin. The host-side trace is
`--one-wire`, the echo discard a shared line requires of any tool driving
it. **9** is the third wire change and the largest: bit opcodes in place of
the letters, one sealed header shape for every command, and a verdict on that
seal before the command runs (*Session* above). It also drops the
running-slot write guard, which the seal makes redundant and which was the
only thing standing between a resident copy and its own slot. Identify is
answered by both knock bytes so version discovery works before the version is
known, which is what keeps a deployed pureboot 8 drivable and self-updatable
to 9.
Every closed generation is tagged in this repo at its era's last commit — the
commit just before the next version bump, so a tag holds everything its
version ever gained — and each tag carries the `libavr/` submodule pinned to
the libavr that loader was built against, as the whole libavr era does commit
by commit. `git checkout v3 && git submodule update --init libavr` followed by
the usual preset build therefore reproduces the v3 loader exactly; the open
generation is `main`.
Collapsing four command bodies into one transfer loop is what paid for the
version: the data space, the host-issued SPM operations and the fuses now share
the loop, the cursor and the argument decode that `R`/`r`/`w` each carried a
copy of. The loader shrank while gaining all three.
The tool carries its own version, free to drift; `--version` prints it and the
window.
@@ -223,7 +495,7 @@ ATmega328P profiles (addresses for its 32 KiB):
| BOOTSZ | BOOTRST | Behavior |
|---|---|---|
| 256 words (512 B) | programmed | *Standalone*: reset always enters the loader; **self-update impossible** (the staging slot lies outside the boot section, where SPM is disabled). |
| 512 words (1 KB) | unprogrammed | *Self-update, app-first*: reset always boots the application, which owns all 31.5 KB and must offer its own jump to 0x7e00 to reach the loader (a virgin chip reaches it by reset across erased flash). Updates are power-fail-safe except mid-rewrite of the resident slot itself (no reset path leads to the staging copy then). |
| 512 words (1 KB) | unprogrammed | *Self-update, app-first*: reset always boots the application, which owns all 31.5 KB and must offer its own jump to 0x7e00 to reach the loader (a virgin chip reaches it by reset across erased flash; *Entering from a running application* under Activation is how that jump is written). Updates are power-fail-safe except mid-rewrite of the resident slot itself (no reset path leads to the staging copy then). |
| 512 words (1 KB) | programmed | *Self-update, loader-first*: reset lands at 0x7c00 — the staging slot, normally erased, so execution walks up into the loader; during an update it is the staging copy itself, so a mid-rewrite power loss recovers by reset. The loss windows move to the staging install/retire page writes instead (page-write scale). The host keeps `[0x7c00, 0x7e00)` clear of application data (`--force` overrides). |
Applications are flashed unmodified here — word 0 stays the application's own
@@ -245,6 +517,18 @@ mega (SPM only executes from the boot section — reflash the .hex), but *runs*
on a patched-vector chip, and the ordinary `--update-loader` flow re-homes it
into the top slot from there (`pureboot.rehome`).
**Fixed-baud on an internal RC oscillator is a deployment risk the build
cannot see.** The factory trim is ±10 % where an 8N1 frame survives about
±4: a part at the edge answers nothing at the built rate, and the symptom —
silence — reads as a wiring fault (a real ATtiny13A measured 5.5 %, outside
every standard rate at its own documented default). The **autobaud build is
the deployment-proof backend**: it has no rate to miss. Where fixed-baud on
RC is wanted anyway, measure first and bake the trim: an autobaud session's
`--info` prints the part's true clock from the loader's own measured bit
period, OSCCAL moves the oscillator about 1 % per step, and `OSCCAL <byte>`
builds the correction in — one buildmeasure iteration converges. A loader
already deployed and silent is diagnosed with `--scan` (Host tool).
## Updating the loader
`pureboot.py --update-loader new_pureboot.bin` replaces the resident loader
@@ -252,11 +536,32 @@ with any pureboot build — a re-timed window, a newer version — using the
loader itself as its own staging loader. The image is the loader's own 512
bytes as a raw binary, or the Intel HEX the build emits beside it.
The preflight refuses an image built for another chip: the info block embedded
in every pureboot binary (signature, page size, loader base, EEPROM size,
flags) must match the device's own, and the error names both. Die revisions
share their base signature and geometry, so their images are interchangeable —
as the silicon is.
One thing the image cannot tell the host: **which link it speaks.** The update
works by entering copies of the *new* image (steps 3 and 4 below), so a build
made for another baud or another backend answers on that one and not on the
session's — and 512 bytes of position-independent code carry no header to read
it from. Where the new image's link differs, name it:
```sh
# a 57600 fixed-baud resident, replaced by an autobaud build
pureboot.py --port … --baud 57600 --update-loader ab.bin --staged-autobaud
# …or by a 38400 build of the same backend
pureboot.py --port … --baud 57600 --update-loader sw38400.bin --staged-baud 38400
```
The host retunes on the open port, so no DTR pulse resets the copy it is talking
to. Omit them against a changed link and the update stops after installing the
staging copy, saying so and naming this as the cause.
An `OSCCAL`-baked image is a link change in effect even at an unchanged rate
on paper: the staging copy shifts the physical clock the moment its `run()`
starts, and from then on speaks exactly what it was built for. Declare it
like any other link change — `--staged-baud` with the new build's rate.
The preflight refuses an image built for another chip: the stamp every pureboot
binary carries must resolve to the device's own geometry, and the error names
both. Die revisions share their base signature and geometry, so their images
are interchangeable — as the silicon is.
1. The staging slot `[base512, base)` is saved to a host-side state file (on
the 1 KB tiny13s that is the whole application, vectors included).
@@ -264,7 +569,7 @@ as the silicon is.
the host composes the slot's last word as a jump to the resident base, so
even an abandoned staging copy times out into a loader. A loader already
sitting whole in the staging slot is left as the staging copy instead —
rewriting it would only meet its own running-slot guard.
rewriting it in place would be a copy overwriting itself as it runs.
3. `J` enters the staging copy, which rewrites the resident slot. Where a
patched reset vector routes through the resident, the host first re-aims
word 0 at the staging copy, so a power loss mid-rewrite still resets into a
@@ -273,10 +578,15 @@ as the silicon is.
content, and the state file is discarded.
Every phase is idempotent and keyed off the actual flash state, so re-running
the same command after any interruption resumes and completes. The state file
carries the only bytes not recoverable from the device; losing it mid-update
still completes the update, and the staging region comes back by reflashing
the application. A boot-sectioned mega needs its fuses for the preflight — read
the same command after any interruption resumes and completes — with one
qualification, which is the link again: from step 2 on, the copy the re-run has
to reach is the *new* image, so a resumed run needs the same `--staged-*` as the
first one. On a patched-vector part step 3 also re-aims word 0 at the staging
copy, so after that point a reset reaches the new image's link and **only** that
one; a re-run on the resident's link finds nothing at all. The state file carries
the only bytes not recoverable from the device; losing it mid-update still
completes the update, and the staging region comes back by reflashing the
application. A boot-sectioned mega needs its fuses for the preflight — read
from the device, or supplied with `--assume-fuses` where reading is impossible
(simulators).
@@ -293,25 +603,73 @@ to reset gets its reset pulse and opens the activation window by itself.
--info --fuses --flash app.hex
Operations run in a fixed order within one session: info, fuses, loader
update, flash (erase / program / read / verify), EEPROM (the same) then the
loader hands over to the application. `--stay` keeps the session alive
update, flash (erase / program / read / verify), EEPROM (the same), then
`--poke` and `--peek` in that order, so one invocation writes and reads the
write back — then the loader hands over to the application. `--stay` keeps the session alive
instead, and a later invocation reconnects into it. `--flash` and `--eeprom`
verify by read-back unless `--no-verify`, and a flash page that reads back
wrong is rewritten up to three times before the run stops (see `W` above).
wrong is rewritten up to three times before the run stops (see the fill above).
`--verify-flash` only reports. Images are raw binary, or Intel HEX by
extension. `--force` overrides the refusable safety checks — today, flashing
application data into a mega's reset walk region.
Readouts come one fact per line: `--info` decodes the info block field by
field, `--fuses` each fuse byte plus, on a boot-sectioned mega, its decoded
meaning. Transfers that take wire time draw a transient progress bar on stderr
`--autobaud` opens with the calibration pulse instead of the plain knock, for a
loader built `SERIAL autobaud`; the rest of the session is identical, at
whatever `--baud` the host chose. Its `--info` adds the **measured clock**
the loader's bit-period unit, decoded and multiplied by the session rate —
which is the number an `OSCCAL` bake or a fixed-baud build for the part is
held against; `--clock <hz>` states the drift against a nominal.
`--one-wire` marks the link as a shared line (*One-wire* above): the tool
reads back and verifies its own echoed bytes, whatever the backend.
It combines with everything, `--scan` included — undiscarded echoes would
answer every rate a scan probes.
`--scan` is the diagnosis once a fixed-baud loader has gone silent: it walks
±10 % around `--baud` in 2 % steps, nearest first, one probe per activation
window — reset the target as each probe announces itself (a board with DTR
wired to reset is pulsed by the probe's own port-open). A loader
off-frequency answers at its oscillator's ratio, and the report gives the
found rate as the session workaround, the offset, the OSCCAL correction's
direction at ~1 % per step, and the autobaud way out. Standalone — no other
operation combines with it.
`--poke ADDR:HEX` and `--peek ADDR[:N]` reach the data space (pureboot 5) —
SRAM, and through the same address space the register file and every I/O
register. Reading an I/O register can have side effects (reading UDR clears its
flags), which is the caller's business to know.
Reads are safe anywhere; **two small regions cannot be written without ending the
session,** because they are what the loader is standing on:
- the **top of SRAM**, where its stack lives — a handful of bytes below RAMEND;
- on an **autobaud** build, the **measured bit period**: two bytes in
GPIOR2:GPIOR1 where the chip has the pair (data `0x32..0x33` on the
t25/45/85, `0x4A..0x4B` from the x8 generation on — such a loader has *no*
static RAM at all), and the two bytes at RAMSTART on the chips without one
(the t13s and classic megas), where they are the whole of the loader's
static RAM. Overwrite either home and the next reply is timed against
garbage — the symptom is a mangled prompt byte rather than any error; the
loader is fine, it simply is no longer speaking the agreed rate.
Both are self-inflicted rather than defects, and a reset clears them. Note also
that `--poke` can write OSCCAL, which does take effect — but a session can only
survive a step or two of it before the clock walks the link out of the rate
autobaud locked to, and OSCCAL reverts on reset regardless.
Readouts come one fact per line: `--info` prints the device's version and
signature and the geometry that follows from them, `--fuses` each fuse byte
plus, on a boot-sectioned mega, its decoded meaning. Transfers that take wire time draw a transient progress bar on stderr
when it is a tty. `-v`/`--verbose` adds the decisions as they happen: knock
counts, the programming plan, update state handling and per-phase page counts.
## Tests
`tools/check.sh` runs every chip's workflow (`--full` adds the reflect-mode
builds of libavr's spot set; `tools/make_presets.py` regenerates the presets).
libavr rides as the `libavr/` submodule (`git submodule update --init libavr`);
`LIBAVR_ROOT` (cache or environment) overrides it for tandem development
against a working tree. `tools/check.sh` runs every chip's workflow (`--full`
adds the reflect-mode builds of libavr's spot set; `tools/make_presets.py`
regenerates the presets).
Per chip preset, `ctest` runs:
- `pureboot.size` — the 510-byte (patched-vector) / 512-byte budget;
@@ -321,23 +679,50 @@ Per chip preset, `ctest` runs:
the fastest clock — where a software UART's per-bit spin outgrows its
one-register delay loop and takes the 16-bit one. That is the largest image
the configuration space produces, and a shape the ladder default (always the
*fastest* rate a clock reaches) never picks. Pins are immediate operands and
the timeout is a constant: neither is an axis;
- `pbm_*.size` — under `--full`, the exhaustive cross product replacing that
compact matrix: every plausible oscillator (the internal ones, the CKDIV8
floor, the plain and the UART crystals) × every rate reachable from it ×
every backend, unreachable combinations dropping out rather than aborting
the configure. Bounded to one chip per size-bearing class — flash
addressing, hand-over shape, page size, USART inventory — since everything
else in the image is chip-independent code;
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, the
info block within the image's first 256 bytes;
*fastest* rate a clock reaches) never picks. Pins are an axis for one reason
only, and it is enough: a bit-banged link on a USART's own pins has to
release that USART, so `pureboot_{sw,autobaud}_on_usart{0,1}` build there
too. The timeout is a constant and is no axis;
- `pureboot_autobaud.size` — the clock-free build, which has no clock or baud
axis of its own: one binary per chip has to serve every point the matrix
below sweeps. `pureboot*osccal*.size` add the `OSCCAL` trim on the stock
shape and on the tightest image in the space (autobaud on a USART's own
pins), holding both of the trim write's addressing encodings to the budget;
- `pureboot_autobaud.unit` — the measured bit period sits where `--info`
reads it (wire contract, not layout accident): in the GPIOR pair, with no
RAM object at all, on the chips that have one; as the loader's only RAM
object at exactly ram_start elsewhere;
- `pbm_*.size` — with `PUREBOOT_FULL_MATRIX=1`, the exhaustive cross product
replacing that compact matrix, on **every** chip: every plausible oscillator
(the internal ones, the CKDIV8 floor, the plain and the UART crystals) ×
every rate reachable from it × every backend, unreachable combinations
dropping out rather than aborting the configure. Thousands of points per
chip, and cheap enough to run rather than reason about;
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, no
flash-resident section but `.text`, and the image byte-identical when linked
at a different base — which is position independence itself rather than a
proxy for it;
- `pureboot.handshake` — the host tool's activation must not hang on a target
that never falls quiet: the drain after a prompt is bounded by the handshake
deadline, and a well-behaved loader still connects;
- `pureboot.updatelink` — an update whose image changes the baud or the backend
must follow the staging copy onto *its* link, since that copy is the new image;
and where nothing was declared, the failure must name the link rather than
report a bare activation timeout, because by then the staging slot is written
and on a 1 KiB tiny that was the application;
- `pureboot.planner` — the host tool's pure logic: programming orders and their
recovery properties, the surgery, the staging composition, the boot-fuse
decode, the update preflight over synthetic fuse bytes, and the repairing
verify against a fake device;
- `pureboot.scan``--scan`'s walk and report logic: the probe order, the
rate arithmetic, and the trim advice's direction. A pty carries bytes at
any termios rate, so the rate physics itself belongs to the hardware
harness, and what the wire would arbitrate is pinned as logic;
- `presets.generated` — CMakePresets.json matches its generator
(`tools/make_presets.py --check`), so a hand edit or a generator change
cannot drift the pair apart;
- `pureboot.protocol` — end to end against a simavr device
(`test/pureboot_device.c`: a hardware USART as a pty, or a cycle-timed
(`test/pureboot_device.cpp`: a hardware USART as a pty, or a cycle-timed
GPIO⇄pty bridge for a software-UART build, plus the SPM/NVM module simavr's
tiny cores lack) driven by the real host tool through knock-from-reset,
program + verify of both memories, session reconnect, an external reset
@@ -354,6 +739,12 @@ Per chip preset, `ctest` runs:
- `pureboot.usart1` (644A) — the same suite over the second hardware USART:
instance selection is compile-checked everywhere, but only a live session
proves the loader polls the USART it claims;
- `pureboot.mute` (328P) — a software link on USART0's own pins, entered from an
application that handed over with that USART still enabled: the loader must
still answer, which it does only because it releases it. The pin ownership is
the runner's, not simavr's — simavr wires a USART through IRQs and never takes
the pin from the port, so without that model the state under test could not
arise at all;
- `pureboot.dirty` (328P) — entering the loader from a running application over
an SPM buffer it deliberately dirtied, the case the loader declines to guard:
a bare verify must see the corruption and the repairing verify must fix it in
@@ -361,7 +752,59 @@ Per chip preset, `ctest` runs:
anywhere, which is what makes the path constructible;
- `pureboot.update` — the full `--update-loader` flow, then every power-fail
phase: the device is killed mid-write, restarted from its flash dump, and a
re-run must complete the update with the application intact.
re-run must complete the update with the application intact;
- `pureboot.osccal` (328P, t85) — a loader built with the `OSCCAL` axis holds
the trim register at the built byte from its first prompt, observed through
the wire on one chip per addressing encoding (`sts` and low-I/O `out`);
- `pureboot.autobaud` (328P, 1284P) — the clock-free build over the GPIO⇄pty
bridge: the calibration handshake, a flash + EEPROM + fuse round trip against
the simulator's own memory, a data-space round trip, the hand-over — then the
same binary again at double the clock, which is the property the backend
exists for. The measured clock `--info` prints is asserted against the
simulator's exact clock, inside the unit encoding's own envelope, at both
points. A lone calibration pulse with no knock behind it must still let
the application boot, so no wait in activation can be unbounded.
`size`, `pi` and `planner` are host logic and run anywhere; the
simulator-driven targets need simavr and a pty, so they are POSIX-only.
`size`, `unit`, `pi`, `planner`, `scan` and `handshake` are host logic and run
anywhere; the simulator-driven targets need simavr and a pty, so they are
POSIX-only.
## Hardware
The suite above proves the protocol on every chip; it cannot prove a *board*.
Two things live only on silicon: an RC oscillator that is not on its nominal, and
a reset edge that has to come from somewhere. `tools/pbrig.py` and
`tools/pbhw.py` cover that, and know nothing per-board — every deployment fact
is a flag or a `PUREBOOT_*` environment variable.
```sh
export PUREBOOT_PROGRAMMER=atmelice_isp PUREBOOT_PART=t13 PUREBOOT_PORT=COM6
tools/pbrig.py backup rig-backup/ # verified, before anything is written
tools/pbhw.py --autobaud --loader build/ab.bin --app build/pbapp.hex --marker APP
```
`pbrig.py` is the primitives — `signature`, `reset`, `flash`, `fuses`, `backup`,
`rate` — and the module `pbhw.py` builds on. Two rig facts are encoded in it
because neither is guessable: an **ISP access is the reset edge** (the part runs
the moment the programmer releases it, which is the only edge available when the
adapter's DTR is not wired to reset, so a session begins with an ISP touch and
knocks immediately after), and **avrdude splits `-U` on colons**, so a Windows
path's drive letter breaks the spec and every file is passed as a bare name with
avrdude run in its own directory.
`pbrig.py rate` is the one that turns "the loader is silent, so the wiring must
be wrong" into a number. Against a fixture built with `PUREBOOT_HEARTBEAT` — a
*fixed* cycles-per-bit transmitter — it sweeps the host rate, and the band where
the marker still decodes brackets the part's true bit rate; with the clock the
image was built for, that is the clock the part is really running at. No
instrument beyond the adapter already attached. An ATtiny13A measured this way
came out at 9.072 MHz against its 9.6 MHz nominal, 5.5 % — inside the
datasheet's ±10 % and outside what an 8N1 frame survives, which is the whole
case for the autobaud backend on such a part.
`pbhw.py` takes its bounds from the info block the loader reports, so one run
covers a 1 KiB tiny and a 128 KiB mega alike: identity, the EEPROM round trip
and erase, an application flashed and verified and then *seen running*, the
application region read and erased, the loader slot proven intact across that
erase by an independent ISP read, and an oversized image refused. It overwrites
the application flash and EEPROM, which is why `backup` comes first.

View File

@@ -1,251 +0,0 @@
# pureboot autobaud — findings, and how the version question settled itself
The autobaud loader measures the host's bit timing **at runtime** from a
calibration pulse, so the image carries no clock: one clock-agnostic binary per
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
for the software-serial deployments — the RC-oscillator parts (the tinies,
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
each needs a per-clock build. Autobaud erases that axis. That is the win —
deployment, not bytes.
This document records how the fit was established, why the two versions that
were under review are both dead, and what the loader looks like now.
## The decision, settled by measurement
Two autobaud loaders were built for review, differing in one tradeoff — the
running-slot write guard against strict purity. `pureboot_autobaud_pure.cpp`
landed at **508 B** on the 1284P (4 B spare) and `pureboot_autobaud_reg.cpp` at
**512** (zero spare).
Then hardware testing found a defect that neither could absorb.
**A single spurious calibration pulse wedged the loader.** `run()` budgeted only
the start-edge wait inside `measure()`; the `rx()` that read the knock behind it
was unbudgeted and blocked forever. One stray low pulse on an unattended device
— EMI, or a host that opens the port and never knocks — held the loader in its
activation loop and the application never ran. On a field device that is a hang,
not a hiccup, and it is exactly the deployment autobaud is for.
The fix is to bound the whole activation: an expired knock budget returns a byte
that cannot be the knock, so control falls back into the budgeted `measure()`,
and a line that stays idle boots the application there. It costs about 22 bytes.
| 1284P, with the activation fix | size | 512 B budget |
|---|---|---|
| `pureboot_autobaud_pure.cpp` | 530 | **over by 18** |
| `pureboot_autobaud_reg.cpp` | 534 | **over by 22** |
| `pureboot_autobaud_uni.cpp` | **464** | **48 B spare** |
Both candidates were unshippable, and the margin they were competing over was
never real — it was the space the missing fix should have occupied. So the
choice is not between them. It is the third loader below, which fits with room
to spare *and* carries features neither had. The two sources stay in the tree
for the record; only the unified one is built.
## The unified loader
The insight that paid was the one that had already paid once: **merging command
bodies removes cost that moving them around only redistributes.** Folding `R`,
`r` and `w` into a single address-and-count path had been worth 14 B earlier.
Pushed further — one read command and one write command over *named spaces*
it is worth far more, because four transfer loops collapse into one.
`pureboot_autobaud_uni.cpp` is pureboot 5. It is strictly pure: no inline
assembly, no global register variable, and **no GPIOR either** — the measured
unit lives in a plain static, so the loader claims no chip resource an
application might want, and the GPIOR-versus-static question disappears along
with the chips that have no GPIOR.
### The protocol
| command | arguments | |
|---|---|---|
| `b` | — | version, then the three signature bytes |
| `J` | addr16 | ack, then jump (word address) |
| `W` | sel8, addr16, page bytes | fill the flash page buffer |
| `G` | sel8, addr16, n8 | read n bytes (0 means 256) |
| `g` | sel8, addr16, n8, then n bytes | write, each byte acked |
`sel` is `space | bank << 4`. The low nibble names the space; the high nibble is
flash's third address byte, so every transfer speaks a **byte** address inside a
64 KiB bank and no command has to carry word addresses. The host must not span a
bank boundary in one transfer — it already chunks by page, so nothing it does
comes close.
| space | | |
|---|---|---|
| 0 | flash | `lpm`/`elpm` |
| 1 | EEPROM | |
| 2 | data | SRAM — and with it the register file and every I/O register, which share the data address space on AVR |
| 3 | fuse and lock | |
| 4 | SPM | write-only: the data byte goes to SPMCSR and fires the instruction at the selected address |
Three things follow from that table that the loader never had:
- **RAM read and write**, the missing feature. In a per-command design it would
have cost a fresh dispatch arm and a fresh loop, ~30 B, on a loader with 4 B
spare. As one more space on a shared loop it is a single `ld`/`st`. It also
hands the host arbitrary **I/O register access** for free, because AVR maps
the peripherals into the same address space.
- **Host-driven SPM.** `W` used to end with a hardcoded erase, write and RWW
re-enable — 42 B. Those are now three writes to the SPM space, reusing the
store path's own address, data byte and ack. The host pays three extra
round-trips per page (18 wire bytes against 256 of data) and gains the ability
to issue *any* SPM operation, lock bits included.
- **`W` on the same footing as everything else.** It takes the same selector and
the same byte address instead of a word address of its own, which made flash
addressing uniform across the protocol *and* was 20 B cheaper than keeping its
private convention.
Read and write are `G` and `g` — the same letter, one bit apart — so the
transfer loop picks its direction with a one-word skip rather than a compare.
### Why the SPM space has to be one primitive
It is tempting to go further and expose a generic "poke this I/O register", from
which the host could drive SPM itself. The hardware forbids it: `out SPMCSR, x`
and the `spm` that follows must issue within four cycles, and EEPROM's
EEMPE→EEPE window is the same shape. A host cannot hit a four-cycle window
across a serial link. **Atomicity is the floor, and the atomic unit must be
resident.** That is the real limit on how low-level a bootloader's primitives
can go — not the byte count.
## Where the bytes went
The starting point was the 508 B pure build, disassembled and attributed:
| phase | bytes |
|---|---|
| `link::rx` + `link::tx` (bit-banged UART) | 102 |
| autobaud measure + knock | 62 |
| reset vector, pin init, WDRF check, jump, ack, EEPROM wait | 50 |
| command loop head + dispatch tree | 42 |
| `W` program flash page | 98 |
| `R` / `r` / `w` / `F` bodies, plus the shared address decode | 122 |
| `b` info, `J` jump | 32 |
**208 B — 41% — is physical layer and activation**, which no protocol change can
touch. The command bodies were the entire addressable surface, and they were
four copies of one idea.
The route from a first attempt to the final loader, all on the 1284P:
| step | size |
|---|---|
| unified `G`/`P`, `load`/`store` outlined, `__uint24` cursor | 600 |
| …`load`/`store` inlined; unit in `.noinit`, not `.bss` | 534 |
| …16-bit cursor with the bank in the selector; direction as a command bit | 510 |
| …erase/write/RWW moved out to the SPM space | 484 |
| …`W` sharing the selector-and-address decode | **464** |
Three of those steps are worth keeping as lessons:
- **Outlining `load`/`store` cost more than the four bodies they replaced.** As
functions they were 110 B against the 106 B of inline bodies — the AVR ABI's
argument marshalling plus prologue ate the entire saving. Inlined into the one
shared loop they cost only their own instructions. The call-site lesson cuts
both ways: *merging* call sites pays, *creating* one does not.
- **A `.bss` static drags in `__do_clear_bss`** — 18 B of startup code to zero a
variable that is always measured before it is read. `.noinit` is correct here
and free.
- **A three-byte cursor taxes every space.** Widening the shared cursor so flash
could reach past 64 KiB put an extra increment on EEPROM and RAM reads that
never need it. Moving the bank into the selector byte kept the cursor at
sixteen bits and cost nothing on the wire.
### What did not work
- **Encoding the space in the command byte** (so dispatch becomes masking rather
than a compare tree) cannot carry the bank's four bits alongside a space. The
cheap half of the idea survived as the direction bit; the rest lost to the
selector byte, which is also more extensible.
- **A generic primitive interpreter** — a loader with no logic at all, driven
entirely by the host — is not reachable on AVR. Harvard architecture means the
program counter cannot fetch from data space, so the classic "upload a flash
algorithm into RAM and jump to it" bootstrap is impossible, and on every
boot-sectioned part SPM only takes effect from the boot section anyway. What
remains is a fixed primitive set: still a protocol, still logic, only at a
different granularity. debugWIRE reaches that design point only because its
interpreter is *in silicon*; it costs the loader nothing because it is not in
the loader.
- Below 512 B the saved bytes are largely unspendable on the boot-sectioned
chips: the 328P's smallest boot section is exactly 512 B, and the 1284P's is
1024 B, of which pureboot already occupies only the top half. The margin
matters as headroom for correctness fixes — as this defect showed — not as
flash returned to the application. On the patch-vector parts, which have no
boot section, it *is* returned: on the ATtiny13 the loader is 43% of a 1 KiB
part, and every byte is real.
## The codegen coupling, still load-bearing
`count >> 2` is exact only because the calibration pulse's bit-count (7, from
the 0xC0 byte) equals the poll loop's cycles per iteration (7 — `sbis` 1,
`rjmp` 2, `adiw` 2, `rjmp` 2). The loop shape survived every restructuring here,
verified in the disassembly, but a toolchain bump that reshapes it would break
the lock silently. `test/pbautobaud.py` is what pins it: a wrong unit fails the
flash verify.
## Sizes — every chip
Budget 510 B on the patch-vector parts, 512 elsewhere. The 1284P is no longer
the tight one: the bank nibble made far flash *cheaper* than the near-flash
arithmetic it replaced.
| size | chips | budget | spare |
|---|---|---|---|
| 444 | ATtiny13, 13A | 510 | 66 |
| 448 | ATmega48, 48A, 48P, 48PA; ATtiny25 | 510 | 62 |
| 452 | ATtiny45, 85 | 510 | 58 |
| 460 | ATmega644, 644A, 644P, 644PA | 512 | 52 |
| 464 | **ATmega1284, 1284P**; ATmega8, 8A, 88, 88A, 88P, 88PA | 512 | 48 |
| 466 | ATmega16, 16A, 32, 32A; 164A/P/PA, 168/A/P/PA, 324A/P/PA, 328, 328P | 512 | 46 |
All 37 chips build and size-test green, plus the 12-preset reflect spot set
(guidance rule 4 — the reflect matrix is never run in full), which matches its
generated counterpart byte for byte on every chip in the set. Worst case across
the whole set is **466 B, 46 under budget**.
## Host tool and simulation
- **`pureboot.py`** speaks both generations. `Info.version >= 5` selects the
unified path; everything below it keeps the four-command protocol, so the
fixed-baud loader is untouched. `--autobaud` sends the 0xC0 pulse and one
knock, then derives full geometry from the signature. New: `--peek ADDR[:N]`
and `--poke ADDR:HEX` reach the data space.
- **`test/pbautobaud.py`** drives the loader over the GPIO⇄pty software-UART
bridge through the calibration handshake, a flash + EEPROM + fuse round-trip
cross-checked against the simulator's own memory, a RAM read/write round-trip,
and a hand-over to the fixture application — then repeats at double the F_CPU
with the same binary, which is the clock-agnostic property autobaud exists
for. Run on the near-flash 328P and the word-addressed 1284P.
- It also **pins the activation hang**: the test sends a lone calibration pulse
with no knock behind it and requires the application to boot. Against the
unfixed loader that assertion never returns.
## What remains
- **Real-hardware acceptance.** A cycle-exact simulator cannot produce what
autobaud exists for: a real RC oscillator at ±10% with drift and jitter.
simavr proves the arithmetic and the fit at exact clocks; only silicon proves
the feature. Drive an internal-oscillator ATtiny at a fixed host baud and
confirm lock plus a full flash and verify.
- **A generic `spm::command()` in libavr.** The SPM space issues a runtime
command through `spm::detail::page_command` where RAMPZ exists, and falls back
to a dispatch over the known operations where it does not — the one
preprocessor branch in the file. A two-line library addition would make it
uniform and save a few bytes on the 36 non-RAMPZ chips, none of which are
tight.
- **Retire or revive the two dead variants.** They are kept only as the record
of the measurement; nothing builds them.
## Files
- `pureboot_autobaud_uni.cpp` — the loader. pureboot 5.
- `pureboot_autobaud_pure.cpp`, `pureboot_autobaud_reg.cpp` — superseded, not
built; 530 and 534 B on the 1284P once the activation hang is fixed.
- `pureboot.py``--autobaud`, the unified transfer path, `--peek`/`--poke`.
- `test/pbautobaud.py` — the end-to-end sim test and the hang regression.
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
probe at 506 B, off-tree and gitignored; a size reference only. The unified
loader is 42 B under it, with features the probe never had.

View File

@@ -1,14 +1,17 @@
// pureboot a serial bootloader on libavr: one C++ source, no inline
// assembly, no global register variables, 512 bytes on every chip libavr
// targets. The device speaks primitives; every composite (verify, erase,
// pureboot - a serial bootloader on libavr: one C++ source, no inline
// assembly, no global register variables, a 512-byte slot on every chip
// libavr targets. The device speaks primitives; every composite (verify, erase,
// reset-vector surgery, self-update) lives in the host tool. Protocol,
// deployment and configuration: README.md next to this file.
//
// The image is position-independent PC-relative control flow, wire
// addresses in, the write guard and the info block both anchored on the
// runtime return address — so the identical binary runs from any slot. That
// is what makes a copy one slot below able to rewrite the resident one, and
// every change here has to keep it (test/check_pi.py).
// The image is position-independent - PC-relative control flow and wire
// addresses in, no absolute address formed anywhere - so the identical binary
// runs from any slot. That is what makes a copy one slot below able to rewrite
// the resident one, and every change here has to keep it (test/check_pi.py).
// It does not need to know *which* slot it is in: nothing here refuses an
// address, so there is no running-slot comparison to anchor.
#include <chrono>
#include <libavr/libavr.hpp>
@@ -22,85 +25,176 @@ namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
// The EEPROM procedure's step 2 - wait until SPMEN clears - guards a flash
// operation still in flight, and this loader never has one when it touches
// the EEPROM: commit() waits every boot-sectioned page operation out before
// the ack, and everywhere else the CPU halts through the operation itself.
// The posture states the omission the datasheet grants for exactly that
// (DS40002061B section 8.6.3).
constexpr auto no_spm = ee::spm_interlock::omitted;
constexpr std::uint8_t ack = '+';
// The refusal, which is the ack inverted: on a link whose whole problem is
// flipped bits, the byte saying "nothing happened" should be as far as a byte
// can be from the one saying "it did", and the complement is all eight bits.
// It is also the only spelling that needs no justifying - every other value
// would be a choice.
constexpr std::uint8_t nak = static_cast<std::uint8_t>(~ack);
// What a command's fields must fold to. Any non-zero constant does: zero is
// what a run of one repeated byte folds to, and a repeated byte is the shape of
// both a line stuck at a level and a page of erased flash arriving where a
// header belongs.
constexpr std::uint8_t seal = 0x5a;
// The opcode, as bits rather than letters. Each is a one-instruction skip,
// where a set of arbitrary values costs a compare and a branch apiece - and
// with the seal deciding what is a command at all, there is nothing left for a
// readable spelling to buy. A transfer is the absence of the other three, and
// its direction is the low bit.
//
// Identify is bit 5 for one reason: 'p' and 'b' both carry it, and those are
// the knock. Two things follow that no other assignment gives. A host cannot
// know which generation it is talking to until something has answered, so the
// command reporting the version has to mean the same thing before the version
// is known - 'b' still asks it. And a knock aimed at a loader that is
// *already* in session has to stay harmless: with no opcode reserved as
// invalid, every byte now starts a command, so a knock that meant nothing to
// earlier generations would otherwise consume the five header bytes behind it
// and put the stream out of step. Answering both knock bytes with the identity
// keeps the reconnect exactly as cheap as it was.
enum : std::uint8_t { op_write = 1, op_fill = 4, op_jump = 8, op_identify = 0x20 };
// Deployment parameters come from the build (pureboot_add_loader()). The
// signature is not one of them: the chip database is the only universal
// source a tiny13A cannot read its own signature row from code.
#if !defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD)
// source - a tiny13A cannot read its own signature row from code. An autobaud
// build carries no clock and no baud at all; it measures both.
#if !defined(PUREBOOT_AUTOBAUD) && (!defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD))
#error \
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud create loader targets with pureboot_add_loader() (README.md)"
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud - create loader targets with pureboot_add_loader(), or PUREBOOT_AUTOBAUD for a clock-free one (README.md)"
#endif
#if !defined(PUREBOOT_AUTOBAUD)
using dev = avr::device<{.clock = avr::hertz_t{PUREBOOT_CLOCK_HZ}}>;
constexpr avr::baud_t wire_baud{PUREBOOT_BAUD};
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{avr::hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return avr::hw::db.field_index(reg, "WDRF");
}
#endif
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
// Chips without a hardware boot section the tinies and the m48s, whose SPM
// runs from anywhere (Atmel-8271 §26) keep the application's relocated
// reset vector in the word under the slot.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
// Chips without a hardware boot section - the tinies and the m48s, whose SPM
// runs from anywhere (Atmel-8271 section 26) - keep the application's relocated
// reset vector in the word under the slot. The size itself is the linker's and
// the host's business: nothing in here needs to know where the slot ends.
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = avr::hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one). A slot is 256 of
// those — one value of a wire address's high byte, where 512 bytes span two.
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint16_t wire_base =
word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base);
// Past 64 KiB one bank of flash does not cover the chip, so a transfer's
// selector byte carries the bank and the wire address stays a byte address
// within it. 'J' is the exception: it is a word address everywhere, because
// that is what the hardware's own jump takes.
constexpr bool banked_flash = spm::flash_bytes > 65536;
// A compile-time window, so the whole EEPROM belongs to the application;
// re-timing a deployed loader is a self-update with a re-timed build.
// re-timing a deployed loader is a self-update with a re-timed build. An
// autobaud build has no clock to convert seconds against and counts polls.
#if !defined(PUREBOOT_TIMEOUT)
#define PUREBOOT_TIMEOUT 8
#endif
constexpr std::uint8_t timeout_seconds = PUREBOOT_TIMEOUT;
// The loader's one identity number. The protocol carries none of its own —
// a version implies it, and the host tool holds that map (README.md).
constexpr std::uint8_t version = 4;
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr avr::uint24_t autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The 'b' reply, byte for byte (layout: README.md). Flash-resident because
// no crt copies a .data image — and flash_table's storage carries the word
// alignment 'b' needs to halve the address on the large chips.
// One wire byte per line: this is the reply's layout, not a list.
// A build may bake a measured oscillator trim (README.md: the RC-oscillator
// deployment answer); the byte is applied at the top of run(). Orthogonal to
// the serial backend - an autobaud build may carry it for the application's
// benefit alone.
#if defined(PUREBOOT_OSCCAL)
static_assert(PUREBOOT_OSCCAL >= 0 && PUREBOOT_OSCCAL <= 0xff, "PUREBOOT_OSCCAL is one OSCCAL byte");
#endif
// The loader's one identity number. The protocol carries none of its own -
// a version implies it, and the host tool holds that map (README.md).
constexpr std::uint8_t version = 9;
// The image's identity stamp, for the host tool rather than for the wire: an
// update image is a bare 512-byte slot, and without this nothing in it says
// which chip it was built for. The tool refuses to install an image whose
// stamp does not match the device - flashing a foreign loader bricks the
// target, and the loader itself cannot check what has already replaced it.
//
// Never read from flash by the loader - 'b' answers out of this array, but at
// constant indices, so those fold to immediates and no runtime address of it
// is ever formed. That folding is a correctness property, not a size one: the
// bytes live in program memory and a formed address would be dereferenced as
// *data* space, which is why this is the raw array rule 36 otherwise bans - a
// `std::array` here stops the read loop unrolling and emits exactly that
// `ld` (measured: +8 B and a wrong answer on the wire). `used` keeps the
// compiler from dropping the copy the host needs and `retain` keeps
// --gc-sections from collecting it.
// clang-format off
inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{
'P',
'B',
version,
[[gnu::used, gnu::retain, gnu::section(".text.stamp")]]
inline constexpr std::uint8_t identity_stamp[]{
'P', 'B', // the magic the host scans an image for
version, // and from here on, exactly what 'b' answers
avr::hw::db.signature[0],
avr::hw::db.signature[1],
avr::hw::db.signature[2],
static_cast<std::uint8_t>(page), // 0 means 256
wire_base & 0xff,
wire_base >> 8,
avr::hw::db.mem.eeprom_size & 0xff,
avr::hw::db.mem.eeprom_size >> 8,
static_cast<std::uint8_t>((boot_section ? 0 : 1) | (word_flash ? 2 : 0)), // patch-vector, word-addressed
}>
info_data;
};
// clang-format on
// Where the identity proper starts: past the magic the host scans for.
constexpr std::uint8_t stamp_identity = 2;
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL,
// defaulting to the chip's USART0 where it has one. The software receiver is
// the polled one: the vector table belongs to the application. Templates on
// the clock, so only the selected backend instantiates. pending() is the
// cheap line test the activation window polls; drain() holds until the last
// frame is off the wire, so a hand-over cannot let the target's re-init clip
// the ack.
// The address spaces a transfer can name, in a selector byte's low nibble.
// Flash is 0 so it is the cheapest to select.
//
// spm_ops is the one that is not memory: naming it hands the count field to
// SPMCSR and fires the instruction at the transfer's address, which is how
// page erase, page write and RWW re-enable reach the wire without the loader
// carrying a command for each. The hardware's four-cycle store-to-SPM window
// is why this is one fused primitive and not a poke of SPMCSR - no host can
// hit that window across a serial link.
//
// It is the one space with no direction: the opcode's write bit is not
// consulted, because a sealed command naming this space says what it means and
// there is nothing for the other direction to denote. Testing the bit anyway
// would cost six bytes to catch a host contradicting itself, which is the same
// trade the running-slot guard lost.
enum : std::uint8_t { sp_flash = 0, sp_eeprom = 1, sp_data = 2, sp_fuse = 3, sp_spm = 4 };
// A selector's high nibble is the flash bank - the address bits above the
// 16-bit wire address, RAMPZ on the chips that have one. Keeping it here
// rather than widening the wire address is what lets one 16-bit cursor serve
// every space: a 24-bit cursor would pay its extra byte on EEPROM and data
// reads that can never need it.
[[gnu::always_inline]] inline std::uint8_t space_of(std::uint8_t selector)
{
return selector & 0x0f;
}
[[gnu::always_inline]] inline std::uint8_t bank_of(std::uint8_t selector)
{
return static_cast<std::uint8_t>(selector >> 4);
}
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL /
// PUREBOOT_AUTOBAUD, defaulting to the chip's USART0 where it has one. The
// software receiver is the polled one: the vector table belongs to the
// application. Templates on the clock, so only the selected backend
// instantiates. pending() is the cheap line test the activation window polls;
// drain() holds until the last frame is off the wire, so a hand-over cannot
// let the target's re-init clip the ack.
#if defined(PUREBOOT_SOFT_SERIAL) && defined(PUREBOOT_USART)
#error "PUREBOOT_SOFT_SERIAL and PUREBOOT_USART select opposing serial backends"
#endif
#if defined(PUREBOOT_AUTOBAUD) && defined(PUREBOOT_USART)
#error "PUREBOOT_AUTOBAUD measures a software link; it cannot drive a hardware USART"
#endif
#if defined(PUREBOOT_HALF_DUPLEX) && (defined(PUREBOOT_SOFT_SERIAL) || defined(PUREBOOT_AUTOBAUD))
#error "PUREBOOT_HALF_DUPLEX is the hardware USART's one-wire mode; a software link goes one-wire by RX == TX"
#endif
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
@@ -108,18 +202,61 @@ inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{
#define PUREBOOT_TX pb1
#endif
#if defined(PUREBOOT_USART)
constexpr char usart_digit = '0' + PUREBOOT_USART;
constexpr int usart_unit = PUREBOOT_USART;
#else
constexpr char usart_digit = '0';
constexpr int usart_unit = 0;
#endif
template <avr::hertz_t C>
struct hardware_link {
using uart = avr::uart::usart<usart_digit, C, {.baud = wire_baud, .max_baud_error = 2.5_pct}>;
// One-wire on the hardware USART (PUREBOOT_HALF_DUPLEX): RXD and TXD tied
// together off-chip, exactly one direction enabled at a time - the library's
// .half_duplex turn-around. The activation window is unchanged; only its
// poll grows the release-line test readable() carries in this mode.
constexpr bool hw_half_duplex =
#if defined(PUREBOOT_HALF_DUPLEX)
true;
#else
false;
#endif
// The compiled idle poll: lds UCSR0A (2), sbrc skipping the exit (2),
// sbiw + sbci + sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 10;
template <avr::hertz_t C, avr::baud_t B>
struct hardware_link {
// The deployment envelope is the build's: pureboot_baud_feasible() holds
// every configured rate within 2.5 %, which the bench has proven across
// the fleet, and the datasheet's stricter per-frame tolerance table would
// refuse the stock 115200 at 16 MHz (+2.1 %) that every deployed board
// runs. .allow_baud_error states that this is meant.
using uart = avr::uart::usart<usart_unit, C,
{
.baud = B,
.allow_baud_error = true,
.half_duplex = hw_half_duplex,
}>;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// the RXC test, then sbiw + sbci + brne (5). The test's cost follows the
// status register's home - a 2-cycle bit-skip where UCSRnA sits in
// bit-addressable I/O (the classic megas), lds + skip (4) in extended
// I/O. Half-duplex polls through readable()'s release-line test, which
// -Os outlines: the rcall (3), the UCSR#B read and not-taken skip with
// the jump over the write (I/O 3, extended 5), the ret (4) - and the
// call in the loop body pushes the countdown into call-saved registers,
// where the uint24_t step is ldi+sub+sbc+sbc (4) instead of sbiw+sbci
// (3). Measured off the built loops: 18 a poll in bit-addressable I/O,
// 22 in extended. A uint32_t countdown pays one more sbci -
// window_polls() adds it where the count forces the wide type. Held per
// chip by the pureboot.window gates. The lookup rides the baud parameter
// so it stays dependent: the trait is an incomplete type on the
// USART-less chips, which parse this template without ever instantiating
// it.
template <avr::baud_t Baud, typename U = avr::hw::usart_of<usart_unit>>
static consteval std::uint8_t poll_cost()
{
if (hw_half_duplex) {
return U::ucsra::addr < 0x40 ? 18 : 22;
}
return U::ucsra::addr < 0x40 ? 7 : 9;
}
static constexpr std::uint8_t poll_cycles = poll_cost<B>();
static void init()
{
@@ -128,7 +265,7 @@ struct hardware_link {
static bool pending()
{
return uart::rx_ready();
return uart::readable();
}
static std::uint8_t rx()
@@ -143,18 +280,26 @@ struct hardware_link {
static void drain()
{
uart::drain();
// A drain here always follows this link's own write - the frame is
// in flight by construction, so the completion the wait needs is
// guaranteed and the bounded default's countdown would be dead bytes.
uart::drain_unbounded();
}
};
template <avr::hertz_t C>
template <avr::hertz_t C, avr::baud_t B>
struct software_link {
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, wire_baud>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, wire_baud>;
// RX == TX is the one-wire deployment: the transmitter becomes a guest
// on the receiver's pull-up line, taking the pin's direction for exactly
// one frame per byte.
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, B>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, B, avr::PUREBOOT_RX == avr::PUREBOOT_TX>;
// The compiled idle poll: sbis skipping the exit (2), sbiw + sbci +
// sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 8;
// The compiled idle poll around the window's narrow (uint24_t) countdown:
// sbis skipping the exit (2), sbiw + sbci + brne (5). A uint32_t
// countdown pays one more sbci - window_polls() adds it where the count
// forces the wide type. Held by the pureboot.window gate.
static constexpr std::uint8_t poll_cycles = 7;
static void init()
{
@@ -182,18 +327,54 @@ struct software_link {
}
};
#if defined(PUREBOOT_USART)
static_assert(avr::uart::has_usart<usart_digit>(), "PUREBOOT_USART selects a hardware USART this chip does not have");
using link = hardware_link<dev::clock>;
// The clock-free link: the bit period is measured from the host's calibration
// pulse instead of derived from a clock, so one image serves every F_CPU and
// every rate. Activation differs in kind from the other two - there is no
// clock to time a window against - so this backend brings its own, below.
struct autobaud_link {
// The unit in GPIOR2:GPIOR1 where the chip has them: the loader owns the
// whole chip while it runs, and the pair costs one word per access where
// the RAM word costs two - six words across the image.
using uart = avr::uart::software_autobaud<avr::PUREBOOT_RX, avr::PUREBOOT_TX, avr::uart::unit_home::gpior>;
static void init()
{
avr::init<uart>();
}
static std::uint8_t rx()
{
return uart::template read<off>();
}
static void tx(std::uint8_t byte)
{
uart::template write<off>(byte);
}
static void drain()
{
// A drain here always follows this link's own write - the frame is
// in flight by construction, so the completion the wait needs is
// guaranteed and the bounded default's countdown would be dead bytes.
uart::drain_unbounded();
}
};
#if defined(PUREBOOT_AUTOBAUD)
using link = autobaud_link;
#elif defined(PUREBOOT_USART)
static_assert(avr::uart::has_usart<usart_unit>(), "PUREBOOT_USART selects a hardware USART this chip does not have");
using link = hardware_link<dev::clock, wire_baud>;
#elif defined(PUREBOOT_SOFT_SERIAL)
using link = software_link<dev::clock>;
using link = software_link<dev::clock, wire_baud>;
#else
using link =
std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock>, software_link<dev::clock>>;
using link = std::conditional_t<avr::uart::has_usart<usart_unit>(), hardware_link<dev::clock, wire_baud>,
software_link<dev::clock, wire_baud>>;
#endif
// The application's entry, pinned by the linker (--defsym): word 0 on a
// boot-sectioned mega, the trampoline at base 2 elsewhere. Reaching it must
// boot-sectioned mega, the trampoline at base - 2 elsewhere. Reaching it must
// not depend on where this copy runs, so the jump goes through a pointer, and
// [[gnu::noipa]] keeps the constant from folding back into a relative call.
extern "C" [[noreturn]] void pureboot_app();
@@ -204,24 +385,71 @@ extern "C" [[noreturn]] void pureboot_app();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
[[noreturn]] void run_app()
{
jump(pureboot_app);
}
// The window as one 32-bit countdown, divided by the backend's counted
// poll-loop cycles. Whole seconds is all it promises.
// Activation: a bounded wait for the host, then the knock. Both forms boot the
// application when the window closes on an idle line, and both bound *every*
// wait - a knock awaited without a deadline would let one stray edge hold an
// unattended device in the loader forever.
#if defined(PUREBOOT_AUTOBAUD)
// The window is a fixed poll budget: with no clock, whole seconds cannot be
// timed. A uint24_t holds it - a fourth byte would cost two words at every
// countdown step for range never used.
void await_host()
{
for (;;) {
if (!link::uart::calibrate(autobaud_budget)) {
run_app();
}
// The calibration pulse has already proven a host is there, so one
// byte activates. A knock that never arrives falls back to calibrate(),
// whose own budget then boots the application.
if (link::uart::template read<off>(autobaud_budget) == 'p') {
return;
}
}
}
#else
// The window as one countdown, divided by the backend's counted poll-loop
// cycles. Whole seconds is all it promises. The per-poll cost depends on the
// countdown's own width (a uint32_t decrement chain is one sbci longer), and
// the width depends on the poll count - solved narrow-first: a count that
// fits 24 bits at the narrow cost keeps the narrow loop, anything else takes
// the wide loop at its own cost. A count fitting 24 bits only at the wide
// cost stays wide, so the choice cannot oscillate on the boundary.
consteval std::uint32_t polls_at(std::uint32_t per_poll)
{
// Whole-window cycles first, then the per-poll division: one truncation
// instead of one per second. Same instructions either way - only the
// countdown's immediate moves.
return static_cast<std::uint32_t>(dev::cycles_for<std::chrono::seconds{timeout_seconds}>() / per_poll);
}
consteval bool narrow_window()
{
return polls_at(link::poll_cycles) <= 0xffffff;
}
consteval std::uint32_t window_polls()
{
return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles);
return polls_at(narrow_window() ? link::poll_cycles : link::poll_cycles + 1u);
}
// The countdown in the narrowest type that holds it: a fourth byte would
// cost a wider decrement chain at every poll for range most windows never
// use (the autobaud budget makes the same choice).
using window_t = std::conditional_t<narrow_window(), avr::uint24_t, std::uint32_t>;
bool pending_before_deadline()
{
std::uint32_t polls = window_polls();
window_t polls = window_polls();
do {
if (link::pending())
if (link::pending()) {
return true;
}
} while (--polls);
return false;
}
@@ -230,11 +458,20 @@ bool pending_before_deadline()
// application runs.
std::uint8_t rx_deadline()
{
if (!pending_before_deadline())
if (!pending_before_deadline()) {
run_app();
}
return link::rx();
}
void await_host()
{
// 'p' then 'b', each under a fresh window; anything else is line noise.
while (rx_deadline() != 'p' || rx_deadline() != 'b') {
}
}
#endif
// Inlined: read across a call, the first byte strands in a call-saved
// register the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
@@ -243,7 +480,7 @@ std::uint8_t rx_deadline()
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
// The wire's byte pair as the word it is AVR is little-endian too, so the
// The wire's byte pair as the word it is - AVR is little-endian too, so the
// cast is the identity a shift-and-or spelling makes the compiler rediscover.
// Callers read into named variables first: the wire order is a sequence of
// reads, not an argument order.
@@ -252,213 +489,227 @@ std::uint8_t rx_deadline()
return std::bit_cast<std::uint16_t>(pair);
}
// Counts arrive in the wire's 8-bit form: 0 means 256. Both streamers fold
// into the one command that reads flash, which is what lets the far one's
// 24-bit cursor sit in the command loop's own call-saved registers.
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
// The 24-bit cursor as the machine holds it — the RAMPZ byte and a 16-bit Z,
// carried apart; the reassembled address folds away inside the far load.
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
// Carrying the wrap is smaller than the flat 32-bit cursor GCC
// builds without it.
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
// Out of line: three sites send it, and a call is shorter than three
// load-immediates.
// Out of line: several sites send it, and a call is shorter than a
// load-immediate at each.
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
// A wire address and its selector's bank as the flash address they name.
[[gnu::always_inline]] inline spm::flash_address_t flash_address([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
{
do
link::tx(ee::read(address++));
while (--count);
if constexpr (banked_flash) {
return (static_cast<spm::flash_address_t>(bank) << 16) | at;
} else {
return at;
}
}
// Host-paced: the ack goes out once the write has begun, so the next byte
// arrives while it completes and nothing is missed without a buffer.
void store_eeprom(std::uint16_t address, std::uint8_t count)
// One byte out of any space. Every accessor shares the transfer's cursor, its
// loop and its call site, so a space costs only its own instruction rather
// than a body, a loop and a dispatch arm of its own.
[[gnu::always_inline]] inline std::uint8_t load(std::uint8_t space, [[maybe_unused]] std::uint8_t bank,
std::uint16_t at)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
if (space == sp_eeprom) {
return ee::read<no_spm>(at);
}
if (space == sp_data) {
return *reinterpret_cast<volatile std::uint8_t *>(at);
}
if (space == sp_fuse) {
return spm::read_fuse<off>(static_cast<spm::fuse>(at));
}
if constexpr (banked_flash) {
return avr::flash_load_far<std::uint8_t>(flash_address(bank, at));
} else {
return avr::flash_load(reinterpret_cast<const std::uint8_t *>(at));
}
}
// One page into the SPM buffer, then erase and program — except the slot
// this code is running in (`slot_high`, from run()), which is drained and
// left alone. A broken host therefore cannot brick the running loader, and a
// copy one slot lower may rewrite the resident one.
// One byte into a writable space. Flash is not one of them - it arrives a
// page at a time through 'W' and is committed by the sealed SPM command - and
// the fuses are not writable at all: SPM reaches flash and boot lock bits only.
[[gnu::always_inline]] inline void store(std::uint8_t space, std::uint16_t at, std::uint8_t value)
{
if (space == sp_data) {
*reinterpret_cast<volatile std::uint8_t *>(at) = value;
return;
}
// Host-paced: the ack goes out once the write has begun, so the next byte
// arrives while it completes and nothing is missed without a buffer.
ee::write<off, no_spm>(at, value);
}
// The irreversible half of the protocol, and the whole of it: page erase, page
// write and the lock bits are one SPM command each, and nothing else the loader
// does outlasts being done again. Reached only from a sealed command (run()),
// so both the byte handed to SPMCSR and the address it fires at are the ones
// the host computed its seal over.
//
// Nothing discards the buffer first: it is write-once per word (§26.2.1), so
// Nothing here refuses an address. A loader that will not write its own slot
// cannot plant anything in it either, and a resident copy able to rewrite its
// own trailing page is what lets a 512-byte boot section - where no staging
// copy can run SPM at all - carry an SPM primitive for an application-side
// installer to drive. The protection that made the guard look necessary is the
// seal: a wire fault can no longer name an address, only a host can, and a host
// that names this one means it.
[[gnu::always_inline]] inline void commit(std::uint8_t bank, std::uint16_t at, std::uint8_t value)
{
spm::command<off>(value, flash_address(bank, at));
// Only a boot-sectioned mega runs on while its RWW section programs;
// everywhere else the CPU halts through erase and write, so the wait
// is already over by the time it returns.
if constexpr (boot_section) {
spm::wait();
}
}
// One page into the SPM buffer, and only that: the erase and the write that
// commit it are host-issued sp_spm stores, which reach the same fused
// store-and-SPM pair through the transfer path's own address and data.
//
// Nothing discards the buffer first: it is write-once per word (section 26.2.1), so
// filling over a refused page or an application's leavings programs stale
// words but a page write auto-erases it (§26.2.1; §19.2 on the tinies), so
// words - but a page write auto-erases it (section 26.2.1; section 19.2 on the tinies), so
// that write clears the condition and the host's read-back rewrites the page.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
void fill_page(std::uint8_t bank, std::uint16_t at)
{
// The address names a page, so its in-page bits are dropped and the walk
// starts at the page base — one induction either way: a byte-addressed
// wire address walks the page itself (the offset bits wrap back to zero),
// while a word one becomes a byte cursor once. The slot index is the wire
// address's high byte — on byte-addressed chips the byte address's, with
// the low bit dropped, since a slot is two of those.
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
// A page is aligned, so it never crosses 64 KiB: RAMPZ is a per-page
// constant and the 16-bit Z's low byte is the whole in-page offset.
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
// starts at the page base; the low byte of the cursor is the whole in-page
// offset, since a page is aligned and never crosses a bank.
std::uint16_t z = at & ~static_cast<std::uint16_t>(page - 1);
// The receipt is the erase-first contract's token and costs nothing here:
// the erase is the host's own sealed SPM command, before or after the fill
// as it chooses (the page write clears a stale buffer either way, above).
const auto open = spm::page::begin<spm::from::boot_section, off>(flash_address(bank, z));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
spm::fill<off>(open, flash_address(bank, z), word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
}
if (page_high != slot_high) {
// Only a boot-sectioned mega runs on while its RWW section programs;
// everywhere else the CPU halts through erase and write.
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
// Programming leaves the RWW section disabled; reads need it back on. The
// same store discards the buffer (§26.2.2), so a boot-sectioned mega never
// meets the stale-word case above.
if constexpr (boot_section)
spm::rww_enable<off>();
}
// The four fuse and lock bytes in the hardware's own Z order: low, lock,
// extended, high.
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
} while (static_cast<std::uint8_t>(z) & (page - 1));
}
[[noreturn]] void run()
{
#if defined(PUREBOOT_OSCCAL)
// The build's oscillator trim, ahead of everything - the WDRF bail
// included - so every path out of reset, the watchdog hand-over to the
// application first among them, runs on the corrected clock.
avr::clock::calibrate(PUREBOOT_OSCCAL);
#endif
// A watchdog reset belongs to the application, whose watchdog stays forced
// on until it clears WDRF no activation window in its way.
if (avr::hw::field_impl<wdrf_field()>::test())
// on until it clears WDRF - no activation window in its way.
if (avr::power::peek_reset_cause().watchdog) {
run_app();
}
link::init();
// The high byte of the slot this copy runs at, which the write guard and
// the info block both follow: the return address is a word address, so its
// high byte is the 256-word slot index, doubled back into byte terms where
// the wire counts bytes. Taken as byteswap's low byte — the builtin already
// swaps the two stacked bytes, and the double swap folds away, where `>> 8`
// would leave the swap materialized.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
// 'p' then 'b', each under a fresh window; anything else is line noise.
while (rx_deadline() != 'p' || rx_deadline() != 'b') {
}
await_host();
for (;;) {
// No prompt while an EEPROM write runs: it blocks SPM and fuse reads
// (§26.2.1), and the prompt is the previous command's completion ack.
ee::wait();
// (section 26.2.1), and the prompt is the previous command's completion ack.
ee::wait<no_spm>();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
if (command & op_identify) {
// Identity: the version, then the three signature bytes. Straight
// out of the stamp, so the wire and the image can never disagree
// about what this loader is. The indices are constant and the array
// is constexpr, so these are immediates, not flash reads: nothing
// here needs the stamp's runtime address. Unsealed, because it
// takes no argument and changes nothing - and because a command
// that cannot be got wrong is what a lost host resynchronises on.
for (std::uint8_t at = stamp_identity; at != sizeof identity_stamp; ++at) {
link::tx(identity_stamp[at]);
}
case 'b': // info block, read relative to the running slot
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for all four: 'b' is a flash read
// whose arguments the loader already knows, so it joins the
// wire-argument three rather than streaming from a call site of its
// own. That leaves one flash streamer in the image, and lets its
// cursor live in this never-returning loop's own call-saved
// registers instead of being saved and restored around a call.
std::uint16_t address;
std::uint8_t count;
if (command == 'b') {
// The block sits in the image's first 256 bytes (check_pi.py
// asserts it) and slots are 512-aligned, so the low byte of its
// link address is its offset in any slot — halved where wire
// units are words. The high byte is runtime data, so no
// absolute address is ever materialized.
const auto link_byte =
static_cast<std::uint8_t>(reinterpret_cast<std::uint16_t>(info_data.storage.data()));
const std::uint8_t low = word_flash ? static_cast<std::uint8_t>(link_byte >> 1) : link_byte;
address = static_cast<std::uint16_t>(low | (slot_high << 8));
count = static_cast<std::uint8_t>(info_data.size());
} else {
address = rx16();
count = link::rx();
// One decode, one cursor and one loop for every space, both
// directions and the jump: a command per memory would carry a copy
// of all three each. The jump - the hand-over and staging transfer
// - carries a selector it ignores so its address rides the same two
// reads as everything else; the page fill joins the same decode
// rather than keeping an address form of its own, so flash
// addressing is uniform across every command that names it. Both
// carry the count they do not use for the same reason: one header
// shape is one decode, and one seal covers a fixed set of bytes.
const std::uint8_t selector = link::rx();
const std::uint8_t space = space_of(selector);
const std::uint8_t bank = bank_of(selector);
std::uint16_t at = rx16();
std::uint8_t count = link::rx();
const std::uint8_t sealed = link::rx();
// The seal: every field that decides what this command does folded
// into one byte the host chose, tested before any of it happens.
//
// Checked here rather than acknowledged afterwards, which is the
// whole point. An ack reports a command that has already run, and
// for the one command that cannot be taken back a report is not a
// defence. Once the running-slot guard is gone the address is as
// fatal as the command byte - a wrong one reaches the loader's own
// page - so the seal covers the act and the place together, and a
// stream that lost or mangled either cannot produce it.
//
// Folded here, after the last read, and never accumulated across
// the reads: every field is still live at this point because the
// command needs it anyway, so the fold costs one xor each and no
// register. An accumulator would have to survive four calls, and
// paying for that in call-saved registers costs more than the whole
// check costs in arithmetic - measured at fourteen bytes, on a
// budget of ten.
std::uint8_t fold = command;
fold ^= selector;
fold ^= static_cast<std::uint8_t>(at);
fold ^= static_cast<std::uint8_t>(at >> 8);
fold ^= count;
fold ^= sealed;
// The verdict, and it is not a courtesy. Every command whose
// payload the host sends without waiting - a page fill, a write
// burst - would otherwise be handed to a loader that has already
// gone back to reading commands, so a *detected* error would
// become the desync the seal exists to prevent: a 128-byte page
// read as command headers is twenty-one more chances at the one in
// two hundred and fifty-six. Answering the seal before the payload
// is what keeps a refusal local to the command that earned it.
//
// An unknown opcode lands here too - every bit pattern is now some
// command, so it is the seal, not a table of valid letters, that
// rejects noise, and the host hears about it either way.
if (fold != seal) {
link::tx(nak);
} else {
tx_ack();
if (command & op_jump) {
link::drain();
jump(reinterpret_cast<void (*)()>(at));
} else if (command & op_fill) {
fill_page(bank, at);
} else if (space == sp_spm) {
// An SPM command is the whole of what this loader can do
// that doing again will not undo, and it is one byte - so
// it rides the count field, inside the seal, rather than
// arriving as data after the seal has been checked. Which
// is also what makes a deliberate lock-bit write
// expressible, where refusing it outright did not.
commit(bank, at, count);
} else {
do {
// Direction is one bit of the opcode, so the loop
// picks it with a one-word skip.
if (command & op_write) {
store(space, at, link::rx());
tx_ack();
} else {
link::tx(load(space, bank, at));
}
++at;
} while (--count);
}
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
@@ -466,4 +717,7 @@ void send_fuses()
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;
// stack::hardware: activation is reset-only, so the reset logic's own
// SP = RAMEND stands wherever the datasheet guarantees it (the classic
// megas still get the write); a 'J' entry runs on the caller's live stack.
template struct avr::startup::entry<pureboot::run, avr::startup::stack::hardware>;

File diff suppressed because it is too large Load Diff

View File

@@ -1,396 +0,0 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 530 B on the 1284P against a 512 B slot. Its 4 B of margin was
// never spare capacity; it was the space the missing fix should have occupied.
// pureboot_autobaud_uni.cpp replaces it at 464 B with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE PURE VERSION — no inline assembly, no global register variables,
// exactly the constraints the fixed-baud loader keeps.
//
// Two source files exist for review (autobaud.md):
// this one, pure, and pureboot_autobaud_reg.cpp, which keeps the running-slot
// write guard at the cost of one global register variable. They differ only in
// where the measured unit lives and whether the guard is present.
//
// What this version trades to fit 512 B in pure C++ (the 1284 at 508), each
// licensed by "the host guarantees safety" (README.md) and the owner's approval
// to simplify the info block:
// - the measured per-bit unit lives in the two general-purpose I/O scratch
// registers (GPIOR) where the chip has them, in a static otherwise —
// reached through libavr's named register surface, so no asm and no global
// register variable; the loader stays pure;
// - the info block is slimmed to the version and the signature — the chip's
// identity — from which the host derives page size, loader base, EEPROM
// size and the addressing flags via its own chip database;
// - no running-slot write guard: the host never programs the loader's own
// slot, and a broken host bricking the target is the host's bug;
// - a single-byte activation knock: the calibration pulse already proves a
// host is present.
//
// Position independence is kept and is in fact total here: control flow is
// PC-relative, the wire carries addresses, and with the slimmed info block and
// no write guard nothing anchors on the runtime address at all.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 4;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations). It lives
// in the two adjacent general-purpose I/O scratch registers (GPIOR1:GPIOR2)
// where the chip has them — in/out reach them in one word where a static's
// lds/sts take two, and there is no .bss to clear — and in a plain static
// otherwise (the t13, m8 and m16/32 have no GPIOR). Both are pure: the named
// register surface, no inline asm, no global register variable.
constexpr bool have_gpior = hw::db.reg_index("GPIOR1") >= 0 && hw::db.reg_index("GPIOR2") >= 0;
std::uint16_t unit_backing;
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
if constexpr (Gpior)
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
else
return unit_backing;
}
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
if constexpr (Gpior) {
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
} else
unit_backing = u;
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program. No running-slot write
// guard: the host guarantees it never targets the loader's own slot (the pure
// version's one dropped safety net, licensed — README.md).
void program_flash(std::uint16_t wire_address)
{
spm::flash_address_t address;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
}
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop — its cursor then lives in the loop's own call-saved
// registers instead of being saved and restored around a call
// (the call-site-count lesson, autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16());
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,371 +0,0 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 534 B on the 1284P against a 512 B slot, so the global register
// variable it broke purity for buys nothing. pureboot_autobaud_uni.cpp replaces
// it at 464 B, strictly pure and with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE REGISTER VERSION — keeps the running-slot write guard, at the cost
// of one global register variable (r4) holding the measured unit. That variable
// is pureboot's single, deliberate break from its no-global-register-variable
// rule, present only in this variant; everything else stays pure C++.
//
// Two source files exist for review (autobaud.md):
// pureboot_autobaud_pure.cpp, fully pure but dropping the write guard, and this
// one. They differ only in where the measured unit lives (a call-saved register
// here, GPIOR/RAM there) and whether the guard is present.
//
// The register buys ~32 B over a RAM home — an outlined rx/tx reads it with one
// move where a static costs an lds — and that is what lets the write guard stay
// while the image still fits 512 B (the 1284 at 512, exactly). The unit is
// written once through a noinline setter so the store lands immediately before a
// ret: GCC otherwise deletes a global-register store whose only readers are
// callees (autobaud.md, upstream bug 6).
//
// Simplifications shared with the pure version, each licensed: a slimmed info
// block (version + signature; the host derives geometry from its chip database)
// and a single-byte activation knock (the calibration pulse already proves a
// host). Position independence is kept: control flow is PC-relative and the
// write guard anchors on the runtime return address, as the fixed-baud loader.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations) in a
// call-saved register that the serial callees read directly, so no wire path
// threads it and rx/tx reach it with a move, not a load. Written only through
// set_unit() below.
register std::uint16_t g_unit asm("r4");
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock; only the software-UART pins are a parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint8_t version = 4;
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The one store into g_unit, isolated so it lands right before the ret: a
// global-register store whose only later readers are callees is dropped
// otherwise (autobaud.md, upstream bug 6).
[[gnu::noinline]] void set_unit(std::uint16_t v)
{
g_unit = v;
}
// The autobaud software link: bit-banged with cycle-counted delays, but the
// per-bit delay is g_unit, measured from the host's calibration pulse.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into g_unit. The host sends 0xC0 — a start bit
// plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so count >> 2 is the bit
// period in _delay_loop_2's four-cycle iterations. 0 means the budget
// expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
set_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
_delay_loop_2(static_cast<std::uint16_t>(g_unit + (g_unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(g_unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
out_t::clear(); // start bit
_delay_loop_2(g_unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(g_unit);
}
out_t::set(); // stop bit
_delay_loop_2(g_unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program — except the slot this
// code is running in (`slot_high`, from run()), which is drained and left
// alone. A broken host therefore cannot brick the running loader, and a copy one
// slot lower may rewrite the resident one.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
{
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
}
if (page_high != slot_high) {
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// The high byte of the slot this copy runs at, which the write guard
// follows: the return address is a word address, so its high byte is the
// 256-word slot index, doubled back into byte terms on a byte-addressed
// chip. Taken as byteswap's low byte — the builtin already swaps the two
// stacked bytes, and the double swap folds away.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
// Measure the calibration pulse into g_unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop (autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,411 +0,0 @@
// pureboot autobaud, the unified-primitive version — VARIANT A, an explicit
// space byte. One read command and one write command carry a space selector, so
// flash, EEPROM, RAM and the fuses share a single cursor, a single transfer loop
// and a single argument decode instead of one command body each.
//
// Strictly pure: no inline assembly, no global register variables, and no GPIOR
// either — the measured unit lives in a plain static, so the loader claims no
// chip resource an application might want. (PUREBOOT_UNIT_GPIOR=1 puts it back
// in the I/O scratch registers, kept only as a measurement axis.)
//
// Against pureboot_autobaud_pure.cpp this version:
// - adds RAM read and write, which the loader has never had. Because AVR maps
// the register file and the whole I/O space into the data address space,
// that one space also gives the host arbitrary peripheral access for free;
// - collapses 'R' (read flash), 'r' (read EEPROM), 'w' (write EEPROM) and 'F'
// (fuses) — four bodies, four loops — into 'G' and 'P' over four spaces;
// - fixes the activation hang: a lone calibration pulse used to leave the
// loader blocked forever in the knock's rx(), so a stray edge on an
// unattended device wedged it in the loader and the application never ran.
//
// Position independence is kept and is total: control flow is PC-relative, the
// wire carries addresses, and nothing anchors on the runtime address.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The unit's home. A plain static by default — pureboot claims no GPIOR, so the
// application keeps both scratch registers. The GPIOR spelling is retained
// behind a macro purely so the two can be measured against each other.
#if !defined(PUREBOOT_UNIT_GPIOR)
#define PUREBOOT_UNIT_GPIOR 0
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 5;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The two SPM commands the loader still has to recognise by value, on the chips
// where it cannot issue a runtime one generically. Taken from the chip's own
// definitions rather than spelled 3 and 5 — though every part pureboot targets
// agrees on those, which is what lets the host send the raw SPMCSR byte.
constexpr std::uint8_t spm_erase = __BOOT_PAGE_ERASE;
constexpr std::uint8_t spm_write = __BOOT_PAGE_WRITE;
// The spaces a transfer can name. Flash is 0 so it is the cheap default.
//
// sp_spm is the one that is not memory: a write there hands its data byte to
// SPMCSR and fires the instruction at the given flash address, so page erase,
// page write and RWW re-enable become host-issued commands instead of a
// hardcoded tail inside 'W'. The store side already owns an address, a data
// byte and an ack, so the whole sequence costs only the fused out/spm pair. It
// also lets the host reach every other SPM operation — lock bits included —
// which the loader previously had no way to expose.
enum : std::uint8_t { sp_flash = 0, sp_eeprom = 1, sp_ram = 2, sp_fuse = 3, sp_spm = 4 };
// A transfer's selector byte is `space | bank << 4`: the low nibble names the
// space, the high nibble carries flash's third address byte (RAMPZ) on the
// chips that have one. Putting the bank here rather than widening the address
// keeps the shared cursor sixteen bits for every space — a three-byte cursor
// costs its extra increment on EEPROM and RAM reads too, which never need it.
// The host must not span a bank boundary in one transfer; it already chunks by
// page, so nothing it does today comes close.
// .noinit, not .bss: the unit is always measured before it is read, so it needs
// no zeroing — and a zeroed .bss would drag in __do_clear_bss, 18 bytes of
// startup code for a variable that is written before its first use.
[[gnu::section(".noinit")]] std::uint16_t unit_backing;
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
#if PUREBOOT_UNIT_GPIOR
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
#else
return unit_backing;
#endif
}
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
#if PUREBOOT_UNIT_GPIOR
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
#else
unit_backing = u;
#endif
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
//
// The shape of the counting loop is load-bearing, not incidental: the >> 2
// is exact only while the pulse's bit-count equals the loop's cycles per
// iteration. Both are 7 here (sbis 1 + rjmp 2 + adiw 2 + rjmp 2). Reshaping
// this loop silently changes the lock; test/pbautobaud.py is what pins it.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. That is the whole hang fix: no wait during activation is unbounded,
// so a stray calibration pulse can no longer wedge the loader.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
// One byte out of any space. The four accessors share the cursor, the loop and
// the call site — the whole point of the unified commands — so each costs only
// its own instruction rather than a body, a loop and a dispatch arm.
[[gnu::always_inline]] inline std::uint8_t load(std::uint8_t space, std::uint8_t bank, std::uint16_t at)
{
if (space == sp_eeprom)
return ee::read(at);
if (space == sp_ram)
return *reinterpret_cast<volatile std::uint8_t *>(at);
if (space == sp_fuse)
return spm::read_fuse<off>(static_cast<spm::fuse>(at));
if constexpr (word_flash)
return avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(bank) << 16) | at);
else
return avr::flash_load(reinterpret_cast<const std::uint8_t *>(at));
}
// One byte into a writable space. Flash is not one of them — it arrives a page
// at a time through 'W' — and the fuses are not writable at all.
[[gnu::always_inline]] inline void store(std::uint8_t space, [[maybe_unused]] std::uint8_t bank, std::uint16_t at,
std::uint8_t value)
{
if (space == sp_ram) {
*reinterpret_cast<volatile std::uint8_t *>(at) = value;
return;
}
if (space == sp_spm) {
// The fused store-and-fire: SPMCSR takes the data byte and the SPM
// issues against Z in the same unscheduled pair the hardware's
// four-cycle window demands, which is exactly why this is one
// primitive and not a poke of SPMCSR followed by a poke of anything
// else. A host cannot hit that window across a serial link.
// The preprocessor rather than `if constexpr` only because
// spm::detail::page_command does not exist at all where there is no
// RAMPZ, and a discarded constexpr branch outside a template is still
// name-checked. A generic spm::command() in libavr would let the
// runtime command through on every chip and retire the dispatch below.
#if defined(RAMPZ)
spm::detail::page_command(value, (static_cast<spm::flash_address_t>(bank) << 16) | at);
#else
if (value == spm_erase)
spm::erase_page<off>(at);
else if (value == spm_write)
spm::write_page<off>(at);
else if constexpr (boot_section)
spm::rww_enable<off>();
#endif
if constexpr (boot_section)
spm::wait();
return;
}
ee::write<off>(at, value);
}
// One page into the SPM buffer, and only that: the erase, the write and the RWW
// re-enable that used to follow are now three host-issued writes to sp_spm,
// which reach the same fused out/spm pair through the store path's own address
// and data. No running-slot write guard: the host guarantees it never targets
// the loader's own slot (licensed — README.md).
void program_flash([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
{
std::uint16_t z = at & ~static_cast<std::uint16_t>(page - 1);
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
#if defined(RAMPZ)
spm::fill<off>((static_cast<spm::flash_address_t>(bank) << 16) | z, word_of({low, high}));
#else
spm::fill<off>(z, word_of({low, high}));
#endif
z += 2;
} while (static_cast<std::uint8_t>(z) & (page - 1));
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock —
// both under the poll budget. An expired budget (no host) boots the
// application; anything but 'p', including the knock timing out, re-measures
// and so returns to the budgeted wait that boots it.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'W': // program one flash page: sel8, addr16, then page bytes
case 'G': // read: sel8, addr16, n8 (0 = 256)
case 'g': { // write: sel8, addr16, n8, then n bytes each acked
// The unified transfer. One decode, one cursor, one loop for every
// space and both directions — the four command bodies this replaces
// each carried their own copy of all three. Read and write are the
// same letter in the two cases, so the direction is bit 5 of the
// command and the loop tests it with a one-word skip. 'W' joins the
// same selector-and-address decode rather than keeping a word
// address of its own, which makes flash addressing uniform across
// every command that names it and costs nothing to share.
const std::uint8_t sel = link::rx();
const std::uint8_t space = sel & 0x0f;
const std::uint8_t bank = static_cast<std::uint8_t>(sel >> 4);
std::uint16_t at = rx16();
if (command == 'W') {
program_flash(bank, at);
break;
}
std::uint8_t count = link::rx();
do {
if (command & 0x20) {
store(space, bank, at, link::rx());
tx_ack();
} else
link::tx(load(space, bank, at));
++at;
} while (--count);
break;
}
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,47 +1,78 @@
#!/usr/bin/env python3
"""Position-independence lint: the two link-time facts that let the identical
image run from any slot, asserted from the built ELF.
"""Position-independence lint: the property that lets the identical image run
from any slot, asserted from the built ELF and its object.
1. No absolute jmp/call -mrelax normally guarantees it, but a branch that
1. No absolute jmp/call - -mrelax normally guarantees it, but a branch that
grows out of relaxation range would break it silently.
2. The info block within the image's first 256 bytes: 'b' rebuilds its
address as (running slot high byte : link address low byte).
2. Nothing flash-resident to address: the image is .text alone, so there is
no table whose runtime address has to be reconstructed.
3. The image is byte-identical when linked at a different base. This is
position independence itself rather than a proxy for it - an absolute
address anywhere in the image would move with the link and show up as a
differing byte.
Usage: check_pi.py <objdump> <nm> <elf> <text_start_hex>
Usage: check_pi.py <objdump> <objcopy> <cxx> <mcu> <elf> <object> <text_start_hex>
"""
import os
import re
import subprocess
import sys
import tempfile
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
objdump, nm, elf, text_start = sys.argv[1:]
objdump, objcopy, cxx, mcu, elf, obj, text_start = sys.argv[1:]
text_start = int(text_start, 0)
listing = subprocess.run([objdump, "-d", elf], capture_output=True, text=True, check=True).stdout
absolute = [
line
for line in listing.splitlines()
if re.search(r"\t(jmp|call)\t", line)
]
absolute = [line for line in listing.splitlines() if re.search(r"\t(jmp|call)\t", line)]
if absolute:
print("FAIL: absolute control flow in the image:")
print("\n".join(absolute))
sys.exit(1)
fail("absolute control flow in the image:\n" + "\n".join(absolute))
symbols = subprocess.run([nm, "-C", elf], capture_output=True, text=True, check=True).stdout
info = [line for line in symbols.splitlines() if "flash_table" in line and "::storage" in line]
if len(info) != 1:
print(f"FAIL: expected one info-block storage symbol, found {len(info)}")
sys.exit(1)
address = int(info[0].split()[0], 16)
offset = address - text_start
if not 0 <= offset < 256:
print(f"FAIL: info block at image offset {offset:#x}, must sit in the first 256 bytes")
sys.exit(1)
# Allocated flash beyond .text would be data the running copy has to find.
# Only ALLOC sections reach the device at all; .comment and the debug
# sections ride along in the ELF container and are never flashed. objdump
# prints each section's flags on the line following its header.
headers = subprocess.run([objdump, "-h", elf], capture_output=True, text=True, check=True).stdout.splitlines()
for index, line in enumerate(headers):
fields = line.split()
if len(fields) < 6 or not fields[0].isdigit():
continue
name, size = fields[1], int(fields[2], 16)
flags = headers[index + 1] if index + 1 < len(headers) else ""
if "ALLOC" not in flags or not size:
continue
if name not in (".text", ".noinit", ".bss"):
fail(f"flash-resident section {name} ({size} bytes): the image must be .text alone")
print(f"PI lint: control flow PC-relative, info block at offset {offset:#x}")
# Relink at a different base and compare the bytes.
with tempfile.TemporaryDirectory() as work:
elsewhere = text_start - 0x200 if text_start >= 0x200 else text_start + 0x200
images = []
for base, tag in ((text_start, "here"), (elsewhere, "there")):
relinked = os.path.join(work, f"{tag}.elf")
binary = os.path.join(work, f"{tag}.bin")
subprocess.run(
[cxx, f"-mmcu={mcu}", "-nostartfiles", f"-Wl,--section-start=.text={base:#x}",
"-Wl,--defsym=pureboot_app=0", "-mrelax", obj, "-o", relinked],
check=True, capture_output=True)
subprocess.run([objcopy, "-O", "binary", relinked, binary], check=True)
images.append(open(binary, "rb").read())
if len(images[0]) != len(images[1]):
fail(f"the image is {len(images[0])} B linked at {text_start:#x} and "
f"{len(images[1])} B at {elsewhere:#x} - relaxation followed the address")
if images[0] != images[1]:
differing = [i for i, (a, b) in enumerate(zip(*images)) if a != b]
fail(f"the image changes when linked at {elsewhere:#x} instead of {text_start:#x}: "
f"{len(differing)} byte(s) differ, first at offset {differing[0]:#x}")
print(f"PI lint: control flow PC-relative, .text only, identical linked at {text_start:#x} and {elsewhere:#x}")
if __name__ == "__main__":

View File

@@ -4,6 +4,9 @@ if(NOT _res EQUAL 0)
endif()
# avr-size line 2 is "<text> <data> <bss> <dec> <hex> <file>".
string(REGEX MATCH "\n[ \t]*([0-9]+)" _m "${_out}")
if(NOT _m)
message(FATAL_ERROR "could not read a .text size out of ${SIZE_TOOL}'s output for ${ELF}:\n${_out}")
endif()
set(_text ${CMAKE_MATCH_1})
if(_text GREATER LIMIT)
message(FATAL_ERROR ".text is ${_text} bytes, over the ${LIMIT}-byte boot section")

36
test/check_unit.cmake Normal file
View File

@@ -0,0 +1,36 @@
# Asserts the autobaud loader's measured unit sits where the host will read
# it (--info's measured clock - the address is wire contract). Two homes: on
# a chip with the GPIOR pair the unit lives there and the image must carry no
# RAM word for it at all; elsewhere it is the first RAM object at SRAM start.
# Run as
# cmake -DOBJDUMP=... -DELF=... -DRAM_START=<data address> [-DGPIOR=<data address>]
# -P check_unit.cmake
execute_process(COMMAND ${OBJDUMP} -t ${ELF} OUTPUT_VARIABLE _syms RESULT_VARIABLE _res)
if(NOT _res EQUAL 0)
message(FATAL_ERROR "${OBJDUMP} -t ${ELF} failed")
endif()
# The symbol line: "00800100 l O .noinit 00000002 <mangled>6m_unitE".
string(REGEX MATCH "\n0*([0-9a-f]+)[^\n]+[ \t][^ \t\n]*6m_unitE\n" _line "${_syms}")
if(GPIOR)
if(_line)
message(FATAL_ERROR "m_unit RAM symbol present although the unit's home is GPIOR ${GPIOR} - "
"the host peeks the pair, and a RAM copy would be dead weight")
endif()
message(STATUS "no m_unit RAM object - the unit lives in the GPIOR pair at ${GPIOR}")
return()
endif()
if(NOT _line)
message(FATAL_ERROR "no m_unit symbol in ${ELF} - is this the autobaud loader?")
endif()
# AVR data-space symbols carry the 0x800000 VMA offset.
math(EXPR _want "0x800000 + ${RAM_START}" OUTPUT_FORMAT HEXADECIMAL)
math(EXPR _have "0x${CMAKE_MATCH_1}" OUTPUT_FORMAT HEXADECIMAL)
if(NOT _have STREQUAL _want)
message(FATAL_ERROR "m_unit sits at ${_have}, ram_start is ${_want} - the host peeks ram_start")
endif()
message(STATUS "m_unit at ${_have} == ram_start")

View File

@@ -7,103 +7,122 @@
// SPM genuinely writes avr->flash on the mega cores, so on exit (or SIGTERM)
// we dump the flash image to a file for a ground-truth cross-check against
// what the client read back through the bootloader.
#include <signal.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <array>
#include <csignal>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <print>
#include <unistd.h>
// The parts headers (uart_pty.h) carry no C++ linkage guards of their own,
// unlike simavr's core headers - the block covers both harmlessly.
extern "C" {
#include "avr_uart.h"
#include "sim_avr.h"
#include "sim_elf.h"
#include "uart_pty.h"
}
static avr_t *avr;
static uart_pty_t uart_pty;
static const char *dump_path;
namespace {
static void finish(int sig)
avr_t *avr;
uart_pty_t uart_pty;
const char *dump_path;
[[noreturn]] void finish(int)
{
(void)sig;
if (dump_path) {
FILE *f = fopen(dump_path, "wb");
std::FILE *f = std::fopen(dump_path, "wb");
if (f) {
fwrite(avr->flash, 1, avr->flashend + 1, f);
fclose(f);
std::fwrite(avr->flash, 1, avr->flashend + 1, f);
std::fclose(f);
}
}
uart_pty_stop(&uart_pty);
_exit(0);
}
} // namespace
int main(int argc, char *argv[])
{
if (argc < 3) {
fprintf(stderr, "usage: %s <tsb.elf> <boot_base_hex> [flash_dump.bin]\n", argv[0]);
std::println(stderr, "usage: {} <tsb.elf> <boot_base_hex> [flash_dump.bin]", argv[0]);
return 2;
}
uint32_t boot_base = (uint32_t)strtoul(argv[2], NULL, 0);
dump_path = argc >= 4 ? argv[3] : NULL;
auto boot_base = static_cast<std::uint32_t>(std::strtoul(argv[2], nullptr, 0));
dump_path = argc >= 4 ? argv[3] : nullptr;
avr = avr_make_mcu_by_name("atmega328p");
if (!avr) {
fprintf(stderr, "device: no ATmega328P core\n");
std::println(stderr, "device: no ATmega328P core");
return 1;
}
avr_init(avr);
avr->frequency = 16000000;
// Real flash powers up erased (0xff); the app region must look erased
// before the bootloader programs it.
memset(avr->flash, 0xff, avr->flashend + 1);
std::memset(avr->flash, 0xff, avr->flashend + 1);
// simavr's ELF loader flattens the flash base to 0 (it expects an app at
// 0x0), but it hands back the boot code in fw.flash; place it at the boot
// section base ourselves and enter there (BOOTRST is not modelled).
elf_firmware_t fw = {0};
elf_firmware_t fw{};
if (elf_read_firmware(argv[1], &fw) != 0) {
fprintf(stderr, "device: cannot read %s\n", argv[1]);
std::println(stderr, "device: cannot read {}", argv[1]);
return 1;
}
memcpy(avr->flash + boot_base, fw.flash, fw.flashsize);
// An image that runs past flash end cannot execute on hardware, and a
// naive copy of it would smash the heap beyond avr->flash - after which
// the simulation misbehaves in ways that point everywhere but here.
// Refuse it loudly instead.
if (boot_base + fw.flashsize > avr->flashend + 1) {
std::println(stderr, "device: {} B at {:#x} runs past flash end {:#x} - image does not fit its slot",
fw.flashsize, boot_base, avr->flashend);
return 1;
}
std::memcpy(avr->flash + boot_base, fw.flash, fw.flashsize);
avr->pc = boot_base;
avr->codeend = avr->flashend;
// Optional: seed the config page (one page below the boot section) with a
// hex byte string, so the password gate and emergency erase can be tested.
// Layout: [appjump lo][appjump hi][timeout][password...][0xff].
const char *cfg = getenv("TSB_CONFIG");
const char *cfg = std::getenv("TSB_CONFIG");
if (cfg) {
uint32_t app_end = boot_base - 128; // config page sits directly below the boot code
std::uint32_t app_end = boot_base - 128; // config page sits directly below the boot code
for (int i = 0; cfg[i] && cfg[i + 1]; i += 2) {
char b[3] = {cfg[i], cfg[i + 1], 0};
avr->flash[app_end + i / 2] = (uint8_t)strtoul(b, NULL, 16);
const std::array pair{cfg[i], cfg[i + 1], '\0'};
avr->flash[app_end + i / 2] = static_cast<std::uint8_t>(std::strtoul(pair.data(), nullptr, 16));
}
}
// POLL_SLEEP makes simavr usleep(1) on every status-register read while the
// UART is idle a host-CPU-saving hack that models no hardware and paces a
// UART is idle - a host-CPU-saving hack that models no hardware and paces a
// tight-polling loader (one that releases TX between bytes, as one-wire does)
// in real time, distorting protocol timing. Clear it so the loader runs at
// true cycle speed.
uint32_t uflags = 0;
std::uint32_t uflags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS('0'), &uflags);
uflags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS('0'), &uflags);
uart_pty_init(avr, &uart_pty);
uart_pty_connect(&uart_pty, '0');
printf("TSB_PTY %s\n", uart_pty.pty.slavename);
fflush(stdout);
std::println("TSB_PTY {}", uart_pty.pty.slavename);
std::fflush(stdout);
signal(SIGTERM, finish);
signal(SIGINT, finish);
std::signal(SIGTERM, finish);
std::signal(SIGINT, finish);
for (;;) {
int state = avr_run(avr);
if (state == cpu_Done || state == cpu_Crashed)
if (state == cpu_Done || state == cpu_Crashed) {
break;
}
}
finish(0);
return 0;
}

View File

@@ -1,16 +1,20 @@
// Test-fixture application for the pureboot protocol tests: prints "APP" on
// the chip's serial link (the same link the loader uses) the proof that
// the chip's serial link (the same link the loader uses) - the proof that
// the loader's hand-over, and on the tinies the host's reset-vector
// surgery, actually launched it. Linked normally (crt, vectors at 0); on
// the tinies its reset vector is the rjmp the host re-homes.
//
// On the hardware-USART link it then listens, and an 'L' makes it jump into
// the resident loader the application-owned loader entry a
// the resident loader - the application-owned loader entry a
// BOOTRST-unprogrammed mega relies on (reset always boots the application
// there), exercised by the self-update tests. The software link idles:
// reset reaches those loaders through the patched vector (or the runner
// models BOOTRST), so the application owes them nothing.
//
// PUREBOOT_HANDOVER drops the listening and jumps straight in, leaving the
// USART enabled behind it - the hand-over state a loader bit-banging on that
// USART's own pins has to survive.
//
// The fixture speaks the deployment its loader was built for: the same
// PUREBOOT_* defines configure it, and without them it assumes the stock
// deployment (the crystal/RC clock table below, the chip's natural link).
@@ -26,10 +30,12 @@ consteval avr::hertz_t clock()
return avr::hertz_t{PUREBOOT_CLOCK_HZ};
#else
auto name = std::string_view{avr::hw::db.name};
if (name.starts_with("ATtiny13"))
if (name.starts_with("ATtiny13")) {
return 9.6_MHz;
if (name.starts_with("ATtiny"))
}
if (name.starts_with("ATtiny")) {
return 8_MHz;
}
return 16_MHz;
#endif
}
@@ -37,6 +43,9 @@ consteval avr::hertz_t clock()
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_USART)
#define PUREBOOT_USART 0
#endif
@@ -46,7 +55,7 @@ consteval bool use_hardware()
#if defined(PUREBOOT_SOFT_SERIAL)
return false;
#else
return avr::hw::db.has_instance("USART0") || avr::hw::db.has_instance("USART");
return avr::uart::has_usart<0>();
#endif
}
@@ -59,29 +68,57 @@ struct link {
#else
static constexpr avr::baud_t baud{115200};
#endif
using tx_t = avr::uart::usart<'0' + PUREBOOT_USART, C, {.baud = baud, .max_baud_error = 2.5_pct}>;
// The fixture speaks whatever rate the loader was built for, stock
// 115200 at 16 MHz included, which sits past the receiver-tolerance
// table's bound - the same deployment envelope the loader itself states.
using tx_t = avr::uart::usart<PUREBOOT_USART, C, {.baud = baud, .allow_baud_error = true}>;
static void init()
{
avr::init<tx_t>();
}
static void tx(char c)
{
tx_t::write(static_cast<std::uint8_t>(c));
}
// The loader sits in the top slot - 512 bytes on every chip. The jump
// takes a word address, which is what makes the >64 KiB chips' entry
// reachable through a 16-bit pointer at all.
static void enter_loader()
{
constexpr std::uint32_t slot = 512;
reinterpret_cast<void (*)()>(static_cast<std::uint16_t>((avr::hw::db.mem.flash_size - slot) / 2))();
}
[[noreturn]] static void idle()
{
// 'L' hands back to the loader in the top slot — 512 bytes on every
// chip. The jump takes a word address, which is what makes the
// >64 KiB chips' entry reachable through a 16-bit pointer at all.
constexpr std::uint32_t slot = 512;
#if defined(PUREBOOT_HANDOVER)
// Hand back at once, with this USART still enabled - the state that
// leaves a bit-banged loader on its pins mute unless the loader
// releases it. Unconditional because there is no command wire to
// wait on: that loader's link is the pins, not this peripheral.
enter_loader();
__builtin_unreachable();
#else
for (;;) {
auto command = tx_t::read_blocking();
if (command == 'L')
reinterpret_cast<void (*)()>(static_cast<std::uint16_t>((avr::hw::db.mem.flash_size - slot) / 2))();
if (command == 'L') {
enter_loader();
}
// 'D' leaves every word of the SPM page buffer dirty, so that a
// following 'L' enters the loader with the buffer it never clears.
// Hardware refuses application-section SPM on a boot-sectioned
// part; simavr dispatches it anyway, which is the whole reason the
// state is constructible - the stated section is the compilable
// fiction that matches what the simulator runs.
if (command == 'D') {
for (std::uint16_t at = 0; at < avr::spm::page_bytes; at += 2)
avr::spm::fill(at, 0xdead);
const auto open = avr::spm::page::begin<avr::spm::from::boot_section>(0);
for (std::uint16_t at = 0; at < avr::spm::page_bytes; at += 2) {
avr::spm::fill(open, at, 0xdead);
}
tx('D');
}
}
#endif
}
};
@@ -92,15 +129,49 @@ struct link<C, false> {
#else
static constexpr avr::baud_t baud{57600};
#endif
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, baud>;
// A shared-pin deployment (RX == TX) banners as a guest on its own line:
// the pull-up input is the released line, the transmitter takes the pin
// for exactly one frame per byte - the shape a real one-wire application
// beside this loader uses.
static constexpr bool one_wire = avr::PUREBOOT_RX == avr::PUREBOOT_TX;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, baud, one_wire>;
static void init()
{
// The guest transmitter configures no pin; the released line - the
// pull-up input a receiver would own - is established here.
if constexpr (one_wire) {
avr::init<avr::io::input<avr::PUREBOOT_TX, avr::io::pull::up>, tx_t>();
} else {
avr::init<tx_t>();
}
}
static void tx(char c)
{
tx_t::write(static_cast<std::uint8_t>(c));
}
[[noreturn]] static void idle()
{
#if defined(PUREBOOT_HEARTBEAT)
// Repeat the banner forever, which turns the fixture into a fixed
// cycles-per-bit transmitter: `tools/pbrig.py rate` sweeps the host rate
// against it to find the part's true bit rate, and from that the clock
// its RC oscillator is really running at. Only the *bit* timing carries
// the measurement - the delay merely spaces the lines out, so its own
// error does not matter. Software link only: the hardware-link idle owes
// the self-update tests a command loop, and a crystal deployment has
// nothing to measure.
while (true) {
tx('A');
tx('P');
tx('P');
tx('\r');
tx('\n');
dev::delay<50_ms>();
}
#else
while (true) {
}
#endif
}
};
@@ -108,9 +179,14 @@ struct link<C, false> {
int main()
{
avr::init<typename link<dev::clock>::tx_t>();
link<dev::clock>::init();
#if !defined(PUREBOOT_HANDOVER)
link<dev::clock>::tx('A');
link<dev::clock>::tx('P');
link<dev::clock>::tx('P');
#endif
// The hand-over fixture stays silent: nothing is listening on the USART it
// brings up - the loader it hands to speaks those pins directly - so its
// banner would be a write into a peer that does not exist.
link<dev::clock>::idle();
}

View File

@@ -1,21 +1,23 @@
#!/usr/bin/env python3
"""End-to-end autobaud test: drive an autobaud loader in simavr through the
calibration handshake and a flash + EEPROM + fuse round-trip, cross-checked
against the simulator's ground-truth memory then repeat at a second F_CPU with
against the simulator's ground-truth memory - then repeat at a second F_CPU with
the *same* loader binary, which is the property autobaud exists for: one
clock-agnostic image that locks onto whatever rate the host sends.
Usage: pbautobaud.py <device_bin> <loader_elf> <mcu> <base_hex> <page>
<app_bin> <app_hz> <app_baud> <tool_py> <workdir>
<app_bin> <app_hz> <app_baud> <tool_py> <workdir> [link]
The loader is a software-serial build on PB0/PB1 (pureboot_add_autobaud's
default), so the runner drives it over the GPIO⇄pty bridge (-l sw:B0,B1). The
app fixture is built for (app_hz, app_baud); the hand-over is checked at that
point, and a second point at half the clock proves the lock is measured, not
baked in.
The loader is a software-serial build, driven over the GPIO<->pty bridge; the
optional link overrides the default -l sw:B0,B1 - RX == TX in it is the
one-wire deployment, and every session then runs with the host's echo
discard on. The app fixture is built for (app_hz, app_baud); the hand-over
is checked at that point, and a second point at half the clock proves the
lock is measured, not baked in.
"""
import os
import re
import sys
import time
@@ -26,8 +28,12 @@ def fail(message):
def main():
(device_bin, elf, mcu, base_hex, page, app_bin, app_hz, app_baud, tool, workdir) = sys.argv[1:]
args = sys.argv[1:]
link = args.pop() if len(args) == 11 else "sw:B0,B1"
(device_bin, elf, mcu, base_hex, page, app_bin, app_hz, app_baud, tool, workdir) = args
base, page, app_hz, app_baud = int(base_hex, 0), int(page), int(app_hz), int(app_baud)
one_wire = re.fullmatch(r"sw:([A-H][0-7]),\1(@[01])?", link) is not None
extra = ("--one-wire",) if one_wire else ()
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
@@ -39,7 +45,7 @@ def main():
open(ee_path, "wb").write(ee_image)
# The geometry the surgery planner needs, from the chip class the runner is
# told the same derivation pbtest.py makes: the boot-sectioned megas need
# told - the same derivation pbtest.py makes: the boot-sectioned megas need
# no vector surgery, the tinies and the boot-section-less m48s do, and the
# large chips speak word addresses.
mega = mcu.startswith("atmega")
@@ -54,19 +60,32 @@ def main():
"""One clock point: reset, calibrate + knock, program, verify against the
simulator's own flash, and (at the app's point) hand over to the fixture."""
dump = os.path.join(workdir, f"flash_{label}.bin")
device = pbsim.Device(device_bin, elf, mcu, str(hz), base_hex, page, baud, dump, link="sw:B0,B1")
device = pbsim.Device(device_bin, elf, mcu, str(hz), base_hex, page, baud, dump, link=link)
try:
# The host tool, in autobaud mode, sends the 0xC0 calibration pulse
# and a single knock at `baud`; the loader locks to it.
out = pbsim.run_tool(tool, device.pty, baud, "--autobaud", "--info", "--fuses",
"--flash", app_bin, "--eeprom", ee_path, "--stay")
out = pbsim.run_tool(tool, device.pty, baud, *extra, "--autobaud", "--info", "--clock", str(hz),
"--fuses", "--flash", app_bin, "--eeprom", ee_path, "--stay")
for needed in ("version", "signature", "fuses", "verify:", "stays"):
if needed not in out:
fail(f"{label}: session output lacks {needed!r}\n{out}")
# The measured clock, decoded from the unit at whichever home this
# version keeps it in. The runner's clock is exact, so the figure
# must land inside the
# encoding's own envelope: the loader floors the bit period to
# 4-cycle spin granules after an 8-cycle discount, and the edge
# poll can shave a few cycles more - one granule of slack below
# the true clock, none above (in cycles per bit, times the rate).
measured = re.search(r"measured\s+(\d+) Hz", out)
if not measured:
fail(f"{label}: --info lacks the measured clock\n{out}")
measured = int(measured.group(1))
if not hz - 19 * baud <= measured <= hz + 4 * baud:
fail(f"{label}: measured clock {measured} Hz is {measured - hz:+d} off the true {hz}")
# Read both memories back over the locked link and check them.
read_flash = os.path.join(workdir, f"rf_{label}.bin")
read_eeprom = os.path.join(workdir, f"re_{label}.bin")
out = pbsim.run_tool(tool, device.pty, baud, "--autobaud", "--verify-flash", app_bin,
out = pbsim.run_tool(tool, device.pty, baud, *extra, "--autobaud", "--verify-flash", app_bin,
"--verify-eeprom", ee_path, "--read-flash", read_flash,
"--read-eeprom", read_eeprom, "--stay")
if out.count("verify:") != 2:
@@ -75,34 +94,37 @@ def main():
fail(f"{label}: EEPROM read-back mismatch")
if hand_over:
# Regression: a calibration pulse with no knock behind it must
# not wedge the loader. The knock's edge wait used to be
# unbudgeted, so one stray low pulse EMI, or a host that opens
# the port and never knocks — held the loader forever and the
# application never ran. The whole activation is bounded now, so
# the window closes and the app boots; the banner is the proof.
# A calibration pulse with no knock behind it must not wedge
# the loader: the whole activation is bounded, including the
# knock's edge wait, so one stray low pulse - EMI, or a host
# that opens the port and never knocks - closes the window and
# boots the application. The banner is the proof.
# (The pause lets the loader reach its measurement loop, so the
# pulse is genuinely seen and the test cannot pass vacuously.)
device.reset()
port = pb.Port(device.pty, baud)
if one_wire:
port = pb.OneWirePort(port)
try:
time.sleep(0.2)
port.write(bytes((pb.CALIBRATE,)))
# Accumulate rather than match exactly: the reset leaves the
# idle line a framing artefact ahead of the banner, which is
# noise here the question is only whether the app ran.
# noise here - the question is only whether the app ran.
seen = b""
deadline = time.monotonic() + 180.0
while b"APP" not in seen and time.monotonic() < deadline:
seen += port.read_available(1.0)
if b"APP" not in seen:
fail(f"{label}: lone calibration pulse wedged the loader app never bannered, saw {seen!r}")
fail(f"{label}: lone calibration pulse wedged the loader - app never bannered, saw {seen!r}")
print(f" {label}: lone calibration pulse does not wedge the loader")
finally:
port.close()
device.reset()
port = pb.Port(device.pty, baud)
if one_wire:
port = pb.OneWirePort(port)
try:
loader = pb.Loader(port)
live = loader.connect_autobaud(15)
@@ -114,13 +136,13 @@ def main():
# the stack at the top. Reading it back over the same
# locked link proves both directions of the new space.
probe = bytes(range(0x30, 0x40))
loader.write_ram(0x0200, probe)
if loader.read_ram(0x0200, len(probe)) != probe:
loader.write_data(0x0200, probe)
if loader.read_data(0x0200, len(probe)) != probe:
fail(f"{label}: RAM round-trip mismatch")
# The register file and the I/O space share the data
# address space on AVR, so the same command reaches a
# peripheral register. SPMCSR reads back as idle here.
verbose_ram = loader.read_ram(0x0200, 4)
verbose_ram = loader.read_data(0x0200, 4)
print(f" {label}: RAM read/write ok ({verbose_ram.hex()})")
loader.run_application()
banner = port.read_exact(3, 5.0)
@@ -141,13 +163,47 @@ def main():
print(f" {label}: locked at {hz} Hz / {baud} Bd, flash+EEPROM verified"
+ (", hand-over ok" if hand_over else ""))
def must_lock(hz, baud, label):
"""The calibration alone, at a tight bit period. Nothing is programmed -
the question is only whether the loader can still measure the pulse."""
dump = os.path.join(workdir, f"flash_{label}.bin")
device = pbsim.Device(device_bin, elf, mcu, str(hz), base_hex, page, baud, dump,
link=link)
try:
port = pb.Port(device.pty, baud)
if one_wire:
port = pb.OneWirePort(port)
try:
live = pb.Loader(port).connect_autobaud(15)
if live.version != pb.NEWEST_LOADER:
fail(f"{label}: loader reports pureboot {live.version}")
finally:
port.close()
finally:
device.stop()
print(f" {label}: locked at {hz} Hz / {baud} Bd ({hz / baud:.0f} cycles a bit)")
# The app fixture is built for one clock; the hand-over banners there. A
# second point at double that clock, same loader binary, proves the lock is
# measured, not baked in the whole point of autobaud. (Doubling keeps the
# measured, not baked in - the whole point of autobaud. (Doubling keeps the
# bit period healthy; halving would drop it below the software UART's floor.)
round_trip(app_hz, app_baud, "clock-a", hand_over=True)
round_trip(app_hz * 2, app_baud, "clock-b", hand_over=False)
print("pbautobaud: calibration lock and flash/EEPROM/fuse round-trip pass at both clocks")
# Both points above sit near 100 cycles a bit, which is comfortable. The
# calibration's real floor is far tighter, and it is worth a gate: measured
# here, the lock is solid down to ~36 cycles a bit and fails outright by ~31
# - a sharp edge, not a fraying one. This pins the tightest standard rate the
# fixture's clock reaches, so a change that raises the floor is caught.
#
# It does *not* bound what a real deployment can use. On silicon the
# oscillator's own jitter costs roughly a factor of two: an ATtiny13A on its
# factory RC trim was reliable at ~118 cycles a bit and already locking only
# 1 attempt in 5 by ~59, which no exact-clock simulation can show. The
# deployable envelope is a README matter; this is the logic's floor.
must_lock(app_hz, app_baud * 2, "tight-bit")
print("pbautobaud: calibration lock and flash/EEPROM/fuse round-trip pass at both clocks, "
"and the tight bit period still locks")
if __name__ == "__main__":

View File

@@ -1,12 +1,12 @@
#!/usr/bin/env python3
"""Dirty-page-buffer acceptance test: with no discard in the loader, a page
filled over words an earlier writer left takes those instead. The whole
contract is asserted a bare verify sees the corruption, the repairing
contract is asserted - a bare verify sees the corruption, the repairing
verify fixes it in one rewrite, and it stays fixed.
The state is reached the one way the loader cannot prevent: an application
dirties the buffer and jumps in with no reset between. Boot-sectioned megas
forbid that outright (SPM runs only from the boot section, Atmel-8271 §26.2),
forbid that outright (SPM runs only from the boot section, Atmel-8271 section 26.2),
but simavr dispatches SPM from anywhere, which is what makes it constructible.
Usage: pbdirty.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
@@ -66,7 +66,7 @@ def main():
fail(f"the read-back failed, but not at verify: {error}")
else:
# Either the fixture no longer dirties the buffer, or the loader
# clears it again in which case this test's premise is gone.
# clears it again - in which case this test's premise is gone.
fail("programming over a dirty page buffer came back clean")
# What the programming path uses: one rewrite settles it, and it stays

124
test/pbglitch.py Normal file
View File

@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""A link that damages bytes, deliberately and reproducibly, against the seal
that exists for it.
The board this was written for loses and mangles bytes on its own serial path,
and the failure that made it matter - a page-fill byte lost, the stream one
byte out, a page-address byte arriving where an SPMCSR value belongs - is not
reachable by asking a healthy link nicely. So the damage is injected here, at
a named byte index rather than a probability: a failing case is a case that
fails again.
Every check is a pair. The same bit flipped in the same field is applied on
one side of the seal and then the other: *after* the host seals the header,
which is a mangled command and must be refused, and *before*, which is a
well-formed command for something else and must be obeyed. Only the pair
proves anything - a test that showed the refusal alone would pass against a
loader that had simply stopped doing SPM, and one that showed the corruption
alone would not say what caught it.
Usage: pbglitch.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <tool_py> <workdir>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def frame(pb, op, space, address, count, damage=None, before_seal=False):
"""A sealed command header, optionally with one byte damaged.
`damage` is (index, mask). Applied before the seal is computed it produces
a valid command for whatever the damaged fields now say; applied after, a
command whose seal no longer matches its own body - which is the shape a
link fault actually has."""
head = bytearray((op, pb.selector(space, address), address & 0xFF,
(address >> 8) & 0xFF, count & 0xFF))
if damage and before_seal:
head[damage[0]] ^= damage[1]
seal = pb.SEAL
for byte in head:
seal ^= byte
out = bytearray(head + bytes((seal,)))
if damage and not before_seal:
out[damage[0]] ^= damage[1]
return bytes(out)
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, tool, workdir = sys.argv[1:]
base, page, baud = int(base_hex, 0), int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "dump.bin")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump)
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
info = loader.connect(25)
if info.version < pb.SEALED_LOADER:
fail(f"this test is for pureboot {pb.SEALED_LOADER} and later, not {info.version}")
# A page of known bytes to watch. Everything below aims at it, so
# "nothing happened" is a readable claim rather than an absence.
marker = bytes((0x40 + (i & 0x3F)) for i in range(page))
loader.write_page(0, marker)
if loader.read_flash(0, page) != marker:
fail("the marker page did not survive an undamaged write")
# Every field of the header, one bit each. A damaged seal must be
# refused, the loader must re-prompt, and the page must be untouched -
# and it is the erase being aimed at it, so a single escape is visible.
for index in range(6):
bad = frame(pb, pb.OP_WRITE, pb.SP_SPM, 0, pb.SPM_ERASE, damage=(index, 0x01))
port.write(bad)
answer = port.read_exact(1, 5.0)
if answer != pb.NAK:
fail(f"a header damaged in byte {index} was not refused (got {answer.hex()})")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail(f"no prompt after refusing a header damaged in byte {index}")
if loader.read_flash(0, page) != marker:
fail(f"damage in byte {index} reached flash - the marker page changed")
# The fill, whose payload is the protocol's one unacked burst: a
# refused fill must be refused *before* the page is sent, or the host
# is left pushing 128 bytes into a loader reading commands. Nothing is
# sent after the verdict here, and the very next command must be
# understood - that is the whole claim.
bad = frame(pb, pb.OP_FILL, pb.SP_FLASH, 0, page, damage=(3, 0x80))
port.write(bad)
if port.read_exact(1, 5.0) != pb.NAK:
fail("a damaged fill header was not refused")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("no prompt after refusing a damaged fill header")
if loader.identity().raw != info.raw:
fail("the loader was out of step after refusing a fill")
# The pair's other half. The identical flip, applied before the seal:
# a well-formed erase of the page one bit away from the one intended.
# It must be obeyed - otherwise the refusals above prove nothing about
# the seal and only that this loader stopped erasing.
port.write(frame(pb, pb.OP_WRITE, pb.SP_SPM, 0, pb.SPM_ERASE,
damage=(2, 0x01), before_seal=True))
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("a correctly sealed erase was refused")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("no prompt after a correctly sealed erase")
if loader.read_flash(0, page) != b"\xff" * page:
fail("the sealed erase did not reach flash - the marker page is intact")
port.close()
finally:
device.stop()
print("pbglitch: damaged headers are refused, the identical damage sealed is obeyed")
main()

79
test/pbmute.py Normal file
View File

@@ -0,0 +1,79 @@
#!/usr/bin/env python3
"""Hand-over with a USART left enabled on the loader's own pins.
A software or autobaud link deployed on a USART's TxD is mute if an
application hands over with that USART still enabled: TXEN keeps the USART
owning the pin, so the bit-banged transmitter's port writes go nowhere and the
loader receives and obeys while answering nothing. The link's init releases it.
The state is reached the way silicon reaches it - an application that sets up
its USART and jumps in with no reset between, so nothing clears UCSRnB for it.
The pin ownership itself is modelled by the device runner: simavr wires a
USART through IRQs alone and never takes the pin from the port, so without
that the mute could not happen here at all (test/pureboot_device.cpp).
Usage: pbmute.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <app_bin> <tool_py> <workdir> <link>
"""
import os
import re
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, app_bin, tool, workdir, link = sys.argv[1:]
page, baud = int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
if "@" not in link:
fail(f"the link {link} names no owning USART - nothing would be under test")
# A shared line (RX == TX) echoes the host's own bytes; discard them the
# way the shipped --one-wire mode does.
one_wire = re.fullmatch(r"sw:([A-H][0-7]),\1@[01]", link) is not None
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "dump.bin")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, link=link)
try:
port = pb.Port(device.pty, baud)
if one_wire:
port = pb.OneWirePort(port)
loader = pb.Loader(port)
loader.connect(25)
resident = loader.info.version
# Install the fixture and let it take over. It brings up the USART
# that owns these pins and jumps straight back in.
pb.op_flash(loader, app_bin, erase=False, verify=True)
loader.run_application()
# The loader is running again with that USART enabled behind it. Only
# the release makes it audible; without it the connect times out.
loader = pb.Loader(port)
try:
loader.connect(25)
except pb.Error as error:
fail(f"the loader never answered after the hand-over - the USART still owns its TX pin ({error})")
if loader.info.version != resident:
fail(f"identity changed across the hand-over: {resident} then {loader.info.version}")
# Answering is not enough: it has to still be a working loader.
pb.verify_pages(loader, pb.plan_flash(open(app_bin, "rb").read(), loader.info))
port.close()
finally:
device.stop()
print("pbmute: a loader on a USART's own pins answers after a hand-over that left it enabled")
if __name__ == "__main__":
main()

46
test/pbosccal.py Normal file
View File

@@ -0,0 +1,46 @@
#!/usr/bin/env python3
"""The build-time OSCCAL trim, observed through the wire: a loader built with
the OSCCAL axis holds the trim register at the built byte from its first
prompt on - the write sits at the top of run(), ahead of the WDRF bail, so
every path out of reset runs on the corrected clock. simavr's clock does not
follow OSCCAL, which is what makes the value assertable at all: the register
is plain state there, and the peek must return exactly what the build
declared rather than whatever the oscillator needed.
Usage: pbosccal.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <osccal_addr> <osccal_value> <tool_py> <workdir>
[link]
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
args = sys.argv[1:]
link = args.pop() if len(args) == 12 else None
(device_bin, elf, mcu, hz, base_hex, page, baud, addr, value, tool, workdir) = args
addr, value, baud = int(addr, 0), int(value, 0), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "flash_dump.bin")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, link=link)
try:
out = pbsim.run_tool(tool, device.pty, baud, "--peek", f"{addr:#x}:1")
want = f"{addr:#06x} {value:02x}"
if want not in out:
fail(f"OSCCAL at {addr:#x} did not read back {value:#04x}:\n{out}")
finally:
device.stop()
print("OK")
if __name__ == "__main__":
main()

View File

@@ -3,11 +3,10 @@
canonical slot must still be a working loader, and the ordinary
--update-loader flow must put a build into the top slot from there.
Two positions. Address 0, a raw .bin handed to a programmer: the staging
install and the word-0 redirect run from copies outside page 0's slot, so the
running-slot guard never blocks them. And the staging slot itself, where a
loader already sitting there IS the staging copy — recognized by its embedded
block and left in place, then streaming the new resident like any staged copy.
Two positions. Address 0, a raw .bin handed to a programmer. And the staging
slot itself, where a loader already sitting there IS the staging copy -
recognized by its embedded block and left in place, then streaming the new
resident like any staged copy.
Usage: pbrehome.py <device_bin> <pureboot_elf> <update_bin> <mcu> <hz>
<base_hex> <page> <baud> <app_bin> <tool_py> <workdir>
@@ -22,7 +21,7 @@ def fail(message):
sys.exit(1)
def rehome_from(pbsim, pb, device_bin, elf, place_hex, guard_probe, update_bin, base, page, baud, app_bin, workdir,
def rehome_from(pbsim, pb, device_bin, elf, place_hex, update_bin, base, page, baud, app_bin, workdir,
mcu, hz):
"""Place the loader at `place_hex`, heal through --update-loader, flash
the application, expect the banner."""
@@ -36,15 +35,7 @@ def rehome_from(pbsim, pb, device_bin, elf, place_hex, guard_probe, update_bin,
loader = pb.Loader(port)
info = loader.connect(25)
if info.base != base:
fail(f"the misplaced copy reports base {info.base:#06x} the info block must stay canonical")
# The accidental slot still guards itself; re-homing rides on the
# canonical slots being writable from it.
probe = int(guard_probe, 0)
before = loader.read_flash(probe, info.page)
loader.write_page(probe, bytes(info.page))
if loader.read_flash(probe, info.page) != before:
fail("the misplaced copy's guard let its own slot change")
fail(f"the misplaced copy reports base {info.base:#06x} - the info block must stay canonical")
# The ordinary update flow puts the build into the top slot.
pb.op_update_loader(loader, 25, update_bin, state, None)
@@ -76,16 +67,15 @@ def main():
os.makedirs(workdir, exist_ok=True)
# Address 0: the raw-.bin-to-a-programmer accident. The guard probe is
# the copy's own page 0.
rehome_from(pbsim, pb, device_bin, elf, "0x0", "0x0", update_bin, base, page, baud, app_bin, workdir, mcu, hz)
# Address 0: the raw-.bin-to-a-programmer accident.
rehome_from(pbsim, pb, device_bin, elf, "0x0", update_bin, base, page, baud, app_bin, workdir, mcu, hz)
print("re-home from address 0: converged")
# The staging slot: erased flash with the loader sitting exactly where
# a staging copy would the tool must leave it in place and let it
# a staging copy would - the tool must leave it in place and let it
# stream the (different) update build into the resident slot.
stage = base - pb.SLOT
rehome_from(pbsim, pb, device_bin, elf, hex(stage), hex(stage), update_bin, base, page, baud, app_bin, workdir,
rehome_from(pbsim, pb, device_bin, elf, hex(stage), update_bin, base, page, baud, app_bin, workdir,
mcu, hz)
print("re-home from the staging slot: converged")

View File

@@ -1,9 +1,8 @@
#!/usr/bin/env python3
"""Position-independence acceptance test: the identical binary, flashed one
slot below the resident, must serve the complete command set from there. The
info block must come back byte-identical, the write guard must refuse the
staged copy's own slot and permit the resident's, and the staged copy must be
able to rewrite the resident verbatim.
info block must come back byte-identical, and the staged copy must be able to
rewrite the resident verbatim - which is the whole of what relocation is for.
Usage: pbreloc.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <tool_py> <workdir>
@@ -64,20 +63,16 @@ def main():
if loader.read_eeprom(0, len(pattern)) != pattern:
fail("EEPROM round-trip through the staged copy")
# The guard, both ways: its own slot refused (drained, unchanged), the
# resident slot writable. The refusal leaves its drained words in the
# SPM buffer, so the write that follows may take them — and clears
# them by writing, so the retry must not.
before = loader.read_flash(stage, page)
loader.write_page(stage, bytes(page))
if loader.read_flash(stage, page) != before:
fail("the staged copy's guard let its own slot change")
# The resident slot, written from the copy standing beside it - the
# whole point of relocating. There is no running-slot guard to probe
# against: nothing here refuses an address, and a copy that erases the
# page it is executing from does not come back to report it.
# pbselfwrite.py gates that direction on a device it is allowed to
# destroy.
marker = bytes((i * 3) & 0xFF for i in range(page))
loader.write_page(base, marker)
if loader.read_flash(base, page) != marker:
loader.write_page(base, marker)
if loader.read_flash(base, page) != marker:
fail("the staged copy could not write the resident slot, even on retry")
fail("the staged copy could not write the resident slot")
# Restore the resident image through the staged copy, then 'J' back
# into it and prove it lives.

102
test/pbselfwrite.py Normal file
View File

@@ -0,0 +1,102 @@
#!/usr/bin/env python3
"""The seal, gated on the one command that proves it: erase the page the
loader is executing from.
pureboot 9 dropped the running-slot write guard, so this command is now
permitted - that is what lets a resident copy plant something in its own slot,
which on a chip whose boot section *is* the loader slot is the only route a
self-update has. Permitted means the loader must actually do it, and the only
honest proof is the flash afterwards.
What stands in the guard's place is the seal, and the two halves are tested
against each other here: the identical destructive command, refused when its
seal is wrong and honoured when it is right. A test that only showed the
refusal would pass just as well against a loader that ignores SPM entirely.
Usage: pbselfwrite.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <tool_py> <workdir>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def sealed_frame(pb, op, space, address, count):
"""A command header and its seal, built here rather than borrowed from the
tool: this test is about what the loader accepts, and a probe that shares
the host's frame builder cannot tell a wrong frame from a wrong loader."""
head = bytes((op, pb.selector(space, address), address & 0xFF,
(address >> 8) & 0xFF, count & 0xFF))
seal = pb.SEAL
for byte in head:
seal ^= byte
return head + bytes((seal,))
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, tool, workdir = sys.argv[1:]
base, page, baud = int(base_hex, 0), int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "dump.bin")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump)
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
info = loader.connect(25)
if info.version < pb.SEALED_LOADER:
fail(f"this test is for pureboot {pb.SEALED_LOADER} and later, not {info.version}")
# Erase the first page of the running slot: the entry stub and the
# command loop are both in it, so a loader that performs this does not
# answer again. Nothing else in the protocol is as sharp a probe.
frame = sealed_frame(pb, pb.OP_WRITE, pb.SP_SPM, base, pb.SPM_ERASE)
# Red: the same command with one bit wrong in its seal. Refused before
# anything happens, and the loader is still there to say so.
broken = bytearray(frame)
broken[-1] ^= 0x01
port.write(bytes(broken))
if port.read_exact(1, 5.0) != pb.NAK:
fail("an unsealed erase of the running page was not refused")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("the loader did not re-prompt after refusing the erase")
alive = loader.read_flash(base, 8)
if alive == b"\xff" * 8:
fail("the refused erase happened anyway - the running page reads erased")
# Green: the identical command, correctly sealed. Nothing is required
# of the link from here on. The verdict is *issued* before the SPM, but
# the erase takes the code that would have finished saying it, and how
# much of it survives is the chip's business - an erase removes one page
# and nothing else, so a loader whose command loop lives past the page
# erased will prompt as usual where one with 128-byte pages goes with
# the stub. None of that is the claim. The claim is that the erase
# reached flash, and the dump is both the only witness for it and a
# better one: it tells "accepted and performed" from "merely answered".
port.write(frame)
try:
port.read_exact(2, 2.0)
except pb.Error:
pass
port.close()
finally:
device.stop()
# Ground truth: the simulator's flash, not the loader's opinion of it.
flash = open(dump, "rb").read()
if flash[base : base + page] != b"\xff" * page:
fail("the sealed erase did not reach flash - the running page is intact")
print("pbselfwrite: the running slot is refused unsealed and erased sealed")
main()

View File

@@ -8,14 +8,17 @@ import subprocess
class Device:
def __init__(self, binary, elf, mcu, hz, base_hex, page, baud, dump, reset_hex=None, resume=None, link=None):
def __init__(self, binary, elf, mcu, hz, base_hex, page, baud, dump, reset_hex=None, resume=None, link=None,
window=False):
cmd = [binary]
if link:
cmd += ["-l", link]
if window:
cmd.append("-w") # report the first-transmit cycle, free-run idle
cmd += [elf, mcu, hz, base_hex, str(page), str(baud), dump]
if reset_hex is not None or resume is not None:
# Chips without a hardware boot section the tinies and the
# m48s reset to address 0 like silicon; the boot-sectioned
# Chips without a hardware boot section - the tinies and the
# m48s - reset to address 0 like silicon; the boot-sectioned
# megas re-vector to the loader base (BOOTRST).
patch = not mcu.startswith("atmega") or mcu.startswith("atmega48")
cmd.append(reset_hex if reset_hex is not None else ("0" if patch else base_hex))
@@ -41,7 +44,7 @@ class Device:
self.proc.send_signal(signal.SIGUSR1)
def power_fail(self):
"""SIGTERM: the runner dumps its flash and exits the image a
"""SIGTERM: the runner dumps its flash and exits - the image a
restart resumes from."""
self.stop()
return self.dump

View File

@@ -11,6 +11,7 @@ loader built off the chip's natural serial default.
"""
import os
import re
import sys
@@ -20,7 +21,7 @@ def fail(message):
def rjmp_decode(word, at, flash_words):
"""Where an rjmp word at word-address `at` lands deliberately written
"""Where an rjmp word at word-address `at` lands - deliberately written
against the instruction-set definition (12-bit signed offset), not with
the host tool's encoder, so an encoding bug cannot verify itself."""
if word & 0xF000 != 0xC000:
@@ -55,6 +56,12 @@ def main():
# the page byte is the wire's 0-means-256.
mega = mcu.startswith("atmega")
patch = not mega or mcu.startswith("atmega48")
# Where SRAM begins: the x8 and x4 megas push it past their extended I/O
# space, everything else starts right after the plain I/O registers. The
# loader keeps no statics and its stack sits at RAMEND, so the first SRAM
# byte is free for the data-space probe below.
classic = mcu in ("atmega8", "atmega8a", "atmega16", "atmega16a", "atmega32", "atmega32a")
ram_base = 0x0100 if mega and not classic else 0x0060
word_flash = base + pb.SLOT > 0x10000
wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0)
@@ -64,21 +71,34 @@ def main():
+ bytes([flags])
)
# A shared-line link (RX == TX in the -l spec) makes the host read every
# byte it sends back off the line; all sessions then discard the echo.
one_wire = bool(link) and re.fullmatch(r"sw:([A-H][0-7]),\1(@[01])?", link) is not None
extra = ("--one-wire",) if one_wire else ()
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, link=link)
try:
# Session 1: knock from reset, identify, program everything, stay.
out = pbsim.run_tool(tool, device.pty, baud, "--info", "--fuses", "--flash", app_bin,
out = pbsim.run_tool(tool, device.pty, baud, *extra, "--info", "--fuses", "--flash", app_bin,
"--eeprom", ee_path, "--stay")
for needed in ("version", "signature", "fuses", "verify:", "stays"):
if needed not in out:
fail(f"session 1 output lacks {needed!r}")
# Session 2: reconnect into the live session, verify, dump, hand over
# is deferred the pty must be reopened for the APP banner first.
out = pbsim.run_tool(tool, device.pty, baud, "--verify-flash", app_bin, "--verify-eeprom", ee_path,
"--read-flash", read_flash, "--read-eeprom", read_eeprom, "--stay")
# Session 2: reconnect into the live session, verify, dump, exercise
# the data space; hand over is deferred - the pty must be reopened for
# the APP banner first.
probe = "c0ffee"
out = pbsim.run_tool(tool, device.pty, baud, *extra, "--verify-flash", app_bin, "--verify-eeprom", ee_path,
"--read-flash", read_flash, "--read-eeprom", read_eeprom,
"--poke", f"{ram_base:#x}:{probe}", "--peek", f"{ram_base:#x}:3", "--stay")
if out.count("verify:") != 2:
fail("session 2 did not verify both memories")
# What went into SRAM must come back out of it: the data space is one
# more selector on the same transfer as flash and EEPROM, so a wrong
# selector decode would show up here and nowhere else.
if probe not in out.replace(" ", ""):
fail(f"data-space round trip at {ram_base:#x} did not read back {probe}\n{out}")
eeprom_back = open(read_eeprom, "rb").read()
if eeprom_back[: len(ee_image)] != ee_image:
@@ -97,11 +117,13 @@ def main():
# land in the application, which banners on the same link.
device.reset()
port = pb.Port(device.pty, baud)
if one_wire:
port = pb.OneWirePort(port)
try:
loader = pb.Loader(port)
live = loader.connect(15)
# The loader built from this tree must report a version the tool
# beside it speaks a bump the tool was never told about is a
# beside it speaks - a bump the tool was never told about is a
# loader it would refuse to talk to. Not equality with the newest:
# the tool now spans two loader generations, the fixed-baud one
# here and the unified autobaud loader that follows it.
@@ -109,15 +131,42 @@ def main():
fail(f"loader reports pureboot {live.version}, the tool speaks "
f"{pb.OLDEST_LOADER}..{pb.NEWEST_LOADER}")
# A W addressed inside a page rather than at its base must still
# consume exactly one page and prompt. The loader's own slot is
# the target — it is drained and never programmed — and the
# payload is erased-state bytes, so the probe can disturb neither
# the image nor the page buffer it leaves behind.
wire = wire_base + 1
port.write(bytes((ord("W"), wire & 0xFF, wire >> 8)) + b"\xff" * page)
# A fill addressed inside a page rather than at its base must still
# consume exactly one page and prompt. The loader's own slot is the
# target and the payload is erased-state bytes, so the probe can
# disturb neither the image nor the page buffer it leaves behind: a
# fill only loads the buffer, and nothing commits it. Hand-built
# rather than through write_page(), which would follow the fill
# with its erase and write; the point here is that the fill alone
# consumes exactly one page whatever the address's low bits say.
# Hand-sealed too - a protocol probe that borrowed the tool's own
# frame builder could not tell a wrong frame from a wrong loader.
wire = base + 1
head = bytes((pb.OP_FILL, pb.selector(pb.SP_FLASH, wire), wire & 0xFF,
(wire >> 8) & 0xFF, page & 0xFF))
seal = pb.SEAL
for byte in head:
seal ^= byte
port.write(head + bytes((seal,)))
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("unaligned W did not return to the prompt")
fail("the loader refused a correctly sealed fill")
port.write(b"\xff" * page)
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("unaligned fill did not return to the prompt")
# And the seal itself, red: one wrong bit in the address of that
# same frame must be refused outright. The verdict has to arrive
# *before* the page would have been sent - that ordering is what
# keeps a refusal from turning into a desync - so the probe sends
# no payload at all and expects the loader straight back at the
# command level.
broken = bytearray(head + bytes((seal,)))
broken[2] ^= 0x01
port.write(bytes(broken))
if port.read_exact(1, 5.0) != pb.NAK:
fail("a header with a broken seal was not refused")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("the loader did not re-prompt after refusing a broken seal")
loader.run_application()
banner = port.read_exact(3, 5.0)
@@ -137,7 +186,7 @@ def main():
# The surgery, decoded independently: the patched vector must land on the
# loader, the trampoline on the application's own entry (patched-vector
# chips only a boot-sectioned mega's word 0 stays the application's).
# chips only - a boot-sectioned mega's word 0 stays the application's).
if patch:
flash_words = (base + pb.SLOT) // 2
app = open(app_bin, "rb").read()

View File

@@ -4,8 +4,8 @@ itself with a re-timed build, and every power-fail phase is rehearsed by
killing the device mid-write, restarting it from its flash dump, and letting
a re-run complete the update.
The boot-sectioned megas run the BOOTRST-unprogrammed profile reset boots
the application, whose 'L' is the application-owned loader entry with
The boot-sectioned megas run the BOOTRST-unprogrammed profile - reset boots
the application, whose 'L' is the application-owned loader entry - with
--assume-fuses standing in for the fuse read simavr cannot model.
Usage: pbupdate.py <device_bin> <pureboot_elf> <update_elf> <mcu> <hz>
@@ -39,8 +39,8 @@ class PowerFail(Exception):
def assumed_fuses(pb, image):
"""Synthetic 'F' bytes for --assume-fuses: the smallest boot section
covering both the resident and the staging slot (two slots what a
self-update needs), BOOTRST unprogrammed the per-chip BOOTSZ ladder
covering both the resident and the staging slot (two slots - what a
self-update needs), BOOTRST unprogrammed - the per-chip BOOTSZ ladder
and fuse byte come from the tool's own table, keyed by the update
image's embedded signature."""
info = pb.image_info(image)
@@ -138,7 +138,7 @@ def main():
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, reset_hex=reset_hex)
final = "v0"
try:
# The application first its planner output is the restore truth.
# The application first - its planner output is the restore truth.
pbsim.run_tool(tool, device.pty, baud, "--flash", app_bin, "--stay")
port, loader = connect(device)
app_pages = pb.plan_flash(open(app_bin, "rb").read(), loader.info)
@@ -167,7 +167,7 @@ def main():
# direction so the flash is never already at its target. The mega's
# mid-resident-rewrite loss is exercised as a host crash instead:
# with BOOTRST unprogrammed and the resident mid-erase, a power loss
# there has no reset path into the staging copy the documented
# there has no reset path into the staging copy - the documented
# cost of that profile (README).
for kill_region, kill_hits, kill_device in (
("stage", 2, True),

137
test/pbwindow.py Normal file
View File

@@ -0,0 +1,137 @@
#!/usr/bin/env python3
"""The activation window as a behavioral duration gate.
The loader's window is a counted poll loop whose per-poll cost is hand-counted
in the source (`link::poll_cycles`) - but the loop compiles in consumer
context, so only the running image can prove the count. This test installs a
real application beside the loader (the host tool's own `plan_flash` supplies
the reset-vector surgery), starts the simulator with the line idle, and reads
the cycle of the first transmit activity: nothing talks until the window
closes and the application banners, so that cycle *is* the window, give or
take a banner lead measured in microseconds. Asserted at +/-2 % - one
mis-counted cycle per poll shifts a window by 10 % and more.
Fixed-baud loaders declare their window in seconds (--seconds, the build's
TIMEOUT). The autobaud loader's window is its calibration poll budget
(--autobaud-polls); the seconds it amounts to are budget x 9 / f_cpu, the
measured cost of the calibrate() wait loop this gate pins.
"""
import argparse
import importlib.util
import pathlib
import select
import sys
import time
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
from pbsim import Device
# The calibrate() budget loop's cycles per poll in the built image - what the
# README's window arithmetic rests on, verified here. A measured fact, not a
# design constant: the wait's exit branches land where the compiler's block
# layout puts them, and the bounded-calibration rework moved the loop from
# ten cycles to nine.
AUTOBAUD_POLL_CYCLES = 9
def load_tool(path):
spec = importlib.util.spec_from_file_location("pureboot", path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def compose_flash(pb, loader_bytes, app_bytes, mcu, base, page):
"""The flash image a completed programming session leaves: application
(with the tinies' vector surgery), loader at base - built through the
host tool's own planner so the surgery is the shipped one, not a copy."""
flash_size = base + pb.SLOT
patch = not mcu.startswith("atmega") or mcu.startswith("atmega48")
word_flash = flash_size > 0x10000
wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0)
raw = bytes((ord("P"), ord("B"), 5, 0, 0, 0, page & 0xFF,
wire_base & 0xFF, wire_base >> 8, 0, 0, flags))
info = pb.Info(raw)
flash = bytearray(b"\xff" * flash_size)
for address, content in pb.plan_flash(app_bytes, info).items():
flash[address:address + len(content)] = content
flash[base:base + len(loader_bytes)] = loader_bytes
return bytes(flash)
def first_tx_cycle(device, deadline):
"""The PB_WINDOW_TX report, or None. The runner prints it once."""
stream = device.proc.stdout
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
return None
ready, _, _ = select.select([stream], [], [], remaining)
if not ready:
return None
line = stream.readline()
if not line:
return None
if line.startswith("PB_WINDOW_TX"):
return int(line.split()[1])
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--device", required=True)
parser.add_argument("--loader", required=True)
parser.add_argument("--mcu", required=True)
parser.add_argument("--hz", type=int, required=True)
parser.add_argument("--base", required=True)
parser.add_argument("--page", type=int, required=True)
parser.add_argument("--baud", type=int, required=True)
parser.add_argument("--app", required=True)
parser.add_argument("--tool", required=True)
parser.add_argument("--workdir", required=True)
parser.add_argument("--link", default=None)
parser.add_argument("--seconds", type=float, default=None)
parser.add_argument("--autobaud-polls", type=int, default=None)
args = parser.parse_args()
if (args.seconds is None) == (args.autobaud_polls is None):
parser.error("exactly one of --seconds / --autobaud-polls")
pb = load_tool(args.tool)
base = int(args.base, 0)
expected = (args.seconds if args.seconds is not None
else args.autobaud_polls * AUTOBAUD_POLL_CYCLES / args.hz)
work = pathlib.Path(args.workdir)
work.mkdir(parents=True, exist_ok=True)
# Every loader target objcopies its slot content beside the ELF (.bin).
loader_bytes = pathlib.Path(args.loader + ".bin").read_bytes()
app_bytes = pathlib.Path(args.app).read_bytes()
flash_file = work / "window-flash.bin"
flash_file.write_bytes(compose_flash(pb, loader_bytes, app_bytes, args.mcu, base, args.page))
device = Device(args.device, args.loader, args.mcu, str(args.hz), args.base, args.page,
args.baud, str(work / "window-dump.bin"), resume=str(flash_file),
link=args.link, window=True)
try:
# Simulation speed is machine-dependent; a few hundred thousand
# cycles per wall second is the pessimistic floor.
budget = max(60.0, expected * args.hz / 300000)
cycle = first_tx_cycle(device, time.monotonic() + budget)
finally:
device.stop()
if cycle is None:
print(f" [FAIL] no transmit activity within {budget:.0f} s wall "
f"(expected a {expected:.2f} s window)")
return 1
measured = cycle / args.hz
error = (measured - expected) / expected
ok = abs(error) <= 0.02
print(f" [{'PASS' if ok else 'FAIL'}] window {measured:.3f} s vs declared "
f"{expected:.3f} s ({error:+.1%}, gate +/-2%)")
return 0 if ok else 1
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -1,461 +0,0 @@
// simavr "device" for the pureboot protocol tests, every chip. Loads the
// boot-linked ELF at the loader base, starts execution there (BOOTRST / the
// patched vector are not what is under test), and exposes the loader's
// serial link as a pty for the real host tool:
//
// - Hardware USART builds: simavr's uart_pty on the selected instance.
// - Software UART builds: an 8N1 bridge between a pty and the GPIO pins,
// timed against the simulated cycle counter (drives the loader's RX,
// decodes its TX).
//
// The link follows the chip's natural default (USART0 on the megas, the
// software UART on PB0/PB1 elsewhere) unless -l overrides it: `-l usart1`
// for the second instance, `-l sw:B5,B1` for a software build's RX,TX pins.
//
// simavr's tiny cores decode the SPM opcode but attach no NVM module — SPM
// is a silent no-op (the mega's boot section has one, avr_flash). The
// missing module is supplied here: the SPM ioctl reads SPMCSR/Z/r1:r0 and
// implements buffer fill, page erase, page write, and CTPB, completing
// instantly. RFLB's LPM diversion (fuse readout) stays unmodeled, so the
// 'F' command answers with flash bytes — the tests assert transport only.
//
// On exit (or SIGTERM) the flash and EEPROM are dumped to files for a
// ground-truth cross-check against what the host read back.
#include <fcntl.h>
#include <pty.h>
#include <signal.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <termios.h>
#include <unistd.h>
#include "avr_eeprom.h"
#include "avr_flash.h"
#include "avr_ioport.h"
#include "avr_uart.h"
#include "sim_avr.h"
#include "sim_elf.h"
#include "sim_io.h"
#include "uart_pty.h"
static avr_t *avr;
static uart_pty_t uart_pty;
static int link_software;
static char uart_digit = '0';
static char sw_rx_port = 'B', sw_tx_port = 'B';
static int sw_rx_bit = 0, sw_tx_bit = 1;
static const char *dump_path;
static uint32_t reset_pc;
static volatile sig_atomic_t reset_requested;
static int parse_link(const char *spec)
{
if (strcmp(spec, "usart0") == 0 || strcmp(spec, "usart1") == 0) {
link_software = 0;
uart_digit = spec[5];
return 0;
}
if (strncmp(spec, "sw", 2) == 0) {
link_software = 1;
if (spec[2] == '\0')
return 0;
if (sscanf(spec + 2, ":%c%d,%c%d", &sw_rx_port, &sw_rx_bit, &sw_tx_port, &sw_tx_bit) == 4)
return 0;
}
return -1;
}
// simavr 1.6's avr_flash PGERS handler erases spm_pagesize bytes starting at
// Z & ~1 instead of the page containing Z (its PGWRT path masks correctly) —
// hardware ignores the in-page bits (§26.8.1), so an erase issued with Z
// anywhere inside the page wipes half the neighbouring page in simulation
// only. Wrap the mega's registered flash ioctl and re-dispatch page erases
// with Z forced to the page boundary; everything else passes through.
//
// A second gap on the boot-section-less m48s: their RWWSRE bit is the
// temporary-buffer discard (Atmel-8271 §26.2/§26.3.1), but the stock model
// gates its RWWSRE branch on AVR_SELFPROG_HAVE_RWW — absent on the m48
// core — so the discard store falls through into the buffer-fill branch and
// plants whatever Z/R1:R0 happen to hold. Perform the silicon's discard
// here instead.
static avr_flash_t *mega_flash;
static int (*mega_flash_ioctl)(avr_io_t *io, uint32_t ctl, void *param);
static int fixed_flash_ioctl(avr_io_t *io, uint32_t ctl, void *param)
{
if (ctl == AVR_IOCTL_FLASH_SPM && avr_regbit_get(io->avr, mega_flash->pgers)) {
uint16_t z = (uint16_t)(io->avr->data[30] | (io->avr->data[31] << 8));
uint16_t masked = (uint16_t)(z & ~(mega_flash->spm_pagesize - 1));
io->avr->data[30] = (uint8_t)masked;
io->avr->data[31] = (uint8_t)(masked >> 8);
int result = mega_flash_ioctl(io, ctl, param);
io->avr->data[30] = (uint8_t)z;
io->avr->data[31] = (uint8_t)(z >> 8);
return result;
}
if (ctl == AVR_IOCTL_FLASH_SPM && !(mega_flash->flags & AVR_SELFPROG_HAVE_RWW) &&
(io->avr->data[mega_flash->r_spm] & 0x11) == 0x11) { // RWWSRE|SELFPRGEN: the m48 buffer discard
for (int i = 0; i < mega_flash->spm_pagesize / 2; i++) {
mega_flash->tmppage[i] = 0xffff;
mega_flash->tmppage_used[i] = 0;
}
avr_regbit_clear(io->avr, mega_flash->selfprgen);
return 0;
}
return mega_flash_ioctl(io, ctl, param);
}
static void fix_mega_flash_erase(void)
{
for (avr_io_t *io = avr->io_port; io; io = io->next) {
if (io->kind && strcmp(io->kind, "flash") == 0) {
mega_flash = (avr_flash_t *)io;
mega_flash_ioctl = io->ioctl;
io->ioctl = fixed_flash_ioctl;
return;
}
}
fprintf(stderr, "device: no flash module to fix — SPM page erases may misalign\n");
}
static void request_reset(int sig)
{
(void)sig;
reset_requested = 1;
}
// ------------------------------------------------------------- tiny NVM ---
typedef struct {
avr_io_t io;
uint8_t buffer[128];
uint8_t used[128]; // a buffer word loads once until erased — like silicon
unsigned page;
} tiny_nvm_t;
static tiny_nvm_t nvm;
static int nvm_ioctl(avr_io_t *io, uint32_t ctl, void *param)
{
(void)param;
if (ctl != AVR_IOCTL_FLASH_SPM)
return -1;
tiny_nvm_t *n = (tiny_nvm_t *)io;
avr_t *mcu = io->avr;
uint8_t command = mcu->data[0x57] & 0x1f; // SPMCSR, both tinies
uint16_t z = (uint16_t)(mcu->data[30] | (mcu->data[31] << 8));
uint32_t page_base = (uint32_t)(z & ~(n->page - 1)) % (mcu->flashend + 1);
if (command == 0x01) { // SPMEN alone: buffer fill from r1:r0
unsigned offset = z & (n->page - 1) & ~1u;
if (!n->used[offset]) { // first write wins until the buffer clears
n->buffer[offset] = mcu->data[0];
n->buffer[offset + 1] = mcu->data[1];
n->used[offset] = 1;
}
} else if (command == 0x03) { // PGERS
memset(mcu->flash + page_base, 0xff, n->page);
} else if (command == 0x05) { // PGWRT: programming only clears bits
for (unsigned i = 0; i < n->page; i++)
mcu->flash[page_base + i] &= n->buffer[i];
memset(n->buffer, 0xff, n->page);
memset(n->used, 0, n->page);
} else if (command == 0x11) { // CTPB
memset(n->buffer, 0xff, n->page);
memset(n->used, 0, n->page);
}
mcu->data[0x57] &= (uint8_t)~0x1f; // the operation completes instantly
return 0;
}
// ----------------------------------------------------------- GPIO bridge ---
static int pty_master = -1;
static avr_irq_t *rx_pin; // the loader's RX (PB0), driven from the pty
static avr_cycle_count_t bit_cycles;
static int tx_level = 1, tx_active, tx_bit;
static uint8_t tx_shift;
static avr_cycle_count_t tx_sample(avr_t *mcu, avr_cycle_count_t when, void *param)
{
(void)mcu;
(void)param;
tx_shift = (uint8_t)((tx_shift >> 1) | (tx_level ? 0x80 : 0));
if (++tx_bit < 8)
return when + bit_cycles;
if (write(pty_master, &tx_shift, 1) != 1)
fprintf(stderr, "device: pty write lost a byte\n");
tx_active = 0;
return 0;
}
static void tx_hook(avr_irq_t *irq, uint32_t value, void *param)
{
(void)irq;
(void)param;
int level = value & 1;
if (!tx_active && tx_level == 1 && level == 0) { // start edge
tx_active = 1;
tx_bit = 0;
avr_cycle_timer_register(avr, bit_cycles + bit_cycles / 2, tx_sample, NULL);
}
tx_level = level;
}
static uint8_t rx_queue[8192];
static unsigned rx_head, rx_tail; // ring: head = next to send
static int rx_active, rx_bit;
static uint8_t rx_byte;
static void rx_start_next(void);
static avr_cycle_count_t rx_step(avr_t *mcu, avr_cycle_count_t when, void *param)
{
(void)mcu;
(void)param;
if (rx_bit < 8) {
avr_raise_irq(rx_pin, (rx_byte >> rx_bit) & 1);
rx_bit++;
return when + bit_cycles;
}
if (rx_bit == 8) { // stop bit, plus one idle bit of margin
avr_raise_irq(rx_pin, 1);
rx_bit++;
return when + 2 * bit_cycles;
}
rx_active = 0;
rx_start_next();
return 0;
}
static void rx_start_next(void)
{
if (rx_active || rx_head == rx_tail)
return;
rx_byte = rx_queue[rx_head];
rx_head = (rx_head + 1) % sizeof(rx_queue);
rx_active = 1;
rx_bit = 0;
avr_raise_irq(rx_pin, 0); // start bit
avr_cycle_timer_register(avr, bit_cycles, rx_step, NULL);
}
// A reset abandons whatever the bridge was mid-transfer: bytes still queued
// for a chip that no longer has the context to receive them meaningfully,
// and a decode in progress on a TX line the reset may have already changed.
// The pending cycle timers must go with the state: avr_reset drops the TX
// output latch, whose falling edge starts a spurious decode before this
// runs, and a stale tx_sample would then interleave with the loader's first
// real answer through the shared shift state, corrupting it.
static void bridge_reset(void)
{
avr_cycle_timer_cancel(avr, tx_sample, NULL);
avr_cycle_timer_cancel(avr, rx_step, NULL);
rx_head = rx_tail = 0;
rx_active = 0;
tx_active = 0;
tx_level = 1;
avr_raise_irq(rx_pin, 1); // idle line
}
static void poll_pty(void)
{
uint8_t chunk[256];
ssize_t got = read(pty_master, chunk, sizeof(chunk));
for (ssize_t i = 0; i < got; i++) {
unsigned next = (rx_tail + 1) % sizeof(rx_queue);
if (next == rx_head)
break; // full: the host will retry on timeout
rx_queue[rx_tail] = chunk[i];
rx_tail = next;
}
if (got > 0)
rx_start_next();
}
// ------------------------------------------------------------------ main ---
static void finish(int sig)
{
(void)sig;
if (dump_path) {
FILE *f = fopen(dump_path, "wb");
if (f) {
fwrite(avr->flash, 1, avr->flashend + 1, f);
fclose(f);
}
avr_eeprom_desc_t ee = {.ee = NULL, .offset = 0, .size = 0};
if (avr_ioctl(avr, AVR_IOCTL_EEPROM_GET, &ee) == 0 && ee.ee && ee.size) {
char path[512];
snprintf(path, sizeof(path), "%s.eeprom", dump_path);
f = fopen(path, "wb");
if (f) {
fwrite(ee.ee, 1, ee.size, f);
fclose(f);
}
}
}
if (!link_software)
uart_pty_stop(&uart_pty);
_exit(0);
}
int main(int argc, char *argv[])
{
int link_given = 0;
for (int opt; (opt = getopt(argc, argv, "l:")) != -1;) {
if (opt != 'l' || parse_link(optarg) != 0) {
fprintf(stderr, "device: bad link spec (usart0, usart1, sw, or sw:B0,B1 as RX,TX)\n");
return 2;
}
link_given = 1;
}
int args = argc - optind;
if (args < 7 || args > 9) {
fprintf(stderr,
"usage: %s [-l link] <pureboot.elf> <mcu> <hz> <base_hex> <page> <baud> <flash_dump>"
" [reset_hex] [resume_flash]\n"
" -l link: usart0 | usart1 | sw[:B0,B1] (RX,TX); default: the chip's own\n"
" reset_hex: reset vector (default: base with a boot section, else 0)\n"
" resume_flash: raw full-flash image loaded instead of the ELF — a prior\n"
" run's dump, for power-fail resume tests\n",
argv[0]);
return 2;
}
argv += optind - 1; // argv[1] is the ELF again, whatever was parsed
const char *mcu_name = argv[2];
uint32_t base = (uint32_t)strtoul(argv[4], NULL, 0);
unsigned page = (unsigned)atoi(argv[5]);
unsigned baud = (unsigned)atoi(argv[6]);
dump_path = argv[7];
int is_mega = strncmp(mcu_name, "atmega", 6) == 0;
if (!link_given)
link_software = !is_mega; // the chips' natural links: USART0, or PB0/PB1
avr = avr_make_mcu_by_name(mcu_name);
if (!avr) {
fprintf(stderr, "device: no %s core\n", mcu_name);
return 1;
}
avr_init(avr);
avr->frequency = (uint32_t)strtoul(argv[3], NULL, 0);
memset(avr->flash, 0xff, avr->flashend + 1); // real flash powers up erased
if (args > 8) {
// Resume: the full flash image of an interrupted prior run.
FILE *f = fopen(argv[9], "rb");
if (!f || fread(avr->flash, 1, avr->flashend + 1, f) == 0) {
fprintf(stderr, "device: cannot read %s\n", argv[9]);
return 1;
}
fclose(f);
} else {
elf_firmware_t fw = {0};
if (elf_read_firmware(argv[1], &fw) != 0) {
fprintf(stderr, "device: cannot read %s\n", argv[1]);
return 1;
}
memcpy(avr->flash + base, fw.flash, fw.flashsize);
}
// The boot-sectioned megas enter the loader in hardware (BOOTRST, not
// modeled — the argument picks the modeled fuse's target); the tinies
// and the boot-section-less m48s reset to word 0 like silicon — erased
// flash walks up into the loader, and after the host's surgery the
// patched vector routes there.
int boot_section = is_mega && strncmp(mcu_name, "atmega48", 8) != 0;
reset_pc = args > 7 ? (uint32_t)strtoul(argv[8], NULL, 0) : (boot_section ? base : 0);
avr->pc = reset_pc;
avr->codeend = avr->flashend;
// Erased EEPROM, as hardware powers up (simavr zeroes it).
uint8_t blank[1024];
memset(blank, 0xff, sizeof(blank));
avr_eeprom_desc_t seed = {.ee = blank, .offset = 0, .size = 0};
if (avr_ioctl(avr, AVR_IOCTL_EEPROM_GET, &seed) == 0 && seed.size <= sizeof(blank)) {
seed.ee = blank;
avr_ioctl(avr, AVR_IOCTL_EEPROM_SET, &seed);
}
// The megas carry simavr's avr_flash module (and its two gaps the wrap
// above fixes); the tinies get the NVM module simavr lacks. Which serial
// bridge runs is the link's business, not the chip class's.
if (is_mega) {
fix_mega_flash_erase();
} else {
nvm.page = page;
memset(nvm.buffer, 0xff, sizeof(nvm.buffer));
nvm.io.kind = "tiny_nvm";
nvm.io.ioctl = nvm_ioctl;
avr_register_io(avr, &nvm.io);
}
if (!link_software) {
// POLL_SLEEP paces an idle-polling loader in host real time (a
// no-hardware CPU-saving hack); clear it so cycles run free.
uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
uart_pty_init(avr, &uart_pty);
uart_pty_connect(&uart_pty, uart_digit);
printf("PB_PTY %s\n", uart_pty.pty.slavename);
} else {
bit_cycles = (avr->frequency + baud / 2) / baud; // matches uart.hpp's own rounding exactly
rx_pin = avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_rx_port), (unsigned)sw_rx_bit);
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_tx_port), (unsigned)sw_tx_bit), tx_hook,
NULL);
avr_raise_irq(rx_pin, 1); // idle line
int slave;
struct termios raw;
cfmakeraw(&raw);
if (openpty(&pty_master, &slave, NULL, &raw, NULL) != 0) {
fprintf(stderr, "device: openpty failed\n");
return 1;
}
fcntl(pty_master, F_SETFL, O_NONBLOCK);
printf("PB_PTY %s\n", ttyname(slave));
}
fflush(stdout);
signal(SIGTERM, finish);
signal(SIGINT, finish);
signal(SIGUSR1, request_reset); // an external reset line, for the tests
long since_poll = 0;
for (;;) {
int state = avr_run(avr);
if (state == cpu_Done || state == cpu_Crashed)
break;
if (reset_requested) {
reset_requested = 0;
avr_reset(avr);
avr->pc = reset_pc;
if (!link_software) { // reset restores the pacing hack; re-clear it
uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
} else {
bridge_reset();
}
}
if (link_software && ++since_poll >= 2000) {
since_poll = 0;
poll_pty();
// An unthrottled idle simulation runs the activation window out
// from under the host's real-time knock cadence: a 1 MHz build's
// 8 s window is 8 M cycles — tens of wall milliseconds — so a
// first knock lost to an in-flight reset misses the window
// entirely. Pace the simulation only while the bridge is fully
// quiet (nothing decoding, nothing queued); transfers keep full
// speed, and a quiet window stretches toward real time.
if (!rx_active && !tx_active && rx_head == rx_tail)
usleep(200);
}
}
finish(0);
return 0;
}

703
test/pureboot_device.cpp Normal file
View File

@@ -0,0 +1,703 @@
// simavr "device" for the pureboot protocol tests, every chip. Loads the
// boot-linked ELF at the loader base, starts execution there (BOOTRST / the
// patched vector are not what is under test), and exposes the loader's
// serial link as a pty for the real host tool:
//
// - Hardware USART builds: simavr's uart_pty on the selected instance.
// - Software UART builds: an 8N1 bridge between a pty and the GPIO pins,
// timed against the simulated cycle counter (drives the loader's RX,
// decodes its TX).
//
// The link follows the chip's natural default (USART0 on the megas, the
// software UART on PB0/PB1 elsewhere) unless -l overrides it: `-l usart1`
// for the second instance, `-l sw:B5,B1` for a software build's RX,TX pins,
// and `-l sw:D0,D1@0` where those pins are a USART's own - see the pin
// ownership the bridge models below.
//
// simavr's tiny cores decode the SPM opcode but attach no NVM module - SPM
// is a silent no-op (the mega's boot section has one, avr_flash). The
// missing module is supplied here: the SPM ioctl reads SPMCSR/Z/r1:r0 and
// implements buffer fill, page erase, page write, and CTPB, completing
// instantly. RFLB's LPM diversion (fuse readout) stays unmodeled, so the
// 'F' command answers with flash bytes - the tests assert transport only.
//
// On exit (or SIGTERM) the flash and EEPROM are dumped to files for a
// ground-truth cross-check against what the host read back.
#include <array>
#include <csignal>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <print>
#include <string_view>
#include <fcntl.h>
#include <pty.h>
#include <termios.h>
#include <unistd.h>
// The parts headers (uart_pty.h) carry no C++ linkage guards of their own,
// unlike simavr's core headers - the block covers both harmlessly.
extern "C" {
#include "avr_eeprom.h"
#include "avr_flash.h"
#include "avr_ioport.h"
#include "avr_uart.h"
#include "sim_avr.h"
#include "sim_elf.h"
#include "sim_io.h"
#include "uart_pty.h"
}
namespace {
avr_t *avr;
uart_pty_t uart_pty;
bool link_software;
avr_uart_t *hw_uart; // the pty-driven USART, for the datasheet-reset fix below
char uart_digit = '0';
char sw_rx_port = 'B', sw_tx_port = 'B';
int sw_rx_bit = 0, sw_tx_bit = 1;
char sw_tx_owner = 0; // the USART whose TXD the software link sits on
const char *dump_path;
std::uint32_t reset_pc;
volatile std::sig_atomic_t reset_requested;
// -w: report the cycle of the first transmit activity, once. What the
// activation-window gate reads - with an idle line and an application
// installed, the first thing that ever talks is the application's banner,
// so this cycle *is* the loader's window plus a banner lead measured in
// microseconds. Idle pacing is skipped in this mode: there is no real-time
// host in the loop, and a paced multi-second window would take hours.
bool window_report;
bool window_tx_seen;
void window_first_tx()
{
if (!window_report || window_tx_seen) {
return;
}
window_tx_seen = true;
std::println("PB_WINDOW_TX {}", avr->cycle);
std::fflush(stdout);
}
void window_uart_hook(avr_irq_t *, std::uint32_t, void *)
{
window_first_tx();
}
// One-wire (RX == TX in the link spec): both directions on one GPIO line
// idling on the firmware's pull-up. The bridge then follows the pin's
// direction the way the real wiring does: it drives only while the
// firmware's DDR bit reads input, decodes transitions as the firmware's
// transmit only while the firmware owns the line, ignores its own raises
// coming back through the shared irq - and echoes every byte it drives back
// to the pty, which is what the host-side FTDI tie does and what the host
// tool's --one-wire mode reads back and discards.
bool link_one_wire;
bool mcu_owns_line;
bool self_drive;
int parse_link(std::string_view spec)
{
if (spec == "usart0" || spec == "usart1") {
link_software = false;
uart_digit = spec[5];
return 0;
}
if (spec.starts_with("sw")) {
link_software = true;
if (spec.size() == 2) {
return 0;
}
char owner = 0;
int fields =
std::sscanf(spec.data() + 2, ":%c%d,%c%d@%c", &sw_rx_port, &sw_rx_bit, &sw_tx_port, &sw_tx_bit, &owner);
if (fields == 4 || fields == 5) {
sw_tx_owner = owner;
link_one_wire = sw_rx_port == sw_tx_port && sw_rx_bit == sw_tx_bit;
return 0;
}
}
return -1;
}
// simavr 1.6's avr_flash PGERS handler erases spm_pagesize bytes starting at
// Z & ~1 instead of the page containing Z (its PGWRT path masks correctly) -
// hardware ignores the in-page bits (section 26.8.1), so an erase issued with Z
// anywhere inside the page wipes half the neighbouring page in simulation
// only. Wrap the mega's registered flash ioctl and re-dispatch page erases
// with Z forced to the page boundary; everything else passes through.
//
// A second gap on the boot-section-less m48s: their RWWSRE bit is the
// temporary-buffer discard (Atmel-8271 section 26.2/section 26.3.1), but the stock model
// gates its RWWSRE branch on AVR_SELFPROG_HAVE_RWW - absent on the m48
// core - so the discard store falls through into the buffer-fill branch and
// plants whatever Z/R1:R0 happen to hold. Perform the silicon's discard
// here instead.
avr_flash_t *mega_flash;
int (*mega_flash_ioctl)(avr_io_t *io, std::uint32_t ctl, void *param);
int fixed_flash_ioctl(avr_io_t *io, std::uint32_t ctl, void *param)
{
if (ctl == AVR_IOCTL_FLASH_SPM && avr_regbit_get(io->avr, mega_flash->pgers)) {
auto z = static_cast<std::uint16_t>(io->avr->data[30] | (io->avr->data[31] << 8));
auto masked = static_cast<std::uint16_t>(z & ~(mega_flash->spm_pagesize - 1));
io->avr->data[30] = static_cast<std::uint8_t>(masked);
io->avr->data[31] = static_cast<std::uint8_t>(masked >> 8);
int result = mega_flash_ioctl(io, ctl, param);
io->avr->data[30] = static_cast<std::uint8_t>(z);
io->avr->data[31] = static_cast<std::uint8_t>(z >> 8);
return result;
}
if (ctl == AVR_IOCTL_FLASH_SPM && !(mega_flash->flags & AVR_SELFPROG_HAVE_RWW) &&
(io->avr->data[mega_flash->r_spm] & 0x11) == 0x11) { // RWWSRE|SELFPRGEN: the m48 buffer discard
for (int i = 0; i < mega_flash->spm_pagesize / 2; i++) {
mega_flash->tmppage[i] = 0xffff;
mega_flash->tmppage_used[i] = 0;
}
avr_regbit_clear(io->avr, mega_flash->selfprgen);
return 0;
}
return mega_flash_ioctl(io, ctl, param);
}
void fix_mega_flash_erase()
{
for (avr_io_t *io = avr->io_port; io; io = io->next) {
if (io->kind && std::string_view{io->kind} == "flash") {
mega_flash = reinterpret_cast<avr_flash_t *>(io);
mega_flash_ioctl = io->ioctl;
io->ioctl = fixed_flash_ioctl;
return;
}
}
std::println(stderr, "device: no flash module to fix - SPM page erases may misalign");
}
void request_reset(int)
{
reset_requested = 1;
}
// ------------------------------------------------------------- tiny NVM ---
struct tiny_nvm_t {
avr_io_t io;
std::array<std::uint8_t, 128> buffer;
std::array<std::uint8_t, 128> used; // a buffer word loads once until erased - like silicon
unsigned page;
};
tiny_nvm_t nvm;
int nvm_ioctl(avr_io_t *io, std::uint32_t ctl, void *)
{
if (ctl != AVR_IOCTL_FLASH_SPM) {
return -1;
}
auto *n = reinterpret_cast<tiny_nvm_t *>(io);
avr_t *mcu = io->avr;
std::uint8_t command = mcu->data[0x57] & 0x1f; // SPMCSR, both tinies
auto z = static_cast<std::uint16_t>(mcu->data[30] | (mcu->data[31] << 8));
std::uint32_t page_base = static_cast<std::uint32_t>(z & ~(n->page - 1)) % (mcu->flashend + 1);
if (command == 0x01) { // SPMEN alone: buffer fill from r1:r0
unsigned offset = z & (n->page - 1) & ~1u;
if (!n->used[offset]) { // first write wins until the buffer clears
n->buffer[offset] = mcu->data[0];
n->buffer[offset + 1] = mcu->data[1];
n->used[offset] = 1;
}
} else if (command == 0x03) { // PGERS
std::memset(mcu->flash + page_base, 0xff, n->page);
} else if (command == 0x05) { // PGWRT: programming only clears bits
for (unsigned i = 0; i < n->page; i++) {
mcu->flash[page_base + i] &= n->buffer[i];
}
std::memset(n->buffer.data(), 0xff, n->page);
std::memset(n->used.data(), 0, n->page);
} else if (command == 0x11) { // CTPB
std::memset(n->buffer.data(), 0xff, n->page);
std::memset(n->used.data(), 0, n->page);
}
mcu->data[0x57] &= static_cast<std::uint8_t>(~0x1f); // the operation completes instantly
return 0;
}
// ----------------------------------------------------------- GPIO bridge ---
int pty_master = -1;
avr_irq_t *rx_pin; // the loader's RX (PB0), driven from the pty
avr_cycle_count_t bit_cycles;
int tx_level = 1, tx_active, tx_bit;
std::uint8_t tx_shift;
avr_cycle_count_t tx_sample(avr_t *, avr_cycle_count_t when, void *)
{
if (tx_bit < 0) {
// Half a bit into the start bit: a real receiver re-samples here and
// abandons a false start. The device's own init produces one - DDR
// drives the pin low for the instructions until the idle level is
// written - and without this check that glitch decodes as a stray
// byte (and would read as first transmit activity under -w).
if (tx_level) {
tx_active = 0;
return 0;
}
window_first_tx();
tx_bit = 0;
return when + bit_cycles;
}
if (tx_bit < 8) {
tx_shift = static_cast<std::uint8_t>((tx_shift >> 1) | (tx_level ? 0x80 : 0));
if (++tx_bit < 8) {
return when + bit_cycles;
}
// The byte is delivered at the stop bit's sampling point (9.5 bit
// times), where a hardware receiver raises its RXC - not sooner: a
// host answering before the stop bit would put its start bit on the
// wire while the device is still driving, which the device,
// transmitting, is not watching for.
return when + bit_cycles;
}
if (write(pty_master, &tx_shift, 1) != 1) {
std::println(stderr, "device: pty write lost a byte");
}
tx_active = 0;
return 0;
}
// A USART owns its TxD pin whenever its transmitter is enabled, and the port
// register cannot drive it (section 20.2 / Atmel-8271 section 19.2) - which is why a
// bit-banged link deployed on those pins is mute until it clears UCSRnB.
// simavr wires a USART entirely through IRQs and never touches the port pin
// model, so the ownership does not exist there and the mute cannot happen:
// supply it, or the very state this models is untestable. The link spec's
// trailing @n names the USART; without one the pins are nobody's.
avr_uart_t *tx_owner;
bool tx_pin_taken()
{
if (!tx_owner) {
return false;
}
if (avr_regbit_get(avr, tx_owner->txen)) {
return true;
}
// One-wire on the USART's RXD: RXEN forces the shared pin's direction to
// input (section 20.7.3), so the firmware's drive goes nowhere until the
// release - the receive-side twin of the TXD hold.
return link_one_wire && avr_regbit_get(avr, tx_owner->rxen);
}
// simavr leaves TXEN set in UCSRnB out of reset, where silicon clears the
// whole register (section 20.11.3) - which would hand the pin to a USART no code has
// enabled, making a freshly reset chip mute for reasons hardware does not
// have. Reset it the way the datasheet does, so the ownership starts from
// nobody's and only an application that really enables the USART takes it.
void reset_tx_owner()
{
if (tx_owner) {
avr_regbit_clear(avr, tx_owner->txen);
}
}
void find_tx_owner()
{
for (avr_io_t *io = avr->io_port; io; io = io->next) {
if (io->kind && std::string_view{io->kind} == "uart" &&
reinterpret_cast<avr_uart_t *>(io)->name == sw_tx_owner) {
tx_owner = reinterpret_cast<avr_uart_t *>(io);
reset_tx_owner();
return;
}
}
std::println(stderr, "device: no USART{} to own the software link's TX pin", sw_tx_owner);
}
void tx_hook(avr_irq_t *, std::uint32_t value, void *)
{
if (link_one_wire && (self_drive || !mcu_owns_line)) {
// The bridge's own drive coming back through the shared irq, or a
// transition while the line is the bridge's - either way not the
// firmware talking: the decoder sees an idle line.
tx_level = 1;
return;
}
if (tx_pin_taken()) { // the USART holds the line; the port write goes nowhere
tx_level = 1;
return;
}
int level = value & 1;
if (!tx_active && tx_level == 1 && level == 0) { // start edge, confirmed mid-bit
tx_active = 1;
tx_bit = -1;
avr_cycle_timer_register(avr, bit_cycles / 2, tx_sample, nullptr);
}
tx_level = level;
}
std::array<std::uint8_t, 8192> rx_queue;
unsigned rx_head, rx_tail; // ring: head = next to send
int rx_active, rx_bit;
std::uint8_t rx_byte;
void rx_start_next();
// Every level the bridge itself puts on the line goes through here, so the
// shared-pin decoder can tell its own drive from the firmware's.
void bridge_drive(int level)
{
self_drive = true;
avr_raise_irq(rx_pin, static_cast<std::uint32_t>(level));
self_drive = false;
}
avr_cycle_count_t rx_step(avr_t *, avr_cycle_count_t when, void *)
{
if (rx_bit < 8) {
bridge_drive((rx_byte >> rx_bit) & 1);
rx_bit++;
return when + bit_cycles;
}
if (rx_bit == 8) { // stop bit, plus one idle bit of margin
bridge_drive(1);
// The host-side tie: an FTDI adapter on a one-wire line reads every
// byte it transmits - supply that echo, which the host tool's
// --one-wire mode consumes as its wiring check.
if (link_one_wire && write(pty_master, &rx_byte, 1) != 1) {
std::println(stderr, "device: pty echo lost a byte");
}
rx_bit++;
return when + 2 * bit_cycles;
}
rx_active = 0;
rx_start_next();
return 0;
}
void rx_start_next()
{
if (rx_active || rx_head == rx_tail) {
return;
}
// The firmware is answering on the shared line: hold the byte - a real
// host's transmission waits out the reply on the wire too. The next
// poll_pty tick retries once the line is handed back.
if (link_one_wire && mcu_owns_line) {
return;
}
rx_byte = rx_queue[rx_head];
rx_head = (rx_head + 1) % rx_queue.size();
rx_active = 1;
rx_bit = 0;
bridge_drive(0); // start bit
avr_cycle_timer_register(avr, bit_cycles, rx_step, nullptr);
}
// The shared pin's direction is the line's ownership: DDR-out is the
// firmware driving a frame, DDR-in hands the line back to the bridge.
void on_ddr(avr_irq_t *, std::uint32_t value, void *)
{
const bool owns = (value >> sw_rx_bit) & 1;
if (mcu_owns_line && !owns) {
bridge_drive(1); // hand-back: a turn-based host idles here, and the cache stays truthful
}
mcu_owns_line = owns;
// A byte held back while the firmware answered starts from the next
// poll_pty tick, never from inside the DDR write itself - the port
// model's own pull-up re-derivation runs right after this notify and
// would erase a start edge raised here.
}
// A reset abandons whatever the bridge was mid-transfer: bytes still queued
// for a chip that no longer has the context to receive them meaningfully,
// and a decode in progress on a TX line the reset may have already changed.
// The pending cycle timers must go with the state: avr_reset drops the TX
// output latch, whose falling edge starts a spurious decode before this
// runs, and a stale tx_sample would then interleave with the loader's first
// real answer through the shared shift state, corrupting it.
void bridge_reset()
{
avr_cycle_timer_cancel(avr, tx_sample, nullptr);
avr_cycle_timer_cancel(avr, rx_step, nullptr);
rx_head = rx_tail = 0;
rx_active = 0;
tx_active = 0;
tx_level = 1;
mcu_owns_line = false; // avr_reset zeroed DDR: every pin reads input again
// Re-drive the idle line through a forced transition: ioport pin irqs are
// IRQ_FLAG_FILTERED, and avr_reset zeroes the port latch while the irq
// keeps its pre-reset cached value - so a plain raise(1) against a cached
// 1 is dropped and the device reads the line stuck low. A loader entering
// calibration on that line measures reset-to-first-edge as one giant
// pulse and mis-locks or boots the application on the first real knock.
// No cycles run between the two raises, so the device only ever sees the
// final idle-high.
bridge_drive(0);
bridge_drive(1);
}
void poll_pty()
{
std::array<std::uint8_t, 256> chunk;
ssize_t got = read(pty_master, chunk.data(), chunk.size());
for (ssize_t i = 0; i < got; i++) {
unsigned next = (rx_tail + 1) % rx_queue.size();
if (next == rx_head) {
break; // full: the host will retry on timeout
}
rx_queue[rx_tail] = chunk[i];
rx_tail = next;
}
// Unconditional: a byte held back while the firmware owned a shared
// line restarts from here once the hand-back has happened.
rx_start_next();
}
// ------------------------------------------------------------------ main ---
[[noreturn]] void finish(int)
{
if (dump_path) {
std::FILE *f = std::fopen(dump_path, "wb");
if (f) {
std::fwrite(avr->flash, 1, avr->flashend + 1, f);
std::fclose(f);
}
avr_eeprom_desc_t ee = {
.ee = nullptr,
.offset = 0,
.size = 0,
};
if (avr_ioctl(avr, AVR_IOCTL_EEPROM_GET, &ee) == 0 && ee.ee && ee.size) {
std::array<char, 512> path;
std::snprintf(path.data(), path.size(), "%s.eeprom", dump_path);
f = std::fopen(path.data(), "wb");
if (f) {
std::fwrite(ee.ee, 1, ee.size, f);
std::fclose(f);
}
}
}
if (!link_software) {
uart_pty_stop(&uart_pty);
}
_exit(0);
}
} // namespace
int main(int argc, char *argv[])
{
bool link_given = false;
for (int opt; (opt = getopt(argc, argv, "l:w")) != -1;) {
if (opt == 'w') {
window_report = true;
continue;
}
if (opt != 'l' || parse_link(optarg) != 0) {
std::println(stderr, "device: bad link spec (usart0, usart1, sw, or sw:B0,B1 as RX,TX)");
return 2;
}
link_given = true;
}
int args = argc - optind;
if (args < 7 || args > 9) {
std::print(stderr,
"usage: {} [-l link] [-w] <pureboot.elf> <mcu> <hz> <base_hex> <page> <baud> <flash_dump>"
" [reset_hex] [resume_flash]\n"
" -l link: usart0 | usart1 | sw[:B0,B1[@0]] (RX,TX, then the USART owning\n"
" them); default: the chip's own\n"
" -w: print PB_WINDOW_TX <cycle> at the first transmit activity and\n"
" free-run idle time (window measurement mode)\n"
" reset_hex: reset vector (default: base with a boot section, else 0)\n"
" resume_flash: raw full-flash image loaded instead of the ELF - a prior\n"
" run's dump, for power-fail resume tests\n",
argv[0]);
return 2;
}
argv += optind - 1; // argv[1] is the ELF again, whatever was parsed
const std::string_view mcu_name = argv[2];
auto base = static_cast<std::uint32_t>(std::strtoul(argv[4], nullptr, 0));
auto page = static_cast<unsigned>(std::atoi(argv[5]));
auto baud = static_cast<unsigned>(std::atoi(argv[6]));
dump_path = argv[7];
const bool is_mega = mcu_name.starts_with("atmega");
if (!link_given) {
link_software = !is_mega; // the chips' natural links: USART0, or PB0/PB1
}
avr = avr_make_mcu_by_name(mcu_name.data());
if (!avr) {
std::println(stderr, "device: no {} core", mcu_name);
return 1;
}
avr_init(avr);
avr->frequency = static_cast<std::uint32_t>(std::strtoul(argv[3], nullptr, 0));
std::memset(avr->flash, 0xff, avr->flashend + 1); // real flash powers up erased
if (args > 8) {
// Resume: the full flash image of an interrupted prior run.
std::FILE *f = std::fopen(argv[9], "rb");
if (!f || std::fread(avr->flash, 1, avr->flashend + 1, f) == 0) {
std::println(stderr, "device: cannot read {}", argv[9]);
return 1;
}
std::fclose(f);
} else {
elf_firmware_t fw{};
if (elf_read_firmware(argv[1], &fw) != 0) {
std::println(stderr, "device: cannot read {}", argv[1]);
return 1;
}
// An image past flash end would smash the simulator's heap and turn
// into phantom peripheral behavior (lessons: believe the size gate
// first) - refuse it loudly instead.
if (base + fw.flashsize > avr->flashend + 1) {
std::println(stderr, "device: {} B at {:#x} runs past flash end {:#x} - image does not fit its slot",
fw.flashsize, base, avr->flashend);
return 1;
}
std::memcpy(avr->flash + base, fw.flash, fw.flashsize);
}
// The boot-sectioned megas enter the loader in hardware (BOOTRST, not
// modeled - the argument picks the modeled fuse's target); the tinies
// and the boot-section-less m48s reset to word 0 like silicon - erased
// flash walks up into the loader, and after the host's surgery the
// patched vector routes there.
const bool boot_section = is_mega && !mcu_name.starts_with("atmega48");
reset_pc = args > 7 ? static_cast<std::uint32_t>(std::strtoul(argv[8], nullptr, 0)) : (boot_section ? base : 0);
avr->pc = reset_pc;
avr->codeend = avr->flashend;
// Erased EEPROM, as hardware powers up (simavr zeroes it).
std::array<std::uint8_t, 1024> blank;
std::memset(blank.data(), 0xff, blank.size());
avr_eeprom_desc_t seed = {
.ee = blank.data(),
.offset = 0,
.size = 0,
};
if (avr_ioctl(avr, AVR_IOCTL_EEPROM_GET, &seed) == 0 && seed.size <= blank.size()) {
seed.ee = blank.data();
avr_ioctl(avr, AVR_IOCTL_EEPROM_SET, &seed);
}
// The megas carry simavr's avr_flash module (and its two gaps the wrap
// above fixes); the tinies get the NVM module simavr lacks. Which serial
// bridge runs is the link's business, not the chip class's.
if (is_mega) {
fix_mega_flash_erase();
} else {
nvm.page = page;
std::memset(nvm.buffer.data(), 0xff, nvm.buffer.size());
nvm.io.kind = "tiny_nvm";
nvm.io.ioctl = nvm_ioctl;
avr_register_io(avr, &nvm.io);
}
if (!link_software) {
// POLL_SLEEP paces an idle-polling loader in host real time (a
// no-hardware CPU-saving hack); clear it so cycles run free.
std::uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
// simavr leaves TXEN set out of reset where silicon clears the whole
// UCSR#B (section 20.11.3). Harmless to a loader that enables TXEN itself -
// but a half-duplex build's receiver-only init then *drops* TXEN,
// and this uart model clears UDRE on that edge and never re-raises
// it on a later enable: the first transmitter after the hand-over
// waits UDRE forever, a wedge silicon does not have. Start from the
// datasheet's zero, as the software bridge's tx-owner model does.
for (avr_io_t *io = avr->io_port; io; io = io->next) {
if (io->kind && std::string_view{io->kind} == "uart" &&
reinterpret_cast<avr_uart_t *>(io)->name == uart_digit) {
hw_uart = reinterpret_cast<avr_uart_t *>(io);
}
}
if (hw_uart) {
avr_regbit_clear(avr, hw_uart->txen);
}
uart_pty_init(avr, &uart_pty);
uart_pty_connect(&uart_pty, uart_digit);
if (window_report) {
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_UART_GETIRQ(uart_digit), UART_IRQ_OUTPUT),
window_uart_hook, nullptr);
}
std::println("PB_PTY {}", uart_pty.pty.slavename);
} else {
bit_cycles = (avr->frequency + baud / 2) / baud; // matches uart.hpp's own rounding exactly
if (sw_tx_owner) {
find_tx_owner();
}
rx_pin = avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_rx_port), static_cast<unsigned>(sw_rx_bit));
avr_irq_register_notify(
avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_tx_port), static_cast<unsigned>(sw_tx_bit)), tx_hook,
nullptr);
if (link_one_wire) {
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_rx_port), IOPORT_IRQ_DIRECTION_ALL),
on_ddr, nullptr);
}
bridge_drive(1); // idle line
int slave;
struct termios raw;
cfmakeraw(&raw);
if (openpty(&pty_master, &slave, nullptr, &raw, nullptr) != 0) {
std::println(stderr, "device: openpty failed");
return 1;
}
fcntl(pty_master, F_SETFL, O_NONBLOCK);
std::println("PB_PTY {}", ttyname(slave));
}
std::fflush(stdout);
std::signal(SIGTERM, finish);
std::signal(SIGINT, finish);
std::signal(SIGUSR1, request_reset); // an external reset line, for the tests
long since_poll = 0;
for (;;) {
int state = avr_run(avr);
if (state == cpu_Done || state == cpu_Crashed) {
break;
}
if (reset_requested) {
reset_requested = 0;
avr_reset(avr);
avr->pc = reset_pc;
if (!link_software) { // reset restores the pacing hack; re-clear it
std::uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
if (hw_uart) { // and simavr's bogus reset TXEN (section 20.11.3: zero)
avr_regbit_clear(avr, hw_uart->txen);
}
} else {
bridge_reset();
reset_tx_owner();
}
}
if (link_software && ++since_poll >= 2000) {
since_poll = 0;
poll_pty();
// An unthrottled idle simulation runs the activation window out
// from under the host's real-time knock cadence: a 1 MHz build's
// 8 s window is 8 M cycles - tens of wall milliseconds - so a
// first knock lost to an in-flight reset misses the window
// entirely. Pace the simulation only while the bridge is fully
// quiet (nothing decoding, nothing queued); transfers keep full
// speed, and a quiet window stretches toward real time.
if (!window_report && !rx_active && !tx_active && rx_head == rx_tail) {
usleep(200);
}
}
}
finish(0);
}

212
test/test_handshake.py Normal file
View File

@@ -0,0 +1,212 @@
#!/usr/bin/env python3
"""Host-tool activation handshake: bounded against a line that misbehaves.
`_handshake` drains the line after it sees a prompt, to absorb a real loader's
trailing bytes before it asks for the identity. That drain must be bounded: a
target that never falls quiet - a board stuck in a reset loop presents exactly
this, ~60 reboots/s of UART-reset garbage in which a stray 0x2b reads as a
prompt - otherwise spins the tool forever. Regression for that hang, plus a
control that a well-behaved loader still connects.
The handshake must also survive its own leftovers: after `--stay` the loader's
final prompt can still be in the USB pipeline when the next invocation opens
the port, and on a board wired to reset on open, that opening starts a fresh
activation window the stale prompt then betrays - the tool commits to an
identity read against a device that never heard its knock, and what it finally
collects is the application's banner. StaleDTRPort is that moment as a port.
Stdlib only, no device: host-tool logic, so it runs on every chip's preset
beside pureboot.planner.
"""
import importlib.util
import pathlib
import threading
import time
PB = pathlib.Path(__file__).resolve().parents[1] / "pureboot" / "pureboot.py"
_spec = importlib.util.spec_from_file_location("pureboot", PB)
pb = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(pb)
P = F = 0
def check(name, ok):
global P, F
P, F = P + (1 if ok else 0), F + (0 if ok else 1)
print(f" [{'PASS' if ok else 'FAIL'}] {name}")
class FloodPort:
"""A line that never falls quiet: read_available always returns bytes, and
they contain a prompt. No identity ever completes."""
def flush_input(self):
pass
def write(self, data):
pass
def read_available(self, wait):
time.sleep(0.01) # a real read waits; keep the busy loop off a core
return b"+\x00\xff"
def read_exact(self, count, timeout):
raise pb.Error("no identity")
class LoaderPort:
"""A well-behaved pureboot 5: a prompt to the knock, then quiet, then the
slim identity (version 5 + m328p signature) and a closing prompt."""
def __init__(self):
self.pending = b""
self.exacts = 0
def flush_input(self):
self.pending = b""
def write(self, data):
if b"p" in data:
self.pending = b"+" # the prompt answers the knock, nothing else
def read_available(self, wait):
data, self.pending = self.pending, b""
return data
def read_exact(self, count, timeout):
self.exacts += 1
return b"\x05\x1e\x95\x0f" if self.exacts == 1 else b"+" # identity, then prompt
class StaleDTRPort:
"""`--stay`, then a fresh invocation on a board that resets when its port
opens. Three facts of that moment, all timed from the open: the previous
session's final prompt is still in transit and lands only after the
opening flush has already run; the reset holds the device off the line
at first, eating anything written before it completes; and the fresh
window is finite - once it expires the application boots and prints a
banner whose bytes are what a pending identity read collects. A
handshake that trusts the stale prompt spends the whole window waiting
on a device that never heard its knock; one that drains the line first
knocks into the real window and connects."""
STALE_AT = 0.02 # the leftover prompt becomes visible (post-flush)
READY_AT = 0.05 # reset complete, activation window opens
WINDOW = 1.0 # window length; expiry boots the application
def __init__(self):
self.t0 = time.monotonic()
# (visible-from, bytes): the line as a timed queue.
self.queue = [(self.t0 + self.STALE_AT, b"+")]
self.armed = False # a 'p' heard inside the window arms 'b'
self.booted = False
def _boot_check(self):
if not self.booted and time.monotonic() > self.t0 + self.READY_AT + self.WINDOW:
self.booted = True
self.queue.append((self.t0 + self.READY_AT + self.WINDOW,
b"W r libavr tempmon\r\n"))
def _visible(self):
self._boot_check()
now = time.monotonic()
return b"".join(d for t, d in self.queue if t <= now)
def _consume(self, n):
now = time.monotonic()
left = []
for t, d in self.queue:
if t <= now and n:
take = min(n, len(d))
d = d[take:]
n -= take
if d:
left.append((t, d))
self.queue = left
def flush_input(self):
self._consume(len(self._visible()))
def write(self, data):
self._boot_check()
now = time.monotonic()
if now < self.t0 + self.READY_AT or self.booted:
return # still in reset, or the application owns the line
if b"p" in data:
self.armed = True
self.queue.append((now + 0.01, b"+"))
if b"b" in data and self.armed:
# The slim identity (version 5 + m328p signature) and a prompt.
self.queue.append((now + 0.01, b"\x05\x1e\x95\x0f+"))
def read_available(self, wait):
deadline = time.monotonic() + wait
while True:
data = self._visible()
if data:
self._consume(len(data))
return data
if time.monotonic() >= deadline:
return b""
time.sleep(0.005)
def read_exact(self, count, timeout):
deadline = time.monotonic() + timeout
data = b""
while len(data) < count:
visible = self._visible()
if visible:
take = visible[:count - len(data)]
self._consume(len(take))
data += take
elif time.monotonic() >= deadline:
raise pb.Error(f"timeout: got {len(data)} of {count} bytes")
else:
time.sleep(0.005)
return data
def terminates(port, wait, budget):
"""Run connect_autobaud in a thread; True if it returns/raises within
`budget` seconds rather than hanging."""
done = threading.Event()
def run():
try:
pb.Loader(port).connect_autobaud(wait)
except Exception:
pass
finally:
done.set()
threading.Thread(target=run, daemon=True).start()
return done.wait(budget)
def main():
# the hang: a flooding target must not spin the drain forever. With wait=0.5
# the whole handshake has to give up well inside a few seconds.
check("flooding target: handshake terminates, drain is bounded",
terminates(FloodPort(), wait=0.5, budget=4.0))
# the control: a real loader still connects and reads identity.
info = pb.Loader(LoaderPort()).connect_autobaud(2.0)
check("well-behaved loader still connects (version 5)", info.version == 5)
# the stale prompt: a --stay leftover plus reset-on-open must not burn the
# fresh window - the pre-knock drain absorbs it and the first real knock
# lands inside the window.
try:
stale_ok = pb.Loader(StaleDTRPort()).connect(2.5).version == 5
except pb.Error as failed:
print(f" ({failed})")
stale_ok = False
check("stale --stay prompt + reset-on-open: connects in the fresh window", stale_ok)
print(f"\n {P} passed, {F} failed")
return 1 if F else 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Host-tool unit tests the planning and policy logic, no simulator:
"""Host-tool unit tests - the planning and policy logic, no simulator:
programming orders and their recovery properties, the reset-vector surgery,
the staging composition, the boot-fuse decode, and the update preflight over
fuse combinations simavr cannot model.
@@ -30,8 +30,14 @@ def info_of(pb, base, page, patch, flash, signature=(0x1E, 0x93, 0x0B), word_fla
scale = 2 if word_flash else 1
wire_base = base // scale
flags = (1 if patch else 0) | (2 if word_flash else 0)
# The EEPROM size comes from the signature, as it must: pureboot 5 derives
# the whole geometry from the signature rather than sending it, so a
# synthetic block that disagreed with its own signature would describe a
# chip that cannot exist.
eeprom = pb.CHIP_GEOMETRY[signature][2]
raw = bytes((0x50, 0x42, pb.NEWEST_LOADER if version is None else version,
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8, 0, 2, flags))
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8,
eeprom & 0xFF, eeprom >> 8, flags))
info = pb.Info(raw)
if info.flash_size != flash:
fail(f"info_of({base:#x}) decodes to {info.flash_size:#x} of flash, not {flash:#x}")
@@ -70,7 +76,7 @@ def main():
"newer tool",
)
# mega_boot: BOOTSZ words and the BOOTRST sense per chip the fuse byte
# mega_boot: BOOTSZ words and the BOOTRST sense per chip - the fuse byte
# index (EXTENDED on the x8 line except the m328s' HIGH, HIGH elsewhere)
# and the per-family ladders (Atmel-2486/2466/2503/2545/8271/DS40002065/
# 8272/8011/2593/42719). Synthetic 'F' replies: only the boot byte
@@ -108,15 +114,15 @@ def main():
fail(f"mega_boot {signature[1]:02x}{signature[2]:02b} unprogrammed: {prog} {at:#07x}")
# Word-addressed info decode: the 1284P's base and page ride the wire
# scaled a 17-bit base halved into the block's two bytes, a 256-byte page
# spelled 0 and its slot is the same 512 bytes as everywhere else, so its
# scaled - a 17-bit base halved into the block's two bytes, a 256-byte page
# spelled 0 - and its slot is the same 512 bytes as everywhere else, so its
# staging slot lands inside the 1 KiB minimum boot section.
big = info_of(pb, 0x1FE00, 0, False, 0x20000, signature=(0x1E, 0x97, 0x05), word_flash=True)
if big.page != 256 or big.base != 0x1FE00 or big.stage != 0x1FC00:
fail(f"word-addressed info decode: page {big.page}, base {big.base:#x}, stage {big.stage:#x}")
# Surgery: word 0 lands on the loader, the trampoline on the original
# entry checked with an independent decoder.
# entry - checked with an independent decoder.
app = bytes((0xC0 | 0x00, 0xC0)) + bytes((0x12,)) * 300 # rjmp .+0x00C0... entry word 0xC0C0
entry = rjmp_decode(app[0] | (app[1] << 8), 0, tiny.flash_size // 2)
pages = pb.plan_flash(app, tiny)
@@ -152,7 +158,7 @@ def main():
# Staging content: the identical image plus the through-word on a
# patched-vector chip; hard size clamps either way.
image = bytes(range(256)) * 2 # 512 B too big for a tiny slot
image = bytes(range(256)) * 2 # 512 B - too big for a tiny slot
expect_error("tiny staging size", lambda: pb.staging_content(image, tiny), "510")
staged = pb.staging_content(image[:508], tiny)
through = staged[510] | (staged[511] << 8)
@@ -162,11 +168,16 @@ def main():
fail("mega staging content should be the bare image")
expect_error("mega staging size", lambda: pb.staging_content(image + b"!", mega), "512")
# The embedded info block: found in a synthetic binary, absent in noise.
binary = bytes((0xAA,)) * 10 + tiny.raw + bytes((0xBB,)) * 10
# The image stamp: found in a synthetic binary, absent in noise. pureboot
# 5 stamps the magic, its version and the signature, and the geometry is
# looked up from there - so what comes back must equal what a live device
# of the same chip reports.
stamp = bytes((0x50, 0x42, pb.NEWEST_LOADER)) + bytes(tiny.signature)
binary = bytes((0xAA,)) * 10 + stamp + bytes((0xBB,)) * 10
found = pb.image_info(binary)
if found is None or found.raw != tiny.raw:
fail("image_info misses the embedded block")
fail(f"image_info misreads the v{pb.NEWEST_LOADER} stamp: "
f"{found.raw.hex() if found else None} != {tiny.raw.hex()}")
if pb.image_info(bytes((0xAA,)) * 40) is not None:
fail("image_info invents a block")
# An older loader's image stays readable, so a deployed build can be
@@ -217,7 +228,7 @@ def main():
# The 1284s' smallest boot section (512 words) is exactly the resident
# slot plus its staging slot, so self-update is possible at the minimum
# BOOTSZ no fuse step up, the 644's geometry. That holds only while a
# BOOTSZ - no fuse step up, the 644's geometry. That holds only while a
# slot is 512 B: at 1 KiB the staging slot would fall outside the section
# and the preflight would refuse.
notes = pb.update_preflight(bytes((0xAA,)) * 8 + big.raw, big, fuses(0xFE))
@@ -280,7 +291,7 @@ def main():
if device.writes != pb.RETRIES + 1:
fail(f"unrepairable page took {device.writes} writes, expected {pb.RETRIES + 1}")
# The knock handshake against a device that is not listening yet the
# The knock handshake against a device that is not listening yet - the
# state a port open leaves behind: it resets the chip into a fresh
# activation window while the previous session's prompt is still in
# flight, so the first knock is lost and a prompt arrives anyway.

71
test/test_scan.py Normal file
View File

@@ -0,0 +1,71 @@
#!/usr/bin/env python3
"""--scan's walk and report logic, no simulator: the probe order, the rate
arithmetic, and the advice's direction. The rate physics itself is not
sim-testable - a pty carries bytes at any termios rate - so what the wire
would arbitrate is pinned here as logic instead.
Usage: test_scan.py <tool_py>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
sys.path.insert(0, os.path.dirname(os.path.abspath(sys.argv[1])))
import pureboot as pb
walk = pb.scan_ratios()
if walk != [0, -2, 2, -4, 4, -6, 6, -8, 8, -10, 10]:
fail(f"probe walk is not built-rate-first, nearest-out: {walk}")
if pb.scan_rate(9600, 4) != 9984 or pb.scan_rate(9600, -4) != 9216:
fail("probe rate arithmetic")
if pb.scan_rate(115200, 0) != 115200:
fail("the built rate must probe unchanged")
# A loader answering fast means a fast oscillator: the trim goes down.
report = "\n".join(pb.scan_report(9600, 4, 6))
for needle in ("9984", "+4 %", "--baud 9984", "4 steps lower", "pureboot 6"):
if needle not in report:
fail(f"+4 % report lacks {needle!r}:\n{report}")
report = "\n".join(pb.scan_report(9600, -6, 6))
if "6 steps higher" not in report:
fail(f"-6 % report advises the wrong direction:\n{report}")
report = "\n".join(pb.scan_report(9600, 0, 6))
if "none" not in report or "steps" in report:
fail(f"an on-rate answer must advise no trim:\n{report}")
report = "\n".join(pb.scan_report(9600, 4, 6, clock=9600000))
if "9984000" not in report:
fail(f"the absolute clock must scale with the found ratio:\n{report}")
# The walk's rates mostly have no termios B-constant, so the POSIX port
# must set them through termios2 - probed on a pty, which accepts the
# ioctl without caring about the speed. Without this every off-nominal
# probe would abort the walk on the platform --scan matters most on.
if os.name == "posix":
import pty
master, slave = pty.openpty()
try:
port = pb.Port(os.ttyname(slave), pb.scan_rate(9600, 4))
port.set_baud(pb.scan_rate(9600, -4))
port.close()
except pb.Error as error:
fail(f"PosixPort refused an off-nominal probe rate: {error}")
finally:
os.close(master)
os.close(slave)
print("OK")
if __name__ == "__main__":
main()

167
test/test_update_link.py Executable file
View File

@@ -0,0 +1,167 @@
#!/usr/bin/env python3
"""Self-update across a link change: the host must follow the staging copy.
`--update-loader` installs the new image in the staging slot and then *enters
it* to have it rewrite the resident. That copy is the new image, so it speaks the
new image's baud and backend - but the host was talking to the *resident*. Where
the two differ, the host kept knocking at the old rate in the old mode, the
staging copy never answered, and the update stranded: staging installed, resident
untouched, and on a 1 KiB tiny the application region (which *is* the staging
slot there) already gone.
The wire cannot be probed for this - 512 bytes of position-independent code carry
no header saying what rate they were built for - so the operator declares it, and
a mismatch with nothing declared has to say so instead of reporting a bare
timeout.
Stdlib only, no device: host-tool logic, so it runs on every chip's preset beside
pureboot.planner.
"""
import importlib.util
import pathlib
PB = pathlib.Path(__file__).resolve().parents[1] / "pureboot" / "pureboot.py"
_spec = importlib.util.spec_from_file_location("pureboot", PB)
pb = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(pb)
IDENTITY = b"\x05\x1e\x95\x0f" # pureboot 5 + m328p signature
P = F = 0
def check(name, ok, detail=""):
global P, F
P, F = P + (1 if ok else 0), F + (0 if ok else 1)
print(f" [{'PASS' if ok else 'FAIL'}] {name}" + (f" - {detail}" if detail else ""))
class TwoLinkPort:
"""A board whose resident and staging copy answer on different links.
Only the rate currently set decides who can be heard, which is the physical
truth: a loader's bit timing is a cycle count, so a copy built for another
rate is unreadable until the host retunes. The knock bytes carry the mode, so
a backend mismatch is caught the same way.
"""
def __init__(self, resident=(57600, False), staged=(38400, False)):
self.resident, self.staged = resident, staged
self.baud = resident[0]
self.entered = False # a 'J' has handed control to the staging copy
self.switches = [] # every retune the host asked for
self.pending = bytearray() # what the device has queued to send
# --- the part under test needs this to exist at all
def set_baud(self, baud):
self.baud = baud
self.switches.append(baud)
def flush_input(self):
self.pending.clear()
def _audible(self, knock=None):
baud, autobaud = self.staged if self.entered else self.resident
if self.baud != baud:
return False
if knock is None:
return True
return knock == (bytes((pb.CALIBRATE, ord("p"))) if autobaud else b"pb")
def write(self, data):
data = bytes(data)
if data[:1] == b"J" and len(data) == 3:
# The resident acks the jump, then control moves to the copy.
if self._audible():
self.pending += pb.PROMPT
self.entered = True
elif data in (b"pb", bytes((pb.CALIBRATE, ord("p")))):
if self._audible(data):
self.pending += pb.PROMPT
elif data == b"b":
if self._audible():
self.pending += IDENTITY + pb.PROMPT
def read_available(self, wait):
out, self.pending = bytes(self.pending), bytearray()
return out
def read_exact(self, count, timeout):
if len(self.pending) < count:
raise pb.Error(f"timeout: got {len(self.pending)} of {count} bytes")
out, self.pending = bytes(self.pending[:count]), self.pending[count:]
return out
def connected(port):
"""A Loader already in session with the resident."""
loader = pb.Loader(port)
loader.connect(2.0)
return loader
def main():
# The control first: where the staged image keeps the resident's link, the
# flow works and needs no retune. This is the case that always passed, and
# it is what made the bug look like "self-update is broken" rather than
# "self-update cannot change the link".
port = TwoLinkPort(resident=(57600, False), staged=(57600, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0)
check("same link: staging copy entered", True)
except pb.Error as error:
check("same link: staging copy entered", False, str(error))
# A baud change, declared. The host must retune before knocking.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(38400, False))
check("baud change declared: entered after retuning", 38400 in port.switches,
f"switches={port.switches}")
except (pb.Error, TypeError) as error:
check("baud change declared: entered after retuning", False, repr(error))
# A backend change, declared: the knock itself has to become the calibration
# pulse, or an autobaud staging copy never hears a thing.
port = TwoLinkPort(resident=(57600, False), staged=(57600, True))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(57600, True))
check("backend change declared: entered as autobaud", True)
except (pb.Error, TypeError) as error:
check("backend change declared: entered as autobaud", False, repr(error))
# Nothing declared against a changed link: it still cannot work, but the
# error has to name the cause. A bare "no answer" sent the operator looking
# at the wiring while the application region sat erased.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 0.3)
check("undeclared mismatch: reported", False, "unexpectedly succeeded")
except pb.Error as error:
text = str(error).lower()
check("undeclared mismatch: error names the link, not just a timeout",
"link" in text or "baud" in text or "backend" in text, str(error))
except TypeError as error:
check("undeclared mismatch: error names the link, not just a timeout",
False, repr(error))
# The resident's own link must be restored for the caller: a declared
# staging link is for the copy, and the tool talks to the new resident after.
port = TwoLinkPort(resident=(57600, False), staged=(38400, False))
loader = connected(port)
try:
loader.enter_copy(0x7C00, 2.0, link=(38400, False))
check("session records the link it is now speaking", loader.baud == 38400,
f"loader.baud={getattr(loader, 'baud', None)}")
except (pb.Error, TypeError, AttributeError) as error:
check("session records the link it is now speaking", False, repr(error))
print(f"\n {P} passed, {F} failed")
return 1 if F else 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -144,7 +144,7 @@ class Host:
self._expect(CONFIRM, "C end")
return echo
# Activation when the config page carries a password: 3×'@' then the
# Activation when the config page carries a password: 3x'@' then the
# password bytes, then the info block + mainloop '!'.
def activate_password(self, password):
self.s.reset_input_buffer()
@@ -167,6 +167,20 @@ class Host:
self._expect(CONFIRM, "emergency mainloop ready")
# A wrong password byte hangs the loader, still draining the line. Two
# things must not happen: it must not activate, and it must not fall
# through to the emergency erase - a byte the gate has already refused
# reaching the erase would let a guess wipe the part.
def refuse_password(self, byte):
self.s.reset_input_buffer()
self.s.write(bytes([KNOCK, KNOCK, KNOCK, byte]))
return self.s.read(1)
def say(self, byte):
self.s.write(bytes([byte]))
return self.s.read(1)
def check(cond, msg):
if not cond:
raise AssertionError(msg)
@@ -181,7 +195,7 @@ PW_BYTES = bytes([0x50, 0x57])
def scenario_roundtrip(host):
"""Activation + info block + flash/EEPROM/config read-write round-trips, on
a device with a blank (erased) config page the usual no-password case."""
a device with a blank (erased) config page - the usual no-password case."""
info = host.activate()
check(info[0:3] == b"TSB", f"magic 'TSB' (got {info[0:3]!r})")
check(info[6:9] == bytes([0x1E, 0x95, 0x0F]), f"signature 1E 95 0F (got {info[6:9].hex()})")
@@ -219,6 +233,15 @@ def scenario_emergency(host):
check(host.read_eeprom(1) == b"\xff" * PAGE, "EEPROM wiped")
def scenario_wrong_password(host):
"""A wrong password byte neither activates the loader nor opens the
emergency erase behind it - the oracle carries a dedicated fix for the
second, and nothing here exercised either half."""
check(host.refuse_password(PW_BYTES[0] ^ 1) == b"", "a wrong password byte draws no reply")
check(host.say(0x00) == b"", "a 0 byte after it does not request the erase")
check(host.say(CONFIRM) == b"", "and neither does a confirm")
def main():
binary, elf, boot_base = sys.argv[1], sys.argv[2], sys.argv[3]
failures = []
@@ -229,6 +252,7 @@ def main():
("round-trip", None, scenario_roundtrip),
("password activation", PW_CONFIG, scenario_password),
("emergency erase", PW_CONFIG, scenario_emergency),
("wrong password", PW_CONFIG, scenario_wrong_password),
]
for name, config, fn in groups:
print(f"--- {name} ---")

View File

@@ -1,39 +1,68 @@
#!/bin/bash
# The port's gate: every chip's generated workflow build, size matrix, and
# The port's gate: every chip's generated workflow - build, size matrix, and
# the simulator-driven protocol suites. --full adds the reflect-spot builds
# (libavr's rule: reflect compiles are bounded to its spot set, never the
# full matrix) and swaps the compact size matrix for the exhaustive
# clock × baud × backend cross product. LIBAVR_ROOT must point at the libavr
# checkout.
# clock x baud x backend cross product. libavr resolves from the `libavr/`
# submodule; LIBAVR_ROOT overrides it for a working tree.
set -e
cd "$(dirname "$0")/.."
full=0
[[ "$1" == "--full" ]] && { full=1; shift; export PUREBOOT_FULL_MATRIX=1; }
CHIPS=(attiny13 attiny13a attiny25 attiny45 attiny85
atmega8 atmega8a atmega16 atmega16a atmega32 atmega32a
atmega48 atmega48a atmega48p atmega48pa
atmega88 atmega88a atmega88p atmega88pa
atmega168 atmega168a atmega168p atmega168pa
atmega328 atmega328p
atmega164a atmega164p atmega164pa
atmega324a atmega324p atmega324pa
atmega644 atmega644a atmega644p atmega644pa
atmega1284 atmega1284p)
REFLECT_SPOT=(attiny13a attiny85 atmega8 atmega16a atmega32a atmega48pa
atmega88 atmega168pa atmega328p atmega164a atmega644p atmega1284)
# The chip lists come from the presets rather than being spelled a second time
# here: a chip added to make_presets.py and missed in a copy of its list would
# be a gate that silently never builds it, which is the one failure mode a gate
# cannot report. tools/make_presets.py is the single source, CMakePresets.json
# is its output, and this reads that.
readarray -t WORKFLOWS < <(python3 -c '
import json, sys
presets = json.load(open("CMakePresets.json"))["workflowPresets"]
print("\n".join(p["name"] for p in presets))')
if ((${#WORKFLOWS[@]} == 0)); then
echo "no workflow presets in CMakePresets.json - run tools/make_presets.py" >&2
exit 1
fi
CHIPS=()
REFLECT_SPOT=()
for workflow in "${WORKFLOWS[@]}"; do
case $workflow in
*-generated) CHIPS+=("${workflow%-generated}") ;;
*-reflect) REFLECT_SPOT+=("${workflow%-reflect}") ;;
esac
done
# Every preset runs even after one goes red, and the gate fails at the end
# naming all of them: stopping at the first failure turns a red - a stale size
# canary above all - into an alibi for every chip behind it, and a loader can
# ship on a chip this gate has not compiled since.
red=()
run_preset() {
echo "==== $1 ===="
cmake --workflow --preset "$1" "${@:2}" || red+=("$1")
}
for chip in "${CHIPS[@]}"; do
echo "==== $chip ===="
cmake --workflow --preset "$chip-generated" "$@"
run_preset "$chip-generated" "$@"
done
if ((full)); then
for chip in "${REFLECT_SPOT[@]}"; do
echo "==== $chip reflect ===="
cmake --workflow --preset "$chip-reflect" "$@"
run_preset "$chip-reflect" "$@"
done
fi
if ((${#red[@]})); then
printf '==== red presets ====\n' >&2
printf ' %s\n' "${red[@]}" >&2
exit 1
fi
# Every tree is freshly built now - the one moment the README's size table
# can be held to what the images measure (a per-preset ctest sees only its
# own chip; the table needs all of them, and ungated it drifts: a
# common-code shave moves every row at once with nothing over budget).
python3 tools/sizes.py check-readme
echo "check: every chip green"

View File

@@ -1,17 +1,20 @@
#!/usr/bin/env python3
"""Regenerate CMakePresets.json one uniform pipeline per chip.
"""Regenerate CMakePresets.json - one uniform pipeline per chip.
Every chip gets generated-mode configure/build/test presets and a workflow
running all three. Reflect-mode presets (configure + build, no tests the
running all three. Reflect-mode presets (configure + build, no tests - the
port's TUs compile identically; the sims prove nothing new there) exist for
libavr's reflect spot set only, mirroring its rule: the full reflect matrix
is never built, one chip per hardware class and pack vintage is.
Run from the repo root: tools/make_presets.py
Run from the repo root: tools/make_presets.py - or with --check, which
verifies the committed file matches this generator and edits nothing (the
ctest entry `presets.generated` runs that, so drift reds the gate).
"""
import json
import os
import sys
CHIPS = [
"attiny13", "attiny13a", "attiny25", "attiny45", "attiny85",
@@ -41,7 +44,7 @@ def main():
"hidden": True,
"generator": "Ninja",
"binaryDir": "${sourceDir}/build/${presetName}",
"toolchainFile": "$env{LIBAVR_ROOT}/cmake/avr-toolchain.cmake",
"toolchainFile": "${sourceDir}/libavr/cmake/avr-toolchain.cmake",
"cacheVariables": {
"CMAKE_BUILD_TYPE": "Release",
"CMAKE_EXPORT_COMPILE_COMMANDS": "ON",
@@ -72,6 +75,9 @@ def main():
for chip in REFLECT_SPOT:
add(chip, "reflect")
# CMake rejects unknown fields in the presets root, $comment included, so
# the file cannot carry a generated-file marker; the --check ctest is the
# whole of rule 10's guard here.
presets = {
"version": 8,
"configurePresets": configure,
@@ -79,12 +85,19 @@ def main():
"testPresets": test,
"workflowPresets": workflows,
}
rendered = json.dumps(presets, indent=1) + "\n"
path = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "CMakePresets.json")
if "--check" in sys.argv[1:]:
current = open(path).read() if os.path.exists(path) else ""
if current != rendered:
print("CMakePresets.json does not match its generator - run tools/make_presets.py")
return 1
return 0
with open(path, "w") as f:
json.dump(presets, f, indent=1)
f.write("\n")
f.write(rendered)
print(f"{len(CHIPS)} chips, {len(REFLECT_SPOT)} reflect: {os.path.normpath(path)}")
return 0
if __name__ == "__main__":
main()
sys.exit(main())

377
tools/pbhw.py Executable file
View File

@@ -0,0 +1,377 @@
#!/usr/bin/env python3
"""Hardware acceptance suite for a pureboot deployment.
`tools/check.sh` proves the protocol under simavr on every chip. This proves one
*board*: that the loader actually installed on it answers, that the memories
round-trip over the real link, that the application it flashes runs afterwards,
and that the refusals which keep a 512-byte slot alive still fire. Run it once
when a board is brought up, and again whenever the deployment moves - a new
clock, a new backend, new pins.
Every check derives its bounds from the info block the loader itself reports, so
nothing here is per-chip: the same run covers a 1 KiB tiny whose application
region is 510 usable bytes and a 128 KiB mega whose flash needs a bank in the
selector.
**This overwrites the board's application flash and EEPROM.** Capture them first
with `pbrig.py backup`, which verifies what it captured.
tools/pbhw.py --programmer atmelice_isp --part t13 --port COM6 \
--autobaud --loader build/ab.bin --app build/pbapp.hex \
--marker APP
"""
from __future__ import annotations
import argparse
import pathlib
import sys
import tempfile
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
import pbrig # noqa: E402
class Suite:
def __init__(self, rig: pbrig.Rig, work: pathlib.Path):
self.rig = rig
self.work = work
self.results: list[tuple[str, bool, str]] = []
def check(self, name: str, ok: bool, detail: str = "") -> bool:
self.results.append((name, ok, detail))
print(f" {'PASS' if ok else 'FAIL'} {name}" + (f" {detail}" if detail else ""))
return ok
@staticmethod
def _brief(text: str, limit: int = 78) -> str:
return " | ".join(l.strip() for l in text.splitlines() if l.strip())[:limit]
# ----------------------------------------------------------------- checks
def identity(self) -> object | None:
"""The info block, which every later check takes its bounds from."""
module = pbrig.load_pureboot(self.rig.d.pureboot)
self.rig.reset()
port = self.rig.open_port() # wrapped for the echo where the line is shared
try:
loader = module.Loader(port)
if self.rig.d.autobaud:
loader.connect_autobaud(self.rig.d.wait)
else:
loader.connect(self.rig.d.wait)
info = loader.info
self.check("identity read", True, info.describe())
return info
except Exception as error: # noqa: BLE001 - a dead link is a result
self.check("identity read", False, str(error)[:70])
return None
finally:
try:
port.close()
except Exception:
pass
def scan(self) -> None:
"""The --scan walk against real termios and a real oscillator: every
probe rate must open a port (the off-nominal rates exist only through
termios2), and one probe must answer - the nominal on a healthy board,
a neighbor on a drifted one. The rig injects the one reset per probe
the operator supplies in the field; this is the rate physics the
simulator cannot arbitrate (a pty carries bytes at any rate), pinned
on silicon."""
module = pbrig.load_pureboot(self.rig.d.pureboot)
found = None
try:
for pct in module.scan_ratios():
rate = module.scan_rate(self.rig.d.baud, pct)
self.rig.reset()
try:
# Same wrap as identity(): on a shared line an undiscarded
# echo answers every rate a scan probes, so the walk would
# report the first one it tried.
port = self.rig.open_port(rate)
except module.Error as error:
self.check("scan opens every probe rate", False, f"{rate} Bd: {error}")
return
try:
module.Loader(port).connect(min(self.rig.d.wait, 6.0))
found = pct
break
except module.Error:
continue
finally:
port.close()
except Exception as error: # noqa: BLE001 - a rig hiccup is a result
self.check("scan walks the probe ladder", False, str(error)[:70])
return
self.check("scan finds the board's rate", found is not None,
"no probe answered" if found is None else f"{found:+d} % of {self.rig.d.baud} Bd")
def eeprom(self, info) -> None:
size = info.eeprom_size
if not size:
print(" skip EEPROM (this part has none)")
return
# A pattern no erase or partial write could produce by accident.
pattern = bytes((i * 7 + 3) & 0xFF for i in range(size))
image = self.work / "ee.bin"
image.write_bytes(pattern)
rc, out = self.rig.pureboot("--eeprom", str(image), "--verify-eeprom", str(image))
self.check(f"EEPROM write + verify ({size} B)", rc == 0, self._brief(out))
back = self.work / "ee-back.bin"
rc, out = self.rig.pureboot("--read-eeprom", str(back))
got = back.read_bytes() if back.exists() else b""
self.check("EEPROM reads back what was written", got == pattern, f"{len(got)} B")
self.rig.pureboot("--erase-eeprom")
erased = self.work / "ee-erased.bin"
self.rig.pureboot("--read-eeprom", str(erased))
got = erased.read_bytes() if erased.exists() else b""
self.check("EEPROM erase leaves 0xff", got == b"\xff" * size, f"{len(got)} B")
def application(self, info, app: pathlib.Path, marker: str,
marker_wait: float = 2.5) -> None:
rc, out = self.rig.pureboot("--flash", str(app), "--verify-flash", str(app))
self.check(f"application flash + verify ({app.name})", rc == 0, self._brief(out))
if marker:
# The tool hands over as it ends its session, so the application is
# already running - but only on a board whose DTR is unwired, where
# opening a port simply listens. Where DTR *is* wired to reset (an
# Arduino, most USB-serial dev boards), this open resets the part
# and the activation window comes first, so a marker emitted once at
# startup happens on the far side of a wait this cannot know the
# length of: the window is a compile-time constant and nothing on
# the wire reports it. Hence --marker-wait, and a fixture that
# repeats its banner (PUREBOOT_HEARTBEAT) rather than saying it once.
data = self.rig.capture(seconds=marker_wait)
seen = marker.encode() in data
sample = "".join(chr(b) if 32 <= b < 127 else "." for b in data[:40])
self.check(f"application runs (emits {marker!r})", seen,
f"|{sample}|" if seen or data else
f"nothing in {marker_wait:g} s - if this board resets when its port "
f"opens, that wait has to outlast the activation window")
back = self.work / "app-back.bin"
rc, out = self.rig.pureboot("--read-flash", str(back))
got = back.read_bytes() if back.exists() else b""
self.check("application flash reads back", rc == 0 and len(got) == info.base,
f"{len(got)} B of {info.base}")
def _witness(self, info, slot_length: int):
"""Read back the erased region and the loader slot: (erased, slot, how).
Prefers ISP, because an independent reader is the only one that can
testify about a loader just asked to erase around itself. Where no
programmer is attached the link answers instead - which is weaker for
exactly the reason it is worth having, a destroyed loader being unable
to report anything at all. The two are never printed under one word:
an absent probe is a fact about the bench, a wrong byte is a verdict on
the loader, and a check that conflates them stops being read.
"""
limit = info.base - 2 if info.patch_vector else info.base
whole = self.work / "whole.bin"
if self.rig.read_memory("flash", whole, "r"):
image = whole.read_bytes()
image += b"\xff" * (info.flash_size - len(image))
return image[0:limit], image[info.base:info.base + slot_length], "ISP"
module = pbrig.load_pureboot(self.rig.d.pureboot)
port = self.rig.open_port()
try:
loader = module.Loader(port)
if self.rig.d.autobaud:
loader.connect_autobaud(self.rig.d.wait)
else:
loader.connect(self.rig.d.wait)
return (loader.read_flash(0, limit),
loader.read_flash(info.base, slot_length),
"the link, no probe attached - the loader's own account")
except Exception as error: # noqa: BLE001 - a dead link is a result
print(f" skip slot checks: no programmer, and the link did not "
f"answer either ({str(error)[:60]})")
return None, None, ""
finally:
try:
port.close()
except Exception: # noqa: BLE001
pass
def erase_and_slot(self, info, loader_image: pathlib.Path | None) -> None:
rc, out = self.rig.pureboot("--erase-flash")
self.check("application region erases", rc == 0, self._brief(out))
want = loader_image.read_bytes() if loader_image and loader_image.exists() else b""
erased, slot, how = self._witness(info, len(want))
if erased is None:
return
# Erased application flash, up to the trampoline word the host composes
# on a patched-vector part.
limit = info.base - 2 if info.patch_vector else info.base
self.check("erased application region is 0xff",
set(erased) <= {0xFF}, f"0x0000..{limit:#06x} via {how}")
if want:
self.check("loader slot survives the erase", slot == want,
f"{len(want)} B at {info.base:#06x} via {how}")
else:
print(" skip loader slot comparison (pass --loader <image.bin>)")
def seal(self, info, rounds: int = 1) -> None:
"""The seal, adversarially, over the real link.
pureboot 9 has no running-slot guard: what stops a mangled command from
erasing the loader is the seal and nothing else. So this aims the worst
command the protocol has - an SPM erase at the loader's own first page -
and damages one header byte at a time. Every one must come back NAK with
the slot untouched and the session still in step.
On a board whose link drops or mangles bytes of its own accord this is
also the stress test: `--seal-rounds` repeats it, and a link fault
during a round is indistinguishable to the loader from the damage being
injected, which is the point.
"""
module = pbrig.load_pureboot(self.rig.d.pureboot)
self.rig.reset()
port = self.rig.open_port()
try:
loader = module.Loader(port)
if self.rig.d.autobaud:
loader.connect_autobaud(self.rig.d.wait)
else:
loader.connect(self.rig.d.wait)
if loader.info.version < module.SEALED_LOADER:
print(f" skip seal checks (loader is pureboot {loader.info.version})")
return
head_of = lambda dmg: self._sealed(module, module.OP_WRITE, module.SP_SPM,
info.base, module.SPM_ERASE, dmg)
refused = 0
attempts = 0
for _ in range(rounds):
for index in range(6):
attempts += 1
port.write(head_of((index, 0x01)))
verdict = port.read_exact(1, 5.0)
if verdict != module.NAK:
self.check(f"damaged byte {index} refused", False,
f"verdict {verdict.hex()}")
return
if port.read_exact(1, 5.0) != module.PROMPT:
self.check(f"re-prompt after byte {index}", False, "no prompt")
return
refused += 1
self.check("damaged headers refused", refused == attempts,
f"{refused}/{attempts}, every header byte")
self.check("session still in step", loader.identity().raw == info.raw)
# And the slot itself, read back over the link: the loader is the
# thing that would have been erased, so its own account of its
# first bytes is a real witness - an erased page reads all 0xff.
head = loader.read_flash(info.base, 16)
self.check("loader slot intact", set(head) != {0xFF}, head[:8].hex())
except Exception as error: # noqa: BLE001 - a dead link is a result
self.check("seal checks", False, str(error)[:70])
finally:
try:
port.close()
except Exception: # noqa: BLE001
pass
@staticmethod
def _sealed(module, op, space, address, count, damage=None):
"""A sealed header, damaged after sealing - the shape a link fault has."""
head = bytearray((op, module.selector(space, address), address & 0xFF,
(address >> 8) & 0xFF, count & 0xFF))
seal = module.SEAL
for byte in head:
seal ^= byte
out = bytearray(head + bytes((seal,)))
if damage:
out[damage[0]] ^= damage[1]
return bytes(out)
def refusals(self, info) -> None:
# One word too many: a patched-vector part spends the slot's last word
# on the trampoline, so its application stops two bytes short.
limit = info.base - 2 if info.patch_vector else info.base
oversized = self.work / "oversized.bin"
oversized.write_bytes(bytes(limit + 2))
rc, out = self.rig.pureboot("--flash", str(oversized))
self.check(f"image over {limit} B refused", rc != 0, self._brief(out))
# ------------------------------------------------------------------- run
def run(self, app: pathlib.Path | None, loader_image: pathlib.Path | None,
marker: str, marker_wait: float = 2.5, seal_rounds: int = 1) -> int:
print("identity")
info = self.identity()
if info is None:
print("\nthe loader never answered; nothing below can be trusted")
return 1
if not self.rig.d.autobaud:
print("\nscan")
self.scan()
print("\nEEPROM")
self.eeprom(info)
if app:
print("\napplication")
self.application(info, app, marker, marker_wait)
else:
print("\nskip application checks (pass --app <image.hex>)")
print("\nerase and the slot boundary")
self.erase_and_slot(info, loader_image)
print("\nthe seal")
self.seal(info, seal_rounds)
print("\nrefusals")
self.refusals(info)
passed = sum(1 for _, ok, _ in self.results if ok)
print(f"\n{passed}/{len(self.results)} passed")
return 0 if passed == len(self.results) else 1
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="hardware acceptance suite for one pureboot deployment",
epilog="overwrites the board's application flash and EEPROM - back them up first")
pbrig.Deployment.add_arguments(parser)
parser.add_argument("--app", type=pathlib.Path,
help="application image to flash (test/pbapp.cpp built for this deployment)")
parser.add_argument("--loader", type=pathlib.Path,
help="the resident loader's .bin, to prove the slot survives an erase")
parser.add_argument("--seal-rounds", type=int, default=1,
help="repeat the adversarial seal sweep N times (a lossy board's stress test)")
parser.add_argument("--marker", default="",
help="text the application emits when it runs, e.g. APP")
parser.add_argument("--marker-wait", type=float, default=2.5,
help="seconds to listen for it. On a board whose DTR is wired to "
"reset, opening the port resets the part, so this must outlast "
"the activation window (default 2.5)")
args = parser.parse_args(argv)
rig = pbrig.Rig(pbrig.Deployment.from_args(args))
print(f"rig: {args.part} on {args.programmer}, link {args.port} at {args.baud} Bd"
f"{' (autobaud)' if args.autobaud else ''}")
print("this overwrites the application flash and EEPROM\n")
with tempfile.TemporaryDirectory(prefix="pbhw-") as temporary:
return Suite(rig, pathlib.Path(temporary)).run(args.app, args.loader, args.marker,
args.marker_wait, args.seal_rounds)
if __name__ == "__main__":
try:
sys.exit(main())
except pbrig.Error as error:
print(f"error: {error}", file=sys.stderr)
sys.exit(2)

445
tools/pbrig.py Executable file
View File

@@ -0,0 +1,445 @@
#!/usr/bin/env python3
"""Hardware rig driver for pureboot: an ISP programmer beside a serial link.
The simulated suites (`test/pb*.py`) prove the protocol; this drives the same
loader on real silicon, where the things a cycle-exact simulator cannot model
live - an RC oscillator off its nominal, a reset edge that has to come from
somewhere, a serial bridge with its own idea of what a baud is.
Nothing here knows a port name, a part or a programmer. Every deployment fact
arrives from the command line or the environment, so the same script serves any
board: see `Deployment`. As a module it is the reset/flash/talk primitives that
`pbhw.py` builds its acceptance suite from; as a command it is the handful of
one-shot operations worth having on a rig - most importantly `backup`, which is
the only thing standing between a fuse experiment and an unrecoverable part.
Two rig facts are encoded here because they are not guessable and cost a
session each to learn:
* **An ISP access resets the part**, and it runs again the moment the programmer
releases it. That is the only reset edge available when the serial adapter's
DTR is not wired to reset - so a loader session begins with an ISP touch and
knocks immediately after, which is what `Rig.pureboot()` does.
* **avrdude splits `-U memory:op:file:format` on colons**, so a Windows path's
drive letter breaks the spec. Every file argument is therefore passed as a
bare filename with avrdude run in that file's own directory.
"""
from __future__ import annotations
import argparse
import dataclasses
import importlib.util
import os
import pathlib
import subprocess
import sys
import time
HERE = pathlib.Path(__file__).resolve().parent
DEFAULT_PUREBOOT = HERE.parent / "pureboot" / "pureboot.py"
# Memories worth capturing before an experiment, and the format each is read in.
# Fuses and lock are per-part: a part without an extended fuse simply fails that
# one read, which `backup` reports and steps over rather than aborting on.
BACKUP_MEMORIES = (
("flash", "i", "hex"),
("flash", "r", "bin"),
("eeprom", "i", "hex"),
("eeprom", "r", "bin"),
("lfuse", "h", "hex"),
("hfuse", "h", "hex"),
("efuse", "h", "hex"),
("lock", "h", "hex"),
("calibration", "h", "hex"),
)
class Error(Exception):
pass
def bitclock_for(hz: int) -> str:
"""A safe ISP bitclock for a part *currently running* at `hz`.
SCK must stay under a quarter of the target clock, so the bitclock follows
the clock in force - not the one about to be fused in. Halving that ceiling
again costs nothing on a link that moves a few hundred bytes and buys margin
against an oscillator that is already known to be off its nominal.
"""
ceiling = hz // 8
rungs = (1000, 4000, 8000, 32000, 125000, 400000)
if ceiling < rungs[0]:
raise Error(f"a part at {hz} Hz is too slow to reach over ISP safely")
best = max(rung for rung in rungs if rung <= ceiling)
return f"{best // 1000}kHz"
@dataclasses.dataclass
class Deployment:
"""Everything about one board. No default names a real device."""
port: str = "" # serial device the loader speaks on
baud: int = 57600 # host rate; for autobaud, the rate to drive
autobaud: bool = False # send the calibration pulse instead of p+b
one_wire: bool = False # shared line: the host discards its own echo
programmer: str = "" # avrdude -c
part: str = "" # avrdude -p
avrdude: str = "avrdude"
bitclock: str = "125kHz" # see bitclock_for()
pureboot: pathlib.Path = DEFAULT_PUREBOOT
wait: int = 12 # seconds the host keeps knocking
@classmethod
def from_env(cls) -> "Deployment":
"""Environment defaults, so a rig's facts live in one place per machine."""
return cls(
port=os.environ.get("PUREBOOT_PORT", ""),
baud=int(os.environ.get("PUREBOOT_BAUD", "57600")),
autobaud=os.environ.get("PUREBOOT_AUTOBAUD", "") not in ("", "0"),
one_wire=os.environ.get("PUREBOOT_ONE_WIRE", "") not in ("", "0"),
programmer=os.environ.get("PUREBOOT_PROGRAMMER", ""),
part=os.environ.get("PUREBOOT_PART", ""),
avrdude=os.environ.get("AVRDUDE", "avrdude"),
bitclock=os.environ.get("PUREBOOT_BITCLOCK", "125kHz"),
pureboot=pathlib.Path(os.environ.get("PUREBOOT_TOOL", str(DEFAULT_PUREBOOT))),
)
@staticmethod
def add_arguments(parser: argparse.ArgumentParser) -> None:
"""Deployment flags, shared by this tool and pbhw.py."""
env = Deployment.from_env()
parser.add_argument("--port", default=env.port, help="serial device the loader speaks on")
parser.add_argument("--baud", type=int, default=env.baud,
help="host rate (for autobaud, the rate to drive)")
parser.add_argument("--autobaud", action="store_true", default=env.autobaud,
help="send the calibration pulse instead of the p+b knock")
parser.add_argument("--one-wire", action="store_true", default=env.one_wire,
help="shared line: pass the tool its echo discard")
parser.add_argument("--programmer", default=env.programmer, help="avrdude -c, e.g. atmelice_isp")
parser.add_argument("--part", default=env.part, help="avrdude -p, e.g. t13 or m328p")
parser.add_argument("--avrdude", default=env.avrdude, help="path to avrdude")
parser.add_argument("--bitclock", default=env.bitclock, help="ISP bitclock, e.g. 125kHz or 8kHz")
parser.add_argument("--pureboot", type=pathlib.Path, default=env.pureboot,
help="path to pureboot.py")
parser.add_argument("--wait", type=int, default=env.wait, help="seconds to keep knocking")
@classmethod
def from_args(cls, args: argparse.Namespace) -> "Deployment":
return cls(port=args.port, baud=args.baud, autobaud=args.autobaud,
one_wire=args.one_wire,
programmer=args.programmer, part=args.part, avrdude=args.avrdude,
bitclock=args.bitclock, pureboot=args.pureboot, wait=args.wait)
def load_pureboot(path: pathlib.Path = DEFAULT_PUREBOOT):
"""The host tool as a module - its Port and Loader, not a subprocess.
Used where a subprocess cannot express what is needed: a poke followed by a
peek in the *same* session, or a raw read at an arbitrary baud.
"""
spec = importlib.util.spec_from_file_location("pureboot", path)
if spec is None or spec.loader is None:
raise Error(f"cannot load the host tool from {path}")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
class Rig:
"""One board: its programmer on one side, its serial link on the other."""
def __init__(self, deployment: Deployment):
self.d = deployment
if not deployment.programmer or not deployment.part:
raise Error("a rig needs --programmer and --part")
# ------------------------------------------------------------- programmer
def avrdude(self, *args: str, cwd: pathlib.Path | None = None,
bitclock: str | None = None, timeout: int = 300) -> subprocess.CompletedProcess:
command = [self.d.avrdude, "-c", self.d.programmer, "-p", self.d.part,
"-B", bitclock or self.d.bitclock, *args]
return subprocess.run(command, capture_output=True, text=True,
cwd=None if cwd is None else str(cwd), timeout=timeout)
@staticmethod
def _ok(result: subprocess.CompletedProcess) -> bool:
return result.returncode == 0
def reset(self, bitclock: str | None = None) -> None:
"""An ISP access, which resets the part; it runs when avrdude exits."""
self.avrdude("-U", "signature:r:-:h", bitclock=bitclock)
def signature(self, bitclock: str | None = None) -> str:
result = self.avrdude("-U", "signature:r:-:h", bitclock=bitclock)
for line in reversed(result.stdout.splitlines()):
if line.strip().startswith("0x"):
return line.strip()
raise Error(f"no signature read: {(result.stderr or result.stdout).strip()[:200]}")
def read_memory(self, memory: str, destination: pathlib.Path, fmt: str = "r",
bitclock: str | None = None) -> bool:
"""Read `memory` into `destination`, whose directory avrdude runs in."""
destination = pathlib.Path(destination).resolve()
destination.parent.mkdir(parents=True, exist_ok=True)
result = self.avrdude("-U", f"{memory}:r:{destination.name}:{fmt}",
cwd=destination.parent, bitclock=bitclock)
# A memory the part does not have (a tiny's extended fuse) leaves avrdude
# happy and the file empty. An empty capture is a miss, not a backup.
return self._ok(result) and destination.exists() and destination.stat().st_size > 0
def write_memory(self, memory: str, source: pathlib.Path, fmt: str = "i",
erase: bool = False, bitclock: str | None = None) -> bool:
source = pathlib.Path(source).resolve()
args = ["-U", f"{memory}:w:{source.name}:{fmt}"]
if erase:
args.insert(0, "-e")
result = self.avrdude(*args, cwd=source.parent, bitclock=bitclock)
return "verified" in (result.stdout + result.stderr)
def flash_hex(self, image: pathlib.Path, erase: bool = True,
bitclock: str | None = None) -> bool:
return self.write_memory("flash", image, "i", erase=erase, bitclock=bitclock)
def read_fuses(self, bitclock: str | None = None) -> dict[str, str]:
out: dict[str, str] = {}
for fuse in ("lfuse", "hfuse", "efuse", "lock"):
result = self.avrdude("-U", f"{fuse}:r:-:h", bitclock=bitclock)
values = [l.strip() for l in result.stdout.splitlines() if l.strip().startswith("0x")]
if values:
out[fuse] = values[-1]
return out
def write_fuses(self, bitclock: str | None = None, **fuses: str) -> bool:
"""Write named fuses. A fuse change moves the clock the *next* access is
timed against, so pass a bitclock safe for both sides of the change."""
args: list[str] = []
for name, value in fuses.items():
args += ["-U", f"{name}:w:{value}:m"]
if not args:
return True
result = self.avrdude(*args, bitclock=bitclock)
text = result.stdout + result.stderr
return "verified" in text or "written" in text
# ------------------------------------------------------------ backup
def backup(self, directory: pathlib.Path, prefix: str = "") -> dict[str, bool]:
"""Capture every memory worth keeping, then prove it by a second read.
A backup nobody verified is a guess. Each memory is read twice and the
two reads compared; a mismatch is reported rather than quietly stored.
"""
directory = pathlib.Path(directory).resolve()
directory.mkdir(parents=True, exist_ok=True)
stem = prefix or self.d.part
status: dict[str, bool] = {}
for memory, fmt, extension in BACKUP_MEMORIES:
name = f"{stem}-{memory}.{extension}"
if not self.read_memory(memory, directory / name, fmt):
status[f"{memory}.{extension}"] = False
continue
if extension == "bin": # only the raw form is worth comparing byte-wise
again = directory / f".{name}.again"
self.read_memory(memory, again, fmt)
same = again.exists() and again.read_bytes() == (directory / name).read_bytes()
again.unlink(missing_ok=True)
status[f"{memory}.{extension}"] = same
else:
status[f"{memory}.{extension}"] = True
return status
# ------------------------------------------------------------ serial link
def pureboot(self, *args: str, reset_first: bool = True, baud: int | None = None,
autobaud: bool | None = None, timeout: int = 300,
bitclock: str | None = None) -> tuple[int, str]:
"""Reset, then knock immediately - see the module docstring.
Returns the host tool's exit status and its combined output, so a caller
can assert on what it printed as well as on whether it succeeded.
"""
if reset_first:
self.reset(bitclock=bitclock)
command = [sys.executable, str(self.d.pureboot), "--port", self.d.port,
"--baud", str(self.d.baud if baud is None else baud),
"--wait", str(self.d.wait)]
if self.d.autobaud if autobaud is None else autobaud:
command.append("--autobaud")
if self.d.one_wire:
command.append("--one-wire")
command += [str(a) for a in args]
try:
result = subprocess.run(command, capture_output=True, text=True, timeout=timeout)
except subprocess.TimeoutExpired as expired:
return 99, f"TIMEOUT after {timeout}s\n{expired.stdout or ''}{expired.stderr or ''}"
return result.returncode, (result.stdout or "") + (result.stderr or "")
def open_port(self, baud: int | None = None):
"""A port opened the way this deployment says to speak to the board.
Everything the rig runs as a *subprocess* gets its flags from
`pureboot()` above; anything that drives the protocol in-process has
to reach the same facts, and until this existed only the subprocess
path could. A shared line is the one where that gap is fatal rather
than untidy: the host reads back every byte it writes, so an
undiscarded echo answers the knock before the device does. Open
through here and a one-wire deployment cannot be silently driven as
a two-wire one.
"""
module = load_pureboot(self.d.pureboot)
port = module.Port(self.d.port, self.d.baud if baud is None else baud)
return module.OneWirePort(port) if self.d.one_wire else port
def capture(self, seconds: float = 2.0, baud: int | None = None) -> bytes:
"""Listen to whatever the board is saying, at an arbitrary rate.
Opening the port does not reset a board whose DTR is unwired, so this can
sample a running application repeatedly without disturbing it - which is
what makes the rate sweep below possible.
"""
module = load_pureboot(self.d.pureboot)
port = module.Port(self.d.port, self.d.baud if baud is None else baud)
try:
data = b""
deadline = time.monotonic() + seconds
while time.monotonic() < deadline:
chunk = port.read_available(0.2)
if chunk:
data += chunk
return data
finally:
try:
port.close()
except Exception:
pass
def measure_rate(rig: Rig, marker: bytes, built_baud: int, nominal_hz: int | None = None,
span_percent: float = 12.0, step_percent: float = 0.5,
seconds: float = 0.75) -> dict:
"""Find a transmitting board's true bit rate, using only the serial port.
The board must be emitting something recognisable at a *fixed* cycles-per-bit
- `test/pbapp.cpp` built with PUREBOOT_HEARTBEAT does. Since its bit timing is
a cycle count, its wire rate scales with its actual clock, so the host rates
at which `marker` still decodes bracket that rate; the centre of the band is
the answer, and with the clock the image was built for it gives the real one.
This is the measurement that turns "the loader is silent, so the wiring must
be wrong" into a number, and it needs no instrument beyond the adapter
already attached.
"""
steps = int(span_percent / step_percent)
clean: list[int] = []
samples: list[tuple[int, int, bool]] = []
for index in range(-steps, steps + 1):
baud = int(round(built_baud * (1 + index * step_percent / 100.0)))
if baud <= 0:
continue
data = rig.capture(seconds=seconds, baud=baud)
hit = marker in data
samples.append((baud, len(data), hit))
if hit:
clean.append(baud)
result: dict = {"samples": samples, "clean": clean, "built_baud": built_baud}
if clean:
low, high = min(clean), max(clean)
centre = (low + high) / 2.0
result |= {"low": low, "high": high, "centre": centre,
"half_width_percent": (high - low) / 2.0 / centre * 100.0,
"error_percent": (centre / built_baud - 1.0) * 100.0}
if nominal_hz:
result["measured_hz"] = nominal_hz * centre / built_baud
return result
# ------------------------------------------------------------------- command
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="pureboot hardware rig: ISP reset/flash beside the serial link")
Deployment.add_arguments(parser)
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser("signature", help="read the part signature over ISP")
sub.add_parser("reset", help="reset the part (an ISP access) and let it run")
sub.add_parser("fuses", help="read the fuse and lock bytes")
p = sub.add_parser("flash", help="program a hex image over ISP")
p.add_argument("image", type=pathlib.Path)
p.add_argument("--no-erase", action="store_true", help="do not chip-erase first")
p = sub.add_parser("backup", help="capture and verify every memory")
p.add_argument("directory", type=pathlib.Path)
p.add_argument("--prefix", default="", help="filename stem (default: the part name)")
p = sub.add_parser("rate", help="measure the board's true bit rate and clock")
p.add_argument("--marker", default="APP", help="text the board emits (default: APP)")
p.add_argument("--built-baud", type=int, required=True,
help="the baud the running image was built for")
p.add_argument("--nominal-hz", type=int, default=0,
help="the clock the image was built for, to report the real one")
p.add_argument("--span", type=float, default=12.0, help="sweep +-this many percent")
p.add_argument("--step", type=float, default=0.5, help="sweep step in percent")
p.add_argument("--verbose", action="store_true", help="print every step")
p = sub.add_parser("bitclock", help="a safe ISP bitclock for a clock in force")
p.add_argument("hz", type=int)
args = parser.parse_args(argv)
if args.command == "bitclock":
print(bitclock_for(args.hz))
return 0
rig = Rig(Deployment.from_args(args))
if args.command == "signature":
print(rig.signature())
elif args.command == "reset":
rig.reset()
print("reset")
elif args.command == "fuses":
for name, value in rig.read_fuses().items():
print(f"{name:<6} {value}")
elif args.command == "flash":
ok = rig.flash_hex(args.image, erase=not args.no_erase)
print(f"{args.image.name}: {'verified' if ok else 'FAILED'}")
return 0 if ok else 1
elif args.command == "backup":
status = rig.backup(args.directory, args.prefix)
for name, ok in status.items():
print(f" {'ok ' if ok else 'FAIL'} {name}")
missing = [n for n, ok in status.items() if not ok]
# Fuses a part does not have are expected misses, not failures.
fatal = [n for n in missing if not n.startswith(("efuse", "calibration"))]
print(f"\n{len(status) - len(missing)}/{len(status)} captured into {args.directory}")
return 1 if fatal else 0
elif args.command == "rate":
result = measure_rate(rig, args.marker.encode(), args.built_baud,
args.nominal_hz or None, args.span, args.step)
if args.verbose:
for baud, size, hit in result["samples"]:
print(f" {baud:7d} Bd {size:5d} B {'MARKER' if hit else ''}")
if not result["clean"]:
print(f"no capture contained {args.marker!r} at any rate - is the board "
f"transmitting, and on the pin this port is wired to?")
return 1
print(f"clean band {result['low']}..{result['high']} Bd")
print(f"centre {result['centre']:.0f} Bd "
f"(+-{result['half_width_percent']:.1f} %)")
print(f"vs built {result['built_baud']} Bd ({result['error_percent']:+.1f} %)")
if "measured_hz" in result:
print(f"true clock {result['measured_hz'] / 1e6:.3f} MHz")
return 0
if __name__ == "__main__":
try:
sys.exit(main())
except Error as error:
print(f"error: {error}", file=sys.stderr)
sys.exit(2)

194
tools/sizes.py Executable file
View File

@@ -0,0 +1,194 @@
#!/usr/bin/env python3
"""What the loader images actually measure, and whether the README still agrees.
The size matrix asserts every image fits its slot; it says nothing about the
numbers the README prints, and those drift. Every row of that table was eight
bytes stale once `startup::caller_page()` landed - common code, so every build
moved at once and no test noticed, because none of them was over budget.
Two questions, both answered from built trees:
sizes.py max the largest image per chip, and anything over budget
sizes.py check-readme the README's per-chip table against what is built
Nothing here knows a chip's geometry. The (image, budget) pairs come from each
build's own `CTestTestfile.cmake` - the same values the gate checks - so the
slot rules stay where they belong, in `pureboot/CMakeLists.txt`, and a chip
added or a budget changed needs no edit here. Only trees a configure preset
still owns are read: a stale directory keeps its last build, and a loader built
before a slot changed will happily report a size that was true once
(`tools/prune-build-trees.sh` in libavr removes them).
Sizes come from `avr-size`, and a target is only as current as its last build -
run the gate first if you want the table checked against today's source.
"""
from __future__ import annotations
import argparse
import pathlib
import re
import shutil
import subprocess
import sys
ROOT = pathlib.Path(__file__).resolve().parents[1]
# add_test(<name>.size ... -DELF=<path> ... -DLIMIT=<n> ...) - the gate's own
# pairing of an image with the budget it must fit.
# ctest writes the name as a bracket argument ([=[name.size]=]) and quotes the
# rest, so the name starts after the bracket and the path ends at the quote.
SIZE_TEST = re.compile(r'add_test\(\s*\[=\[(?P<name>[^\]]+?)\.size\]=\][^\n]*?'
r'-DELF=(?P<elf>[^"\s]+)[^\n]*?-DLIMIT=(?P<limit>\d+)')
def avr_size() -> str:
for env in (ROOT / "../../toolchain").resolve().glob("avr-gcc-*/bin/avr-size"):
if env.is_file():
return str(env)
found = shutil.which("avr-size")
if not found:
sys.exit("no avr-size found (build the toolchain, or put it on PATH)")
return found
def preset_dirs() -> list[pathlib.Path]:
"""Build trees a configure preset still owns, newest-listed first."""
listing = subprocess.run(["cmake", "--list-presets"], cwd=ROOT, capture_output=True, text=True)
names = re.findall(r'^\s*"(.+)"$', listing.stdout, re.MULTILINE)
if not names:
sys.exit("cmake --list-presets returned nothing - run from a configured checkout")
return [d for d in (ROOT / "build" / n for n in names) if (d / "CTestTestfile.cmake").is_file()]
def measure(paths: list[str], tool: str) -> dict[str, int]:
""".text per ELF, in one avr-size call per batch."""
sizes: dict[str, int] = {}
for start in range(0, len(paths), 400):
batch = [p for p in paths[start:start + 400] if pathlib.Path(p).is_file()]
if not batch:
continue
out = subprocess.run([tool, *batch], capture_output=True, text=True).stdout
for line in out.splitlines()[1:]:
fields = line.split()
if len(fields) >= 6 and fields[0].isdigit():
sizes[fields[5]] = int(fields[0])
return sizes
def collect() -> dict[str, list[tuple[str, int, int]]]:
"""chip -> [(target, text, limit)], from every owned build tree."""
tool = avr_size()
found: dict[str, list[tuple[str, str, int]]] = {}
for tree in preset_dirs():
chip = tree.name.split("-")[0]
for match in SIZE_TEST.finditer((tree / "CTestTestfile.cmake").read_text()):
found.setdefault(chip, []).append((match["name"], match["elf"], int(match["limit"])))
sizes = measure([elf for rows in found.values() for _, elf, _ in rows], tool)
# A chip's generated and reflect trees must answer with the same bytes
# (the identity invariant), so the same target measuring two sizes means
# a stale tree - or an identity breach. Either is a finding; picking one
# silently is how a gate reports another build's numbers as today's.
for chip, rows in found.items():
seen: dict[str, tuple[int, str]] = {}
for name, elf, _ in rows:
if elf not in sizes:
continue
if name in seen and seen[name][0] != sizes[elf]:
sys.exit(f"{chip} {name}: {seen[name][0]} B in {seen[name][1]} but "
f"{sizes[elf]} B in {elf} - a stale tree (rebuild or remove it) "
f"or a cross-mode identity breach")
seen.setdefault(name, (sizes[elf], elf))
measured = {
chip: sorted(((name, sizes[elf], limit) for name, elf, limit in rows if elf in sizes),
key=lambda row: -row[1])
for chip, rows in sorted(found.items())
}
# A configured-but-unbuilt preset registers its tests with no images behind
# them; it is not a chip with nothing to say, it is a chip not built yet.
return {chip: rows for chip, rows in measured.items() if rows}
def cmd_max(args) -> int:
measured = collect()
if not measured:
sys.exit("nothing built - configure and build a preset first")
over = []
print(f"{'chip':<13} {'largest image':<34} {'.text':>6} {'budget':>7} headroom")
for chip, rows in measured.items():
name, text, limit = rows[0]
flag = "OVER" if text > limit else f"{limit - text:>5} B"
print(f"{chip:<13} {name:<34} {text:>6} {limit:>7} {flag}")
over += [(chip, n, t, l) for n, t, l in rows if t > l]
total = sum(len(rows) for rows in measured.values())
print(f"\n{total} images across {len(measured)} chips")
if over:
print("\nOVER BUDGET:")
for chip, name, text, limit in over:
print(f" {chip} {name}: {text} > {limit}")
return 1
tightest = min(((chip, n, t, l) for chip, rows in measured.items() for n, t, l in rows),
key=lambda row: row[3] - row[2])
chip, name, text, limit = tightest
print(f"tightest fit: {chip} {name} - {text} of {limit}, {limit - text} B spare")
return 0
def cmd_check_readme(args) -> int:
"""The README's per-chip table, against the stock build and the worst
autobaud configuration (OSCCAL baked, plus the USART-pin release where
the chip has a USART; the one-wire fold of the same build is its twin
and competes for the same cell) - the config the Autobaud column
documents."""
readme = (ROOT / "pureboot" / "README.md").read_text()
measured = collect()
rows = re.findall(r"^\|\s*(AT\w+[^|]*?)\s*\|[^|]*\|[^|]*\|[^|]*\|\s*(\d+) B\s*\|\s*(\d+) B\s*\|$",
readme, re.MULTILINE)
if not rows:
sys.exit("no size table found in pureboot/README.md")
bad = skipped = 0
for chips, stock_doc, auto_doc in rows:
# "ATmega48, 48A, 48P, 48PA" plus any footnote mark - the first name
# is the family's base, and the sub strips the rest.
chip = re.sub(r"[^a-z0-9]", "", chips.split(",")[0].strip().lower())
built = {name: text for name, text, _ in measured.get(chip, [])}
# The on-USART pair defines the column where the chip has a USART;
# the default-pin pair is the whole space elsewhere. Whichever twin
# measures larger is the number the cell must state.
candidates = [name for name in ("pureboot_autobaud_osccal_on_usart0",
"pureboot_1w_autobaud_osccal_on_usart0") if name in built]
if not candidates:
candidates = [name for name in ("pureboot_autobaud_osccal",
"pureboot_1w_autobaud_osccal") if name in built]
worst = max(candidates, key=lambda name: built[name], default="pureboot_autobaud_osccal")
for target, documented in (("pureboot", stock_doc), (worst, auto_doc)):
if target not in built:
skipped += 1
continue
if built[target] != int(documented):
print(f" {chip:<12} {target:<18} README says {documented} B, built is {built[target]} B")
bad += 1
if bad:
print(f"\n{bad} row(s) stale - update pureboot/README.md")
return 1
# A row whose target was not built is only skipped, so a chip-name change
# or a build tree that holds nothing would otherwise skip every row and
# report a match over an empty comparison.
if skipped == 2 * len(rows):
sys.exit(f"none of the {len(rows)} README rows matched a built image - "
f"refusing to report a match over nothing")
print(f"README size table matches every built image ({len(rows)} rows"
+ (f", {skipped} not built" if skipped else "") + ")")
return 0
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
subs = parser.add_subparsers(dest="cmd", required=True)
subs.add_parser("max", help="largest image per chip, and anything over budget")
subs.add_parser("check-readme", help="the README's size table against what is built")
args = parser.parse_args()
return {"max": cmd_max, "check-readme": cmd_check_readme}[args.cmd](args)
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -1,14 +1,15 @@
// TinySafeBoot on libavr tier 3: full feature parity in the 512-byte boot
// TinySafeBoot on libavr - tier 3: full feature parity in the 512-byte boot
// section, in C++ except where the C ABI itself is the cost.
//
// The complete TinySafeBoot feature set watchdog-reset bail, one-wire
// The complete TinySafeBoot feature set - watchdog-reset bail, one-wire
// half-duplex UART, a config-page activation timeout, the password gate,
// emergency erase, and config/flash/EEPROM read-write — at 510 bytes in the
// 512-byte BOOTSZ=11 section the hand-written oracle occupies (500 B). This tier used to be one
// monolithic inline-asm routine; it is now the tricks tier's C++ (same
// register protocol, same structure see tsb_tricks.cpp, including the
// global-register miscompile rules) with exactly two routines kept in
// assembly, the two whose remaining cost *is* the calling convention:
// emergency erase, and config/flash/EEPROM read-write - inside the 512-byte
// BOOTSZ=11 section the hand-written oracle occupies (oracle/README.md holds
// what each tier measures, in one table rather than four). The
// body is the tricks tier's C++ (same register protocol, same structure - see
// tsb_tricks.cpp, including the global-register miscompile rules) with exactly
// two routines kept in assembly, the two whose remaining cost *is* the calling
// convention:
//
// rx the bounded receive: C++ must re-floor the timeout window on every
// call (the global-register-store miscompile) and split it across
@@ -16,12 +17,12 @@
// countdown.
// store the page-store loop: C++ cannot hold the receive byte pair and the
// walked Z pointer across the rx calls without call-saved staging
// (push/pop + a YZ copy per word); the asm calls rx knowing exactly
// (push/pop + a Y->Z copy per word); the asm calls rx knowing exactly
// which registers it touches and walks Z live across the whole page.
//
// Everything else bring-up, activation, password gate, emergency erase,
// Everything else - bring-up, activation, password gate, emergency erase,
// dispatch, every SPM/EEPROM/flash primitive, every geometry/baud/info
// constant is C++ on libavr, and the two asm routines splice into the same
// constant - is C++ on libavr, and the two asm routines splice into the same
// global-register protocol the C++ uses (g_addr in Y, g_cnt in r16, g_window
// in r7, g_receiving in r6), so calls cross the boundary with no marshalling.
//
@@ -41,10 +42,16 @@ namespace hw = avr::hw;
namespace tsb {
namespace {
// The loader is purely polled it never enables interrupts so every SPM and
// The loader is purely polled - it never enables interrupts - so every SPM and
// EEPROM lock folds to nothing under this posture.
constexpr auto off = avr::irq::guard_policy::unused;
// Strict request/response: every SPM operation is waited out before the next
// byte moves, so no flash operation is ever in flight at an EEPROM access -
// the write procedure's step 2 has nothing to guard, the omission the
// datasheet grants (DS40002061B section 8.6.3).
constexpr auto no_spm = ee::spm_interlock::omitted;
constexpr std::uint8_t confirm = '!';
constexpr std::uint8_t request = '?';
constexpr std::uint8_t knock = '@';
@@ -57,28 +64,34 @@ constexpr std::uint16_t boot_bytes = 512;
constexpr std::uint16_t app_end = spm::flash_bytes - boot_bytes - page;
constexpr std::uint16_t eeprom_end = avr::hw::db.mem.eeprom_size - 1;
// Lockout-proof floor for the activation window (the oracle's F_CPU/1MHz).
constexpr std::uint8_t act_min = 16;
// Lockout-proof floor for the activation window: the oracle's F_CPU/1MHz, so
// it follows the clock rather than restating it (rule 41).
constexpr auto act_min = static_cast<std::uint8_t>((16_MHz).hz / 1'000'000);
// Post-activation window: the host gets seconds, not milliseconds, mid-session.
constexpr std::uint8_t comm_window = 200;
constexpr std::uint16_t build_date = 26 * 512 + 7 * 32 + 20;
// Fixed 115200 8N1; the library solves UBRR + U2X from clock and baud.
constexpr auto baud = avr::uart::detail::solve_baud(16_MHz, 115200_Bd);
constexpr auto baud = avr::uart::solve_baud(16_MHz, 115200_Bd, 8, avr::uart::parity::none);
// One bit time on the wire: the turn-around a shared-line peer needs to stop
// driving before this one starts. Derived from the solved rate, so it follows
// the link rather than a count measured against one.
constexpr auto guard_cycles = static_cast<std::uint32_t>((16_MHz).hz / baud.actual);
// The 16-byte device-info block, streamed out on activation.
// clang-format off
[[gnu::progmem]] constexpr std::uint8_t info[16] = {
[[gnu::progmem]] constexpr auto info = std::to_array<std::uint8_t>({
'T', 'S', 'B',
build_date & 0xFF, build_date >> 8,
0xF3, // status: native-UART fixed-baud lineage
0x1E, 0x95, 0x0F, // ATmega328P signature
avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2],
page / 2, // page size in words
(app_end / 2) & 0xFF, (app_end / 2) >> 8,
eeprom_end & 0xFF, eeprom_end >> 8,
0xAA, 0xAA,
};
});
// clang-format on
register std::uint16_t g_addr asm("r28");
@@ -94,7 +107,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
// Bounded byte receive (asm 1 of 2): release the one-wire line on a direction
// change, poll RXC0 under the oracle's nested X-register countdown seeded from
// g_window (floored against lockout), byte or 0-on-silence in r24. Z survives
// the property the store's word loop rides on.
// - the property the store's word loop rides on.
[[gnu::noinline, gnu::noclone]] std::uint8_t rx()
{
std::uint8_t byte;
@@ -115,7 +128,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
" brne 3b \n\t"
" sbiw r26, 1 \n\t"
" brcc 2b \n\t"
" clr %[b] \n\t" // silence 0, which no compare accepts
" clr %[b] \n\t" // silence -> 0, which no compare accepts
" rjmp 5f \n\t"
"4: lds %[b], %[udr0] \n\t"
"5: \n\t"
@@ -129,14 +142,13 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
// One-wire transmit: take the line (TXEN0 alone) on a direction change with a
// turn-around guard, put the byte out, hold the line until the whole frame is
// out (TXC0, not UDRE0), W1C TXC0 by storing the sampled status back (keeps
// U2X0). Plain C++ it compiles *smaller* than the oracle's routine.
// U2X0). Plain C++ - it compiles *smaller* than the oracle's routine.
[[gnu::noinline, gnu::noclone]] void tx(std::uint8_t byte)
{
if (g_receiving) {
g_receiving = 0;
hw::ucsr0b::write(hw::ucsr0b::txen0(1));
for (std::uint8_t guard = 46; guard; --guard)
;
avr::delay::cycles<guard_cycles>();
}
hw::udr0::write(byte);
std::uint8_t status;
@@ -153,7 +165,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
return rx();
}
// One flash byte [g_addr++] (the advance right before ret the
// One flash byte <- [g_addr++] (the advance right before ret - the
// global-register rule, see tsb_tricks.cpp).
[[gnu::noinline, gnu::noclone]] std::uint8_t sflash()
{
@@ -162,18 +174,18 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
return byte;
}
// One EEPROM byte [g_addr++].
// One EEPROM byte <- [g_addr++].
[[gnu::noinline, gnu::noclone]] std::uint8_t eerd()
{
std::uint8_t byte = ee::read(g_addr);
std::uint8_t byte = ee::read<no_spm>(g_addr);
++g_addr;
return byte;
}
// One EEPROM byte [g_addr++].
// One EEPROM byte -> [g_addr++].
[[gnu::noinline, gnu::noclone]] void eewr(std::uint8_t byte)
{
ee::write<off>(g_addr, byte);
ee::write<off, no_spm>(g_addr, byte);
++g_addr;
}
@@ -185,7 +197,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
} while (--g_cnt);
}
// Wait out a running SPM op, then re-open the RWW section after every page
// Wait out a running SPM op, then re-open the RWW section - after every page
// op and before handing over, as the oracle does.
[[gnu::noinline, gnu::noclone]] void settle()
{
@@ -201,12 +213,12 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
tsb_app();
}
// Step g_addr one page down and erase that page (the decrement lives here
// Step g_addr one page down and erase that page (the decrement lives here -
// the global-register rule).
[[gnu::noinline, gnu::noclone]] void erase_below()
{
g_addr -= page;
spm::erase_page<off>(g_addr);
spm::command<off>(spm::op::erase, g_addr);
settle();
}
@@ -221,7 +233,7 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
}
// Stream one host page into the erased flash page at g_addr (asm 2 of 2): the
// word pair stages in r0:r1 straight from rx (whose register set is known
// word pair stages in r0:r1 straight from rx (whose register set is known -
// the cross-call liveness C++ cannot express), Z walks the page and PGWRT
// programs it. g_addr is left at the next page base.
[[gnu::noinline, gnu::noclone]] void store_flash()
@@ -256,14 +268,15 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
{
// A watchdog reset hands straight back to the application, as the
// reference loader does, rather than re-entering the bootloader.
if (hw::mcusr::wdrf.test())
if (hw::mcusr::wdrf.test()) {
appjump();
}
// Lean bring-up from reset state: UCSR0C already reads 8N1, UBRR0H reads
// 0, and rx()/tx() raise RXEN0/TXEN0 on first use only the divisor low
// 0, and rx()/tx() raise RXEN0/TXEN0 on first use - only the divisor low
// byte and U2X0 need a store. The library still does the datasheet work.
static_assert(baud.u2x && baud.ubrr < 256, "lean bring-up writes UBRR0L only, with U2X0");
hw::reg<"UBRR0">::write(static_cast<std::uint8_t>(baud.ubrr));
hw::ubrr0::write(static_cast<std::uint8_t>(baud.ubrr));
hw::ucsr0a::write(hw::ucsr0a::u2x0(1));
// General-purpose registers are undefined at power-on (no crt zeroes them);
// the direction latch must start "not receiving" so the first rx() enables
@@ -271,13 +284,15 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
// same reason.
g_receiving = 0;
// Activation: 3×'@', each inside the config page's timeout window (rx
// floors it so a corrupt page cannot lock the loader out); anything else
// including silence hands over.
// Activation: 3x'@', each inside the config page's timeout window (rx
// floors it so a corrupt page cannot lock the loader out); anything else -
// including silence - hands over.
g_window = avr::flash_load(flash_ptr(app_end + 2));
for (std::uint8_t k = 3; k; --k)
if (rx() != knock)
for (std::uint8_t k = 3; k; --k) {
if (rx() != knock) {
appjump();
}
}
g_window = comm_window;
// Password gate (config page from app_end+3, 0xff-terminated; a blank
@@ -291,17 +306,19 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
std::uint8_t expected = avr::flash_load(flash_ptr(g_addr)) & mask;
++g_addr;
if (expected == 0xff) {
g_addr = reinterpret_cast<std::uint16_t>(&info[0]);
g_addr = reinterpret_cast<std::uint16_t>(info.data());
g_cnt = sizeof(info);
sendf();
break;
}
std::uint8_t got = rx();
if (got == 0) {
if (mask == 0)
if (mask == 0) {
continue;
if (rcnf() != confirm || rcnf() != confirm)
}
if (rcnf() != confirm || rcnf() != confirm) {
appjump();
}
erase_application(); // leaves g_addr = 0 for the EEPROM walk
do {
eewr(0xff);
@@ -310,9 +327,10 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
erase_below();
break;
}
if (got != expected)
if (got != expected) {
mask = 0;
}
}
for (;;) {
tx(confirm); // Mainloop ready
@@ -320,23 +338,27 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
switch (rx()) {
case 'f': // read application flash, one page per host '!'
for (;;) {
if (rx() != confirm)
if (rx() != confirm) {
break;
}
g_cnt = page;
sendf();
if (g_addr >= app_end)
if (g_addr >= app_end) {
break;
}
}
break;
case 'F': // erase the application, then take pages behind '?'
erase_application(); // leaves g_addr = 0, the write start
while (rcnf() == confirm)
while (rcnf() == confirm) {
store_flash();
}
break;
case 'e': // read EEPROM, one page per host '!', until the host stops
for (;;) {
if (rx() != confirm)
if (rx() != confirm) {
break;
}
g_cnt = page;
do {
tx(eerd());
@@ -358,8 +380,9 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
sendf();
break;
case 'C': // replace the config page, then echo it back to verify
if (rcnf() != confirm)
if (rcnf() != confirm) {
break;
}
g_addr = app_end + page;
erase_below(); // leaves g_addr = app_end, the store target
store_flash();
@@ -373,14 +396,6 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
} // namespace
} // namespace tsb
// Reset lands here: BOOTRST vectors to the boot section base and .vectors is
// laid first, so this is the first instruction executed. No crt ran, so set
// the stack pointer before anything is called.
extern "C" [[gnu::naked, gnu::used, gnu::section(".vectors")]] void __boot_entry()
{
SP = RAMEND;
// The one line of crt this loader needs: compiled code assumes
// __zero_reg__ (r1) is 0, and power-on registers are undefined.
asm volatile("clr __zero_reg__");
tsb::run();
}
// Reset lands at the boot section base (BOOTRST): the entry stub in .vectors
// is laid first and does the one line of crt a crt-less image needs.
template struct avr::startup::entry<tsb::run, avr::startup::stack::hardware>;

329
tsb/tsb_policy.cpp Normal file
View File

@@ -0,0 +1,329 @@
// TinySafeBoot on libavr - the policy floor: pureboot's rules, measured.
//
// The full TinySafeBoot feature set - watchdog bail, one-wire half-duplex,
// config-page activation timeout, password gate, emergency erase, and
// config/flash/EEPROM read-write - under philosophy #5 exactly as pureboot
// obeys it: no assembly, no register variables; code, attributes, and flags
// only. Every lesson pureboot's development produced is applied - the
// library's half-duplex serial and startup entry, lean bring-up from reset
// state, one merged send loop over both memories, oracle-shaped loop bounds,
// locals threaded through noinline primitives, pureboot's codegen flags -
// and the result sits below the idiomatic tier and above the 512 B boot
// section the tricks/asm tiers reach with the banned mechanisms
// (oracle/README.md holds all four). This tier exists to keep that gap an
// artifact
// rather than a claim: the gap to 512 is the rent of policy-clean C++ -
// helpers that hold a cursor across rx()/tx() pay push/pop and argument
// threading where a global-register protocol pays nothing, and both
// control-flow merges tried (a parametrized paged session, a merged store
// loop) measured larger than the split cases they replaced. TSB's wire fixes
// the per-command loop shapes on the device, so pureboot 5's one-transfer-
// loop collapse has no purchase here.
//
// The wire protocol is strict request/response, which is what makes the
// shared line safe: the device drives it only between a received command and
// its reply, and releases it (the library's half-duplex choreography)
// whenever it waits.
#include <libavr/libavr.hpp>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
using dev = avr::device<{.clock = 16_MHz}>;
// One-wire: RX and TX share the line, exactly as the native-UART TSB expects.
// 115200 at 16 MHz lands +2.1 % off, past the receiver-tolerance table the
// solver holds rates to - the oracle's own deployment has run there for a
// decade, so the override states that it is meant.
using serial_t = dev::uart0<{
.baud = 115200_Bd,
.allow_baud_error = true,
.half_duplex = true,
}>;
inline constexpr serial_t serial{};
namespace tsb {
namespace {
// The loader is purely polled - it never enables interrupts - so every SPM and
// EEPROM lock folds to nothing under this posture.
constexpr auto off = avr::irq::guard_policy::unused;
// Strict request/response: every SPM operation is waited out before the next
// byte moves, so no flash operation is ever in flight at an EEPROM access -
// the write procedure's step 2 has nothing to guard, the omission the
// datasheet grants (DS40002061B section 8.6.3).
constexpr auto no_spm = ee::spm_interlock::omitted;
// The handshake bytes, identical across every TSB host.
constexpr std::uint8_t confirm = '!';
constexpr std::uint8_t request = '?';
constexpr std::uint8_t knock = '@';
// Boot geometry for the 1 KB boot section (BOOTSZ=10); the page size and the
// flash/EEPROM extents are the chip database's to know. app_end is the config
// page (TSB's LASTPAGE), one page below the boot section.
constexpr std::uint16_t page = spm::page_bytes;
constexpr std::uint16_t boot_bytes = 1024;
constexpr std::uint16_t app_end = spm::flash_bytes - boot_bytes - page;
constexpr std::uint16_t eeprom_end = avr::hw::db.mem.eeprom_size - 1;
// Lockout-proof floor for the activation window: the oracle's F_CPU/1MHz, so
// it follows the clock rather than restating it (rule 41).
constexpr auto act_min = static_cast<std::uint8_t>(dev::clock.hz / 1'000'000);
// Post-activation window: the host gets seconds, not milliseconds, mid-session.
constexpr std::uint8_t comm_window = 200;
// Firmware version stamp: YY*512 + MM*32 + DD, the encoding the host decodes.
constexpr std::uint16_t build_date = 26 * 512 + 7 * 32 + 27;
// The 16-byte device-info block, streamed out on activation.
// clang-format off
[[gnu::progmem]] constexpr auto info = std::to_array<std::uint8_t>({
'T', 'S', 'B',
build_date & 0xFF, build_date >> 8,
0xF3, // status: native-UART fixed-baud lineage
avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2],
page / 2, // page size in words
(app_end / 2) & 0xFF, (app_end / 2) >> 8, // app-flash boundary, words
eeprom_end & 0xFF, eeprom_end >> 8,
0xAA, 0xAA, // ATmega processor-type marker (bytes 14 == 15)
});
// clang-format on
// The receive window, pre-floored where it is set. In .noinit: there is no
// crt to clear a .bss image, and run() stores it before the first receive.
[[gnu::section(".noinit")]] std::uint8_t window;
const std::uint8_t *flash_ptr(std::uint16_t addr)
{
return reinterpret_cast<const std::uint8_t *>(addr);
}
// Bounded byte receive: poll under nested countdowns, 0 on silence. The 0
// then falls through every compare - not a knock, not a confirm, not a
// command - so a silent host unwinds the loader to the application from
// anywhere, and a mid-session cable pull cannot wedge it. The line release on
// a direction change is the serial backend's.
[[gnu::noinline]] std::uint8_t rx()
{
std::uint16_t outer = static_cast<std::uint16_t>(window) << 8;
do {
std::uint8_t fine = 0;
do {
if (auto byte = serial.read()) {
return *byte;
}
} while (--fine);
} while (--outer);
return 0;
}
// One-wire transmit: the backend takes the line with a turn-around guard and
// holds it until the whole frame is out.
[[gnu::noinline]] void tx(std::uint8_t byte)
{
serial.write(byte);
}
// '?', then hand back the host's reply for the callers' one-byte compare.
[[gnu::noinline]] std::uint8_t rcnf()
{
tx(request);
return rx();
}
// The one send loop: the info block, the config page, application flash and
// EEPROM pages all stream through here.
[[gnu::noinline]] void send_block(bool eep, std::uint16_t at, std::uint8_t count)
{
do {
tx(eep ? ee::read<no_spm>(at) : avr::flash_load(flash_ptr(at)));
++at;
} while (--count);
}
// One EEPROM byte in - shared by the emergency wipe and the 'E' stream.
[[gnu::noinline]] void eeput(std::uint16_t at, std::uint8_t value)
{
ee::write<off, no_spm>(at, value);
}
// Wait out a running SPM op, then re-open the RWW section - after every page
// op and before handing over, as the oracle does.
[[gnu::noinline]] void settle()
{
spm::wait();
spm::rww_enable<off>();
}
// One host page straight into the erased flash page at `at` - through the SPM
// word buffer (low byte then high), no SRAM staging - then committed. `at`
// names a page base, so the cursor's low byte reaching the boundary ends the
// walk.
[[gnu::noinline]] void store_flash_page(std::uint16_t at)
{
const auto open = spm::page::begin<spm::from::boot_section, off>(at);
do {
std::uint8_t low = rx();
std::uint8_t high = rx();
spm::fill<off>(open, at, std::bit_cast<std::uint16_t>(std::array{low, high}));
at += 2;
} while (static_cast<std::uint8_t>(at) & (page - 1));
spm::command<off>(spm::op::write, at - page);
settle();
}
extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --defsym=tsb_app=0
[[noreturn]] void appjump()
{
settle();
tsb_app();
}
// Step one page down and erase it - the erase shared by the whole-app walk,
// the config rewrite and the emergency wipe; hands the stepped address back.
[[gnu::noinline]] std::uint16_t erase_below(std::uint16_t at)
{
at -= page;
spm::command<off>(spm::op::erase, at);
settle();
return at;
}
// Erase the whole application, top-down like the oracle: the loop bound is a
// compare with zero, and the returned 0 is the address every caller wants
// next.
[[gnu::noinline]] std::uint16_t erase_application()
{
std::uint16_t at = app_end;
do {
at = erase_below(at);
} while (at != 0);
return at;
}
[[noreturn]] void run()
{
// A watchdog reset hands straight back to the application, as the
// reference loader does, rather than re-entering the bootloader.
if (avr::hw::mcusr::wdrf.test()) {
appjump();
}
// Lean bring-up from reset state: UCSR0C already reads 8N1, UBRR0H reads
// 0, and the half-duplex write()/read() raise TXEN0/RXEN0 on first use -
// only the divisor low byte and U2X0 need a store. The solver still does
// the datasheet work; the asserts pin the reset-state assumptions.
{
constexpr auto sol = avr::uart::solve_baud(dev::clock, 115200_Bd, 8, avr::uart::parity::none);
static_assert(sol.u2x && sol.ubrr < 256, "lean bring-up writes UBRR0L only, with U2X0");
avr::hw::ubrr0::write(static_cast<std::uint8_t>(sol.ubrr));
avr::hw::ucsr0a::write(avr::hw::ucsr0a::u2x0(1));
}
// Activation: 3x'@', each inside the config page's timeout window
// (floored so a corrupt page cannot lock the loader out); anything else -
// including silence - hands over.
window = avr::flash_load(flash_ptr(app_end + 2)) | act_min;
for (std::uint8_t k = 3; k; --k) {
if (rx() != knock) {
appjump();
}
}
window = comm_window;
// Password gate (config page from app_end+3, 0xff-terminated; a blank
// page is no password). A wrong byte blanks the comparison and drains the
// line forever, so a wrong password can never fall through; a 0 requests
// emergency erase behind two confirms. On pass the info block goes out;
// the emergency path skips it and drops into the command loop.
std::uint16_t at = app_end + 3;
std::uint8_t mask = 0xff;
for (;;) {
std::uint8_t expected = avr::flash_load(flash_ptr(at)) & mask;
++at;
if (expected == 0xff) {
send_block(false, reinterpret_cast<std::uint16_t>(info.data()), info.size());
break;
}
std::uint8_t got = rx();
if (got == 0) {
if (mask == 0) {
continue;
}
if (rcnf() != confirm || rcnf() != confirm) {
appjump();
}
std::uint16_t a = erase_application();
do {
eeput(a, 0xff);
} while (++a <= eeprom_end);
erase_below(app_end + page);
break;
}
if (got != expected) {
mask = 0;
}
}
for (;;) {
tx(confirm); // Mainloop ready
const std::uint8_t command = rx();
switch (command) {
case 'f': // read application flash, one page per host '!'
for (std::uint16_t a = 0; a < app_end; a += page) {
if (rx() != confirm) {
break;
}
send_block(false, a, page);
}
break;
case 'e': // read EEPROM, one page per host '!', until the host stops
for (std::uint16_t a = 0;; a += page) {
if (rx() != confirm) {
break;
}
send_block(true, a, page);
}
break;
case 'F': { // erase the application, then take pages behind '?'
std::uint16_t a = erase_application();
for (; rcnf() == confirm; a += page) {
store_flash_page(a);
}
break;
}
case 'E': // take EEPROM pages behind '?', each write host-paced
for (std::uint16_t a = 0; rcnf() == confirm;) {
std::uint8_t count = page;
do {
eeput(a, rx());
++a;
} while (--count);
}
break;
case 'c': // read the config page
read_config:
send_block(false, app_end, page);
break;
case 'C': // replace the config page, then echo it back to verify
if (rcnf() != confirm) {
break;
}
store_flash_page(erase_below(app_end + page));
goto read_config;
default: // 'q' or any other byte runs the application
appjump();
}
}
}
} // namespace
} // namespace tsb
// Reset lands at the boot section base (BOOTRST): the entry stub in .vectors
// is laid first and does the one line of crt a crt-less image needs.
template struct avr::startup::entry<tsb::run, avr::startup::stack::hardware>;

View File

@@ -1,10 +1,10 @@
// TinySafeBoot on libavr tier 1: pure, idiomatic C++.
// TinySafeBoot on libavr - tier 1: pure, idiomatic C++.
//
// A serial flash bootloader for the ATmega328P boot section, reimplementing the
// TinySafeBoot native-UART fixed-baud protocol on libavr with the full feature
// set of the hand-written oracle: a watchdog-reset bail, one-wire half-duplex,
// a config-page activation timeout, the password gate, emergency erase, and
// config/flash/EEPROM read-write. This variant is written for clarity
// config/flash/EEPROM read-write. This variant is written for clarity -
// well-factored functions, no compiler-specific size hacks, no inline assembly.
// The one-wire wiring, the flash-resident info block and every SPM/EEPROM lock
// are libavr's to handle; the only attribute is the naked reset entry that
@@ -20,16 +20,29 @@ namespace ee = avr::eeprom;
using dev = avr::device<{.clock = 16_MHz}>;
// One-wire: RX and TX share the line, exactly as the native-UART TSB expects.
using serial_t = dev::uart0<{.baud = 115200_Bd, .max_baud_error = 3_pct, .half_duplex = true}>;
// 115200 at 16 MHz lands +2.1 % off, past the receiver-tolerance table the
// solver holds rates to - the oracle's own deployment has run there for a
// decade, so the override states that it is meant.
using serial_t = dev::uart0<{
.baud = 115200_Bd,
.allow_baud_error = true,
.half_duplex = true,
}>;
inline constexpr serial_t serial{};
namespace tsb {
namespace {
// The loader is purely polled it never enables interrupts so every SPM and
// The loader is purely polled - it never enables interrupts - so every SPM and
// EEPROM lock folds to nothing under this posture.
constexpr auto off = avr::irq::guard_policy::unused;
// Strict request/response: every SPM operation is waited out before the next
// byte moves, so no flash operation is ever in flight at an EEPROM access -
// the write procedure's step 2 has nothing to guard, the omission the
// datasheet grants (DS40002061B section 8.6.3).
constexpr auto no_spm = ee::spm_interlock::omitted;
// The handshake bytes, identical across every TSB host.
constexpr std::uint8_t confirm = '!';
constexpr std::uint8_t request = '?';
@@ -50,26 +63,48 @@ constexpr std::uint16_t build_date = 26 * 512 + 7 * 32 + 20;
// The 16-byte device-info block the host reads on activation. A flash_table
// keeps it in progmem with no .data image (there is no crt to copy one).
// clang-format off
inline constexpr std::array<std::uint8_t, 16> info_data = {
inline constexpr auto info_data = std::to_array<std::uint8_t>({
'T', 'S', 'B',
build_date & 0xFF, build_date >> 8,
0xF3, // status byte (native-UART fixed-baud lineage)
0x1E, 0x95, 0x0F, // ATmega328P signature
avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2],
page / 2, // page size in words
(app_end / 2) & 0xFF, (app_end / 2) >> 8, // app-flash boundary, words
eeprom_end & 0xFF, eeprom_end >> 8,
0xAA, 0xAA, // ATmega processor-type marker (bytes 14 == 15)
};
});
// clang-format on
using info = avr::flash_table<info_data>;
// Blocking byte read/write over the one-wire line: read() releases the line to
// the receiver, write() takes it and holds it until the frame is out.
// The lockout-proof floor for the receive window: the oracle's F_CPU/1MHz, so
// it follows the clock rather than restating it.
constexpr auto act_min = static_cast<std::uint8_t>(dev::clock.hz / 1'000'000);
// The receive window, pre-floored where it is set. In .noinit: there is no crt
// to clear a .bss image, and run() stores it before the first receive.
[[gnu::section(".noinit")]] std::uint8_t window;
// Bounded byte read over the one-wire line - read() releases the line to the
// receiver - answering 0 on silence. That 0 falls through every compare below:
// not a knock, not a confirm, not a command, so a silent host unwinds the
// loader to the application from anywhere and a mid-session cable pull cannot
// wedge it. The oracle lists that timeout among its own fixes, and a blocking
// read is how a tier loses it.
std::uint8_t rx()
{
return serial.read_blocking();
std::uint16_t outer = static_cast<std::uint16_t>(window) << 8;
do {
std::uint8_t fine = 0;
do {
if (auto byte = serial.read()) {
return *byte;
}
} while (--fine);
} while (--outer);
return 0;
}
// write() takes the line and holds it until the frame is out.
void tx(std::uint8_t byte)
{
serial.write(byte);
@@ -83,14 +118,16 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
// Stream `count` bytes to the host, from flash (LPM) or from EEPROM.
void send_flash(std::uint16_t addr, std::uint8_t count)
{
while (count--)
while (count--) {
tx(avr::flash_load(flash_ptr(addr++)));
}
}
void send_eeprom(std::uint16_t addr, std::uint8_t count)
{
while (count--)
tx(ee::read(addr++));
while (count--) {
tx(ee::read<no_spm>(addr++));
}
}
// Prompt the host with '?' and report whether it answered '!'.
@@ -101,32 +138,33 @@ bool request_confirm()
}
// Stream one page from the host straight into the already-erased flash page at
// `addr`, filling the SPM word buffer low byte then high no SRAM staging, so
// `addr`, filling the SPM word buffer low byte then high - no SRAM staging, so
// receiving and programming are the same loop.
void store_flash_page(std::uint16_t addr)
{
const auto open = spm::page::begin<spm::from::boot_section, off>(addr);
for (std::uint16_t i = 0; i < page; i += 2) {
std::uint8_t lo = rx();
std::uint8_t hi = rx();
spm::fill<off>(addr + i, static_cast<std::uint16_t>(lo | (hi << 8)));
spm::fill<off>(open, addr + i, static_cast<std::uint16_t>(lo | (hi << 8)));
}
spm::write_page<off>(addr);
spm::wait();
spm::write_page<spm::from::boot_section, off>(addr); // blocking: waits the write out
}
// Stream one page from the host straight into EEPROM, byte by byte.
void store_eeprom_page(std::uint16_t addr)
{
for (std::uint16_t i = 0; i < page; ++i)
ee::write<off>(addr + i, rx());
for (std::uint16_t i = 0; i < page; ++i) {
ee::write<off, no_spm>(addr + i, rx());
}
}
// Erase one flash page and wait it out the erase step shared by the whole-app
// erase, the config-page rewrite and the emergency wipe.
// Erase one flash page, waited out by the blocking spelling - the erase step
// shared by the whole-app erase, the config-page rewrite and the emergency
// wipe.
void erase_page(std::uint16_t addr)
{
spm::erase_page<off>(addr);
spm::wait();
spm::erase_page<spm::from::boot_section, off>(addr);
}
// Erase the whole application, one page at a time, top-down as the reference
@@ -157,8 +195,9 @@ extern "C" [[noreturn]] void tsb_app();
void read_flash()
{
for (std::uint16_t a = 0; a < app_end; a += page) {
if (rx() != confirm)
if (rx() != confirm) {
return;
}
send_flash(a, page);
}
}
@@ -167,8 +206,9 @@ void read_flash()
void read_eeprom()
{
for (std::uint16_t a = 0;; a += page) {
if (rx() != confirm)
if (rx() != confirm) {
return;
}
send_eeprom(a, page);
}
}
@@ -178,22 +218,25 @@ void read_eeprom()
void write_flash()
{
erase_application();
for (std::uint16_t a = 0; request_confirm(); a += page)
for (std::uint16_t a = 0; request_confirm(); a += page) {
store_flash_page(a);
}
}
// 'E': take pages the host offers behind '?' into EEPROM.
void write_eeprom()
{
for (std::uint16_t a = 0; request_confirm(); a += page)
for (std::uint16_t a = 0; request_confirm(); a += page) {
store_eeprom_page(a);
}
}
// 'C': replace the config page, then echo it back for the host to verify.
void write_config()
{
if (!request_confirm())
if (!request_confirm()) {
return;
}
erase_page(app_end);
store_flash_page(app_end);
spm::rww_enable<off>();
@@ -206,8 +249,9 @@ void write_config()
void emergency_erase()
{
erase_application();
for (std::uint16_t a = 0; a <= eeprom_end; ++a)
ee::write<off>(a, 0xff);
for (std::uint16_t a = 0; a <= eeprom_end; ++a) {
ee::write<off, no_spm>(a, 0xff);
}
erase_page(app_end);
spm::rww_enable<off>();
}
@@ -222,45 +266,54 @@ gate password_gate()
{
for (const std::uint8_t *pw = flash_ptr(app_end + 3);; ++pw) {
std::uint8_t expected = avr::flash_load(pw);
if (expected == 0xff)
if (expected == 0xff) {
return gate::pass;
}
std::uint8_t got = rx();
if (got == 0)
if (got == 0) {
return gate::emergency;
if (got != expected)
for (;;)
}
if (got != expected) {
for (;;) {
rx();
}
}
}
}
[[noreturn]] void run()
{
// A watchdog reset hands straight back to the application, as the reference
// loader does, rather than re-entering the bootloader.
if (avr::hw::mcusr::wdrf.test())
if (avr::hw::mcusr::wdrf.test()) {
appjump();
}
avr::init<serial_t>();
// Activation: the host knocks three '@' inside a window whose length is the
// config page's timeout byte (floored so a corrupt page can never lock the
// loader out). An idle port times out and boots the application.
__uint24 idle = static_cast<__uint24>(avr::flash_load(flash_ptr(app_end + 2)) | 16) << 16;
// config page's timeout byte, floored so a corrupt page can never lock the
// loader out. An idle port times out and boots the application; the same
// window then bounds every receive of the session.
window = avr::flash_load(flash_ptr(app_end + 2)) | act_min;
__uint24 idle = static_cast<__uint24>(window) << 16;
std::uint8_t knocks = 0;
while (knocks < 3) {
if (auto byte = serial.read())
if (auto byte = serial.read()) {
knocks = *byte == knock ? knocks + 1 : 0;
else if (--idle == 0)
} else if (--idle == 0) {
appjump();
}
}
switch (password_gate()) {
case gate::pass:
send_flash(reinterpret_cast<std::uint16_t>(info::storage.data()), info::size());
break;
case gate::emergency:
if (!request_confirm() || !request_confirm())
if (!request_confirm() || !request_confirm()) {
appjump();
}
emergency_erase();
break;
}
@@ -295,14 +348,6 @@ gate password_gate()
} // namespace
} // namespace tsb
// Reset lands here: BOOTRST vectors to the boot section base and .vectors is
// laid first, so this is the first instruction executed. No crt ran, so set the
// stack pointer before anything is called.
extern "C" [[gnu::naked, gnu::used, gnu::section(".vectors")]] void __boot_entry()
{
SP = RAMEND;
// The one line of crt this loader needs: compiled code assumes
// __zero_reg__ (r1) is 0, and power-on registers are undefined.
asm volatile("clr __zero_reg__");
tsb::run();
}
// Reset lands at the boot section base (BOOTRST): the entry stub in .vectors
// is laid first and does the one line of crt a crt-less image needs.
template struct avr::startup::entry<tsb::run, avr::startup::stack::hardware>;

View File

@@ -1,16 +1,16 @@
// TinySafeBoot on libavr tier 2: C++ with compiler trickery, no assembly.
// TinySafeBoot on libavr - tier 2: C++ with compiler trickery, no assembly.
//
// The full TinySafeBoot feature set watchdog bail, one-wire half-duplex,
// The full TinySafeBoot feature set - watchdog bail, one-wire half-duplex,
// config-page activation timeout, password gate, emergency erase, and
// config/flash/EEPROM read-write in pure C++, 526 bytes: 14 over the 512-byte
// boot section the hand-written oracle fits, from 168 over at this tier's first
// floor. The structure mirrors the oracle's: a handful of tiny noinline
// config/flash/EEPROM read-write - in pure C++, a little over the 512-byte boot
// section the hand-written oracle fits (oracle/README.md holds what each tier
// measures). The structure mirrors the oracle's: a handful of tiny noinline
// primitives sharing one whole-loader register allocation, expressed as global
// register variables so no helper ever saves, spills, or reloads any of it.
//
// The register protocol (all call-saved, so calls preserve them by ABI):
// Y (r28:r29) g_addr the walked flash/EEPROM address adiw-able
// r16 g_cnt byte countdown of the running block ldi-able
// Y (r28:r29) g_addr the walked flash/EEPROM address - adiw-able
// r16 g_cnt byte countdown of the running block - ldi-able
// r7 g_window rx timeout, roughly 30 ms units at 16 MHz
// r6 g_receiving one-wire direction latch, cleared at bring-up
// (power-on registers are undefined)
@@ -18,10 +18,10 @@
// GCC 16.1 miscompiles stores into global register variables: an update whose
// remaining uses all hide inside callees is deleted whenever a CALL follows it
// before any jump/ret (the backend's liveness walk lumps fixed registers with
// call-clobbered ones minimal repro in libavr's
// local/scratch/probes/gcc-avr-globalreg-repro.cpp, lessons.md entry). Every
// call-clobbered ones - minimal repro in libavr's
// test/upstream/gcc-avr-globalreg-repro.cpp). Every
// g_* update below therefore sits where a *local* read or a jump/ret follows
// it the helpers advance g_addr immediately before returning, and rx()
// it - the helpers advance g_addr immediately before returning, and rx()
// re-floors the window on every call instead of storing the floored value
// once. The layout is load-bearing; do not "simplify" it.
//
@@ -41,10 +41,16 @@ namespace hw = avr::hw;
namespace tsb {
namespace {
// The loader is purely polled it never enables interrupts so every SPM and
// The loader is purely polled - it never enables interrupts - so every SPM and
// EEPROM lock folds to nothing under this posture.
constexpr auto off = avr::irq::guard_policy::unused;
// Strict request/response: every SPM operation is waited out before the next
// byte moves, so no flash operation is ever in flight at an EEPROM access -
// the write procedure's step 2 has nothing to guard, the omission the
// datasheet grants (DS40002061B section 8.6.3).
constexpr auto no_spm = ee::spm_interlock::omitted;
constexpr std::uint8_t confirm = '!';
constexpr std::uint8_t request = '?';
constexpr std::uint8_t knock = '@';
@@ -57,28 +63,34 @@ constexpr std::uint16_t boot_bytes = 1024;
constexpr std::uint16_t app_end = spm::flash_bytes - boot_bytes - page;
constexpr std::uint16_t eeprom_end = avr::hw::db.mem.eeprom_size - 1;
// Lockout-proof floor for the activation window (the oracle's F_CPU/1MHz).
constexpr std::uint8_t act_min = 16;
// Lockout-proof floor for the activation window: the oracle's F_CPU/1MHz, so
// it follows the clock rather than restating it (rule 41).
constexpr auto act_min = static_cast<std::uint8_t>((16_MHz).hz / 1'000'000);
// Post-activation window: the host gets seconds, not milliseconds, mid-session.
constexpr std::uint8_t comm_window = 200;
constexpr std::uint16_t build_date = 26 * 512 + 7 * 32 + 20;
// Fixed 115200 8N1; the library solves UBRR + U2X from clock and baud.
constexpr auto baud = avr::uart::detail::solve_baud(16_MHz, 115200_Bd);
constexpr auto baud = avr::uart::solve_baud(16_MHz, 115200_Bd, 8, avr::uart::parity::none);
// One bit time on the wire: the turn-around a shared-line peer needs to stop
// driving before this one starts. Derived from the solved rate, so it follows
// the link rather than a count measured against one.
constexpr auto guard_cycles = static_cast<std::uint32_t>((16_MHz).hz / baud.actual);
// The 16-byte device-info block, streamed out on activation.
// clang-format off
[[gnu::progmem]] constexpr std::uint8_t info[16] = {
[[gnu::progmem]] constexpr auto info = std::to_array<std::uint8_t>({
'T', 'S', 'B',
build_date & 0xFF, build_date >> 8,
0xF3, // status: native-UART fixed-baud lineage
0x1E, 0x95, 0x0F, // ATmega328P signature
avr::hw::db.signature[0], avr::hw::db.signature[1], avr::hw::db.signature[2],
page / 2, // page size in words
(app_end / 2) & 0xFF, (app_end / 2) >> 8,
eeprom_end & 0xFF, eeprom_end >> 8,
0xAA, 0xAA,
};
});
// clang-format on
register std::uint16_t g_addr asm("r28");
@@ -93,8 +105,8 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
// Bounded byte receive, the oracle's shape: release the one-wire line on a
// direction change, poll RXC0 under nested countdowns, 0 on silence. The 0
// then falls through every compare not a knock, not a confirm, not a
// command so a silent host unwinds the loader to the application from
// then falls through every compare - not a knock, not a confirm, not a
// command - so a silent host unwinds the loader to the application from
// anywhere, and a mid-session cable pull cannot wedge it.
[[gnu::noinline, gnu::noclone]] std::uint8_t rx()
{
@@ -102,24 +114,25 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
g_receiving = 1;
hw::ucsr0b::write(hw::ucsr0b::rxen0(1)); // RXEN0 alone: release and listen
}
// act_min ORs in here, per call, not once into g_window at setup the
// act_min ORs in here, per call, not once into g_window at setup - the
// one placement the global-register-store miscompile cannot delete.
std::uint16_t outer = static_cast<std::uint16_t>(g_window | act_min) << 8;
do {
std::uint8_t fine = 0;
do {
auto status = hw::ucsr0a::read();
if (status & hw::ucsr0a::rxc0(1).value)
if (status & hw::ucsr0a::rxc0(1).value) {
return hw::udr0::read();
}
} while (--fine);
} while (--outer);
return 0;
}
// One-wire transmit: take the line (TXEN0 alone the receiver must be off
// One-wire transmit: take the line (TXEN0 alone - the receiver must be off
// while driving) on a direction change, with a turn-around guard so a shorted
// peer can switch first; then hold the line until the whole frame is out
// (TXC0, not UDRE0 the stop bit must be on the wire before a caller may
// (TXC0, not UDRE0 - the stop bit must be on the wire before a caller may
// release the line), and W1C TXC0 by storing the sampled status back, which
// keeps U2X0.
[[gnu::noinline, gnu::noclone]] void tx(std::uint8_t byte)
@@ -127,8 +140,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
if (g_receiving) {
g_receiving = 0;
hw::ucsr0b::write(hw::ucsr0b::txen0(1));
for (std::uint8_t guard = 46; guard; --guard)
;
avr::delay::cycles<guard_cycles>();
}
hw::udr0::write(byte);
std::uint8_t status;
@@ -145,7 +157,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
return rx();
}
// One flash byte [g_addr++] (the advance right before ret see header).
// One flash byte <- [g_addr++] (the advance right before ret - see header).
[[gnu::noinline, gnu::noclone]] std::uint8_t sflash()
{
std::uint8_t byte = avr::flash_load(flash_ptr(g_addr));
@@ -153,18 +165,18 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
return byte;
}
// One EEPROM byte [g_addr++].
// One EEPROM byte <- [g_addr++].
[[gnu::noinline, gnu::noclone]] std::uint8_t eerd()
{
std::uint8_t byte = ee::read(g_addr);
std::uint8_t byte = ee::read<no_spm>(g_addr);
++g_addr;
return byte;
}
// One EEPROM byte [g_addr++].
// One EEPROM byte -> [g_addr++].
[[gnu::noinline, gnu::noclone]] void eewr(std::uint8_t byte)
{
ee::write<off>(g_addr, byte);
ee::write<off, no_spm>(g_addr, byte);
++g_addr;
}
@@ -176,7 +188,7 @@ const std::uint8_t *flash_ptr(std::uint16_t addr)
} while (--g_cnt);
}
// Wait out a running SPM op, then re-open the RWW section after every page
// Wait out a running SPM op, then re-open the RWW section - after every page
// op and before handing over, as the oracle does.
[[gnu::noinline, gnu::noclone]] void settle()
{
@@ -198,12 +210,12 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
[[gnu::noinline, gnu::noclone]] void erase_below()
{
g_addr -= page;
spm::erase_page<off>(g_addr);
spm::command<off>(spm::op::erase, g_addr);
settle();
}
// Erase the whole application, top-down like the oracle: the loop bound is a
// compare with zero, and g_addr = 0 the value every caller wants next is
// compare with zero, and g_addr = 0 - the value every caller wants next - is
// handed back for free.
[[gnu::noinline, gnu::noclone]] void erase_application()
{
@@ -214,18 +226,19 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
}
// Stream one host page into the erased flash page at g_addr (SPM word buffer,
// low byte then high) no SRAM staging, receive and program are one loop.
// low byte then high) - no SRAM staging, receive and program are one loop.
// g_addr is left at the next page base.
[[gnu::noinline, gnu::noclone]] void store_flash()
{
const auto open = spm::page::begin<spm::from::boot_section, off>(g_addr);
g_cnt = page / 2;
do {
std::uint16_t word = rx();
word |= static_cast<std::uint16_t>(rx()) << 8;
spm::fill<off>(g_addr, word);
spm::fill<off>(open, g_addr, word);
g_addr += 2;
} while (--g_cnt);
spm::write_page<off>(g_addr - page);
spm::command<off>(spm::op::write, g_addr - page);
settle();
}
@@ -233,14 +246,15 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
{
// A watchdog reset hands straight back to the application, as the
// reference loader does, rather than re-entering the bootloader.
if (hw::mcusr::wdrf.test())
if (hw::mcusr::wdrf.test()) {
appjump();
}
// Lean bring-up from reset state: UCSR0C already reads 8N1, UBRR0H reads
// 0, and rx()/tx() raise RXEN0/TXEN0 on first use only the divisor low
// 0, and rx()/tx() raise RXEN0/TXEN0 on first use - only the divisor low
// byte and U2X0 need a store. The library still does the datasheet work.
static_assert(baud.u2x && baud.ubrr < 256, "lean bring-up writes UBRR0L only, with U2X0");
hw::reg<"UBRR0">::write(static_cast<std::uint8_t>(baud.ubrr));
hw::ubrr0::write(static_cast<std::uint8_t>(baud.ubrr));
hw::ucsr0a::write(hw::ucsr0a::u2x0(1));
// General-purpose registers are undefined at power-on (no crt zeroes them);
// the direction latch must start "not receiving" so the first rx() enables
@@ -248,13 +262,15 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
// same reason.
g_receiving = 0;
// Activation: 3×'@', each inside the config page's timeout window (rx
// floors it so a corrupt page cannot lock the loader out); anything else
// including silence hands over.
// Activation: 3x'@', each inside the config page's timeout window (rx
// floors it so a corrupt page cannot lock the loader out); anything else -
// including silence - hands over.
g_window = avr::flash_load(flash_ptr(app_end + 2));
for (std::uint8_t k = 3; k; --k)
if (rx() != knock)
for (std::uint8_t k = 3; k; --k) {
if (rx() != knock) {
appjump();
}
}
g_window = comm_window;
// Password gate (config page from app_end+3, 0xff-terminated; a blank
@@ -268,17 +284,19 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
std::uint8_t expected = avr::flash_load(flash_ptr(g_addr)) & mask;
++g_addr;
if (expected == 0xff) {
g_addr = reinterpret_cast<std::uint16_t>(&info[0]);
g_addr = reinterpret_cast<std::uint16_t>(info.data());
g_cnt = sizeof(info);
sendf();
break;
}
std::uint8_t got = rx();
if (got == 0) {
if (mask == 0)
if (mask == 0) {
continue;
if (rcnf() != confirm || rcnf() != confirm)
}
if (rcnf() != confirm || rcnf() != confirm) {
appjump();
}
erase_application(); // leaves g_addr = 0 for the EEPROM walk
do {
eewr(0xff);
@@ -287,9 +305,10 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
erase_below();
break;
}
if (got != expected)
if (got != expected) {
mask = 0;
}
}
for (;;) {
tx(confirm); // Mainloop ready
@@ -297,23 +316,27 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
switch (rx()) {
case 'f': // read application flash, one page per host '!'
for (;;) {
if (rx() != confirm)
if (rx() != confirm) {
break;
}
g_cnt = page;
sendf();
if (g_addr >= app_end)
if (g_addr >= app_end) {
break;
}
}
break;
case 'F': // erase the application, then take pages behind '?'
erase_application(); // leaves g_addr = 0, the write start
while (rcnf() == confirm)
while (rcnf() == confirm) {
store_flash();
}
break;
case 'e': // read EEPROM, one page per host '!', until the host stops
for (;;) {
if (rx() != confirm)
if (rx() != confirm) {
break;
}
g_cnt = page;
do {
tx(eerd());
@@ -335,8 +358,9 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
sendf();
break;
case 'C': // replace the config page, then echo it back to verify
if (rcnf() != confirm)
if (rcnf() != confirm) {
break;
}
g_addr = app_end + page;
erase_below(); // leaves g_addr = app_end, the store target
store_flash();
@@ -350,14 +374,6 @@ extern "C" [[noreturn]] void tsb_app(); // the application's reset vector: --def
} // namespace
} // namespace tsb
// Reset lands here: BOOTRST vectors to the boot section base and .vectors is
// laid first, so this is the first instruction executed. No crt ran, so set
// the stack pointer before anything is called.
extern "C" [[gnu::naked, gnu::used, gnu::section(".vectors")]] void __boot_entry()
{
SP = RAMEND;
// The one line of crt this loader needs: compiled code assumes
// __zero_reg__ (r1) is 0, and power-on registers are undefined.
asm volatile("clr __zero_reg__");
tsb::run();
}
// Reset lands at the boot section base (BOOTRST): the entry stub in .vectors
// is laid first and does the one line of crt a crt-less image needs.
template struct avr::startup::entry<tsb::run, avr::startup::stack::hardware>;