18 Commits

Author SHA1 Message Date
7d103ca957 pureboot: review-pass fixes to the host tool and device runner
pureboot.py: reject an empty image file with a clear error instead of
an IndexError deep in the vector-surgery planner; tighten the erase
docstring (order is irrelevant there — every target byte is the same
value, unlike a real flash where page 0 must go last).

pureboot_device.c: the GPIO bridge's bit_cycles used plain truncating
division where the firmware computes its own bit period with
round-to-nearest (uart.hpp: (Clock.hz + Baud.bd/2)/Baud.bd) — one
cycle off per bit on both tinies, harmless in practice but needless
drift against a firmware built to a different constant. Matched
exactly. Also clear the queued-bytes/decode-in-progress bridge state
on the test-only reset signal, so a future reset-mid-transfer scenario
can't feed a freshly reset chip bytes queued for its previous life.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AReSwkWkPX2A9Ym6grxRAh
2026-07-20 10:45:19 +02:00
eca7a41051 pureboot: gitignore python bytecode cache
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AReSwkWkPX2A9Ym6grxRAh
2026-07-20 10:43:56 +02:00
7314f7ab3b pureboot: stop tracking the python bytecode cache
A stray __pycache__/*.pyc from a local test run got swept into the
previous commit's git add. Untracked and gitignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AReSwkWkPX2A9Ym6grxRAh
2026-07-20 10:43:36 +02:00
5b361904ab pureboot: host tool and end-to-end protocol tests, all three chips
pureboot.py (Python stdlib only): images as raw binary or Intel HEX,
flash and EEPROM programming with read-back verify, erase composites,
fuse and info readout, activation-timeout configuration, and the
tinies' reset-vector surgery — the trampoline word below the loader,
page 0 written last.

The test spawns a simavr device (pureboot_device.c) — the mega's USART
as a pty; on the tinies a cycle-timed GPIO<->pty bridge for the polled
software UART plus the NVM module simavr's tiny cores lack (their SPM
opcode ioctls into a void and silently does nothing) — and drives it
with the real tool: knock from reset (erased-flash walk on the tinies),
program and verify both memories, timeout write, session reconnect, an
external reset through the patched vector, hand-over, and the fixture
application's banner. Results are cross-checked against ground-truth
memory dumps and an independent decode of the surgery's rjmp words,
red-verified against a sabotaged encoder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AReSwkWkPX2A9Ym6grxRAh
2026-07-20 05:54:15 +02:00
833e134e01 pureboot: the device — one pure C++ source, 512 bytes, every chip
No inline assembly, no global register variables; libavr does the
datasheet work. The device speaks primitives — flash read/page-program,
EEPROM read/write, fuse read, info block, EEPROM-resident activation
timeout, hand-over — and verify, erase, reset-vector surgery, and
timeout configuration live in the host tool. 490 B on the ATtiny13A,
510 B on the ATtiny85, 484 B on the ATmega328P, each linked into the
top 512 bytes of flash; per-chip size tests gate all three.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AReSwkWkPX2A9Ym6grxRAh
2026-07-20 05:33:28 +02:00
da730b7bb5 tsb: third size pass — restructure to the oracle's shape
The second pass concluded the 168 B tricks->asm gap was per-call ABI
cost. Most of it was structure. Rebuilt around the oracle's own shape —
argless noinline primitives over a whole-loader call-saved register
protocol (g_addr in Y, count r16, window r7, direction latch r6), a
top-down erase_below whose loop tests against zero and hands callers
g_addr = 0 for free, bounded rx everywhere (a silent host unwinds to
the app from any state, as the oracle does), and a named tsb_app entry
that --pmem-wrap-around=32k relaxes to the wrapped rjmp:

  tsb_asm    510 B in the 512 B section (oracle: 500), C++ except rx
             and the page-store loop — the two routines whose remaining
             cost is the calling convention itself (~30 asm lines, was
             ~280)
  tsb_tricks 526 B, no assembly at all (was 666)
  tsb_pure   836 B, still one readable function per command (was 842)

Every g_* update placement works around a GCC 16.1 wrong-code bug
(stores into global register variables deleted when only callees read
them — repro and rules in libavr dev/lessons.md). Also fixes two
latent hardware bugs all earlier tiers carried, masked by simavr's
zeroed register file: the crt-less entries never established
__zero_reg__ = 0, and the direction latch was read before written —
power-on registers are undefined.

All tiers full oracle feature parity, protocol tests green in both
libavr modes, .text byte-identical across modes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JYufebsiWvGkAJ2fLAB1gT
2026-07-20 01:00:27 +02:00
c351bee257 tsb: beat the first-pass size floors (tricks 666, pure 842)
tricks 778->666: always_inline every single-call handler into the
[[noreturn]] reset entry (which pays no prologue, so their push/pop of
call-saved registers vanishes), walk the page pointer in Y (adiw, base
recovered as g_addr-page) instead of recomputing Z=base+offset, bring
the UART up in the two registers that are not already at their reset
value, and seed the activation counter as __uint24.

pure 896->842: TU-local internal linkage (proper hygiene, and it lets
the compiler inline the one-call handlers), a byte-wide activation
count, __uint24 timeout. Still one readable function per command.

asm unchanged at 498: its C++-expressible parts are already C++; the
core stays asm (the 666 B all-tricks tier is 168 B over — per-call ABI
tax, not a feature). All three cross-mode byte-identical, protocol green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 20:15:49 +02:00
34e7f1be34 tsb: drive each tier to its size floor
asm 502->498 B (below the oracle's 500): the stack bring-up moves to plain C++,
and a register is reserved for the config-page high byte instead of reloading it
at each app-flash-boundary compare. tricks 808->778 B: shared erase/rww helpers
plus the libavr half-duplex W1C fix. pure 950->896 B and no SRAM: streams
rx->SPM/EEPROM instead of staging a 128 B page buffer. All three keep full oracle
feature parity and stay byte-identical across modes; protocol tests (round-trip +
password + emergency erase) green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 18:47:38 +02:00
445e187722 tsb: document the three tiers at full parity in the build file
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 16:50:18 +02:00
250aba5cfb tsb: protocol test covers the password gate and emergency erase
Each scenario group now runs on its own freshly-reset device: the round-trip
on a blank config page, plus a password-config device that must be sent the
password after the knock to activate, and an emergency-erase device where a
0-byte + two confirms wipes flash, EEPROM and the config page (verified by
reading all three back as 0xff). All three tiers pass every group.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 16:49:20 +02:00
5c900720e3 tsb: pure and tricks tiers reach full oracle feature parity
Both tiers gain the features the asm tier already carries — one-wire
half-duplex (via libavr's new .half_duplex), the config-page activation
timeout, and emergency erase (password \0 + double-confirm wipes flash,
EEPROM and the config page) — on top of the watchdog bail, password gate and
config/flash/EEPROM read-write they already had. pure stays idiomatic
(flash_table info block, one function per command) at 950 B; tricks keeps its
compiler trickery (call-saved global-register page walk, unified runtime-flag
paths pinned noinline/noclone, streaming stores, arithmetic command decode)
at 808 B. Both byte-identical across generated and reflect modes; the size
gradient across the three tiers is now 502 / 808 / 950 B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 16:47:21 +02:00
7d6ef959b2 tsb: asm tier reaches full oracle feature parity at 502 B
Rewrite the inline-asm tier so it matches the hand-written fixed-baud oracle's
feature set inside the 512 B boot section: watchdog-reset bail, one-wire
half-duplex (RXEN/TXEN toggled per direction, TX turnaround guard),
config-page activation timeout, the password gate (wrong byte hangs draining
the UART), emergency erase (password \0 + double-confirm wipes flash, EEPROM
and the config page), and config/flash/EEPROM read-write. Every geometry,
baud and info-block constant comes from libavr consteval; only the dense
control flow is hand-written. 502 B, byte-identical across generated and
reflect modes.

Test harness: seed the config page from TSB_CONFIG so the password and
emergency-erase paths are exercisable, and clear simavr's AVR_UART_FLAG_POLL_
SLEEP — a host-CPU-saving usleep(1)-per-idle-poll hack that models no hardware
and paces a one-wire loader (which releases TX between bytes) in real time,
distorting protocol timing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 16:23:16 +02:00
11ffbce2e2 tsb: vendor the fixed-baud assembly oracle as the size/feature bar
The Seed Robotics native-UART fixed-baud TinySafeBoot (GPLv3), reference
only — not built. Assembles to 500 B with the full feature set, proving
≤512 B and full feature parity are simultaneously reachable. Also drops the
stale empty stk500v2/ leftover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHeP42XU3wf6RfhyuxBTE5
2026-07-19 15:29:58 +02:00
f32a27ff15 tsb: use the named register surface
Direct register access now reads through the named surface
(hw::mcusr::wdrf.test(), hw::ucsr0b::write(...)) instead of the string form,
matching how libavr itself is written. Zero-overhead: pure 740 B, tricks 658 B,
asm 508 B unchanged, all byte-identical across modes, protocol green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:55:01 +02:00
57d94cf631 tsb: refactor the pure tier onto libavr sugar
The showcase tier now leans on the helpers it fed back instead of reaching under
them: the info block is an avr::flash_table (no raw [[gnu::progmem]]), a page is
filled with spm::fill(addr, span) (no hand-packed lo|hi<<8 loop), and the
WDT-reset bail reads field<"MCUSR","WDRF">::test() (no read() & {}(1).value).

Zero-overhead throughout: .text stays 740 B, byte-identical across generated and
reflect modes, protocol test green. The info block streams through the existing
address-based send_flash rather than a range-for over the flash_table — the
range-for is a distinct loop that cannot share the loader's one flash streamer,
so it would add 14 B for no functional gain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:33:53 +02:00
8203a24f33 tsb: slim the port branch to the libavr reimplementation
main carried the whole pre-libavr tree beside the port: the other-bootloader
directories (blink, stk500v2), the Atmel Studio solution/project, and — dead in
the tsb dir itself — four submodule links to the superseded io/flash/uart/type
libraries the libavr sources never include. None are build inputs; CMake drives
the three variants through FetchContent. master keeps the full legacy tree
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:12:56 +02:00
2906da3272 tsb: drop the local -O3 strip, now handled by the libavr toolchain
The -O3 leak is fixed upstream (cmake/release-os.cmake via CMAKE_PROJECT_INCLUDE),
so the port no longer needs its own string(REPLACE); a Release build is -Os
through the toolchain file. Verified: all three variants build at their sizes
(508/658/740) and pass the size + protocol ctest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 10:52:13 +02:00
64c1e484b5 tsb: reimplement TinySafeBoot on libavr in three size tiers
The native-UART fixed-baud TinySafeBoot protocol, ported onto libavr as a
crt-free boot-section loader, in three variants that trade clarity for size:

  tsb_pure   740 B  idiomatic C++: SRAM page buffer, separate flash/EEPROM
                    leaves, shared framing; the polled `unused` guard posture.
  tsb_tricks 658 B  unified runtime-flag paths (noinline/noclone), call-saved
                    global-register page walk — attributes only, no asm.
  tsb_asm    508 B  streaming store + hand-rolled UART/SPM/EEPROM/erase loops;
                    fits the 512 B boot section (BOOTSZ=11). Trims the optional
                    password gate and WDT-reset bail — unreachable in C++ with
                    both (hand-asm is ~15 % denser). Tiers 1-2 keep them and
                    live in the 1 KB section they fit.

All three are .text byte-identical across libavr's generated and reflect modes.
The CMake build strips the leaked -O3 (a Release build is silently -O3, not the
-Os this loader is measured against) and gates each variant's size against its
section. A simavr harness (test/device.c + test/tsbtest.py) drives the real wire
protocol over a pty and flashes the device; the size and protocol tests run in
ctest. Verified byte-for-byte against the reference tsbloader_adv (C#/mono):
activate, read info, flash write + verify.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 05:00:51 +02:00
25 changed files with 590 additions and 6916 deletions

3
.gitmodules vendored
View File

@@ -1,3 +0,0 @@
[submodule "libavr"]
path = libavr
url = ../libavr.git

View File

@@ -8,9 +8,6 @@ include(FetchContent)
if(NOT LIBAVR_ROOT AND DEFINED ENV{LIBAVR_ROOT})
set(LIBAVR_ROOT $ENV{LIBAVR_ROOT})
endif()
if(NOT LIBAVR_ROOT)
set(LIBAVR_ROOT ${CMAKE_CURRENT_SOURCE_DIR}/libavr)
endif()
if(LIBAVR_ROOT)
FetchContent_Declare(libavr SOURCE_DIR ${LIBAVR_ROOT})
else()
@@ -54,20 +51,6 @@ if(PROJECT_IS_TOP_LEVEL)
endif()
endif()
# The ELF is only a container (symbols, section headers) and is never flashed —
# and the host tool's load_image() dispatches on extension, so handing it one
# would silently program the header bytes. Every loader image therefore gets
# both flashable forms beside it at link time: .hex for avrdude, and .bin for
# the host tool's raw path (which is what the reloc and update tests convert to
# on the fly). .eeprom is dropped — EEPROM content is its own update.
function(add_image_outputs name)
add_custom_command(TARGET ${name} POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O ihex -R .eeprom
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.hex
COMMAND ${CMAKE_OBJCOPY} -O binary -R .eeprom
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.bin)
endfunction()
# The TinySafeBoot protocol reimplemented on libavr in three variants that trade
# clarity for size. Each links into the ATmega328P boot section (BOOTSZ selects
# its size; BOOTRST vectors a reset to its base) with -nostartfiles — a polled
@@ -107,7 +90,6 @@ function(add_tsb_variant name bytes)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${base_hex}
-Wl,--defsym=tsb_app=0 -Wl,--pmem-wrap-around=32k)
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
add_image_outputs(${name})
if(PROJECT_IS_TOP_LEVEL)
add_test(NAME ${name}.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:${name}>
@@ -129,35 +111,52 @@ if(LIBAVR_MCU STREQUAL "atmega328p")
endif()
# pureboot — the pure-constraint port (see pureboot/README.md): one source,
# no inline assembly, no global register variables, every libavr chip,
# fitting each chip's smallest boot sector. The geometry and the
# pureboot_add_loader() deployment function live in pureboot/CMakeLists.txt —
# the unit a downstream project consumes; everything below is this port's
# own build: the stock loaders, their tests, and the size matrix. The
# distinct binary dir keeps the `pureboot` target's output name free.
add_subdirectory(pureboot pureboot-cmake)
# no inline assembly, no global register variables, every libavr chip, 512
# bytes each. The loader owns the top 512 bytes of flash on every chip; the
# application entry symbol is address 0 on the mega (reset re-vectors to the
# loader through BOOTRST, so word 0 stays the application's own vector) and
# the trampoline word just below the loader on the tinies (host-side vector
# surgery points it at the application). --pmem-wrap-around models AVR's
# modulo-flash PC where the flash is big enough to need it.
if(LIBAVR_MCU STREQUAL "attiny13a")
set(_pb_flash 1024)
set(_pb_wrap "")
set(_pb_page 32)
set(_pb_hz 9600000)
set(_pb_baud 57600)
set(_pb_eeprom 64)
elseif(LIBAVR_MCU STREQUAL "attiny85")
set(_pb_flash 8192)
set(_pb_wrap -Wl,--pmem-wrap-around=8k)
set(_pb_page 64)
set(_pb_hz 8000000)
set(_pb_baud 57600)
set(_pb_eeprom 512)
else()
set(_pb_flash 32768)
set(_pb_wrap -Wl,--pmem-wrap-around=32k)
set(_pb_page 128)
set(_pb_hz 16000000)
set(_pb_baud 115200)
set(_pb_eeprom 1024)
endif()
math(EXPR _pb_base "${_pb_flash} - 512")
math(EXPR _pb_base_hex "${_pb_base}" OUTPUT_FORMAT HEXADECIMAL)
if(LIBAVR_MCU STREQUAL "atmega328p")
set(_pb_app 0)
else()
math(EXPR _pb_app "${_pb_base} - 2")
endif()
# The stock loader: the family-default deployment (crystal/RC clock, the
# chip's natural link, default pins). The activation window stays a cache
# variable — re-timing a deployed loader is a self-update with a re-timed
# build. pureboot9 is that re-timed build, and what the update test installs.
set(PUREBOOT_TIMEOUT 8 CACHE STRING "pureboot activation window, seconds")
pureboot_add_loader(pureboot TIMEOUT ${PUREBOOT_TIMEOUT})
add_executable(pureboot pureboot/pureboot.cpp)
target_link_libraries(pureboot PRIVATE libavr)
target_link_options(pureboot PRIVATE -nostartfiles -Wl,--section-start=.text=${_pb_base_hex}
-Wl,--defsym=pureboot_app=${_pb_app} ${_pb_wrap})
add_custom_command(TARGET pureboot POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:pureboot>)
if(PROJECT_IS_TOP_LEVEL)
get_target_property(_pb_stock_hz pureboot PUREBOOT_HZ)
get_target_property(_pb_stock_baud pureboot PUREBOOT_BAUD)
add_test(NAME pureboot.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:pureboot>
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
if(Python3_FOUND)
add_test(NAME pureboot.pi
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/check_pi.py
${CMAKE_OBJDUMP} ${CMAKE_NM} $<TARGET_FILE:pureboot> ${PUREBOOT_BASE_HEX})
add_test(NAME pureboot.planner
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/test_planner.py
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py)
endif()
-DLIMIT=512 -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
# The protocol test flashes this fixture through the loader with the real
# host tool and expects its banner after the hand-over; a normally linked
@@ -169,257 +168,10 @@ if(PROJECT_IS_TOP_LEVEL)
COMMAND ${CMAKE_OBJCOPY} -O binary $<TARGET_FILE:pbapp> $<TARGET_FILE:pbapp>.bin)
add_test(NAME pureboot.protocol
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbtest.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud} ${PUREBOOT_EEPROM}
$<TARGET_FILE:pbapp>.bin ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${LIBAVR_MCU} ${_pb_hz} ${_pb_base_hex}
${_pb_page} ${_pb_baud} ${_pb_eeprom} $<TARGET_FILE:pbapp>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbtest-work)
set_tests_properties(pureboot.protocol PROPERTIES TIMEOUT 180)
# The position-independence acceptance test: the identical image,
# installed one slot lower, must serve the full command set.
add_test(NAME pureboot.reloc
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbreloc.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud}
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbreloc-work)
set_tests_properties(pureboot.reloc PROPERTIES TIMEOUT 180
ENVIRONMENT "PB_OBJCOPY=${CMAKE_OBJCOPY}")
# Entering the loader from a running application with no reset
# between, over a page buffer the application dirtied — the case the
# loader declines to guard and the host repairs. Hardware forbids the
# state here (SPM runs only from the boot section); simavr does not,
# which is what makes it constructible.
if(LIBAVR_MCU STREQUAL "atmega328p")
add_test(NAME pureboot.dirty
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbdirty.py
${PB_DEVICE} $<TARGET_FILE:pureboot> ${PUREBOOT_SIM_MCU} ${_pb_stock_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_pb_stock_baud}
$<TARGET_FILE:pbapp>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbdirty-work)
set_tests_properties(pureboot.dirty PROPERTIES TIMEOUT 180)
endif()
# Re-homing: a loader mistakenly programmed at address 0 (a raw .bin
# handed to a programmer) or sitting in the staging slot must heal
# into the canonical slot through the ordinary --update-loader flow.
# Patched-vector behavior, so one representative chip carries it.
if(LIBAVR_MCU STREQUAL "attiny85")
add_test(NAME pureboot.rehome
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbrehome.py
${PB_DEVICE} $<TARGET_FILE:pureboot> $<TARGET_FILE:pureboot9>.bin
${PUREBOOT_SIM_MCU} ${_pb_stock_hz} ${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE}
${_pb_stock_baud} $<TARGET_FILE:pbapp>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbrehome-work)
set_tests_properties(pureboot.rehome PROPERTIES TIMEOUT 180)
endif()
# The self-update end-to-end: the re-timed build (same source, only
# the timeout differs — a byte-different image) replaces the resident
# through --update-loader, with every power-fail phase rehearsed from
# the runner's flash dumps.
pureboot_add_loader(pureboot9 TIMEOUT 9)
add_test(NAME pureboot.update
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbupdate.py
${PB_DEVICE} $<TARGET_FILE:pureboot> $<TARGET_FILE:pureboot9>
${PUREBOOT_SIM_MCU} ${_pb_stock_hz} ${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE}
${_pb_stock_baud} $<TARGET_FILE:pbapp>.bin
${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbupdate-work)
set_tests_properties(pureboot.update PROPERTIES TIMEOUT 600
ENVIRONMENT "PB_OBJCOPY=${CMAKE_OBJCOPY}")
endif()
# The size matrix: every configuration axis that could move the image
# size — the serial backend (different code), the USART instance
# (different registers), the clock (different constants), and the baud
# through the shapes its bit timing takes — each combination must still
# fit the chip's slot budget. Pins are size-neutral (port and bit are
# immediate operands) and the timeout is a constant, so neither adds an
# axis. The stock build is one point of this matrix and already has its
# test.
function(pureboot_size_variant name)
pureboot_add_loader(${name} ${ARGN})
add_test(NAME ${name}.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:${name}>
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
# The autobaud loader (pureboot/autobaud.md): one clock-agnostic image per
# chip, no clock × baud axis, size-tested against the same per-chip budget on
# every chip.
#
# Only the unified version is built. Its two predecessors —
# pureboot_autobaud_pure.cpp at 508 B and pureboot_autobaud_reg.cpp at 512 —
# had 4 B and 0 B of margin on the 1284P, and the fix for the activation hang
# costs ~22, which puts them at 530 and 534. Neither can ship, so the choice
# the branch existed to offer is settled by measurement rather than taste.
# The sources stay for the record; autobaud.md carries the numbers.
function(pureboot_autobaud_variant name source)
pureboot_add_autobaud(${name} ${source})
add_test(NAME ${name}.size
COMMAND ${CMAKE_COMMAND} -DSIZE_TOOL=${CMAKE_SIZE} -DELF=$<TARGET_FILE:${name}>
-DLIMIT=${PUREBOOT_LIMIT} -P ${CMAKE_CURRENT_SOURCE_DIR}/test/check_size.cmake)
endfunction()
pureboot_autobaud_variant(pureboot_autobaud_uni pureboot_autobaud_uni.cpp)
# One point of the exhaustive matrix, named from its resolved parameters
# so the enumeration cannot collide with itself. Unreachable rates drop
# out here rather than aborting the configure.
function(pureboot_matrix_point hz baud link)
if(link STREQUAL "software")
pureboot_baud_feasible(${hz} ${baud} 1 _ok)
set(_args SERIAL software)
else()
pureboot_baud_feasible(${hz} ${baud} 0 _ok)
set(_args USART ${link})
endif()
if(_ok)
pureboot_size_variant(pbm_${hz}_${baud}_${link} CLOCK ${hz} BAUD ${baud} ${_args})
endif()
endfunction()
# Clock points: the shipped-fuse floor (CKDIV8), the calibrated RC, and
# the crystal the stock build assumes (the tiny13's ladder is its own RC
# menu — it has no crystal option).
if(LIBAVR_MCU MATCHES "^attiny13")
set(_matrix_clocks 1200000 4800000 9600000)
set(_full_clocks 128000 600000 1200000 4800000 9600000)
else()
set(_matrix_clocks 1000000 8000000 16000000)
set(_full_clocks 128000 1000000 1843200 2000000 3686400 4000000 7372800 8000000
11059200 12000000 14745600 16000000 18432000 20000000)
endif()
# The exhaustive cross product: every clock a deployment plausibly runs
# — the internal oscillators, the shipped CKDIV8 floor, the plain
# crystals and the UART crystals — against every rate, against every
# backend. Beyond the ladder the list carries the slow rates a
# sub-megahertz oscillator is left with, which no ladder rate reaches
# (16000 Bd is the only rate the 128 kHz oscillator holds exactly); at
# the fast clocks those same rates also select the software UART's
# 16-bit _delay_loop_2 bit spin (two words more setup at each of its five
# sites), the largest image the space produces and a shape the ladder
# default — always the *fastest* rate a clock reaches — never picks.
#
# Bounded to one chip per size-bearing class: flash addressing (the
# word-addressed 1284), hand-over shape (the patched vector on the tinies
# and m48s), page size, and USART inventory. Everything else in the image
# is chip-independent code, so a further chip buys builds and no
# coverage; every chip outside the set carries the compact matrix.
get_property(_full_bauds GLOBAL PROPERTY PUREBOOT_BAUD_LADDER)
list(APPEND _full_bauds 16000 4800 2400 1200)
set(_matrix_spot attiny13a attiny85 atmega48pa atmega8a atmega168pa
atmega328p atmega164a atmega644a atmega1284p)
if(DEFINED ENV{PUREBOOT_FULL_MATRIX} AND LIBAVR_MCU IN_LIST _matrix_spot)
foreach(_matrix_hz IN LISTS _full_clocks)
foreach(_matrix_baud IN LISTS _full_bauds)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} software)
if(PUREBOOT_HAS_USART)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 0)
endif()
if(PUREBOOT_HAS_USART1)
pureboot_matrix_point(${_matrix_hz} ${_matrix_baud} 1)
endif()
endforeach()
endforeach()
else()
foreach(_matrix_hz IN LISTS _matrix_clocks)
math(EXPR _matrix_khz "${_matrix_hz} / 1000")
if(PUREBOOT_HAS_USART OR NOT _matrix_hz EQUAL _pb_stock_hz)
pureboot_size_variant(pureboot_sw_${_matrix_khz}k CLOCK ${_matrix_hz} SERIAL software)
endif()
if(PUREBOOT_HAS_USART AND NOT _matrix_hz EQUAL _pb_stock_hz)
pureboot_size_variant(pureboot_hw_${_matrix_khz}k CLOCK ${_matrix_hz} SERIAL hardware)
endif()
if(PUREBOOT_HAS_USART1 AND NOT _matrix_hz EQUAL _pb_stock_hz)
pureboot_size_variant(pureboot_usart1_${_matrix_khz}k CLOCK ${_matrix_hz} USART 1)
endif()
endforeach()
list(GET _matrix_clocks -1 _matrix_top_hz)
pureboot_size_variant(pureboot_sw_wide CLOCK ${_matrix_top_hz} BAUD 9600 SERIAL software)
endif()
if(PUREBOOT_HAS_USART1)
pureboot_size_variant(pureboot_usart1 USART 1)
endif()
# One configured deployment end to end — a real board's shape rather
# than the stock assumption: the ATmega328P on its shipped 1 MHz fuses,
# the software UART on hand-picked pins (TX = PB1, RX = PB5), the ladder
# baud (9600). The full protocol suite runs against it, fixture
# application included, over the runner's GPIO bridge — proving the
# configuration plumbing produces a working loader, not just one that
# fits.
if(LIBAVR_MCU STREQUAL "atmega328p" AND DEFINED PB_DEVICE)
pureboot_size_variant(pureboot_custom CLOCK 1000000 SERIAL software RX pb5 TX pb1)
get_target_property(_custom_hz pureboot_custom PUREBOOT_HZ)
get_target_property(_custom_baud pureboot_custom PUREBOOT_BAUD)
get_target_property(_custom_link pureboot_custom PUREBOOT_LINK)
add_executable(pbapp_custom test/pbapp.cpp)
target_link_libraries(pbapp_custom PRIVATE libavr)
target_compile_definitions(pbapp_custom PRIVATE PUREBOOT_CLOCK_HZ=${_custom_hz}
PUREBOOT_BAUD=${_custom_baud} PUREBOOT_SOFT_SERIAL PUREBOOT_TX=pb1)
add_custom_command(TARGET pbapp_custom POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_custom> $<TARGET_FILE:pbapp_custom>.bin)
add_test(NAME pureboot.custom
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbtest.py
${PB_DEVICE} $<TARGET_FILE:pureboot_custom> ${PUREBOOT_SIM_MCU} ${_custom_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_custom_baud} ${PUREBOOT_EEPROM}
$<TARGET_FILE:pbapp_custom>.bin ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbcustom-work ${_custom_link})
set_tests_properties(pureboot.custom PROPERTIES TIMEOUT 180)
endif()
# The second USART, driven for real on one chip: instance selection is
# compile-checked everywhere, but only a live session proves the loader
# initialized and polls the USART it claims to. The fixture application
# banners on the same instance.
if(LIBAVR_MCU STREQUAL "atmega644a" AND DEFINED PB_DEVICE)
get_target_property(_usart1_hz pureboot_usart1 PUREBOOT_HZ)
get_target_property(_usart1_baud pureboot_usart1 PUREBOOT_BAUD)
add_executable(pbapp_usart1 test/pbapp.cpp)
target_link_libraries(pbapp_usart1 PRIVATE libavr)
target_compile_definitions(pbapp_usart1 PRIVATE PUREBOOT_CLOCK_HZ=${_usart1_hz}
PUREBOOT_BAUD=${_usart1_baud} PUREBOOT_USART=1)
add_custom_command(TARGET pbapp_usart1 POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_usart1> $<TARGET_FILE:pbapp_usart1>.bin)
add_test(NAME pureboot.usart1
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbtest.py
${PB_DEVICE} $<TARGET_FILE:pureboot_usart1> ${PUREBOOT_SIM_MCU} ${_usart1_hz}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} ${_usart1_baud} ${PUREBOOT_EEPROM}
$<TARGET_FILE:pbapp_usart1>.bin ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbusart1-work usart1)
set_tests_properties(pureboot.usart1 PROPERTIES TIMEOUT 180)
endif()
# The autobaud variants driven end to end over the software-UART bridge (both
# under review — pureboot/autobaud.md): the host sends the 0xC0 calibration
# pulse, the loader times it, locks, and programs. Run on the near-flash 328P
# and the word-addressed 1284P — the two flash-addressing classes — and each
# at two clocks with the one binary, which is the clock-agnostic property
# autobaud exists for (test/pbautobaud.py). The fixture application banners
# over the same software link at the first clock's rate.
if(LIBAVR_MCU MATCHES "^atmega(328p|1284p)$" AND DEFINED PB_DEVICE)
add_executable(pbapp_autobaud test/pbapp.cpp)
target_link_libraries(pbapp_autobaud PRIVATE libavr)
target_compile_definitions(pbapp_autobaud PRIVATE PUREBOOT_CLOCK_HZ=1000000
PUREBOOT_BAUD=9600 PUREBOOT_SOFT_SERIAL PUREBOOT_TX=pb1)
add_custom_command(TARGET pbapp_autobaud POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O binary
$<TARGET_FILE:pbapp_autobaud> $<TARGET_FILE:pbapp_autobaud>.bin)
foreach(_variant uni)
add_test(NAME pureboot.autobaud_${_variant}
COMMAND ${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/test/pbautobaud.py
${PB_DEVICE} $<TARGET_FILE:pureboot_autobaud_${_variant}> ${PUREBOOT_SIM_MCU}
${PUREBOOT_BASE_HEX} ${PUREBOOT_PAGE} $<TARGET_FILE:pbapp_autobaud>.bin
1000000 9600 ${CMAKE_CURRENT_SOURCE_DIR}/pureboot/pureboot.py
${CMAKE_BINARY_DIR}/pbautobaud-${_variant}-work)
set_tests_properties(pureboot.autobaud_${_variant} PROPERTIES TIMEOUT 240)
endforeach()
endif()
endif()

File diff suppressed because it is too large Load Diff

1
libavr

Submodule libavr deleted from edc77ca43f

View File

@@ -1,347 +0,0 @@
# pureboot as a consumable CMake unit: the per-chip geometry, the default baud
# ladder, and pureboot_add_loader() — the one way a loader target is created.
# A downstream project brings its usual libavr setup (the `libavr` target and
# the LIBAVR_MCU toolchain preset), adds this directory, and states its
# deployment; every argument is optional (README.md):
#
# add_subdirectory(bootloader/pureboot)
# pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5)
# Per-family geometry, deployment defaults, and the linker wrap the PC modulo
# needs. The slot is 512 bytes on every chip. The USART flags mirror the
# hardware inventory the loader's own static asserts check — the plain 644 is
# the x4 family's one single-USART die (Atmel-2593).
set(_pb_has_usart 1)
set(_pb_has_usart1 0)
if(LIBAVR_MCU MATCHES "^attiny13a?$")
set(_pb_flash 1024)
set(_pb_wrap "")
set(_pb_page 32)
set(_pb_hz 9600000)
set(_pb_eeprom 64)
set(_pb_has_usart 0)
elseif(LIBAVR_MCU STREQUAL "attiny25")
set(_pb_flash 2048)
set(_pb_wrap "")
set(_pb_page 32)
set(_pb_hz 8000000)
set(_pb_eeprom 128)
set(_pb_has_usart 0)
elseif(LIBAVR_MCU STREQUAL "attiny45")
set(_pb_flash 4096)
set(_pb_wrap "")
set(_pb_page 64)
set(_pb_hz 8000000)
set(_pb_eeprom 256)
set(_pb_has_usart 0)
elseif(LIBAVR_MCU STREQUAL "attiny85")
set(_pb_flash 8192)
set(_pb_wrap -Wl,--pmem-wrap-around=8k)
set(_pb_page 64)
set(_pb_hz 8000000)
set(_pb_eeprom 512)
set(_pb_has_usart 0)
elseif(LIBAVR_MCU MATCHES "^atmega48(a|p|pa)?$")
set(_pb_flash 4096)
set(_pb_wrap "")
set(_pb_page 64)
set(_pb_hz 16000000)
set(_pb_eeprom 256)
elseif(LIBAVR_MCU MATCHES "^atmega8a?$" OR LIBAVR_MCU MATCHES "^atmega88(a|p|pa)?$")
set(_pb_flash 8192)
set(_pb_wrap -Wl,--pmem-wrap-around=8k)
set(_pb_page 64)
set(_pb_hz 16000000)
set(_pb_eeprom 512)
elseif(LIBAVR_MCU MATCHES "^atmega16a?$" OR LIBAVR_MCU MATCHES "^atmega168(a|p|pa)?$")
set(_pb_flash 16384)
set(_pb_wrap -Wl,--pmem-wrap-around=16k)
set(_pb_page 128)
set(_pb_hz 16000000)
set(_pb_eeprom 512)
elseif(LIBAVR_MCU MATCHES "^atmega164(a|p|pa)$")
set(_pb_flash 16384)
set(_pb_wrap -Wl,--pmem-wrap-around=16k)
set(_pb_page 128)
set(_pb_hz 16000000)
set(_pb_eeprom 512)
set(_pb_has_usart1 1)
elseif(LIBAVR_MCU MATCHES "^atmega32a?$" OR LIBAVR_MCU MATCHES "^atmega328p?$")
set(_pb_flash 32768)
set(_pb_wrap -Wl,--pmem-wrap-around=32k)
set(_pb_page 128)
set(_pb_hz 16000000)
set(_pb_eeprom 1024)
elseif(LIBAVR_MCU MATCHES "^atmega324(a|p|pa)$")
set(_pb_flash 32768)
set(_pb_wrap -Wl,--pmem-wrap-around=32k)
set(_pb_page 128)
set(_pb_hz 16000000)
set(_pb_eeprom 1024)
set(_pb_has_usart1 1)
elseif(LIBAVR_MCU MATCHES "^atmega644(a|p|pa)?$")
# 64 KiB is exactly the 16-bit byte space, so plain LPM still reaches
# everything and the wire stays byte-addressed. The plain 644 is the
# family's one single-USART die.
set(_pb_flash 65536)
set(_pb_wrap -Wl,--pmem-wrap-around=64k)
set(_pb_page 256)
set(_pb_hz 16000000)
set(_pb_eeprom 2048)
if(NOT LIBAVR_MCU STREQUAL "atmega644")
set(_pb_has_usart1 1)
endif()
elseif(LIBAVR_MCU MATCHES "^atmega1284p?$")
# 128 KiB: wire addresses are words, reads go through ELPM, and the PC's
# modulo wrap exceeds what --pmem-wrap-around models.
set(_pb_flash 131072)
set(_pb_wrap "")
set(_pb_page 256)
set(_pb_hz 16000000)
set(_pb_eeprom 4096)
set(_pb_has_usart1 1)
else()
message(FATAL_ERROR "pureboot: no geometry for ${LIBAVR_MCU}")
endif()
set(_pb_slot 512)
math(EXPR _pb_base "${_pb_flash} - ${_pb_slot}")
math(EXPR _pb_base_hex "${_pb_base}" OUTPUT_FORMAT HEXADECIMAL)
# Patched-vector chips hand over through the trampoline word below the slot,
# which is also the slot's own last word — their budget is slot 2.
if(LIBAVR_MCU MATCHES "^atmega" AND NOT LIBAVR_MCU MATCHES "^atmega48")
set(_pb_app 0)
set(_pb_limit ${_pb_slot})
else()
math(EXPR _pb_app "${_pb_base} - 2")
math(EXPR _pb_limit "${_pb_slot} - 2")
endif()
# simavr names its cores after the base dies; the A revisions run on them
# (the 644PA on the 644P core).
set(_pb_sim_mcu ${LIBAVR_MCU})
if(LIBAVR_MCU MATCHES "^atmega(8|16|32|48|88|164|168|644)a$")
string(REGEX REPLACE "a$" "" _pb_sim_mcu ${LIBAVR_MCU})
elseif(LIBAVR_MCU STREQUAL "atmega644pa")
set(_pb_sim_mcu atmega644p)
endif()
# The function runs in its caller's scope, so everything it needs crosses
# scopes as global properties.
set_property(GLOBAL PROPERTY PUREBOOT_BASE_HEX ${_pb_base_hex})
set_property(GLOBAL PROPERTY PUREBOOT_APP ${_pb_app})
set_property(GLOBAL PROPERTY PUREBOOT_WRAP "${_pb_wrap}")
set_property(GLOBAL PROPERTY PUREBOOT_DEFAULT_HZ ${_pb_hz})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART ${_pb_has_usart})
set_property(GLOBAL PROPERTY PUREBOOT_HAS_USART1 ${_pb_has_usart1})
# The port's own build (tests, the size matrix) reads the geometry from the
# parent scope; a downstream consumer gets the same variables for free.
set(PUREBOOT_BASE_HEX ${_pb_base_hex} PARENT_SCOPE)
set(PUREBOOT_PAGE ${_pb_page} PARENT_SCOPE)
set(PUREBOOT_SLOT ${_pb_slot} PARENT_SCOPE)
set(PUREBOOT_LIMIT ${_pb_limit} PARENT_SCOPE)
set(PUREBOOT_EEPROM ${_pb_eeprom} PARENT_SCOPE)
set(PUREBOOT_DEFAULT_HZ ${_pb_hz} PARENT_SCOPE)
set(PUREBOOT_HAS_USART ${_pb_has_usart} PARENT_SCOPE)
set(PUREBOOT_HAS_USART1 ${_pb_has_usart1} PARENT_SCOPE)
set(PUREBOOT_SIM_MCU ${_pb_sim_mcu} PARENT_SCOPE)
# The rates a default may pick, fastest first.
set_property(GLOBAL PROPERTY PUREBOOT_BAUD_LADDER 115200 57600 38400 19200 9600)
# Whether <baud> is reachable from <clock> within 2.5 %, by the same
# best-of-U2X-and-plain divisor search libavr's solve_baud runs, so a build
# never trips the compile-time error it is checked against. A software build
# also needs the polled receiver's 100-cycles-a-bit floor: at low clocks the
# U2X divisor reaches rates the bit-banged sampler cannot.
function(pureboot_baud_feasible clock baud software outvar)
set(${outvar} 0 PARENT_SCOPE)
math(EXPR _cycles "${clock} / ${baud}")
if(software AND _cycles LESS 100)
return()
endif()
foreach(divisor 8 16)
math(EXPR _step "${divisor} * ${baud}")
math(EXPR _n "(${clock} + ${_step} / 2) / ${_step}")
if(_n LESS 1 OR _n GREATER 4096)
continue()
endif()
math(EXPR _actual "${clock} / (${divisor} * ${_n})")
math(EXPR _delta "${_actual} - ${baud}")
if(_delta LESS 0)
math(EXPR _delta "-(${_delta})")
endif()
math(EXPR _error_bp "${_delta} * 10000 / ${baud}")
if(_error_bp LESS_EQUAL 250)
set(${outvar} 1 PARENT_SCOPE)
return()
endif()
endforeach()
endfunction()
# The fastest ladder rate the clock reaches.
function(pureboot_default_baud clock software outvar)
get_property(_ladder GLOBAL PROPERTY PUREBOOT_BAUD_LADDER)
foreach(baud ${_ladder})
pureboot_baud_feasible(${clock} ${baud} ${software} _ok)
if(_ok)
set(${outvar} ${baud} PARENT_SCOPE)
return()
endif()
endforeach()
message(FATAL_ERROR "pureboot: no standard baud rate fits a ${clock} Hz clock within 2.5 % "
"— pass BAUD <rate> to deploy a non-standard one")
endfunction()
# pureboot_add_loader(<name> [CLOCK <hz>] [BAUD <bd>]
# [SERIAL auto|hardware|software] [USART <n>]
# [RX <pin>] [TX <pin>] [TIMEOUT <s>])
#
# The loader target plus its flashable images (<name>.hex for a programmer,
# <name>.bin for --update-loader). The resolved deployment is stamped on the
# target as PUREBOOT_HZ / PUREBOOT_BAUD / PUREBOOT_LINK (the link spelled
# usart0, usart1 or sw:<RX>,<TX>) — what a test harness speaks to it with.
function(pureboot_add_loader name)
cmake_parse_arguments(PB "" "CLOCK;BAUD;SERIAL;USART;RX;TX;TIMEOUT" "" ${ARGN})
if(PB_UNPARSED_ARGUMENTS)
message(FATAL_ERROR "pureboot_add_loader(${name}): unknown arguments ${PB_UNPARSED_ARGUMENTS}")
endif()
get_property(_hz GLOBAL PROPERTY PUREBOOT_DEFAULT_HZ)
get_property(_base_hex GLOBAL PROPERTY PUREBOOT_BASE_HEX)
get_property(_app GLOBAL PROPERTY PUREBOOT_APP)
get_property(_wrap GLOBAL PROPERTY PUREBOOT_WRAP)
get_property(_usart GLOBAL PROPERTY PUREBOOT_HAS_USART)
get_property(_usart1 GLOBAL PROPERTY PUREBOOT_HAS_USART1)
if(NOT PB_CLOCK)
set(PB_CLOCK ${_hz})
endif()
if(NOT PB_TIMEOUT)
set(PB_TIMEOUT 8)
endif()
if(NOT PB_SERIAL)
set(PB_SERIAL auto)
endif()
if(DEFINED PB_USART AND PB_SERIAL STREQUAL "software")
message(FATAL_ERROR "pureboot_add_loader(${name}): USART ${PB_USART} contradicts SERIAL software")
endif()
if(DEFINED PB_USART)
set(PB_SERIAL hardware)
elseif(PB_SERIAL STREQUAL "hardware")
set(PB_USART 0)
endif()
set(_serial_defines "")
if(PB_SERIAL STREQUAL "hardware")
if(PB_USART EQUAL 1 AND NOT _usart1)
message(FATAL_ERROR "pureboot_add_loader(${name}): ${LIBAVR_MCU} has no USART1")
elseif(NOT _usart)
message(FATAL_ERROR "pureboot_add_loader(${name}): ${LIBAVR_MCU} has no hardware USART")
endif()
set(_serial_defines PUREBOOT_USART=${PB_USART})
set(_link usart${PB_USART})
else()
if(PB_SERIAL STREQUAL "auto")
if(_usart AND (PB_RX OR PB_TX))
message(WARNING "pureboot_add_loader(${name}): RX/TX apply to the software UART, "
"which auto does not pick on ${LIBAVR_MCU} — SERIAL software to force it")
endif()
if(_usart)
set(_link usart0)
else()
set(PB_SERIAL software)
endif()
endif()
if(PB_SERIAL STREQUAL "software")
if(NOT PB_RX)
set(PB_RX pb0)
endif()
if(NOT PB_TX)
set(PB_TX pb1)
endif()
foreach(_pin ${PB_RX} ${PB_TX})
if(NOT _pin MATCHES "^p[a-h][0-7]$")
message(FATAL_ERROR "pureboot_add_loader(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
set(_serial_defines PUREBOOT_SOFT_SERIAL PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
# sw:<RX>,<TX> as port letter and bit, upcased.
string(SUBSTRING ${PB_RX} 1 2 _rx_pin)
string(SUBSTRING ${PB_TX} 1 2 _tx_pin)
string(TOUPPER "sw:${_rx_pin},${_tx_pin}" _link)
string(REPLACE "SW" "sw" _link ${_link})
endif()
endif()
if(NOT PB_BAUD)
if(PB_SERIAL STREQUAL "software")
pureboot_default_baud(${PB_CLOCK} 1 PB_BAUD)
else()
pureboot_default_baud(${PB_CLOCK} 0 PB_BAUD)
endif()
endif()
set(_defines PUREBOOT_CLOCK_HZ=${PB_CLOCK} PUREBOOT_BAUD=${PB_BAUD} PUREBOOT_TIMEOUT=${PB_TIMEOUT}
${_serial_defines})
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/pureboot.cpp)
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE ${_defines})
# Codegen shaping for the loader TU only, worth 1436 B depending on the
# chip. At -Os GCC otherwise rewrites the byte-stream loops' counters into
# end-pointer forms that cost registers (-fno-ivopts,
# -fno-split-wide-types), leaves register pressure on the table with the
# default allocator (-fira-algorithm=priority), and keeps loop-invariant
# immediates and expression temporaries in registers
# (-fno-move-loop-invariants, -fno-tree-ter) — but every loop body here
# contains a call, so a register held across it costs more than the
# load-immediate it saves.
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
# The ELF is a container, never flashed: .hex for a programmer, .bin (the
# slot's bare bytes) for --update-loader.
add_custom_command(TARGET ${name} POST_BUILD
COMMAND ${CMAKE_OBJCOPY} -O ihex -R .eeprom
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.hex
COMMAND ${CMAKE_OBJCOPY} -O binary -R .eeprom
$<TARGET_FILE:${name}> $<TARGET_FILE:${name}>.bin)
set_target_properties(${name} PROPERTIES PUREBOOT_HZ ${PB_CLOCK} PUREBOOT_BAUD ${PB_BAUD}
PUREBOOT_LINK ${_link})
endfunction()
# pureboot_add_autobaud(<name> <source> [RX <pin>] [TX <pin>])
#
# An autobaud software-serial loader from <source> (pureboot_autobaud_*.cpp).
# Autobaud measures the host's bit timing at runtime, so the image carries no
# clock and no baud — one binary per chip runs at any F_CPU. Same per-chip
# geometry, link and codegen flags as pureboot_add_loader(); only the clock and
# baud axes fall away. Two source files are under review (autobaud.md):
# pureboot_autobaud_pure.cpp and pureboot_autobaud_reg.cpp.
function(pureboot_add_autobaud name source)
cmake_parse_arguments(PB "" "RX;TX" "" ${ARGN})
if(NOT PB_RX)
set(PB_RX pb0)
endif()
if(NOT PB_TX)
set(PB_TX pb1)
endif()
foreach(_pin ${PB_RX} ${PB_TX})
if(NOT _pin MATCHES "^p[a-h][0-7]$")
message(FATAL_ERROR "pureboot_add_autobaud(${name}): pin '${_pin}' is not of the form pb1")
endif()
endforeach()
get_property(_base_hex GLOBAL PROPERTY PUREBOOT_BASE_HEX)
get_property(_app GLOBAL PROPERTY PUREBOOT_APP)
get_property(_wrap GLOBAL PROPERTY PUREBOOT_WRAP)
add_executable(${name} ${CMAKE_CURRENT_FUNCTION_LIST_DIR}/${source})
target_link_libraries(${name} PRIVATE libavr)
target_compile_definitions(${name} PRIVATE PUREBOOT_RX=${PB_RX} PUREBOOT_TX=${PB_TX})
target_compile_options(${name} PRIVATE
-fno-ivopts -fira-algorithm=priority -fno-move-loop-invariants -fno-tree-ter -fno-split-wide-types)
target_link_options(${name} PRIVATE -nostartfiles -Wl,--section-start=.text=${_base_hex}
-Wl,--defsym=pureboot_app=${_app} ${_wrap})
add_custom_command(TARGET ${name} POST_BUILD COMMAND ${CMAKE_SIZE} $<TARGET_FILE:${name}>)
endfunction()

View File

@@ -2,366 +2,122 @@
A serial bootloader on [libavr](https://git.blackmark.me/avr/libavr), pure by
constraint: one C++ source, no inline assembly, no global register variables
(attributes and compiler flags allowed), **512 bytes on every chip libavr
targets — all 37**. The device speaks primitives; every composite — verify,
erase, reset-vector surgery, updating the loader itself — lives in the host
tool (`pureboot.py`).
(attributes allowed), built for every chip libavr targets, **512 bytes on
each** — 490 B on the ATtiny13A, 510 B on the ATtiny85, 484 B on the
ATmega328P. The device speaks primitives; every composite — verify, erase,
reset-vector surgery, timeout configuration — lives in the host tool
(`pureboot.py`).
The image is **position-independent**: control flow is PC-relative, the
read/write paths take wire addresses, the write guard protects the slot the
code is *running* in (from the runtime return address), the info block is
addressed from that same anchor, and the application jump is an indirect call
to an absolute entry. The identical binary therefore runs from any slot with
every command intact, which makes pureboot **its own staging loader**: the
host installs the same binary one slot below the resident, jumps into it, and
lets it rewrite the resident.
## Link
## Chips
| Chip | Serial | Baud | Clock assumed |
|---|---|---|---|
| ATmega328P | USART0, RXD/TXD = PD0/PD1 | 115200 8N1 | 16 MHz crystal |
| ATtiny85 | software UART, RX = PB0, TX = PB1 | 57600 8N1 | 8 MHz internal RC |
| ATtiny13A | software UART, RX = PB0, TX = PB1 | 57600 8N1 | 9.6 MHz internal RC |
Sizes are the default configuration: the hardware USART0 at 115200 8N1 on a
16 MHz crystal, or the software UART on RX = PB0 / TX = PB1 at 57600 8N1 on
the tinies' RC oscillator (9.6 MHz on the t13s, 8 MHz above). Every axis moves
per build — see *Configuration*; the largest image any of them produces is a
software UART at a slow baud, which on the 1284s is 494 B, the tightest fit in
the whole matrix at 18 B spare.
| Chip | Flash | Loader at | Link | Size |
|---|---|---|---|---|
| ATtiny13, ATtiny13A † | 1 KiB | 0x0200 | software | 416 B |
| ATtiny25 † | 2 KiB | 0x0600 | software | 420 B |
| ATtiny45 † | 4 KiB | 0x0e00 | software | 424 B |
| ATtiny85 † | 8 KiB | 0x1e00 | software | 424 B |
| ATmega8, 8A | 8 KiB | 0x1e00 | USART0 | 396 B |
| ATmega16, 16A | 16 KiB | 0x3e00 | USART0 | 400 B |
| ATmega32, 32A | 32 KiB | 0x7e00 | USART0 | 400 B |
| ATmega48, 48A, 48P, 48PA † | 4 KiB | 0x0e00 | USART0 | 414 B |
| ATmega88, 88A, 88P, 88PA | 8 KiB | 0x1e00 | USART0 | 434 B |
| ATmega168, 168A, 168P, 168PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega328, 328P | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega164A, 164P, 164PA | 16 KiB | 0x3e00 | USART0 | 438 B |
| ATmega324A, 324P, 324PA | 32 KiB | 0x7e00 | USART0 | 438 B |
| ATmega644, 644A, 644P, 644PA | 64 KiB | 0xfe00 | USART0 | 432 B |
| ATmega1284, 1284P | 128 KiB | 0x1fe00 | USART0 | 478 B |
† No hardware boot section: the host patches the reset vector, and the budget
is 510 bytes, since the slot's last word is the trampoline.
The 1284s are the heaviest because they alone carry the far-flash machinery —
ELPM reads, RAMPZ page commands, a word-addressed wire.
The software UART enables the RX pull-up; TX idles high. All multi-byte wire
quantities are little-endian.
## Configuration
Every deployment axis is a build parameter of `pureboot_add_loader()` (in
`pureboot/CMakeLists.txt`) — the one way a loader target is created, by this
repo's build and by a downstream project alike:
| Argument | Meaning | Default |
|---|---|---|
| `CLOCK <hz>` | the clock the board runs | 16 MHz megas, 8 MHz t25/45/85, 9.6 MHz t13s |
| `BAUD <bd>` | the wire rate | the ladder below |
| `SERIAL auto\|hardware\|software` | the link backend | `auto`: the hardware USART where the chip has one |
| `USART <n>` | the USART instance (x4 megas carry two) | 0 |
| `RX <pin>`, `TX <pin>` | software-UART pins | `pb0`, `pb1` |
| `TIMEOUT <s>` | the activation window | 8 |
The default baud is the fastest of 115200/57600/38400/19200/9600 the clock
reaches within 2.5 % — the same U2X-included divisor search libavr's baud
solver runs — and on a software build additionally within the polled
receiver's 100-cycles-a-bit floor. Whatever is picked or overridden is
re-checked in the compile: an infeasible combination, or a USART the chip does
not have, fails with a named static assert.
A downstream project brings its usual libavr setup (the `libavr` target, the
chip via the `LIBAVR_MCU` toolchain preset), consumes this directory, and
states its deployment — an ATmega328P on its shipped 1 MHz fuses with the
software UART on hand-picked pins, say:
```cmake
FetchContent_Declare(bootloader GIT_REPOSITORY git@git.blackmark.me:avr/bootloader.git GIT_TAG main)
FetchContent_MakeAvailable(bootloader)
add_subdirectory(${bootloader_SOURCE_DIR}/pureboot pureboot)
pureboot_add_loader(myboot CLOCK 1000000 SERIAL software TX pb1 RX pb5)
```
The function emits the ELF plus `myboot.hex` (the programmer artifact) and
`myboot.bin` (the self-update image), prints the size, and stamps the resolved
deployment on the target as the `PUREBOOT_HZ`, `PUREBOOT_BAUD` and
`PUREBOOT_LINK` properties — what a flashing script or test harness needs to
speak to the build. This exact deployment runs the full protocol suite in CI
(`pureboot.custom`).
The tiny RX pin has its pull-up enabled; TX idles high. All multi-byte
quantities on the wire are little-endian.
## Activation
Reset enters the loader (BOOTRST on the boot-sectioned megas, the patched
reset vector elsewhere) — except a watchdog reset, which hands straight to the
application with no activation window, since the application owns its watchdog.
This is deliberate: it lets an application reboot itself instantly rather than
sit through the window. The application must clear WDRF itself (libavr's
`watchdog::disable()` does). **Gotcha:** WDRF is sticky (cleared only by
software, not by a later reset), so an application that watchdog-resets and
never clears it diverts *every* subsequent reset — external ones included —
past the window too, and the loader becomes reachable only through an external
programmer until the flag is cleared. A serial recovery path therefore assumes
the application clears WDRF on its own reset path.
Reset enters the loader (BOOTRST on the mega, the patched reset vector on the
tinies) — except a watchdog reset, which hands straight to the application
(the application owns its watchdog; it must clear WDRF itself, which also
releases the WDRF-forced WDE).
The host then knocks `p` then `b`, each awaited byte under a fresh activation
window; any other byte is discarded and awaited again, so line noise can delay
the loader but never lock it. A window expiring on an idle line boots the
application.
The host then has one activation window per awaited byte to knock: `p` then
`b`. Each awaited byte gets a fresh window; any other byte is discarded and
awaited again (line noise cannot lock the loader, only delay it). A window
expiring with an idle line boots the application.
The window is a compile-time constant (`TIMEOUT`, 8 s by default), so the whole
EEPROM belongs to the application — pureboot keeps no state of its own.
Re-timing a deployed loader is a self-update with a re-timed build.
The window length in seconds is the **last EEPROM cell** (address
`eeprom_size - 1`); `0x00` and the erased `0xff` both mean the 4 s default,
so a full EEPROM erase resets the timeout rather than maxing it. The host
changes it with the ordinary EEPROM-write command.
## Session
After the knock the loader stays in its command loop until `J` jumps away or
the chip resets. Before reading each command it waits for any pending EEPROM
write and sends the prompt `+` (0x2b), which is therefore also the previous
command's completion ack. A session is: await `+`, send a command, read its
reply, repeat.
On chips whose flash exceeds 64 KiB (the 1284s — info-block flag bit 1) the
`R`/`W` flash addresses are **word** addresses; everywhere else they are byte
addresses (the 644s' 64 KiB is exactly the 16-bit byte space). EEPROM
addresses and all counts are bytes.
The loader trusts the host to keep addresses in range: it does not bound them
against the info block. **Gotcha:** a `w` (or `r`) that runs past `E2END` wraps
— EEAR is only as wide as the array, so an address past the end truncates onto
low EEPROM and the write silently overwrites it. Keeping writes within the
advertised sizes is the host's job (the shipped tool does); the flash budget
is better spent on features than on re-checking a bound the host already holds.
After the knock the loader stays in its command loop until `G` or a reset.
Before reading each command it waits for any pending EEPROM write to finish
and sends the prompt `+` (0x2b) — the prompt is therefore also the completion
ack of the previous command. A session is: await `+`, send a command, read
its reply, repeat.
| Cmd | Arguments | Reply |
|---|---|---|
| `b` | — | the 12-byte info block |
| `R` | addr16, n8 | n flash bytes (n = 0 means 256) |
| `W` | addr16 (any address in the page), then one page of data | — (completion = next prompt) |
| `W` | addr16, then one page of data | — (completion = next prompt) |
| `r` | addr16, n8 | n EEPROM bytes (n = 0 means 256) |
| `w` | addr16, n8, then n data bytes | `+` per byte, sent once its write has begun |
| `F` | — | 4 bytes: low fuse, lock, extended fuse, high fuse |
| `J` | word address (16-bit) | `+`, then execution continues there |
| `G` | — | `+`, then the application runs |
| other | — | ignored; the loop re-prompts (send a junk byte, await `+`, to resync) |
`W` streams exactly one SPM page (size from the info block) into the buffer,
then erases and programs — except pages inside the 512-byte slot
the loader is *running* in, which are drained and left alone, so a broken host
cannot brick the running copy and a staged copy may rewrite the resident.
The loader never clears the SPM buffer before a fill, so **one `W` may program
the wrong bytes, and the host is what fixes it**. The buffer is write-once per
word until cleared, and two things leave words in it: a refused page, and —
where SPM runs from anywhere, the tinies and the m48s — an application that
self-programmed before entering. The next `W` takes those stale words and
clears them, since a page write auto-erases the buffer (§26.2.1; §19.2 on the
tinies), so repeating it programs correctly. The host therefore verifies every
page it writes and rewrites what comes back wrong (three retries, then it
stops).
`w` is host-paced: send the next byte only after the previous byte's `+`. `F`
returns the bytes in the hardware's Z order; on a chip without an extended
fuse byte that slot carries no meaning. Fuse *writing* does not exist — SPM
reaches flash and boot lock bits only.
`J` is the one control-transfer primitive: it runs the application (word 0 or
the trampoline word, both known from the info block) and moves between loader
copies during a self-update. A jump to a slot's base re-enters that copy's own
startup, which must then be knocked afresh.
then erases and programs; the address must be page-aligned. Pages inside the
loader's own 512 bytes are drained but never programmed — a broken host
cannot brick the chip. `w` is host-paced: send the next byte only after the
previous byte's `+`. `F` returns the bytes in the hardware's Z order; on a chip without an
extended fuse byte (the ATtiny13A) that slot carries no meaning. Fuse *writing* does not
exist: SPM reaches flash (and, on the mega, lock bits) only — fuse bytes are
external-programming territory by hardware.
The info block (`b`):
| Offset | Content |
|---|---|
| 02 | `'P'`, `'B'`, pureboot version (3) |
| 02 | `'P'`, `'B'`, protocol version (1) |
| 35 | device signature |
| 6 | SPM page size in bytes (0 means 256) |
| 78 | loader base — application flash ends here (a word address when bit 1 is set) |
| 6 | SPM page size in bytes |
| 78 | loader base — application flash ends here |
| 910 | EEPROM size |
| 11 | bit 0: host must patch the reset vector (no hardware boot section); bit 1: flash wire addresses are word addresses |
| 11 | bit 0 set: host must patch the reset vector (no hardware boot section) |
## Version
The info block's third byte is the **pureboot version** — the loader's one
identity number, and the only way to tell what a deployed loader is. Nothing
else is numbered: the wire protocol has no version, a pureboot version implies
it, and the host tool holds that map. The tool states the window of loader
versions it speaks (`OLDEST_LOADER`/`NEWEST_LOADER` in `pureboot.py`), and a
version that changes the protocol becomes the new floor there. None has so
far: 1 through 3 speak the identical session. A loader newer than the tool is
refused by name rather than decoded on the assumption that nothing moved.
The tool carries its own version, free to drift; `--version` prints it and the
window.
Composites are the host's job: verify = read back and compare, erase =
write `0xff` (per page for flash, per byte for EEPROM), timeout = EEPROM
write to the last cell.
## Deployment
The build leaves three artifacts per chip. The ELF is a container for the
tests and objcopy, never flashed. The **.hex is the programmer artifact**: it
carries its own addresses and lands the loader in its top slot, touching
nothing else. The **.bin is the self-update image** — the slot's bare bytes.
**ATmega328P**: program the loader at 0x7e00 with an external programmer;
fuses BOOTSZ = 11 (256 words) and BOOTRST programmed. Applications are
flashed unmodified — reset re-vectors to the loader in hardware, word 0
stays the application's own reset vector, and `G` jumps to 0.
**Boot-sectioned megas**: program the loader at `flash 512` with an external
programmer. Every such mega has a BOOTSZ step whose boot section is exactly
the 512-byte slot — the second-smallest step on the 8 KiB and 16 KiB chips,
the smallest on the 32 KiB ones — so the ATmega328P profiles below apply to
every one of them with its own addresses; the per-chip BOOTSZ ladders live in
the host tool (`BOOT_FUSE`).
The **644s and 1284s** are the geometry's sweet spot: their smallest boot
section (512 words = 1 KiB) is exactly *two* slots, so the resident and its
staging slot both live inside the minimum section. Self-update needs no fuse
step up, and the standalone profile does not exist — reset lands one erased
slot below the loader (0xfc00 / 0x1fc00) and walks up into it.
ATmega328P profiles (addresses for its 32 KiB):
| BOOTSZ | BOOTRST | Behavior |
|---|---|---|
| 256 words (512 B) | programmed | *Standalone*: reset always enters the loader; **self-update impossible** (the staging slot lies outside the boot section, where SPM is disabled). |
| 512 words (1 KB) | unprogrammed | *Self-update, app-first*: reset always boots the application, which owns all 31.5 KB and must offer its own jump to 0x7e00 to reach the loader (a virgin chip reaches it by reset across erased flash). Updates are power-fail-safe except mid-rewrite of the resident slot itself (no reset path leads to the staging copy then). |
| 512 words (1 KB) | programmed | *Self-update, loader-first*: reset lands at 0x7c00 — the staging slot, normally erased, so execution walks up into the loader; during an update it is the staging copy itself, so a mid-rewrite power loss recovers by reset. The loss windows move to the staging install/retire page writes instead (page-write scale). The host keeps `[0x7c00, 0x7e00)` clear of application data (`--force` overrides). |
Applications are flashed unmodified here — word 0 stays the application's own
reset vector, and the hand-over jumps to 0.
**Patched-vector chips — the tinies and the m48s** (no boot section; the m48s'
SPM runs from the entire flash, Atmel-8271 §26): program the loader at
`flash 512`; erased flash below it walks up into the loader, so a virgin
chip activates. Flashing an application then takes reset-vector surgery: word
0 becomes an `rjmp` to the loader base, and the application's own entry is
re-encoded as a trampoline `rjmp` in the word just below the loader
(`base 2`, where the hand-over jumps). Every other vector stays the
application's. The patched page 0 and the trampoline page are written *first*
and an erase runs top-down, so from the first write on an interruption still
resets into the loader.
A .bin programmed at address 0 by mistake is dead weight on a boot-sectioned
mega (SPM only executes from the boot section — reflash the .hex), but *runs*
on a patched-vector chip, and the ordinary `--update-loader` flow re-homes it
into the top slot from there (`pureboot.rehome`).
## Updating the loader
`pureboot.py --update-loader new_pureboot.bin` replaces the resident loader
with any pureboot build — a re-timed window, a newer version — using the
loader itself as its own staging loader. The image is the loader's own 512
bytes as a raw binary, or the Intel HEX the build emits beside it.
The preflight refuses an image built for another chip: the info block embedded
in every pureboot binary (signature, page size, loader base, EEPROM size,
flags) must match the device's own, and the error names both. Die revisions
share their base signature and geometry, so their images are interchangeable —
as the silicon is.
1. The staging slot `[base512, base)` is saved to a host-side state file (on
the 1 KB tiny13s that is the whole application, vectors included).
2. The resident installs the update image there. On the patched-vector chips
the host composes the slot's last word as a jump to the resident base, so
even an abandoned staging copy times out into a loader. A loader already
sitting whole in the staging slot is left as the staging copy instead —
rewriting it would only meet its own running-slot guard.
3. `J` enters the staging copy, which rewrites the resident slot. Where a
patched reset vector routes through the resident, the host first re-aims
word 0 at the staging copy, so a power loss mid-rewrite still resets into a
loader; on the tiny13s the staging slot carries the reset vector itself.
4. `J` enters the new resident, which restores the staging slot's saved
content, and the state file is discarded.
Every phase is idempotent and keyed off the actual flash state, so re-running
the same command after any interruption resumes and completes. The state file
carries the only bytes not recoverable from the device; losing it mid-update
still completes the update, and the staging region comes back by reflashing
the application. A boot-sectioned mega needs its fuses for the preflight — read
from the device, or supplied with `--assume-fuses` where reading is impossible
(simulators).
**Tinies** (no boot section): program the loader at `flash - 512`; erased
flash below it walks up into the loader, so a virgin chip activates. When
flashing an application the host performs reset-vector surgery: the
application's own `rjmp` target is re-encoded as a trampoline `rjmp` in the
word just below the loader (`base - 2`, where `G` jumps), and word 0 is
rewritten to `rjmp` to the loader base. Every other vector stays the
application's. Page 0 is written last, so an interrupted flash leaves word 0
erased and the chip still falls through to the loader on the next reset.
## Host tool
`pureboot.py` — Python 3, standard library only. The port layer is the one
platform-specific part: termios drives any tty on POSIX (a USB adapter as well
as a simavr pty), the Win32 serial API through `ctypes` drives a COM port on
Windows (`--port COM6`; the `\\.\` form for two-digit ports is supplied by the
tool). Opening the port asserts DTR and RTS on both, so a board that wires DTR
to reset gets its reset pulse and opens the activation window by itself.
`pureboot.py` — Python 3, standard library only (termios drives any tty,
a USB adapter as well as a simavr pty):
pureboot.py --port /dev/ttyUSB0 --baud 57600 \
--info --fuses --flash app.hex
--info --fuses --flash app.hex --timeout 10
Operations run in a fixed order within one session: info, fuses, loader
update, flash (erase / program / read / verify), EEPROM (the same) — then the
loader hands over to the application. `--stay` keeps the session alive
instead, and a later invocation reconnects into it. `--flash` and `--eeprom`
verify by read-back unless `--no-verify`, and a flash page that reads back
wrong is rewritten up to three times before the run stops (see `W` above).
`--verify-flash` only reports. Images are raw binary, or Intel HEX by
extension. `--force` overrides the refusable safety checks — today, flashing
application data into a mega's reset walk region.
Readouts come one fact per line: `--info` decodes the info block field by
field, `--fuses` each fuse byte plus, on a boot-sectioned mega, its decoded
meaning. Transfers that take wire time draw a transient progress bar on stderr
when it is a tty. `-v`/`--verbose` adds the decisions as they happen: knock
counts, the programming plan, update state handling and per-phase page counts.
Operations run in a fixed order within one session: info, fuses, flash
(erase / program / read / verify), EEPROM (erase / program / read / verify),
timeout — then the loader hands over to the application; `--stay` keeps the
session alive instead, and a later invocation reconnects into it (the knock
converges there too). `--flash` and `--eeprom` verify by read-back unless
`--no-verify`; images are raw binary, or Intel HEX by extension.
## Tests
`tools/check.sh` runs every chip's workflow (`--full` adds the reflect-mode
builds of libavr's spot set; `tools/make_presets.py` regenerates the presets).
Per chip preset, `ctest` runs:
- `pureboot.size` — the 510-byte (patched-vector) / 512-byte budget;
- `pureboot_*.size` — the size matrix: the serial backends × the clock ladder
(1/8/16 MHz; the t13s' own RC menu), the USART1 instance across that same
ladder on the x4 chips, and `pureboot_sw_wide`, the slowest ladder rate at
the fastest clock — where a software UART's per-bit spin outgrows its
one-register delay loop and takes the 16-bit one. That is the largest image
the configuration space produces, and a shape the ladder default (always the
*fastest* rate a clock reaches) never picks. Pins are immediate operands and
the timeout is a constant: neither is an axis;
- `pbm_*.size` — under `--full`, the exhaustive cross product replacing that
compact matrix: every plausible oscillator (the internal ones, the CKDIV8
floor, the plain and the UART crystals) × every rate reachable from it ×
every backend, unreachable combinations dropping out rather than aborting
the configure. Bounded to one chip per size-bearing class — flash
addressing, hand-over shape, page size, USART inventory — since everything
else in the image is chip-independent code;
- `pureboot.pi` — the position-independence lint: no absolute `jmp`/`call`, the
info block within the image's first 256 bytes;
- `pureboot.planner` — the host tool's pure logic: programming orders and their
recovery properties, the surgery, the staging composition, the boot-fuse
decode, the update preflight over synthetic fuse bytes, and the repairing
verify against a fake device;
- `pureboot.protocol` — end to end against a simavr device
(`test/pureboot_device.c`: a hardware USART as a pty, or a cycle-timed
GPIO⇄pty bridge for a software-UART build, plus the SPM/NVM module simavr's
tiny cores lack) driven by the real host tool through knock-from-reset,
program + verify of both memories, session reconnect, an external reset
through the patched vector, and the hand-over to a fixture application whose
banner proves the launch — cross-checked against the simulator's
ground-truth memory dumps and an independent decode of the surgery;
- `pureboot.reloc` — the identical image one slot below the resident serves the
complete command set from there;
- `pureboot.rehome` (t85) — a loader programmed at address 0 or in the staging
slot re-homes into the top slot through the ordinary update flow;
- `pureboot.custom` (328P) — the configuration example's 1 MHz software-serial
build driving the full protocol suite, proving the plumbing produces a
working loader and not just one that fits;
- `pureboot.usart1` (644A) — the same suite over the second hardware USART:
instance selection is compile-checked everywhere, but only a live session
proves the loader polls the USART it claims;
- `pureboot.dirty` (328P) — entering the loader from a running application over
an SPM buffer it deliberately dirtied, the case the loader declines to guard:
a bare verify must see the corruption and the repairing verify must fix it in
one rewrite. Hardware forbids the state here, but simavr dispatches SPM from
anywhere, which is what makes the path constructible;
- `pureboot.update` — the full `--update-loader` flow, then every power-fail
phase: the device is killed mid-write, restarted from its flash dump, and a
re-run must complete the update with the application intact.
`size`, `pi` and `planner` are host logic and run anywhere; the
simulator-driven targets need simavr and a pty, so they are POSIX-only.
Per chip preset, `ctest` runs the 512-byte size gate and the end-to-end
protocol test: a simavr device (`test/pureboot_device.c` — the mega's USART
as a pty; on the tinies a cycle-timed GPIO⇄pty bridge for the software UART,
plus the SPM/NVM module simavr's tiny cores lack) driven by the real host
tool through knock-from-reset, program + verify of both memories, timeout
configuration, session reconnect, an external reset through the patched
vector, and the hand-over to a fixture application whose banner proves the
launch — cross-checked against the simulator's ground-truth memory dumps and
an independent decode of the surgery's rjmp words.

View File

@@ -1,251 +0,0 @@
# pureboot autobaud — findings, and how the version question settled itself
The autobaud loader measures the host's bit timing **at runtime** from a
calibration pulse, so the image carries no clock: one clock-agnostic binary per
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
for the software-serial deployments — the RC-oscillator parts (the tinies,
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
each needs a per-clock build. Autobaud erases that axis. That is the win —
deployment, not bytes.
This document records how the fit was established, why the two versions that
were under review are both dead, and what the loader looks like now.
## The decision, settled by measurement
Two autobaud loaders were built for review, differing in one tradeoff — the
running-slot write guard against strict purity. `pureboot_autobaud_pure.cpp`
landed at **508 B** on the 1284P (4 B spare) and `pureboot_autobaud_reg.cpp` at
**512** (zero spare).
Then hardware testing found a defect that neither could absorb.
**A single spurious calibration pulse wedged the loader.** `run()` budgeted only
the start-edge wait inside `measure()`; the `rx()` that read the knock behind it
was unbudgeted and blocked forever. One stray low pulse on an unattended device
— EMI, or a host that opens the port and never knocks — held the loader in its
activation loop and the application never ran. On a field device that is a hang,
not a hiccup, and it is exactly the deployment autobaud is for.
The fix is to bound the whole activation: an expired knock budget returns a byte
that cannot be the knock, so control falls back into the budgeted `measure()`,
and a line that stays idle boots the application there. It costs about 22 bytes.
| 1284P, with the activation fix | size | 512 B budget |
|---|---|---|
| `pureboot_autobaud_pure.cpp` | 530 | **over by 18** |
| `pureboot_autobaud_reg.cpp` | 534 | **over by 22** |
| `pureboot_autobaud_uni.cpp` | **464** | **48 B spare** |
Both candidates were unshippable, and the margin they were competing over was
never real — it was the space the missing fix should have occupied. So the
choice is not between them. It is the third loader below, which fits with room
to spare *and* carries features neither had. The two sources stay in the tree
for the record; only the unified one is built.
## The unified loader
The insight that paid was the one that had already paid once: **merging command
bodies removes cost that moving them around only redistributes.** Folding `R`,
`r` and `w` into a single address-and-count path had been worth 14 B earlier.
Pushed further — one read command and one write command over *named spaces*
it is worth far more, because four transfer loops collapse into one.
`pureboot_autobaud_uni.cpp` is pureboot 5. It is strictly pure: no inline
assembly, no global register variable, and **no GPIOR either** — the measured
unit lives in a plain static, so the loader claims no chip resource an
application might want, and the GPIOR-versus-static question disappears along
with the chips that have no GPIOR.
### The protocol
| command | arguments | |
|---|---|---|
| `b` | — | version, then the three signature bytes |
| `J` | addr16 | ack, then jump (word address) |
| `W` | sel8, addr16, page bytes | fill the flash page buffer |
| `G` | sel8, addr16, n8 | read n bytes (0 means 256) |
| `g` | sel8, addr16, n8, then n bytes | write, each byte acked |
`sel` is `space | bank << 4`. The low nibble names the space; the high nibble is
flash's third address byte, so every transfer speaks a **byte** address inside a
64 KiB bank and no command has to carry word addresses. The host must not span a
bank boundary in one transfer — it already chunks by page, so nothing it does
comes close.
| space | | |
|---|---|---|
| 0 | flash | `lpm`/`elpm` |
| 1 | EEPROM | |
| 2 | data | SRAM — and with it the register file and every I/O register, which share the data address space on AVR |
| 3 | fuse and lock | |
| 4 | SPM | write-only: the data byte goes to SPMCSR and fires the instruction at the selected address |
Three things follow from that table that the loader never had:
- **RAM read and write**, the missing feature. In a per-command design it would
have cost a fresh dispatch arm and a fresh loop, ~30 B, on a loader with 4 B
spare. As one more space on a shared loop it is a single `ld`/`st`. It also
hands the host arbitrary **I/O register access** for free, because AVR maps
the peripherals into the same address space.
- **Host-driven SPM.** `W` used to end with a hardcoded erase, write and RWW
re-enable — 42 B. Those are now three writes to the SPM space, reusing the
store path's own address, data byte and ack. The host pays three extra
round-trips per page (18 wire bytes against 256 of data) and gains the ability
to issue *any* SPM operation, lock bits included.
- **`W` on the same footing as everything else.** It takes the same selector and
the same byte address instead of a word address of its own, which made flash
addressing uniform across the protocol *and* was 20 B cheaper than keeping its
private convention.
Read and write are `G` and `g` — the same letter, one bit apart — so the
transfer loop picks its direction with a one-word skip rather than a compare.
### Why the SPM space has to be one primitive
It is tempting to go further and expose a generic "poke this I/O register", from
which the host could drive SPM itself. The hardware forbids it: `out SPMCSR, x`
and the `spm` that follows must issue within four cycles, and EEPROM's
EEMPE→EEPE window is the same shape. A host cannot hit a four-cycle window
across a serial link. **Atomicity is the floor, and the atomic unit must be
resident.** That is the real limit on how low-level a bootloader's primitives
can go — not the byte count.
## Where the bytes went
The starting point was the 508 B pure build, disassembled and attributed:
| phase | bytes |
|---|---|
| `link::rx` + `link::tx` (bit-banged UART) | 102 |
| autobaud measure + knock | 62 |
| reset vector, pin init, WDRF check, jump, ack, EEPROM wait | 50 |
| command loop head + dispatch tree | 42 |
| `W` program flash page | 98 |
| `R` / `r` / `w` / `F` bodies, plus the shared address decode | 122 |
| `b` info, `J` jump | 32 |
**208 B — 41% — is physical layer and activation**, which no protocol change can
touch. The command bodies were the entire addressable surface, and they were
four copies of one idea.
The route from a first attempt to the final loader, all on the 1284P:
| step | size |
|---|---|
| unified `G`/`P`, `load`/`store` outlined, `__uint24` cursor | 600 |
| …`load`/`store` inlined; unit in `.noinit`, not `.bss` | 534 |
| …16-bit cursor with the bank in the selector; direction as a command bit | 510 |
| …erase/write/RWW moved out to the SPM space | 484 |
| …`W` sharing the selector-and-address decode | **464** |
Three of those steps are worth keeping as lessons:
- **Outlining `load`/`store` cost more than the four bodies they replaced.** As
functions they were 110 B against the 106 B of inline bodies — the AVR ABI's
argument marshalling plus prologue ate the entire saving. Inlined into the one
shared loop they cost only their own instructions. The call-site lesson cuts
both ways: *merging* call sites pays, *creating* one does not.
- **A `.bss` static drags in `__do_clear_bss`** — 18 B of startup code to zero a
variable that is always measured before it is read. `.noinit` is correct here
and free.
- **A three-byte cursor taxes every space.** Widening the shared cursor so flash
could reach past 64 KiB put an extra increment on EEPROM and RAM reads that
never need it. Moving the bank into the selector byte kept the cursor at
sixteen bits and cost nothing on the wire.
### What did not work
- **Encoding the space in the command byte** (so dispatch becomes masking rather
than a compare tree) cannot carry the bank's four bits alongside a space. The
cheap half of the idea survived as the direction bit; the rest lost to the
selector byte, which is also more extensible.
- **A generic primitive interpreter** — a loader with no logic at all, driven
entirely by the host — is not reachable on AVR. Harvard architecture means the
program counter cannot fetch from data space, so the classic "upload a flash
algorithm into RAM and jump to it" bootstrap is impossible, and on every
boot-sectioned part SPM only takes effect from the boot section anyway. What
remains is a fixed primitive set: still a protocol, still logic, only at a
different granularity. debugWIRE reaches that design point only because its
interpreter is *in silicon*; it costs the loader nothing because it is not in
the loader.
- Below 512 B the saved bytes are largely unspendable on the boot-sectioned
chips: the 328P's smallest boot section is exactly 512 B, and the 1284P's is
1024 B, of which pureboot already occupies only the top half. The margin
matters as headroom for correctness fixes — as this defect showed — not as
flash returned to the application. On the patch-vector parts, which have no
boot section, it *is* returned: on the ATtiny13 the loader is 43% of a 1 KiB
part, and every byte is real.
## The codegen coupling, still load-bearing
`count >> 2` is exact only because the calibration pulse's bit-count (7, from
the 0xC0 byte) equals the poll loop's cycles per iteration (7 — `sbis` 1,
`rjmp` 2, `adiw` 2, `rjmp` 2). The loop shape survived every restructuring here,
verified in the disassembly, but a toolchain bump that reshapes it would break
the lock silently. `test/pbautobaud.py` is what pins it: a wrong unit fails the
flash verify.
## Sizes — every chip
Budget 510 B on the patch-vector parts, 512 elsewhere. The 1284P is no longer
the tight one: the bank nibble made far flash *cheaper* than the near-flash
arithmetic it replaced.
| size | chips | budget | spare |
|---|---|---|---|
| 444 | ATtiny13, 13A | 510 | 66 |
| 448 | ATmega48, 48A, 48P, 48PA; ATtiny25 | 510 | 62 |
| 452 | ATtiny45, 85 | 510 | 58 |
| 460 | ATmega644, 644A, 644P, 644PA | 512 | 52 |
| 464 | **ATmega1284, 1284P**; ATmega8, 8A, 88, 88A, 88P, 88PA | 512 | 48 |
| 466 | ATmega16, 16A, 32, 32A; 164A/P/PA, 168/A/P/PA, 324A/P/PA, 328, 328P | 512 | 46 |
All 37 chips build and size-test green, plus the 12-preset reflect spot set
(guidance rule 4 — the reflect matrix is never run in full), which matches its
generated counterpart byte for byte on every chip in the set. Worst case across
the whole set is **466 B, 46 under budget**.
## Host tool and simulation
- **`pureboot.py`** speaks both generations. `Info.version >= 5` selects the
unified path; everything below it keeps the four-command protocol, so the
fixed-baud loader is untouched. `--autobaud` sends the 0xC0 pulse and one
knock, then derives full geometry from the signature. New: `--peek ADDR[:N]`
and `--poke ADDR:HEX` reach the data space.
- **`test/pbautobaud.py`** drives the loader over the GPIO⇄pty software-UART
bridge through the calibration handshake, a flash + EEPROM + fuse round-trip
cross-checked against the simulator's own memory, a RAM read/write round-trip,
and a hand-over to the fixture application — then repeats at double the F_CPU
with the same binary, which is the clock-agnostic property autobaud exists
for. Run on the near-flash 328P and the word-addressed 1284P.
- It also **pins the activation hang**: the test sends a lone calibration pulse
with no knock behind it and requires the application to boot. Against the
unfixed loader that assertion never returns.
## What remains
- **Real-hardware acceptance.** A cycle-exact simulator cannot produce what
autobaud exists for: a real RC oscillator at ±10% with drift and jitter.
simavr proves the arithmetic and the fit at exact clocks; only silicon proves
the feature. Drive an internal-oscillator ATtiny at a fixed host baud and
confirm lock plus a full flash and verify.
- **A generic `spm::command()` in libavr.** The SPM space issues a runtime
command through `spm::detail::page_command` where RAMPZ exists, and falls back
to a dispatch over the known operations where it does not — the one
preprocessor branch in the file. A two-line library addition would make it
uniform and save a few bytes on the 36 non-RAMPZ chips, none of which are
tight.
- **Retire or revive the two dead variants.** They are kept only as the record
of the measurement; nothing builds them.
## Files
- `pureboot_autobaud_uni.cpp` — the loader. pureboot 5.
- `pureboot_autobaud_pure.cpp`, `pureboot_autobaud_reg.cpp` — superseded, not
built; 530 and 534 B on the 1284P once the activation hang is fixed.
- `pureboot.py``--autobaud`, the unified transfer path, `--peek`/`--poke`.
- `test/pbautobaud.py` — the end-to-end sim test and the hang regression.
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
probe at 506 B, off-tree and gitignored; a size reference only. The unified
loader is 42 B under it, with features the probe never had.

View File

@@ -1,14 +1,17 @@
// pureboot — a serial bootloader on libavr: one C++ source, no inline
// assembly, no global register variables, 512 bytes on every chip libavr
// targets. The device speaks primitives; every composite (verify, erase,
// reset-vector surgery, self-update) lives in the host tool. Protocol,
// deployment and configuration: README.md next to this file.
// pureboot — a serial bootloader on libavr, pure by constraint: one C++
// source with no inline assembly and no global register variables, built for
// every chip libavr targets, 512 bytes on each. The device speaks primitives
// — read/program flash, read/write EEPROM, fuse bytes, an info block, run —
// and everything composite (verify, erase, reset-vector surgery on the
// tinies, timeout configuration) lives in the host tool. Protocol reference:
// README.md next to this file.
//
// The image is position-independent — PC-relative control flow, wire
// addresses in, the write guard and the info block both anchored on the
// runtime return address — so the identical binary runs from any slot. That
// is what makes a copy one slot below able to rewrite the resident one, and
// every change here has to keep it (test/check_pi.py).
// Entry: reset lands in avr::startup::entry below (BOOTRST on the mega; the
// patched reset vector — or erased flash walking up into the loader — on the
// tinies). A watchdog reset hands straight to the application. Otherwise the
// host has one activation window — EEPROM's last cell, in seconds — to knock
// ("pb"); an idle line boots the application. A session then stays in the
// command loop until 'G' hands over or the chip resets.
#include <libavr/libavr.hpp>
@@ -19,107 +22,84 @@ namespace ee = avr::eeprom;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
// Purely polled interrupts stay off, every guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Deployment parameters come from the build (pureboot_add_loader()). The
// signature is not one of them: the chip database is the only universal
// source — a tiny13A cannot read its own signature row from code.
#if !defined(PUREBOOT_CLOCK_HZ) || !defined(PUREBOOT_BAUD)
#error \
"PUREBOOT_CLOCK_HZ and PUREBOOT_BAUD select this build's clock and baud — create loader targets with pureboot_add_loader() (README.md)"
#endif
using dev = avr::device<{.clock = avr::hertz_t{PUREBOOT_CLOCK_HZ}}>;
constexpr avr::baud_t wire_baud{PUREBOOT_BAUD};
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
// Per-chip personality, from the chip database: the clocks the dogfood
// boards run (16 MHz crystal on the mega, calibrated RC on the tinies) and
// the device signature (compile-time data — the tiny13A cannot even read its
// signature row from code).
consteval avr::hertz_t clock()
{
auto reg = std::string_view{avr::hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return avr::hw::db.field_index(reg, "WDRF");
if (avr::hw::db.name == "ATtiny13A")
return 9.6_MHz;
if (avr::hw::db.name == "ATtiny85")
return 8_MHz;
return 16_MHz;
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
// Chips without a hardware boot section — the tinies and the m48s, whose SPM
// runs from anywhere (Atmel-8271 §26) — keep the application's relocated
// reset vector in the word under the slot.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
consteval std::array<std::uint8_t, 3> signature()
{
if (avr::hw::db.name == "ATtiny13A")
return {0x1e, 0x90, 0x07};
if (avr::hw::db.name == "ATtiny85")
return {0x1e, 0x93, 0x0b};
return {0x1e, 0x95, 0x0f};
}
using dev = avr::device<{.clock = clock()}>;
// Geometry: the loader owns the top 512 bytes of flash; the byte below it is
// the trampoline word (the application's relocated reset vector) on chips
// without a hardware boot section. The RWWSRE bit marks a separate boot
// section — on classic AVR the two capabilities coincide.
constexpr std::uint16_t boot_bytes = 512;
constexpr std::uint16_t base = static_cast<std::uint16_t>(spm::flash_bytes - boot_bytes);
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = avr::hw::curated::has_boot_section();
constexpr bool boot_section = avr::hw::db.field_index("SPMCSR", "RWWSRE") >= 0;
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one). A slot is 256 of
// those — one value of a wire address's high byte, where 512 bytes span two.
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint16_t wire_base =
word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base);
// The activation timeout lives in EEPROM's last cell, in seconds; the host
// rewrites it with the ordinary EEPROM-write command. An unprogrammed cell —
// 0x00 or the erased 0xff — means the 4 s default: a stray value can never
// floor the window to nothing and lock the loader out, and erasing the whole
// EEPROM resets the timeout instead of maxing it to 255 s.
constexpr std::uint16_t timeout_cell = avr::hw::db.mem.eeprom_size - 1;
constexpr std::uint8_t default_seconds = 4;
// A compile-time window, so the whole EEPROM belongs to the application;
// re-timing a deployed loader is a self-update with a re-timed build.
#if !defined(PUREBOOT_TIMEOUT)
#define PUREBOOT_TIMEOUT 8
#endif
constexpr std::uint8_t timeout_seconds = PUREBOOT_TIMEOUT;
// The loader's one identity number. The protocol carries none of its own —
// a version implies it, and the host tool holds that map (README.md).
constexpr std::uint8_t version = 4;
// The 'b' reply, byte for byte (layout: README.md). Flash-resident because
// no crt copies a .data image — and flash_table's storage carries the word
// alignment 'b' needs to halve the address on the large chips.
// One wire byte per line: this is the reply's layout, not a list.
// clang-format off
inline constexpr avr::flash_table<std::array<std::uint8_t, 12>{
// The 12-byte info block the host reads with the 'b' command; flash-resident
// (there is no crt to copy a .data image).
inline constexpr std::array<std::uint8_t, 12> info_data = {
'P',
'B',
version,
avr::hw::db.signature[0],
avr::hw::db.signature[1],
avr::hw::db.signature[2],
static_cast<std::uint8_t>(page), // 0 means 256
wire_base & 0xff,
wire_base >> 8,
1, // magic, protocol version
signature()[0],
signature()[1],
signature()[2],
static_cast<std::uint8_t>(page),
base & 0xff,
base >> 8, // app flash ends here; loader base
avr::hw::db.mem.eeprom_size & 0xff,
avr::hw::db.mem.eeprom_size >> 8,
static_cast<std::uint8_t>((boot_section ? 0 : 1) | (word_flash ? 2 : 0)), // patch-vector, word-addressed
}>
info_data;
// clang-format on
boot_section ? 0 : 1, // bit 0: host must patch the reset vector (no hardware boot section)
};
using info = avr::flash_table<info_data>;
// The serial link, per the build's PUREBOOT_USART / PUREBOOT_SOFT_SERIAL,
// defaulting to the chip's USART0 where it has one. The software receiver is
// the polled one: the vector table belongs to the application. Templates on
// the clock, so only the selected backend instantiates. pending() is the
// cheap line test the activation window polls; drain() holds until the last
// frame is off the wire, so a hand-over cannot let the target's re-init clip
// the ack.
#if defined(PUREBOOT_SOFT_SERIAL) && defined(PUREBOOT_USART)
#error "PUREBOOT_SOFT_SERIAL and PUREBOOT_USART select opposing serial backends"
#endif
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
#if defined(PUREBOOT_USART)
constexpr char usart_digit = '0' + PUREBOOT_USART;
#else
constexpr char usart_digit = '0';
#endif
// The serial link: the hardware USART where the chip has one, the polled
// software UART (no vector — the table belongs to the application) on PB0/PB1
// elsewhere. Both are class templates on the clock so only the selected
// backend is ever instantiated. pending() is the cheap line test the
// activation window polls; rx() then picks the byte up.
template <avr::hertz_t C>
consteval std::int16_t rxc_field()
{
return avr::hw::db.field_index("UCSR0A", "RXC0");
}
template <avr::hertz_t C>
struct hardware_link {
using uart = avr::uart::usart<usart_digit, C, {.baud = wire_baud, .max_baud_error = 2.5_pct}>;
// The compiled idle poll: lds UCSR0A (2), sbrc skipping the exit (2),
// sbiw + sbci + sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 10;
using uart = avr::uart::usart0<C, {.baud = 115200_Bd, .max_baud_error = 2.5_pct}>;
static void init()
{
@@ -128,7 +108,7 @@ struct hardware_link {
static bool pending()
{
return uart::rx_ready();
return avr::hw::field_impl<rxc_field<C>()>::test();
}
static std::uint8_t rx()
@@ -140,21 +120,12 @@ struct hardware_link {
{
uart::write(byte);
}
static void drain()
{
uart::drain();
}
};
template <avr::hertz_t C>
struct software_link {
using rx_t = avr::uart::software_rx_polled<C, avr::PUREBOOT_RX, wire_baud>;
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, wire_baud>;
// The compiled idle poll: sbis skipping the exit (2), sbiw + sbci +
// sbci + brne (6).
static constexpr std::uint8_t poll_cycles = 8;
using rx_t = avr::uart::software_rx_polled<C, avr::pb0, 57600_Bd>;
using tx_t = avr::uart::software_tx<C, avr::pb1, 57600_Bd>;
static void init()
{
@@ -163,7 +134,7 @@ struct software_link {
static bool pending()
{
return rx_t::start_pending();
return !avr::io::input<avr::pb0>::read(); // a start bit has begun
}
static std::uint8_t rx()
@@ -175,123 +146,74 @@ struct software_link {
{
tx_t::template write<off>(byte);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
#if defined(PUREBOOT_USART)
static_assert(avr::uart::has_usart<usart_digit>(), "PUREBOOT_USART selects a hardware USART this chip does not have");
using link = hardware_link<dev::clock>;
#elif defined(PUREBOOT_SOFT_SERIAL)
using link = software_link<dev::clock>;
#else
using link =
std::conditional_t<avr::uart::has_usart<usart_digit>(), hardware_link<dev::clock>, software_link<dev::clock>>;
#endif
using link = std::conditional_t<avr::hw::db.has_reg("UDR0"), hardware_link<dev::clock>, software_link<dev::clock>>;
// The application's entry, pinned by the linker (--defsym): word 0 on a
// boot-sectioned mega, the trampoline at base 2 elsewhere. Reaching it must
// not depend on where this copy runs, so the jump goes through a pointer, and
// [[gnu::noipa]] keeps the constant from folding back into a relative call.
// The application's entry: the linker pins pureboot_app to 0x0000 on the
// mega (reset re-vectors here through BOOTRST, so address 0 stays the
// application's own vector) and to the trampoline word at base - 2 on the
// tinies (--defsym in CMakeLists.txt).
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
[[noreturn]] void run_app()
{
target();
__builtin_unreachable();
pureboot_app();
}
[[gnu::noinline, noreturn]] void run_app()
// One activation tick is 65536 pending() polls — a pin (or flag) test plus a
// 16-bit countdown, about 8 cycles. Whole-second precision is all the
// timeout cell promises; the seconds count stays a loop bound (a runtime
// multiply would drag libgcc's __mulhi3 into the MUL-less tinies).
consteval std::uint16_t ticks_per_second()
{
jump(pureboot_app);
return static_cast<std::uint16_t>(dev::clock.hz / (65536ull * 8u));
}
static_assert(ticks_per_second() >= 1);
// The window as one 32-bit countdown, divided by the backend's counted
// poll-loop cycles. Whole seconds is all it promises.
consteval std::uint32_t window_polls()
bool pending_before(std::uint8_t seconds)
{
return timeout_seconds * static_cast<std::uint32_t>(dev::clock.hz / link::poll_cycles);
}
bool pending_before_deadline()
{
std::uint32_t polls = window_polls();
do {
std::uint16_t ticks = ticks_per_second();
do {
std::uint16_t spins = 0; // wraps first, so 65536 polls per tick
do {
if (link::pending())
return true;
} while (--polls);
} while (--spins);
} while (--ticks);
} while (--seconds);
return false;
}
// A knock byte under the deadline: an idle window means no host, so the
// application runs.
std::uint8_t rx_deadline()
// A knock byte under the activation deadline: an idle line means no host is
// there, and the application runs.
std::uint8_t rx_deadline(std::uint8_t seconds)
{
if (!pending_before_deadline())
if (!pending_before(seconds))
run_app();
return link::rx();
}
// Inlined: read across a call, the first byte strands in a call-saved
// register the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
std::uint16_t rx16()
{
std::uint16_t low = link::rx();
std::uint8_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
// The wire's byte pair as the word it is — AVR is little-endian too, so the
// cast is the identity a shift-and-or spelling makes the compiler rediscover.
// Callers read into named variables first: the wire order is a sequence of
// reads, not an argument order.
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
const std::uint8_t *flash_ptr(std::uint16_t address)
{
return std::bit_cast<std::uint16_t>(pair);
return reinterpret_cast<const std::uint8_t *>(address);
}
// Counts arrive in the wire's 8-bit form: 0 means 256. Both streamers fold
// into the one command that reads flash, which is what lets the far one's
// 24-bit cursor sit in the command loop's own call-saved registers.
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
// The streamers take the count in the wire's 8-bit form: 0 means 256.
void send_flash(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
link::tx(avr::flash_load(flash_ptr(address++)));
while (--count);
}
// The 24-bit cursor as the machine holds it — the RAMPZ byte and a 16-bit Z,
// carried apart; the reassembled address folds away inside the far load.
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
// Carrying the wrap is smaller than the flat 32-bit cursor GCC
// builds without it.
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
// Out of line: three sites send it, and a call is shorter than three
// load-immediates.
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
@@ -299,164 +221,99 @@ void send_eeprom(std::uint16_t address, std::uint8_t count)
while (--count);
}
// Host-paced: the ack goes out once the write has begun, so the next byte
// arrives while it completes and nothing is missed without a buffer.
// EEPROM write, host-paced: each ack goes out once the byte's write has
// begun, so the next byte arrives while it completes and the following
// write's own ready-wait sees an idle line. Nothing is ever missed, on
// either serial backend, without a buffer.
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
link::tx(ack);
} while (--count);
}
// One page into the SPM buffer, then erase and program — except the slot
// this code is running in (`slot_high`, from run()), which is drained and
// left alone. A broken host therefore cannot brick the running loader, and a
// copy one slot lower may rewrite the resident one.
//
// Nothing discards the buffer first: it is write-once per word (§26.2.1), so
// filling over a refused page or an application's leavings programs stale
// words — but a page write auto-erases it (§26.2.1; §19.2 on the tinies), so
// that write clears the condition and the host's read-back rewrites the page.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
// One flash page: stream the bytes into the SPM buffer as little-endian
// words, then erase and program. Addresses in the loader's own 512 bytes
// are drained but never programmed — a broken host cannot brick the chip.
// On the mega the RWW section is re-enabled so reads work immediately.
void program_flash(std::uint16_t address)
{
// The address names a page, so its in-page bits are dropped and the walk
// starts at the page base — one induction either way: a byte-addressed
// wire address walks the page itself (the offset bits wrap back to zero),
// while a word one becomes a byte cursor once. The slot index is the wire
// address's high byte — on byte-addressed chips the byte address's, with
// the low bit dropped, since a slot is two of those.
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
// A page is aligned, so it never crosses 64 KiB: RAMPZ is a per-page
// constant and the 16-bit Z's low byte is the whole in-page offset.
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
for (std::uint16_t i = 0; i < page; i += 2) {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
spm::fill<off>(address + i, static_cast<std::uint16_t>(low | (high << 8)));
}
if (page_high != slot_high) {
// Only a boot-sectioned mega runs on while its RWW section programs;
// everywhere else the CPU halts through erase and write.
if (address < base) {
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
// Programming leaves the RWW section disabled; reads need it back on. The
// same store discards the buffer (§26.2.2), so a boot-sectioned mega never
// meets the stale-word case above.
if constexpr (boot_section)
spm::rww_enable<off>();
}
}
// The four fuse and lock bytes in the hardware's own Z order: low, lock,
// extended, high.
// The four fuse/lock bytes in the hardware's own Z order: low, lock,
// extended, high. Writing fuses is not a thing self-programming can do on
// AVR — SPM reaches flash (and boot lock bits) only.
void send_fuses()
{
std::uint8_t which = 0;
do
for (std::uint8_t which = 0; which < 4; ++which)
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
// A watchdog reset belongs to the application, whose watchdog stays forced
// on until it clears WDRF — no activation window in its way.
if (avr::hw::field_impl<wdrf_field()>::test())
// A watchdog reset belongs to the application (whose watchdog stays
// forced on until it clears WDRF) — no activation window in its way.
if (avr::hw::mcusr::wdrf.test())
run_app();
link::init();
// The high byte of the slot this copy runs at, which the write guard and
// the info block both follow: the return address is a word address, so its
// high byte is the 256-word slot index, doubled back into byte terms where
// the wire counts bytes. Taken as byteswap's low byte — the builtin already
// swaps the two stacked bytes, and the double swap folds away, where `>> 8`
// would leave the swap materialized.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
std::uint8_t seconds = ee::read(timeout_cell);
if (seconds == 0 || seconds == 0xff)
seconds = default_seconds;
// 'p' then 'b', each under a fresh window; anything else is line noise.
while (rx_deadline() != 'p' || rx_deadline() != 'b') {
// The knock: 'p' then 'b', each under a fresh window; any other byte is
// line noise and waits again. Falling out of a window runs the app.
while (rx_deadline(seconds) != 'p' || rx_deadline(seconds) != 'b') {
}
for (;;) {
// No prompt while an EEPROM write runs: it blocks SPM and fuse reads
// (§26.2.1), and the prompt is the previous command's completion ack.
// No prompt while an EEPROM write runs: a pending write blocks SPM
// and fuse reads (§26.2.1), and the ack tells the host all is done.
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // info block, read relative to the running slot
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for all four: 'b' is a flash read
// whose arguments the loader already knows, so it joins the
// wire-argument three rather than streaming from a call site of its
// own. That leaves one flash streamer in the image, and lets its
// cursor live in this never-returning loop's own call-saved
// registers instead of being saved and restored around a call.
std::uint16_t address;
std::uint8_t count;
if (command == 'b') {
// The block sits in the image's first 256 bytes (check_pi.py
// asserts it) and slots are 512-aligned, so the low byte of its
// link address is its offset in any slot — halved where wire
// units are words. The high byte is runtime data, so no
// absolute address is ever materialized.
const auto link_byte =
static_cast<std::uint8_t>(reinterpret_cast<std::uint16_t>(info_data.storage.data()));
const std::uint8_t low = word_flash ? static_cast<std::uint8_t>(link_byte >> 1) : link_byte;
address = static_cast<std::uint16_t>(low | (slot_high << 8));
count = static_cast<std::uint8_t>(info_data.size());
} else {
address = rx16();
count = link::rx();
}
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
link::tx(ack);
switch (link::rx()) {
case 'b': // info block
send_flash(reinterpret_cast<std::uint16_t>(info::storage.data()), info::size());
break;
case 'R': { // read flash: addr16, n8 (0 = 256)
std::uint16_t address = rx16();
send_flash(address, link::rx());
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
program_flash(rx16());
break;
case 'r': { // read EEPROM: addr16, n8
std::uint16_t address = rx16();
send_eeprom(address, link::rx());
break;
}
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
std::uint16_t address = rx16();
store_eeprom(address, link::rx());
break;
}
case 'F': // fuse and lock bytes
send_fuses();
break;
case 'G': // hand over to the application
link::tx(ack);
run_app();
default: // unknown bytes are ignored; the loop re-acks
break;
}

File diff suppressed because it is too large Load Diff

View File

@@ -1,396 +0,0 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 530 B on the 1284P against a 512 B slot. Its 4 B of margin was
// never spare capacity; it was the space the missing fix should have occupied.
// pureboot_autobaud_uni.cpp replaces it at 464 B with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE PURE VERSION — no inline assembly, no global register variables,
// exactly the constraints the fixed-baud loader keeps.
//
// Two source files exist for review (autobaud.md):
// this one, pure, and pureboot_autobaud_reg.cpp, which keeps the running-slot
// write guard at the cost of one global register variable. They differ only in
// where the measured unit lives and whether the guard is present.
//
// What this version trades to fit 512 B in pure C++ (the 1284 at 508), each
// licensed by "the host guarantees safety" (README.md) and the owner's approval
// to simplify the info block:
// - the measured per-bit unit lives in the two general-purpose I/O scratch
// registers (GPIOR) where the chip has them, in a static otherwise —
// reached through libavr's named register surface, so no asm and no global
// register variable; the loader stays pure;
// - the info block is slimmed to the version and the signature — the chip's
// identity — from which the host derives page size, loader base, EEPROM
// size and the addressing flags via its own chip database;
// - no running-slot write guard: the host never programs the loader's own
// slot, and a broken host bricking the target is the host's bug;
// - a single-byte activation knock: the calibration pulse already proves a
// host is present.
//
// Position independence is kept and is in fact total here: control flow is
// PC-relative, the wire carries addresses, and with the slimmed info block and
// no write guard nothing anchors on the runtime address at all.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 4;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations). It lives
// in the two adjacent general-purpose I/O scratch registers (GPIOR1:GPIOR2)
// where the chip has them — in/out reach them in one word where a static's
// lds/sts take two, and there is no .bss to clear — and in a plain static
// otherwise (the t13, m8 and m16/32 have no GPIOR). Both are pure: the named
// register surface, no inline asm, no global register variable.
constexpr bool have_gpior = hw::db.reg_index("GPIOR1") >= 0 && hw::db.reg_index("GPIOR2") >= 0;
std::uint16_t unit_backing;
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
if constexpr (Gpior)
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
else
return unit_backing;
}
template <bool Gpior = have_gpior>
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
if constexpr (Gpior) {
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
} else
unit_backing = u;
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program. No running-slot write
// guard: the host guarantees it never targets the loader's own slot (the pure
// version's one dropped safety net, licensed — README.md).
void program_flash(std::uint16_t wire_address)
{
spm::flash_address_t address;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
}
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop — its cursor then lives in the loop's own call-saved
// registers instead of being saved and restored around a call
// (the call-site-count lesson, autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16());
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,371 +0,0 @@
// SUPERSEDED — kept for the record, not built. With the activation hang fixed
// (a lone calibration pulse used to wedge the loader, see autobaud.md) this
// version is 534 B on the 1284P against a 512 B slot, so the global register
// variable it broke purity for buys nothing. pureboot_autobaud_uni.cpp replaces
// it at 464 B, strictly pure and with more features.
//
// pureboot autobaud — the software-serial variant that measures the host's bit
// timing at runtime, so one binary runs at any F_CPU: the image carries no
// clock. THE REGISTER VERSION — keeps the running-slot write guard, at the cost
// of one global register variable (r4) holding the measured unit. That variable
// is pureboot's single, deliberate break from its no-global-register-variable
// rule, present only in this variant; everything else stays pure C++.
//
// Two source files exist for review (autobaud.md):
// pureboot_autobaud_pure.cpp, fully pure but dropping the write guard, and this
// one. They differ only in where the measured unit lives (a call-saved register
// here, GPIOR/RAM there) and whether the guard is present.
//
// The register buys ~32 B over a RAM home — an outlined rx/tx reads it with one
// move where a static costs an lds — and that is what lets the write guard stay
// while the image still fits 512 B (the 1284 at 512, exactly). The unit is
// written once through a noinline setter so the store lands immediately before a
// ret: GCC otherwise deletes a global-register store whose only readers are
// callees (autobaud.md, upstream bug 6).
//
// Simplifications shared with the pure version, each licensed: a slimmed info
// block (version + signature; the host derives geometry from its chip database)
// and a single-byte activation knock (the calibration pulse already proves a
// host). Position independence is kept: control flow is PC-relative and the
// write guard anchors on the runtime return address, as the fixed-baud loader.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
// The measured per-bit delay (in _delay_loop_2 four-cycle iterations) in a
// call-saved register that the serial callees read directly, so no wire path
// threads it and rx/tx reach it with a move, not a load. Written only through
// set_unit() below.
register std::uint16_t g_unit asm("r4");
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock; only the software-UART pins are a parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
constexpr bool word_flash = spm::flash_bytes > 65536;
constexpr std::uint8_t version = 4;
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The one store into g_unit, isolated so it lands right before the ret: a
// global-register store whose only later readers are callees is dropped
// otherwise (autobaud.md, upstream bug 6).
[[gnu::noinline]] void set_unit(std::uint16_t v)
{
g_unit = v;
}
// The autobaud software link: bit-banged with cycle-counted delays, but the
// per-bit delay is g_unit, measured from the host's calibration pulse.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into g_unit. The host sends 0xC0 — a start bit
// plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so count >> 2 is the bit
// period in _delay_loop_2's four-cycle iterations. 0 means the budget
// expired.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
set_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
_delay_loop_2(static_cast<std::uint16_t>(g_unit + (g_unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(g_unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. Without this the knock's edge wait was unbounded, so a single
// spurious calibration pulse (EMI, or a host that opens the port and never
// knocks) wedged the loader and the application never ran.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
out_t::clear(); // start bit
_delay_loop_2(g_unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(g_unit);
}
out_t::set(); // stop bit
_delay_loop_2(g_unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_near(std::uint16_t address, std::uint8_t count)
{
do
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
while (--count);
}
[[maybe_unused, gnu::always_inline]] inline void send_flash_far(std::uint16_t address, std::uint8_t count)
{
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
do {
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
if (++z == 0)
++rampz;
} while (--count);
}
[[gnu::always_inline]] inline void send_flash(std::uint16_t address, std::uint8_t count)
{
if constexpr (word_flash)
send_flash_far(address, count);
else
send_flash_near(address, count);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
void send_eeprom(std::uint16_t address, std::uint8_t count)
{
do
link::tx(ee::read(address++));
while (--count);
}
void store_eeprom(std::uint16_t address, std::uint8_t count)
{
do {
ee::write<off>(address++, link::rx());
tx_ack();
} while (--count);
}
// One page into the SPM buffer, then erase and program — except the slot this
// code is running in (`slot_high`, from run()), which is drained and left
// alone. A broken host therefore cannot brick the running loader, and a copy one
// slot lower may rewrite the resident one.
void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
{
spm::flash_address_t address;
std::uint8_t page_high;
if constexpr (word_flash) {
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1) & ~static_cast<std::uint16_t>(page - 1);
std::uint16_t z = z0;
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z, word_of({low, high}));
z += 2;
} while (static_cast<std::uint8_t>(z));
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
page_high = static_cast<std::uint8_t>(wire_address >> 8);
} else {
address = static_cast<spm::flash_address_t>(wire_address & ~static_cast<std::uint16_t>(page - 1));
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
spm::fill<off>(address, word_of({low, high}));
address += 2;
} while (static_cast<std::uint8_t>(address) & (page - 1));
address -= 2; // back inside the page — erase and write ignore the word bits
page_high = static_cast<std::uint8_t>(address >> 8) & 0xfe;
}
if (page_high != slot_high) {
spm::erase_page<off>(address);
if constexpr (boot_section)
spm::wait();
spm::write_page<off>(address);
if constexpr (boot_section)
spm::wait();
}
if constexpr (boot_section)
spm::rww_enable<off>();
}
void send_fuses()
{
std::uint8_t which = 0;
do
link::tx(spm::read_fuse<off>(static_cast<spm::fuse>(which)));
while (++which != 4);
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// The high byte of the slot this copy runs at, which the write guard
// follows: the return address is a word address, so its high byte is the
// 256-word slot index, doubled back into byte terms on a byte-addressed
// chip. Taken as byteswap's low byte — the builtin already swaps the two
// stacked bytes, and the double swap folds away.
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
const std::uint8_t ra_high = static_cast<std::uint8_t>(std::byteswap(ra_words));
const std::uint8_t slot_high = word_flash ? ra_high : static_cast<std::uint8_t>(ra_high << 1);
// Measure the calibration pulse into g_unit, then take one 'p' knock. An
// expired budget (no host) boots the application; a pulse that decodes to
// anything but 'p' re-measures.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'R': // read flash: addr16, n8 (0 = 256)
case 'r': // read EEPROM: addr16, n8
case 'w': { // write EEPROM: addr16, n8, then n bytes each acked
// One address-and-count path for the three, so the flash streamer
// keeps a single call site and inlines into this never-returning
// loop (autobaud.md).
const std::uint16_t address = rx16();
const std::uint8_t count = link::rx();
if (command == 'r')
send_eeprom(address, count);
else if (command == 'w')
store_eeprom(address, count);
else
send_flash(address, count);
break;
}
case 'W': // program one flash page: addr16, page bytes
program_flash(rx16(), slot_high);
break;
case 'F': // fuse and lock bytes
send_fuses();
break;
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,411 +0,0 @@
// pureboot autobaud, the unified-primitive version — VARIANT A, an explicit
// space byte. One read command and one write command carry a space selector, so
// flash, EEPROM, RAM and the fuses share a single cursor, a single transfer loop
// and a single argument decode instead of one command body each.
//
// Strictly pure: no inline assembly, no global register variables, and no GPIOR
// either — the measured unit lives in a plain static, so the loader claims no
// chip resource an application might want. (PUREBOOT_UNIT_GPIOR=1 puts it back
// in the I/O scratch registers, kept only as a measurement axis.)
//
// Against pureboot_autobaud_pure.cpp this version:
// - adds RAM read and write, which the loader has never had. Because AVR maps
// the register file and the whole I/O space into the data address space,
// that one space also gives the host arbitrary peripheral access for free;
// - collapses 'R' (read flash), 'r' (read EEPROM), 'w' (write EEPROM) and 'F'
// (fuses) — four bodies, four loops — into 'G' and 'P' over four spaces;
// - fixes the activation hang: a lone calibration pulse used to leave the
// loader blocked forever in the knock's rx(), so a stray edge on an
// unattended device wedged it in the loader and the application never ran.
//
// Position independence is kept and is total: control flow is PC-relative, the
// wire carries addresses, and nothing anchors on the runtime address.
#include <libavr/libavr.hpp>
#include <util/delay_basic.h>
using namespace avr::literals;
namespace spm = avr::spm;
namespace ee = avr::eeprom;
namespace hw = avr::hw;
namespace pureboot {
namespace {
// Purely polled: every interrupt guard folds to nothing.
constexpr auto off = avr::irq::guard_policy::unused;
constexpr std::uint8_t ack = '+';
// Autobaud carries no clock, so PUREBOOT_CLOCK_HZ / PUREBOOT_BAUD are not
// consulted; only the software-UART pins are a deployment parameter.
#if !defined(PUREBOOT_RX)
#define PUREBOOT_RX pb0
#endif
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
// The unit's home. A plain static by default — pureboot claims no GPIOR, so the
// application keeps both scratch registers. The GPIOR spelling is retained
// behind a macro purely so the two can be measured against each other.
#if !defined(PUREBOOT_UNIT_GPIOR)
#define PUREBOOT_UNIT_GPIOR 0
#endif
// The watchdog reset flag's home: MCUSR, or the classic megas' MCUCSR.
consteval std::int16_t wdrf_field()
{
auto reg = std::string_view{hw::db.regs[static_cast<std::size_t>(avr::power::detail::reset_reg())].name};
return hw::db.field_index(reg, "WDRF");
}
// The loader owns the top 512 bytes; a staging copy goes in the slot below.
constexpr std::uint16_t slot_bytes = 512;
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
constexpr std::uint16_t page = spm::page_bytes;
constexpr bool boot_section = hw::curated::has_boot_section();
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so flash
// addresses there are word addresses ('J' always was one).
constexpr bool word_flash = spm::flash_bytes > 65536;
// The loader's one identity number (README.md); the slimmed info block carries
// it and the signature, and the host maps a version to its protocol.
constexpr std::uint8_t version = 5;
// The activation window as a fixed poll budget: with no clock, whole seconds
// cannot be timed. A __uint24 (AVR's three-byte type) holds it — a fourth byte
// would cost two words at each countdown step for range never used.
#if !defined(PUREBOOT_AUTOBAUD_POLLS)
#define PUREBOOT_AUTOBAUD_POLLS 4000000
#endif
constexpr __uint24 autobaud_budget = PUREBOOT_AUTOBAUD_POLLS;
// The two SPM commands the loader still has to recognise by value, on the chips
// where it cannot issue a runtime one generically. Taken from the chip's own
// definitions rather than spelled 3 and 5 — though every part pureboot targets
// agrees on those, which is what lets the host send the raw SPMCSR byte.
constexpr std::uint8_t spm_erase = __BOOT_PAGE_ERASE;
constexpr std::uint8_t spm_write = __BOOT_PAGE_WRITE;
// The spaces a transfer can name. Flash is 0 so it is the cheap default.
//
// sp_spm is the one that is not memory: a write there hands its data byte to
// SPMCSR and fires the instruction at the given flash address, so page erase,
// page write and RWW re-enable become host-issued commands instead of a
// hardcoded tail inside 'W'. The store side already owns an address, a data
// byte and an ack, so the whole sequence costs only the fused out/spm pair. It
// also lets the host reach every other SPM operation — lock bits included —
// which the loader previously had no way to expose.
enum : std::uint8_t { sp_flash = 0, sp_eeprom = 1, sp_ram = 2, sp_fuse = 3, sp_spm = 4 };
// A transfer's selector byte is `space | bank << 4`: the low nibble names the
// space, the high nibble carries flash's third address byte (RAMPZ) on the
// chips that have one. Putting the bank here rather than widening the address
// keeps the shared cursor sixteen bits for every space — a three-byte cursor
// costs its extra increment on EEPROM and RAM reads too, which never need it.
// The host must not span a bank boundary in one transfer; it already chunks by
// page, so nothing it does today comes close.
// .noinit, not .bss: the unit is always measured before it is read, so it needs
// no zeroing — and a zeroed .bss would drag in __do_clear_bss, 18 bytes of
// startup code for a variable that is written before its first use.
[[gnu::section(".noinit")]] std::uint16_t unit_backing;
[[gnu::always_inline]] inline std::uint16_t get_unit()
{
#if PUREBOOT_UNIT_GPIOR
return static_cast<std::uint16_t>(hw::reg_impl<hw::db.reg_index("GPIOR1")>::read() |
(hw::reg_impl<hw::db.reg_index("GPIOR2")>::read() << 8));
#else
return unit_backing;
#endif
}
[[gnu::always_inline]] inline void put_unit(std::uint16_t u)
{
#if PUREBOOT_UNIT_GPIOR
hw::reg_impl<hw::db.reg_index("GPIOR1")>::write(static_cast<std::uint8_t>(u));
hw::reg_impl<hw::db.reg_index("GPIOR2")>::write(static_cast<std::uint8_t>(u >> 8));
#else
unit_backing = u;
#endif
}
// The autobaud software link: bit-banged with cycle-counted delays like the
// fixed-baud software backend, but the per-bit delay is the measured unit, not
// a consteval constant. rx/tx load the unit once into a local, so each bit
// spins from a register with no per-bit reload.
struct link {
using in_t = avr::io::input<avr::PUREBOOT_RX, avr::io::pull::up>;
using out_t = avr::io::output<avr::PUREBOOT_TX>;
static void init()
{
avr::init<in_t, out_t>();
out_t::set(); // idle high
}
// Time the calibration pulse into the unit. The host sends 0xC0 — a start
// bit plus six zero data bits are one low pulse of seven bit-times — and the
// counted poll loop is seven cycles an iteration, so the count is the pulse
// length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2 is that
// period in _delay_loop_2's four-cycle iterations. Waits for the start edge
// under the poll budget; 0 (returned, and stored) means the budget expired.
//
// The shape of the counting loop is load-bearing, not incidental: the >> 2
// is exact only while the pulse's bit-count equals the loop's cycles per
// iteration. Both are 7 here (sbis 1 + rjmp 2 + adiw 2 + rjmp 2). Reshaping
// this loop silently changes the lock; test/pbautobaud.py is what pins it.
static std::uint16_t measure(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
std::uint16_t count = 0;
while (!in_t::read())
++count;
const std::uint16_t u = static_cast<std::uint16_t>(count >> 2);
put_unit(u);
return u;
}
// The eight data bits, entered just past the falling edge of the start bit.
// Split out of rx() so the activation path can wait for that edge under a
// budget while the command loop waits for it indefinitely.
static std::uint8_t sample()
{
const std::uint16_t unit = get_unit();
_delay_loop_2(static_cast<std::uint16_t>(unit + (unit >> 1))); // 1.5 bits to the LSB centre
std::uint8_t value = 0;
for (std::uint8_t bit = 0; bit < 8; ++bit) {
value >>= 1;
if (in_t::read())
value |= 0x80;
_delay_loop_2(unit);
}
return value;
}
static std::uint8_t rx()
{
while (in_t::read()) // await the start edge
;
return sample();
}
// A byte under the poll budget, for the activation knock. An expired budget
// returns 0, which is not the knock, so the caller falls back into the
// budgeted measure() — and a line that stays idle boots the application
// there. That is the whole hang fix: no wait during activation is unbounded,
// so a stray calibration pulse can no longer wedge the loader.
static std::uint8_t rx_bounded(__uint24 budget)
{
while (in_t::read())
if (!--budget)
return 0;
return sample();
}
static void tx(std::uint8_t value)
{
const std::uint16_t unit = get_unit();
out_t::clear(); // start bit
_delay_loop_2(unit);
for (std::uint8_t bit = 0; bit < 8; ++bit) {
out_t::write(value & 1);
value >>= 1;
_delay_loop_2(unit);
}
out_t::set(); // stop bit
_delay_loop_2(unit);
}
static void drain()
{
// The software transmitter returns only after the stop bit.
}
};
extern "C" [[noreturn]] void pureboot_app();
[[gnu::noipa, noreturn]] void jump(void (*target)())
{
target();
__builtin_unreachable();
}
[[gnu::noinline, noreturn]] void run_app()
{
jump(pureboot_app);
}
// Inlined: read across a call, the first byte strands in a call-saved register
// the caller has to push and pop.
[[gnu::always_inline]] inline std::uint16_t rx16()
{
std::uint16_t low = link::rx();
return static_cast<std::uint16_t>(low | (link::rx() << 8));
}
[[gnu::always_inline]] inline std::uint16_t word_of(std::array<std::uint8_t, 2> pair)
{
return std::bit_cast<std::uint16_t>(pair);
}
[[gnu::noinline]] void tx_ack()
{
link::tx(ack);
}
// One byte out of any space. The four accessors share the cursor, the loop and
// the call site — the whole point of the unified commands — so each costs only
// its own instruction rather than a body, a loop and a dispatch arm.
[[gnu::always_inline]] inline std::uint8_t load(std::uint8_t space, std::uint8_t bank, std::uint16_t at)
{
if (space == sp_eeprom)
return ee::read(at);
if (space == sp_ram)
return *reinterpret_cast<volatile std::uint8_t *>(at);
if (space == sp_fuse)
return spm::read_fuse<off>(static_cast<spm::fuse>(at));
if constexpr (word_flash)
return avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(bank) << 16) | at);
else
return avr::flash_load(reinterpret_cast<const std::uint8_t *>(at));
}
// One byte into a writable space. Flash is not one of them — it arrives a page
// at a time through 'W' — and the fuses are not writable at all.
[[gnu::always_inline]] inline void store(std::uint8_t space, [[maybe_unused]] std::uint8_t bank, std::uint16_t at,
std::uint8_t value)
{
if (space == sp_ram) {
*reinterpret_cast<volatile std::uint8_t *>(at) = value;
return;
}
if (space == sp_spm) {
// The fused store-and-fire: SPMCSR takes the data byte and the SPM
// issues against Z in the same unscheduled pair the hardware's
// four-cycle window demands, which is exactly why this is one
// primitive and not a poke of SPMCSR followed by a poke of anything
// else. A host cannot hit that window across a serial link.
// The preprocessor rather than `if constexpr` only because
// spm::detail::page_command does not exist at all where there is no
// RAMPZ, and a discarded constexpr branch outside a template is still
// name-checked. A generic spm::command() in libavr would let the
// runtime command through on every chip and retire the dispatch below.
#if defined(RAMPZ)
spm::detail::page_command(value, (static_cast<spm::flash_address_t>(bank) << 16) | at);
#else
if (value == spm_erase)
spm::erase_page<off>(at);
else if (value == spm_write)
spm::write_page<off>(at);
else if constexpr (boot_section)
spm::rww_enable<off>();
#endif
if constexpr (boot_section)
spm::wait();
return;
}
ee::write<off>(at, value);
}
// One page into the SPM buffer, and only that: the erase, the write and the RWW
// re-enable that used to follow are now three host-issued writes to sp_spm,
// which reach the same fused out/spm pair through the store path's own address
// and data. No running-slot write guard: the host guarantees it never targets
// the loader's own slot (licensed — README.md).
void program_flash([[maybe_unused]] std::uint8_t bank, std::uint16_t at)
{
std::uint16_t z = at & ~static_cast<std::uint16_t>(page - 1);
do {
std::uint8_t low = link::rx();
std::uint8_t high = link::rx();
#if defined(RAMPZ)
spm::fill<off>((static_cast<spm::flash_address_t>(bank) << 16) | z, word_of({low, high}));
#else
spm::fill<off>(z, word_of({low, high}));
#endif
z += 2;
} while (static_cast<std::uint8_t>(z) & (page - 1));
}
[[noreturn]] void run()
{
if (hw::field_impl<wdrf_field()>::test())
run_app();
link::init();
// Measure the calibration pulse into the unit, then take one 'p' knock —
// both under the poll budget. An expired budget (no host) boots the
// application; anything but 'p', including the knock timing out, re-measures
// and so returns to the budgeted wait that boots it.
for (;;) {
if (link::measure(autobaud_budget) == 0)
run_app();
if (link::rx_bounded(autobaud_budget) == 'p')
break;
}
for (;;) {
ee::wait();
tx_ack();
const std::uint8_t command = link::rx();
switch (command) {
case 'J': { // jump to a wire word address: hand-over and staging transfer
auto target = reinterpret_cast<void (*)()>(rx16());
tx_ack();
link::drain();
jump(target);
}
case 'b': // chip identity: version then the three signature bytes
link::tx(version);
link::tx(hw::db.signature[0]);
link::tx(hw::db.signature[1]);
link::tx(hw::db.signature[2]);
break;
case 'W': // program one flash page: sel8, addr16, then page bytes
case 'G': // read: sel8, addr16, n8 (0 = 256)
case 'g': { // write: sel8, addr16, n8, then n bytes each acked
// The unified transfer. One decode, one cursor, one loop for every
// space and both directions — the four command bodies this replaces
// each carried their own copy of all three. Read and write are the
// same letter in the two cases, so the direction is bit 5 of the
// command and the loop tests it with a one-word skip. 'W' joins the
// same selector-and-address decode rather than keeping a word
// address of its own, which makes flash addressing uniform across
// every command that names it and costs nothing to share.
const std::uint8_t sel = link::rx();
const std::uint8_t space = sel & 0x0f;
const std::uint8_t bank = static_cast<std::uint8_t>(sel >> 4);
std::uint16_t at = rx16();
if (command == 'W') {
program_flash(bank, at);
break;
}
std::uint8_t count = link::rx();
do {
if (command & 0x20) {
store(space, bank, at, link::rx());
tx_ack();
} else
link::tx(load(space, bank, at));
++at;
} while (--count);
break;
}
default: // unknown bytes are ignored; the loop re-acks
break;
}
}
}
} // namespace
} // namespace pureboot
template struct avr::startup::entry<pureboot::run>;

View File

@@ -1,48 +0,0 @@
#!/usr/bin/env python3
"""Position-independence lint: the two link-time facts that let the identical
image run from any slot, asserted from the built ELF.
1. No absolute jmp/call — -mrelax normally guarantees it, but a branch that
grows out of relaxation range would break it silently.
2. The info block within the image's first 256 bytes: 'b' rebuilds its
address as (running slot high byte : link address low byte).
Usage: check_pi.py <objdump> <nm> <elf> <text_start_hex>
"""
import re
import subprocess
import sys
def main():
objdump, nm, elf, text_start = sys.argv[1:]
text_start = int(text_start, 0)
listing = subprocess.run([objdump, "-d", elf], capture_output=True, text=True, check=True).stdout
absolute = [
line
for line in listing.splitlines()
if re.search(r"\t(jmp|call)\t", line)
]
if absolute:
print("FAIL: absolute control flow in the image:")
print("\n".join(absolute))
sys.exit(1)
symbols = subprocess.run([nm, "-C", elf], capture_output=True, text=True, check=True).stdout
info = [line for line in symbols.splitlines() if "flash_table" in line and "::storage" in line]
if len(info) != 1:
print(f"FAIL: expected one info-block storage symbol, found {len(info)}")
sys.exit(1)
address = int(info[0].split()[0], 16)
offset = address - text_start
if not 0 <= offset < 256:
print(f"FAIL: info block at image offset {offset:#x}, must sit in the first 256 bytes")
sys.exit(1)
print(f"PI lint: control flow PC-relative, info block at offset {offset:#x}")
if __name__ == "__main__":
main()

View File

@@ -1,19 +1,8 @@
// Test-fixture application for the pureboot protocol tests: prints "APP" on
// the chip's serial link (the same link the loader uses) — the proof that
// the loader's hand-over, and on the tinies the host's reset-vector
// surgery, actually launched it. Linked normally (crt, vectors at 0); on
// the tinies its reset vector is the rjmp the host re-homes.
//
// On the hardware-USART link it then listens, and an 'L' makes it jump into
// the resident loader — the application-owned loader entry a
// BOOTRST-unprogrammed mega relies on (reset always boots the application
// there), exercised by the self-update tests. The software link idles:
// reset reaches those loaders through the patched vector (or the runner
// models BOOTRST), so the application owes them nothing.
//
// The fixture speaks the deployment its loader was built for: the same
// PUREBOOT_* defines configure it, and without them it assumes the stock
// deployment (the crystal/RC clock table below, the chip's natural link).
// Test-fixture application for the pureboot protocol test: prints "APP" on
// the chip's serial link (the same link the loader uses) and idles — the
// proof that the loader's hand-over, and on the tinies the host's
// reset-vector surgery, actually launched it. Linked normally (crt, vectors
// at 0); on the tinies its reset vector is the rjmp the host re-homes.
#include <libavr/libavr.hpp>
using namespace avr::literals;
@@ -22,86 +11,31 @@ namespace {
consteval avr::hertz_t clock()
{
#if defined(PUREBOOT_CLOCK_HZ)
return avr::hertz_t{PUREBOOT_CLOCK_HZ};
#else
auto name = std::string_view{avr::hw::db.name};
if (name.starts_with("ATtiny13"))
if (avr::hw::db.name == "ATtiny13A")
return 9.6_MHz;
if (name.starts_with("ATtiny"))
if (avr::hw::db.name == "ATtiny85")
return 8_MHz;
return 16_MHz;
#endif
}
#if !defined(PUREBOOT_TX)
#define PUREBOOT_TX pb1
#endif
#if !defined(PUREBOOT_USART)
#define PUREBOOT_USART 0
#endif
consteval bool use_hardware()
{
#if defined(PUREBOOT_SOFT_SERIAL)
return false;
#else
return avr::hw::db.has_instance("USART0") || avr::hw::db.has_instance("USART");
#endif
}
using dev = avr::device<{.clock = clock()}>;
template <avr::hertz_t C, bool Hardware = use_hardware()>
template <avr::hertz_t C, bool Hardware = avr::hw::db.has_reg("UDR0")>
struct link {
#if defined(PUREBOOT_BAUD)
static constexpr avr::baud_t baud{PUREBOOT_BAUD};
#else
static constexpr avr::baud_t baud{115200};
#endif
using tx_t = avr::uart::usart<'0' + PUREBOOT_USART, C, {.baud = baud, .max_baud_error = 2.5_pct}>;
using tx_t = avr::uart::usart0<C, {.baud = 115200_Bd, .max_baud_error = 2.5_pct}>;
static void tx(char c)
{
tx_t::write(static_cast<std::uint8_t>(c));
}
[[noreturn]] static void idle()
{
// 'L' hands back to the loader in the top slot — 512 bytes on every
// chip. The jump takes a word address, which is what makes the
// >64 KiB chips' entry reachable through a 16-bit pointer at all.
constexpr std::uint32_t slot = 512;
for (;;) {
auto command = tx_t::read_blocking();
if (command == 'L')
reinterpret_cast<void (*)()>(static_cast<std::uint16_t>((avr::hw::db.mem.flash_size - slot) / 2))();
// 'D' leaves every word of the SPM page buffer dirty, so that a
// following 'L' enters the loader with the buffer it never clears.
if (command == 'D') {
for (std::uint16_t at = 0; at < avr::spm::page_bytes; at += 2)
avr::spm::fill(at, 0xdead);
tx('D');
}
}
}
};
template <avr::hertz_t C>
struct link<C, false> {
#if defined(PUREBOOT_BAUD)
static constexpr avr::baud_t baud{PUREBOOT_BAUD};
#else
static constexpr avr::baud_t baud{57600};
#endif
using tx_t = avr::uart::software_tx<C, avr::PUREBOOT_TX, baud>;
using tx_t = avr::uart::software_tx<C, avr::pb1, 57600_Bd>;
static void tx(char c)
{
tx_t::write(static_cast<std::uint8_t>(c));
}
[[noreturn]] static void idle()
{
while (true) {
}
}
};
} // namespace
@@ -112,5 +46,6 @@ int main()
link<dev::clock>::tx('A');
link<dev::clock>::tx('P');
link<dev::clock>::tx('P');
link<dev::clock>::idle();
while (true) {
}
}

View File

@@ -1,154 +0,0 @@
#!/usr/bin/env python3
"""End-to-end autobaud test: drive an autobaud loader in simavr through the
calibration handshake and a flash + EEPROM + fuse round-trip, cross-checked
against the simulator's ground-truth memory — then repeat at a second F_CPU with
the *same* loader binary, which is the property autobaud exists for: one
clock-agnostic image that locks onto whatever rate the host sends.
Usage: pbautobaud.py <device_bin> <loader_elf> <mcu> <base_hex> <page>
<app_bin> <app_hz> <app_baud> <tool_py> <workdir>
The loader is a software-serial build on PB0/PB1 (pureboot_add_autobaud's
default), so the runner drives it over the GPIO⇄pty bridge (-l sw:B0,B1). The
app fixture is built for (app_hz, app_baud); the hand-over is checked at that
point, and a second point at half the clock proves the lock is measured, not
baked in.
"""
import os
import sys
import time
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
(device_bin, elf, mcu, base_hex, page, app_bin, app_hz, app_baud, tool, workdir) = sys.argv[1:]
base, page, app_hz, app_baud = int(base_hex, 0), int(page), int(app_hz), int(app_baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
ee_image = bytes(range(0xA0, 0xB0))
ee_path = os.path.join(workdir, "ee.bin")
open(ee_path, "wb").write(ee_image)
# The geometry the surgery planner needs, from the chip class the runner is
# told — the same derivation pbtest.py makes: the boot-sectioned megas need
# no vector surgery, the tinies and the boot-section-less m48s do, and the
# large chips speak word addresses.
mega = mcu.startswith("atmega")
patch = not mega or mcu.startswith("atmega48")
word_flash = base + pb.SLOT > 0x10000
wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0)
ground_truth = pb.Info(bytes([ord("P"), ord("B"), pb.NEWEST_LOADER, 0, 0, 0, page & 0xFF,
wire_base & 0xFF, wire_base >> 8, 0, 0, flags]))
def round_trip(hz, baud, label, hand_over):
"""One clock point: reset, calibrate + knock, program, verify against the
simulator's own flash, and (at the app's point) hand over to the fixture."""
dump = os.path.join(workdir, f"flash_{label}.bin")
device = pbsim.Device(device_bin, elf, mcu, str(hz), base_hex, page, baud, dump, link="sw:B0,B1")
try:
# The host tool, in autobaud mode, sends the 0xC0 calibration pulse
# and a single knock at `baud`; the loader locks to it.
out = pbsim.run_tool(tool, device.pty, baud, "--autobaud", "--info", "--fuses",
"--flash", app_bin, "--eeprom", ee_path, "--stay")
for needed in ("version", "signature", "fuses", "verify:", "stays"):
if needed not in out:
fail(f"{label}: session output lacks {needed!r}\n{out}")
# Read both memories back over the locked link and check them.
read_flash = os.path.join(workdir, f"rf_{label}.bin")
read_eeprom = os.path.join(workdir, f"re_{label}.bin")
out = pbsim.run_tool(tool, device.pty, baud, "--autobaud", "--verify-flash", app_bin,
"--verify-eeprom", ee_path, "--read-flash", read_flash,
"--read-eeprom", read_eeprom, "--stay")
if out.count("verify:") != 2:
fail(f"{label}: did not verify both memories\n{out}")
if open(read_eeprom, "rb").read()[: len(ee_image)] != ee_image:
fail(f"{label}: EEPROM read-back mismatch")
if hand_over:
# Regression: a calibration pulse with no knock behind it must
# not wedge the loader. The knock's edge wait used to be
# unbudgeted, so one stray low pulse — EMI, or a host that opens
# the port and never knocks — held the loader forever and the
# application never ran. The whole activation is bounded now, so
# the window closes and the app boots; the banner is the proof.
# (The pause lets the loader reach its measurement loop, so the
# pulse is genuinely seen and the test cannot pass vacuously.)
device.reset()
port = pb.Port(device.pty, baud)
try:
time.sleep(0.2)
port.write(bytes((pb.CALIBRATE,)))
# Accumulate rather than match exactly: the reset leaves the
# idle line a framing artefact ahead of the banner, which is
# noise here — the question is only whether the app ran.
seen = b""
deadline = time.monotonic() + 180.0
while b"APP" not in seen and time.monotonic() < deadline:
seen += port.read_available(1.0)
if b"APP" not in seen:
fail(f"{label}: lone calibration pulse wedged the loader — app never bannered, saw {seen!r}")
print(f" {label}: lone calibration pulse does not wedge the loader")
finally:
port.close()
device.reset()
port = pb.Port(device.pty, baud)
try:
loader = pb.Loader(port)
live = loader.connect_autobaud(15)
if not pb.OLDEST_LOADER <= live.version <= pb.NEWEST_LOADER:
fail(f"{label}: loader reports pureboot {live.version}")
if loader.unified:
# pureboot 5's data space. 0x0200 is clear of the
# loader's own .noinit unit at the bottom of SRAM and of
# the stack at the top. Reading it back over the same
# locked link proves both directions of the new space.
probe = bytes(range(0x30, 0x40))
loader.write_ram(0x0200, probe)
if loader.read_ram(0x0200, len(probe)) != probe:
fail(f"{label}: RAM round-trip mismatch")
# The register file and the I/O space share the data
# address space on AVR, so the same command reaches a
# peripheral register. SPMCSR reads back as idle here.
verbose_ram = loader.read_ram(0x0200, 4)
print(f" {label}: RAM read/write ok ({verbose_ram.hex()})")
loader.run_application()
banner = port.read_exact(3, 5.0)
if banner != b"APP":
fail(f"{label}: application banner was {banner!r}")
finally:
port.close()
finally:
device.stop()
# Ground truth (read after the runner exits and writes its dump): what
# the tool programmed must be what the simulator actually holds.
pages = pb.plan_flash(open(app_bin, "rb").read(), ground_truth)
flash_true = open(dump, "rb").read()
for address, data in pages.items():
if flash_true[address : address + page] != data:
fail(f"{label}: simulator flash differs from the programmed image at {address:#06x}")
print(f" {label}: locked at {hz} Hz / {baud} Bd, flash+EEPROM verified"
+ (", hand-over ok" if hand_over else ""))
# The app fixture is built for one clock; the hand-over banners there. A
# second point at double that clock, same loader binary, proves the lock is
# measured, not baked in — the whole point of autobaud. (Doubling keeps the
# bit period healthy; halving would drop it below the software UART's floor.)
round_trip(app_hz, app_baud, "clock-a", hand_over=True)
round_trip(app_hz * 2, app_baud, "clock-b", hand_over=False)
print("pbautobaud: calibration lock and flash/EEPROM/fuse round-trip pass at both clocks")
if __name__ == "__main__":
main()

View File

@@ -1,88 +0,0 @@
#!/usr/bin/env python3
"""Dirty-page-buffer acceptance test: with no discard in the loader, a page
filled over words an earlier writer left takes those instead. The whole
contract is asserted — a bare verify sees the corruption, the repairing
verify fixes it in one rewrite, and it stays fixed.
The state is reached the one way the loader cannot prevent: an application
dirties the buffer and jumps in with no reset between. Boot-sectioned megas
forbid that outright (SPM runs only from the boot section, Atmel-8271 §26.2),
but simavr dispatches SPM from anywhere, which is what makes it constructible.
Usage: pbdirty.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <app_bin> <tool_py> <workdir>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, app_bin, tool, workdir = sys.argv[1:]
page, baud = int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
dump = os.path.join(workdir, "dump.bin")
# Reset boots the application on a BOOTRST-unprogrammed mega; its 'L' is
# the loader entry this test needs, reached without a reset.
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, reset_hex="0")
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
loader.connect(25)
# Install the application and hand over to it.
pb.op_flash(loader, app_bin, erase=False, verify=True)
loader.run_application()
if port.read_exact(3, 5.0) != b"APP":
fail("the application did not start")
port.write(b"D")
if port.read_exact(1, 5.0) != b"D":
fail("the application did not acknowledge dirtying the page buffer")
port.write(b"L")
loader = pb.Loader(port)
loader.connect(25)
# Program by hand, so the corruption is observable before anything
# repairs it.
pages = pb.plan_flash(open(app_bin, "rb").read(), loader.info)
for address in sorted(pages):
loader.write_page(address, pages[address])
try:
pb.verify_pages(loader, pages)
except pb.Error as error:
if "verify failed" not in str(error):
fail(f"the read-back failed, but not at verify: {error}")
else:
# Either the fixture no longer dirties the buffer, or the loader
# clears it again — in which case this test's premise is gone.
fail("programming over a dirty page buffer came back clean")
# What the programming path uses: one rewrite settles it, and it stays
# settled.
pb.verify_pages(loader, pages, repair=True)
pb.verify_pages(loader, pages)
# Ground truth beyond the loader's own read-back.
loader.run_application()
if port.read_exact(3, 5.0) != b"APP":
fail("the application did not start after the recovered write")
port.close()
finally:
device.stop()
print("pbdirty: a dirty page buffer is caught by verify and cleared by the retry")
if __name__ == "__main__":
main()

View File

@@ -1,96 +0,0 @@
#!/usr/bin/env python3
"""Re-homing acceptance test: an image programmed somewhere other than its
canonical slot must still be a working loader, and the ordinary
--update-loader flow must put a build into the top slot from there.
Two positions. Address 0, a raw .bin handed to a programmer: the staging
install and the word-0 redirect run from copies outside page 0's slot, so the
running-slot guard never blocks them. And the staging slot itself, where a
loader already sitting there IS the staging copy — recognized by its embedded
block and left in place, then streaming the new resident like any staged copy.
Usage: pbrehome.py <device_bin> <pureboot_elf> <update_bin> <mcu> <hz>
<base_hex> <page> <baud> <app_bin> <tool_py> <workdir>
"""
import os
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def rehome_from(pbsim, pb, device_bin, elf, place_hex, guard_probe, update_bin, base, page, baud, app_bin, workdir,
mcu, hz):
"""Place the loader at `place_hex`, heal through --update-loader, flash
the application, expect the banner."""
dump = os.path.join(workdir, f"dump-{place_hex}.bin")
state = os.path.join(workdir, f"rehome-{place_hex}.pbstate")
if os.path.exists(state):
os.unlink(state)
device = pbsim.Device(device_bin, elf, mcu, hz, place_hex, page, baud, dump, reset_hex="0")
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
info = loader.connect(25)
if info.base != base:
fail(f"the misplaced copy reports base {info.base:#06x} — the info block must stay canonical")
# The accidental slot still guards itself; re-homing rides on the
# canonical slots being writable from it.
probe = int(guard_probe, 0)
before = loader.read_flash(probe, info.page)
loader.write_page(probe, bytes(info.page))
if loader.read_flash(probe, info.page) != before:
fail("the misplaced copy's guard let its own slot change")
# The ordinary update flow puts the build into the top slot.
pb.op_update_loader(loader, 25, update_bin, state, None)
update = open(update_bin, "rb").read()
if loader.read_flash(base, len(update)) != update:
fail("the canonical slot does not hold the update image")
# An application flashed through the healed resident overwrites the
# stale copy (surgery included) and launches.
pages = pb.plan_flash(open(app_bin, "rb").read(), loader.info)
for address in pb.covered(pages, loader.info, skip_blank=False):
loader.write_page(address, pages[address])
pb.verify_pages(loader, pages)
loader.run_application()
if port.read_exact(3, 5.0) != b"APP":
fail(f"application does not banner after the re-home from {place_hex}")
port.close()
finally:
device.stop()
def main():
(device_bin, elf, update_bin, mcu, hz, base_hex, page, baud, app_bin, tool, workdir) = sys.argv[1:]
base, page, baud = int(base_hex, 0), int(page), int(baud)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
# Address 0: the raw-.bin-to-a-programmer accident. The guard probe is
# the copy's own page 0.
rehome_from(pbsim, pb, device_bin, elf, "0x0", "0x0", update_bin, base, page, baud, app_bin, workdir, mcu, hz)
print("re-home from address 0: converged")
# The staging slot: erased flash with the loader sitting exactly where
# a staging copy would — the tool must leave it in place and let it
# stream the (different) update build into the resident slot.
stage = base - pb.SLOT
rehome_from(pbsim, pb, device_bin, elf, hex(stage), hex(stage), update_bin, base, page, baud, app_bin, workdir,
mcu, hz)
print("re-home from the staging slot: converged")
print("pbrehome: a misplaced loader re-homes through the ordinary update flow")
if __name__ == "__main__":
main()

View File

@@ -1,96 +0,0 @@
#!/usr/bin/env python3
"""Position-independence acceptance test: the identical binary, flashed one
slot below the resident, must serve the complete command set from there. The
info block must come back byte-identical, the write guard must refuse the
staged copy's own slot and permit the resident's, and the staged copy must be
able to rewrite the resident verbatim.
Usage: pbreloc.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <tool_py> <workdir>
"""
import os
import subprocess
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def main():
device_bin, elf, mcu, hz, base_hex, page, baud, tool, workdir = sys.argv[1:]
base, page, baud = int(base_hex, 0), int(page), int(baud)
stage = None # derived from the device's own info (slot-sized) below
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
objcopy = os.environ.get("PB_OBJCOPY", "avr-objcopy")
image_path = os.path.join(workdir, "pureboot.bin")
subprocess.run([objcopy, "-O", "binary", elf, image_path], check=True)
image = open(image_path, "rb").read()
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, os.path.join(workdir, "dump.bin"))
try:
port = pb.Port(device.pty, baud)
loader = pb.Loader(port)
info = loader.connect(25)
if info.base != base:
fail(f"info reports base {info.base:#06x}")
resident_info = info.raw
# Install the staging copy exactly as the update flow would.
stage = info.stage
staged = pb.staging_content(image, info)
pb.write_differing(loader, stage, staged)
# Enter it; from here on, every command runs in the relocated copy.
staged_info = loader.enter_copy(stage, 25)
if staged_info.raw != resident_info:
fail(f"staged info {staged_info.raw.hex()} != resident info {resident_info.hex()}")
# 'R' from the staged copy already proved itself in the install
# verify; 'F' must answer 4 bytes (values are unmodeled in simavr).
if len(loader.read_fuses()) != 4:
fail("fuse read from the staged copy")
# EEPROM round-trip through the staged copy.
pattern = bytes(range(0x50, 0x60))
loader.write_eeprom(0, pattern)
if loader.read_eeprom(0, len(pattern)) != pattern:
fail("EEPROM round-trip through the staged copy")
# The guard, both ways: its own slot refused (drained, unchanged), the
# resident slot writable. The refusal leaves its drained words in the
# SPM buffer, so the write that follows may take them — and clears
# them by writing, so the retry must not.
before = loader.read_flash(stage, page)
loader.write_page(stage, bytes(page))
if loader.read_flash(stage, page) != before:
fail("the staged copy's guard let its own slot change")
marker = bytes((i * 3) & 0xFF for i in range(page))
loader.write_page(base, marker)
if loader.read_flash(base, page) != marker:
loader.write_page(base, marker)
if loader.read_flash(base, page) != marker:
fail("the staged copy could not write the resident slot, even on retry")
# Restore the resident image through the staged copy, then 'J' back
# into it and prove it lives.
resident = image + b"\xff" * (pb.SLOT - len(image))
pb.write_differing(loader, base, resident)
back_info = loader.enter_copy(base, 25)
if back_info.raw != resident_info:
fail("the restored resident does not serve its info block")
port.close()
finally:
device.stop()
print("pbreloc: the relocated copy serves the full command set")
if __name__ == "__main__":
main()

View File

@@ -1,68 +0,0 @@
"""Shared simavr harness for the pureboot tests: spawn the device runner,
hand out its pty, restart it from a flash dump (the power-fail path), and
keep its chatter out of undrained pipes."""
import os
import signal
import subprocess
class Device:
def __init__(self, binary, elf, mcu, hz, base_hex, page, baud, dump, reset_hex=None, resume=None, link=None):
cmd = [binary]
if link:
cmd += ["-l", link]
cmd += [elf, mcu, hz, base_hex, str(page), str(baud), dump]
if reset_hex is not None or resume is not None:
# Chips without a hardware boot section — the tinies and the
# m48s — reset to address 0 like silicon; the boot-sectioned
# megas re-vector to the loader base (BOOTRST).
patch = not mcu.startswith("atmega") or mcu.startswith("atmega48")
cmd.append(reset_hex if reset_hex is not None else ("0" if patch else base_hex))
if resume is not None:
cmd.append(resume)
self.log = open(dump + ".log", "a")
self.proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=self.log, text=True)
self.dump = dump
self.pty = None
for _ in range(50):
line = self.proc.stdout.readline()
if not line:
break
if line.startswith("PB_PTY"):
self.pty = line.split()[1]
break
if not self.pty:
self.stop()
raise RuntimeError("device did not report a pty")
def reset(self):
"""The external reset line: SIGUSR1 re-enters at the reset vector."""
self.proc.send_signal(signal.SIGUSR1)
def power_fail(self):
"""SIGTERM: the runner dumps its flash and exits — the image a
restart resumes from."""
self.stop()
return self.dump
def stop(self):
self.proc.terminate()
try:
self.proc.wait(timeout=5)
except subprocess.TimeoutExpired:
self.proc.kill()
self.log.close()
def run_tool(tool, pty, baud, *args, timeout=180):
result = subprocess.run(
[os.environ.get("PYTHON", "python3"), tool, "--port", pty, "--baud", str(baud), "--wait", "25", *args],
capture_output=True,
text=True,
timeout=timeout,
)
print(result.stdout, end="")
if result.returncode != 0:
raise RuntimeError(f"tool exited {result.returncode}: {result.stderr.strip()}")
return result.stdout

View File

@@ -1,17 +1,19 @@
#!/usr/bin/env python3
"""End-to-end protocol test: drive the simavr device with the real host tool
over its pty through flash, EEPROM, fuse and hand-over scenarios, and
cross-check the tool's view against the simulator's ground-truth dumps.
"""End-to-end pureboot protocol test: spawn the simavr device, then drive it
with the real host tool (pureboot.py, as a subprocess over the device's pty)
through flash + EEPROM + timeout + fuse + hand-over scenarios, and cross-check
the tool's view against the simulator's ground-truth memory dumps.
Usage: pbtest.py <device_bin> <pureboot_elf> <mcu> <hz> <base_hex> <page>
<baud> <eeprom_size> <app_bin> <tool_py> <workdir> [link]
The optional link is the runner's -l spec (usart1, sw:B5,B1, ...), for a
loader built off the chip's natural serial default.
<baud> <eeprom_size> <app_bin> <tool_py> <workdir>
Exits 0 if every scenario passes.
"""
import os
import signal
import subprocess
import sys
import time
def fail(message):
@@ -31,14 +33,53 @@ def rjmp_decode(word, at, flash_words):
return (at + 1 + offset) % flash_words
class Device:
def __init__(self, binary, elf, mcu, hz, base, page, baud, dump):
self.proc = subprocess.Popen(
[binary, elf, mcu, hz, base, str(page), str(baud), dump],
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True,
)
self.dump = dump
self.pty = None
deadline = time.time() + 5
while time.time() < deadline:
line = self.proc.stdout.readline()
if not line:
break
if line.startswith("PB_PTY"):
self.pty = line.split()[1]
break
if not self.pty:
self.stop()
raise RuntimeError("device did not report a pty")
def stop(self):
self.proc.terminate()
try:
self.proc.wait(timeout=3)
except subprocess.TimeoutExpired:
self.proc.kill()
def run_tool(tool, pty, baud, *args):
result = subprocess.run(
[sys.executable, tool, "--port", pty, "--baud", str(baud), "--wait", "20", *args],
capture_output=True,
text=True,
timeout=120,
)
print(result.stdout, end="")
if result.returncode != 0:
fail(f"tool exited {result.returncode}: {result.stderr.strip()}")
return result.stdout
def main():
args = sys.argv[1:]
link = args.pop() if len(args) == 12 else None
(device_bin, elf, mcu, hz, base_hex, page, baud, eeprom_size, app_bin, tool, workdir) = args
(device_bin, elf, mcu, hz, base_hex, page, baud, eeprom_size, app_bin, tool, workdir) = sys.argv[1:]
base, page, baud, eeprom_size = int(base_hex, 0), int(page), int(baud), int(eeprom_size)
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
os.makedirs(workdir, exist_ok=True)
@@ -49,33 +90,25 @@ def main():
read_flash = os.path.join(workdir, "readback_flash.bin")
read_eeprom = os.path.join(workdir, "readback_eeprom.bin")
# The geometry the host will discover, for computing the expected image:
# the boot-sectioned megas need no vector surgery (the tinies and the
# boot-section-less m48s do), the large chips speak word addresses, and
# the page byte is the wire's 0-means-256.
mega = mcu.startswith("atmega")
patch = not mega or mcu.startswith("atmega48")
word_flash = base + pb.SLOT > 0x10000
wire_base = base // 2 if word_flash else base
flags = (1 if patch else 0) | (2 if word_flash else 0)
# The geometry the host will discover, for computing the expected image.
info = pb.Info(
bytes([ord("P"), ord("B"), pb.NEWEST_LOADER, 0, 0, 0, page & 0xFF])
+ bytes([wire_base & 0xFF, wire_base >> 8, eeprom_size & 0xFF, eeprom_size >> 8])
+ bytes([flags])
bytes([ord("P"), ord("B"), 1, 0, 0, 0, page])
+ bytes([base & 0xFF, base >> 8, eeprom_size & 0xFF, eeprom_size >> 8])
+ bytes([0 if mcu == "atmega328p" else 1])
)
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, link=link)
device = Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump)
try:
# Session 1: knock from reset, identify, program everything, stay.
out = pbsim.run_tool(tool, device.pty, baud, "--info", "--fuses", "--flash", app_bin,
"--eeprom", ee_path, "--stay")
for needed in ("version", "signature", "fuses", "verify:", "stays"):
out = run_tool(tool, device.pty, baud, "--info", "--fuses", "--flash", app_bin,
"--eeprom", ee_path, "--timeout", "8", "--stay")
for needed in ("device: signature", "fuses:", "verify:", "activation timeout: 8 s", "stays"):
if needed not in out:
fail(f"session 1 output lacks {needed!r}")
# Session 2: reconnect into the live session, verify, dump, hand over
# is deferred — the pty must be reopened for the APP banner first.
out = pbsim.run_tool(tool, device.pty, baud, "--verify-flash", app_bin, "--verify-eeprom", ee_path,
out = run_tool(tool, device.pty, baud, "--verify-flash", app_bin, "--verify-eeprom", ee_path,
"--read-flash", read_flash, "--read-eeprom", read_eeprom, "--stay")
if out.count("verify:") != 2:
fail("session 2 did not verify both memories")
@@ -83,6 +116,8 @@ def main():
eeprom_back = open(read_eeprom, "rb").read()
if eeprom_back[: len(ee_image)] != ee_image:
fail("EEPROM read-back mismatch")
if eeprom_back[-1] != 8:
fail(f"timeout cell reads {eeprom_back[-1]}, expected 8")
# The expected post-surgery flash, straight from the tool's planner.
pages = pb.plan_flash(open(app_bin, "rb").read(), info)
@@ -93,33 +128,16 @@ def main():
# An external reset re-enters through the patched word 0 (tinies; the
# runner resets them to address 0 like silicon) or BOOTRST (mega).
# The loader must answer a fresh knock, and the 'J' hand-over must
# land in the application, which banners on the same link.
device.reset()
# The loader must answer a fresh knock, and 'G' must land in the
# application, which banners on the same link.
device.proc.send_signal(signal.SIGUSR1)
port = pb.Port(device.pty, baud)
try:
loader = pb.Loader(port)
live = loader.connect(15)
# The loader built from this tree must report a version the tool
# beside it speaks — a bump the tool was never told about is a
# loader it would refuse to talk to. Not equality with the newest:
# the tool now spans two loader generations, the fixed-baud one
# here and the unified autobaud loader that follows it.
if not pb.OLDEST_LOADER <= live.version <= pb.NEWEST_LOADER:
fail(f"loader reports pureboot {live.version}, the tool speaks "
f"{pb.OLDEST_LOADER}..{pb.NEWEST_LOADER}")
# A W addressed inside a page rather than at its base must still
# consume exactly one page and prompt. The loader's own slot is
# the target — it is drained and never programmed — and the
# payload is erased-state bytes, so the probe can disturb neither
# the image nor the page buffer it leaves behind.
wire = wire_base + 1
port.write(bytes((ord("W"), wire & 0xFF, wire >> 8)) + b"\xff" * page)
loader.connect(15)
port.write(b"G")
if port.read_exact(1, 5.0) != pb.PROMPT:
fail("unaligned W did not return to the prompt")
loader.run_application()
fail("no ack for G")
banner = port.read_exact(3, 5.0)
if banner != b"APP":
fail(f"application banner was {banner!r}")
@@ -136,10 +154,9 @@ def main():
fail("loader region looks erased in the ground-truth dump")
# The surgery, decoded independently: the patched vector must land on the
# loader, the trampoline on the application's own entry (patched-vector
# chips only — a boot-sectioned mega's word 0 stays the application's).
if patch:
flash_words = (base + pb.SLOT) // 2
# loader, the trampoline on the application's own entry.
if mcu != "atmega328p":
flash_words = (base + 512) // 2
app = open(app_bin, "rb").read()
word0 = flash_true[0] | (flash_true[1] << 8)
if rjmp_decode(word0, 0, flash_words) != base // 2:
@@ -151,7 +168,7 @@ def main():
ee_true_path = dump + ".eeprom"
if os.path.exists(ee_true_path):
ee_true = open(ee_true_path, "rb").read()
if ee_true[: len(ee_image)] != ee_image:
if ee_true[: len(ee_image)] != ee_image or ee_true[-1] != 8:
fail("ground-truth EEPROM does not match what was programmed")
print("pbtest: all scenarios pass")

View File

@@ -1,222 +0,0 @@
#!/usr/bin/env python3
"""Self-update end-to-end: an application is flashed, the loader replaces
itself with a re-timed build, and every power-fail phase is rehearsed by
killing the device mid-write, restarting it from its flash dump, and letting
a re-run complete the update.
The boot-sectioned megas run the BOOTRST-unprogrammed profile — reset boots
the application, whose 'L' is the application-owned loader entry — with
--assume-fuses standing in for the fuse read simavr cannot model.
Usage: pbupdate.py <device_bin> <pureboot_elf> <update_elf> <mcu> <hz>
<base_hex> <page> <baud> <app_bin> <tool_py> <workdir>
"""
import os
import subprocess
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def rjmp_decode(word, at, flash_words):
"""Written against the instruction-set definition, not with the tool's
encoder, so an encoding bug cannot verify itself."""
if word & 0xF000 != 0xC000:
fail(f"word at {at * 2:#06x} is {word:#06x}, not an rjmp")
offset = word & 0x0FFF
if offset >= 0x800:
offset -= 0x1000
return (at + 1 + offset) % flash_words
class PowerFail(Exception):
pass
def assumed_fuses(pb, image):
"""Synthetic 'F' bytes for --assume-fuses: the smallest boot section
covering both the resident and the staging slot (two slots — what a
self-update needs), BOOTRST unprogrammed — the per-chip BOOTSZ ladder
and fuse byte come from the tool's own table, keyed by the update
image's embedded signature."""
info = pb.image_info(image)
which, ladder = pb.BOOT_FUSE[bytes(info.signature[1:3])]
bits = min((b for b in ladder if ladder[b] * 2 >= 2 * pb.SLOT), key=lambda b: ladder[b])
fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF))
fuses[which] = 0xF8 | (bits << 1) | 1
return bytes(fuses)
def make_fault_loader(pb, base, slot, kill_region, kill_hits, device):
"""A Loader whose write_page kills the device (or, with device=None,
just the host) at the Nth write into a region; the sequence
stage->resident->stage distinguishes the install from the restore."""
class FaultLoader(pb.Loader):
def __init__(self, port):
super().__init__(port)
self.seen_resident = False
self.hits = 0
def write_page(self, address, data):
if address >= base:
phase = "resident"
self.seen_resident = True
elif address >= base - slot:
phase = "stage_restore" if self.seen_resident else "stage"
else:
phase = "app"
if phase == kill_region:
self.hits += 1
if self.hits == kill_hits:
if device is not None:
device.power_fail()
raise PowerFail(f"{kill_region} write {kill_hits}")
super().write_page(address, data)
return FaultLoader
def main():
(device_bin, elf, update_elf, mcu, hz, base_hex, page, baud, app_bin, tool, workdir) = sys.argv[1:]
base, page, baud = int(base_hex, 0), int(page), int(baud)
mega = mcu.startswith("atmega")
# The m48s are megas without a boot section: patched vector, no fuse
# preflight, and the same reset-to-0 the tinies get.
patch = not mega or mcu.startswith("atmega48")
reset_hex = "0" if mega else None # the boot-sectioned mega runs BOOTRST-unprogrammed here
sys.path.insert(0, os.path.dirname(os.path.abspath(tool)))
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import pbsim
import pureboot as pb
slot = pb.SLOT
os.makedirs(workdir, exist_ok=True)
objcopy = os.environ.get("PB_OBJCOPY", "avr-objcopy")
images = {}
for name, source in (("v0", elf), ("v9", update_elf)):
path = os.path.join(workdir, name + ".bin")
subprocess.run([objcopy, "-O", "binary", source, path], check=True)
images[name] = open(path, "rb").read()
if images["v0"] == images["v9"]:
fail("the update image is byte-identical to the resident build")
dump = os.path.join(workdir, "dump.bin")
state = os.path.join(workdir, "update.pbstate")
fuses = assumed_fuses(pb, images["v0"]) if mega and not patch else None
def connect(device):
port = pb.Port(device.pty, baud)
if mega:
# Reset boots the application here; its 'L' is the loader entry.
# To a live loader the same byte is an ignored command.
port.read_available(0.5)
port.write(b"L")
loader = pb.Loader(port)
loader.connect(25)
return port, loader
def padded(image):
return image + b"\xff" * (slot - len(image))
def resident_bytes(loader):
return loader.read_flash(base, slot)
def assert_state(loader, image, app_pages):
if resident_bytes(loader) != padded(image):
fail("resident loader does not match the update image")
stage = base - slot
got = loader.read_flash(stage, slot)
for address, data in app_pages.items():
if stage <= address < base:
if got[address - stage : address - stage + page] != data:
fail(f"staging region page {address:#06x} not restored")
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump, reset_hex=reset_hex)
final = "v0"
try:
# The application first — its planner output is the restore truth.
pbsim.run_tool(tool, device.pty, baud, "--flash", app_bin, "--stay")
port, loader = connect(device)
app_pages = pb.plan_flash(open(app_bin, "rb").read(), loader.info)
port.close()
# A clean CLI update, resident -> v9.
args = ["--update-loader", os.path.join(workdir, "v9.bin"), "--state", state, "--stay"]
if fuses:
args += ["--assume-fuses", fuses.hex()]
out = pbsim.run_tool(tool, device.pty, baud, *args)
if "loader updated" not in out:
fail("update did not report success")
if os.path.exists(state):
fail("state file survived a completed update")
port, loader = connect(device)
assert_state(loader, images["v9"], app_pages)
loader.run_application()
if port.read_exact(3, 5.0) != b"APP":
fail("application does not banner after the update")
port.close()
final = "v9"
print("clean update: resident replaced, staging restored, application intact")
# Power-fail rehearsal: kill mid-phase, restart from the dump,
# re-run, and the update must still complete. Each round flips the
# direction so the flash is never already at its target. The mega's
# mid-resident-rewrite loss is exercised as a host crash instead:
# with BOOTRST unprogrammed and the resident mid-erase, a power loss
# there has no reset path into the staging copy — the documented
# cost of that profile (README).
for kill_region, kill_hits, kill_device in (
("stage", 2, True),
("resident", 1, patch),
("stage_restore", 2, True),
):
device.reset() # the previous round left the application running
port, loader = connect(device)
target = "v9" if resident_bytes(loader) == padded(images["v0"]) else "v0"
image_path = os.path.join(workdir, target + ".bin")
injected = make_fault_loader(pb, base, slot, kill_region, kill_hits, device if kill_device else None)(port)
injected.info = loader.info
try:
pb.op_update_loader(injected, 25, image_path, state, fuses)
fail(f"{kill_region}: fault never triggered")
except PowerFail as event:
print(f"power fail injected: {event}")
port.close()
if kill_device:
device = pbsim.Device(device_bin, elf, mcu, hz, base_hex, page, baud, dump,
reset_hex=reset_hex, resume=dump)
port, loader = connect(device)
pb.op_update_loader(loader, 25, image_path, state, fuses)
assert_state(loader, images[target], app_pages)
loader.run_application()
if port.read_exact(3, 5.0) != b"APP":
fail(f"{kill_region}: application lost after the resumed update")
port.close()
final = target
print(f"resumed after {kill_region} loss: update completed, application intact")
finally:
device.stop()
# Ground truth: the simulator's own flash against the final state, and
# on the patched-vector chips an independent decode of the reset routing.
flash = open(dump, "rb").read()
if flash[base : base + slot] != padded(images[final]):
fail("ground-truth resident region does not match the final image")
if patch:
flash_words = (base + slot) // 2
word0 = flash[0] | (flash[1] << 8)
if rjmp_decode(word0, 0, flash_words) != base // 2:
fail("ground-truth reset vector does not land on the loader")
app = open(app_bin, "rb").read()
trampoline = flash[base - 2] | (flash[base - 1] << 8)
if rjmp_decode(trampoline, (base - 2) // 2, flash_words) != rjmp_decode(app[0] | (app[1] << 8), 0, flash_words):
fail("ground-truth trampoline does not land on the application entry")
print("pbupdate: clean update + all power-fail phases recovered")
if __name__ == "__main__":
main()

View File

@@ -1,16 +1,12 @@
// simavr "device" for the pureboot protocol tests, every chip. Loads the
// boot-linked ELF at the loader base, starts execution there (BOOTRST / the
// patched vector are not what is under test), and exposes the loader's
// simavr "device" for the pureboot protocol tests, all three chips. Loads
// the boot-linked ELF at the loader base, starts execution there (BOOTRST /
// the patched vector are not what is under test), and exposes the loader's
// serial link as a pty for the real host tool:
//
// - Hardware USART builds: simavr's uart_pty on the selected instance.
// - Software UART builds: an 8N1 bridge between a pty and the GPIO pins,
// timed against the simulated cycle counter (drives the loader's RX,
// decodes its TX).
//
// The link follows the chip's natural default (USART0 on the megas, the
// software UART on PB0/PB1 elsewhere) unless -l overrides it: `-l usart1`
// for the second instance, `-l sw:B5,B1` for a software build's RX,TX pins.
// - ATmega328P: the hardware USART0 through simavr's uart_pty.
// - Tinies: an 8N1 bridge between a pty and the GPIO software UART
// (drives PB0, the loader's RX; decodes PB1, its TX), timed against the
// simulated cycle counter.
//
// simavr's tiny cores decode the SPM opcode but attach no NVM module — SPM
// is a silent no-op (the mega's boot section has one, avr_flash). The
@@ -42,84 +38,11 @@
static avr_t *avr;
static uart_pty_t uart_pty;
static int link_software;
static char uart_digit = '0';
static char sw_rx_port = 'B', sw_tx_port = 'B';
static int sw_rx_bit = 0, sw_tx_bit = 1;
static int use_uart_pty;
static const char *dump_path;
static uint32_t reset_pc;
static volatile sig_atomic_t reset_requested;
static int parse_link(const char *spec)
{
if (strcmp(spec, "usart0") == 0 || strcmp(spec, "usart1") == 0) {
link_software = 0;
uart_digit = spec[5];
return 0;
}
if (strncmp(spec, "sw", 2) == 0) {
link_software = 1;
if (spec[2] == '\0')
return 0;
if (sscanf(spec + 2, ":%c%d,%c%d", &sw_rx_port, &sw_rx_bit, &sw_tx_port, &sw_tx_bit) == 4)
return 0;
}
return -1;
}
// simavr 1.6's avr_flash PGERS handler erases spm_pagesize bytes starting at
// Z & ~1 instead of the page containing Z (its PGWRT path masks correctly) —
// hardware ignores the in-page bits (§26.8.1), so an erase issued with Z
// anywhere inside the page wipes half the neighbouring page in simulation
// only. Wrap the mega's registered flash ioctl and re-dispatch page erases
// with Z forced to the page boundary; everything else passes through.
//
// A second gap on the boot-section-less m48s: their RWWSRE bit is the
// temporary-buffer discard (Atmel-8271 §26.2/§26.3.1), but the stock model
// gates its RWWSRE branch on AVR_SELFPROG_HAVE_RWW — absent on the m48
// core — so the discard store falls through into the buffer-fill branch and
// plants whatever Z/R1:R0 happen to hold. Perform the silicon's discard
// here instead.
static avr_flash_t *mega_flash;
static int (*mega_flash_ioctl)(avr_io_t *io, uint32_t ctl, void *param);
static int fixed_flash_ioctl(avr_io_t *io, uint32_t ctl, void *param)
{
if (ctl == AVR_IOCTL_FLASH_SPM && avr_regbit_get(io->avr, mega_flash->pgers)) {
uint16_t z = (uint16_t)(io->avr->data[30] | (io->avr->data[31] << 8));
uint16_t masked = (uint16_t)(z & ~(mega_flash->spm_pagesize - 1));
io->avr->data[30] = (uint8_t)masked;
io->avr->data[31] = (uint8_t)(masked >> 8);
int result = mega_flash_ioctl(io, ctl, param);
io->avr->data[30] = (uint8_t)z;
io->avr->data[31] = (uint8_t)(z >> 8);
return result;
}
if (ctl == AVR_IOCTL_FLASH_SPM && !(mega_flash->flags & AVR_SELFPROG_HAVE_RWW) &&
(io->avr->data[mega_flash->r_spm] & 0x11) == 0x11) { // RWWSRE|SELFPRGEN: the m48 buffer discard
for (int i = 0; i < mega_flash->spm_pagesize / 2; i++) {
mega_flash->tmppage[i] = 0xffff;
mega_flash->tmppage_used[i] = 0;
}
avr_regbit_clear(io->avr, mega_flash->selfprgen);
return 0;
}
return mega_flash_ioctl(io, ctl, param);
}
static void fix_mega_flash_erase(void)
{
for (avr_io_t *io = avr->io_port; io; io = io->next) {
if (io->kind && strcmp(io->kind, "flash") == 0) {
mega_flash = (avr_flash_t *)io;
mega_flash_ioctl = io->ioctl;
io->ioctl = fixed_flash_ioctl;
return;
}
}
fprintf(stderr, "device: no flash module to fix — SPM page erases may misalign\n");
}
static void request_reset(int sig)
{
(void)sig;
@@ -131,7 +54,6 @@ static void request_reset(int sig)
typedef struct {
avr_io_t io;
uint8_t buffer[128];
uint8_t used[128]; // a buffer word loads once until erased — like silicon
unsigned page;
} tiny_nvm_t;
@@ -149,21 +71,16 @@ static int nvm_ioctl(avr_io_t *io, uint32_t ctl, void *param)
uint32_t page_base = (uint32_t)(z & ~(n->page - 1)) % (mcu->flashend + 1);
if (command == 0x01) { // SPMEN alone: buffer fill from r1:r0
unsigned offset = z & (n->page - 1) & ~1u;
if (!n->used[offset]) { // first write wins until the buffer clears
n->buffer[offset] = mcu->data[0];
n->buffer[offset + 1] = mcu->data[1];
n->used[offset] = 1;
}
} else if (command == 0x03) { // PGERS
memset(mcu->flash + page_base, 0xff, n->page);
} else if (command == 0x05) { // PGWRT: programming only clears bits
for (unsigned i = 0; i < n->page; i++)
mcu->flash[page_base + i] &= n->buffer[i];
memset(n->buffer, 0xff, n->page);
memset(n->used, 0, n->page);
} else if (command == 0x11) { // CTPB
memset(n->buffer, 0xff, n->page);
memset(n->used, 0, n->page);
}
mcu->data[0x57] &= (uint8_t)~0x1f; // the operation completes instantly
return 0;
@@ -245,14 +162,8 @@ static void rx_start_next(void)
// A reset abandons whatever the bridge was mid-transfer: bytes still queued
// for a chip that no longer has the context to receive them meaningfully,
// and a decode in progress on a TX line the reset may have already changed.
// The pending cycle timers must go with the state: avr_reset drops the TX
// output latch, whose falling edge starts a spurious decode before this
// runs, and a stale tx_sample would then interleave with the loader's first
// real answer through the shared shift state, corrupting it.
static void bridge_reset(void)
{
avr_cycle_timer_cancel(avr, tx_sample, NULL);
avr_cycle_timer_cancel(avr, rx_step, NULL);
rx_head = rx_tail = 0;
rx_active = 0;
tx_active = 0;
@@ -297,42 +208,23 @@ static void finish(int sig)
}
}
}
if (!link_software)
if (use_uart_pty)
uart_pty_stop(&uart_pty);
_exit(0);
}
int main(int argc, char *argv[])
{
int link_given = 0;
for (int opt; (opt = getopt(argc, argv, "l:")) != -1;) {
if (opt != 'l' || parse_link(optarg) != 0) {
fprintf(stderr, "device: bad link spec (usart0, usart1, sw, or sw:B0,B1 as RX,TX)\n");
if (argc != 8) {
fprintf(stderr, "usage: %s <pureboot.elf> <mcu> <hz> <base_hex> <page> <baud> <flash_dump>\n", argv[0]);
return 2;
}
link_given = 1;
}
int args = argc - optind;
if (args < 7 || args > 9) {
fprintf(stderr,
"usage: %s [-l link] <pureboot.elf> <mcu> <hz> <base_hex> <page> <baud> <flash_dump>"
" [reset_hex] [resume_flash]\n"
" -l link: usart0 | usart1 | sw[:B0,B1] (RX,TX); default: the chip's own\n"
" reset_hex: reset vector (default: base with a boot section, else 0)\n"
" resume_flash: raw full-flash image loaded instead of the ELF — a prior\n"
" run's dump, for power-fail resume tests\n",
argv[0]);
return 2;
}
argv += optind - 1; // argv[1] is the ELF again, whatever was parsed
const char *mcu_name = argv[2];
uint32_t base = (uint32_t)strtoul(argv[4], NULL, 0);
unsigned page = (unsigned)atoi(argv[5]);
unsigned baud = (unsigned)atoi(argv[6]);
dump_path = argv[7];
int is_mega = strncmp(mcu_name, "atmega", 6) == 0;
if (!link_given)
link_software = !is_mega; // the chips' natural links: USART0, or PB0/PB1
use_uart_pty = strcmp(mcu_name, "atmega328p") == 0;
avr = avr_make_mcu_by_name(mcu_name);
if (!avr) {
@@ -343,29 +235,16 @@ int main(int argc, char *argv[])
avr->frequency = (uint32_t)strtoul(argv[3], NULL, 0);
memset(avr->flash, 0xff, avr->flashend + 1); // real flash powers up erased
if (args > 8) {
// Resume: the full flash image of an interrupted prior run.
FILE *f = fopen(argv[9], "rb");
if (!f || fread(avr->flash, 1, avr->flashend + 1, f) == 0) {
fprintf(stderr, "device: cannot read %s\n", argv[9]);
return 1;
}
fclose(f);
} else {
elf_firmware_t fw = {0};
if (elf_read_firmware(argv[1], &fw) != 0) {
fprintf(stderr, "device: cannot read %s\n", argv[1]);
return 1;
}
memcpy(avr->flash + base, fw.flash, fw.flashsize);
}
// The boot-sectioned megas enter the loader in hardware (BOOTRST, not
// modeled — the argument picks the modeled fuse's target); the tinies
// and the boot-section-less m48s reset to word 0 like silicon — erased
// flash walks up into the loader, and after the host's surgery the
// patched vector routes there.
int boot_section = is_mega && strncmp(mcu_name, "atmega48", 8) != 0;
reset_pc = args > 7 ? (uint32_t)strtoul(argv[8], NULL, 0) : (boot_section ? base : 0);
// The mega enters the loader in hardware (BOOTRST, not modeled); the
// tinies reset to word 0 like silicon — erased flash walks up into the
// loader, and after the host's surgery the patched vector routes there.
reset_pc = use_uart_pty ? base : 0;
avr->pc = reset_pc;
avr->codeend = avr->flashend;
@@ -378,34 +257,26 @@ int main(int argc, char *argv[])
avr_ioctl(avr, AVR_IOCTL_EEPROM_SET, &seed);
}
// The megas carry simavr's avr_flash module (and its two gaps the wrap
// above fixes); the tinies get the NVM module simavr lacks. Which serial
// bridge runs is the link's business, not the chip class's.
if (is_mega) {
fix_mega_flash_erase();
if (use_uart_pty) {
// POLL_SLEEP paces an idle-polling loader in host real time (a
// no-hardware CPU-saving hack); clear it so cycles run free.
uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS('0'), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS('0'), &flags);
uart_pty_init(avr, &uart_pty);
uart_pty_connect(&uart_pty, '0');
printf("PB_PTY %s\n", uart_pty.pty.slavename);
} else {
nvm.page = page;
memset(nvm.buffer, 0xff, sizeof(nvm.buffer));
nvm.io.kind = "tiny_nvm";
nvm.io.ioctl = nvm_ioctl;
avr_register_io(avr, &nvm.io);
}
if (!link_software) {
// POLL_SLEEP paces an idle-polling loader in host real time (a
// no-hardware CPU-saving hack); clear it so cycles run free.
uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
uart_pty_init(avr, &uart_pty);
uart_pty_connect(&uart_pty, uart_digit);
printf("PB_PTY %s\n", uart_pty.pty.slavename);
} else {
bit_cycles = (avr->frequency + baud / 2) / baud; // matches uart.hpp's own rounding exactly
rx_pin = avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_rx_port), (unsigned)sw_rx_bit);
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ(sw_tx_port), (unsigned)sw_tx_bit), tx_hook,
NULL);
rx_pin = avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ('B'), 0);
avr_irq_register_notify(avr_io_getirq(avr, AVR_IOCTL_IOPORT_GETIRQ('B'), 1), tx_hook, NULL);
avr_raise_irq(rx_pin, 1); // idle line
int slave;
@@ -433,27 +304,18 @@ int main(int argc, char *argv[])
reset_requested = 0;
avr_reset(avr);
avr->pc = reset_pc;
if (!link_software) { // reset restores the pacing hack; re-clear it
if (use_uart_pty) { // reset restores the pacing hack; re-clear it
uint32_t flags = 0;
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS(uart_digit), &flags);
avr_ioctl(avr, AVR_IOCTL_UART_GET_FLAGS('0'), &flags);
flags &= ~AVR_UART_FLAG_POLL_SLEEP;
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS(uart_digit), &flags);
avr_ioctl(avr, AVR_IOCTL_UART_SET_FLAGS('0'), &flags);
} else {
bridge_reset();
}
}
if (link_software && ++since_poll >= 2000) {
if (!use_uart_pty && ++since_poll >= 2000) {
since_poll = 0;
poll_pty();
// An unthrottled idle simulation runs the activation window out
// from under the host's real-time knock cadence: a 1 MHz build's
// 8 s window is 8 M cycles — tens of wall milliseconds — so a
// first knock lost to an in-flight reset misses the window
// entirely. Pace the simulation only while the bridge is fully
// quiet (nothing decoding, nothing queued); transfers keep full
// speed, and a quiet window stretches toward real time.
if (!rx_active && !tx_active && rx_head == rx_tail)
usleep(200);
}
}
finish(0);

View File

@@ -1,349 +0,0 @@
#!/usr/bin/env python3
"""Host-tool unit tests — the planning and policy logic, no simulator:
programming orders and their recovery properties, the reset-vector surgery,
the staging composition, the boot-fuse decode, and the update preflight over
fuse combinations simavr cannot model.
Usage: test_planner.py <tool_py>
"""
import sys
def fail(message):
print(f"FAIL: {message}")
sys.exit(1)
def expect_error(what, fn, *needles):
try:
fn()
except Exception as error:
for needle in needles:
if needle not in str(error):
fail(f"{what}: error lacks {needle!r}: {error}")
return
fail(f"{what}: no error raised")
def info_of(pb, base, page, patch, flash, signature=(0x1E, 0x93, 0x0B), word_flash=False, version=None):
scale = 2 if word_flash else 1
wire_base = base // scale
flags = (1 if patch else 0) | (2 if word_flash else 0)
raw = bytes((0x50, 0x42, pb.NEWEST_LOADER if version is None else version,
*signature, page & 0xFF, wire_base & 0xFF, wire_base >> 8, 0, 2, flags))
info = pb.Info(raw)
if info.flash_size != flash:
fail(f"info_of({base:#x}) decodes to {info.flash_size:#x} of flash, not {flash:#x}")
return info
def rjmp_decode(word, at, flash_words):
if word & 0xF000 != 0xC000:
fail(f"not an rjmp: {word:#06x}")
offset = word & 0x0FFF
if offset >= 0x800:
offset -= 0x1000
return (at + 1 + offset) % flash_words
def main():
import os
sys.path.insert(0, os.path.dirname(os.path.abspath(sys.argv[1])))
import pureboot as pb
tiny = info_of(pb, 0x1E00, 64, True, 0x2000)
mega = info_of(pb, 0x7E00, 128, False, 0x8000, signature=(0x1E, 0x95, 0x0F))
# Versioning: the block's third byte is the loader's version, and the tool
# speaks a window of them. Every version in the window decodes, so an older
# deployed loader stays usable; one above the window is refused by name,
# since which version changed the protocol is knowledge only the tool
# holds, and it holds none about a version it has never heard of.
for version in range(pb.OLDEST_LOADER, pb.NEWEST_LOADER + 1):
if info_of(pb, 0x1E00, 64, True, 0x2000, version=version).version != version:
fail(f"pureboot {version} does not decode")
expect_error(
"unknown loader version",
lambda: info_of(pb, 0x1E00, 64, True, 0x2000, version=pb.NEWEST_LOADER + 1),
f"pureboot {pb.NEWEST_LOADER + 1}",
"newer tool",
)
# mega_boot: BOOTSZ words and the BOOTRST sense per chip — the fuse byte
# index (EXTENDED on the x8 line except the m328s' HIGH, HIGH elsewhere)
# and the per-family ladders (Atmel-2486/2466/2503/2545/8271/DS40002065/
# 8272/8011/2593/42719). Synthetic 'F' replies: only the boot byte
# carries meaning.
cases = (
((0x1E, 0x93, 0x07), 0x2000, 3, {0b11: 0x1F00, 0b10: 0x1E00, 0b01: 0x1C00, 0b00: 0x1800}), # m8
((0x1E, 0x94, 0x03), 0x4000, 3, {0b11: 0x3F00, 0b10: 0x3E00, 0b01: 0x3C00, 0b00: 0x3800}), # m16
((0x1E, 0x95, 0x02), 0x8000, 3, {0b11: 0x7E00, 0b10: 0x7C00, 0b01: 0x7800, 0b00: 0x7000}), # m32
((0x1E, 0x93, 0x0A), 0x2000, 2, {0b11: 0x1F00, 0b10: 0x1E00, 0b01: 0x1C00, 0b00: 0x1800}), # m88
((0x1E, 0x93, 0x0F), 0x2000, 2, {0b11: 0x1F00, 0b10: 0x1E00, 0b01: 0x1C00, 0b00: 0x1800}), # m88P
((0x1E, 0x94, 0x06), 0x4000, 2, {0b11: 0x3F00, 0b10: 0x3E00, 0b01: 0x3C00, 0b00: 0x3800}), # m168/168A
((0x1E, 0x94, 0x0B), 0x4000, 2, {0b11: 0x3F00, 0b10: 0x3E00, 0b01: 0x3C00, 0b00: 0x3800}), # m168P
((0x1E, 0x95, 0x14), 0x8000, 3, {0b11: 0x7E00, 0b10: 0x7C00, 0b01: 0x7800, 0b00: 0x7000}), # m328
((0x1E, 0x95, 0x0F), 0x8000, 3, {0b11: 0x7E00, 0b10: 0x7C00, 0b01: 0x7800, 0b00: 0x7000}), # m328P
((0x1E, 0x94, 0x0F), 0x4000, 3, {0b11: 0x3F00, 0b10: 0x3E00, 0b01: 0x3C00, 0b00: 0x3800}), # m164A
((0x1E, 0x94, 0x0A), 0x4000, 3, {0b11: 0x3F00, 0b10: 0x3E00, 0b01: 0x3C00, 0b00: 0x3800}), # m164P
((0x1E, 0x95, 0x15), 0x8000, 3, {0b11: 0x7E00, 0b10: 0x7C00, 0b01: 0x7800, 0b00: 0x7000}), # m324A
((0x1E, 0x96, 0x09), 0x10000, 3, {0b11: 0xFC00, 0b10: 0xF800, 0b01: 0xF000, 0b00: 0xE000}), # m644
((0x1E, 0x96, 0x0A), 0x10000, 3, {0b11: 0xFC00, 0b10: 0xF800, 0b01: 0xF000, 0b00: 0xE000}), # m644P
((0x1E, 0x97, 0x06), 0x20000, 3, {0b11: 0x1FC00, 0b10: 0x1F800, 0b01: 0x1F000, 0b00: 0x1E000}), # 1284
((0x1E, 0x97, 0x05), 0x20000, 3, {0b11: 0x1FC00, 0b10: 0x1F800, 0b01: 0x1F000, 0b00: 0x1E000}), # 1284P
)
for signature, flash, which, ladder in cases:
chip = info_of(pb, flash - pb.SLOT, 128 if flash < 0x20000 else 0, False, flash,
signature=signature, word_flash=flash > 0x10000)
for bits, start in ladder.items():
fuses = bytearray((0xFF, 0xFF, 0xFF, 0xFF))
fuses[which] = (0xF8 | (bits << 1)) & ~1
prog, at = pb.mega_boot(chip, bytes(fuses))
if not prog or at != start:
fail(f"mega_boot {signature[1]:02x}{signature[2]:02x} BOOTSZ={bits:02b} programmed: {prog} {at:#07x}")
fuses[which] |= 1
prog, at = pb.mega_boot(chip, bytes(fuses))
if prog or at != start:
fail(f"mega_boot {signature[1]:02x}{signature[2]:02b} unprogrammed: {prog} {at:#07x}")
# Word-addressed info decode: the 1284P's base and page ride the wire
# scaled — a 17-bit base halved into the block's two bytes, a 256-byte page
# spelled 0 — and its slot is the same 512 bytes as everywhere else, so its
# staging slot lands inside the 1 KiB minimum boot section.
big = info_of(pb, 0x1FE00, 0, False, 0x20000, signature=(0x1E, 0x97, 0x05), word_flash=True)
if big.page != 256 or big.base != 0x1FE00 or big.stage != 0x1FC00:
fail(f"word-addressed info decode: page {big.page}, base {big.base:#x}, stage {big.stage:#x}")
# Surgery: word 0 lands on the loader, the trampoline on the original
# entry — checked with an independent decoder.
app = bytes((0xC0 | 0x00, 0xC0)) + bytes((0x12,)) * 300 # rjmp .+0x00C0... entry word 0xC0C0
entry = rjmp_decode(app[0] | (app[1] << 8), 0, tiny.flash_size // 2)
pages = pb.plan_flash(app, tiny)
word0 = pages[0][0] | (pages[0][1] << 8)
if rjmp_decode(word0, 0, tiny.flash_size // 2) != tiny.base // 2:
fail("surgery: patched word 0 misses the loader")
tp = pages[tiny.base - 64]
tramp = tp[62] | (tp[63] << 8)
if rjmp_decode(tramp, (tiny.base - 2) // 2, tiny.flash_size // 2) != entry:
fail("surgery: trampoline misses the original entry")
expect_error("non-rjmp vector", lambda: pb.plan_flash(bytes((0x0C, 0x94)) + app[2:], tiny), "not an rjmp")
looped = bytearray(app)
word = pb.rjmp_to(0, tiny.base // 2, tiny.flash_size // 2)
looped[0], looped[1] = word & 0xFF, word >> 8
expect_error("read-back image", lambda: pb.plan_flash(bytes(looped), tiny), "read-back")
expect_error("oversize image", lambda: pb.plan_flash(bytes(0x1DFF), tiny), "application flash ends")
# Ordering: patched vector puts page 0 first and the trampoline second;
# a boot section puts page 0 last. Blank pages drop only when erased.
order = pb.covered(pages, tiny, skip_blank=False)
if order[0] != 0 or order[1] != tiny.base - 64:
fail(f"tiny order starts {order[:2]}, want page 0 then trampoline page")
if sorted(order[2:]) != order[2:]:
fail("tiny order tail not ascending")
mega_pages = pb.plan_flash(bytes((0xFF,)) * 600, mega)
morder = pb.covered(mega_pages, mega, skip_blank=False)
if morder[-1] != 0 or sorted(morder[:-1]) != morder[:-1]:
fail(f"mega order {morder}, want ascending with page 0 last")
blanky = {0: pages[0], 64: bytes((0xFF,)) * 64, 128: pages[128], tiny.base - 64: tp}
slim = pb.covered(blanky, tiny, skip_blank=True)
if 64 in slim or 0 not in slim or tiny.base - 64 not in slim:
fail(f"skip_blank order wrong: {slim}")
# Staging content: the identical image plus the through-word on a
# patched-vector chip; hard size clamps either way.
image = bytes(range(256)) * 2 # 512 B — too big for a tiny slot
expect_error("tiny staging size", lambda: pb.staging_content(image, tiny), "510")
staged = pb.staging_content(image[:508], tiny)
through = staged[510] | (staged[511] << 8)
if rjmp_decode(through, (tiny.base - 2) // 2, tiny.flash_size // 2) != tiny.base // 2:
fail("through-word misses the resident base")
if pb.staging_content(image, mega) != image:
fail("mega staging content should be the bare image")
expect_error("mega staging size", lambda: pb.staging_content(image + b"!", mega), "512")
# The embedded info block: found in a synthetic binary, absent in noise.
binary = bytes((0xAA,)) * 10 + tiny.raw + bytes((0xBB,)) * 10
found = pb.image_info(binary)
if found is None or found.raw != tiny.raw:
fail("image_info misses the embedded block")
if pb.image_info(bytes((0xAA,)) * 40) is not None:
fail("image_info invents a block")
# An older loader's image stays readable, so a deployed build can be
# identified and installed like any other.
old = info_of(pb, 0x1E00, 64, True, 0x2000, version=pb.OLDEST_LOADER)
found_old = pb.image_info(bytes((0xAA,)) * 10 + old.raw)
if found_old is None or found_old.version != pb.OLDEST_LOADER:
fail("image_info misses an older loader's block")
# loader_image must peel a padded image down to the slot content: a raw
# .bin padded from address 0 (or a whole-flash read-back with the loader
# resident at base) yields the same bytes as the bare slot image.
import tempfile
slot_image = bytes((0xAA,)) * 10 + tiny.raw + bytes((0xCC,)) * 40
padded = bytes((0xFF,)) * tiny.base + slot_image
with tempfile.NamedTemporaryFile(suffix=".bin", delete=False) as f:
f.write(padded)
padded_path = f.name
try:
if pb.loader_image(padded_path) != slot_image:
fail("loader_image does not peel a padded image to the slot content")
finally:
os.unlink(padded_path)
# Update preflight: the full fuse matrix, plus target mismatch.
other = info_of(pb, 0x1E00, 32, True, 0x2000)
expect_error("wrong-target image", lambda: pb.update_preflight(binary, other, None), "another target")
expect_error("mega needs fuses", lambda: pb.update_preflight(bytes((0xAA,)) * 8 + mega.raw, mega, None),
"--assume-fuses")
mega_image = bytes((0xAA,)) * 8 + mega.raw
def fuses(high):
return bytes((0xFF, 0xFF, 0xFF, high))
expect_error("BOOTSZ 512 B", lambda: pb.update_preflight(mega_image, mega, fuses(0xFE)),
"cannot self-update", "BOOTSZ")
notes = pb.update_preflight(mega_image, mega, fuses(0xFD)) # 1 KB, BOOTRST unprogrammed
if not any("BOOTRST unprogrammed" in n for n in notes):
fail(f"1K/unprogrammed notes: {notes}")
notes = pb.update_preflight(mega_image, mega, fuses(0xFC)) # 1 KB, BOOTRST programmed
if not any("staging slot" in n for n in notes):
fail(f"1K/programmed notes: {notes}")
notes = pb.update_preflight(mega_image, mega, fuses(0xFA)) # 2 KB, BOOTRST programmed
if not any("application flash" in n for n in notes):
fail(f"2K/programmed notes: {notes}")
if pb.update_preflight(bytes((0xAA,)) * 8 + tiny.raw, tiny, None) != []:
fail("tiny preflight should pass without fuses")
# The 1284s' smallest boot section (512 words) is exactly the resident
# slot plus its staging slot, so self-update is possible at the minimum
# BOOTSZ — no fuse step up, the 644's geometry. That holds only while a
# slot is 512 B: at 1 KiB the staging slot would fall outside the section
# and the preflight would refuse.
notes = pb.update_preflight(bytes((0xAA,)) * 8 + big.raw, big, fuses(0xFE))
if not any("staging slot" in n for n in notes):
fail(f"1284 minimum-BOOTSZ notes: {notes}")
# The walk-region refusal: BOOTRST aimed below the loader plus app data
# in the walk span errors without --force; erased spans and unprogrammed
# BOOTRST pass.
deep = {0x7800: bytes((1,)) * 128}
expect_error("walk region", lambda: pb.check_walk_region(deep, mega, fuses(0xFA), False), "--force")
pb.check_walk_region(deep, mega, fuses(0xFA), True)
pb.check_walk_region(deep, mega, fuses(0xFB), False) # BOOTRST unprogrammed
pb.check_walk_region({0x7800: bytes((0xFF,)) * 128}, mega, fuses(0xFA), False)
pb.check_walk_region(deep, mega, None, False) # fuses unknown: no check
# The repairing verify: a mismatched page is rewritten rather than raised,
# bounded so a fault that is not self-clearing cannot spin.
class FakeLoader:
"""A device whose first `bad` writes of any page land wrong."""
def __init__(self, info, bad):
self.info = info
self.bad = bad
self.flash = {}
self.writes = 0
def write_page(self, address, data):
self.writes += 1
self.flash[address] = bytes(len(data)) if self.bad > 0 else bytes(data)
self.bad -= 1
def read_flash(self, address, count):
return self.flash.get(address, bytes(count))
want = {0: bytes((i * 5) & 0xFF for i in range(128))}
# One bad write, then good: repaired in place, and the caller never sees
# an error. The rewrite is counted, so a silent no-op cannot pass.
device = FakeLoader(info_of(pb, 0x7E00, 128, False, 0x8000), bad=1)
device.write_page(0, want[0])
pb.verify_pages(device, want, repair=True)
if device.writes != 2:
fail(f"repairing verify made {device.writes} writes, expected 2")
# Without repair the same state raises, so the repair is what fixed it.
device = FakeLoader(info_of(pb, 0x7E00, 128, False, 0x8000), bad=1)
device.write_page(0, want[0])
expect_error("verify without repair", lambda: pb.verify_pages(device, want), "verify failed")
# A page that never comes good stops after RETRIES rewrites, and says so.
device = FakeLoader(info_of(pb, 0x7E00, 128, False, 0x8000), bad=99)
device.write_page(0, want[0])
expect_error(
"unrepairable page",
lambda: pb.verify_pages(device, want, repair=True),
"verify failed",
f"after {pb.RETRIES} retries",
)
if device.writes != pb.RETRIES + 1:
fail(f"unrepairable page took {device.writes} writes, expected {pb.RETRIES + 1}")
# The knock handshake against a device that is not listening yet — the
# state a port open leaves behind: it resets the chip into a fresh
# activation window while the previous session's prompt is still in
# flight, so the first knock is lost and a prompt arrives anyway.
class FakePort:
"""A loader in its activation window, plus `lost` leading writes the
reset swallows and one stale prompt still on the wire."""
def __init__(self, info_raw, lost=0, stale=b"", active=False):
self.info_raw = info_raw
self.lost = lost
self.inflight = bytearray(stale)
self.rx = bytearray()
self.active = active
self.last = None
def flush_input(self):
self.rx.clear()
def write(self, data):
if self.lost:
self.lost -= 1
return
for byte in bytes(data):
if not self.active:
self.active = self.last == ord("p") and byte == ord("b")
self.last = byte
if self.active:
self.rx += pb.PROMPT
elif byte == ord("b"):
self.rx += self.info_raw + pb.PROMPT
else:
self.rx += pb.PROMPT
def read_available(self, wait):
self.rx = self.inflight + self.rx # the stale prompt lands late
self.inflight.clear()
out, self.rx = bytes(self.rx), bytearray()
return out
def read_exact(self, count, timeout):
if len(self.rx) < count:
raise pb.Error(f"timeout: got {len(self.rx)} of {count} bytes")
out, self.rx = bytes(self.rx[:count]), self.rx[count:]
return out
raw = info_of(pb, 0x7E00, 128, False, 0x8000).raw
for what, port in (
("clean window", FakePort(raw)),
("stale prompt over a lost knock", FakePort(raw, lost=1, stale=pb.PROMPT)),
("live session", FakePort(raw, active=True)),
):
info = pb.Loader(port).connect(5)
if info.raw != raw:
fail(f"connect ({what}) returned {info.raw.hex()}")
# A device that never answers still says so, and a version the tool cannot
# speak is reported as such rather than retried into a timeout.
expect_error("dead device", lambda: pb.Loader(FakePort(raw, lost=99)).connect(0), "no answer")
old = bytes(raw[:2]) + bytes((pb.NEWEST_LOADER + 1,)) + bytes(raw[3:])
expect_error("unspeakable version", lambda: pb.Loader(FakePort(old)).connect(5), "needs a newer tool")
print("test_planner: all planner and policy checks pass")
if __name__ == "__main__":
main()

View File

@@ -1,39 +0,0 @@
#!/bin/bash
# The port's gate: every chip's generated workflow — build, size matrix, and
# the simulator-driven protocol suites. --full adds the reflect-spot builds
# (libavr's rule: reflect compiles are bounded to its spot set, never the
# full matrix) and swaps the compact size matrix for the exhaustive
# clock × baud × backend cross product. LIBAVR_ROOT must point at the libavr
# checkout.
set -e
cd "$(dirname "$0")/.."
full=0
[[ "$1" == "--full" ]] && { full=1; shift; export PUREBOOT_FULL_MATRIX=1; }
CHIPS=(attiny13 attiny13a attiny25 attiny45 attiny85
atmega8 atmega8a atmega16 atmega16a atmega32 atmega32a
atmega48 atmega48a atmega48p atmega48pa
atmega88 atmega88a atmega88p atmega88pa
atmega168 atmega168a atmega168p atmega168pa
atmega328 atmega328p
atmega164a atmega164p atmega164pa
atmega324a atmega324p atmega324pa
atmega644 atmega644a atmega644p atmega644pa
atmega1284 atmega1284p)
REFLECT_SPOT=(attiny13a attiny85 atmega8 atmega16a atmega32a atmega48pa
atmega88 atmega168pa atmega328p atmega164a atmega644p atmega1284)
for chip in "${CHIPS[@]}"; do
echo "==== $chip ===="
cmake --workflow --preset "$chip-generated" "$@"
done
if ((full)); then
for chip in "${REFLECT_SPOT[@]}"; do
echo "==== $chip reflect ===="
cmake --workflow --preset "$chip-reflect" "$@"
done
fi
echo "check: every chip green"

View File

@@ -1,90 +0,0 @@
#!/usr/bin/env python3
"""Regenerate CMakePresets.json — one uniform pipeline per chip.
Every chip gets generated-mode configure/build/test presets and a workflow
running all three. Reflect-mode presets (configure + build, no tests — the
port's TUs compile identically; the sims prove nothing new there) exist for
libavr's reflect spot set only, mirroring its rule: the full reflect matrix
is never built, one chip per hardware class and pack vintage is.
Run from the repo root: tools/make_presets.py
"""
import json
import os
CHIPS = [
"attiny13", "attiny13a", "attiny25", "attiny45", "attiny85",
"atmega8", "atmega8a", "atmega16", "atmega16a", "atmega32", "atmega32a",
"atmega48", "atmega48a", "atmega48p", "atmega48pa",
"atmega88", "atmega88a", "atmega88p", "atmega88pa",
"atmega168", "atmega168a", "atmega168p", "atmega168pa",
"atmega328", "atmega328p",
"atmega164a", "atmega164p", "atmega164pa",
"atmega324a", "atmega324p", "atmega324pa",
"atmega644", "atmega644a", "atmega644p", "atmega644pa",
"atmega1284", "atmega1284p",
]
# libavr's REFLECT_SPOT (tools/check.sh): one chip per hardware class and
# pack vintage.
REFLECT_SPOT = [
"attiny13a", "attiny85", "atmega8", "atmega16a", "atmega32a",
"atmega48pa", "atmega88", "atmega168pa", "atmega328p", "atmega164a",
"atmega644p", "atmega1284",
]
def main():
configure = [{
"name": "base",
"hidden": True,
"generator": "Ninja",
"binaryDir": "${sourceDir}/build/${presetName}",
"toolchainFile": "$env{LIBAVR_ROOT}/cmake/avr-toolchain.cmake",
"cacheVariables": {
"CMAKE_BUILD_TYPE": "Release",
"CMAKE_EXPORT_COMPILE_COMMANDS": "ON",
"CMAKE_COLOR_DIAGNOSTICS": "ON",
},
}]
build, test, workflows = [], [], []
def add(chip, mode):
name = f"{chip}-{mode}"
configure.append({
"name": name,
"inherits": "base",
"cacheVariables": {
"LIBAVR_MCU": chip,
"LIBAVR_REFLECT": "ON" if mode == "reflect" else "OFF",
},
})
build.append({"name": name, "configurePreset": name})
steps = [{"type": "configure", "name": name}, {"type": "build", "name": name}]
if mode == "generated":
test.append({"name": name, "configurePreset": name, "output": {"outputOnFailure": True}})
steps.append({"type": "test", "name": name})
workflows.append({"name": name, "steps": steps})
for chip in CHIPS:
add(chip, "generated")
for chip in REFLECT_SPOT:
add(chip, "reflect")
presets = {
"version": 8,
"configurePresets": configure,
"buildPresets": build,
"testPresets": test,
"workflowPresets": workflows,
}
path = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "CMakePresets.json")
with open(path, "w") as f:
json.dump(presets, f, indent=1)
f.write("\n")
print(f"{len(CHIPS)} chips, {len(REFLECT_SPOT)} reflect: {os.path.normpath(path)}")
if __name__ == "__main__":
main()