pureboot: the 1284P rides a 1 KiB slot — its own boot-sector minimum
The far machinery (ELPM reads, RAMPZ page commands, wire-word math) costs ~46 B over the m328P's 504, and the tsb-calibrated C++-to-asm gap says no implementation of this feature set reaches 512 on this chip — a boundary its hardware does not have anyway: the 1284P's smallest boot sector is 1 KiB. The slot therefore becomes per-geometry (512 B, or 1 KiB past 64 KiB), which the host derives from the word-addressing flag; slot arithmetic unifies (the index is the wire high byte with its low bit dropped in either unit), the update preflight demands a two-slot boot section in the chip's own terms, and pbapp's hand-back jumps to the real slot base. libavr's far primitives split their RAMPZ/Z asm operands (a page never crosses 64 KiB, so callers keep a byte and a 16-bit cursor — the 32-bit address folds away; flash_load_far's byte form becomes the out-RAMPZ+elpm pair avr-libc's pgm_read_byte_far rebuilds per call), and the host splits reads at 64 KiB boundaries. All ten chips pass the full suite — the 1284P at 558 B including protocol, relocation, and the power-fail self-update — with pureboot byte-identical across generated and reflect modes everywhere, and the original three chips' images unchanged to the byte (488/502/504). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -2,22 +2,28 @@
|
||||
|
||||
A serial bootloader on [libavr](https://git.blackmark.me/avr/libavr), pure by
|
||||
constraint: one C++ source, no inline assembly, no global register variables
|
||||
(attributes allowed), built for every chip libavr targets, **512 bytes on
|
||||
each** — 488 B on the ATtiny13A, 502 B on the ATtiny85, 504 B on the
|
||||
ATmega328P. The device speaks primitives; every composite — verify, erase,
|
||||
(attributes allowed), built for every chip libavr targets, **fitting each
|
||||
chip's smallest boot sector**: 512 bytes everywhere — 488 B on the
|
||||
ATtiny13A, 502 B on the ATtiny85, 466–504 B across the megas — except the
|
||||
ATmega1284P, whose smallest boot sector is 1 KiB and whose far-flash
|
||||
machinery (ELPM reads, RAMPZ page commands, word-addressed wire) lands at
|
||||
558 B in a 1 KiB slot: the 512-byte figure is a hardware boundary that chip
|
||||
simply does not have, and no implementation of this feature set fits it
|
||||
there. The device speaks primitives; every composite — verify, erase,
|
||||
reset-vector surgery, updating the loader itself — lives in the host tool
|
||||
(`pureboot.py`).
|
||||
|
||||
The image is **position-independent**: control flow is PC-relative, the
|
||||
read/write paths take wire addresses, the write guard protects the 512-byte
|
||||
slot the code is *running* in (from the runtime return address), the info
|
||||
block is addressed from that same anchor, and the application jump is an
|
||||
indirect call to an absolute entry. The identical binary therefore runs from
|
||||
any 512-byte slot with every command intact — which makes pureboot **its own
|
||||
staging loader**: the host installs the same binary one slot below the
|
||||
resident, jumps into it, and lets it rewrite the resident. On the tinies the
|
||||
budget is 510, not 512: a slot's last word belongs to the host-managed
|
||||
trampoline (below).
|
||||
read/write paths take wire addresses, the write guard protects the slot the
|
||||
code is *running* in (from the runtime return address), the info block is
|
||||
addressed from that same anchor, and the application jump is an indirect
|
||||
call to an absolute entry. The identical binary therefore runs from any
|
||||
slot with every command intact — which makes pureboot **its own staging
|
||||
loader**: the host installs the same binary one slot below the resident,
|
||||
jumps into it, and lets it rewrite the resident. The slot is 512 bytes
|
||||
(1 KiB on the word-addressed large chips, matching their boot-sector
|
||||
minimum); on the tinies the budget is 510, not 512: a slot's last word
|
||||
belongs to the host-managed trampoline (below).
|
||||
|
||||
## Link
|
||||
|
||||
|
||||
@@ -60,13 +60,16 @@ consteval std::int16_t wdrf_field()
|
||||
return avr::hw::db.field_index(reg, "WDRF");
|
||||
}
|
||||
|
||||
// Geometry: the resident loader owns the top 512 bytes of flash; the word
|
||||
// below it is the trampoline (the application's relocated reset vector) on
|
||||
// chips without a hardware boot section. The RWWSRE bit marks a separate
|
||||
// boot section — on classic AVR the two capabilities coincide (the m8/m32
|
||||
// packs spell its register SPMCR).
|
||||
constexpr std::uint16_t boot_bytes = 512;
|
||||
constexpr std::uint32_t base = spm::flash_bytes - boot_bytes;
|
||||
// Geometry: the resident loader owns the top slot of flash — 512 bytes,
|
||||
// except on the >64 KiB chips whose own smallest boot sector is 1 KiB (the
|
||||
// 1284P): there the slot is 1 KiB, matching the hardware boundary the
|
||||
// 512-byte figure comes from everywhere else. The word below the slot is
|
||||
// the trampoline (the application's relocated reset vector) on chips
|
||||
// without a hardware boot section. The RWWSRE bit marks a separate boot
|
||||
// section — on classic AVR the two capabilities coincide (the m8/m32 packs
|
||||
// spell its register SPMCR).
|
||||
constexpr std::uint16_t slot_bytes = spm::flash_bytes > 65536 ? 1024 : 512;
|
||||
constexpr std::uint32_t base = spm::flash_bytes - slot_bytes;
|
||||
constexpr std::uint16_t page = spm::page_bytes;
|
||||
constexpr bool boot_section = [] {
|
||||
for (auto reg : {"SPMCSR", "SPMCR"})
|
||||
@@ -77,8 +80,10 @@ constexpr bool boot_section = [] {
|
||||
|
||||
// Past 64 KiB a byte address no longer fits the wire's 16 bits, so on the
|
||||
// large chips every flash address on the wire — and all slot arithmetic —
|
||||
// is a word address instead ('J' always was one). Slots stay 512 bytes =
|
||||
// 256 words, so a slot is one high byte in either unit.
|
||||
// is a word address instead ('J' always was one). A slot spans the same
|
||||
// wire-high-byte pair in either unit (512 B = 2 x 256 bytes, 1 KiB =
|
||||
// 2 x 256 words), so the slot index is the high byte with its low bit
|
||||
// dropped everywhere.
|
||||
constexpr bool word_flash = spm::flash_bytes > 65536;
|
||||
constexpr std::uint16_t wire_base = word_flash ? static_cast<std::uint16_t>(base / 2) : static_cast<std::uint16_t>(base);
|
||||
constexpr std::uint16_t wire_page_mask = word_flash ? (page / 2 - 1) : (page - 1);
|
||||
@@ -277,10 +282,18 @@ std::uint16_t rx16()
|
||||
[[gnu::noinline]] void send_flash(std::uint16_t address, std::uint8_t count)
|
||||
{
|
||||
if constexpr (word_flash) {
|
||||
auto byte_address = static_cast<std::uint32_t>(address) << 1;
|
||||
do
|
||||
link::tx(avr::flash_load_far<std::uint8_t>(byte_address++));
|
||||
while (--count);
|
||||
// The 24-bit cursor as the machine holds it: the RAMPZ byte and a
|
||||
// 16-bit Z, carried explicitly (the reassembled 32-bit address
|
||||
// folds away inside the inlined far load). A single read never
|
||||
// crosses a 64 KiB boundary — the protocol forbids it and the host
|
||||
// splits its chunks there — so RAMPZ holds for the whole run.
|
||||
std::uint8_t rampz = static_cast<std::uint8_t>(address >> 15);
|
||||
std::uint16_t z = static_cast<std::uint16_t>(address << 1);
|
||||
do {
|
||||
link::tx(avr::flash_load_far<std::uint8_t>((static_cast<std::uint32_t>(rampz) << 16) | z));
|
||||
if (++z == 0)
|
||||
++rampz; // robustness for a host that reads across 64 KiB
|
||||
} while (--count);
|
||||
} else {
|
||||
do
|
||||
link::tx(avr::flash_load(reinterpret_cast<const std::uint8_t *>(address++)));
|
||||
@@ -335,16 +348,22 @@ void program_flash(std::uint16_t wire_address, std::uint8_t slot_high)
|
||||
spm::flash_address_t address;
|
||||
std::uint8_t page_high;
|
||||
if constexpr (word_flash) {
|
||||
address = static_cast<spm::flash_address_t>(static_cast<std::uint32_t>(wire_address) << 1);
|
||||
const auto start = address;
|
||||
// Pages are aligned, so one page never crosses a 64 KiB boundary:
|
||||
// RAMPZ is a per-page constant and the fill cursor is a 16-bit Z
|
||||
// whose low byte is the whole in-page offset (256-byte pages). The
|
||||
// slot index is simply the wire word address's high byte.
|
||||
const std::uint8_t rampz = static_cast<std::uint8_t>(wire_address >> 15);
|
||||
const std::uint16_t z0 = static_cast<std::uint16_t>(wire_address << 1);
|
||||
std::uint16_t z = z0;
|
||||
do {
|
||||
std::uint8_t low = link::rx();
|
||||
std::uint8_t high = link::rx();
|
||||
spm::fill<off>(address, static_cast<std::uint16_t>(low | (high << 8)));
|
||||
address += 2;
|
||||
} while (static_cast<std::uint8_t>(address));
|
||||
address = start;
|
||||
page_high = static_cast<std::uint8_t>(static_cast<std::uint16_t>(address >> 8) >> 1);
|
||||
spm::fill<off>((static_cast<spm::flash_address_t>(rampz) << 16) | z,
|
||||
static_cast<std::uint16_t>(low | (high << 8)));
|
||||
z += 2;
|
||||
} while (static_cast<std::uint8_t>(z));
|
||||
address = (static_cast<spm::flash_address_t>(rampz) << 16) | z0;
|
||||
page_high = static_cast<std::uint8_t>(wire_address >> 8) & 0xfe;
|
||||
} else {
|
||||
address = static_cast<spm::flash_address_t>(wire_address);
|
||||
do {
|
||||
@@ -397,8 +416,8 @@ void send_fuses()
|
||||
// program_flash refuses this one slot and the info block is addressed
|
||||
// from it, so both follow wherever the code was flashed.
|
||||
const std::uint16_t ra_words = reinterpret_cast<std::uint16_t>(__builtin_return_address(0));
|
||||
const std::uint8_t slot_high =
|
||||
word_flash ? static_cast<std::uint8_t>(ra_words >> 8) : static_cast<std::uint8_t>((ra_words >> 8) << 1);
|
||||
const std::uint8_t slot_high = word_flash ? static_cast<std::uint8_t>(ra_words >> 8) & 0xfe
|
||||
: static_cast<std::uint8_t>((ra_words >> 8) << 1);
|
||||
|
||||
// The knock: 'p' then 'b', each under a fresh window; any other byte is
|
||||
// line noise and waits again. Falling out of a window runs the app.
|
||||
|
||||
@@ -36,7 +36,7 @@ else:
|
||||
|
||||
PROMPT = b"+"
|
||||
PROTOCOL_VERSION = 1
|
||||
SLOT = 512 # the loader slot size; also the self-update staging distance
|
||||
SLOT = 512 # the loader slot on byte-addressed chips; word-addressed ones (>64 KiB) use 1 KiB — their own smallest boot sector
|
||||
|
||||
|
||||
class Error(Exception):
|
||||
@@ -278,8 +278,9 @@ class Info:
|
||||
scale = 2 if self.word_flash else 1
|
||||
self.base = (raw[7] | (raw[8] << 8)) * scale
|
||||
self.eeprom_size = raw[9] | (raw[10] << 8)
|
||||
self.flash_size = self.base + SLOT
|
||||
self.stage = self.base - SLOT # where a staging copy of the loader goes
|
||||
self.slot = 1024 if self.word_flash else SLOT
|
||||
self.flash_size = self.base + self.slot
|
||||
self.stage = self.base - self.slot # where a staging copy of the loader goes
|
||||
# The hand-over target, as the word address 'J' takes: the trampoline
|
||||
# below the loader (tinies), or word 0 (mega — the application's own
|
||||
# reset vector; BOOTRST re-vectors a reset into the loader instead).
|
||||
@@ -349,10 +350,18 @@ class Loader:
|
||||
def read_flash(self, address, count):
|
||||
if not self.info.word_flash:
|
||||
return self._stream_read("R", address, count)
|
||||
# Word-addressed wire: widen to even bounds, read, trim.
|
||||
# Word-addressed wire: widen to even bounds and never let one read
|
||||
# cross a 64 KiB boundary (the device holds RAMPZ for a whole run).
|
||||
start = address & ~1
|
||||
span = (address + count + 1 & ~1) - start
|
||||
data = self._stream_read("R", start, span, address_scale=2)
|
||||
data = b""
|
||||
at = start
|
||||
remaining = span
|
||||
while remaining:
|
||||
chunk = min(remaining, 0x10000 - (at & 0xFFFF))
|
||||
data += self._stream_read("R", at, chunk, address_scale=2)
|
||||
at += chunk
|
||||
remaining -= chunk
|
||||
return data[address - start : address - start + count]
|
||||
|
||||
def read_eeprom(self, address, count):
|
||||
@@ -564,12 +573,13 @@ def staging_content(image, info):
|
||||
that word, which for a staging copy is the slot's own last word: an rjmp
|
||||
to the resident base. The staging copy's fall-through and 'J'-free exit
|
||||
both land in a loader instead of garbage."""
|
||||
if len(image) > (SLOT - 2 if info.patch_vector else SLOT):
|
||||
raise Error(f"loader image is {len(image)} B, the slot holds {SLOT - 2 if info.patch_vector else SLOT}")
|
||||
content = bytearray(image) + bytearray([0xFF] * (SLOT - len(image)))
|
||||
slot = info.slot
|
||||
if len(image) > (slot - 2 if info.patch_vector else slot):
|
||||
raise Error(f"loader image is {len(image)} B, the slot holds {slot - 2 if info.patch_vector else slot}")
|
||||
content = bytearray(image) + bytearray([0xFF] * (slot - len(image)))
|
||||
if info.patch_vector:
|
||||
through = rjmp_to((info.base - 2) // 2, info.base // 2, info.flash_size // 2)
|
||||
content[SLOT - 2], content[SLOT - 1] = through & 0xFF, through >> 8
|
||||
content[slot - 2], content[slot - 1] = through & 0xFF, through >> 8
|
||||
return bytes(content)
|
||||
|
||||
|
||||
@@ -592,8 +602,8 @@ def update_preflight(image, info, fuse_bytes):
|
||||
raise Error(
|
||||
f"cannot self-update: the staging slot {info.stage:#06x} lies below the "
|
||||
f"boot section ({bls_start:#06x}) where SPM is disabled "
|
||||
f"— a boot section of at least 1 KB (BOOTSZ) is required, and only an "
|
||||
f"external programmer can change fuses"
|
||||
f"— a boot section of at least two slots ({2 * info.slot} B, BOOTSZ) is "
|
||||
f"required, and only an external programmer can change fuses"
|
||||
)
|
||||
if not bootrst:
|
||||
warnings.append(
|
||||
@@ -634,7 +644,7 @@ class UpdateState:
|
||||
self.data = {
|
||||
"signature": info.signature.hex(),
|
||||
"base": info.base,
|
||||
"staging": loader.read_flash(info.stage, SLOT).hex(),
|
||||
"staging": loader.read_flash(info.stage, info.slot).hex(),
|
||||
"page0": loader.read_flash(0, info.page).hex() if info.patch_vector else "",
|
||||
}
|
||||
with open(self.path, "w") as f:
|
||||
@@ -690,7 +700,7 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes):
|
||||
for warning in update_preflight(image, info, fuse_bytes):
|
||||
print(f"note: {warning}")
|
||||
staged = staging_content(image, info)
|
||||
resident = bytes(image) + bytes([0xFF] * (SLOT - len(image)))
|
||||
resident = bytes(image) + bytes([0xFF] * (info.slot - len(image)))
|
||||
page = info.page
|
||||
|
||||
state = UpdateState(state_path)
|
||||
@@ -700,7 +710,7 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes):
|
||||
# address 0 (the 1 KB tiny13A), its first page carries the reset vector:
|
||||
# written last, so any earlier interruption still resets into the old
|
||||
# resident, and from then on resets enter the staging copy.
|
||||
order = list(range(0, SLOT, page))
|
||||
order = list(range(0, info.slot, page))
|
||||
if info.stage == 0:
|
||||
order = order[1:] + [0]
|
||||
if write_differing(loader, info.stage, staged, order):
|
||||
@@ -723,7 +733,7 @@ def op_update_loader(loader, wait, path, state_path, fuse_bytes):
|
||||
loader.enter_copy(info.base, wait)
|
||||
if redirect:
|
||||
write_differing(loader, 0, state.page0)
|
||||
order = list(range(0, SLOT, page))
|
||||
order = list(range(0, info.slot, page))
|
||||
if info.stage == 0:
|
||||
order = [0] + order[1:]
|
||||
write_differing(loader, info.stage, state.staging, order)
|
||||
|
||||
Reference in New Issue
Block a user