pureboot: autobaud variant, two versions for review

Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one
clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments
no longer need a per-clock build.

Two source files, differing only in the write-guard/purity tradeoff:
pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch
registers, running-slot write guard dropped, 508 B on the 1284) stays strictly
pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard
kept, 512 B) keeps every feature at the cost of that single GRV. Both fit
512/510 on all 37 chips and share two licensed simplifications: a slimmed info
block (version + signature; the host derives geometry from the chip database)
and a single-byte activation knock.

pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that
set the target, and the compiler-knob path to it. Size-tested on every chip via
pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the
host calibration handshake, and real-hardware acceptance remain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-07-23 22:15:38 +02:00
parent c78bf2841a
commit eecf9673b9
5 changed files with 952 additions and 0 deletions

193
pureboot/autobaud.md Normal file
View File

@@ -0,0 +1,193 @@
# pureboot autobaud — findings and the version decision
The autobaud variant measures the host's bit timing **at runtime** from a
calibration pulse, so the image carries no clock: one clock-agnostic binary per
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
for the software-serial deployments — the RC-oscillator parts (the tinies,
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
each needs a per-clock build. Autobaud erases that axis. That is the win —
deployment, not bytes.
This document records how the fit was established, why there are **two source
files under review**, and what remains before either can ship.
## The decision to make
Two complete autobaud loaders sit side by side for review. They implement the
identical wire protocol and the identical command set; they differ in exactly
two things, and the choice between them is a single tradeoff — **the running-slot
write guard vs. strict purity.**
| | `pureboot_autobaud_pure.cpp` | `pureboot_autobaud_reg.cpp` |
|---|---|---|
| measured unit lives in | GPIOR I/O scratch regs (RAM where absent) | one global register variable (`r4`) |
| running-slot write guard | **dropped** | **kept** |
| purity (no asm, no GRV) | **yes** | one GRV — the single break |
| ATmega1284P size | **508 B** (4 B spare) | **512 B** (exact) |
| all 37 chips | 440508 B | 452512 B |
Both also take two shared, licensed simplifications (below): a slimmed info
block and a single-byte activation knock.
**The pure version** keeps pureboot's defining constraint — *"one C++ source, no
inline assembly, no global register variables"* (README.md) — intact, and pays
for it by dropping the write guard, which the host compensates for. **The
register version** keeps the guard by pinning the measured unit to a call-saved
register, which is only expressible as a global register variable — the one
thing the purity rule forbids.
Recommendation deferred to the owner. The pure version aligns with the stated
constraints (purity mandated; relaxing host-guaranteed safety licensed); the
register version keeps every feature at the cost of the purity claim and leaves
the 1284 with zero margin.
## The mechanism
The host sends calibration byte **0xC0** — a start bit plus six zero data bits
form a single low pulse of seven bit-times. The loader counts that pulse in a
poll loop the compiler emits at **seven cycles an iteration**, so the count is
the pulse length in cycles ÷ 7 × 7 = one bit period in cycles, and `count >> 2`
is that period in `_delay_loop_2`'s four-cycle iterations — the delay unit fed
to every rx/tx bit. Numerically sound: the per-bit delay tracks the true period
to **< 0.2 %** from 116 MHz at 9600 (worst seen 2.8 % at 128 kHz/1200, still
inside UART tolerance).
**The catch — codegen coupling.** `count >> 2` is exact only because the
calibration pulse's bit-count (7) equals the poll loop's cycles/iter (7). That
couples the on-wire calibration byte to what the compiler emits; a toolchain bump
that reshapes the loop breaks it silently. This must be pinned by a sim timing
test that fails on drift (see *What remains*).
The activation window is a **fixed poll budget** (`PUREBOOT_AUTOBAUD_POLLS`,
default 4,000,000), not `timeout × clock` — with no clock, whole seconds cannot
be timed. A `__uint24` holds it; a `std::uint32_t`'s fourth byte would cost two
words at each countdown step.
## How the target was found — the hand-assembly floor
Pure C++ first came out **560 B** on the 1284 (full info + guard). To learn
whether that was a hard floor or C++ overhead, a hand-optimal assembly loader was
written with the **full v4 protocol** (all commands, the full 12-byte info block,
the write guard, position independence) plus autobaud — `local/scratch/autobaud/`
in the libavr checkout, off-tree, a size probe.
**The hand-asm floor is 506 B** — it fits 512 with the *full* info block *and*
the guard, no concessions. That proved a fit was possible and gave a target. The
hand-asm achieves it two ways pure C++ cannot express directly:
1. **The measured unit in a call-saved register (Y).** rx/tx/spin read it
directly — no argument, no `movw`. In C++ an *outlined* rx/tx can only receive
the unit as a parameter (a `movw` at each of ~19 call sites, +38 B) or read it
from RAM (an `lds`), unless it is a global register variable. That register is
worth ~32 B and is exactly the impurity at issue.
2. **A cheap PC-relative slot anchor** (`rcall .+0; pop; pop`, ~10 B) where the
C++ `__builtin_return_address(0)` costs ~1826 B (GCC re-derives the stack
slot and byteswaps).
Measured on identical C++ (the same loader, unit in a register vs. RAM):
| unit home | 1284 size |
|---|---|
| global register variable | 528 |
| RAM static | 560 |
| hand-asm (register + cheap anchor) | 506 |
So even *with* a GRV, C++ carries ~22 B over hand-asm; with RAM, ~54 B. The
compiler flag axis was spent (the tuned `-f` set is optimal; `-fipa-ra`,
`-fipa-icf`, `-ftree-tail-merge`, `-flto`, `-fwhole-program` all gave nothing —
GCC's AVR ABI passes arguments in fixed registers regardless).
## The pure lever — GPIOR
The register cost is the RAM home's `lds`/`sts` (two words each) plus `.bss` and
its `__do_clear_bss`. The general-purpose I/O scratch registers (**GPIOR1:GPIOR2**,
adjacent) are reached by `in`/`out` — one word — through libavr's named register
surface: **no inline asm, no global register variable, fully pure.** Storing the
unit there sidesteps `.bss` and halves each access. Where a chip has no GPIOR
(the t13, m8, m16/32) the pure version falls back to a plain static; those chips
have ample headroom.
GPIOR alone did not close the gap (the write guard's `slot_high` machinery is the
bulk). The pure version fits by combining GPIOR storage with the licensed
simplifications; the register version fits by pinning the unit and so affording
the guard.
### The ablation (1284, all pure except the GRV row)
| config | GPIOR unit | GRV unit |
|---|---|---|
| full info + guard + 2-knock | 550 | 528 |
| slim info + guard + 2-knock | 540 | 518 |
| slim info + guard + 1-knock | — | **512** ✓ |
| slim info + no guard + 2-knock | 514 | — |
| slim info + no guard + 1-knock | **508** ✓ | — |
The pivotal cost is the write guard's `slot_high` (`__builtin_return_address`,
~26 B), removable only when the info block is slim (so nothing else needs
`slot_high`) *and* the guard is gone. That is why the pure version drops the
guard and the register version keeps it.
## Sizes — every chip, both versions
Both build and size-test green on all 37 (patched-vector budget 510, else 512):
| chip class | pure | register | budget |
|---|---|---|---|
| ATtiny13/13A | 446 | 452 | 510 |
| ATtiny25/45/85 | 440444 | 456460 | 510 |
| ATmega8/8A | 472 | 476 | 512 |
| ATmega16/32 | 476 | 478 | 512 |
| ATmega48 family | 440 | 456 | 510 |
| ATmega88/168/328 | 462466 | 476478 | 512 |
| ATmega164/324/644 | 466 | 478 | 512 |
| **ATmega1284/P** | **508** | **512** | 512 |
The 1284 is the tight one — the far-flash machinery (ELPM/RAMPZ, word-addressed
wire) it alone carries. The `pureboot_autobaud_*.size` tests run per chip in the
port's build.
## The shared simplifications, and why each is licensed
- **Slimmed info block.** `'b'` returns the version and the three signature bytes
— the loader's identity and the chip's — as immediates, instead of the full
12-byte block read from flash. The host derives page size, loader base, EEPROM
size and the addressing flags from the signature via its own chip database.
This drops the flash-resident table and, with the guard also gone in the pure
version, the entire `slot_high` anchor. Owner-approved: "keep version + chip
type, simplify the rest; the host derives geometry."
- **Single-byte knock.** One `'p'` activates; the calibration pulse has already
proven a host is present. (Fixed-baud pureboot needs `p`+`b` because it has no
prior proof.)
- **No running-slot write guard** *(pure version only)*. The guard stops a
*buggy* host from programming the loader's own slot; the host is guaranteed
never to do so (it knows the loader's location). Licensed by "the host
guarantees safety" (README.md). Self-update still works — it is the safety net
against host bugs that is lost, not a functional path.
## What remains (either version)
The decider — does pure C++ fit 512 on every chip — is answered (yes). Not yet
done, and required before shipping:
1. **Sim-validate the lock.** Run the chosen variant over the GPIO⇄pty bridge at
≥3 exact F_CPU (1/8/16 MHz, the tiny13's 9.6 MHz): confirm it measures the
right unit and completes a flash + verify. This test is also what pins the
codegen-coupled calibration constant (0xC0 vs the 7-cycle loop) against a
toolchain bump.
2. **Host tool + protocol.** `pureboot.py` learns the calibration handshake
(send the 0xC0 lead, then knock at the locked rate), the slim info format
(derive geometry from the signature), and — for the pure version — that the
loader carries no write guard. README protocol section updated.
3. **Real-hardware acceptance.** Autobaud exists for what a cycle-exact simulator
cannot produce: a real RC oscillator at ±10 % with drift and jitter. Drive an
internal-oscillator ATtiny from the host at a fixed baud; confirm lock plus a
full flash + verify (Windows environment, hardware present — confirm first).
## Files
- `pureboot_autobaud_pure.cpp` — the pure version (GPIOR/RAM unit, no guard).
- `pureboot_autobaud_reg.cpp` — the register version (GRV unit, guard kept).
- `CMakeLists.txt``pureboot_add_autobaud()` builds either; the port's build
size-tests both on every chip.
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
probe, off-tree and gitignored; not sim-verified, a size reference only.