Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments no longer need a per-clock build. Two source files, differing only in the write-guard/purity tradeoff: pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch registers, running-slot write guard dropped, 508 B on the 1284) stays strictly pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard kept, 512 B) keeps every feature at the cost of that single GRV. Both fit 512/510 on all 37 chips and share two licensed simplifications: a slimmed info block (version + signature; the host derives geometry from the chip database) and a single-byte activation knock. pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that set the target, and the compiler-knob path to it. Size-tested on every chip via pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the host calibration handshake, and real-hardware acceptance remain. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.7 KiB
pureboot autobaud — findings and the version decision
The autobaud variant measures the host's bit timing at runtime from a calibration pulse, so the image carries no clock: one clock-agnostic binary per chip runs at any F_CPU and locks onto whatever baud the host sends. It exists for the software-serial deployments — the RC-oscillator parts (the tinies, internal-oscillator megas) whose exact clock is uncertain and drifts, so today each needs a per-clock build. Autobaud erases that axis. That is the win — deployment, not bytes.
This document records how the fit was established, why there are two source files under review, and what remains before either can ship.
The decision to make
Two complete autobaud loaders sit side by side for review. They implement the identical wire protocol and the identical command set; they differ in exactly two things, and the choice between them is a single tradeoff — the running-slot write guard vs. strict purity.
pureboot_autobaud_pure.cpp |
pureboot_autobaud_reg.cpp |
|
|---|---|---|
| measured unit lives in | GPIOR I/O scratch regs (RAM where absent) | one global register variable (r4) |
| running-slot write guard | dropped | kept |
| purity (no asm, no GRV) | yes | one GRV — the single break |
| ATmega1284P size | 508 B (4 B spare) | 512 B (exact) |
| all 37 chips | 440–508 B | 452–512 B |
Both also take two shared, licensed simplifications (below): a slimmed info block and a single-byte activation knock.
The pure version keeps pureboot's defining constraint — "one C++ source, no inline assembly, no global register variables" (README.md) — intact, and pays for it by dropping the write guard, which the host compensates for. The register version keeps the guard by pinning the measured unit to a call-saved register, which is only expressible as a global register variable — the one thing the purity rule forbids.
Recommendation deferred to the owner. The pure version aligns with the stated constraints (purity mandated; relaxing host-guaranteed safety licensed); the register version keeps every feature at the cost of the purity claim and leaves the 1284 with zero margin.
The mechanism
The host sends calibration byte 0xC0 — a start bit plus six zero data bits
form a single low pulse of seven bit-times. The loader counts that pulse in a
poll loop the compiler emits at seven cycles an iteration, so the count is
the pulse length in cycles ÷ 7 × 7 = one bit period in cycles, and count >> 2
is that period in _delay_loop_2's four-cycle iterations — the delay unit fed
to every rx/tx bit. Numerically sound: the per-bit delay tracks the true period
to < 0.2 % from 1–16 MHz at 9600 (worst seen −2.8 % at 128 kHz/1200, still
inside UART tolerance).
The catch — codegen coupling. count >> 2 is exact only because the
calibration pulse's bit-count (7) equals the poll loop's cycles/iter (7). That
couples the on-wire calibration byte to what the compiler emits; a toolchain bump
that reshapes the loop breaks it silently. This must be pinned by a sim timing
test that fails on drift (see What remains).
The activation window is a fixed poll budget (PUREBOOT_AUTOBAUD_POLLS,
default 4,000,000), not timeout × clock — with no clock, whole seconds cannot
be timed. A __uint24 holds it; a std::uint32_t's fourth byte would cost two
words at each countdown step.
How the target was found — the hand-assembly floor
Pure C++ first came out 560 B on the 1284 (full info + guard). To learn
whether that was a hard floor or C++ overhead, a hand-optimal assembly loader was
written with the full v4 protocol (all commands, the full 12-byte info block,
the write guard, position independence) plus autobaud — local/scratch/autobaud/
in the libavr checkout, off-tree, a size probe.
The hand-asm floor is 506 B — it fits 512 with the full info block and the guard, no concessions. That proved a fit was possible and gave a target. The hand-asm achieves it two ways pure C++ cannot express directly:
- The measured unit in a call-saved register (Y). rx/tx/spin read it
directly — no argument, no
movw. In C++ an outlined rx/tx can only receive the unit as a parameter (amovwat each of ~19 call sites, +38 B) or read it from RAM (anlds), unless it is a global register variable. That register is worth ~32 B and is exactly the impurity at issue. - A cheap PC-relative slot anchor (
rcall .+0; pop; pop, ~10 B) where the C++__builtin_return_address(0)costs ~18–26 B (GCC re-derives the stack slot and byteswaps).
Measured on identical C++ (the same loader, unit in a register vs. RAM):
| unit home | 1284 size |
|---|---|
| global register variable | 528 |
| RAM static | 560 |
| hand-asm (register + cheap anchor) | 506 |
So even with a GRV, C++ carries ~22 B over hand-asm; with RAM, ~54 B. The
compiler flag axis was spent (the tuned -f set is optimal; -fipa-ra,
-fipa-icf, -ftree-tail-merge, -flto, -fwhole-program all gave nothing —
GCC's AVR ABI passes arguments in fixed registers regardless).
The pure lever — GPIOR
The register cost is the RAM home's lds/sts (two words each) plus .bss and
its __do_clear_bss. The general-purpose I/O scratch registers (GPIOR1:GPIOR2,
adjacent) are reached by in/out — one word — through libavr's named register
surface: no inline asm, no global register variable, fully pure. Storing the
unit there sidesteps .bss and halves each access. Where a chip has no GPIOR
(the t13, m8, m16/32) the pure version falls back to a plain static; those chips
have ample headroom.
GPIOR alone did not close the gap (the write guard's slot_high machinery is the
bulk). The pure version fits by combining GPIOR storage with the licensed
simplifications; the register version fits by pinning the unit and so affording
the guard.
The ablation (1284, all pure except the GRV row)
| config | GPIOR unit | GRV unit |
|---|---|---|
| full info + guard + 2-knock | 550 | 528 |
| slim info + guard + 2-knock | 540 | 518 |
| slim info + guard + 1-knock | — | 512 ✓ |
| slim info + no guard + 2-knock | 514 | — |
| slim info + no guard + 1-knock | 508 ✓ | — |
The pivotal cost is the write guard's slot_high (__builtin_return_address,
~26 B), removable only when the info block is slim (so nothing else needs
slot_high) and the guard is gone. That is why the pure version drops the
guard and the register version keeps it.
Sizes — every chip, both versions
Both build and size-test green on all 37 (patched-vector budget 510, else 512):
| chip class | pure | register | budget |
|---|---|---|---|
| ATtiny13/13A | 446 | 452 | 510 |
| ATtiny25/45/85 | 440–444 | 456–460 | 510 |
| ATmega8/8A | 472 | 476 | 512 |
| ATmega16/32 | 476 | 478 | 512 |
| ATmega48 family | 440 | 456 | 510 |
| ATmega88/168/328 | 462–466 | 476–478 | 512 |
| ATmega164/324/644 | 466 | 478 | 512 |
| ATmega1284/P | 508 | 512 | 512 |
The 1284 is the tight one — the far-flash machinery (ELPM/RAMPZ, word-addressed
wire) it alone carries. The pureboot_autobaud_*.size tests run per chip in the
port's build.
The shared simplifications, and why each is licensed
- Slimmed info block.
'b'returns the version and the three signature bytes — the loader's identity and the chip's — as immediates, instead of the full 12-byte block read from flash. The host derives page size, loader base, EEPROM size and the addressing flags from the signature via its own chip database. This drops the flash-resident table and, with the guard also gone in the pure version, the entireslot_highanchor. Owner-approved: "keep version + chip type, simplify the rest; the host derives geometry." - Single-byte knock. One
'p'activates; the calibration pulse has already proven a host is present. (Fixed-baud pureboot needsp+bbecause it has no prior proof.) - No running-slot write guard (pure version only). The guard stops a buggy host from programming the loader's own slot; the host is guaranteed never to do so (it knows the loader's location). Licensed by "the host guarantees safety" (README.md). Self-update still works — it is the safety net against host bugs that is lost, not a functional path.
What remains (either version)
The decider — does pure C++ fit 512 on every chip — is answered (yes). Not yet done, and required before shipping:
- Sim-validate the lock. Run the chosen variant over the GPIO⇄pty bridge at ≥3 exact F_CPU (1/8/16 MHz, the tiny13's 9.6 MHz): confirm it measures the right unit and completes a flash + verify. This test is also what pins the codegen-coupled calibration constant (0xC0 vs the 7-cycle loop) against a toolchain bump.
- Host tool + protocol.
pureboot.pylearns the calibration handshake (send the 0xC0 lead, then knock at the locked rate), the slim info format (derive geometry from the signature), and — for the pure version — that the loader carries no write guard. README protocol section updated. - Real-hardware acceptance. Autobaud exists for what a cycle-exact simulator cannot produce: a real RC oscillator at ±10 % with drift and jitter. Drive an internal-oscillator ATtiny from the host at a fixed baud; confirm lock plus a full flash + verify (Windows environment, hardware present — confirm first).
Files
pureboot_autobaud_pure.cpp— the pure version (GPIOR/RAM unit, no guard).pureboot_autobaud_reg.cpp— the register version (GRV unit, guard kept).CMakeLists.txt—pureboot_add_autobaud()builds either; the port's build size-tests both on every chip.local/scratch/autobaud/floor_1284.S(libavr checkout) — the hand-asm floor probe, off-tree and gitignored; not sim-verified, a size reference only.