Files
bootloader/pureboot/autobaud.md
BlackMark eecf9673b9 pureboot: autobaud variant, two versions for review
Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one
clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments
no longer need a per-clock build.

Two source files, differing only in the write-guard/purity tradeoff:
pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch
registers, running-slot write guard dropped, 508 B on the 1284) stays strictly
pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard
kept, 512 B) keeps every feature at the cost of that single GRV. Both fit
512/510 on all 37 chips and share two licensed simplifications: a slimmed info
block (version + signature; the host derives geometry from the chip database)
and a single-byte activation knock.

pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that
set the target, and the compiler-knob path to it. Size-tested on every chip via
pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the
host calibration handshake, and real-hardware acceptance remain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 22:15:38 +02:00

194 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# pureboot autobaud — findings and the version decision
The autobaud variant measures the host's bit timing **at runtime** from a
calibration pulse, so the image carries no clock: one clock-agnostic binary per
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
for the software-serial deployments — the RC-oscillator parts (the tinies,
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
each needs a per-clock build. Autobaud erases that axis. That is the win —
deployment, not bytes.
This document records how the fit was established, why there are **two source
files under review**, and what remains before either can ship.
## The decision to make
Two complete autobaud loaders sit side by side for review. They implement the
identical wire protocol and the identical command set; they differ in exactly
two things, and the choice between them is a single tradeoff — **the running-slot
write guard vs. strict purity.**
| | `pureboot_autobaud_pure.cpp` | `pureboot_autobaud_reg.cpp` |
|---|---|---|
| measured unit lives in | GPIOR I/O scratch regs (RAM where absent) | one global register variable (`r4`) |
| running-slot write guard | **dropped** | **kept** |
| purity (no asm, no GRV) | **yes** | one GRV — the single break |
| ATmega1284P size | **508 B** (4 B spare) | **512 B** (exact) |
| all 37 chips | 440508 B | 452512 B |
Both also take two shared, licensed simplifications (below): a slimmed info
block and a single-byte activation knock.
**The pure version** keeps pureboot's defining constraint — *"one C++ source, no
inline assembly, no global register variables"* (README.md) — intact, and pays
for it by dropping the write guard, which the host compensates for. **The
register version** keeps the guard by pinning the measured unit to a call-saved
register, which is only expressible as a global register variable — the one
thing the purity rule forbids.
Recommendation deferred to the owner. The pure version aligns with the stated
constraints (purity mandated; relaxing host-guaranteed safety licensed); the
register version keeps every feature at the cost of the purity claim and leaves
the 1284 with zero margin.
## The mechanism
The host sends calibration byte **0xC0** — a start bit plus six zero data bits
form a single low pulse of seven bit-times. The loader counts that pulse in a
poll loop the compiler emits at **seven cycles an iteration**, so the count is
the pulse length in cycles ÷ 7 × 7 = one bit period in cycles, and `count >> 2`
is that period in `_delay_loop_2`'s four-cycle iterations — the delay unit fed
to every rx/tx bit. Numerically sound: the per-bit delay tracks the true period
to **< 0.2 %** from 116 MHz at 9600 (worst seen 2.8 % at 128 kHz/1200, still
inside UART tolerance).
**The catch — codegen coupling.** `count >> 2` is exact only because the
calibration pulse's bit-count (7) equals the poll loop's cycles/iter (7). That
couples the on-wire calibration byte to what the compiler emits; a toolchain bump
that reshapes the loop breaks it silently. This must be pinned by a sim timing
test that fails on drift (see *What remains*).
The activation window is a **fixed poll budget** (`PUREBOOT_AUTOBAUD_POLLS`,
default 4,000,000), not `timeout × clock` — with no clock, whole seconds cannot
be timed. A `__uint24` holds it; a `std::uint32_t`'s fourth byte would cost two
words at each countdown step.
## How the target was found — the hand-assembly floor
Pure C++ first came out **560 B** on the 1284 (full info + guard). To learn
whether that was a hard floor or C++ overhead, a hand-optimal assembly loader was
written with the **full v4 protocol** (all commands, the full 12-byte info block,
the write guard, position independence) plus autobaud — `local/scratch/autobaud/`
in the libavr checkout, off-tree, a size probe.
**The hand-asm floor is 506 B** — it fits 512 with the *full* info block *and*
the guard, no concessions. That proved a fit was possible and gave a target. The
hand-asm achieves it two ways pure C++ cannot express directly:
1. **The measured unit in a call-saved register (Y).** rx/tx/spin read it
directly — no argument, no `movw`. In C++ an *outlined* rx/tx can only receive
the unit as a parameter (a `movw` at each of ~19 call sites, +38 B) or read it
from RAM (an `lds`), unless it is a global register variable. That register is
worth ~32 B and is exactly the impurity at issue.
2. **A cheap PC-relative slot anchor** (`rcall .+0; pop; pop`, ~10 B) where the
C++ `__builtin_return_address(0)` costs ~1826 B (GCC re-derives the stack
slot and byteswaps).
Measured on identical C++ (the same loader, unit in a register vs. RAM):
| unit home | 1284 size |
|---|---|
| global register variable | 528 |
| RAM static | 560 |
| hand-asm (register + cheap anchor) | 506 |
So even *with* a GRV, C++ carries ~22 B over hand-asm; with RAM, ~54 B. The
compiler flag axis was spent (the tuned `-f` set is optimal; `-fipa-ra`,
`-fipa-icf`, `-ftree-tail-merge`, `-flto`, `-fwhole-program` all gave nothing —
GCC's AVR ABI passes arguments in fixed registers regardless).
## The pure lever — GPIOR
The register cost is the RAM home's `lds`/`sts` (two words each) plus `.bss` and
its `__do_clear_bss`. The general-purpose I/O scratch registers (**GPIOR1:GPIOR2**,
adjacent) are reached by `in`/`out` — one word — through libavr's named register
surface: **no inline asm, no global register variable, fully pure.** Storing the
unit there sidesteps `.bss` and halves each access. Where a chip has no GPIOR
(the t13, m8, m16/32) the pure version falls back to a plain static; those chips
have ample headroom.
GPIOR alone did not close the gap (the write guard's `slot_high` machinery is the
bulk). The pure version fits by combining GPIOR storage with the licensed
simplifications; the register version fits by pinning the unit and so affording
the guard.
### The ablation (1284, all pure except the GRV row)
| config | GPIOR unit | GRV unit |
|---|---|---|
| full info + guard + 2-knock | 550 | 528 |
| slim info + guard + 2-knock | 540 | 518 |
| slim info + guard + 1-knock | — | **512** ✓ |
| slim info + no guard + 2-knock | 514 | — |
| slim info + no guard + 1-knock | **508** ✓ | — |
The pivotal cost is the write guard's `slot_high` (`__builtin_return_address`,
~26 B), removable only when the info block is slim (so nothing else needs
`slot_high`) *and* the guard is gone. That is why the pure version drops the
guard and the register version keeps it.
## Sizes — every chip, both versions
Both build and size-test green on all 37 (patched-vector budget 510, else 512):
| chip class | pure | register | budget |
|---|---|---|---|
| ATtiny13/13A | 446 | 452 | 510 |
| ATtiny25/45/85 | 440444 | 456460 | 510 |
| ATmega8/8A | 472 | 476 | 512 |
| ATmega16/32 | 476 | 478 | 512 |
| ATmega48 family | 440 | 456 | 510 |
| ATmega88/168/328 | 462466 | 476478 | 512 |
| ATmega164/324/644 | 466 | 478 | 512 |
| **ATmega1284/P** | **508** | **512** | 512 |
The 1284 is the tight one — the far-flash machinery (ELPM/RAMPZ, word-addressed
wire) it alone carries. The `pureboot_autobaud_*.size` tests run per chip in the
port's build.
## The shared simplifications, and why each is licensed
- **Slimmed info block.** `'b'` returns the version and the three signature bytes
— the loader's identity and the chip's — as immediates, instead of the full
12-byte block read from flash. The host derives page size, loader base, EEPROM
size and the addressing flags from the signature via its own chip database.
This drops the flash-resident table and, with the guard also gone in the pure
version, the entire `slot_high` anchor. Owner-approved: "keep version + chip
type, simplify the rest; the host derives geometry."
- **Single-byte knock.** One `'p'` activates; the calibration pulse has already
proven a host is present. (Fixed-baud pureboot needs `p`+`b` because it has no
prior proof.)
- **No running-slot write guard** *(pure version only)*. The guard stops a
*buggy* host from programming the loader's own slot; the host is guaranteed
never to do so (it knows the loader's location). Licensed by "the host
guarantees safety" (README.md). Self-update still works — it is the safety net
against host bugs that is lost, not a functional path.
## What remains (either version)
The decider — does pure C++ fit 512 on every chip — is answered (yes). Not yet
done, and required before shipping:
1. **Sim-validate the lock.** Run the chosen variant over the GPIO⇄pty bridge at
≥3 exact F_CPU (1/8/16 MHz, the tiny13's 9.6 MHz): confirm it measures the
right unit and completes a flash + verify. This test is also what pins the
codegen-coupled calibration constant (0xC0 vs the 7-cycle loop) against a
toolchain bump.
2. **Host tool + protocol.** `pureboot.py` learns the calibration handshake
(send the 0xC0 lead, then knock at the locked rate), the slim info format
(derive geometry from the signature), and — for the pure version — that the
loader carries no write guard. README protocol section updated.
3. **Real-hardware acceptance.** Autobaud exists for what a cycle-exact simulator
cannot produce: a real RC oscillator at ±10 % with drift and jitter. Drive an
internal-oscillator ATtiny from the host at a fixed baud; confirm lock plus a
full flash + verify (Windows environment, hardware present — confirm first).
## Files
- `pureboot_autobaud_pure.cpp` — the pure version (GPIOR/RAM unit, no guard).
- `pureboot_autobaud_reg.cpp` — the register version (GRV unit, guard kept).
- `CMakeLists.txt``pureboot_add_autobaud()` builds either; the port's build
size-tests both on every chip.
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
probe, off-tree and gitignored; not sim-verified, a size reference only.