Measure the host's bit timing at runtime from a 0xC0 calibration pulse, so one clock-agnostic image per chip runs at any F_CPU — the RC-oscillator deployments no longer need a per-clock build. Two source files, differing only in the write-guard/purity tradeoff: pureboot_autobaud_pure.cpp (the measured unit in the GPIOR I/O scratch registers, running-slot write guard dropped, 508 B on the 1284) stays strictly pure; pureboot_autobaud_reg.cpp (unit in one global register variable, guard kept, 512 B) keeps every feature at the cost of that single GRV. Both fit 512/510 on all 37 chips and share two licensed simplifications: a slimmed info block (version + signature; the host derives geometry from the chip database) and a single-byte activation knock. pureboot/autobaud.md records the decision, the hand-assembly floor (506 B) that set the target, and the compiler-knob path to it. Size-tested on every chip via pureboot_add_autobaud(); the fixed-baud loader is untouched. Sim validation, the host calibration handshake, and real-hardware acceptance remain. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
194 lines
9.7 KiB
Markdown
194 lines
9.7 KiB
Markdown
# pureboot autobaud — findings and the version decision
|
||
|
||
The autobaud variant measures the host's bit timing **at runtime** from a
|
||
calibration pulse, so the image carries no clock: one clock-agnostic binary per
|
||
chip runs at any F_CPU and locks onto whatever baud the host sends. It exists
|
||
for the software-serial deployments — the RC-oscillator parts (the tinies,
|
||
internal-oscillator megas) whose exact clock is uncertain and drifts, so today
|
||
each needs a per-clock build. Autobaud erases that axis. That is the win —
|
||
deployment, not bytes.
|
||
|
||
This document records how the fit was established, why there are **two source
|
||
files under review**, and what remains before either can ship.
|
||
|
||
## The decision to make
|
||
|
||
Two complete autobaud loaders sit side by side for review. They implement the
|
||
identical wire protocol and the identical command set; they differ in exactly
|
||
two things, and the choice between them is a single tradeoff — **the running-slot
|
||
write guard vs. strict purity.**
|
||
|
||
| | `pureboot_autobaud_pure.cpp` | `pureboot_autobaud_reg.cpp` |
|
||
|---|---|---|
|
||
| measured unit lives in | GPIOR I/O scratch regs (RAM where absent) | one global register variable (`r4`) |
|
||
| running-slot write guard | **dropped** | **kept** |
|
||
| purity (no asm, no GRV) | **yes** | one GRV — the single break |
|
||
| ATmega1284P size | **508 B** (4 B spare) | **512 B** (exact) |
|
||
| all 37 chips | 440–508 B | 452–512 B |
|
||
|
||
Both also take two shared, licensed simplifications (below): a slimmed info
|
||
block and a single-byte activation knock.
|
||
|
||
**The pure version** keeps pureboot's defining constraint — *"one C++ source, no
|
||
inline assembly, no global register variables"* (README.md) — intact, and pays
|
||
for it by dropping the write guard, which the host compensates for. **The
|
||
register version** keeps the guard by pinning the measured unit to a call-saved
|
||
register, which is only expressible as a global register variable — the one
|
||
thing the purity rule forbids.
|
||
|
||
Recommendation deferred to the owner. The pure version aligns with the stated
|
||
constraints (purity mandated; relaxing host-guaranteed safety licensed); the
|
||
register version keeps every feature at the cost of the purity claim and leaves
|
||
the 1284 with zero margin.
|
||
|
||
## The mechanism
|
||
|
||
The host sends calibration byte **0xC0** — a start bit plus six zero data bits
|
||
form a single low pulse of seven bit-times. The loader counts that pulse in a
|
||
poll loop the compiler emits at **seven cycles an iteration**, so the count is
|
||
the pulse length in cycles ÷ 7 × 7 = one bit period in cycles, and `count >> 2`
|
||
is that period in `_delay_loop_2`'s four-cycle iterations — the delay unit fed
|
||
to every rx/tx bit. Numerically sound: the per-bit delay tracks the true period
|
||
to **< 0.2 %** from 1–16 MHz at 9600 (worst seen −2.8 % at 128 kHz/1200, still
|
||
inside UART tolerance).
|
||
|
||
**The catch — codegen coupling.** `count >> 2` is exact only because the
|
||
calibration pulse's bit-count (7) equals the poll loop's cycles/iter (7). That
|
||
couples the on-wire calibration byte to what the compiler emits; a toolchain bump
|
||
that reshapes the loop breaks it silently. This must be pinned by a sim timing
|
||
test that fails on drift (see *What remains*).
|
||
|
||
The activation window is a **fixed poll budget** (`PUREBOOT_AUTOBAUD_POLLS`,
|
||
default 4,000,000), not `timeout × clock` — with no clock, whole seconds cannot
|
||
be timed. A `__uint24` holds it; a `std::uint32_t`'s fourth byte would cost two
|
||
words at each countdown step.
|
||
|
||
## How the target was found — the hand-assembly floor
|
||
|
||
Pure C++ first came out **560 B** on the 1284 (full info + guard). To learn
|
||
whether that was a hard floor or C++ overhead, a hand-optimal assembly loader was
|
||
written with the **full v4 protocol** (all commands, the full 12-byte info block,
|
||
the write guard, position independence) plus autobaud — `local/scratch/autobaud/`
|
||
in the libavr checkout, off-tree, a size probe.
|
||
|
||
**The hand-asm floor is 506 B** — it fits 512 with the *full* info block *and*
|
||
the guard, no concessions. That proved a fit was possible and gave a target. The
|
||
hand-asm achieves it two ways pure C++ cannot express directly:
|
||
|
||
1. **The measured unit in a call-saved register (Y).** rx/tx/spin read it
|
||
directly — no argument, no `movw`. In C++ an *outlined* rx/tx can only receive
|
||
the unit as a parameter (a `movw` at each of ~19 call sites, +38 B) or read it
|
||
from RAM (an `lds`), unless it is a global register variable. That register is
|
||
worth ~32 B and is exactly the impurity at issue.
|
||
2. **A cheap PC-relative slot anchor** (`rcall .+0; pop; pop`, ~10 B) where the
|
||
C++ `__builtin_return_address(0)` costs ~18–26 B (GCC re-derives the stack
|
||
slot and byteswaps).
|
||
|
||
Measured on identical C++ (the same loader, unit in a register vs. RAM):
|
||
|
||
| unit home | 1284 size |
|
||
|---|---|
|
||
| global register variable | 528 |
|
||
| RAM static | 560 |
|
||
| hand-asm (register + cheap anchor) | 506 |
|
||
|
||
So even *with* a GRV, C++ carries ~22 B over hand-asm; with RAM, ~54 B. The
|
||
compiler flag axis was spent (the tuned `-f` set is optimal; `-fipa-ra`,
|
||
`-fipa-icf`, `-ftree-tail-merge`, `-flto`, `-fwhole-program` all gave nothing —
|
||
GCC's AVR ABI passes arguments in fixed registers regardless).
|
||
|
||
## The pure lever — GPIOR
|
||
|
||
The register cost is the RAM home's `lds`/`sts` (two words each) plus `.bss` and
|
||
its `__do_clear_bss`. The general-purpose I/O scratch registers (**GPIOR1:GPIOR2**,
|
||
adjacent) are reached by `in`/`out` — one word — through libavr's named register
|
||
surface: **no inline asm, no global register variable, fully pure.** Storing the
|
||
unit there sidesteps `.bss` and halves each access. Where a chip has no GPIOR
|
||
(the t13, m8, m16/32) the pure version falls back to a plain static; those chips
|
||
have ample headroom.
|
||
|
||
GPIOR alone did not close the gap (the write guard's `slot_high` machinery is the
|
||
bulk). The pure version fits by combining GPIOR storage with the licensed
|
||
simplifications; the register version fits by pinning the unit and so affording
|
||
the guard.
|
||
|
||
### The ablation (1284, all pure except the GRV row)
|
||
|
||
| config | GPIOR unit | GRV unit |
|
||
|---|---|---|
|
||
| full info + guard + 2-knock | 550 | 528 |
|
||
| slim info + guard + 2-knock | 540 | 518 |
|
||
| slim info + guard + 1-knock | — | **512** ✓ |
|
||
| slim info + no guard + 2-knock | 514 | — |
|
||
| slim info + no guard + 1-knock | **508** ✓ | — |
|
||
|
||
The pivotal cost is the write guard's `slot_high` (`__builtin_return_address`,
|
||
~26 B), removable only when the info block is slim (so nothing else needs
|
||
`slot_high`) *and* the guard is gone. That is why the pure version drops the
|
||
guard and the register version keeps it.
|
||
|
||
## Sizes — every chip, both versions
|
||
|
||
Both build and size-test green on all 37 (patched-vector budget 510, else 512):
|
||
|
||
| chip class | pure | register | budget |
|
||
|---|---|---|---|
|
||
| ATtiny13/13A | 446 | 452 | 510 |
|
||
| ATtiny25/45/85 | 440–444 | 456–460 | 510 |
|
||
| ATmega8/8A | 472 | 476 | 512 |
|
||
| ATmega16/32 | 476 | 478 | 512 |
|
||
| ATmega48 family | 440 | 456 | 510 |
|
||
| ATmega88/168/328 | 462–466 | 476–478 | 512 |
|
||
| ATmega164/324/644 | 466 | 478 | 512 |
|
||
| **ATmega1284/P** | **508** | **512** | 512 |
|
||
|
||
The 1284 is the tight one — the far-flash machinery (ELPM/RAMPZ, word-addressed
|
||
wire) it alone carries. The `pureboot_autobaud_*.size` tests run per chip in the
|
||
port's build.
|
||
|
||
## The shared simplifications, and why each is licensed
|
||
|
||
- **Slimmed info block.** `'b'` returns the version and the three signature bytes
|
||
— the loader's identity and the chip's — as immediates, instead of the full
|
||
12-byte block read from flash. The host derives page size, loader base, EEPROM
|
||
size and the addressing flags from the signature via its own chip database.
|
||
This drops the flash-resident table and, with the guard also gone in the pure
|
||
version, the entire `slot_high` anchor. Owner-approved: "keep version + chip
|
||
type, simplify the rest; the host derives geometry."
|
||
- **Single-byte knock.** One `'p'` activates; the calibration pulse has already
|
||
proven a host is present. (Fixed-baud pureboot needs `p`+`b` because it has no
|
||
prior proof.)
|
||
- **No running-slot write guard** *(pure version only)*. The guard stops a
|
||
*buggy* host from programming the loader's own slot; the host is guaranteed
|
||
never to do so (it knows the loader's location). Licensed by "the host
|
||
guarantees safety" (README.md). Self-update still works — it is the safety net
|
||
against host bugs that is lost, not a functional path.
|
||
|
||
## What remains (either version)
|
||
|
||
The decider — does pure C++ fit 512 on every chip — is answered (yes). Not yet
|
||
done, and required before shipping:
|
||
|
||
1. **Sim-validate the lock.** Run the chosen variant over the GPIO⇄pty bridge at
|
||
≥3 exact F_CPU (1/8/16 MHz, the tiny13's 9.6 MHz): confirm it measures the
|
||
right unit and completes a flash + verify. This test is also what pins the
|
||
codegen-coupled calibration constant (0xC0 vs the 7-cycle loop) against a
|
||
toolchain bump.
|
||
2. **Host tool + protocol.** `pureboot.py` learns the calibration handshake
|
||
(send the 0xC0 lead, then knock at the locked rate), the slim info format
|
||
(derive geometry from the signature), and — for the pure version — that the
|
||
loader carries no write guard. README protocol section updated.
|
||
3. **Real-hardware acceptance.** Autobaud exists for what a cycle-exact simulator
|
||
cannot produce: a real RC oscillator at ±10 % with drift and jitter. Drive an
|
||
internal-oscillator ATtiny from the host at a fixed baud; confirm lock plus a
|
||
full flash + verify (Windows environment, hardware present — confirm first).
|
||
|
||
## Files
|
||
|
||
- `pureboot_autobaud_pure.cpp` — the pure version (GPIOR/RAM unit, no guard).
|
||
- `pureboot_autobaud_reg.cpp` — the register version (GRV unit, guard kept).
|
||
- `CMakeLists.txt` — `pureboot_add_autobaud()` builds either; the port's build
|
||
size-tests both on every chip.
|
||
- `local/scratch/autobaud/floor_1284.S` (libavr checkout) — the hand-asm floor
|
||
probe, off-tree and gitignored; not sim-verified, a size reference only.
|