4 Commits

Author SHA1 Message Date
eb213e1025 one-wire: a lost echo and a dead line are not the same report
The blind-write path said "lost to the device's ack" for any missing echo, and
the count is what distinguishes two different faults. Some bytes lost is the
device's ack winning the line against the host's series resistor — ordinary,
and what the knock retry absorbs. *Every* byte lost is nothing coming back at
all, which means the line is not free: a pin held low, a wedge, or an RX that
is not on it.

Found pointing the wrong way on purpose-built hardware. This rig's LED demo
ends by driving every port pin low, and one of them is the shared link — so a
knock into a finished demo got no echo whatsoever and was told the device had
acked, when nothing had answered and nothing could. Same retry either way, but
blaming an ack that never happened sends the reader to the protocol when the
answer is a pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:32:18 +02:00
3e4bfbaf48 pbhw: the marker check assumed a board whose port-open is not a reset
Its comment said "opening the port does not reset a board whose DTR is
unwired, so this simply listens" — true of the tiny it was written against,
false of an Arduino, and this is the generic harness. Where DTR is wired to
reset, that open resets the part and the activation window comes first, so a
fixture emitting its banner once says it on the far side of a wait the suite
cannot know the length of: the window is a compile-time constant and nothing
on the wire reports it. The suite read the silence as an application that
never ran, on a board where it demonstrably had.

So --marker-wait, defaulting to the 2.5 s that was hardcoded, and a failure
that names the window as the candidate rather than leaving the next person to
suspect the loader. The other half is the fixture: PUREBOOT_HEARTBEAT makes
the observation independent of when the listener arrives, which is what the
rig's own builds now pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 17:12:08 +02:00
579ca81b27 one-wire: the knock's lost byte is the wiring, and the diagnosis was unreachable
Measured on an ATtiny13A with the link folded onto PB3 and the FTDI's TX
reaching it through 1 k: a knock aimed at a loader already in session loses
its second byte every time, 8 runs of 8, never intermittently. The first byte
draws a prompt while the second is still going out and the device's push-pull
ack wins the line against the resistor, so that byte is destroyed rather than
delayed — which is what the README predicted and the sim bridge cannot show,
since it arbitrates the line by queueing.

The recovery for it existed and could not run. Two defects:

OneWirePort.write read its echo with read_exact, whose contract is to raise, so
the "one-wire echo missing — is the adapter's RX tied to the line?" message was
unreachable on any line that simply fell quiet, and a bare "timeout: got 0 of 1
bytes" surfaced in its place. The one message the class exists to produce could
never be produced. The read is speculative and is now read_available.

And any raise from write aborted _handshake before the retry loop that exists
to absorb exactly this, whose docstring already claimed it "converges into an
already-live session" — true on a pty, impossible on real wiring. The knock is
now the one write marked blind: a missing echo there is a property of the
shared line, counted and reported under -v rather than raised. Every other
write is ack-paced and cannot collide, so a missing echo there still means an
RX that is not on the line, and still raises.

Both gates green on Windows (31/31 m328p, 16/16 t13a); on hardware the
reconnect now converges on the first knock, the surviving prompt being all the
handshake needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 16:29:17 +02:00
5d520a1ff9 pbhw: --one-wire never reached the suite's own sessions
The flag was plumbed through pbrig.Deployment to the host-tool subprocess
calls and nowhere else, so identity() and scan() opened a raw port and drove
a shared line as though it were two wires. On real one-wire hardware the
adapter's echo answers the knock before the device does, so the suite would
have died at its very first check — "the loader never answered; nothing below
can be trusted" — for the one deployment the flag exists to test, and every
result after it is gated on that check passing.

Both now open through pbrig.Rig.open_port(), which applies the deployment's
link mode. The gap underneath was that only the subprocess path could reach
those facts at all; anything driving the protocol in-process had to restate
them, and did not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:25:46 +02:00
3 changed files with 101 additions and 22 deletions

View File

@@ -454,11 +454,26 @@ class OneWirePort:
def __init__(self, port):
self._port = port
self._pending = b""
self.lost_echoes = 0
def __getattr__(self, name):
return getattr(self._port, name)
def write(self, data):
def write(self, data, blind=False):
"""Put `data` on the line and consume its echo.
`blind` marks the protocol's one multi-byte write with no ack between
its bytes — the knock. Aimed at a loader already in session, its first
byte draws a prompt while the second is still going out, and on real
wiring the device's push-pull ack **wins the line** against the host's
1 k series resistor: that second byte is *destroyed, not delayed*, and
its echo never comes. Measured on an ATtiny13A at 57600 — the loader
answers a single byte perfectly and loses the knock's second every
time. So on a blind write a missing echo is a property of the wiring
rather than a fault in it, and the caller's retry is what deals with
it. Every other write is ack-paced and cannot collide, so a missing
echo there really is an RX that is not on the line.
"""
data = bytes(data)
self._port.write(data)
# The echo arrives at line rate — 10 bits a byte — plus adapter
@@ -466,16 +481,39 @@ class OneWirePort:
# of the error path.
deadline = time.monotonic() + 10 * len(data) / self._port.baud + 0.5
remaining = data
while remaining:
budget = deadline - time.monotonic()
if budget <= 0:
raise Error(f"one-wire echo missing after {len(data) - len(remaining)} of "
f"{len(data)} byte(s) — is the adapter's RX tied to the line?")
byte = self._port.read_exact(1, budget)
if byte == remaining[:1]:
while remaining and time.monotonic() < deadline:
# Speculative, so it cannot be read_exact, whose contract is to
# raise: doing that made the diagnosis below unreachable on every
# quiet line and surfaced a bare "timeout: got 0 of 1 bytes" in
# its place — the one message this class exists to replace.
for byte in self._port.read_available(0.02):
if remaining and byte == remaining[0]:
remaining = remaining[1:]
else:
self._pending += byte
self._pending += bytes((byte,))
if not remaining:
return
if not blind:
raise Error(f"one-wire echo missing after {len(data) - len(remaining)} of "
f"{len(data)} byte(s) — is the adapter's RX tied to the line?")
self.lost_echoes += len(remaining)
# Which loss this is matters, and the count says it. *Some* bytes lost is
# the device's ack winning the line against the host's series resistor —
# ordinary, and what the retry absorbs. *Every* byte lost is nothing
# coming back at all, which is a line that is not free: an application
# holding the shared pin low (this rig's LED demo ends that way), a
# wedge, or an RX that is not on the line. Same retry either way, but
# blaming an ack that never happened sends the reader to the wrong place.
if len(remaining) == len(data):
verbose(f"one-wire: none of {len(data)} byte(s) echoed — the line is not "
f"coming back. Held low by something? (a pin driven low, a wedge, "
f"or an RX not on the line)")
else:
verbose(f"one-wire: {len(remaining)} of {len(data)} knock byte(s) lost to the "
f"device's ack; retrying")
def write_blind(self, data):
self.write(data, blind=True)
def read_exact(self, count, timeout):
taken, self._pending = self._pending[:count], self._pending[count:]
@@ -658,9 +696,16 @@ class Loader:
break
knocks = 0
refusal = None
# The knock is the only write in the protocol with no ack between its
# bytes, so on a shared line it is the only one whose echo may
# legitimately not come back — the device's ack collides with it and
# wins (OneWirePort.write). Losing a byte here is what the retry below
# is for; raising instead aborted the loop before it ever ran, which on
# real wiring made every reconnect into a live session fail.
knock_out = getattr(self.port, "write_blind", self.port.write)
while True:
self.port.flush_input()
self.port.write(knock)
knock_out(knock)
knocks += 1
if PROMPT in self.port.read_available(0.4):
# Settle: absorb a real loader's trailing bytes before asking

View File

@@ -53,7 +53,7 @@ class Suite:
"""The info block, which every later check takes its bounds from."""
module = pbrig.load_pureboot(self.rig.d.pureboot)
self.rig.reset()
port = module.Port(self.rig.d.port, self.rig.d.baud)
port = self.rig.open_port() # wrapped for the echo where the line is shared
try:
loader = module.Loader(port)
if self.rig.d.autobaud:
@@ -87,7 +87,10 @@ class Suite:
rate = module.scan_rate(self.rig.d.baud, pct)
self.rig.reset()
try:
port = module.Port(self.rig.d.port, rate)
# Same wrap as identity(): on a shared line an undiscarded
# echo answers every rate a scan probes, so the walk would
# report the first one it tried.
port = self.rig.open_port(rate)
except module.Error as error:
self.check("scan opens every probe rate", False, f"{rate} Bd: {error}")
return
@@ -129,18 +132,28 @@ class Suite:
got = erased.read_bytes() if erased.exists() else b""
self.check("EEPROM erase leaves 0xff", got == b"\xff" * size, f"{len(got)} B")
def application(self, info, app: pathlib.Path, marker: str) -> None:
def application(self, info, app: pathlib.Path, marker: str,
marker_wait: float = 2.5) -> None:
rc, out = self.rig.pureboot("--flash", str(app), "--verify-flash", str(app))
self.check(f"application flash + verify ({app.name})", rc == 0, self._brief(out))
if marker:
# The tool hands over as it ends its session, so the application is
# already running; opening the port does not reset a board whose DTR
# is unwired, so this simply listens.
data = self.rig.capture(seconds=2.5)
# already running — but only on a board whose DTR is unwired, where
# opening a port simply listens. Where DTR *is* wired to reset (an
# Arduino, most USB-serial dev boards), this open resets the part
# and the activation window comes first, so a marker emitted once at
# startup happens on the far side of a wait this cannot know the
# length of: the window is a compile-time constant and nothing on
# the wire reports it. Hence --marker-wait, and a fixture that
# repeats its banner (PUREBOOT_HEARTBEAT) rather than saying it once.
data = self.rig.capture(seconds=marker_wait)
seen = marker.encode() in data
sample = "".join(chr(b) if 32 <= b < 127 else "." for b in data[:40])
self.check(f"application runs (emits {marker!r})", seen, f"|{sample}|")
self.check(f"application runs (emits {marker!r})", seen,
f"|{sample}|" if seen or data else
f"nothing in {marker_wait:g} s — if this board resets when its port "
f"opens, that wait has to outlast the activation window")
back = self.work / "app-back.bin"
rc, out = self.rig.pureboot("--read-flash", str(back))
@@ -187,7 +200,7 @@ class Suite:
# ------------------------------------------------------------------- run
def run(self, app: pathlib.Path | None, loader_image: pathlib.Path | None,
marker: str) -> int:
marker: str, marker_wait: float = 2.5) -> int:
print("identity")
info = self.identity()
if info is None:
@@ -203,7 +216,7 @@ class Suite:
if app:
print("\napplication")
self.application(info, app, marker)
self.application(info, app, marker, marker_wait)
else:
print("\nskip application checks (pass --app <image.hex>)")
@@ -229,6 +242,10 @@ def main(argv: list[str] | None = None) -> int:
help="the resident loader's .bin, to prove the slot survives an erase")
parser.add_argument("--marker", default="",
help="text the application emits when it runs, e.g. APP")
parser.add_argument("--marker-wait", type=float, default=2.5,
help="seconds to listen for it. On a board whose DTR is wired to "
"reset, opening the port resets the part, so this must outlast "
"the activation window (default 2.5)")
args = parser.parse_args(argv)
rig = pbrig.Rig(pbrig.Deployment.from_args(args))
@@ -236,7 +253,8 @@ def main(argv: list[str] | None = None) -> int:
f"{' (autobaud)' if args.autobaud else ''}")
print("this overwrites the application flash and EEPROM\n")
with tempfile.TemporaryDirectory(prefix="pbhw-") as temporary:
return Suite(rig, pathlib.Path(temporary)).run(args.app, args.loader, args.marker)
return Suite(rig, pathlib.Path(temporary)).run(args.app, args.loader, args.marker,
args.marker_wait)
if __name__ == "__main__":

View File

@@ -281,6 +281,22 @@ class Rig:
return 99, f"TIMEOUT after {timeout}s\n{expired.stdout or ''}{expired.stderr or ''}"
return result.returncode, (result.stdout or "") + (result.stderr or "")
def open_port(self, baud: int | None = None):
"""A port opened the way this deployment says to speak to the board.
Everything the rig runs as a *subprocess* gets its flags from
`pureboot()` above; anything that drives the protocol in-process has
to reach the same facts, and until this existed only the subprocess
path could. A shared line is the one where that gap is fatal rather
than untidy: the host reads back every byte it writes, so an
undiscarded echo answers the knock before the device does. Open
through here and a one-wire deployment cannot be silently driven as
a two-wire one.
"""
module = load_pureboot(self.d.pureboot)
port = module.Port(self.d.port, self.d.baud if baud is None else baud)
return module.OneWirePort(port) if self.d.one_wire else port
def capture(self, seconds: float = 2.0, baud: int | None = None) -> bytes:
"""Listen to whatever the board is saying, at an arbitrary rate.