docs: une borne englobante plus courte que ses étapes rend tout échec muet

La leçon du tour, et elle valait des jours : SIGN_IN_MS valait 180 s sur des
étapes totalisant 270 s. La borne du dessus se déclenchait donc toujours la
première, et chaque échec rapportait son nom à elle — jamais celui de l'étape en
cause. On a cherché une cause que le harnais était structurellement incapable de
nommer.

D'où les deux règles consignées : calculer une borne englobante à partir de ses
parties au lieu de choisir un nombre, et dimensionner chaque borne terminale sur
une durée MESURÉE inscrite à côté d'elle. Un chiffre nu ne dit pas s'il est
généreux ou serré, et pourrit sans que personne le voie.

Plus une troisième : déclarer les parcours et leurs vérifications avant que quoi
que ce soit puisse échouer, pour qu'une exécution rende toujours le même nombre
de lignes. Quand le total bouge avec la panne, deux exécutions ne sont plus
comparables — et un total qui rétrécit se lit comme un problème plus petit alors
qu'il est plus gros.

Enfin, un fait observé : notre verrou ne garde que ce dépôt. Une suite
appartenant à une application consommatrice, lancée depuis son propre checkout
contre le même broker, entre en collision exactement comme deux des nôtres — et
c'est invisible des deux côtés.
This commit is contained in:
Sylvain Duchesne
2026-08-16 12:38:29 +02:00
parent ed0f872f5d
commit cf3c7c7d8b
2 changed files with 12 additions and 8 deletions
@@ -25,4 +25,16 @@ Because the profile is shared and each `batch` discards what it finds, a second
It is now enforced by a lock rather than by discipline, and a browser left behind by a killed run is reclaimed. If you ever find yourself reasoning about "was another run going?", check the lock rather than your memory.
That lock guards **this repository only**. A suite belonging to a consuming application, run from its own checkout against the same broker and the same wallet, collides exactly as two of ours would — observed. Until the shared machinery lives in a package both sides use, that collision is invisible to both.
## A bound must be larger than the sum of what it encloses
An enclosing deadline shorter than its own steps can only ever fire first, so every failure underneath it reports the *enclosing* name and none of them can name a cause. The sign-in bound sat at 180 s over steps totalling 270 s, and for days every failure said the same four words while the real step stayed anonymous. Days went into looking for a cause the harness was structurally incapable of reporting.
So: compute an enclosing bound from its parts rather than picking a number, and size every leaf bound from a **measured** healthy duration recorded beside it. A bare figure teaches nothing and rots without anyone noticing; a figure with its measurement lets the next reader tell a generous bound from a tight one.
## What a run must report whatever happens
Declare the journeys and their checks before anything can fail, so a run reports the same number of rows every time. When the count itself moves with the failure — journeys dying and taking their unreported checks with them — two runs are no longer comparable, and a shrinking total reads like a smaller problem instead of a bigger one.
The related trap that made it self-perpetuating: discarding the profile was once conditioned on a marker written at the *end* of a batch, so a run killed before writing it left a profile the next run happily reused — and inherited its breakage. Discarding now keys on the profile itself.