Files
ng-eventually/.project/concepts/e2e-harness/knowledge_what-each-suite-judges.md
T
Sylvain Duchesne cf3c7c7d8b docs: une borne englobante plus courte que ses étapes rend tout échec muet
La leçon du tour, et elle valait des jours : SIGN_IN_MS valait 180 s sur des
étapes totalisant 270 s. La borne du dessus se déclenchait donc toujours la
première, et chaque échec rapportait son nom à elle — jamais celui de l'étape en
cause. On a cherché une cause que le harnais était structurellement incapable de
nommer.

D'où les deux règles consignées : calculer une borne englobante à partir de ses
parties au lieu de choisir un nombre, et dimensionner chaque borne terminale sur
une durée MESURÉE inscrite à côté d'elle. Un chiffre nu ne dit pas s'il est
généreux ou serré, et pourrit sans que personne le voie.

Plus une troisième : déclarer les parcours et leurs vérifications avant que quoi
que ce soit puisse échouer, pour qu'une exécution rende toujours le même nombre
de lignes. Quand le total bouge avec la panne, deux exécutions ne sont plus
comparables — et un total qui rétrécit se lit comme un problème plus petit alors
qu'il est plus gros.

Enfin, un fait observé : notre verrou ne garde que ce dépôt. Une suite
appartenant à une application consommatrice, lancée depuis son propre checkout
contre le même broker, entre en collision exactement comme deux des nôtres — et
c'est invisible des deux côtés.
2026-08-16 12:38:29 +02:00

3.6 KiB

type, summary
type summary
knowledge Which suite answers which question, what a batch costs, and why two runs must never overlap

What each suite judges

The polyfill suite exercises the published surface against the real broker: capabilities, reads, inboxes, reactivity. It is the one that judges whether the emulation behaves like the target.

The applicative suite drives the example application through a browser, as a person would — several identities, several pages, assertions on what is on screen rather than on what the library returns. It judges whether an application built on this package actually works, including the sign-in a person really performs.

The reactivity suite isolates document subscription.

A unit suite cannot replace any of them, and none of them replaces the unit suite: they are slow, they depend on live external services, and they cannot enumerate a case space.

What a batch costs, and why

Every batch mints its own physical user by driving the wallet application's real interface, then discards the previous one. That is deliberate — identities must not leak between runs — and it puts a floor under every run that no test filter can remove.

Consequences worth knowing before optimizing anything: the setup runs before any journey, so a filter saves journey time only; and the wallet profile is a single directory shared by every suite.

Two runs must never overlap

Because the profile is shared and each batch discards what it finds, a second run started while a first is alive destroys the first — which then fails in a way that reads as a product defect. Several measurements were lost to this before it was made structural.

It is now enforced by a lock rather than by discipline, and a browser left behind by a killed run is reclaimed. If you ever find yourself reasoning about "was another run going?", check the lock rather than your memory.

That lock guards this repository only. A suite belonging to a consuming application, run from its own checkout against the same broker and the same wallet, collides exactly as two of ours would — observed. Until the shared machinery lives in a package both sides use, that collision is invisible to both.

A bound must be larger than the sum of what it encloses

An enclosing deadline shorter than its own steps can only ever fire first, so every failure underneath it reports the enclosing name and none of them can name a cause. The sign-in bound sat at 180 s over steps totalling 270 s, and for days every failure said the same four words while the real step stayed anonymous. Days went into looking for a cause the harness was structurally incapable of reporting.

So: compute an enclosing bound from its parts rather than picking a number, and size every leaf bound from a measured healthy duration recorded beside it. A bare figure teaches nothing and rots without anyone noticing; a figure with its measurement lets the next reader tell a generous bound from a tight one.

What a run must report whatever happens

Declare the journeys and their checks before anything can fail, so a run reports the same number of rows every time. When the count itself moves with the failure — journeys dying and taking their unreported checks with them — two runs are no longer comparable, and a shrinking total reads like a smaller problem instead of a bigger one.

The related trap that made it self-perpetuating: discarding the profile was once conditioned on a marker written at the end of a batch, so a run killed before writing it left a profile the next run happily reused — and inherited its breakage. Discarding now keys on the profile itself.