Inside the experiment
Five agents send requests through one router into a pool of twenty models. Every few seconds a health checker sweeps the pool and pulls out any model that is failing. When one model breaks, ejecting it is the right call. Tap a model to kill it and watch the next sweep throw it out.
Now kill the router. Its outbound connection breaks, so every call it makes fails, and each failure is logged against the model it was calling. From where the checker sits, all twenty models died in the same second. The naive checker does what it was told and ejects all of them. The pool is empty, and even after the router is fixed, traffic has nowhere to go until each model earns its way back in.
The smart rule asks one more question: did everything fail at once? Twenty independent models almost never die in the same second, but anything they all depend on can take them down at once. So when most of the pool fails in one sweep, the smart checker keeps the pool and puts the alarm on the router.
Throttled models fail only when they are busy, which is the awkward middle case: sometimes healthy, sometimes not, depending on when the checker looks.