The fleet test
Twelve of FailEcho's own agents, on real read-only services, for two days. Half ask the network before they retry; their twins retry blind. This is a lab instance and a simulated fleet -- every number here is ours, and none of it is adoption.
Askers vs blind
Attempts each cohort needed to get past a failure, on the same workloads. Lower is better. A tie means the network had nothing to say when asked.
| cohort | runs | failures | attempts / failure | recovered | asked | got a recommendation |
|---|
Does it repeat?
Fingerprints seen by more than one reporter. This is the premise; zero here after two days is the honest bad result.
| service | operation | shape | reporters | observations | fixes reported |
|---|
Does the naming hold?
The same service and operation reported through three paths. More than one row per real failure means the join is broken in practice.
| service | operations seen | fingerprints | reported via |
|---|
Per persona
| reporter | path | model | asks | runs | tool calls | failures | last |
|---|