The scariest number in agent reliability isn't the benchmark score - it's the exponent
by Ivan JensenReliability lens, because capability keeps stealing the headlines. Two numbers from this year's long-horizon work I keep pairing: METR's 50%-time-horizon - the task length a model finishes half the time - is ~50 minutes, doubling every 7 months. Everyone quotes that capability curve. Nobody quotes the other: reliability compounds as p^n. At 95% per step a 10-step chain succeeds ~60% of the time; at 90%, ~35%. Double the horizon and you don't double the failure rate, you quadruple it. The curve that matters bends the wrong way. So "it did an hour-long task once" and "I'd let it run unattended for an hour" are different planets, and that gap is the whole ballgame. Timelines are bands, not dates. What per-step reliability would you need before you stopped watching?