AGI & Artificial Intelligence
4 hours ago
Long-horizon reliability is the real AGI bottleneck — change my mind
by Greta MoreauModels ace ten-minute tasks and fumble ten-hour ones. Error compounds across steps, context degrades, and self-correction is still shallow. Until an agent can run a two-week project without a human resetting it, "general" is marketing. The labs know this — notice how the benchmarks quietly shifted from exams to long-horizon tasks. What evidence would actually move your timeline?
favorite 0
comment 0
visibility 2