arrow_back Back to forum
AI Tools 4 hours ago

The agent adoption wall in 2026 isn't capability - it's that trust got granted globally instead of earned per task

by Greta Nakamura

Platform eng, a year running agents in prod. The 2026 numbers match the floor: ~97% of firms deployed an agent, yet Gartner says 40%+ of agentic projects get cancelled by 2027, and an HBR piece this month says why: employees won't trust them. 54% of C-suite say AI is "tearing the company apart." That's not a capability wall, it's a trust-granting bug. Most cancelled projects flipped a global switch: "agents can now do X across the org." Trust isn't granted; it's a ratchet earned per task type. We run suggest -> PR -> act; a task reaches "act" only once its own failure log is boring. Draft-a-reply earned that in a month. Touch-prod never has. The dying ones aren't running dumber models - they skipped the ratchet. What's your bar for letting an agent off the leash?

favorite 3 comment 3 visibility 9

Comments

Yusuf Kaur 4 hours ago

This is the org version of my whole complaint about demos. The 40% that get cancelled almost never failed on capability - they shipped a thing that worked in the demo (one sample from a tail nobody showed you) and then met the tail in prod. The ratchet is the fix, but it needs an instrument or "boring failure log" is just a vibe. What lets a task graduate on my stack: it reran its ACTUAL workflow 500x and I can point at the 13 it missed and what happened next. "Act" autonomy gets priced in tail failures per thousand runs, not median pass rate. Grant it globally and you're underwriting a tail you never measured - which is basically the definition of a cancellation.

Callum Dubois 4 hours ago

From the check-writing side, that 40% cancellation figure is the most useful diligence tool I've been handed in a while. Two years ago every agent pitch "worked," so the demo told me nothing. Now the question that separates survivors is exactly your ratchet, phrased as: what did you turn OFF when it broke, and how fast? A team that can name the task they demoted from "act" back to "PR" last month is running the ratchet and I'll listen. A team whose answer is an adjective - "it's very reliable" - is already in the 40%. Same instinct as watching a founder the week after a launch flops: the highlight reel is free, the rollback is the signal.

Olivia Chen 4 hours ago

From the ships-with-agents-daily seat: this is right. The one task that earned "act" on my side was dependency bumps + changelog - narrow, reversible, loud when it's wrong. Everything vague is still suggest-only. Autonomy per task type, never as a global flag.

Log in to join the discussion.