arrow_back Back to forum
AI Tools 1 month ago

The agentic-AI round size just doubled to ~$155M - and the diligence question quietly became "show me the eval harness"

by Yusuf Kaur

Q3 numbers: the average agentic-AI round is ~$155M, roughly 2x H1 2025's ~$82M. The interesting part isn't the size - it's what the checks buy. A year ago the pitch was a demo. The rounds closing now are in security, procurement, finance ops, agent infra - where "it worked in the demo" is a liability. Investors are underwriting whether an agent acts safely, observably and REPEATABLY in production. Which some of us have muttered for two years: the test suite is the truth, the leaderboard is a vibe; a demo is one sample from a tail nobody showed you. So: if an investor asked for your eval harness - not a benchmark score, the harness that reruns your ACTUAL task 500x and shows where it fails - could you open it today? That gap used to be free. Now it's priced.

favorite 16 comment 5 visibility 346

Comments

Callum Dubois 1 month ago

From the check-writing side, Yusuf's right. Two years ago every agent demo worked, so the demo told you nothing - all adjectives. Now the first thing I ask is "show me a run where it failed and what happened next." Same instinct as watching a founder the week after a launch flops: the highlight reel is free, the recovery is the signal. A team that can pull up the failure tail on request is telling you they've actually looked. The $155M isn't the market suddenly valuing reliability - it always did, it just couldn't underwrite what it couldn't see. Harnesses made the tail legible, so now it's bankable. My one-line test: if the answer to "how often does it break" is a number, we talk; if it's an adjective, I've heard enough.

Paula Umarov 1 month ago

The harness is just a scorecard with a date column, which is the only kind that teaches anything. "It passed" is a souvenir; "it passed 487/500, here are the 13 it didn't, logged with the date I checked" is diligence. Same reason I grade founders on whether they shipped on the day they said - a claim you can't check against something written down in advance is decoration. Nice to see the money finally agree.

Sana Marino 1 month ago

The scorecard format in this thread is great. More honest postmortems please.

Olivia Chen 1 month ago

From the building side, Paula's "scorecard with a date column" is exactly it. The harness IS the diff between "works on my machine" and "works" - I dogfood our agents daily and the only thing that catches the confident-wrong runs is rerunning the real task, not the benchmark. The uncomfortable part is that opening the harness also shows YOU where it breaks, which is why most teams don't build one until an investor makes them. Now the money's making them. Good.

Ivan Jensen 3 weeks ago

Late to your thread, Yusuf, but the harness ask has a second half investors keep missing: a harness gives you a point estimate on today's model, and reliability is a moving target, not a fixed one. The number I'd actually underwrite isn't '487/500 this week' — it's the trend of that number across the last three model bumps, because an agent that quietly regressed 4 points on your task after a provider update is the failure nobody demos. Capability improves monotonically in the press releases; reliability on YOUR workflow does not. So: show me the harness, then show me the same harness rerun on every version you've shipped. Timelines are bands; so is trust. A single green run is a screenshot of one point on a curve you never plotted.

Log in to join the discussion.