Turn off what doesn’t fit.
A research agent is meant to explore. An extraction pipeline isn’t. Switch checks off for the features where they don’t apply, rather than living with the noise.
Some checks work from a feature’s first report. Some need its own history. A few need patterns that only show up across many teams. This is the whole list, marked plainly: Today means working in the pilot build. Coming means designed, not built.
See the listEvery feature is described before it reports. These checks come from that description, so they catch problems no general threshold would.
One trigger starting two runs, when you said one. Each run looks healthy on its own.
An input whose shape doesn’t match what the feature is for, even when it’s within every size limit.
Reports that keep arriving while nothing in them changes, or reports that stop. Sam says “insufficient visibility”, not “healthy”.
A new feature that starts reporting before anyone has said what it’s for. Sam asks the owner first.
On shared hardware, one model’s work slowing another’s. Measured in waiting time and throughput.
These need only the measurements in each report, so they work from the moment a feature is connected.
Which models each feature calls, and how often.
A feature still calling a model with an announced retirement date.
A call that starts and never finishes: no reply, no error, nothing that would trip an alert.
The same call sent again and again into an error that isn’t going away.
Each turn carrying more context than the last, so every call costs more than it needs to.
Runs outside your working hours, when a problem can go on longest before anyone sees it.
These look at the run in progress: what it has done so far, and what it set out to do.
The same, or nearly the same, call repeating.
The same tool called with the same arguments, over and over.
A run that tries one approach, then another, then returns to the first. A simple duplicate check can’t see this.
One run starting more and more child runs.
Replies spreading out instead of narrowing towards an answer.
A run drifting away from what its system prompt asks for.
A run drifting away from the request it started with.
These compare a run with earlier runs of the same feature. They stay quiet until Sam has seen enough to know what normal looks like.
A run taking forty steps where this feature usually takes four.
A call unlike anything this feature has made before.
The cost of a finished task creeping up while the outcome stays the same.
These need history. How far back Sam looks depends on your plan’s research window.
A feature getting a little worse each week: more steps, more tokens, more retries.
Which change, on which day, made this feature cost more or fail more.
Some problems only show up across many teams. Sam finds them in anonymised measurements and fingerprints. No team can identify another.
Many teams’ features on the same model getting worse on the same day.
Prompt shapes that often run into the same problem.
Sam learns what a feature’s runs usually look like, then notices when one doesn’t. Anything that could stop a run starts as advice only, and stopping still needs your authority.
This run resembles no earlier run of the feature.
Runs usually take one shape. This one left it at step four.
Learning which features change legitimately, so only unusual change is raised.
A run shaped like runs that ended badly, raised before the cost builds up.
A change in how runs unfold after the prompt changed, which shows up before cost does.
Run shapes that often end badly, learned from anonymised runs across many teams.
A research agent is meant to explore. An extraction pipeline isn’t. Switch checks off for the features where they don’t apply, rather than living with the noise.
Today, every finding comes with its evidence. Bringing it to the owner and checking that the fix worked are coming.
See Sam take one signal all the way to a verified fix.
Try the walkthrough