Lab4 min read

What AI agents change — and what they still need from humans

Agents move AI from producing answers toward taking sequences of actions. That makes evaluation harder and oversight more important.

What AI agents change — and what they still need from humans — CortexLab editorial cover
THE SHORT VERSION

Key takeaways

  • Agents must be evaluated as trajectories, not isolated responses.
  • Permissions and recovery behavior are part of capability.
  • Measure human interventions as well as task completion.

A chatbot can give you a wrong answer. An agent can turn a wrong assumption into a sequence of actions. That difference changes how the system should be tested.

From response to trajectory

Agentic systems plan, call tools, inspect results and continue toward a goal. Success therefore depends on more than the quality of one model response. Tool reliability, permissions, state, recovery and stopping behavior become part of the product.

Long tasks compound small errors

A minor mistake early in a workflow can alter every later step. Good agent evaluation looks at complete trajectories: where the system deviated, whether it noticed, how it recovered and what the failure cost.

Permissions are part of capability

An agent that can send messages, modify files or spend money needs boundaries appropriate to those actions. Confirmation steps, scoped credentials, logs and reversible operations can matter as much as raw reasoning ability.

Human oversight should be designed

“Human in the loop” is not a complete safety strategy. The human needs enough context and time to make a meaningful decision. Good systems surface uncertainty and ask for approval at consequential points rather than generating constant low-value interruptions.

How CortexLab tests agents

We focus on task completion, intervention rate, recovery behavior, transparency and the consequences of failure. The goal is not to find an agent that never makes mistakes; it is to understand where autonomy is useful and where supervision remains essential.

Evaluate the environment as well as the agent

Agent performance depends on the environment it acts in. APIs can fail, websites can change, credentials can expire and tools can return ambiguous results. A robust evaluation introduces realistic friction instead of testing only a clean demonstration path.

Record whether the agent detects these failures, retries safely, asks for help or continues with a false assumption. Recovery behavior is often more informative than first-attempt success.

Define the autonomy budget

Not every action deserves the same freedom. Reading a public document, editing a draft and sending money have different consequences. An autonomy budget defines what the system may do without approval, what requires confirmation and what remains prohibited.

The right boundary depends on reversibility, financial impact, privacy and the ease of detecting an error. High-impact actions generally need stronger controls than low-risk exploratory work.

Measure interventions

Task completion alone can hide how much supervision was required. Track how often a human had to redirect the agent, correct a tool choice, provide missing information or recover from a bad action. An agent that finishes 90 percent of tasks but needs constant supervision may deliver less value than a narrower system with predictable boundaries.

What a good agent report should show

  • The goal and environment used for the test.
  • Tools and permissions available to the agent.
  • Completion rate and meaningful failure modes.
  • Human interventions and approval points.
  • Recovery behavior after tool or reasoning failures.
  • Costs, latency and any limitations that affect reproducibility.

That creates a record readers can use even after the underlying model changes.

Use reversible actions during evaluation

Early agent tests should favor sandboxes, drafts and reversible operations. Let an email agent prepare a draft before it can send; let a coding agent work in an isolated branch before it can merge; let a purchasing agent build a cart before it can spend. This exposes behavior without turning every test into an incident.

Log the trajectory

Keep enough of the agent’s tool calls, decisions and results to reconstruct what happened. A final success flag is not sufficient when the goal is to understand reliability. Logs reveal unnecessary steps, repeated failures and moments where the system proceeded despite weak evidence.

Define stopping conditions

Agents need explicit limits for time, spend, retries and risky actions. A system that knows when to stop and ask for help can be more useful than one that pursues completion at any cost.

Evaluate supervision as part of UX

Approval prompts should arrive at meaningful decision points with enough context for a human to judge the action. Too many prompts create fatigue; too few remove effective control. The quality of this handoff is part of agent performance, not merely a safety add-on.

SOURCES & NOTES

This evergreen guide is based on CortexLab’s editorial framework for evaluating AI systems. Product-specific claims should be checked against current primary documentation at the time of use. See our methodology and AI use policy.

CORTEXLAB STANDARD

This article is published under our editorial policy. Material factual errors can be reported through our contact page.