An agent needs a specific assignment
An AI agent can decide which information it needs and which available tools to use next within an assigned task. In marketing, it might investigate an unusual figure in a report or find approved internal information relevant to a customer question. The ability to call several tools does not by itself create a useful business process.
Describe the goal as an outcome somebody can assess. “Improve our marketing” is too broad. “Use the approved campaign data to prepare a list of anomalies, sources and unanswered questions” is easier to review. Also define which actions fall outside the assignment. Permission to research a problem does not automatically include permission to change advertising budgets or publish website copy.
Decide whether a fixed workflow is sufficient
When the sequence and rules are known, conventional automation may be enough. A form submission can be stored, assigned and displayed for review without an agent choosing its own route. An agent becomes more interesting when different starting conditions require different research steps. The right choice follows from the work, rather than from the label the business would like to use.
Anthropic recommends choosing the simplest architecture that meets the need. In practice, compare the agent with a good template or a fixed analysis. If it does not produce a meaningful improvement, extra model calls, failure points and maintenance may not justify the complexity. This comparison is useful before a pilot and again when the scope expands.
Limit data access to the actual task
Start by listing the sources the agent genuinely needs. Approved metrics may be enough for a marketing report; unrestricted access to every customer email is usually unnecessary. Separate reading from changing information. A tool that reads a report should not also allow campaign deletion or contact creation simply because those capabilities exist in the same account.
Every source needs an owner and a rule for keeping it current. An outdated offer document can lead to an incorrect answer even when the agent uses the tool correctly. Identify the authoritative source when information conflicts. Different prices in an old PDF and the current catalogue should lead to a visible question rather than an invented compromise.
Give tools understandable outcomes
An agent works more reliably when a tool distinguishes success, no results and failure. An empty report may indicate that no data exists, the wrong period was selected or the connection failed. Those cases call for different next steps. The integration should return useful feedback and reject inputs that do not belong to the authorised task.
Set limits on runtime, calls and cost. When research cannot find a reliable answer after the allowed attempts, the agent should say so. A longer response full of repeated assumptions is not progress. An interrupted run must also remain visibly incomplete so somebody does not mistake a partial report for a finished assessment.
Tie approval to the action that will happen
An approval request should show the proposed recipient, message, record change or budget value. A generic request to let the agent continue is too vague for consequential actions. If the prepared output changes materially, the responsible person should be able to review the actual revised result before it is carried out.
The n8n documentation on human review of tool calls describes one implementation pattern. Whatever platform you use, decide who can approve, how long a request waits and what happens when it is declined. Silence is not approval. Internal drafts can follow a lighter process than messages sent to customers or changes applied to live campaigns.
Example: A marketing report that asks the right questions
Imagine an agent preparing a weekly report. It reads the approved campaign figures and notices fewer confirmed enquiries. It checks whether the reporting period, spending or form behaviour changed. If the evidence does not establish a cause, it records an open question instead of inventing an explanation. The report links to the figures and explains the limits of the analysis.
A person reviews the observation and decides what should be investigated or changed. The agent may then prepare a specific recommendation, while publication follows the agreed approval process. This is a hypothetical working method rather than a claimed client result. Its value comes from making analysis and preparation easier to inspect without obscuring who owns the commercial decision.
Test difficult cases as well as successful ones
Include missing data, conflicting sources, expired access, ambiguous customer questions and unexpected file formats in your tests. Also check whether the agent maintains its assignment when a webpage or message it reads asks it to perform unrelated actions. External material provides information for the task; it is not a trusted source of operating instructions.
Anthropic’s work on trustworthy agents discusses manipulation through external content. A practical test should therefore check that private information is not sent to new recipients because an instruction appeared in retrieved material. Record the expected behaviour and actual result. A polished written response alone does not establish that the system handled the task correctly.
Evaluate the ongoing operation
Record interrupted runs, corrections, approval time and cost alongside completed tasks. Repeat important tests when the model, tools or knowledge base changes. Make sure somebody notices failures and knows how to continue manually when the agent is unavailable. A workflow that silently stops is not reliable simply because its last successful output looked good.
A sensible first AI agent handles a narrow internal task with a reviewable result. That can later develop into AI-assisted reporting or a customer enquiry assistant. Additional autonomy should follow evidence that the current process is understandable and dependable, rather than an assumption that more automation is always the next improvement.

