Explainer · AI Agents & Automation
What is an AI agent?
Short answer
An AI agent is a system in which a language model decides its own next step towards a goal: it plans, calls a tool (search, code, an API), checks the result and repeats until done or until it needs you. Unlike a workflow, its steps are not fixed in advance. Agents suit open-ended tasks such as coding or customer support, cost more than simpler setups and need guardrails against prompt injection and excessive permissions.

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.
An AI agent is a program in which a language model does not just answer, but decides what to do next: it picks a tool, uses it, looks at the result and repeats until the task is done or it needs your input. A chatbot replies to one message at a time. An agent works towards a goal over several steps.
Two definitions from the companies building them
Google Cloud defines AI agents as "software systems that use AI to pursue goals and complete tasks on behalf of users", showing reasoning, planning and memory, with "a level of autonomy to make decisions, learn, and adapt".
Anthropic draws a line that is useful in practice, between two kinds of "agentic systems":
- Workflows: language models and tools "orchestrated through predefined code paths". The steps are decided in advance by a developer.
- Agents: systems where the model "dynamically direct[s] [its] own processes and tool usage, maintaining control over how [it accomplishes] tasks".
Both definitions agree on the core: a goal, tools, and a loop in which the model chooses the next action.
Agent, assistant or bot?
Google Cloud separates three things that are often mixed up:
| Bot | AI assistant | AI agent | |
|---|---|---|---|
| Autonomy | Lowest: follows pre-programmed rules | Needs your input and direction | Highest: makes decisions to reach a goal |
| Typical task | Fixed questions and answers | Helping you, step by step | Multi-step tasks and workflows |
The same product can be all three. A chat assistant becomes an agent when it is allowed to plan and use tools on its own, such as searching, writing files or running code, before coming back to you.
How an agent works, step by step
- It receives a goal, as a command or in a conversation with you.
- It plans what to do first.
- It acts by calling a tool: a search, an API, a database query, a code run.
- It observes the result. Anthropic stresses that agents need "ground truth" from their environment at each step, such as tool results or code execution, to judge progress.
- It decides the next step, repeating until the goal is met, a stopping condition is hit (for example, a maximum number of steps) or it needs your judgement.
This reason-act-observe loop goes back to research such as ReAct (Yao and colleagues, 2022), which combined reasoning and acting in language models, and which Google Cloud cites as the basis of agents' key features.
Where the tools come from
An agent is only as useful as the tools it can reach. Increasingly those tools are connected through the Model Context Protocol (MCP), an open standard that lets one tool server work with any compatible AI app.
Real examples
- Coding agents that read a codebase, edit several files and run the tests. Anthropic reports that its agents can solve real GitHub issues from the SWE-bench Verified benchmark from the issue description alone, while noting that "human review remains crucial".
- Customer support agents that combine a chat interface with actions such as looking up an order or issuing a refund.
- Research agents that search, read and summarize several sources before answering.
Anthropic's description of where agents have worked best fits all three: tasks that need "both conversation and action", have clear success criteria and keep a human in the loop.
When you do not need an agent
Anthropic's advice to developers is to find "the simplest solution possible" and add complexity only when needed: "this might mean not building agentic systems at all". Agents trade speed and cost for better results on open-ended tasks, and their autonomy brings "the potential for compounding errors", which is why Anthropic recommends "extensive testing in sandboxed environments, along with the appropriate guardrails". For a fixed, repeatable process, a predefined workflow is often cheaper and more predictable.
The main risks
The OWASP Top 10 for LLM applications (2025) lists two risks that apply directly to agents:
- Prompt injection (LLM01): instructions hidden in a web page, email or document the agent reads can change what it does.
- Excessive agency (LLM06): giving the agent more permissions or autonomy than the task needs, so a mistake or an injected instruction can do real damage.
The practical answer is the same in both cases: give the agent only the tools and permissions the task needs, require confirmation before irreversible actions and log what it does.
For costs, use cases and how to decide whether an agent is worth it for your company, see AI agents for business.
What we checked
- Google Cloud defines AI agents as software systems that use AI to pursue goals and complete tasks on behalf of users, with reasoning, planning, memory and a level of autonomy, and distinguishes agents, assistants and bots by autonomy. (Google Cloud, )
- Anthropic distinguishes workflows (predefined code paths) from agents (models directing their own processes and tool use), recommends the simplest solution possible, warns of higher costs and compounding errors, and recommends sandboxed testing and guardrails. (Anthropic, )
- Anthropic reports agents solving SWE-bench Verified GitHub issues from the description alone, with human review still crucial, and describes customer support agents that issue refunds or update tickets. (Anthropic, )
- ReAct (Yao et al., 2022) combined reasoning and acting in language models. (arXiv, )
- Prompt injection (LLM01) and excessive agency (LLM06) are in OWASP's 2025 Top 10 for LLM applications. (OWASP GenAI Security Project, )
- MCP is an open standard for connecting AI applications to tools and data. (Model Context Protocol, )
What may change
- The vocabulary is still settling: vendors use 'agent' for different levels of autonomy.
- Agent capabilities and benchmark results change with each new model.
Frequently asked questions
What is the difference between an AI agent and a chatbot?
A chatbot answers one message at a time. An agent pursues a goal over several steps, choosing tools and actions by itself and checking the results, and only comes back to you when it is done or needs a decision.
What is the difference between an agent and a workflow?
In Anthropic's terms, a workflow runs models and tools through steps predefined in code, while an agent lets the model decide its own steps and tool use. Workflows are more predictable; agents are more flexible.
What are examples of AI agents?
Coding agents that edit files and run tests, customer support agents that look up orders and issue refunds, and research agents that search and summarize several sources.
Are AI agents safe?
They carry specific risks, such as prompt injection and excessive agency (OWASP's 2025 Top 10). Limit their permissions to what the task needs, require confirmation for irreversible actions and keep logs.
Sources
- What is an AI agent? Definition, features and types, Google Cloud. Accessed October 1, 2026.
- Building effective agents, Anthropic. Accessed October 1, 2026.
- ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al., 2022), arXiv. Accessed October 1, 2026.
- OWASP Top 10 for LLM Applications 2025, OWASP GenAI Security Project. Accessed October 1, 2026.
- What is the Model Context Protocol (MCP)?, Model Context Protocol. Accessed October 1, 2026.
Spotted an error or an outdated price? Tell us and we will fix it.
Change history
- : First published, with definitions from Google Cloud and Anthropic and risks from OWASP's 2025 list.
Next review: .