AI agents expand what AI systems can do by moving beyond responses to taking action toward a goal.
For professional teams, that shifts the focus from how convincing a response sounds to whether the system can complete useful tasks reliably. Since AI agents can take actions across multiple steps, use tools, and interact with other systems, it also raises new questions about what the system can do, how much authority it should have, and what happens when something goes wrong.
Understanding those differences starts with what makes an AI system an agent in the first place.
What are AI agents?
An AI agent is a system that can take in information, make decisions, and carry out actions toward a goal. It may work with documents, software applications, data, or even physical environments. What makes it an agent is its ability to use what it observes to decide what to do next.
Modern agents often use a large language model (LLM) to interpret instructions, reason about tasks, and coordinate tools across multiple steps. Unlike a system that follows only a fixed sequence of instructions, an agent can adjust what it does based on the results of previous actions.
The amount of autonomy an agent has depends on how it is designed, which tools it can access, and what permissions it has.
Understanding how agents do this requires a closer look at the components that make them work.
What are the core parts of AI agent architecture?
AI agent architecture connects the information an agent receives with the decisions and actions it can take. The exact design varies, but most modern agents combine several core components: instructions and context, a model or planner, memory or task state, tools, and controls that determine what the system is allowed to do.
The process starts with the agent’s goal, instructions, and available information. A document-review agent, for example, needs the correct files and review criteria rather than a vague instruction to “find problems.”
The model or planner uses that context to decide what to do next. It may break the task into steps, choose a tool, or decide that more information is needed before continuing.
Task context and memory help the agent keep track of information across those steps. Short-term context holds information needed for the current interaction, while saved task state or retrieved prior information can support longer workflows. These forms of memory don't change the model’s learned parameters or weights.
Tools allow the agent to act beyond generating text. Depending on the system, an agent might search a database, retrieve a document, run code, update software, or call another service.
Orchestration code connects these components and coordinates the workflow by managing model calls, storing progress, passing tool results back to the model, invoking approved tools and deciding when the workflow should continue or stop.
Controls & permissions determine which proposed actions are actually allowed. For example, a model may propose sending a message while the surrounding system requires a person to approve it first.
At a high level, the components connect as follows:
- Goal and inputs → model or planner
- Model or planner ↔ task context and memoryProposed action → permission check → approved tool
- Tool result → feedback check → model or planner
- High-impact action → human approval or stop
Together, these components determine how an agent turns a goal into action and where controls enter the process. How those components are arranged also helps distinguish different types of AI agents.
What are the different types of AI agents?
Agents can be categorized based on how they choose actions, respond to information, and learn from feedback. A commonly used classification
include 6 main types of AI agents, each representing different action-selection and learning mechanisms rather than a progression of product maturity. Multi-agent AI represents a separate dimension that's focused on how multiple agents work together.
Simple reflex agents
Simple reflex agents respond to the current observation using condition-action rules without retaining history, such as a thermostat that activates heat below a threshold. Because these agents do not retain information about what happened previously, they work best in predictable environments where the required responses can be defined in advance.
Model-based reflex agents
Model-based reflex agents maintain an internal representation of their environment rather than reacting only to the current input. They update that representation as conditions change and use it to guide their next action.
A cleaning robot, for example, might remember obstacles it has encountered while applying navigation rules. Here, “model-based” refers to the agent's internal representation of its environment, not specifically to a large language model.
Goal-based agents
Goal-based agents make choices that are guided by a stated objective. A route planner, for example, might consider which actions will lead to a destination. Reaching the goal alone doesn't determine how the system should weigh competing preferences such as cost and speed.
Utility-based agents
Utility-based agents go a step further by comparing possible outcomes according to a defined measure of value. For example, a route planner could consider preferences such as travel time, tolls, and fuel use to guide its choice, although incomplete information may prevent it from finding the best real-world option.
Learning agents
Learning agents use feedback to improve how they choose actions over time. They may adjust their behavior based on previous outcomes, errors, or rewards rather than relying only on fixed rules.
Learning can support other agent types as well. A goal-based or utility-based agent, for example, may use feedback to improve future decisions.
Multi-agent systems
Multi-agent systems use more than one agent to complete a task. Different agents may specialize in separate parts of the workflow, hand work off to one another, or review each other's results.
For example, one agent might gather research while another checks the supporting evidence. This can make complex workflows easier to divide, but it can also introduce additional coordination problems and failure points.
These categories are not always mutually exclusive. A single system may combine goal-directed behavior, utility-based decisions, learning, and coordination with other agents. The categories are most useful as a way to understand how an agent makes decisions and how complex its behavior can become.
How do AI agents work in 5 steps
AI agents work by taking in information, deciding what to do next, acting through available tools, and evaluating the results. Modern LLM-based agents often repeat this cycle until they complete the task or reach a defined limit.
Some agents plan one step at a time, while others map out more of the workflow upfront. Memory retains relevant context throughout the process, while permissions and approval gates determine which actions the system can take.
Consider a lawyer using an agent to review approved contracts and prepare a sourced issues list. The following 5 steps show how that workflow might unfold in practice.
Step 1: Perceive the environment
The agent begins by gathering the information it needs to understand the task. For a contract review, that includes the request, approved contracts, review instructions, and relevant saved context. Its environment can include files, application state, and tool responses, so the agent must first establish which documents and review criteria apply.
Step 2: Reason and plan the next steps
Once the model has clearly identified the relevant context, it then determines what information is still missing and which permitted actions could move the task forward. It might locate key provisions, compare them with the review criteria, and collect supporting citations. The plan can change as new results come in.
Step 3: Take action and use tools
The agent then puts that plan into motion through approved search and document tools. These tools can retrieve clauses and help assemble a cited issues list, while the surrounding application controls which proposed actions are actually executed. Changing contract terms or sending documents remains outside the agent’s authority.
Step 4: Observe the results
After each action, the agent checks the results before deciding what to do next. It verifies that retrieved passages belong to the correct contract and support the proposed issue. A failed search or missing document should lead to another permitted approach or escalation rather than an invented citation.
Step 5: Repeat the process
The agent uses those results to update the task state and decide whether another step is needed. The cycle continues until the deliverable meets the criteria or reaches a time, cost, or permission limit. Any unresolved issues and consequential judgments are escalated to a lawyer. Other professional workflows follow a similar loop with different tools, permissions, deliverables, and review requirements.
What are some common AI agent use cases?
AI agents can support professional workflows that involve multiple steps, tools, and information sources. The most useful applications tend to give the agent a defined task while keeping important decisions and high-impact actions under human control. For example:
- Law: An agent can search approved contracts and review playbooks, identify relevant clauses, and prepare an issues list with document citations. A lawyer reviews the analysis and decides what to address, while negotiation and contract changes remain outside the agent’s authority.
- Finance: Approved figures and spreadsheets can support a reconciliation workflow that identifies discrepancies and links them to supporting records. An analyst checks the calculations and assumptions before accepting the work, and the agent doesn't post transactions or authorize payments.
- Medicine: Authorized records can be gathered and organized into a draft chronology for clinical review. A clinician checks the result for completeness and accuracy before using it and remains responsible for diagnosis, prescribing, and changes to a patient’s care plan.
- Software engineering: An AI coding agent can inspect a repository, propose a patch, run tests in an isolated environment, and prepare changes for review. An engineer evaluates the proposed changes before deployment, since passing tests alone doesn't authorize a release.
- Consulting: Approved research and supplied datasets can give an agent the context needed to compare findings and prepare a sourced draft presentation. A consultant reviews the analysis and resolves questions such as whether conflicting figures actually measure the same metrics.
- Education: Course materials and a supplied rubric can guide an agent as it drafts exercises and suggested feedback. An educator reviews the material before sharing it with students and is responsible for consequential grading decisions.
- Cybersecurity: During incident analysis, an agent can review authorized alerts and logs, connect related events, and draft a timeline with supporting evidence. A security analyst validates the findings and retains approval over major response actions, such as disabling accounts or isolating production systems.
Across these workflows, agents can reduce some of the manual effort involved in progressing from one step to the next. Whether that creates real value depends on the quality of the results, the amount of review required, and the cost of running the workflow.
What are the benefits of using AI agents?
AI agents can reduce some of the manual work involved in completing multistep tasks. When they perform reliably within a suitable workflow, they can help teams move work between tools, identify relevant information more quickly, and handle routine sequences that would otherwise require repeated human intervention. Automation and productivity are potentially beneficial, but they don’t guarantee outcomes.
In practice, the advantages of using AI agents can show up in a few ways:
- Fewer manual handoffs: Agents can move between approved tools, documents, and data sources within a bounded workflow instead of requiring a person to transfer information at every step.
- Faster access to relevant context: Agents can retrieve information when it's needed and carry that context into drafting or analysis. In the contract-review workflow, the agent could present a relevant clause alongside the issue it identified for easier review.
- More capacity for repetitive multistep work: Agents can handle routine sequences while professionals focus more of their time on review, judgment, and work that requires domain expertise.
Teams still need to measure whether these benefits improve the overall workflow. It's important to compare total cost, turnaround time, and reviewer effort with the existing process rather than focusing only on how quickly the agent generates an output. Faster drafting provides little value if checking and correcting the work takes longer than completing the task manually.
What are the risks and limitations of using AI agents?
AI agents can introduce risks at several points in a workflow because they rely on models, external information, tools, and surrounding software to complete tasks. An error at one step can affect later decisions or actions, making it important to evaluate both the final result and the process that produced it.
Errors can compound as an agent moves through a multistep task. For example, the system might select an outdated file, misread a requirement, or treat incomplete information as sufficient. A polished final output doesn't prove that the evidence behind it was correct or that every step was completed reliably.
External content creates additional security concerns. Prompt injection occurs when malicious instructions are hidden in sources such as documents or webpages. If an agent treats those instructions as authoritative, it may attempt an unauthorized action or expose sensitive information through connected tools or stored memory. Mercor’s enterprise architecture discussion explains how surrounding controls can separate what a model requests from what the system actually permits.
Agents can also fail during execution even when the underlying reasoning appears sound. A tool may become unavailable, repeated attempts may fail to resolve the task, or a handoff between agents may introduce inconsistent information.
Multi-agent systems can also share the same underlying weaknesses, so adding another agent doesn't necessarily provide independent judgment. Time and operating costs may also become difficult to predict when workflows require repeated tool calls or retries.
These problems can originate in the model or elsewhere in the system, including the data, integrations, and orchestration layer. Managing them requires clear ownership and controls over access, review, and escalation. The next challenge is making sure the organization itself can support these requirements as agent use expands.
What challenges do enterprises face in their adoption of AI agents?
Enterprise AI agent adoption can be limited by fragmented data, unclear ownership, review demands, and ongoing operating requirements. Technical capability alone isn't enough; organizations also need the systems and people to support agent workflows over time.
The main challenges tend to fall into 3 areas:
- Data and integration: Relevant information may be spread across incompatible systems with unclear ownership or permissions. Teams need to determine which sources are authoritative, how agents can access them, and who can approve that access. A deeper look at enterprise agent architecture can help separate deployment choices from policy and execution controls.
- Review capacity and ownership: Agent workflows still require qualified people to review outputs, handle exceptions, and investigate questionable results. Without clear ownership, a pilot may simply shift the effort from completing work to checking it.
- Cost, privacy, and compliance: The business case should account for integration, review, and ongoing operating costs, not just model usage. A McKinsey August 2026 survey reports that cost remains a constraint on AI usage overall, not just agent-specific adoption. Privacy and compliance requirements can also limit which data and workflows are appropriate for agent use.
Starting with one bounded workflow and assigning clear ownership gives teams a more manageable way to address these challenges. From there, they can establish the controls needed to test, supervise, and monitor agent performance over time.
How to use AI agents effectively: 5 best practices
Effective AI agent use starts with defining a clear task and setting boundaries for how the system can complete it. Teams also need reliable data, realistic testing, appropriate human oversight, and ongoing monitoring.
1. Define clear business goals
Specify the deliverable, the outcome it should improve, and the limits of the agent’s role. For example, specifying “prepare a sourced contract issues list for lawyer review” is testable, while “make legal work more efficient” isn't. Use the current process as a baseline and identify what the agent mustn't do.
2. Use reliable, authorized data
Give the agent current and relevant data sources with clear ownership. Preserve document versions and citations so reviewers can trace conclusions. Teams should also define what happens when evidence is missing or contradictory rather than allowing confident wording to fill in the gaps.
3. Test realistic workflows and set clear criteria
Test the agent under conditions it's likely to face, including routine tasks, ambiguous instructions, and tool failures. Define acceptable performance and disqualifying behavior before deployment. NVIDIA’s evaluation guidance treats both tool behavior and the steps taken as evidence alongside the final output.
4. Keep humans in the loop
Testing should identify where professional judgment or approval is required. Assign qualified reviewers, enforce permissions outside the model, and require approval before consequential actions. Users should also be able to pause execution, reject changes, or recover from partial completion when needed.
5. Monitor results and record agent actions
After deployment, track accessed sources, tool calls, outputs, approvals, failures, and costs under appropriate data controls. Recurring failures can become regression tests that help teams catch similar problems after updates. Mercor’s guide to building agent evaluation systems explains how that feedback process can support ongoing evaluation.
These practices help teams put practical boundaries around agent use, but controls alone don't show whether the system performs reliably. Teams still need realistic evaluation to determine whether the complete agent can meet task requirements across repeated runs.
How to evaluate AI agents with real-world benchmarks
Evaluating an AI agent requires testing the complete system on realistic tasks rather than judging isolated model responses. Teams need to examine whether the agent completes the work correctly, stays within its permissions, and performs reliably across repeated runs. NVIDIA distinguishes agent evaluation from model testing because isolated answers can miss failures in tool use, context handling, and multistep execution.
For the contract-review workflow, a scorecard might ask:
- Does the issues list meet the lawyer’s substantive review criteria and include correct source citations?
- Did the agent stay within the approved tools and permissions?
- Does it perform reliably across repeated runs within cost and turnaround limits?
These are illustrative criteria, not reported benchmark results. Teams should test representative cases, inspect individual failures, and document the configuration used for each evaluation, including the model, prompts, tools, permissions, and test conditions.
Real-world benchmarks add another layer of evidence by testing agents on structured professional tasks under defined conditions. Mercor's APEX-Agents research paper describes our methodology for how experts created environments and tasks in investment banking, management consulting, and corporate law. Files and communications provide the working context, while explicit grading criteria define what successful performance looks like. APEX-Agents evaluates both task completion and also considers repeatability rather than treating one successful run as proof of success.
Public benchmarks can help teams compare performance under standardized conditions, but they don't replace testing for a specific organization or workflow. Teams still need their own evaluations to determine whether an agent can meet the requirements of the work they plan to assign it. Mercor examines that process in more detail in its guide on evaluating AI agents.
Choose AI agents with greater confidence
Move beyond demonstrations with evaluations built around the work your agents need to perform. Mercor combines domain expertise and realistic workflows to help teams assess agent quality and cost.
Explore enterprise evalsFrequently Asked Questions
How are AI agents different from LLMs and chatbots?+−
An LLM processes and generates language, while a chatbot is an interface people use to interact with a system. An AI agent can go further by choosing actions and using tools to work toward a goal. Some chatbots include agent capabilities with varying levels of model-directed autonomy.
How are AI agents trained?+−
The models behind AI agents can be trained with methods that include demonstrations or reinforcement learning. As Anthropic explains, building an agent may also involve adding instructions, tools, memory, and controls to an existing model. Creating an agent doesn't always require new model training, and updating its memory doesn't retrain the underlying model.
What industries are leading in agent adoption?+−
In McKinsey’s August 2026 survey, respondents in technology, media, and telecommunications were among the most likely to report scaling agents within business functions. Scaling differs from running a pilot, and adoption rates don't show that agents are ready for every workflow.
What's the difference between an AI agent and generative AI?+−
Generative AI creates content such as text, images, or code. AI agents choose and carry out actions toward a goal, and they may use generative AI as part of that process. Not every generative AI application is an agent, and some traditional agents don't use generative AI at all.
Do AI agents work autonomously?+−
AI agents can complete some steps on their own within defined tasks and permissions, but higher-impact actions may still require human approval. Mercor's guide to autonomous AI agents explains why autonomy doesn't mean unlimited authority.
What tools can AI agents use?+−
AI agents can use tools such as search engines, document and spreadsheet software, databases, application programming interfaces, code execution, and browsers. What an agent can actually access depends on its integrations and permissions.
Are AI agents safe to use in business?+−
AI agents aren't equally suitable for every business workflow. Safety depends on factors such as the data involved, system permissions, verified performance, and the level of human oversight required. Teams should use limited access, realistic testing, and approval controls when undertaking consequential actions or high-stakes work.
How do AI agents differ from traditional automation?+−
Traditional automation usually follows predefined rules and paths. Modern LLM-based agents can choose or revise steps based on changing context. Many systems combine both approaches, and predictable tasks may not require model-directed decision-making.
