Viktar Patotski Viktar Patotski · · AI  · 12 min read

Build an AI Agent with Spring AI 2.0: A Working Example with Guardrails

A support-triage AI agent in Java with Spring AI 2.0.1: one structured workflow step, a bounded tool loop, tenant data the model never sees, and refunds that wait for a human. Plus the guardrails Spring AI now ships and the ones you still write yourself.

A support-triage AI agent in Java with Spring AI 2.0.1: one structured workflow step, a bounded tool loop, tenant data the model never sees, and refunds that wait for a human. Plus the guardrails Spring AI now ships and the ones you still write yourself.

TL;DR: Spring AI 2.0.1 has no Agent class, and you do not need one. An AI agent in Spring is a ChatClient call with tools attached: the framework runs the loop where the model calls a tool, reads the result and decides what to do next. Since 2.0.1 that loop has a built-in tool-call cap (40 per tool and 150 per turn by default, far too high for a support bot) and only the tools you attach to a request can run. What you still write yourself: a token budget, approval for anything that writes, tenant isolation and evals. The worked example below does it in three small classes.

What founders mean by “AI agent”

“Agent” is the loudest word in AI right now and the one a vertical SaaS needs least often. Anthropic’s Building Effective Agents draws the line that matters. Workflows are “systems where LLMs and tools are orchestrated through predefined code paths.” Agents are “systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.”

In a workflow, your code owns the path. In an agent, the LLM (large language model) owns it at runtime. That buys flexibility and costs you predictability, testability and money. Anthropic’s own advice is to start at the bottom: “we recommend finding the simplest solution possible, and only increasing complexity when needed,” because agentic systems “trade latency and cost for better task performance.”

Do you need an agent at all?

Pick the lowest row that solves your problem.

PatternWho owns the pathModel calls per taskUse it for
Single callYour code1Classify, extract, summarize
Prompt chainYour codeFixed NIntake, normalize, then route
RoutingYour code2Support triage to the right prompt or model
Parallel checksYour codeN at onceSeveral independent checks on one document
Evaluator loopYour code2 per roundDraft, critique, revise to a bar
Orchestrator and workersThe modelDynamicWork you cannot split up in advance
Tool-using agentThe modelUp to your cap”Handle this request end to end”

In five of the seven rows your code owns the path. They are cheaper, deterministic and unit-testable. Spring keeps runnable versions of the workflow patterns in spring-ai-examples. Reach for the last row only when the steps depend on what the model finds along the way. Customer support is that kind of task, which is why the example below uses one workflow step and one bounded agent step, not a free-running agent.

What Spring AI 2.0.1 gives you, and what it does not

Spring AI 2.0.1 is the current release (21 August 2026; 2.0.0 went GA on 12 June). It needs Spring Boot 4 and Java 17 or later. For an agent, the parts that matter:

  • The tool loop. Attach @Tool methods to a ChatClient request and the built-in ToolCallingAdvisor runs call, tool, result, call until the model stops asking for tools. This is the agent.
  • Tool-call limits, new in 2.0.1. The loop now stops after 40 calls to any one tool or 150 tool calls in one turn. Before 2.0.1, in the words of the upgrade notes, “there was no limit at all.” Set your own limits with spring.ai.tools.limits.*.
  • Request-scoped tools, also new in 2.0.1. The tools attached to a request “are the only tools that can be executed for that request.” In 2.0.0 a model could call any tool callback bean in the application context by name. Leave the fallback switched off.
  • ToolContext. Per-request data, such as a tenant ID, reaches your tools without ever going to the model.
  • Structured output. .entity(Triage.class) turns a reply into a Java record.
  • Memory, RAG (retrieval-augmented generation) and MCP (Model Context Protocol) as advisors and starters, and a token-usage metric (gen_ai.client.token.usage) through Micrometer.

What it does not give you: an Agent type, an approval step for tools that write, a token or dollar budget, and an eval harness (it has evaluator building blocks such as RelevancyEvaluator, not a test suite). The tool-call cap is the one most tutorials miss, and its own design notes are clear about what it is not: “Not a token or monetary cost budget; that’s a separate concern.” It is “a coarse safety net.”

AI agent example: a support-triage agent in Spring AI

The task: a customer writes to support in a field-service scheduling product. The assistant answers from their tickets and the help center, and it may suggest a refund, but a person approves every refund. Each customer belongs to a tenant, and the model must never read another tenant’s data.

Setup: Spring Boot 4.1.1, the spring-ai-bom at 2.0.1, and one model starter. This example uses spring-ai-starter-model-anthropic; switching provider is a starter and a property change, which is worth keeping that way (why it matters).

The limits, in configuration

spring.ai.tools.limits.max-total-tool-calls=6
spring.ai.tools.limits.max-calls-per-tool-default=3
spring.ai.tools.limits.on-limit-exceeded=THROW

Six tool calls is enough to look up tickets, search the help center twice and propose a refund. A support turn that needs more is a turn a person should read. Keep THROW, the default: the other mode, RETURN_ERROR_RESPONSE, tells the model about the limit and lets the conversation continue, so it does not bound how many times you pay for a model call.

The tools

@Component
public class SupportTools {

    private final TicketRepository tickets;
    private final HelpCenter helpCenter;

    public SupportTools(TicketRepository tickets, HelpCenter helpCenter) {
        this.tickets = tickets;
        this.helpCenter = helpCenter;
    }

    @Tool(description = "List the five most recent support tickets of the customer in this conversation")
    public List<TicketSummary> recentTickets(ToolContext ctx) {
        return tickets.recent(tenantId(ctx), customerId(ctx), 5);
    }

    @Tool(description = "Search the help center for articles that answer the question")
    public List<Article> searchHelpCenter(@ToolParam(description = "Search query") String query, ToolContext ctx) {
        return helpCenter.search(tenantId(ctx), query, 3);
    }

    @Tool(description = "Propose a refund for a human to review. Does NOT issue a refund.")
    public String proposeRefund(@ToolParam(description = "Amount in USD") BigDecimal amountUsd,
                                @ToolParam(description = "Why the refund is justified") String reason,
                                ToolContext ctx) {
        RefundProposals proposals = (RefundProposals) ctx.getContext().get("proposals");
        proposals.add(new RefundProposals.Proposal(customerId(ctx), amountUsd, reason));
        return "Proposal recorded. A person on the team will review it. Do not tell the customer it is approved.";
    }

    private static String tenantId(ToolContext ctx) {
        return (String) ctx.getContext().get("tenantId");
    }

    private static String customerId(ToolContext ctx) {
        return (String) ctx.getContext().get("customerId");
    }
}

Three decisions in this class do most of the security work:

  • The model never chooses whose data it reads. recentTickets takes no customer ID argument. Tenant and customer come from ToolContext, which your code fills and the model never sees. Give a tool a customerId parameter and a prompt injection can ask for someone else’s tickets.
  • The write tool does not write. proposeRefund records a proposal in a per-request object and returns. No money moves inside the loop.
  • Reads stay narrow. Five tickets, three articles, scoped by tenant in the repository query, the same way your REST API already scopes them.

RefundProposals is a small holder class: a thread-safe list of Proposal(customerId, amountUsd, reason) records with add, all and isEmpty.

The agent

@Service
public class SupportAgent {

    public sealed interface Reply {
        record Answer(String text) implements Reply {}
        record NeedsApproval(String draft, RefundProposals proposals) implements Reply {}
        record HandOff(String reason) implements Reply {}
    }

    private final ChatClient chat;
    private final SupportTools tools;

    public SupportAgent(ChatClient.Builder builder, SupportTools tools) {
        this.chat = builder
                .defaultSystem("""
                        You are the support assistant for a field-service scheduling product.
                        Answer from the customer's tickets and the help center only.
                        Never promise a refund; propose one with the tool if it is justified.
                        """)
                .build();
        this.tools = tools;
    }

    public Reply handle(String tenantId, String customerId, String message) {
        // 1. Workflow step: one structured call, your code decides what happens next.
        Triage triage = chat.prompt()
                .user("Classify this support message: " + message)
                .call()
                .entity(Triage.class);
        if (triage == null || triage.category() == Triage.Category.OTHER) {
            return new Reply.HandOff("Not something the assistant should handle");
        }

        // 2. Bounded agent step: the model picks tools, inside limits you set.
        RefundProposals proposals = new RefundProposals();
        ChatResponse response = chat.prompt()
                .user(message)
                .tools(tools)
                .toolContext(Map.of(
                        "tenantId", tenantId,
                        "customerId", customerId,
                        "proposals", proposals))
                .call()
                .chatResponse();

        // 3. Guardrails checked in Java, not in the prompt.
        if (response == null) {
            return new Reply.HandOff("No response from the model");
        }
        String finishReason = response.getResult().getMetadata().getFinishReason();
        if (ToolCallLimitExceededException.FINISH_REASON.equals(finishReason)) {
            return new Reply.HandOff("Tool-call cap reached");
        }
        String text = response.getResult().getOutput().getText();
        if (!proposals.isEmpty()) {
            return new Reply.NeedsApproval(text, proposals);
        }
        return new Reply.Answer(text);
    }
}

Triage is a record with a Category enum (BILLING, HOW_TO, BUG, OTHER) and a one-line summary.

Read handle top to bottom and the split between workflow and agent is visible. Step 1 is a single call whose result your code branches on: anything the classifier cannot place goes to a person before the agent ever runs. Step 2 is the agent: the model decides which tools to call and in what order, inside the six-call cap and with only these three tools available. Step 3 is plain Java.

Two checks in step 3 are not optional:

  • The cap check. When the cap trips, Spring AI ends the loop and returns a response whose finish reason is toolCallLimitExceeded and whose text is the limit error. Skip the check and that error text goes to your customer.
  • The proposal check. It reads the holder, not the reply text. Spring AI has a returnDirect flag that ends the loop when a tool runs, but it only takes effect when every tool called in that step has the flag. If the model proposes a refund and searches the help center in the same step, the loop continues, and the reply may say whatever the model decided. The holder records the proposal either way, and a person sees it before anything happens.

I ran this code on Spring AI 2.0.1 and Spring Boot 4.1.1 against Anthropic’s API, with stub repositories, on 5 October 2026:

TestTool callsResult
”How do I export my invoices to CSV?“1 help-center searchAnswer with the export steps
Same agent, cap set to 1, a multi-step request1 ticket lookup, then the capHandOff("Tool-call cap reached"), a person takes over
Duplicate-charge refund plus a how-to question1 search and 1 refund proposalNeedsApproval, proposal recorded for a person to review

Run the same three against your own model before you ship: the hand-off and approval paths are the ones a demo never exercises.

The guardrails that matter, in order

1. A tool-call cap you chose. Built in since 2.0.1, set in configuration above. With the cap at six and THROW, the agent step makes at most seven model calls, eight with the triage call. That is the number to put into your cost model.

2. A token budget. Not built in. The tool-call cap counts calls, not tokens, and a single call with a long ticket history can cost more than three short ones. Track gen_ai.client.token.usage per tenant from the first day and alert on it; if you need a hard stop, the Spring AI docs list budget enforcement as a use for a custom ToolCallingAdvisor. How the bill grows with calls per task, from a single call to an agent loop, is in what an LLM feature really costs.

3. Tenant context outside the prompt. ToolContext for every identifier that decides whose data a tool touches. Never let the model pass it as an argument.

4. Propose, do not execute. Any tool that writes, sends or charges records an intent. Your code, or a person, carries it out.

5. Timeouts. Set a timeout on the model client’s HTTP calls, so one slow provider response cannot hold a request thread for minutes.

6. Evals before you change the prompt. A unit test that passes once proves little when the output varies between runs. Keep a set of real support messages with expected categories and expected tool choices, and run it in CI before any prompt or model change.

The security risk is the tools

Giving an agent tools gives prompt injection a path to real actions. OWASP (the Open Worldwide Application Security Project) published a Top 10 for Agentic Applications in December 2025, and three of its risks map straight onto the code above: hijacking the agent’s goal through injected content, misusing a legitimate tool, and abusing the identity or privileges the agent runs with. Request-scoped tools, tenant data in ToolContext and write tools that only propose are the direct answers to those three. If you expose your product to agents from the other side, through MCP, the same rules apply; see building an MCP server in Java with Spring AI.

Other ways to build an AI agent in Java

  • Spring AI 2.0.1: primitives plus a bounded loop; the right default on a Spring Boot stack. Choosing between it and LangChain4j is covered in Spring AI vs LangChain4j.
  • LangChain4j 1.21.0: stable core; its agent module (langchain4j-agentic) is still a beta release.
  • Embabel 1.5.2: from Rod Johnson, the creator of Spring; Kotlin-first, Spring-integrated, with a goal-oriented planner. Young, worth watching if you need planning beyond a tool loop.
  • Google ADK for Java 1.11.0: has a first-class LlmAgent class.
  • spring-ai-agent-utils 0.13.0: a community toolkit inspired by coding agents (skills, sub-agents, to-do tools). Not part of Spring AI, and pre-1.0.

If your agent answers from your own documents, it needs retrieval underneath; that part is explained in RAG explained.

Sources


Your AI proof of concept works in the demo and nobody trusts it in production? The gap is the guardrails above, not the model. I help SaaS teams ship AI features with the cost, approval and tenant isolation in place: AI Enablement consulting, or book a free 30-minute call.

Related Posts

View All Posts »
AI Viktar Patotski Viktar Patotski · 13 min read

RAG Explained: What It Is and When You Actually Need It

RAG means handing your AI the right page from your own documents before it answers. It is powerful when your knowledge base is large, changing, or per-customer. It is also something a lot of teams build too early. Here is the honest decision.

Back to Blog