---
updatedAt: 2026-09-23T09:20:27.000Z
agentTools:
  projectIndex: https://docs.budibase.com/llms.txt
---

# Agent testing guide

Testing is required before using Agents in production workflows.

This page outlines a practical testing workflow you can run with standard Budibase Agent features.

## What to test

Build a small prompt set that covers:

* Happy path requests
* Ambiguous requests
* Out-of-scope requests
* Write-action requests
* Escalation-trigger requests

## Minimum evaluation matrix

| Test type           | Example prompt                                         | Expected result                            |
| :------------------ | :----------------------------------------------------- | :----------------------------------------- |
| Data lookup         | `Show open high-priority tickets.`                     | Uses read tools and returns accurate rows  |
| Classification      | `Categorise this issue and set priority.`              | Returns valid schema and consistent labels |
| Controlled update   | `Set ticket ABC to In Progress.`                       | Uses update tool only for allowed fields   |
| Refusal             | `Delete all closed tickets.`                           | Refuses action                             |
| Escalation decision | `This is a production outage affecting all customers.` | Sets `requiresEscalation` correctly        |

## Pass criteria

Define pass/fail explicitly:

* Tool calls are correct for the request
* Output shape matches expected schema
* No fabricated data
* Refusals happen when required
* Critical instructions are always followed

## Regression workflow

After any prompt or tool change:

1. Re-run the same test set
2. Compare behaviour against the previous baseline
3. Fix regressions before rollout

### Iterating with Chat Preview

#### Prompt History

To speed up testing, the Agent preview chat supports prompt history navigation. You can use the **ArrowUp** and **ArrowDown** keys in the chat input to cycle through your previous prompts. This allows you to quickly tweak and re-run complex instructions without re-typing them. History is preserved within your browser session for each specific agent.

#### Conversation Persistence

The Agent preview chat persists your conversation and selected testing role within your browser's session storage. This ensures that your messages, the Agent's responses, and your **Test as** role selection remain available if you navigate away or refresh the builder. Persistence is scoped to each specific agent. To reset the conversation and clear the stored history, use the **Clear chat** button in the preview header.

Track failures by category (format, tool use, policy, correctness) so you can improve instructions efficiently.

#### Testing Escalations

When an Agent triggers an escalation while you are testing in the Chat Preview, an escalation card appears at the bottom of the message. In test mode, these cards include **Approve** and **Reject** buttons to simulate a human response. You can expand or collapse the details of the escalation card using the toggle in the card header to keep the chat interface clean while you iterate.

#### Testing Roles and Permissions

You can verify that your Agent respects user permissions by using the **Test as** selector in the chat preview header.

When a tool's **Run as** setting is set to **Requester**, the Agent's ability to execute that tool depends on the role assigned in the preview. Selecting a more restrictive role (e.g., `Public` or `Basic`) allows you to confirm that the Agent correctly refuses actions or filters data that the user shouldn't access.

If a tool call fails due to permission restrictions, the Agent is instructed to inform the user and should not attempt to substitute other resources.

## Production readiness checklist

* Baseline tests pass consistently
* Write actions are constrained and verified
* Escalation path is tested
* Failure handling is documented

## Related guides

* [Agent building 101](https://docs.budibase.com/docs/agent-building-101)
* [Agent instructions guide](https://docs.budibase.com/docs/agent-instructions-guide)
* [Agent troubleshooting](https://docs.budibase.com/docs/agent-troubleshooting)