Introducing LLM Evaluation on your tests
A new layer of testing for evaluating meaning, not just exact responses.
Not every test needs an exact text match.
When testing conversational experiences, sometimes you need to verify that the system returns a specific response. In other cases, what matters is whether the response meets a particular intent or criterionβeven if the wording is different.
That's why we've introduced LLM Evaluation in Bespoken.
Exact matching when exact wording matters. LLM Evaluation when meaning matters.
π― From exact responses to expected outcomes
Different tests require different ways of defining correctness.
If a test requires a specific response, you can continue using an Expected Prompt to define exactly what the system should say.
For example, in the test below, even a small change in wording can cause the test to fail when using an exact match.
But sometimes, the exact wording isn't what matters. You may want to know whether the system greeted the user and asked what they were calling about, regardless of how it phrased the response.
That's where LLM Evaluation comes in.
Instead of defining the exact response, you can describe the expected intent and use an assertion to evaluate whether the actual response meets that criterion.
This gives QA teams more flexibility to define what "correct" means for each test:
Exact matching when exact wording matters. LLM Evaluation when meaning matters.
How LLM Evaluation works
LLM Evaluation can be added directly to your test suite.
For each step, you can choose whether to evaluate the response using:
- Prompt β when you need to validate a specific response.
- LLM Evaluation β when you care about whether the response fulfills a particular intent or criterion.
- Both β when you want to validate both the expected response and the underlying intent.
This gives QA teams more flexibility when defining what "correct" means for their conversational AI.
And you still get the full response transcription in the test results, so you can see exactly what the system said alongside the evaluation.
Evaluate conversations based on intent
The biggest advantage of LLM Evaluation is that you don't have to predict every way an AI might correctly respond.
Instead of writing: "Your reservation has been cancelled."
You can define the expected behavior as: "The system confirmed that the user's reservation was cancelled."
The system can use different wording and still passβas long as it satisfies the criterion.
This makes automated testing much better suited to the inherently variable nature of conversational AI.
Evaluate the entire call
LLM Evaluation isn't limited to individual responses.
You can also add LLM Call Evaluation as a post-processing step to evaluate the conversation as a whole.
For example:
Bespoken evaluates each criterion against the conversation and reports whether the requirements were met.
This opens up new possibilities for testing things that are difficult to capture with traditional step-by-step assertions.
You can evaluate whether:
- The conversation followed the expected intent
- Required information was collected
- Specific conditions were met throughout the call
- The conversation stayed within a defined language or behavior
From testing to monitoring
LLM Evaluation also works with Monitoring.
The same criteria you use in your automated tests can help evaluate conversations running in production, giving teams another way to understand whether their conversational AI is behaving as expected over time.
Evaluation results are available in History, so you can review how individual calls performed against your defined criteria.
More flexible testing for AI systems
Conversational experiences don't always need to be tested against a single exact response.
Sometimes the exact wording is the requirement. Other times, the intent or outcome is what matters. LLM Evaluation gives QA teams a way to test for both.
Exact matching when exact wording matters. LLM Evaluation when meaning matters.
π€ TL;DR
- Use exact matching when specific wording matters.
- Use LLM Evaluation when the intent or meaning matters.
- Evaluate individual test steps or entire conversations with LLM Call Evaluation.
- Combine both approaches when you need to validate exact responses and broader outcomes.
- Use the same approach for monitoring and review results in History.
Getting Started
LLM Evaluation is available directly in your Bespoken test suites.
Start by adding an evaluation criterion describing what you expect the system to accomplish, then let Bespoken evaluate whether the response or conversation meets it.
Ready to test AI based on what it meansβnot just what it says?
Book a demo to see how Bespoken helps teams build more flexible, reliable conversational AI tests.