Retail Copilot / Tracing / VX-Concierge
VX-Concierge

Get started with Evaluators

Start from a template

Sensitive Data

Whether outputs expose personal or account details.

Prompt Tampering

Whether inputs contain instruction-override attempts.

Harmful Language

Whether outputs contain abusive or unsafe wording.

Skew & Fairness

Whether outputs favour or disadvantage a group.

Fabrication

Whether an answer invents facts absent from the source.

Faithfulness

Whether an answer matches the reference material.

Perceived Failure

Whether the user believed the agent made a mistake.

Sentiment Lift

Whether the user ends the conversation satisfied.

Show all templates

Create from scratch

Model-as-Judge Evaluator

Write a scoring rubric from scratch and grade your traces.

Code Evaluator

Write a custom Python or TypeScript scoring function.