Sensitive Data
Whether outputs expose personal or account details.
Whether outputs expose personal or account details.
Whether inputs contain instruction-override attempts.
Whether outputs contain abusive or unsafe wording.
Whether outputs favour or disadvantage a group.
Whether an answer invents facts absent from the source.
Whether an answer matches the reference material.
Whether the user believed the agent made a mistake.
Whether the user ends the conversation satisfied.
Write a scoring rubric from scratch and grade your traces.
Write a custom Python or TypeScript scoring function.