Prove your agent still follows its own playbook

Test whether your conversational AI agent still honors the commitments you defined for it, before the next release ships.

Your agent changed.
Did its promises?

Prompt changed

Instructions, context, or guardrails were updated.

Model changed

A new model or version now interprets those instructions.

Workflow changed

Tools, routing, or surrounding product logic evolved.

The commitments did not.

Reply rate will not catch it. A missing disclosure, a forbidden claim, or a question asked twice can all look like a successful conversation.

A score tells you how it performed. A finding tells you what broke.

DoppelGegner tests your agent against its playbook and shows the conversation evidence behind each result.

Quality dimensionContextual understanding
Score (0–1)0.62

Illustrative example. Not a stored production result.

Message indexMessage 5 · from the agent under test
Verbatim excerpt
“Before we continue, could you confirm your notice period?”
Explanation

The recipient had already answered this in message 4. That turn lost context and lowered the contextual-understanding assessment.

Detection sourceModel-judged quality dimension
SeverityCRITICAL
PatternA1 Question repetition
Message indexMessage 5 · from the agent under test
Verbatim excerpt
“Before we go further, could you confirm your current notice period?”
Explanation

The recipient answered in message 4 — “I'm on a 30-day notice.” The agent re-asked without acknowledging the answer.

Detection sourceModel-judged conversational check
SeverityHIGH
PatternA3 System error reached the recipient
Message indexMessage 3 · from the agent under test
Verbatim excerpt
“Your attachment was rejected — file must be PDF.”
Explanation

System-voiced error text was sent to the recipient instead of being handled internally.

Detection sourceDeterministic detector

The result has to show its work

Rules outrank judgment

Deterministic checks run before any model. If a rule fails, no model verdict can override it.

Evidence is required

Every finding cites its message and quotes the text. Without both, it does not ship.

We test the agent we didn't build

Your agent runs on its own platform. We reach it from outside, through a test channel, and evaluate what comes back.

Monitoring tells you what happened. We run before the send.