Shadow Evaluation Guide
A structured, no-cost process for payer operations teams to establish a behavioral baseline for incoming AI voice calls — without changing anything in your call flow.
What This Is
- No cost, and no vendor changes. Observe-only: it does not sit in the call path. Your privacy, security, and contractual obligations still apply.
- You run it in shadow mode alongside your current operations.
- Goal: produce a written behavioral baseline showing how AI voice agents currently disclose, escalate, and log calls to your administrative lines.
- This is not a pilot program. There is no certification, no SLA, and no commitment beyond sharing anonymized findings with the community if you choose.
What You Need
- Access to call logs or recordings for a sample of incoming administrative calls (eligibility, claim status, prior authorization).
- Staff time to read this guide and agree the scope internally.
- Willingness to share anonymized findings if you choose to.
Where a shadow evaluation sits
The word that does the work is observe-only. An evaluation reads call traces you already hold; nothing it produces re-enters the call. The diagram is one-way on purpose — there is no arrow back into the live path, because there is no mechanism for one.
How an evaluation unfolds
A sequence, not a schedule. The stages run in this order because each depends on the one before it; how long each takes is yours to decide, and an evaluation can stop after any of them.
Behavioral baseline
Log incoming calls. Identify which callers are AI voice agents. Record disclosure timing, escalation behavior, and audit trail presence.
Gap analysis
Evaluate identified AI calls against the five v1.3 controls (IDG-01, PDX-01, DBC-01, EIT-01, ATR-01), plus the supplemental bot‑to‑bot rule. Quantify what passes, what fails, what is ambiguous.
Written assessment
Compile findings into a short written assessment. Share the anonymized results with the community to help inform the next version of the proposal.
What an Observe-Only Run Actually Returns
Shadow evaluation is a proposition about volume: you find out how often the controls would have acted before anything is allowed to act. This is that measurement over a committed corpus — the decisions the engine returned, with nothing routed, blocked or escalated to produce them.
The Controls You Are Observing
| Control | What it requires |
|---|---|
| IDG-01 | Disclose AI identity at call start, before any data exchange. |
| PDX-01 | No PHI or credentials exchanged before identity disclosure is confirmed. |
| DBC-01 | No deceptive audio artifacts — no fake breathing, no human-name openers. |
| EIT-01 | Immediate, clean escalation path to a human on request. |
| ATR-01 | Basic audit log — enough to reconstruct when disclosure occurred. |
How to Start
Email us — or open a GitHub Discussion. Email contact@nhid-clinical.org with the subject line "Shadow Evaluation". Either way, include your role and the workflows you want to observe.
Related Resources
- Tier 0 Shadow Pilot Kit → — capture schema, measurement script, 2‑4 week plan, report template
- For Payers guide → — the payer-side framing
- Evidence Pack → — guarantees, failure trace, audit model
- Reference demonstration line → — hear the disclosure controls in a real call
Shadow mode in one sentence
Observe today's AI voice traffic against IDG-01, PDX-01, DBC-01, EIT-01, and ATR-01 — log what passes, what fails, what is ambiguous — then decide what to require in the next contract cycle.
Email us → Open a discussion → Download PDF →Open for feedback
Questions? Reach out directly.
Whether you have a process question, a concern about the controls, or you want to share what you're seeing in your own call traffic — feedback from payer operations teams is what shapes this work.
For payers
Establish a behavioral baseline for AI voice agent transparency — no vendor changes, observe-only. It runs alongside live traffic without sitting in the call path; your privacy, security, and contractual obligations still apply. Observe, collect evidence, decide what to require later.
Start here — a pilot in three steps
The Tier 0 Shadow Pilot Kit runs on your own call logs. Observe-only, no vendor changes, and it produces usable numbers without altering the production call flow.
1 Get the kit
Download the Tier 0 Shadow Pilot Kit — a minimal event schema, a measurement script, and a report template.
2 Run it on your logs
Map a sample of your call records to the minimal event schema, then run measure_pilot.py. Nothing touches production.
3 Read your numbers
You get impersonation-latency and disclosure metrics on your own traffic in a ready-to-share report. Use it to decide what to require of vendors.
What You Gain
| Metric | Today (typical) | Target with NHID-Clinical |
|---|---|---|
| Verification latency | 3–5 min or call terminated | < 5 seconds |
| Audit effort per vendor | Manual call review (hours) | ~2 minutes (test suite) |
| RFP disclosure language | Custom per vendor | One standard clause |
| Escalation response time | Untracked | ≤ 2 seconds, logged |
How a payer-side evaluation works
1 Add RFP language
Insert this clause into your next voice AI vendor RFP or BAA amendment:
"The vendor's AI agent SHALL produce NHID-Clinical v1.3 JSON trace logs for all B2B administrative calls, including disclosure timestamps and opt-out handling. The payer may run the open-source conformance test suite against vendor output at any time."
Existing contract? Send as a formal amendment request.
2 Vendor sandbox testing
Ask your vendor to run the open-source suite — results in under 5 minutes.
git clone https://github.com/NHID-Clinical/NHID-Clinical.gitpip install -r requirements.txtpython -m pytest tests/ -v- Send full terminal output (+ optional sample traces)
3 Validate logs yourself
- Place vendor JSON traces in
traces/ - Run
python -m pytest tests/ -v - Verify: disclosure before NPI/member ID; no deceptive artifacts; escalation ≤ 2s on request
4 Measure impact
- Verification latency (target < 5s)
- Escalation volume from identity uncertainty (target > 30% reduction)
5 Decide Next Steps
- Tests pass + metrics improve → consider requiring conformance in future contracts
- Tests fail or flat metrics → remediation or disqualify from future bids
Evaluation Resources
- Executive Brief → — one-page overview for leadership and procurement
- Shadow Evaluation Guide → — the full process
- Evidence Pack → — guarantees, failure trace, audit model
- Regulatory Alignment → — CMS-0057-F, MACPAC, DOJ FCA, state AI laws
- Reference demonstration line → — hear the disclosure controls in a real call
Get involved
Read the specification. Share what you think.
Whether it is right, wrong, incomplete, or misses the real problem — that feedback shapes the next version.
What transparent disclosure sounds like
Side-by-side examples of AI voice agent openings — the kind of interaction this proposal is trying to encourage, and the kind it is trying to reduce.
The core idea is simple: an AI voice agent should say it is not human before asking for or sharing any operational information. The examples below show what that looks like — and what it looks like when it does not happen.
Good examples — disclosure upfront
What we'd like to see more of
These openings work because the AI identifies itself immediately — before asking for anything. The representative knows what they're dealing with from the first sentence.
General payer operations inquiry
Automated nature disclosed immediately. No data exchanged before identification.
Dental payer — claim status
Non-human status disclosed before any data is requested.
Prior authorization follow-up
Clear, early, unambiguous. The representative can make an informed decision about how to proceed.
Eligibility verification
Discloses first, then asks permission. Gives the representative control.
Problematic examples — what this proposal aims to reduce
Patterns that create impersonation latency
These examples illustrate the problem. Each one involves data being exchanged before the representative knows they are talking to an automated system. The scenarios are based on real patterns observed in healthcare payer–provider calls.
The representative shared claim status and engaged in a full workflow exchange before knowing they were talking to automation. NPI and member ID were already exchanged. There is no record of when — or whether — this would have been disclosed without a direct challenge.
Disclosure comes after the purpose and claim details are established. A busy representative may have already started pulling up records.
Breathing sounds, hesitation pauses, and typing are added to create the impression of a human caller. Even if disclosure eventually comes, the framing is designed to mislead.
Disclosure is contingent on being asked. This is the core pattern this proposal aims to change.
A note on scope
These examples cover B2B administrative voice workflows — AI systems calling payer offices on behalf of providers or vendors. They do not apply to patient-facing calls, secure messaging, or internal systems.
The examples above are illustrative, not exhaustive. If you encounter patterns in real calls that are not covered here, we want to hear about them.
Get involved
Seen something different in the real world?
Real call patterns from payer operations staff are the most valuable input to this proposal. Share what you've seen.
The demo line
A scripted, deterministic phone call — no AI variance — run through the exact same NHID-Clinical evaluation pipeline that production traffic goes through. Only the speech is fake.
1. Call the Live Demo Line (Beacon)
Dial +1 (717) 670-6772 from any phone.
After the call ends, evaluation results appear below in real time through the same pipeline production traffic uses.
2. Or get a call directly from Beacon
Beacon — our outbound demo agent — calls your phone and the conversation is evaluated against the same five controls once the call ends. This costs us real calling minutes per request, so it's CAPTCHA- and rate-limited.
3. Watch it evaluated live
Every turn of the call — disclosure, PHI requests, escalation — is run through adapters/call_progress_adapter.py + src/nhid_policy_engine_v1.py exactly like a real vendor webhook, and shown below as it happens.
How this works
This is the same evaluation pipeline production traffic uses.
The two call scripts are the only fake part. The conformance evaluation — IDG-01, PDX-01, DBC-01, EIT-01 — is real.
Going further: Part III of the NHID-Clinical Playbook covers the full evaluation method, governance and data-handling considerations, and what it cannot establish.