How is Vera Joy doing?
A quick health check of your AI pharmacy agent. Run tests, review the results, and track progress toward the launch goal.
Result mix
Recent activity
| ID | Scenario | Category | Severity | v3 Status | Patterns | Update | Review |
|---|
| ID | Pattern | v3 Section | Status | Description | Fix |
|---|
Two views are available. Original shows the baked-in recommendations from the initial review. AI-refreshed regenerates the list using Claude against the current saved prompt and live Pass/Fail state — click Refresh AI recommendations after marking new Pass/Fails or updating the prompt.
| ID | Priority | Title | Section | What to Change | Fixes | Status | Apply |
|---|
A scenario is auto-marked Pass only if its Issue Identified or Additional Notes contains the phrase "No issues" or "No major issues". Everything else defaults to To Confirm. Use the dropdowns in the Coverage Matrix tab to manually set Pass / Fail / To Confirm. Roll-ups: a sub-category passes only if every scenario passes; a category passes only if every sub-category passes; any Fail propagates upward.
By Category
Top-level categories rolled up from sub-categories
By Sub-Category
Sub-categories rolled up from individual scenarios
By Sub-Sub-Category
Individual scenarios (the leaf level)
Prompt update results
Outcomes of the autonomous Prompt update sessions — how many failing use cases the loop drove to a pass, how many still need review, and how many couldn’t be tested.
Outcome share
Pass / Fail by Category
Per-category counts of individual scenarios (not roll-ups). Sorted by Fail count descending.
Call timing by use case (Human: Pass calls only)
Average call duration across only the Human: Pass calls for each use case (AI-only and unconfirmed runs are excluded — a use case appears here once it has at least one human-confirmed Pass). Click any column header to sort — including Category and Sub-category to group the timing by area.
Every scenario's Issue Identified field, checked against the saved Vera Joy prompt by AI. If the prompt already addresses an issue, it's marked Resolved with a green check and a quote from the prompt as evidence; otherwise it stays Open. Results are cached per-browser — re-run the check after updating the prompt.
Open Issues
A shared list of issues anyone on the team can raise. Attach one or more call transcripts to an issue (each with its own note), flag the ones that need attention, and leave comments. Everything here is shared with the whole team.
Patient calls to Vera Joy
Real calls made to the production Vera Joy phone number, pulled from Vapi — every patient engagement, alongside the platform tests. Each call is auto-mapped to its closest use case and judged. New calls are captured automatically once the PROD assistant's webhook is connected; use Import recent calls to pull history on demand.
Call outcomes
Every call — test and production — classified into an outcome: Successful completion (Vera handled it), Success w/ Esc (correctly transferred), Failure w/ Esc (tried, failed, then escalated), Direct Esc (required direct escalation), or Other. Broken down by use case. Timing shows the Vera-Joy leg today; the full escalation leg (until the human hangs up in AWS) fills in once AWS is connected.
Call Analysis
Analyze the calls in a date range: volume and duration, which use cases were covered, and a full transfer breakdown — how many calls were handed to a human, which department they went to, and which transfers should not have happened: an agent malfunction (e.g. a failed profile lookup ending in “I can transfer you to customer service”) or a request the agent should have completed itself. A caller asking for a human is tracked separately and counts as acceptable.
Fail transcripts — biggest areas needing improvement
Transcripts of past failed calls, grouped into the priority improvement areas so you can read exactly where Vera Joy is going wrong. A call appears under every area its use case relates to.
Rulebook — use cases by focus area
Every use case grouped into the 10 focus areas. Each row states what success means and what failure means — compare against the prompt, then mark it Pass, Fail, or Ready. Marks are shared with the whole team.
Drift Monitor
String-detectable scan of Vera Joy's spoken turns for safety-critical and banned-phrase drift (MLOps runbook §3). Zero LLM cost — pure regex. Runs nightly on a cron; run it on demand below. CRITICAL hits are an incident; WARN hits go to weekly review.
Data-Contract Tests
Verifies the guard-proxy tool endpoints return the payload shapes the prompt assumes (MLOps runbook §4). Each failure is a bug Vera would otherwise surface to a live patient. Runs weekly on a cron and on every backend deploy; run it on demand below.
MLOps Runbook
The operating procedure for shipping and monitoring Vera Joy — release units, the per-release checklist, drift monitors, data contracts, the incident→regression loop, and standing compliance items.
Evolution — growth in complexity
How the platform and the Vera Joy prompt have grown — use cases mapped, focus areas, prompt sections and words, engineering activity, and test volume. Use this to validate the value of the work.
Loading…
All comments added in the Coverage Matrix, grouped by Category → Sub-category → Scenario. Comments are stored in this browser's localStorage. Use Export comments to share with the team.
Commentary
Every comment left during manual / human checks, grouped by use case. Pulled from the verdict notes saved with each Human, phone, and Agent-to-Agent confirmation.
This is the Prompt (DEV) — the working script Vera Joy is tested against. The AI updates it automatically when fixing a failing use case (via Apply → in the Improve tab), and you can edit it directly here. Admins can push it straight to the live Vera Joy assistant in Vapi with Push to Vapi (Vera Joy). The server stores each version with author + timestamp.
Every tested use case mapped to the prompt section that governs it (from each scenario's v3_reference). Use this to see which part of the prompt addresses which use case, and where coverage is weak. Sections are ordered as they appear in the prompt; a use case that cites multiple sections appears under each.
The promotion engine moves validated prompt language from DEV to PROD, one prompt section at a time. A use case is ready after 3 consecutive confirmed passes (a Fail resets the streak). A prompt section can be Upgraded to PROD once every use case mapped to it is ready. Updating the DEV prompt resets the streaks for the use cases on any section whose text changed.
The PROD prompt is what powers the production Vera Joy agent. It starts as a snapshot of the DEV prompt and is updated section-by-section as use cases are promoted from the Promotion tab. The diff below shows where PROD currently differs from DEV.
The Greenlit Prompt contains only the use cases that have earned 3 Human : Pass scores (the Ready column in the Kanban) — plus all the shared machinery those scenarios need to work (identity verification, MCP / tool calls, pronunciation, routing, rules, guardrails, persona). It’s a known-good, fully-validated prompt. As new use cases reach 3 Human : Pass, click Update to fold them in.
AI-driven prompt optimizer. Runs Claude over every scenario currently marked Fail in the Coverage Matrix, compares them against the saved Vera Joy prompt, and produces prioritized P0/P1/P2 recommendations. Results are cached in this browser — click Run AI Analysis to refresh.
Success rate by day — all use cases · Phase 1 only · ▲ new prompt version
My profile
Every use case flagged Confirm with DCA from the Auto-Test page — runs whose result was uncertain and parked for DCA to confirm. Resolve each as Pass or Fail (or leave it for DCA). Anyone signed in can resolve these.
Use cases promoted to Ready (set the status dropdown on any use case in Tests → Use Cases or the Use Case Library). Click a use case to see its history and transcripts.
Vera Joy — Ready Use Cases
Status of use cases cleared for production, for the DCA board.
| Set | Title | # Scenarios | Summary |
|---|
Summary Updates
A digest of activity on verajoy.teleperson.com — tests run, AI & human verdicts, prompt updates, and failures awaiting review — emailed at 10:00 PM Eastern. Choose the cadence (nightly on the days you pick, or weekly on one day) and who receives it.
Appearance
Choose how the console looks. System follows your operating system and updates live.