How it works
AI Agent Rx™ by MindBotics puts realistic prospects in front of your AI agents, scores how they handle the conversation, and tells you exactly what to change to score higher next time.
This page walks through what AI Rx tests, how scoring works, and how the rollback + updated-prompt flow keeps the agent improving week over week.
What AI Rx tests
AI Rx tests your client's live AI agents — the same web chat or voice agent a real customer would talk to. You add the client's website (so AI Rx learns what the business actually offers), give it the agent's system prompt (for web chat) or phone number (for voice), pick the scenarios you want covered, and AI Rx takes over.
Each scenario simulates a different kind of customer: a new lead with a specific need, a price-shopper, an after-hours urgency, an objection-hesitant browser, an off-topic curve ball, a booking request. The AI tester stays in character, sends realistic messages, and only ends the conversation when it's either captured what the customer needed or hit a natural stop. Scenarios can be the built-in set or a business-specific set tailored to what this client actually does — more on that next.
Grounding tests in the business
A test is only useful if it knows what the business actually does. When you add a client, AI Rx captures their website (required at onboarding) and turns it into a structured business profile — the real services and products, brand voice, hours, pricing, and policies. You review and approve that profile, and from then on it grounds every test.
Why it matters: grounded tests stay accurate and on-model. The agent is judged against what the business really offers — so a correct "we don't do that" answer scores well instead of being penalized, and the updated prompt never invents a service the business doesn't provide.
- Modern websites just work. Static sites are read instantly; JavaScript / React / Base44 sites are rendered automatically so their real content comes through.
- Google Business Profile. One click enriches the profile with verified hours, phone, rating, and the themes customers mention most in reviews — especially handy for local service businesses.
- Business-specific scenarios. Generate test scenarios from the profile so they probe this business's real customer journeys — sizing and shipping for a product shop, booking and quotes for a trades business — instead of a generic set.
- No prompt yet? Draft one. For a brand-new agent, AI Rx can draft a complete starting prompt straight from the approved profile — grounded only in real facts — that you review and refine.
Nothing grounds a test until you approve the profile, so a bad scrape can never quietly skew your results.
How a test works
- Pick the client, the channel (web chat, voice, or both), and the scenarios to run.
- For each scenario, an AI "tester" plays the customer and talks to the client's real agent — for voice, this is a real outbound phone call placed via our Twilio number to the agent's number.
- When the conversation ends, the full transcript is graded by Claude across five scoring dimensions (relevance, accuracy, tone, lead capture, goal completion) on a 0–100 scale.
- After every scenario finishes, AI Rx synthesizes a single updated system prompt that incorporates every finding and targets the "Optimized" tier on a re-test.
- You see scores, strengths, issues, and a per-scenario recommendation — plus the proposed updated prompt with a one-click Apply button.
How scoring works
Each conversation is graded on five 0–100 dimensions, then rolled into an overall score and a performance tier.
| Dimension | What we look for |
|---|---|
| Relevance | Stayed on point and addressed exactly what the customer asked. No off-topic excursions. |
| Accuracy | No made-up facts. Honest about limits. Only states information the business has actually provided. |
| Tone | Warm, professional, on-brand. Not robotic, pushy, or aggressive. |
| Lead capture | Captured the customer's name AND contact (phone or email) within the conversation, with qualifying questions tied to the customer's goal. |
| Goal completion | Moved the conversation toward the customer's goal (booking, purchase, resolution) and ended with a specific named next step. |
Performance tiers
- Optimized 90 and above — the agent is converting and on-brand; minor polish only.
- Sub-Optimal 50–89 — working but losing leads. The updated prompt focuses on closing specific gaps to push every dimension into the 90s.
- Marginal Below 50 — the agent needs a serious rewrite; treat the report as the checklist.
What's in the report
- Overall score and tier for the whole run.
- Per-scenario cards with the five dimension scores, strengths, specific issues, and a concrete recommendation.
- Transcripts of every conversation (web) or the call summary plus recording link (voice).
- An updated system prompt synthesized from all the findings, designed to score 90+ on a re-test, with a one-click Apply that auto-versions the current prompt.
- Downloadable PDF branded for sharing with your client, plus a permanent public share link (scores + findings only — no transcripts).
The updated prompt
After every run, AI Rx gives Claude the agent's current system prompt plus every scenario's per-dimension scores and findings, and asks for a single complete rewrite that would score 90+ on every dimension on a re-test.
The rewrite preserves your client's brand voice, hard rules, named contacts, prices, and any business facts the current prompt contains. It rewrites aggressively wherever a scoring dimension fell below 90 — and leaves the rest alone.
It also reconciles the prompt against the approved business profile — folding in real services, pricing, and hours the prompt was missing, and correcting anything out of date — while never inventing something the business doesn't offer.
Two buttons on the "Updated prompt" card:
- Copy — copies the full prompt to your clipboard for review or pasting elsewhere.
- Apply to this client — confirms, then swaps the live prompt and snapshots the prior one as a rollback-able version (see below). After Apply, a "Re-run these scenarios to verify ≥90" link surfaces so you can confirm the rewrite actually hit the target.
Versions and rollback
Every save of a client's web-agent prompt — whether you edit it manually or apply an updated prompt from a test report — auto-snapshots the prior value as a numbered version (v1, v2, v3…). On the client detail page, the "Prompt versions" panel lists every prior version with timestamps and a one-click Rollback button.
Rollback is fully reversible: it auto-snapshots the current prompt as a new version before swapping, so any rollback can itself be rolled forward. You can experiment with an updated prompt safely — if it underperforms, one click restores the previous one.
Voice testing
Voice tests place real outbound phone calls from a Twilio number to your client's voice agent. Each call lasts up to three minutes; AI Rx shows a confirmation before placing any calls so you don't fire 25 of them by accident. Voice scenarios are tested one at a time to avoid hitting concurrent-call limits.
Voice testing is included on Pro and Agency plans. Voice findings appear in the per-scenario report; the synthesized updated prompt focuses on the web chat (Vapi voice configuration lives in a different surface).
Plans & limits
Each plan includes a set number of test runs per billing cycle — one run is one scenario tested against one agent. Voice runs place real phone calls, so they have a separate, lower limit and also count toward the total.
| Plan | Price / mo | Clients | Test runs / mo | Voice runs / mo |
|---|---|---|---|---|
| Starter | $129 | 5 | 30 | — |
| Pro | $279 | 25 | 150 | 50 |
| Agency | $499 | 75 | 500 | 200 |
| Enterprise | Custom | 75+ | Custom | Custom |
Need more than 75 clients, higher run volume, SSO, an SLA, or invoicing? That's our Enterprise plan — contact sales for custom pricing.
Limits reset at the start of each billing cycle. By default, when you reach a limit new runs pause until your next cycle or you upgrade — no surprise charges. The app shows a clear message the moment you hit a cap, so you're never charged for something you didn't choose. If you want runs to keep going past a quota, an admin can turn on overage billing in Billing: extra runs are billed at $0.50 per web run and $2.00 per voice run, capped at 2× your quota.
FAQ
- How does AI Rx know what my client actually offers?
- From the client's website (and, optionally, their Google Business Profile). AI Rx builds a reviewed business profile of their real services, hours, pricing, and policies, and grounds every test in it — so the agent is judged against what the business really does, not a generic template.
- Do I have to provide the client's website?
- Yes — it's required when you add a client, because it powers the business profile that keeps tests accurate. You review and approve the profile before it's used, and you can edit anything that didn't come through right.
- What if my client's site is a React / JavaScript app?
- It's handled automatically. AI Rx reads static sites instantly and renders JavaScript-heavy sites (React, Base44, and similar) so their real content comes through. If a page looks under-captured, a one-click full render pulls the rest.
- Can AI Rx write a prompt for an agent that doesn't have one yet?
- Yes. For a brand-new agent, AI Rx can draft a complete starting prompt straight from the approved business profile — grounded only in real facts, with lead-capture and clear next steps built in — that you review, edit, and apply.
- Does AI Rx store my clients' agent prompts?
- Yes — the current prompt and every prior version are stored against the client record in your tenant, and used only by you and your team. They are not shared with other tenants. Each request is scoped by organization ID end to end.
- What happens to my data if I cancel?
- Your data stays in place during the trial and after cancellation; you can resume by reactivating your subscription. Reach out if you need an export or deletion.
- What if the test agent says something off-brand?
- The tester plays the customer, not the agent — so it asks questions and pushes back, but it never speaks as your client. If the report flags a tone issue, the synthesized updated prompt will tighten the agent's voice rules.
- Can I share a report with a client without exposing the transcript?
- Yes. The public share link and the PDF (when set to share mode) show scores + findings only. Internal copies show the full transcript and the updated prompt.
- Does AI Rx rewrite my prompt on its own?
- Only when you click Apply. The synthesized updated prompt is generated automatically with every run, but it doesn't touch your live agent until you confirm — and even then, the prior prompt is auto-saved as a rollback-able version.
