GPT-5 for customer support: what to test
A practical evaluation plan for using GPT-5 in customer support, covering grounded answers, tool calls, multi-turn cases, cost, latency, rollout, and rollback.
- Updated
- Reading time
- 7 min read
10 sections
- Do not start with a headline benchmark
- Build a support-specific test set
- Test tool calls as transactions
GPT-5 introduced configurable reasoning and new tool-use options to OpenAI's model family. Those capabilities are relevant to support work, but a general benchmark cannot tell you whether a model will apply your refund policy correctly or call your billing tool with the right account.
OpenAI now describes the original GPT-5 as a previous model and points developers to newer models in its current model documentation. That makes the durable question more useful than a launch-day comparison: how should a support team test any GPT-5 generation before changing live traffic?
In brief
- Test the model on your sources, policies, tools, and handoff rules.
- Separate retrieval failure from answer-model failure.
- Measure valid tool calls and prevented unsafe actions, not only fluent replies.
- Compare quality, latency, and cost on the same cases before a staged release.
Do not start with a headline benchmark
The original GPT-5 developer announcement reports results on public and internal benchmarks. Those results describe the tested setup. They do not establish your support resolution rate, customer satisfaction, or hallucination rate.
A support system adds variables that model benchmarks do not cover:
- Your knowledge sources and chunking
- Retrieval and reranking
- System instructions and refusal rules
- Tool schemas, authentication, and authorization
- Conversation history
- Human handoff policy
Treat model selection as a product evaluation, not a leaderboard decision.
Build a support-specific test set
Use real questions after removing personal and account data, or write representative cases from approved policies. Include routine questions and the cases that make your team nervous.
Grounded questions
Ask questions whose answers appear clearly in one approved source. Label the supporting passage. Check whether the response:
- Gives the correct answer
- Cites or points to the right source
- Avoids adding terms that are not in the source
- Says it cannot answer when the source is missing
Run retrieval separately. If the correct passage never reaches the model, changing the answer model may not fix the problem.
Conflicting or outdated sources
Include a current policy and an older page with different terms. The system should follow the source priority or refuse when that priority is unclear. A confident blend of both policies is a failure.
Multi-turn cases
Test conversations where the user changes one fact, corrects the agent, or refers back to an earlier message. Examples include:
- Changing a product version midway through troubleshooting
- Asking a follow-up about a different invoice
- Correcting a shipping country
- Returning after a handoff was offered
Score whether the model uses the right fact at the right turn. A long context window does not prove correct context use.
No-answer and escalation cases
Include legal requests, account access without authentication, unsupported refunds, security reports, and questions outside the source set. The right result may be a refusal or a human handoff.
An answer rate close to 100 percent is a warning if some questions should not be answered.
Test tool calls as transactions
Tool calling matters when the agent can read an order, create a ticket, or change data. A syntactically valid call is not enough.
For each tool case, verify:
- The user is authenticated when account data is involved
- The user can access the selected resource
- Required fields come from the conversation or a trusted system
- The model does not invent missing IDs or amounts
- A destructive action needs explicit confirmation
- Retries use an idempotency key
- Tool errors lead to a clear message or handoff
Use a sandbox or a fake tool implementation during evaluation. Record the proposed arguments and expected arguments. Do not let a model test make real account changes.
type ToolCase = {
conversation: Message[]
expectedTool: string | null
expectedArguments: Record<string, unknown> | null
requiresConfirmation: boolean
}Score exact fields that matter. A correct tool name with the wrong subscription ID is still a failed case.
Use a plain scoring rubric
A compact rubric is easier to review than one combined "accuracy" score.
| Check | Pass condition |
|---|---|
| Source support | Every material claim is backed by an approved passage |
| Policy use | The response applies the current rule without adding exceptions |
| Tool choice | The tool matches the user's request and allowed action |
| Tool arguments | IDs, dates, amounts, and flags match trusted inputs |
| Safety | The model refuses or hands off when authorization or evidence is missing |
| Tone | The reply is direct, respectful, and does not promise an outcome it cannot control |
Keep severe failures separate. One unauthorized account action should not disappear inside a high average score.
Measure latency by stage
Record retrieval, model time to first token, full generation, and tool duration separately. A slower answer may come from a new reranker or tool, not the model.
For chat, compare:
- Time until the first useful text appears
- Time until citations and tool results are complete
- Timeout and retry rate
- Handoff time after a failed action
Use the same region, prompts, sources, concurrency, and output limit for both models.
Calculate cost from the evaluated traffic
Do not use an old price table in a model-selection document. Read the current provider price when you run the evaluation.
Calculate expected cost from measured token use:
monthly input cost = input tokens per case × monthly cases × input price
monthly output cost = output tokens per case × monthly cases × output price
monthly tool cost = tool calls per case × monthly cases × tool priceAdd retrieval, reranking, storage, observability, and human review. A model with a lower token price can cost more if it needs longer prompts or more retries. A stronger model can still be a bad purchase if the current one already passes the acceptance set.
Review prompts without rewriting everything
Run the current system prompt first. Then inspect failures. Change one instruction at a time and rerun the complete set.
Watch for:
- Repeated instructions that conflict
- Broad permission to infer missing account facts
- Vague handoff conditions
- Tool descriptions that omit authorization limits
- Output rules that hide uncertainty
Do not optimize the prompt against a handful of examples until it fails the rest of the suite.
Release in controlled steps
A sensible rollout looks like this:
- Pass the offline acceptance set.
- Run tools against a sandbox or read-only environment.
- Send a small share of eligible conversations to the new model.
- Review unsupported claims, handoffs, tool errors, latency, and cost.
- Increase traffic only when the live sample matches the offline result.
- Keep a model switch and the previous prompt available for rollback.
Define rollback conditions before the release. Examples include an unauthorized tool attempt, a rise in unsupported policy answers, or a tail-latency breach. The exact thresholds should come from your normal baseline and risk tolerance.
When not to upgrade
Stay on the current model when it already passes the support set, the new model adds latency without fixing a measured failure, or your main problem is missing content and retrieval.
Also wait if you cannot observe tool arguments, source use, and handoffs. Changing a model inside an opaque system makes the next failure harder to diagnose.
The model name is not the outcome. A safe upgrade is one that improves a defined set of support tasks and can be rolled back when it does not.
For the surrounding system, read RAG versus fine-tuning and preventing unsupported answers.
Frequently asked questions
Does GPT-5 guarantee fewer hallucinations in customer support?
No. Provider benchmarks do not guarantee results on your sources and workflow. Test grounded answers, missing-source cases, and conflicting policies in your complete system.
Should I use GPT-5 for every support question?
Not automatically. Compare models by question type and route only when the added complexity has a measured benefit.
How should I test GPT-5 tool calls?
Use sandboxed tools and labeled cases. Check authentication, authorization, exact arguments, confirmation, idempotency, and error handling.
What is the safest rollout plan?
Pass an offline set, use a small traffic sample, review high-risk failures, and keep a tested switch back to the previous model and prompt.