Deployment · Controlled Rollout Plan
Deployment pathShadow → assisted pilot → measured expansion
Stage 0 · PrepareBaseline and gateGolden dataset, policy index, permissions, metrics, rollback
→
Stage 1 · ShadowNo customer exposureCompare copilot decisions with expert agent outcomes
→
Stage 2 · AssistAgent approval requiredLimited intents, high-value returning customers, chat and email
→
Stage 3 · ExpandGate by evidenceAdd intents and regions only after quality and safety thresholds pass
Deployment controlsSafe by default
- Release unitVersioned prompt, model, retrieval index, policies, and rule set
- ApprovalSupport, policy, security/privacy, and engineering owners
- CanarySmall agent cohort and limited eligible intents
- RollbackOne-click disable generation; retain deterministic search and APIs
- AuditLog evidence, recommendation, edits, action, outcome, and versions
Deployment readinessMust pass before exposure
- Quality≥95% grounded responses; severe errors <1%
- Safety100% critical-action authorization tests
- Reliabilityp95 ≤8 seconds; availability ≥99.5%; safe fallback proven
- PrivacyLeast privilege, masking, retention, and access audit verified
- OperationsDashboards, alerts, incident owner, runbook, and rollback drill ready
Problem statement
Help agents resolve e-commerce issues quickly, accurately, and consistently—without automating away accountability.
Today, agents manually combine order history, policy documents, and prior interactions. The copilot understands intent, retrieves the right context, drafts a grounded response, and escalates edge cases and high-risk decisions to people.
≥12%
Verified FCR lift
−20%
Median handle time
≥95%
Grounded responses
≥4.3
CSAT / 5
<1%
Severe errors
≤8s
p95 latency
Part 1 · Product Thinking (30–35%)
1 · Clarify the problemStart strong
- What are we solving?Slow, inconsistent resolution caused by fragmented evidence and manual interpretation
- Primary userFrontline e-commerce support agent
- Current painRepeated lookup, context switching, policy ambiguity, and inconsistent drafts
- SuccessMore correct first-contact resolutions with lower effort and no safety regression
- AmbiguityAuthority limits, policy ownership, data access, channel scope, and baseline performance
2 · Assumptions & alignmentValidate before build
- MarketMulti-region e-commerce; policies vary by market, seller, and product
- User behaviorAgents review recommendations; customers expect fast, accurate, empathetic help
- Business modelCustomer retention and operating efficiency both matter
- AccessPermissioned order, CRM, policy, and interaction data are available
- Business goalOptimize correct resolution—not contact deflection alone
3 · Supporting KPIsNorth Star plus diagnostics
- EngagementAgent adoption, suggestion acceptance, and meaningful-use rate
- RetentionSeven-day repeat-contact rate and 30-day repeat purchase
- QualityGroundedness, policy accuracy, severe-error rate, CSAT
- EfficiencyHandle time, escalation rate, cost per resolved contact
Out of scopeWhat we will not optimize
- AutonomyNo unsupervised refunds, credits, cancellations, or account changes
- RiskNo fraud appeals, legal threats, safety incidents, or regulated advice
- ChannelsNo voice automation or public social replies in the pilot
- Model workNo foundation-model training or premature fine-tuning
- Metric gamingNo deflection target that weakens resolution quality
4 · Behavioral segmentationChoose one primary segment
| Segment | Behavior | North Star impact | Feasibility | Decision |
|---|---|---|---|---|
| New customers | First purchase; little history | Medium | High | Later |
| High-value returning | Frequent orders and repeat post-purchase contacts | High | High | Primary |
| Casual returning | Infrequent orders and contacts | Medium | High | Later |
| Exception-heavy | Fraud, legal, or repeated overrides | High | Low | Human only |
5 · Problem selectionList 3–4; choose one
| Problem | Impact | AI suitability | Decision |
|---|---|---|---|
| Agents search multiple systems for order and policy context | High | High with grounded retrieval | Selected |
| Agents repeatedly rewrite similar responses | Medium | High | Supporting capability |
| Policy exceptions require judgment | High | Medium; high risk | Human escalation |
| Carrier status must be current | Medium | Low | Deterministic API |
6 · SolutioningTwo focused capabilities; avoid over-design
01 · UnderstandIntent and entitiesIssue, order, item, market, requested outcome
→
02 · RetrieveEvidence packetOrder facts, policy clauses, relevant history
→
03 · AssistDraft + next actionCited answer and actions within authority
→
04 · GovernApprove or escalateAgent sends, edits, or routes to specialist
Part 2 · AI System Design (65–70%)
7 · Is this an AI problem?Hybrid system
| Capability | Approach | Why |
|---|---|---|
| Authentication, order status, eligibility, refund limits, action execution | Deterministic rules + APIs | Exact, auditable, fail-closed behavior |
| Intent, entities, summarization | Probabilistic model | Language varies and may contain linked needs |
| Policy discovery | Hybrid retrieval | Semantic meaning plus exact market/version filters |
| Response draft | Grounded LLM | Natural language after facts and allowed actions are fixed |
| Material exceptions | Human judgment | Accountability for financial and relationship decisions |
8 · Model choiceLLM + rules + retrieval
- LLMEnterprise hosted model with structured output, tool use, and strong instruction following
- Traditional MLOptional fast intent/router model when volume justifies it
- LatencySmall model for routing; stronger model only for complex drafts
- CostContext limits, caching, and complexity-based routing
- InterpretabilityCitations, rule decisions, confidence signals, and audit logs
- Data availabilityStart with prompts/RAG; learn from reviewed production examples
9 · Pre-training considerationsDo not train a foundation model
- Base modelBegin with a capable GPT-class or Claude-class enterprise model
- GeneralizationUseful for varied language, multilingual contacts, and unseen phrasing
- SpecializationSupply company facts through retrieval and tools—not model memory
- Open sourceConsider later if privacy, unit economics, or deployment control outweigh operations burden
- Selection testBenchmark candidates on the same golden dataset for quality, latency, and cost
10 · Post-training / adaptationWhen to use what
| Method | Use | Decision |
|---|---|---|
| Prompt engineering | Role, workflow, tone, structured output, citation, and abstention rules | Use now |
| RAG | Dynamic knowledge, large policy corpus, factual grounding | Use now |
| Fine-tuning | Stable patterns, domain language, or repeated errors not solved by workflow/prompt changes | Defer until reviewed data proves need |
11–13 · RAG, Data, and Precision
11 · RAG decisionYes
- External knowledge?Yes: order state, policy, promotions, logistics, and customer history
- Hallucination risk?High: unsupported promises create financial and trust harm
- SourcesPolicy CMS, order/logistics APIs, CRM history, approved macros, exception directory
- ChunkingPolicy clause/decision unit with title, market, product, version, authority, effective date
- RetrievalMetadata filter + keyword/vector search + reranking
- GroundingNo material claim without cited evidence or deterministic fact
12 · Data strategyMinimize and govern
- Internal logsDe-identified conversations, dispositions, escalation reasons, reopens, CSAT
- User interactionsLive order and permissioned CRM context only when needed
- Synthetic dataRare cases, adversarial wording, missing context, conflicting policies
- Golden datasetConversation + evidence + approved intent/action/escalation/response
- Expert reviewSupport QA and policy owner; dual review for high-risk examples
- GovernancePurpose limits, residency, RBAC, access logs, retention, deletion
13 · Precision versus recall: precision first. For policy conclusions, commitments, and actions, a wrong answer is more harmful than an omitted suggestion. Retrieval can favor recall by finding a broad candidate set; reranking and generation must favor precision. Insufficient or conflicting evidence causes abstention or escalation.
14–15 · Evaluation and Human Oversight
14 · Evaluation & testingOffline → shadow → online
| Layer | Measures | Decision gate |
|---|---|---|
| Offline | Intent accuracy/macro-F1; entity F1; retrieval recall@k and precision@k; citation precision; action-rule accuracy. BLEU/ROUGE are secondary only because overlap does not prove correctness. | 100% critical-action rules; ≥95% citation precision |
| Human evaluation | Correctness, relevance, evidence support, completeness, tone, safe escalation | ≥4.5/5 correctness; severe errors <1% |
| Shadow | Compare recommendation with expert resolution; review high-risk disagreements | No customer exposure until gates pass |
| Online A/B | Verified FCR, task completion, CSAT, handle time, reopen, escalation, acceptance, edit distance | North Star improves with no safety/CSAT regression |
15 · Human interventionRisk-based
- Pilot draftsAgent reviews before sending
- Mandatory escalationFraud, safety, legal, vulnerable customer, policy conflict, high-value exception
- Confidence triggerMissing/conflicting evidence, unknown intent, or out-of-distribution request
- Authority triggerAction exceeds agent role or monetary limit
Feedback systemRLHF-ready evidence
- ExplicitThumbs up/down plus reason code
- ImplicitEdit distance, rejected action, escalation, reopen, repeat contact
- Expert loopWeekly QA of failures and random samples
- Release loopConfirmed failures become regression tests before changes ship
16 · High-Level System Architecture
Input → model → output + retrieval + feedbackHuman governed
Agent WorkspaceConversation and authenticated context
Intent RouterClassify, extract, assess risk
Tools + RAGLive APIs and policy evidence
GuardrailsAuthority, PII, citation gates
LLM DraftCited response and action
Review + FeedbackApprove, edit, escalate, learn
Bonus design: Cache versioned, non-personal policy retrieval; invalidate on policy publication. Parallelize tool calls, stream generation, route by complexity, enforce timeouts, and fall back to deterministic facts plus standard search.
17–18 · Operationalization, Safety, and Guardrails
17 · OperationalizationReal-time assist
- LatencyParallel retrieval, streaming, p95 draft ≤8 seconds
- Token usageRetrieve only relevant chunks; summarize long histories; cap context
- Inference costSmall models for routing; stronger model for complex cases
- ScalabilityQueues, rate limits, autoscaling, circuit breakers, graceful degradation
- Traffic spikesPrioritize active-agent requests and deterministic fallbacks
- MonitoringLatency, cost, failures, groundedness, safety, adoption, business outcomes
- DriftPolicy changes, intent mix, abstention, edits, and outcome degradation
18 · Safety & guardrailsDefense in depth
- HallucinationCitation-required claims, source validation, abstention, output verification
- Prompt injectionTreat retrieved/customer text as data; isolate instructions; restrict tools
- ToxicityInput/output moderation and professional-tone contract
- BiasParity tests across language, region, customer tier, and issue type
- PrivacyLeast privilege, masking, encryption, retention controls, audit logs
- ActionsAllowlisted functions with server-side validation and human approval
19–20 · Go / No-Go and Trade-offs
19 · Go / no-go frameworkPrecommitted thresholds
| Dimension | Launch / expand if | Do not launch / roll back if |
|---|---|---|
| Value | Verified FCR improves ≥12%; handle time improves ≥20% | No significant FCR gain or CSAT declines >0.1 |
| Quality | Grounded-response rate ≥95% | Severe factual or policy errors ≥1% |
| Safety | 100% critical actions pass authorization tests | Unauthorized action, privacy incident, or missed critical escalation |
| Reliability | p95 ≤8 seconds and availability ≥99.5% | Persistent failure without safe fallback |
| Economics | Cost per resolved eligible contact below manual baseline | Model/tool cost erases measurable benefit |
20 · Explicit trade-offsRecommended stance
| Trade-off | Decision | Reason |
|---|---|---|
| Accuracy vs latency | Accept modest latency for verified evidence | A fast unsupported promise is worse than a grounded response |
| Cost vs quality | Route by complexity | Classification does not require the strongest generation model |
| Automation vs control | Draft and recommend; people approve material actions | Preserves accountability while collecting evidence |
| Recall vs precision | Retrieve broadly; answer narrowly | Unsupported claims are costlier than escalation |
| Personalization vs privacy | Use only context needed for the current issue | Data minimization reduces exposure |
RecommendationLaunch a limited agent-facing pilot for high-value returning customers and the highest-volume post-purchase intents. Begin in shadow mode, require agent approval, and expand only after the value, quality, safety, reliability, and economic gates are met.
Evaluation Criteria Coverage
Clarity of thinkingOne segment, one problem, one North Star
Depth in AIHybrid design, RAG, data, evaluation, adaptation
Trade-off awarenessFive explicit decisions and rationale
Structured communicationAll 20 requested sections mapped
Drive the discussionOpen questions and measurable launch gates