VJRA.US · Week 4 Assignment

Deploy_AI-Powered Customer Support Copilot

Product thinking · AI system design · human-governed deployment
← All Projects

Deployment · Controlled Rollout Plan

Deployment pathShadow → assisted pilot → measured expansion
Stage 0 · PrepareBaseline and gateGolden dataset, policy index, permissions, metrics, rollback
→
Stage 1 · ShadowNo customer exposureCompare copilot decisions with expert agent outcomes
→
Stage 2 · AssistAgent approval requiredLimited intents, high-value returning customers, chat and email
→
Stage 3 · ExpandGate by evidenceAdd intents and regions only after quality and safety thresholds pass
Deployment controlsSafe by default
  • Release unitVersioned prompt, model, retrieval index, policies, and rule set
  • ApprovalSupport, policy, security/privacy, and engineering owners
  • CanarySmall agent cohort and limited eligible intents
  • RollbackOne-click disable generation; retain deterministic search and APIs
  • AuditLog evidence, recommendation, edits, action, outcome, and versions
Deployment readinessMust pass before exposure
  • Quality≥95% grounded responses; severe errors <1%
  • Safety100% critical-action authorization tests
  • Reliabilityp95 ≤8 seconds; availability ≥99.5%; safe fallback proven
  • PrivacyLeast privilege, masking, retention, and access audit verified
  • OperationsDashboards, alerts, incident owner, runbook, and rollback drill ready
Problem statement

Help agents resolve e-commerce issues quickly, accurately, and consistently—without automating away accountability.

Today, agents manually combine order history, policy documents, and prior interactions. The copilot understands intent, retrieves the right context, drafts a grounded response, and escalates edge cases and high-risk decisions to people.

≥12%
Verified FCR lift
−20%
Median handle time
≥95%
Grounded responses
≥4.3
CSAT / 5
<1%
Severe errors
≤8s
p95 latency

Part 1 · Product Thinking (30–35%)

1 · Clarify the problemStart strong
  • What are we solving?Slow, inconsistent resolution caused by fragmented evidence and manual interpretation
  • Primary userFrontline e-commerce support agent
  • Current painRepeated lookup, context switching, policy ambiguity, and inconsistent drafts
  • SuccessMore correct first-contact resolutions with lower effort and no safety regression
  • AmbiguityAuthority limits, policy ownership, data access, channel scope, and baseline performance
2 · Assumptions & alignmentValidate before build
  • MarketMulti-region e-commerce; policies vary by market, seller, and product
  • User behaviorAgents review recommendations; customers expect fast, accurate, empathetic help
  • Business modelCustomer retention and operating efficiency both matter
  • AccessPermissioned order, CRM, policy, and interaction data are available
  • Business goalOptimize correct resolution—not contact deflection alone
3 · Supporting KPIsNorth Star plus diagnostics
  • EngagementAgent adoption, suggestion acceptance, and meaningful-use rate
  • RetentionSeven-day repeat-contact rate and 30-day repeat purchase
  • QualityGroundedness, policy accuracy, severe-error rate, CSAT
  • EfficiencyHandle time, escalation rate, cost per resolved contact
Out of scopeWhat we will not optimize
  • AutonomyNo unsupervised refunds, credits, cancellations, or account changes
  • RiskNo fraud appeals, legal threats, safety incidents, or regulated advice
  • ChannelsNo voice automation or public social replies in the pilot
  • Model workNo foundation-model training or premature fine-tuning
  • Metric gamingNo deflection target that weakens resolution quality
4 · Behavioral segmentationChoose one primary segment
SegmentBehaviorNorth Star impactFeasibilityDecision
New customersFirst purchase; little historyMediumHighLater
High-value returningFrequent orders and repeat post-purchase contactsHighHighPrimary
Casual returningInfrequent orders and contactsMediumHighLater
Exception-heavyFraud, legal, or repeated overridesHighLowHuman only
5 · Problem selectionList 3–4; choose one
ProblemImpactAI suitabilityDecision
Agents search multiple systems for order and policy contextHighHigh with grounded retrievalSelected
Agents repeatedly rewrite similar responsesMediumHighSupporting capability
Policy exceptions require judgmentHighMedium; high riskHuman escalation
Carrier status must be currentMediumLowDeterministic API
6 · SolutioningTwo focused capabilities; avoid over-design
01 · UnderstandIntent and entitiesIssue, order, item, market, requested outcome
→
02 · RetrieveEvidence packetOrder facts, policy clauses, relevant history
→
03 · AssistDraft + next actionCited answer and actions within authority
→
04 · GovernApprove or escalateAgent sends, edits, or routes to specialist

Part 2 · AI System Design (65–70%)

7 · Is this an AI problem?Hybrid system
CapabilityApproachWhy
Authentication, order status, eligibility, refund limits, action executionDeterministic rules + APIsExact, auditable, fail-closed behavior
Intent, entities, summarizationProbabilistic modelLanguage varies and may contain linked needs
Policy discoveryHybrid retrievalSemantic meaning plus exact market/version filters
Response draftGrounded LLMNatural language after facts and allowed actions are fixed
Material exceptionsHuman judgmentAccountability for financial and relationship decisions
8 · Model choiceLLM + rules + retrieval
  • LLMEnterprise hosted model with structured output, tool use, and strong instruction following
  • Traditional MLOptional fast intent/router model when volume justifies it
  • LatencySmall model for routing; stronger model only for complex drafts
  • CostContext limits, caching, and complexity-based routing
  • InterpretabilityCitations, rule decisions, confidence signals, and audit logs
  • Data availabilityStart with prompts/RAG; learn from reviewed production examples
9 · Pre-training considerationsDo not train a foundation model
  • Base modelBegin with a capable GPT-class or Claude-class enterprise model
  • GeneralizationUseful for varied language, multilingual contacts, and unseen phrasing
  • SpecializationSupply company facts through retrieval and tools—not model memory
  • Open sourceConsider later if privacy, unit economics, or deployment control outweigh operations burden
  • Selection testBenchmark candidates on the same golden dataset for quality, latency, and cost
10 · Post-training / adaptationWhen to use what
MethodUseDecision
Prompt engineeringRole, workflow, tone, structured output, citation, and abstention rulesUse now
RAGDynamic knowledge, large policy corpus, factual groundingUse now
Fine-tuningStable patterns, domain language, or repeated errors not solved by workflow/prompt changesDefer until reviewed data proves need

11–13 · RAG, Data, and Precision

11 · RAG decisionYes
  • External knowledge?Yes: order state, policy, promotions, logistics, and customer history
  • Hallucination risk?High: unsupported promises create financial and trust harm
  • SourcesPolicy CMS, order/logistics APIs, CRM history, approved macros, exception directory
  • ChunkingPolicy clause/decision unit with title, market, product, version, authority, effective date
  • RetrievalMetadata filter + keyword/vector search + reranking
  • GroundingNo material claim without cited evidence or deterministic fact
12 · Data strategyMinimize and govern
  • Internal logsDe-identified conversations, dispositions, escalation reasons, reopens, CSAT
  • User interactionsLive order and permissioned CRM context only when needed
  • Synthetic dataRare cases, adversarial wording, missing context, conflicting policies
  • Golden datasetConversation + evidence + approved intent/action/escalation/response
  • Expert reviewSupport QA and policy owner; dual review for high-risk examples
  • GovernancePurpose limits, residency, RBAC, access logs, retention, deletion
13 · Precision versus recall: precision first. For policy conclusions, commitments, and actions, a wrong answer is more harmful than an omitted suggestion. Retrieval can favor recall by finding a broad candidate set; reranking and generation must favor precision. Insufficient or conflicting evidence causes abstention or escalation.

14–15 · Evaluation and Human Oversight

14 · Evaluation & testingOffline → shadow → online
LayerMeasuresDecision gate
OfflineIntent accuracy/macro-F1; entity F1; retrieval recall@k and precision@k; citation precision; action-rule accuracy. BLEU/ROUGE are secondary only because overlap does not prove correctness.100% critical-action rules; ≥95% citation precision
Human evaluationCorrectness, relevance, evidence support, completeness, tone, safe escalation≥4.5/5 correctness; severe errors <1%
ShadowCompare recommendation with expert resolution; review high-risk disagreementsNo customer exposure until gates pass
Online A/BVerified FCR, task completion, CSAT, handle time, reopen, escalation, acceptance, edit distanceNorth Star improves with no safety/CSAT regression
15 · Human interventionRisk-based
  • Pilot draftsAgent reviews before sending
  • Mandatory escalationFraud, safety, legal, vulnerable customer, policy conflict, high-value exception
  • Confidence triggerMissing/conflicting evidence, unknown intent, or out-of-distribution request
  • Authority triggerAction exceeds agent role or monetary limit
Feedback systemRLHF-ready evidence
  • ExplicitThumbs up/down plus reason code
  • ImplicitEdit distance, rejected action, escalation, reopen, repeat contact
  • Expert loopWeekly QA of failures and random samples
  • Release loopConfirmed failures become regression tests before changes ship

16 · High-Level System Architecture

Input → model → output + retrieval + feedbackHuman governed
Agent WorkspaceConversation and authenticated context
Intent RouterClassify, extract, assess risk
Tools + RAGLive APIs and policy evidence
GuardrailsAuthority, PII, citation gates
LLM DraftCited response and action
Review + FeedbackApprove, edit, escalate, learn
Bonus design: Cache versioned, non-personal policy retrieval; invalidate on policy publication. Parallelize tool calls, stream generation, route by complexity, enforce timeouts, and fall back to deterministic facts plus standard search.

17–18 · Operationalization, Safety, and Guardrails

17 · OperationalizationReal-time assist
  • LatencyParallel retrieval, streaming, p95 draft ≤8 seconds
  • Token usageRetrieve only relevant chunks; summarize long histories; cap context
  • Inference costSmall models for routing; stronger model for complex cases
  • ScalabilityQueues, rate limits, autoscaling, circuit breakers, graceful degradation
  • Traffic spikesPrioritize active-agent requests and deterministic fallbacks
  • MonitoringLatency, cost, failures, groundedness, safety, adoption, business outcomes
  • DriftPolicy changes, intent mix, abstention, edits, and outcome degradation
18 · Safety & guardrailsDefense in depth
  • HallucinationCitation-required claims, source validation, abstention, output verification
  • Prompt injectionTreat retrieved/customer text as data; isolate instructions; restrict tools
  • ToxicityInput/output moderation and professional-tone contract
  • BiasParity tests across language, region, customer tier, and issue type
  • PrivacyLeast privilege, masking, encryption, retention controls, audit logs
  • ActionsAllowlisted functions with server-side validation and human approval

19–20 · Go / No-Go and Trade-offs

19 · Go / no-go frameworkPrecommitted thresholds
DimensionLaunch / expand ifDo not launch / roll back if
ValueVerified FCR improves ≥12%; handle time improves ≥20%No significant FCR gain or CSAT declines >0.1
QualityGrounded-response rate ≥95%Severe factual or policy errors ≥1%
Safety100% critical actions pass authorization testsUnauthorized action, privacy incident, or missed critical escalation
Reliabilityp95 ≤8 seconds and availability ≥99.5%Persistent failure without safe fallback
EconomicsCost per resolved eligible contact below manual baselineModel/tool cost erases measurable benefit
20 · Explicit trade-offsRecommended stance
Trade-offDecisionReason
Accuracy vs latencyAccept modest latency for verified evidenceA fast unsupported promise is worse than a grounded response
Cost vs qualityRoute by complexityClassification does not require the strongest generation model
Automation vs controlDraft and recommend; people approve material actionsPreserves accountability while collecting evidence
Recall vs precisionRetrieve broadly; answer narrowlyUnsupported claims are costlier than escalation
Personalization vs privacyUse only context needed for the current issueData minimization reduces exposure
RecommendationLaunch a limited agent-facing pilot for high-value returning customers and the highest-volume post-purchase intents. Begin in shadow mode, require agent approval, and expand only after the value, quality, safety, reliability, and economic gates are met.

Evaluation Criteria Coverage

Clarity of thinkingOne segment, one problem, one North Star
Depth in AIHybrid design, RAG, data, evaluation, adaptation
Trade-off awarenessFive explicit decisions and rationale
Structured communicationAll 20 requested sections mapped
Drive the discussionOpen questions and measurable launch gates