Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Cluster Retail Shoppers for Personalization

Easy
Machine LearningCross-ValidationUnsupervised LearningFeature Engineering

Problem

Business Context

ShopSphere, an e-commerce marketplace with 2.4M monthly active users, wants to segment shoppers into behavior-based groups for lifecycle marketing and merchandising. There are no labeled segment definitions today, so the team needs an unsupervised clustering solution that is stable, interpretable, and usable in production.

Dataset

The modeling table is built at the customer-month level from the last 12 months of activity.

Feature GroupCountExamples
Purchase behavior8orders_30d, avg_order_value, discount_share, return_rate
Browsing behavior6sessions_30d, product_views, search_queries, dwell_time
Category affinity10pct_spend_electronics, pct_spend_home, pct_spend_fashion
Engagement & tenure5days_since_last_order, account_age_days, email_click_rate
Geography / device4region, device_type, app_vs_web_share
  • Size: 420K customer-month rows, 33 features
  • Missing data: ~12% missing in email engagement for unsubscribed users, ~4% missing in browsing metrics due to tracking gaps
  • Data characteristics: mixed numerical and categorical features, strong skew in spend/order variables, many low-activity users

Success Criteria

A good solution should produce 4-8 actionable clusters with:

  • silhouette score >= 0.20 after preprocessing
  • cluster stability (Adjusted Rand Index across resamples) >= 0.75
  • clear business interpretation for marketing and merchandising teams

Constraints

  • Segments must be explainable to non-technical stakeholders
  • Batch scoring must finish in under 30 minutes weekly
  • The team prefers simple maintenance over highly complex deep learning approaches

Deliverables

  1. Explain what clustering is and when it is appropriate instead of classification or regression.
  2. Build a clustering pipeline for customer segmentation, including preprocessing and feature engineering.
  3. Compare at least two clustering algorithms and justify the final choice.
  4. Evaluate cluster quality using quantitative metrics and qualitative interpretation.
  5. Describe how you would deploy, monitor, and refresh segments in production.

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Next questions
Cluster Retail Customers for SegmentationEasyTiger AnalyticsClassify and Segment Retail CustomersEasyMidwest Employers CasualtySegment Shoppers and Predict PurchasesEasy