Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Dataford
Popular roles
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Browse
Browse All RolesEvery role hub, from analyst to MLBrowse All CompaniesCompany-specific interview loopsAll Interview GuidesThe full guide library
Top questions by role
Software EngineerData AnalystData ScientistData EngineerBusiness AnalystAI EngineerMachine Learning EngineerProduct Manager
Top questions by skill
SQLPythonStatisticsMachine LearningA/B TestingSystem DesignGenerative AIProduct SenseMetricsBehavioral
Browse all questions →Try a mock interview
Experiences
Practice
Mock InterviewsTimed interview simulations with feedbackSuccess PathYour 6-week structured planModulesCurated lessons by topicWebinarsTalks from ex-Big Tech data leadsPlaygroundA free-form scratch editor
Learn
BlogInterview strategy and career adviceTech Job Market ReportHiring trends across data and AI rolesFor UniversitiesDataford for career centersAbout DatafordWho we are and how we build
Pricing
Build my plan
Implementing a Classification Algorithm
00:00
5 left

Implementing a Classification Algorithm

MediumPython

Problem

Meta Platforms uses feature vectors to represent content such as Instagram Reels. Implement a k-nearest neighbors classifier that predicts each validation item's label from labeled training vectors, then computes validation accuracy.

Formal Specification

Given train_features, a list of numeric vectors, and train_labels, the corresponding string labels, classify every vector in validation_features. Use squared Euclidean distance, which avoids an unnecessary square root. For each validation vector, select the k closest training vectors. The predicted label is the label with the highest frequency among those neighbors. If multiple labels tie, return the lexicographically smallest label.

Return a dictionary with predictions, a list of predicted labels in validation order, and accuracy, the fraction of predictions equal to validation_labels.

Constraints

  • 1 <= len(train_features) <= 2 * 10^3
  • 1 <= len(validation_features) <= 500
  • All vectors have the same dimension
  • 1 <= dimension <= 50
  • 1 <= k <= len(train_features)
  • Labels are non-empty strings

Function Signature

def knn_classify(train_features, train_labels, validation_features, validation_labels, k):
Interviewer

Your question is Implementing a Classification Algorithm. Start with the requirements in the Question tab.

Run and submit as often as you like. When you're ready, talk me through your approach or go straight to the code.

You need to log in / sign up to run or submit.
CodePython 3
Sign up free to run your codeLog inLn 2
Run your code to see test output here.