Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started
Cosine Similarity Vector Search
00:00
5 left

Cosine Similarity Vector Search

MediumPython

Problem

Tata Consultancy Services (North America) needs a basic semantic retrieval component that ranks document embeddings against a query embedding. Implement vector search from scratch using cosine similarity, without machine learning or numerical libraries.

Given documents, a list of equal-length numeric vectors, a numeric query vector, and k, return the indices of the k documents with the highest cosine similarity to the query.

Formal Specification

  • Input: documents, a list of lists of numbers; query, a list of numbers with the same dimension; and integer k.
  • Output: A list of at most k integer indices, sorted by descending cosine similarity. If scores tie, return the smaller index first.
  • Cosine similarity is dot(a, b) / (||a|| * ||b||). If either vector has zero magnitude, define its similarity as 0.0.

Constraints

  • 1 <= len(documents) <= 10^4
  • 1 <= len(query) <= 100
  • Every document has the same dimension as query
  • 0 <= k <= len(documents)
  • Vector components are finite numbers in [-10^6, 10^6]

Function Signature

def vector_search(documents, query, k):
Interviewer

Your question is Cosine Similarity Vector Search. Start with the requirements in the Question tab.

Run and submit as often as you like. When you're ready, talk me through your approach or go straight to the code.

You need to log in / sign up to run or submit.
CodePython 3
You need to log in / sign up to run or submit.Ln 2
Run your code to see test output here.