Job Description & Details
This gig is a senior‑level data science contract focused on tightening up search relevance for a big e‑commerce player. You’ll be knee‑deep in Learn‑to‑Rank models, vector similarity, and large‑scale clickstream data—nothing trivial, and the impact is measured directly in revenue‑moving search metrics.
What You'll Actually Be Doing
You’ll design and ship LTR pipelines that ingest raw clickstream, generate features, and feed them into ranking models built on XGBoost or deep nets. Expect to spend a lot of time tuning vector indexes in Elasticsearch or Faiss, writing Spark jobs to keep the data flowing, and collaborating with engineers to push the models into production with low latency. The day ends with you presenting A/B test results to product and UX, proving that the new ranking actually moves the needle.
The Core Tech Stack
The role lives at the intersection of Python‑centric ML (scikit‑learn, XGBoost, TensorFlow/PyTorch) and search infrastructure (Elasticsearch, Solr, Vespa, plus ANN libs like Faiss). You must be comfortable turning raw clickstream into dense embeddings, then stitching those embeddings into a real‑time retrieval pipeline. Spark/Hadoop experience isn’t a must‑have but helps when you need to process terabytes of logs overnight.
Interview Expectations
- “Walk me through how you would build a Learn‑to‑Rank pipeline from raw click data to a deployed model.” They’re looking for a end‑to‑end story: feature extraction, label generation, model choice, offline evaluation (NDCG, MAP), and online serving with low latency. 2. “Explain the trade‑offs between using Faiss IVF‑PQ vs. HNSW for billion‑scale vector search.” Expect you to discuss index build time, recall vs. latency, memory footprint, and how those choices affect real‑time personalization.
Application Advice
Tailor your resume to highlight “Learn to Rank”, “vector similarity search”, “Elasticsearch/Faiss”, and “large‑scale data pipelines (Spark/Hadoop)”. Use exact phrasing from the JD—e.g., “e‑commerce search relevance”, “clickstream feature engineering”, “retrieval algorithms”. Quantify impact where possible (e.g., “improved NDCG by 12% in A/B test”). A short cover note that mentions you can start immediately will also help you pass the ATS and catch the recruiter’s eye.