Job Description & Details
This contract Data Scientist role on Walmart's Catalog Team puts you right in the middle of massive-scale e-commerce challenges. You will be building machine learning and NLP pipelines that directly impact how millions of products are classified, searched, and displayed. If you like turning messy enterprise data into structured intelligence using modern ML techniques, this is a solid gig.
What You'll Actually Be Doing
Your day-to-day will involve designing and deploying ML models to handle thorny problems like product deduplication, attribute extraction, and semantic matching across billions of catalog items. You'll spend a significant amount of time writing clean Python and Spark jobs to process large-scale retail datasets, ensuring the data quality remains high before it hits production systems. Expect to collaborate closely with product and engineering teams to push your NLP and GenAI solutions live.
The Core Tech Stack
To survive here, you need to be completely fluent in Python and SQL, alongside heavy hands-on experience with NLP workflows like text classification, embeddings, and transformer models. Because you are dealing with Walmart-scale inventory, proficiency in Spark or PySpark for distributed data processing is non-negotiable. You'll also need a firm grasp of deep learning frameworks like PyTorch or TensorFlow, and familiarity with LLMs or GenAI is a major plus for modernizing their catalog intelligence.
Interview Expectations
Expect the hiring team to grill you on how you handle text classification and embedding generation at scale. They'll likely ask you to explain a time you optimized a slow Spark query or handled massive data drift in a production ML pipeline. The interviewer is secretly looking for your ability to balance academic ML theory with pragmatic engineering constraints, so make sure you emphasize latency and scalability in your answers.
Application Advice
Make sure your resume doesn't just list technologies, but explicitly highlights production deployments of NLP and machine learning models. You need to weave keywords like PySpark, text classification, embeddings, Python, and SQL naturally into your experience bullets so you clear the automated screening. If you've worked on retail catalog data, search relevance, or large-scale deduplication before, bring that front and center.