Nebius Senior Technical Product Manager Interview Experience
Nebius · Senior · Product Manager
Interview process
I had a good recruiter round. In the second product case round with HM, they asked me deep technical questions on AI
Interview rounds · 2
- 1
Recruiter screen
BehavioralCross-Functional- Q1. Tell me about your work in AI
- Q2. Tell me about the most difficult product you shipped
- 2
Final / onsite round
Product DesignProduct StrategySystem DesignQ1. Design an inference batching system for a single GPU that can handle up to 100 inputs per batch while users wait synchronously, maximizing utilization under compute constraints.
How they answeredClarifying questions:
Digital-native Companies
Foundation model providers are not the target
Geo: US, EU, UK
Time : 1 Q
Resourcing : couple of engineers
Vision / Mission of Nebius:
To provide AI cloud services in a scalable and vertically integrated manner to various types of customers ( foundational providers, enterprise, digital (startups incl), AI startups)
Stage of product (Token Factory):
Hypergrowth stage
Goal of product (Token Factory):
Goal of Batch Inference in Token Factory is to increase adoption of token factory amongst digital startups and digital companies.
More details: To create a scalable, reliable, easy to use (low code / no code) service that can run process data and inference in batches
Why does Nebius pursue Batch inference?
Many digital companies have use cases where
Cost is a limitation (can onboard customers with limited token spent)
GPU Capacity is a limitation
Use cases can tolerate latency
Why can we give better cost for Batch inference? How to calculate the discount?
GPU Capacity (idle GPUs are used)
Segments
Digital companies and digital startups
Pain points
Cost spent per use case (has limits) as companies have limited budget H H
UX ( product should be usable by personas such as PM / PM/ BA, not just only MLE ) needs to low-code /no code and user friendly VH H
Quality of responses M M
Solution
Product & User Experience
Graphical user interface ( drag and drop style ) to enable everyone in digital companies to create
Select open source LLM model with guidance on cost and TTFT
Implement prefix caching / semantic caching
Call batch inference as a module inside the UI (Call GenAI gateways) - MVP
Vector store from embeddings - MVP
Prompt-based training of the batch inference
Test output anytime using a chatbot-style UI
Technical & Operational Strategy
Observability using another UI
Allow them to measure E2E time for every request
Allow user to see the response and the input
Allow user to see tokens consumed
Go to Market & Execution
Adoption metric
Number of users using batch inference UI product for min x days
Number of live use cases running built using batch inference UI product
Tips from the candidate
- Deep-dive on AI tools and concepts a lot
- Brush up GPU optimisation and cloud platforms
- Know difference between training and inference patterns
Company culture
Its a company where everyone is an ML engineer, irrespective of title (PM / Management all are ML engineers)