Data-driven Product Thinker building scalable growth & experimentation systems

De facto Product Manager for WorkIndia's Candidate Growth team — owning the CLM roadmap end-to-end. I prioritise with an effort-vs-impact framework, partner with engineering and design to ship, and run the experiments that prove whether we moved acquisition, activation, or retention. I don't just analyse; I hypothesise, experiment, and close the loop.

Measurable Impact

Numbers that reflect decisions made, not just tasks completed.

Product Case Studies

How I've approached real problems — from hypothesis to impact.

WorkIndia - OTP Onboarding & Activation

Activation Funnel Diagnosis Vendor Strategy

Verification was the biggest activation drop-off - so logins skip OTP entirely on a saved token, and registrations run a four-channel fallback ladder.

96.3%OTP success
96.6%5-sec delivery
98.2%login segment
1Diagnose 2Ladder 3Tokens 42nd vendor

WorkIndia - Job Discovery & Apply Conversion

Conversion User Research Behavioural Analytics Ranking

The survey said WhatsApp, and the CTA won +30% - but apply rate was still short, so I built a click heatmap and found the real cause.

+30%Lands-to-Apply
+6%applies (density)
+8.1%applies (ranking)
1Ask 2Ship 3Heatmap 4Fix

WorkIndia - CLM Channel Efficiency & Lifecycle Growth

Growth Bandits Cost Efficiency Lifecycle

Spend was outgrowing DAU, so send-time became a contextual bandit rewarded on applications rather than opens - plus algorithm reuse and a new email channel.

−23%cost per DAU
+17%DAU
₹1.2Linfra saved
1Analyse 2Bandit 3Reuse 4Email

WorkIndia - Acquisition Funnel Experimentation

Experimentation Statistics Acquisition

A 1M+ daily visitor funnel was being changed on intuition, so I installed a hypothesis → power → guardrail loop and gated rollout on significance.

+6.2%Clicks-to-Leads
1M+daily visitors
Stat-sigrollout gate
1Hypothesise 2Power 3Test 4Gate

Hevo Data - Funnel Conversion Engine

Product Design Conversion Self-serve

Technical evaluators couldn't estimate cost before talking to sales - so I gave them a calculator, and a chatbot on the content pages that fed it.

+9%qualified leads
80+ETL pipelines
50+blogs instrumented
1Map funnel 2Build calculator 3Deploy bot

Samsung Prism - LLM Safety Guardian

ML Systems NLP Anomaly Detection

Unexplained moderation flags were expensive to review - so every flag got an explanation grounded in the exact policy clause it allegedly violated.

86.3%accuracy
−40%false positives
32,000+documents
1Extract 2Classify 3Explain

Experience

Where I've worked and the systems behind the results.

Data Analyst - Growth & Product (CLM)
WorkIndia
May 2025 - Present

De facto Product Manager for the Candidate Growth team - owning the CLM roadmap end-to-end, prioritising on effort-vs-impact, and running the experiments that prove whether we moved acquisition, activation, or retention.

Activation & Onboarding
  • Lifted OTP success 89.6% → 96.3% by diagnosing verification as the largest activation bottleneck and rebuilding it as a four-channel recovery ladder.
  • 98.2% on the login segment (~56% of volume) via one-tap saved-token re-auth, plus a secondary SMS vendor for delivery redundancy.
Discovery & Conversion
  • +30% Lands-to-Apply - surveyed candidates (45% wanted a WhatsApp CTA vs 18% calling), then shipped a WhatsApp HR CTA with design and engineering.
  • +6% then +8.1% Applies - a self-built click heatmap exposed pagination CTR at ~2x the Apply rate, so density went 10 → 20 jobs/page and ranking became relevancy-based.
  • +22% notification CTR from consumer-app design patterns: dynamic backgrounds, countdown timers, keyword highlighting, multi-CTA layouts.
Lifecycle & Growth (CLM)
  • −23% Cost-per-DAU and ₹1.2L infra saved - the Notification Time Affinity Model, a contextual Thompson Sampling send-time system rewarded on applications, not opens.
  • −13% Cost-per-Download, +16% CTR by tracing why in-app recommendation beat retargeting and reusing the stronger algorithm via Kafka.
  • +17% DAU from lifecycle initiatives plus a new Email channel at ₹0.17 cost-per-DAU (vs ₹1.3), now ~23% of CLM DAU.
  • +6.2% Clicks-to-Leads - owned experimentation strategy for a 1M+ daily visitor acquisition funnel.
A/B Testing Experiment Design Click Heatmap Thompson Sampling SQL / Athena Python Kafka WhatsApp / SMS / Email / Push BRD / PRD / ARD Metabase GitHub Projects
Technical Research Analyst
Hevo Data
Jun 2024 - Dec 2024

Data infrastructure, AI product tooling, and content-driven conversion - systems that turned technical evaluators into qualified leads.

Data Infrastructure
  • Investigated and resolved issues across 80+ distributed ETL pipelines (real-time and batch), improving reliability for downstream analytics and reporting.
  • Log-level and metric-driven RCA across Snowflake, Redshift, and BigQuery destinations, plus Snowflake credit pricing in depth - warehouse sizing, query optimisation, per-pipeline cost attribution.
Product & Conversion
  • +9% qualified leads from a Snowflake-based pricing tool - a self-serve calculator embedded at mid-funnel content pages.
  • Mapped the technical evaluator → MQL funnel, identifying pricing ambiguity as the primary drop-off driver - which motivated the calculator.
AI Product Tooling
  • Built a context-sensitive support chatbot (LangChain + Gemini 1.5 Flash + FAISS) on 50+ technical blogs with sub-second retrieval.
  • Structured-output prompt engineering to hold brand voice and token limits, with usage instrumented to surface top unanswered queries for the content roadmap.
LangChain Gemini 1.5 Flash FAISS Snowflake REST APIs Python ETL/ELT Redshift
Research & Development Intern
Samsung Prism
Nov 2023 – May 2024

ML systems for app policy violation detection - transformer classification and summarization, anomaly detection, and RAG explainability to cut false positives at scale.

ML Modelling
  • Fine-tuned Roberta-Large for multi-class policy violation classification on Samsung's app review dataset - 86.3% accuracy held out.
  • Transformer summarization models balancing accuracy, throughput, and deployment efficiency, plus anomaly monitoring over 10K+ interactions via sliding-window z-scores.
Data Pipelines & RAG
  • Processed 32,000+ documents through scalable pipelines for structured extraction, feeding both the classifier and the retrieval index.
  • −40% false positives from a RAG pipeline (Llama 3.2 1B) grounding every flag in a dense index of policy clauses - moderation became auditable, not hallucinated.
Roberta-Large Llama 3.2 RAG PyTorch HuggingFace Python Summarization Anomaly Detection

Experimentation & Growth Systems

Infrastructure that makes decisions repeatable, not one-off.

Notification Time Affinity Model (Contextual Thompson Sampling)

Send-time as a contextual bandit - rewarded on downstream value, bounded by fatigue caps.

A/B Design on a 1M+ Visitor Funnel

Stated hypothesis, powered sample size, pre-declared guardrails, significance gate before rollout.

Progressive Fallback Ladder (OTP)

Channel recovery as a measured sequence - each rung ordered by evidence, measured on its own.

Diagnose → Ship → Re-diagnose

The loop for when your first fix works and the metric still doesn't move.

Composite & Efficiency Metric Design

Metrics built so a campaign can't win by sacrificing user trust, and so channels compare on one number.

Segment Refresh → Copy Regeneration Loop

A 14-day cycle that doesn't just re-bucket users - it regenerates the copy those users receive.

Independent Technical Projects

Explorations in modeling and simulation outside of work.

Evolved > Programmed

EvoDrive - Genetic Algorithm Simulation

Can evolutionary algorithms solve pathfinding better than rule-based systems?

Fine-tuned LLM for SQL

Text-to-SQL (StarCoder2 Fine-tune)

How cheap can you make an LLM that's actually good at SQL generation?

RAG Search Engine

Perplexa

A conversational engine capable of grounded, real-time web retrieval.

CV + Analytics Pipeline

Vehicle Movement Analysis & Insight Generation

Extracting structured, queryable data from unstructured video streams of vehicle traffic.

How I Think About Products

Mental models I use to make better product decisions.

01

Optimise for downstream metrics, not vanity metrics

Open rate is a proxy; application rate is an outcome. I trace the metric chain to what actually matters, and guardrail the rest.

02

Balance engagement with user fatigue

More notifications ≠ more engagement. Every decision spends an attention budget, so I model unsubscribe cost alongside click gain.

03

Prefer systems over one-off solutions

A script that runs once is technical debt. A pipeline that reruns, logs itself, and handles edge cases is a product. I build for the second run.

04

Use data to guide, not dictate decisions

Data shows what happened, rarely why. I pair the numbers with behavioural context instead of outsourcing the decision to a metric.

05

Ship experiments, not assumptions

Every strong opinion is a hypothesis in disguise. Small fast experiments before large builds minimise the cost of being wrong.

06

Make the invisible legible

The best insight is the one nobody sees yet - buried in event logs, latency spikes, drop-off points. Data exploration is a product skill.

Education

KIIT University

B.Tech in Computer Science (2021-2025)
CGPA: 8.9

Sri Chaitanya Techno School

All India Senior School Certificate Examination (2020-2021)
Percentage: 93%

St. Patricks HS School

Indian Certificate of Secondary Education Examination (2018-2019)
Percentage: 92%

Publications

Precision Agriculture: Digital Twins and Advanced Crop Recommendation

IEEE ICOCT 2025 (Feb 2025)

Authors: Sayan Banerjee, Aniruddha Mukherjee, Suket Kamboj
DOI: 10.48550/arXiv.2502.04054

Efficient Waste Collection and Filtration using IOT

IJSREM (Jan 2023)

Authors: Sayan Banerjee, Rahul Naugariya, Shubham Patel, Shubham Kumar
DOI: 10.55041/IJSREM17403

Skills

Achievements

1st Position in Eureka Innovation

IIT Kharagpur (Feb 2024)

Machine Learning Specialization

DeepLearning.ai (Mar 2023)

Applied Python

Udemy (Dec 2022)

Extra-Curricular

Volunteering

Writing & Blogs

View Medium Profile →

I write about product management, experimentation systems, and translating data science into growth strategy.