Scaling distributed machine learning pipelines across millions of rows using Hadoop, PySpark, and SQL analytical data warehouse engines.
With the rise of Big Data, Rohit scaled machine learning algorithms to distributed compute clusters, leveraging Hadoop MapReduce and PySpark to train models on terabytes of enterprise data.
Engineered distributed PySpark pipelines capable of vectorizing 50M+ customer interaction rows in under 20 minutes.
Transitioned row-oriented databases to Parquet columnar storage, accelerating analytical query speeds by 12x.
Configured PySpark MLlib gradient boosted trees across distributed multi-node Hadoop clusters.
Schedule a 1-on-1 technical session directly with AI & Data Science Consultant Rohit.
AI & Data Science Consultant
2+ Decades AI ExperienceFirst project in AI & ANN in 2004 at IIT Roorkee under the mentorship of Dr. Sunil Padhi (HOD, Electrical Department), writing neural network backpropagation in C language to predict solar sunspots. Today designing stateful Agentic AI networks at rcode.in.
View All 12 Milestones