# Harsh Bhanushali > AI/ML & Full-Stack Systems Engineer building production-grade LLM architectures, model quantization pipelines, and end-to-end SaaS platforms. Creator of the top-downloaded 4-bit GPTQ MedGemma 27B model on Hugging Face (1,500+ downloads) and architect of enterprise platforms including Pulscribe.ai and WanguCRM. ## Core Identity & Overview - **Name:** Harsh Bhanushali - **Title:** AI/ML & Full-Stack Systems Engineer - **Official Website:** https://www.harshbhanushali.in - **Location:** Gujarat, India - **Education:** B.Tech in Information Technology, Charotar University of Science and Technology (CSPIT) - **Primary Specializations:** Large Language Model (LLM) Quantization, vLLM High-Throughput Production Inference, Continuous Batching, PagedAttention, Local RAG Systems, Full-Stack SaaS Architecture (Next.js, FastAPI, PostgreSQL, Kubernetes), Computer Vision (YOLOv8 research published at ICRAIC'24). ## Key Achievements & Verified Metrics - **MedGemma 27B Quantization:** Quantized Google's 54GB MedGemma 27B into a 15GB 4-bit GPTQ model optimized for vLLM with Marlin CUDA kernels. Surpassed **1,500+ downloads** on Hugging Face with 600+ tokens/sec throughput on a single GPU. - **Published Research:** First author on "Human Detection & Counting Using YOLOv8" published in AIP Conference Proceedings (ICRAIC'24), DOI: 10.1063/5.0254163. - **Enterprise Healthcare AI (Pulscribe.ai):** Built an all-in-one HIPAA-compliant AI medical scribing, EHR/EMR sync, and ICD-10 automated coding platform on Kubernetes microservices. - **WhatsApp SaaS Platform (WanguCRM):** Designed and deployed a complete CRM platform leveraging the official WhatsApp Business Cloud API. - **Gujarati F5-TTS Voice Model:** Trained and open-sourced an expressive regional Gujarati text-to-speech voice model on Hugging Face. ## Technical Case Studies & Blog Articles - [Compressing a 54GB Medical AI Brain into 15GB: How I Quantized MedGemma 27B for Production Inference](https://www.harshbhanushali.in/blog/compressing-medgemma-27b-quantization): Complete engineering breakdown of quantizing Google's MedGemma 27B to 4-bit GPTQ with domain calibration, Marlin kernel injection, and 600+ tokens/sec benchmarks. - [vLLM vs. Ollama: The Hard Numbers on Production LLM Inference](https://www.harshbhanushali.in/blog/vllm-vs-ollama-production-inference): Comprehensive benchmark analysis across 9 production dimensions, showing why vLLM achieves 19.3x higher system throughput (793.4 TPS vs 41.2 TPS) and 80ms vs 673ms P99 latency compared to Ollama under multi-user concurrency. ## Open-Source Models on Hugging Face - [HarshBhanushali7705/medgemma-27b-text-it-GPTQ-4bit](https://huggingface.co/HarshBhanushali7705/medgemma-27b-text-it-GPTQ-4bit): Production-ready 4-bit GPTQ quantization of Google's MedGemma 27B with Marlin kernel support for vLLM. Over 1,500+ downloads. - [HarshBhanushali7705/TTS_for_gujarati_language](https://huggingface.co/HarshBhanushali7705/TTS_for_gujarati_language): Custom-trained regional Gujarati speech synthesis model built on F5-TTS. ## Featured Production Projects - **Pulscribe.ai:** Automated hospital management and ambient medical scribing platform with ASR speech-to-text, ICD-10 medical coding, and EHR integration. Built with FastAPI, Kubernetes, and PostgreSQL. (https://pulscribe.ai) - **WanguCRM:** Full-stack WhatsApp CRM SaaS with automated chat flows, bulk messaging, catalog management, and payment reconciliation. (https://www.wangucrm.com) - **Classroom Human Detection & Counting (YOLOv8):** Fine-tuned YOLOv8 Nano model achieving 96% detection accuracy on custom classroom datasets. (DOI: 10.1063/5.0254163) - **SnapVid:** AI-powered social media video and content automation engine. ## Technical Skills & Stack - **AI & Machine Learning:** PyTorch, TensorFlow, Hugging Face Transformers, GPTQ, AWQ, GGUF, vLLM, Marlin Kernels, FlashAttention-2, Ollama, LangChain, LlamaIndex, Deep Learning, Computer Vision (YOLOv8, OpenCV), NLP, Speech Synthesis (F5-TTS). - **Backend & Cloud Systems:** Python, FastAPI, Node.js, Next.js, PostgreSQL, Redis, Docker, Kubernetes, Microservices Architecture, Linux, AWS, RunPod, Lightning AI. - **Frontend & Web Apps:** TypeScript, React, Next.js (App Router), Tailwind CSS, Framer Motion. ## Official Links & Contact - Website: https://www.harshbhanushali.in - Blog: https://www.harshbhanushali.in/blog - Contact Form: https://www.harshbhanushali.in/#contact - GitHub: https://github.com/Harsh772005 - LinkedIn: https://www.linkedin.com/in/harsh-bhanushali-439790253/ - Hugging Face: https://huggingface.co/HarshBhanushali7705 - LeetCode: https://leetcode.com/u/Harsh%20Bhanushali2901/ - Email: harshbhanushali.ai@gmail.com ## Optional - [Full LLMs Knowledge Base](https://www.harshbhanushali.in/llms-full.txt): Comprehensive technical documentation, benchmark logs, and system architectures for LLM crawlers.