— Index

Projects
& Work

15 entries Core ML · 4 Agentic AI · 5 Full-Stack · 4 Fine-Tuning · 2

Every entry below expands in place. Open one to read what it is, how it works, and why it was worth building.

15 / 15 shown

What It Is

A non-autoregressive text classifier that fine-tunes Qwen2.5-0.5B with LoRA to resolve 15+ task families in a single forward pass — trading generative decoding for fast, calibrated label prediction.

How It Works

A LoRA adapter is trained on top of the Qwen2.5-0.5B backbone. Instead of decoding tokens one at a time, the pooled representation feeds a unified classification head that emits task-family logits in one pass, then calibrates them for reliable confidence.

Key Highlights

  • Qwen2.5-0.5B backbone adapted with LoRA for parameter-efficient fine-tuning
  • Non-autoregressive single-pass inference across 15+ task families
  • 8.9x batch throughput speedup versus generative decoding baselines
  • Strong calibration with a measured 0.030 test ECE
  • Unified classification head replacing per-task bespoke pipelines

Why It Matters

Shows how to turn a generative LLM into a fast, calibrated classifier — the kind of efficiency work that makes models cheap enough to ship.

Stack

Python PyTorch HuggingFace LoRA

What It Is

A 57.5M-parameter masked diffusion language model that generates text by iteratively denoising masked tokens in parallel over 128 steps — an alternative to left-to-right autoregression.

How It Works

A diffusion transformer predicts masked tokens across a sequence, then remasks and refines over 128 denoising steps until the text resolves. Training runs on HuggingFace Accelerate, and a Flask demo exposes live generation.

Key Highlights

  • 57.5M-parameter transformer trained as a masked diffusion model
  • Parallel denoising across 128 diffusion steps instead of token-by-token decoding
  • Distributed training pipeline built on HuggingFace Accelerate
  • Interactive Flask demo for live text generation
  • Non-autoregressive sampling exploring diffusion for language

Why It Matters

Diffusion for language is frontier territory — building one from scratch proves the ability to work beyond the standard autoregressive playbook.

Stack

Python PyTorch Accelerate Flask

What It Is

Built a full large language model from the ground up — no shortcuts, no abstractions, just pure understanding. Every component from tokenizer to pretraining loop was implemented independently.

How It Works

Custom BPE tokenizer, data loader, multi-head self-attention with causal masking, full transformer architecture (LayerNorm, FFN, residuals, positional embeddings), and a complete pretraining loop — all without high-level abstractions.

Key Highlights

  • Custom BPE tokenizer built from scratch
  • Multi-head self-attention with causal masking
  • Full transformer block independently implemented
  • End-to-end pretraining, fine-tuning, and inference pipeline
  • Zero reliance on HuggingFace Trainer or pre-built blocks

Why It Matters

If you can build it from scratch, you truly understand it. This proves depth of knowledge at the architecture level.

Stack

Python PyTorch

What It Is

A production-grade modular chatbot with memory, RAG, streaming, tool-calling, and human-in-the-loop — all in one system. Feels less like a demo and more like a real product.

How It Works

LangGraph orchestrates the entire flow. SQLite handles persistent conversation threads. FAISS powers semantic document retrieval. MCP enables external tool-calling. Streamlit streams responses live. Human-in-the-loop checkpoints pause the agent for user approval before proceeding.

Key Highlights

  • Threaded conversations with SQLite persistence across sessions
  • Dual-layer memory: short-term context + long-term semantic store
  • FAISS RAG pipeline for document-grounded answers
  • Streaming responses via Streamlit for real-time UX
  • MCP tool-calling: invoke external APIs from within the agent loop
  • Human-in-the-loop (HITL) support

Why It Matters

Demonstrates the ability to architect a complete, production-style agentic system — not just a script, but a real modular application.

Stack

Python LangGraph FAISS MCP

What It Is

Applies compiler architecture to software generation. Give it a plain English description of an app — it runs through four deterministic stages, validates the output against a typed contract, and self-repairs any broken fields before returning the result. Nothing about it is one-shot.

How It Works

The pipeline is a StateGraph (LangGraph) with five registered nodes: intent → design → schema → validate → repair. Every node reads from and writes to a shared AppState TypedDict. add_conditional_edges off the validate node — if Pydantic's ValidationError fires, the exact error gets stored in state["errors"] and routed to the repair node. The repair → validate back-edge is what makes this a real loop, not a retry. A clean_json_output() strips markdown fences before json.loads(). FastAPI exposes POST /api/generate; React + Vite renders results across four tabbed JSON views.

Key Highlights

  • StateGraph with typed AppState TypedDict through all five nodes — no global state, no side effects between stages
  • AppSchema(BaseModel) enforces { ui, api, database, auth } — Pydantic catches field-level failures with exact error messages
  • add_conditional_edges routes to repair or END based on state["validated"] — graph controls flow, not app code
  • repair → validate back-edge: model sees its own broken output + exact ValidationError, not a generic retry
  • clean_json_output() handles the markdown fence problem that breaks json.loads() on raw LLM responses
  • ChatGroq at temperature 0.5 — consistent JSON structure without degenerate outputs

Why It Matters

Most LLM pipelines are glorified prompt chains — they call the model and hope the output is right. This one treats the model like an unreliable compiler pass: run it, validate against a contract, and feed failures back with surgical precision. The repair loop isn't a fallback — it's a first-class part of the architecture.

Stack

Python LangGraph FastAPI React (Vite) Groq Pydantic

What It Is

A three-phase autonomous AI engine — the cognitive core of a multi-bot social platform. It decides which bots respond to a post, generates original content on a schedule, and defends its persona against prompt injection attacks in live arguments.

How It Works

Phase 1 uses FAISS + cosine similarity to route posts to relevant bots. Phase 2 runs a LangGraph state machine that searches the web and drafts opinionated posts. Phase 3 reconstructs the full argument thread as RAG context and fires back in-character.

Key Highlights

  • FAISS persona router — cosine similarity decides bot engagement, not brute-force
  • LangGraph 3-node autonomous pipeline: decide → search → draft
  • Full-thread RAG engine for context-aware multi-turn replies
  • Prompt injection hardening — override attempts trigger intensified in-persona responses
  • Strict JSON output enforcement across all autonomous content

Why It Matters

This isn't a chatbot wrapper — it's a full agentic decision loop. It shows understanding of autonomous systems, vector-based routing, LangGraph orchestration, and production-level prompt security all in one project.

Stack

Python LangGraph RAG

What It Is

A comprehensive reference library of LangGraph workflows covering every major pattern for real-world agent design — basic, sequential, conditional, parallel, subgraph, persistence, RAG, and tool-use pipelines, plus a Streamlit frontend.

How It Works

Each pipeline is a standalone runnable LangGraph graph with typed state management. Conditional edges handle branching. Parallel nodes handle concurrent tasks. Subgraphs enable modular composition. The Streamlit frontend ties them into an accessible UI.

Key Highlights

  • 8+ pipeline patterns covering the full LangGraph design space
  • Fully typed state management across all graph flows
  • Subgraph composition for modular, reusable agent blocks
  • Streamlit frontend for visual interaction
  • A living reference for real-world agentic AI architecture

Why It Matters

Shows systematic mastery of LangGraph — not just one workflow, but the entire design space explored methodically.

Stack

Python LangGraph Jupyter Notebook

What It Is

Replicated Google's Gemma-3 small language model architecture from scratch in PyTorch to deeply understand its design decisions over earlier transformer models.

How It Works

Studied the Gemma-3 technical report and translated architectural specs into a clean PyTorch implementation — including Grouped-Query Attention (GQA), RoPE positional embeddings, GeGLU activations, and pre-normalization. Full pretraining loop included.

Key Highlights

  • Full Gemma-3 architecture reproduced from the technical paper
  • Grouped-Query Attention (GQA) implemented from scratch
  • RoPE (Rotary Positional Embeddings) for length generalization
  • GeGLU activations replacing standard FFN activations
  • Pre-normalization plus a from-base pretraining loop with loss monitoring

Why It Matters

Demonstrates the ability to read ML research papers and translate them into working implementations — a critical skill for frontier AI work.

Stack

Python PyTorch
Repository private Full case study

What It Is

An end-to-end pipeline for teaching a base language model to follow instructions using supervised fine-tuning — the exact process behind ChatGPT and Claude.

How It Works

Formats datasets into instruction-response pairs. Loss is computed only on response tokens (not the prompt). Training loop handles gradient accumulation, learning rate scheduling, and checkpoint saving.

Key Highlights

  • Instruction dataset formatting with proper prompt-completion separation
  • Loss masking — only response tokens contribute to gradient updates
  • Custom PyTorch training loop with LR scheduling and gradient accumulation
  • Evaluation metrics for instruction-following quality
  • End-to-end: raw dataset in, instruction-tuned model out

Why It Matters

Covers the exact pipeline used to convert raw LLMs into assistants — the same process behind the models people use every day.

Stack

Python PyTorch

What It Is

Fine-tuned a pretrained transformer for binary text classification using transfer learning — fully custom PyTorch implementation, no Trainer APIs.

How It Works

Pretrained transformer weights loaded, LM head replaced with a classification head, fine-tuned end-to-end. Custom PyTorch DataLoaders handle batching and tokenization. Cross-entropy loss with per-epoch accuracy tracking.

Key Highlights

  • Transfer learning: pretrained transformer fine-tuned for SMS spam detection
  • Custom classification head replacing the LM head
  • Full end-to-end PyTorch training — no Trainer abstractions
  • Train / validation / test split with custom data loading and batching
  • Per-epoch accuracy and loss evaluation

Why It Matters

Showcases the fundamentals of NLP fine-tuning — the same technique powering most production text classifiers today.

Stack

Python PyTorch

What It Is

A full-stack app that lets you ask natural language questions about any YouTube video. Point it at a video, ask a question, get a grounded answer — powered by transcript-based RAG.

How It Works

Fetches YouTube transcripts, chunks and embeds them using BAAI/bge-small-en into FAISS, then answers questions via a LangChain RAG chain backed by Groq (Llama-3). Next.js 16 + React 19 frontend.

Key Highlights

  • YouTube transcript ingestion and semantic chunking
  • BAAI/bge-small-en embeddings for high-quality semantic search
  • FAISS vector store for fast similarity retrieval
  • LangChain RAG chain with Groq LLM for sub-second answers
  • Next.js 16 + React 19 modern frontend
  • Full-stack: Python backend + TypeScript/React frontend

Why It Matters

A complete RAG product covering the entire pipeline — from ingestion to retrieval to generation to UI.

Stack

Python Next.js LangChain Groq

What It Is

Type "10 USD to INR" or "convert 250 Swiss francs to yen" — the LLM parses your intent, fetches live rates, and returns the result. Natural language replaces dropdowns entirely.

How It Works

LangChain routes user query to Groq (Llama-3.3-70b) which extracts structured intent (source currency, target currency, amount). That output queries ExchangeRate-API for live rates. Result returned to React frontend via FastAPI.

Key Highlights

  • Natural language interface — no dropdowns, just plain English
  • Groq Llama-3.3-70b for fast, accurate intent extraction
  • LangChain orchestration between LLM and API calls
  • Live exchange rates from ExchangeRate-API
  • FastAPI backend with clean REST endpoints
  • React + Vite frontend for a snappy UI

Why It Matters

Demonstrates LLM tool-use in a real product context — turning unstructured language into structured API calls.

Stack

Python FastAPI Groq

What It Is

A real-time collaboration tool where multiple users can work together with instant synchronization — Mohit's full-stack web project outside AI, showing range as a developer.

How It Works

Next.js + TypeScript frontend. Dedicated Node.js server handles live sync logic using real-time event broadcasting. State changes from any connected client propagate to all others instantly.

Key Highlights

  • Real-time multi-user synchronization
  • Dedicated Node.js server for live event handling
  • Next.js + TypeScript frontend for type-safe, fast UI
  • Event-driven architecture for instant state propagation
  • Clean separation: frontend logic vs sync server

Why It Matters

Shows versatility — Mohit isn't just an AI developer, he can build production full-stack web systems too.

Stack

TypeScript Next.js

What It Is

A reference implementation of ten distinct Retrieval-Augmented Generation patterns — from naive RAG to hybrid, re-ranking and self-query — each modeled as a LangGraph workflow over a shared ChromaDB store.

How It Works

Every pattern is an explicit LangGraph state machine reading from one ChromaDB vector store. A FastAPI backend serves pattern selection and query execution; a React front end lets you compare pattern behavior side by side.

Key Highlights

  • Ten RAG patterns — Agentic, Self, Adaptive, Corrective, Hybrid-KG and more — as discrete workflows
  • Each pattern authored as an explicit LangGraph state machine
  • ChromaDB vector store shared across every retrieval strategy
  • FastAPI backend serving pattern selection and query execution
  • React interface for comparing pattern behavior side by side

Why It Matters

RAG is rarely one algorithm — it's a design space. Implementing ten patterns as swappable graphs shows systematic command of retrieval architecture, not just a single recipe.

Stack

Python LangGraph ChromaDB FastAPI React

What It Is

A Keras neural network that predicts customer churn at ~85% accuracy over 10K records, exported to ONNX for portable inference and shipped as a Dockerized Streamlit application.

How It Works

Records are encoded and scaled, then fed to a feed-forward Keras classifier. The trained model is exported to ONNX for framework-agnostic inference and served through a Streamlit app packaged in Docker for one-command deployment.

Key Highlights

  • Feed-forward Keras classifier reaching ~85% accuracy on 10K customer records
  • Full preprocessing pipeline with categorical encoding and feature scaling
  • ONNX export for framework-agnostic, portable inference
  • Dockerized Streamlit app for one-command deployment
  • Live probability scoring for individual customer inputs

Why It Matters

Covers the full applied-ML loop — train, export, containerize, deploy — the difference between a notebook and a product someone can actually use.

Stack

Python Keras ONNX Streamlit Docker