Naman Gupta · AI developer

I teach machines to readso people don't have to guess.

AI devloper & coder building ML models, RAG pipelines, and agentic systems. I ship research-grade code that makes intelligence useful.

Selected work

Every project below shipped as a full pipeline. Data in, decisions out. Architecture, trade-offs and honest numbers included.

01Mar 2026

Chat-With-Documents

A production-grade RAG pipeline for multilingual document intelligence.

PythonFlaskFAISSHugging Face
Challenge
Documents arrive in many formats and languages; naive retrieval returns noise, and multi-turn questions lose context without memory.
Solution
A custom preprocessing system (chunking, embedding, indexing across 2+ formats) feeding a FAISS semantic index, with a Flask backend that routes queries across Summarize and Q&A modes and carries conversation memory for coherent multi-turn dialogue. Made this without using LAngchain or other frameworks.

2+

document formats supported

Multi

lingual semantic Q&A

↓

average query latency

Built end-to-end: document loaders, a custom chunking strategy tuned for retrieval quality, Hugging Face transformer embeddings, FAISS similarity search, and conversation memory all served through a mode-routing Flask API.

02Dec 2025

Fake News Detection

Four deep architectures, one honest benchmark on 72K articles.

PythonTensorFlowNLTKGloVeScikit-Learn
Challenge
Headline accuracy numbers hide trade-offs which architecture actually minimises false positives on misinformation at scale?
Solution
Designed and benchmarked four deep classifiers with a full NLP preprocessing pipeline, then ran a systematic ablation across architectures and hyperparameters, documenting precision/recall trade-offs for each.

98%+

classification accuracy

4

architectures benchmarked

BiGRU

best F1, fewest false positives

A study in doing evaluation properly: identical preprocessing across models, controlled hyperparameter sweeps, and per-class error analysis. BiGRU won on F1 with the fewest false positives.

032026

Text-to-SQL Chatbot

Ask questions about a MySQL database in plain English.

PythonGeminiLangGraphLangChainMySQLGradio
Challenge
Database questions usually require SQL knowledge and manual schema inspection.
Solution
Built a Gradio interface backed by a cached LangGraph ReAct agent. The agent uses LangChain's SQL toolkit to inspect a configured MySQL schema, generate SQL, execute the query, and return a natural-language answer.

1

database engine supported

NL

database querying interface

Cached

agent initialization

A lightweight natural-language database interface with environment-based credentials and cached agent and database initialization. The current version is a working prototype focused on turning plain-English questions into database answers.

042026

Dense Encoder Fine-Tuning for WikiQA

A domain-specific sentence encoder that makes open-domain QA retrieval sharper.

PythonSentence TransformersHugging FacePandas
Challenge
Generic embedding models can miss the relationship between question wording and answer context, while duplicate documents create conflicting in-batch negatives during contrastive training.
Solution
Fine-tuned BAAI/bge-small-en-v1.5 on cleaned WikiQA question-document pairs with MultipleNegativesRankingLoss, deduplicating contexts before training to prevent representation collapse and using the modern SentenceTransformerTrainer API.

+10pp

Recall@5 improvement

0.89

post-fine-tune Recall@5

MNRL

contrastive training objective

An experiment in adapting a lightweight dense encoder for retrieval-heavy QA. The resulting model, nmngpt0/bge-qa-lilsmall01, turns the WikiQA pipeline into a reusable embedding model for comparing questions with relevant passages.

About

Developer, Learner and GPU burner. I build AI models, and agentic systems that make intelligence useful. I ship research-grade code that makes intelligence useful.

Naman Gupta

Naman Gupta

B.E. Computer Science, AI Specialization

University
Bennett University, Greater Noida
Graduating
Expected July 2027
Based in
Greater Noida, India
Status
Available
Beginning01 / 03

The Journey Starts

Passionate about building intelligent systems that solve real problems. Started exploring AI and ML to understand how machines can learn and adapt.

Learning02 / 03

Deep Dive into ML

Spent countless hours studying algorithms, neural networks, and deep learning frameworks. Built projects that combine research concepts with practical applications.

Building03 / 03

From Theory to Production

Now focused on shipping research-grade code. Working on RAG systems, agentic AI, and ML pipelines that make a real difference.

Capabilities

No progress bars, no five-star ratings. Select a domain. Every node tells you how it's actually been used.

Retrieval, generation and everything that makes language compute.

RAGProduction pipeline in Chat-With-Documents
LangChainOrchestration, memory, agent tooling
Hugging FaceTransformer embeddings & fine-tuning
Prompt EngineeringQuery routing across Summarize / Q&A modes
Agentic AITool-using agent experiments
NLTKTokenisation & preprocessing
spaCyNER & dependency parsing
Fine-tuningTask adaptation of pretrained models

● core · hover any node for context

Journey

Real work, real datasets, real stakeholders. The roles that shaped how I build.

Jun 2024 to Jul 2024

AI/ML Intern

Edu-net Foundation · Virtual Internship

Built data-quality tooling and studied the end-to-end ML lifecycle with IBM AI tools.

  • Built an automated data validation & auditing tool that scans uploaded datasets for missing values, duplicates, outliers and schema inconsistencies, generating structured quality reports with actionable fix recommendations via a Streamlit dashboard.
  • Reduced manual data-review effort and ensured accuracy before downstream analysis.
  • Documented model experiments and results to understand performance trade-offs across architectures.
  • Explored the full ML lifecycle, preprocessing, training, evaluation and deployment, using IBM AI tools.
PythonStreamlitPandasIBM AI Tools

next chapter, loading… could be with your team

Research

Credentials

Every credential links to its live verification page. Click any card to expand.

Contact

Open to AI/ML internships & research roles