Mehul Sharma
Reading

Portfolio · 2026 edition Open to AI · Backend · Systems roles

Mehul
Sharma.

AI & Backend EngineerAI agents · LLM systems · secure APIs

I build the backend that makes AI dependable: retrieval that cites its sources, models you can monitor after launch, and APIs that keep sensitive data in the right hands.

NowAI Engineer at BosonQ Psi (BQP), architecting an AI agent that turns plain-language engineering requests into simulation and optimization runs.

Core stack
Python, FastAPI, PostgreSQL, LangChain
Focus
RAG, evaluation, MLOps, secure backends
Education
Computing & Software Systems, University of Melbourne
Worked in
India and Australia

You're reading the two-minute version. Switch to In depth above for full case studies and career detail.

You're reading the full version. Switch to At a glance for the two-minute summary; you'll stay on the same section.

Selected work

Four systems built to be trusted with real data.

Two from my jobs, two I built to learn production AI properly. Each one is about the same question: how do you know the system is doing the right thing?

  1. 01Multi-tenant RAG platformPersonal build · LLM systems
  2. 02ML platform with drift detectionPersonal build · MLOps
  3. 03Role-aware employee assistantDrish Infotech · Enterprise AI
  4. 04Clinical test backendCurtin Health · Backend

Case 01Personal buildRetrieval & LLM systems

A multi-tenant RAG platform that has to show its sources

Answers stream back with citations, fit a strict context budget, and pass regression-tracked evaluation before anything changes.

PythonFastAPIPostgreSQL + pgvectorRedisLangChainDockerPrometheus

  • Hybrid retrieval: BM25 keyword search plus vector search, reranked by a cross-encoder.
  • Citation-enforced, streaming answers with strict context budgeting.
  • Evaluation for faithfulness, relevance and completeness, tracked for regressions.
  • Structured logging, tracing and Prometheus metrics from day one.
  1. 01IngestAsync document ingestion per tenant
  2. 02Indexpgvector embeddings + BM25
  3. 03RetrieveHybrid keyword + vector search
  4. 04RerankCross-encoder orders candidates
  5. 05AnswerStreamed, cited, within budget
  6. 06EvaluateFaithfulness · relevance · completeness
Request path, left to right, top to bottom.

The problem

RAG chat is easy to demo and hard to trust. When several customers share one system, three things break first: data leaking between tenants, answers that sound right but aren't grounded, and quality drifting quietly each time a prompt or index changes.

What I built

A FastAPI backend with async ingestion into PostgreSQL and pgvector, retrieval scoped to each tenant, and an LLM layer that streams answers and ties every claim to a retrieved passage. An evaluation harness and observability stack sit around it so changes can be measured.

Where it stands

The retrieval, generation, evaluation and observability layers all run end to end. The evaluation suite is the part I'd point a reviewer to first: it's what turns "seems better" into a number.

Key decisions

  1. Hybrid over pure vectorsKeyword search catches exact terms such as IDs and names that embeddings blur. The cross-encoder then decides the final order.
  2. Citations are enforced, not requestedResponses must reference retrieved passages, so a reader can check every claim against its source.
  3. A hard context budgetRetrieved text is fitted to a fixed token budget, which keeps cost and latency predictable as documents grow.
  4. Evaluation as a regression testQuality scores are tracked across changes, so a change that makes answers worse shows up before it ships.

Case 02Personal buildMLOps

An ML platform that notices when its model goes stale

Resume–job matching run as a production system: trained, versioned, shadow-deployed, watched for drift and retrained on feedback.

PythonFastAPIscikit-learnXGBoostSentence TransformersMLflowPostgreSQLPrometheus

  • MLflow experiment tracking and model versioning for reproducible training and safe promotion.
  • FastAPI inference with schema validation, versioned model loading and shadow deployments.
  • Data and prediction drift exposed to Prometheus, feeding retraining.
  • Models evaluated on ROC-AUC and Precision@K.

Where it startedAI Resume Sorter, my earlier TF-IDF and cosine-similarity matcher: 85% matching accuracy, 100+ resumes scored in under 2 seconds, served as a Dockerized FastAPI service.

  1. TrainStructured signals + text embeddings
  2. VersionEvery run tracked in MLflow
  3. ShadowNew model scores live traffic silently
  4. PromoteVersioned load, schema-checked inputs
  5. MonitorData + prediction drift in Prometheus
  6. Retrain ↺Feedback returns the loop to step 01

The problem

Most matching models are trained once and left alone. The data they see changes, their accuracy falls, and nobody notices until the results stop making sense.

What I built

Feature pipelines that combine structured candidate signals with sentence-transformer embeddings; scikit-learn and XGBoost models tracked in MLflow; and a FastAPI inference service with the safety rails a production model needs.

Where it stands

The full loop runs end to end, from training through controlled redeployment. It's the project where I learned that shipping a model is the start of the work, not the end.

Key decisions

  1. Precision@K next to ROC-AUCPeople read the top of a ranked list, so quality at the top matters more than average separation.
  2. Shadow before promoteA candidate model runs beside the current one on real requests before it's allowed to answer them.
  3. Drift as a metric, not a reportDrift signals go to Prometheus like any other service metric, so they can drive alerts and retraining.

Case 03Drish InfotechAI EngineerEnterprise AI

An internal assistant that knows who's asking

An employee AI assistant over company policies and SOPs that shapes each answer to the person's role, department and permissions.

PythonLangChainRAGAgentsModel Context Protocol

  • Built with RAG, LangChain and an agent-based architecture over internal knowledge.
  • Integrated Model Context Protocol (MCP) to pass secure user context: role, department, permissions.
  • Designed ingestion and retrieval pipelines for internal policies and SOPs.
  • Improved response accuracy and reduced manual support effort.
// context attached to a request (illustrative)
{
  "user": {
    "role": "team_lead",
    "department": "operations"
  },
  "permissions": [
    "policies.hr.read",
    "sop.operations.read"
  ],
  "question": "How do I approve overtime?"
}
Illustrative shape only. No production data.

The problem

Employees ask the same policy questions again and again, and the right answer often depends on who is asking. A generic chatbot either over-shares or stays too vague to act on.

What I did

Built the assistant's retrieval and agent layers with LangChain, designed the ingestion pipeline for policy documents and SOPs, and integrated MCP so the model gets a controlled description of the user with every request.

Result

More accurate answers for internal questions and less manual support effort for the teams that used to field them.

Key decisions

  1. Identity through a protocol, not the promptPassing role and permissions through MCP keeps personalization inside defined limits instead of trusting free-text instructions.
  2. Personalize the retrieval, not only the wordingUser context shapes which knowledge is in scope, so answers change in substance, not just tone.

Case 04Curtin HealthMelbourneBackend

A clinical backend for nurses, doctors and patients

One system for health test requests across multiple facilities, built on a normalized data model and strict access control.

REST APIsPostgreSQLJWTRole-based access control

500+health test requests handled daily

15+clinical test workflows integrated

20+tables in a normalized PostgreSQL schema

40%less data redundancy

The problem

Each clinical test had its own data shape, and three very different groups needed access to it: nurses, doctors and patients, across more than one facility.

What I did

Standardized 15+ test workflows onto shared data models, designed a normalized PostgreSQL schema of 20+ tables, and built REST APIs secured with JWT authentication and role-based access control.

Result

A backend handling 500+ test requests a day across facilities, with 40% less redundant data than before.

Key decisions

  1. Standardize the data before the endpointsShared models across test types are what removed the duplication.
  2. Access decided by role, in the APINurses, doctors and patients each see only what their role allows, enforced on every request.

Also built

  • Prompt-to-print apparel platform

    A solo product that turned natural-language prompts into custom T-shirt designs using text-to-image diffusion models. I tested it with real customers, then wound it down.

    Python · FastAPI · diffusion models · 2025
  • AI Resume Sorter

    A resume–job matching microservice using TF-IDF and cosine similarity: 85% accuracy, 100+ resumes in under 2 seconds, documented with OpenAPI.

    Python · TensorFlow · FastAPI · PostgreSQL · Docker

Career story

From schemas to systems to agents.

I started on the data layer, then owned a product end to end, then moved into enterprise AI. Now I'm building an AI agent for engineering simulation.

  1. Backend & data
  2. Whole product
  3. Enterprise AI
  4. Applied AI R&D
  1. 2026 – PresentApplied AI R&D

    AI EngineerNow

    BosonQ Psi (BQP) · Quantum-powered engineering simulation

    Own the architecture of an AI agent that takes engineering requests in plain language, drives the team's surrogate models and optimizers, and reports results back.

    • Designing the agent end to end: interpreting a request, choosing the right surrogate model or optimizer, running it and reporting progress to the user.
    • Building the proof of concept I will present to the company, taking a design request through to a completed optimization run.
    • Working where LLMs meet simulation and optimization, in a company whose core product is simulation software.
  2. Dec 2025 – PresentEnterprise AI

    AI Engineer

    Drish Infotech

    Built a role-aware employee assistant with RAG, agents and MCP.

    • Built an internal AI assistant using RAG, LangChain and agent-based architectures for personalized, role-aware answers.
    • Integrated Model Context Protocol to pass secure user context (role, department, permissions).
    • Designed ingestion and retrieval pipelines for policies and SOPs, improving accuracy and cutting manual support effort.
    Read case study →
  3. Feb 2025 – Sep 2025Whole product

    Founder & Solo Engineer

    Independent AI product · Melbourne

    Built and tested a prompt-to-T-shirt design platform on diffusion models, end to end.

    • Built the whole platform: users describe a design in plain language and get a printable T-shirt design from text-to-image diffusion models.
    • Designed the Python and FastAPI backend for prompt processing, image-generation orchestration and asset storage.
    • Ran product experiments with real customers to test demand and technical feasibility, then chose to wind it down.
  4. Jun 2024 – Jan 2025Backend & data

    Backend Engineer

    Curtin Health · Melbourne

    Backend for 500+ daily health test requests across multiple facilities.

    • Built a backend serving nurses, doctors and patients across multiple facilities, handling 500+ test requests a day.
    • Integrated 15+ clinical test workflows with standardized data models and a normalized PostgreSQL schema (20+ tables), cutting redundancy by 40%.
    • Built secure REST APIs with JWT authentication and role-based access control for patient data.
    Read case study →

Education

  1. 2022 – 2025

    Computing & Software Systems

    University of Melbourne · Australia

    University College resident (2022–2023). Played for the cricket and football teams.

  2. 2019 – 2021

    International Baccalaureate

    National Public School · Singapore

    Football team and quiz team.

Capabilities

Problems I can take off your plate.

Grouped by the problem, not the tool. Each one links to the work that shows it.

  • Make an LLM answer from your data, and show its sources

    Hybrid retrieval, reranking and citation rules that keep answers grounded in documents you control.

    • RAG
    • BM25 + vectors
    • Cross-encoders
    • pgvector
    • LangChain
    • Streaming

    ProofRAG platformDrish assistant

  • Tell whether an AI system is actually getting better

    Evaluation suites and ranking metrics that catch regressions before users do.

    • Faithfulness
    • Relevance
    • Completeness
    • ROC-AUC
    • Precision@K
    • Regression tracking

    ProofRAG platformML platform

  • Keep a model healthy after launch

    Versioned models, shadow deployments and drift signals wired into the same monitoring as the rest of the service.

    • MLflow
    • Shadow deploys
    • Drift detection
    • Prometheus
    • Tracing

    ProofML platformRAG platform

  • Build APIs that handle sensitive data safely

    Normalized schemas, authentication and role-based access for data like patient records.

    • FastAPI
    • REST
    • PostgreSQL design
    • JWT
    • RBAC
    • OpenAPI

    ProofCurtin Health

  • Personalize AI without losing control

    User identity and permissions passed to the model through a protocol, so personalization stays inside defined limits.

    • Model Context Protocol
    • Agents
    • Role-aware context

    ProofDrish assistant

  • Ship it and keep it running

    Containerized services, CI pipelines and cloud deploys, with tests and structured logs.

    • Docker
    • GitHub Actions
    • AWS EC2 / S3
    • Unit testing
    • Microservices

    ProofRAG platformML platform

Toolkit

Languages
Python, JavaScript, TypeScript, SQL, Java
Backend
FastAPI, Node.js, Express
AI & ML
LangChain, TensorFlow, scikit-learn, XGBoost, Sentence Transformers, MLflow
Data
PostgreSQL (pgvector), Redis, MongoDB, SQLite
Infrastructure
Docker, GitHub Actions, AWS (EC2, S3), Prometheus, Vercel
Frontend
React, Next.js, Tailwind CSS
Practice
Microservices, design patterns (GoF, GRASP), Agile/Scrum

About me

How I work.

I'd rather ship a small system I can measure than a big one I have to take on faith.

I finished school in Singapore, studied computing in Melbourne, and now work from India. Each move meant learning a new context quickly, and that's still how I start any project: understand the domain before touching the code.

I learn by building. When I pick up something new, like engineering simulation at BQP, I make a small working version first and read the theory against it. That habit is why both of my platform projects begin with evaluation: I decide how I'll measure a system before I start tuning it.

I've played cricket, football and quiz on teams for most of my life, and I work on engineering teams the same way. I learn my position, cover for others when it counts, and speak up early when something looks off.

Contact

Let's build something dependable.

I'm open to AI engineering, backend and systems roles. Email is the fastest way to reach me, and I'm happy to walk through any of these projects in detail.

AI · Backend · Systems