✦ AI / ML Engineer

Anagha P.
Krishna

I turn machine learning into real products.
From data and modeling to evaluation and deployment, end to end.

MS in Artificial Intelligence · Boston University

AI & MLSpecialization
MS @ BUArtificial Intelligence
Anagha P. Krishna

About

I'm a Machine Learning Engineer and MS student in Artificial Intelligence at Boston University, with a Bachelor's in Machine Learning. As a Data Analyst at EVERSANA I worked hands-on with large-scale pharmaceutical data, which sparked a lasting interest in applying ML to healthcare and medicine. I build ML systems end-to-end, from data and feature engineering through modeling, evaluation, and the deployed API, across classical ML, deep learning, and LLM systems.

What sets my work apart is a bias toward rigor and internals: I implement algorithms from scratch to understand how they work (PPO in pure NumPy), I build the evaluation and monitoring layers most demos skip (drift detection, LLM-as-a-judge, statistical significance testing), and I hold my systems to measurable standards: grounded answers, verified citations, honest refusal, and reproducible metrics. I'd rather ship one model I can trust and explain than ten I can't.

Selected Work

01
RAG · Clinical AI

Clinical RAG Assistant

A clinical RAG that won't lie to a doctor, and proves it.

Local-first RAG over clinical notes with verifiable citations, honest refusal, and grounded extraction.

RAGQdrantBGE embeddingsOllama · Llama 3.1FastAPIStreamlitMLflow
View project
02
Agentic AI · MLOps

Autonomous ML Pipeline with AI Agents

LangGraph agents that run an entire ML workflow autonomously.

Specialized AI agents that explore, model, evaluate, and explain, coordinated by an LLM super-agent.

LangGraphAgentic AIXGBoostSHAPOllamaScikit-learn
View project
03
LLM Evaluation · NLP

LLM Evaluation & Hallucination Analysis

An LLM-as-a-Judge framework that measures faithfulness and catches hallucinations.

A reusable evaluation suite scoring grounding and classifying hallucination type per response.

LLM-as-a-JudgeNLPLangChainOllamaEvaluation
View project
04
NLP Research · LLM Evaluation

Cross-Lingual Hallucination Drift in LLMs

Does an LLM hallucinate more in some languages, and does it depend on the task?

An NLP research study measuring how hallucination rate shifts across languages and task types, with a confound-controlled design and LLM-as-a-judge validation.

NLP ResearchLLM-as-a-JudgeAya Expanse 8BGPT-4o-miniHuggingFaceSciPyStreamlit
View project
05
MLOps · Monitoring

Model Monitoring & Drift Detection

Five drift scenarios versus five statistical detectors.

An MLOps toolkit mapping detection methods to drift types and quantifying performance impact.

MLOpsDrift DetectionScikit-learnStatisticsRandomForest
View project
06
Reinforcement Learning

PPO Reinforcement Learning from Scratch

PPO implemented in pure NumPy, put to work tuning a real classifier.

A from-scratch PPO agent that learns the optimal decision threshold for an imbalanced classifier: RL internals, not library glue.

Reinforcement LearningPPONumPyPolicy GradientsFrom scratch
View project

Skills

LLM / AI

RAGAgentic AILangChainLangGraphTransformersPrompt EngineeringLLM-as-a-JudgeSHAP

Machine Learning

PyTorchTensorFlowXGBoostScikit-learnCNN / RNNReinforcement Learning

Engineering & MLOps

PythonC++FastAPIDockerAWSCI/CDMLflow

Data

SQLMongoDBQdrantFAISSPandasTableau