AI researcher and engineer. I measure whether a model’s explanation is actually true.

Creator ofGlassbox AIOpen to opportunities
Scroll
About

A model’s explanation for its own output is often made up after the fact.

I reverse-engineer transformers with causal circuit analysis (activation patching, ablation, sufficiency and comprehensiveness scoring against ERASER metrics) to test whether a stated explanation is causally load-bearing, not just fluent.

Experience

Machine Learning Engineer

BaseMotion AI, Berlin

Nov 2025 – Apr 2026

Applied mechanistic circuit analysis (activation patching, ablation) to GPT-2 Small, scoring explanations against ERASER sufficiency/comprehensiveness metrics. Built the pipeline in PyTorch and TransformerLens with an automated test suite and CI for reproducible runs. Result: r = 0.009 between model confidence and explanation faithfulness: confidence carries no signal. Separately measured that fluent LLM-written explanations score 99% on perceived quality (ERASER), up from 60%, with no corresponding gain in mechanistic faithfulness.

Independent Researcher, Project Thesis & AI Safety

Hochschule Trier

Apr 2024 – Sep 2025

Synthesized 25+ papers across AI alignment, scalable oversight, and interpretability into structured technical reports, benchmarking robustness methodologies and cataloguing frontier LLM failure modes to establish baseline quantitative metrics for the field.

AI Engineer

ICEICO Technologies, Nagpur

Feb 2021 – Mar 2023

Built PyTorch training pipelines and tuned hyperparameters for internal prediction models, establishing reproducible baselines for testing. Shipped a Retrieval-Augmented Generation pipeline with LangChain to automate knowledge retrieval across production datasets.

Projects

Glassbox AI

Feb 2026 – Mar 2026

An open-source mechanistic interpretability toolkit that finds the causal circuit behind a transformer prediction and turns it into EU AI Act Annex IV documentation automatically. 1,000+ PyPI downloads a month, live HuggingFace demo, CI passing on 9 architecture families.

View on GitHub →

Azure ML Pipeline

Feb 2026 – Jun 2026

End-to-end MLOps pipeline on Azure ML SDK v2 for heart disease prediction: a four-stage pipeline (data prep, training, evaluation, registration) running on auto-scaling compute clusters, MLflow tracking every run, and automatic promotion to a managed REST endpoint once a model clears the accuracy threshold.

View on GitHub →

Azure AI RAG System

Jan 2026 – Jun 2026

Production RAG pipeline on Azure OpenAI and Azure AI Search: documents chunked at 1,000 tokens with 200-token overlap, embedded at 1536 dimensions, retrieved with hybrid vector plus BM25 search, and answered by GPT-4o-mini constrained to the retrieved chunks only. Every response carries a citation back to its source document.

View on GitHub →

Causally-Grounded Mechanistic Interpretability

M.Sc. Thesis · Dec 2025 – May 2026

Evaluates a model circuit the same way you'd evaluate an explanation: does removing it break the output (sufficiency), and does keeping everything else still break it (comprehensiveness)? On GPT-2's IOI task the discovered circuit beat an attention-only baseline on 75% of prompts, scored 100% sufficiency, but only 22% comprehensiveness. Sufficiency alone oversells how reliable an explanation is.

View on GitHub →

Enterprise Agentic AI Platform

May 2025 – Jun 2025

A multi-agent system on LangGraph coordinating three MCP servers (Postgres, documents, notifications) with hybrid RAG, built to take a rough MVP into something enterprise-shaped. FastAPI streamed data over WebSockets to a Next.js frontend, Celery handled async work, Langfuse and Prometheus/Grafana gave it real observability, all containerized with Docker Compose.

View on GitHub →
Current Work

Unpublished. Links go live on release.

Explanation Multiplicity and the Instability of Conformity Claims

Target arXiv · Sep 2026

Circuits recovered by mechanistic interpretability tools are not unique: the same model, task, and method can return different circuits depending on ablation operator, corruption distribution, and metric choice. This measures how far an EU AI Act Annex IV conformity claim built from a discovered circuit moves across the space of defensible analytic specifications, and tests whether that movement is distinguishable from a size-matched random circuit baseline.

In progress

Causal Attribution for Agentic Decisions

Target arXiv · Dec 2026

Multi-agent systems produce decisions that no single deployer can fully attribute: observability tooling logs what happened but not what caused it. This tests whether observability-derived attribution actually tracks causally-measured attribution across real multi-agent traces, and whether it's systematically biased toward recency and verbosity.

In progress

What Monitorability Metrics Miss

Target ICLR 2027 · Sep 2026

The published chain-of-thought monitorability autorater was validated against synthetic degradations, a broken reasoning trace. That checks whether the metric notices a chain of thought falling apart, not whether it notices a well-formed one that conceals its true cause, which is the actual safety concern. This runs the validation the metric's own authors flagged as outstanding.

In progress
Finding

Confidence doesn’t predict faithfulness.

Model confidence predicts explanation faithfulness almost not at all: r = 0.009, measured with causal circuit analysis on GPT-2 Small (Indirect Object Identification task). The 6-head IOI circuit explains 61.4% of the logit difference and is sufficient on its own, but removing it still leaves 78% of the signal intact through backup heads: it isn’t necessary.

Tool

Glassbox AI: audits that write their own paperwork.

Glassbox AI finds the causal circuit behind a transformer prediction with activation patching, scores it for sufficiency and comprehensiveness, and writes EU AI Act Annex IV documentation as structured, content-hashed JSON, about 1.8s per analysis on GPT-2 Small (Apple M2 Pro), 15–37× faster than ACDC depending on circuit size. 932 tests passing in CI across 9 architecture families, GPT-2 through Llama-3.

Find the causal circuit once;
ship EU AI Act documentation with every audit.

Skills

Programming

PythonSQL

ML & Research

PyTorchTransformerLensHugging Facescikit-learnNumPyPandasXGBoostLightGBMSHAP

LLM & Agents

LLM EvaluationRAG (vector + BM25 hybrid)LangGraphModel Context Protocol (MCP)QdrantLlamaIndexPrompt Engineering

Cloud & MLOps

Azure ML (SDK v2)Azure OpenAIAzure AI SearchMLflowFastAPIREST APIsDockerDocker ComposeGitGitHub ActionsCI/CDStreamlit

Focus

Mechanistic InterpretabilityCausal Circuit AnalysisModel Evaluation & ObservabilityReproducible ML PipelinesEU AI Act Annex IV DocumentationPython Packaging (PyPI)

Languages

English (C1)German (B1)Hindi (Native)French (Beginner)
Proof

Glassbox AI ships with 932 tests passing across 9 architecture families and 1,000 PyPI downloads a month. Every number here is checked against the repo, not rounded up.

About

M.Sc candidate in AI/ML (Hochschule Trier, defending 2026), researching mechanistic interpretability and LLM faithfulness evaluation. I build evaluation tooling that turns interpretability findings into audit-ready evidence, not slide-deck claims.

Trier, Germany. Open to full-time ML engineer / AI safety roles and PhD positions, available from June 2026, open to relocate, no sponsorship required.

Education

M.Sc. Interdisciplinary Engineering — AI & ML

Hochschule Trier · 2023–2026 · Grade 1.8, Thesis 1.0

B.Tech, Automotive Engineering

Government College of Engineering and Research · 2018–2021 · Grade 1.8

Certifications
  • Microsoft Certified: Azure AI Engineer Associate, Azure AI Fundamentals
  • BlueDot Impact, Technical AI Safety fellowship
  • Google, AI Essentials, UX Research

Let’s talk.

Availability

If you’re hiring for interpretability, AI safety, or eval work, let’s talk.

Resume →
General inquiries

Press, collaboration, or curiosity. All welcome.

Say hello →