I build retrieval and LLM systems that stay reliable in production.
I own the serving and reliability layers of the internal LLM platform behind AI and search in six Adobe products, and I publish research on when a model's output can be trusted.
6Adobe products on the platform
8sole-authored papers
~7×Lightroom search click-through
Highlights
I own the serving and reliability layers of the internal LLM platform behind AI and search in six Adobe products.
Replaced five per-team LLM integrations with one shared service. New-model integration went from weeks to days.
Built AutoPrompt, which turns a team's own evaluation set into a working prompt baseline.
Shipped Lightroom semantic search. Search click-through rose roughly seven-fold.
First-named inventor of six on a pending Adobe patent, cited as prior art by eight USPTO examiners.
Three of my methods have been reimplemented in open source, among them CERT-FLOW.
Sole-authored papers at ACM Multimedia (Brave New Ideas), ANNPR (Springer LNAI), and CVPR, ICLR and ICML workshops.
Best Paper Award, INCECT 2026, for work on when speculative decoding backfires.
Fellow of BCS, The Chartered Institute for IT.
Program committee for ICML, ICCS and VAND at CVPR. Reviewer for NeurIPS, ICLR and ACL.
Selected work
LLM-serving & retrieval platform (ILUP)
The shared serving and retrieval layer product teams build on instead of rolling their own. It backs AI and search across six Adobe products; I own its serving and reliability layers.
When Can Conformal Risk Control Certify LLM Outputs?
preprint
Sole-authored. When a distribution-free guarantee on an LLM output is reachable, and a proof of when it is not; no tested adaptation rule restored the emitted-risk target on 14 of 16 cross-dataset transfers. Details →
The Generalization Gap in Named Entity Recognition
ANNPR 2026 · main track
Sole-authored. Taggers scoring 89-92% F1 on CoNLL-2003 collapse to 39-43% on novel entity compositions, a ~50-point gap the standard split never reveals. Details →
Retrieval-Augmented Generation for Domain-Specific Question Answering
AAAI 2024 · SDU Workshop
Adobe's production RAG method. Co-author. Cited more than 50 times.
EVICT: evidence-sufficiency verification for visually-grounded QA
CVPR · GRAIL-V
Sole-authored. A training-free probe that catches answers not grounded in the evidence.
PASC: pipeline-aware conformal prediction for multi-stage NLP
ICML · EIML
Sole-authored. Distribution-free coverage guarantees for a whole NLP pipeline, not each stage in isolation.