Research
One line of work: the reliability of retrieval and LLM systems. Knowing when a generated answer can be trusted, and getting there without overspending compute. It runs from a co-authored production RAG method to sole-authored work on calibrated abstention, cross-model reliability, evaluation, and model selection. Select any paper for the summary and a link.
Publications
When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation
preprint
When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation
preprintAsks when conformal risk control can place a distribution-free guarantee on an LLM output, and proves when it cannot. Across 716 configurations spanning six open-weight models, eight datasets and six uncertainty scores, heuristic abstention rules violate their stated risk targets on 7.5% to 12.5% of evaluated settings, and under distribution shift neither static conformal risk control nor any tested adaptation rate held the target on 14 of 16 cross-dataset transfers. The paper adds an impossibility result, a strict hierarchy of bounds, and a closed-form feasibility test for whether a chosen risk target is even reachable.
Sole-authored.
Retrieval-Augmented Generation for Domain-Specific Question Answering
AAAI 2024 · SDU Workshop
Retrieval-Augmented Generation for Domain-Specific Question Answering
AAAI 2024 · SDU WorkshopA retrieval-augmented generation approach for closed, domain-specific question answering that uses user interaction signals to improve retrieval and reduce ungrounded answers. It underpins production question answering at Adobe and has been built on by independent academic and industry groups.
Co-author (one of eight). Cited 50+ times.
EVICT: Evidence-Sufficiency Verification via Counterfactual Dropout for Visually-Grounded Selective QA
CVPR · GRAIL-V
EVICT: Evidence-Sufficiency Verification via Counterfactual Dropout for Visually-Grounded Selective QA
CVPR · GRAIL-VVision-language models often answer confidently while relying on the wrong evidence. EVICT tests this directly: it masks the image region the model claims to depend on, then re-runs the same question. If the answer does not change, the model was not actually using the evidence it cited, and the answer is flagged as unverified.
The probe needs no training and no ground-truth labels, so it is cheap to run as a reliability guardrail on top of an existing model. Its honest limitation: it detects evidence-independence, not correctness. An answer can be genuinely grounded and still wrong.
Sole-authored. Published at the CVPR 2026 GRAIL-V workshop.
PASC: Pipeline-Aware Conformal Prediction for Multi-Stage NLP Pipelines
ICML · EIML
PASC: Pipeline-Aware Conformal Prediction for Multi-Stage NLP Pipelines
ICML · EIMLIn a multi-stage system (NER → disambiguation → typing) errors compound, so calibrating each stage alone under-covers while a Bonferroni union bound over-covers. PASC reduces joint coverage to a single conformal problem on the pipeline's maximum nonconformity score.
On a three-stage pipeline over CoNLL-2003 it reaches 96.4% end-to-end coverage versus 93.4% (Bonferroni) and 86.5% (independent calibration), at the same prediction-set size, and empirically holds target coverage under distribution shift, where independent calibration collapses to 59%.
Reimplemented from the paper in open source.
Sole-authored.
PromptPort: A Reliability Layer for Cross-Model Structured Extraction
preprint
PromptPort: A Reliability Layer for Cross-Model Structured Extraction
preprintFormalizes "format collapse," where one prompt yields clean JSON on one model and malformed output on another, and adds a canonicalization and verification layer so strict parsers stop rejecting correct extractions.
It repairs form, not meaning: it can rescue a malformed-but-correct extraction, but it will not catch a confidently wrong one.
Sole-authored.
Forecasting Model Success at Inference Time: Calibrated Probabilistic Forecasts for Cost-Optimal LLM Cascades
ICML · Forecast
Forecasting Model Success at Inference Time: Calibrated Probabilistic Forecasts for Cost-Optimal LLM Cascades
ICML · ForecastPredicts, per query, whether a smaller model will succeed, and uses that calibrated forecast to escalate only the queries it is likely to fail. Because the forecast is calibrated, the routing threshold maps directly to a chosen cost/quality operating point. On a 75,000-query production named-entity workload, calibration cut expected calibration error from 0.12 to 0.03, and the resulting cascade reached 0.91 micro-F1 at 31% lower cost than always using the large model.
Sole-authored.
Two Wrongs, No Right: Opposing Measurement Failures in LLM Annotators for Civic Discourse
ICML · AI4GOOD
Two Wrongs, No Right: Opposing Measurement Failures in LLM Annotators for Civic Discourse
ICML · AI4GOODWhen LLMs annotate contested social and political text they fail in opposite directions at once, one model over-flags where another under-flags, and they can underestimate how much opposition a population holds by 24 to 40 points. Worse, aggregate accuracy can look near-perfect through “accidental cancellation” while both directional errors stay large.
Sole-authored.
Not All Queries Need Rewriting: When Prompt-Only LLM Refinement Helps and Hurts Dense Retrieval
ICLR · CAO
Not All Queries Need Rewriting: When Prompt-Only LLM Refinement Helps and Hurts Dense Retrieval
ICLR · CAOShows that rewriting a query before retrieval is strongly domain-dependent, it helped on TREC-COVID but hurt on FiQA, because rewrites that swap out domain-specific terms degrade queries that already matched well.
Sole-authored.
The Generalization Gap in Named Entity Recognition: Static Benchmarks Overestimate the Transferable Performance of Neural Pattern Recognizers
ANNPR 2026 · main track
The Generalization Gap in Named Entity Recognition: Static Benchmarks Overestimate the Transferable Performance of Neural Pattern Recognizers
ANNPR 2026 · main trackStandard NER benchmarks reuse the same entities across train and test, so a high score can reflect memorization rather than generalization. This measures how far those scores overstate performance once entity types recombine into novel, unseen compositions, the conditions production systems actually face, and argues for evaluation that reflects them.
Accepted for oral presentation at ANNPR 2026 (Springer proceedings); also presented as a poster at the ICML 2026 CompLearn workshop.
Sole-authored.
Architecture-Homogeneous Model Selection for Representational Alignment
ICLR 2026 · Re4-Align Challenge
Architecture-Homogeneous Model Selection for Representational Alignment
ICLR 2026 · Re4-Align ChallengeWhen you compare two models by their internal representations, which models you pick can drive the result. This selects architecture-homogeneous models so an alignment score reflects the representations themselves, not architectural confounds.
Sole-authored.
The Modality Neglect Problem: Measuring Visual Reliance in Vision-Language Models
ACM Multimedia 2026 · Brave New IdeasThe Unverifiable Output Problem: Why Scaling Cannot Fix Self-Verification in Vision-Language Models
ACM Multimedia 2026 · Brave New IdeasUnified LLM-Orchestrated Data-Engineering Pipelines
IEEE GCWCN 2025Generating Answers to Contextual Queries within a Closed Domain
applicationThe two ACM Multimedia DOIs resolve when the proceedings publish.
@inproceedings{kotte2026modality,
title = {The Modality Neglect Problem: Measuring Visual Reliance
in Vision-Language Models},
author = {Kotte, Varun},
booktitle = {Proceedings of the 34th ACM International Conference on
Multimedia (Brave New Ideas)},
publisher = {Association for Computing Machinery},
year = {2026},
doi = {10.1145/3767308.3832561}
}
@inproceedings{kotte2026unverifiable,
title = {The Unverifiable Output Problem: Why Scaling Cannot Fix
Self-Verification in Vision-Language Models},
author = {Kotte, Varun},
booktitle = {Proceedings of the 34th ACM International Conference on
Multimedia (Brave New Ideas)},
publisher = {Association for Computing Machinery},
year = {2026},
doi = {10.1145/3767308.3832562}
}
@inproceedings{kotte2027generalization,
title = {The Generalization Gap in Named Entity Recognition: Static
Benchmarks Overestimate the Transferable Performance of
Neural Pattern Recognizers},
author = {Kotte, Varun},
booktitle = {Artificial Neural Networks in Pattern Recognition (ANNPR 2026)},
editor = {Dimitri, Giovanna Maria and others},
series = {Lecture Notes in Artificial Intelligence},
volume = {16978},
pages = {1--12},
publisher = {Springer Nature Switzerland},
year = {2027},
doi = {10.1007/978-3-032-39028-8_40}
}
@article{kotte2026crc,
title = {When Can Conformal Risk Control Certify LLM Outputs?
Bounds, Impossibility, and Adaptation},
author = {Kotte, Varun},
journal = {arXiv preprint arXiv:2606.29054},
year = {2026}
}
@article{kotte2026pasc,
title = {PASC: Pipeline-Aware Conformal Prediction with Joint Coverage
Guarantees for Multi-Stage NLP and LLM Pipelines},
author = {Kotte, Varun},
journal = {arXiv preprint arXiv:2605.18812},
year = {2026}
}
@article{sharma2024rag,
title = {Retrieval Augmented Generation for Domain-specific Question
Answering},
author = {Sharma, Sanat and Yoon, David Seunghyun and Dernoncourt, Franck
and Sultania, Dewang and Bagga, Karishma and Zhang, Mengjiao
and Bui, Trung and Kotte, Varun},
journal = {arXiv preprint arXiv:2404.14760},
year = {2024},
note = {AAAI 2024 Workshop on Scientific Document Understanding}
}