About
I'm a machine learning engineer working on computer vision, multimodal models and large language models. I build systems that understand images and video, and language models that answer from evidence rather than from memory.
I care most about reliability: whether a model still works outside its benchmark, and whether a person can check what it says. That shapes how I build models and how I evaluate them.
I'm open to research collaborations in these areas. If you're working on something related, I'd be glad to hear about it.
Research interests
- Computer vision and video understanding
- Vision-language and multimodal models
- Large language models and grounded generation
- Reliable, verifiable AI
Across all of these, I'm drawn to one question: do models still behave reliably once they leave the benchmark?
News
-
2026
Paper
Submitted Joint Prediction of UHPC Properties Using a Feature-Tokenizer Transformer to Applied AI Letters. It's under review.
-
Jan 2025
Career
Joined Accelx Inc. as a Machine Learning Engineer.
-
Jan 2021
Thesis
Graduated from AUST with a thesis on arrhythmia classification using 2-D CNNs.
Selected projects
All projects →Overcooked AAR
I had one Video LLM review real human–human Overcooked play twice, from the event log and from the gameplay video, and checked all 1,037 of its claims against the log.
DocuVision
I fine-tuned LayoutLMv3-large and built a learned question→answer linker that turns scanned forms into structured key–value pairs, then released the model on Hugging Face.
Speech Translator
I built a translator that speaks your words in Spanish, German or Japanese in your own cloned voice, and chose every model by testing it on FLEURS.
Jarvis
I built a voice assistant that runs fully on my own hardware and uses sandboxed tools, then tested its tool calls on a held-out set written after tuning.
Let's collaborate
I'm open to research collaborations, especially on vision, video, or language models. A short email about what you're working on is the best way to start.