About
I'm a machine learning researcher working on computer vision, video understanding, and large language models. I build systems that see, understand, and explain what is happening in images, videos, and text.
I care most about reliability: making sure models keep working in the real world, not just on benchmarks. My work covers real-time video analysis, multimodal learning, and language models that give grounded, trustworthy answers.
I'm always happy to collaborate on research in these areas.
Research interests
- Computer Vision
- Real-Time Video Understanding
- Large Language Models
- Generative AI
- Multimodal Learning
- Human-Centered Computing
Across all of these, I'm drawn to one question: do models still behave reliably once they leave the benchmark?
News
-
2026
Paper
Submitted Joint Prediction of UHPC Properties Using a Feature-Tokenizer Transformer to Applied AI Letters. It's under review.
-
Jan 2025
Career
Joined Accelx Inc. as a Machine Learning Engineer.
-
2023
Course
Completed Machine Learning on Coursera.
-
Jan 2021
Thesis
Graduated from AUST with a thesis on arrhythmia classification using 2-D CNNs.
Selected projects
All projects →Overcooked AAR
After-action reviews of real human–human Overcooked play, written by an LLM from event logs and by a VLM from gameplay video, then blind-rated to find where the two disagree.
Sentinel
Multi-camera, real-time anomaly detection. YOLOv8 + BoT-SORT tracking feeds a rule engine for falls, loitering, abandoned objects, crowding, and wrong-way motion.
DocuRoute
Routes financial questions to hybrid sparse–dense retrieval or sandboxed text-to-SQL over SEC 10-K filings, with cited answers and a custom evaluation harness.
DepthCraft
Lifts Depth Anything V2 metric depth into a 3D point cloud and measures real distances with uncertainty, from a single photo.
Let's collaborate
I'm open to research collaborations, especially on vision, video, or language models. A short email about what you're working on is the best way to start.