About
I'm a machine learning engineer working on video understanding, vision-language models, and language models that explain themselves. I build systems that see what is happening and say what they saw.
I care most about reliability: whether a model keeps working outside the benchmark, and whether what it reports can be checked. That question runs through my work on human behaviour in video, on Video LLMs, and on grounded LLM output.
I'm always happy to collaborate on research in these areas.
Research interests
- Understanding human behaviour from video
- Video LLMs
- Grounded AI explanations
Across all of these, I'm drawn to one question: do models still behave reliably once they leave the benchmark?
News
-
2026
Paper
Submitted Joint Prediction of UHPC Properties Using a Feature-Tokenizer Transformer to Applied AI Letters. It's under review.
-
Jan 2025
Career
Joined Accelx Inc. as a Machine Learning Engineer.
-
2023
Course
Completed Machine Learning on Coursera.
-
Jan 2021
Thesis
Graduated from AUST with a thesis on arrhythmia classification using 2-D CNNs.
Selected projects
All projects →Overcooked AAR
The same model writes an after-action review of real human–human Overcooked play twice — from the event log and from the gameplay video — and every claim it makes is checked against the log.
Sentinel
Multi-camera, real-time anomaly detection. YOLOv8 + BoT-SORT tracking feeds a rule engine for falls, loitering, abandoned objects, crowding, and wrong-way motion.
DocuRoute
Routes financial questions to hybrid sparse–dense retrieval or sandboxed text-to-SQL over SEC 10-K filings, with cited answers and a custom evaluation harness.
DepthCraft
Lifts Depth Anything V2 metric depth into a 3D point cloud and measures real distances with uncertainty, from a single photo.
Let's collaborate
I'm open to research collaborations, especially on vision, video, or language models. A short email about what you're working on is the best way to start.