When a detector reaches 0.998 macro-F1, has it learned the task, or just the traces of how the benchmark was made?
Khalequzzaman Likhon
Machine learning engineer building reliable vision, video, and language systems.
Research interests
What I work on
- Computer Vision
- Real-Time Video Understanding
- Large Language Models
- Generative AI
- Multimodal Learning
- Human-Centered Computing
Across all of these, I care most about reliability: how models behave under real-world distribution shift, when they should abstain, and whether their outputs stay grounded in evidence that can be checked.
Questions I'm pursuing
For detecting fights in surveillance video, which works better: appearance-based spatiotemporal features or explicit human-pose dynamics?
See: RWF-2000 comparison, Sentinel
How can small local models (3B-parameter LLMs, 2B VLMs) be made trustworthy enough for high-stakes decision support?
See: DocuRoute, SentinelRAG
How do we keep generated output faithful to its source, whether that's generated code to its tests or a dubbed voice to its speaker?
See: Research→Code, voice-dub
When does fusing another modality actually help, and would it beat a control where that modality is shuffled?
When should an AI system stop and hand the decision back to a person?
About
A short introduction
I'm a Machine Learning Engineer at Accelx Inc., where I've built safety-critical perception and language systems since January 2025. These include real-time weapon, violence, and fall detection on live multi-camera video, and risk-assessment systems that write grounded, cited alerts. Watching these systems fail in the real world is why I focus on reliability.
My research asks whether strong results mean what they seem to. I've compared pose-based and appearance-based models for violence recognition, built LLM pipelines where every citation is checked, and written negative-result studies that test models against their shortcuts. One manuscript is under review at Applied AI Letters.
I hold a B.Sc. in Computer Science and Engineering from Ahsanullah University of Science and Technology. I'm now looking for PhD opportunities.
News
Recent updates
-
2026
Paper
Submitted Joint Prediction of UHPC Properties Using a Feature-Tokenizer Transformer to Applied AI Letters. It's under review.
-
2026
Paper
Preparing Out-of-Focus Is Not Obliteration, a perturbation analysis of fingerprint alteration detectors.
-
2026
Paper
Preparing Does Geometry Help?, a negative-result study on post-earthquake damage classification.
-
Jan 2025
Career
Joined Accelx Inc. as a Machine Learning Engineer.
-
2023
Course
Completed Machine Learning on Coursera.
-
Jan 2021
Thesis
Graduated from AUST with a thesis on arrhythmia classification using 2-D CNNs.
-
2019
Service
Served on the organizing committee of AUST Codeware.
Selected projects
Things I've built
Sentinel
Multi-camera, real-time anomaly detection. YOLOv8 + BoT-SORT tracking feeds rules for falls, loitering, abandoned objects, and wrong-way motion, and an LSTM autoencoder flags what the rules miss.
DocuRoute
Routes financial questions to hybrid sparse–dense retrieval or sandboxed text-to-SQL over SEC 10-K filings, with cited answers and a custom evaluation harness.
Research→Code
Researcher, coder, and reviewer agents on LangGraph with MCP tools. Code is tested in a sandbox, retries are capped, and a human approves the result.
DepthCraft
Fuses monocular metric depth into 3D meshes for distance measurement, with uncertainty estimates.
Let's talk
I'd love to hear from faculty recruiting PhD students and from researchers working on vision, video, or language models.