Undergrad thesis — vision-language models
Spatial reasoning in VLMs and the mechanistic interpretability behind it: how these models represent spatial relations, and where in the network that computation happens.
Final-year CS undergrad at BUET · AI researcher
I research vision-language models — spatial reasoning and the mechanistic interpretability behind it — and build agentic systems for evaluation and software testing. Before research I spent two years on full-stack and DevOps: microservice platforms, CI/CD, and cloud-native infrastructure.
// about
I'm a final-year Computer Science student at Bangladesh University of Engineering and Technology. I started in full-stack web development, then spent two years on DevOps and DevSecOps — automation pipelines, cloud-native deployments, and secure delivery — alongside competitive CTFs.
My work now is research. My thesis is on spatial reasoning in vision-language models and the mechanistic interpretability behind it. I also worked on agentic systems for exploratory software testing, Bangla LLM-text detection, and hallucination and deepfake detection.
// research
Spatial reasoning in VLMs and the mechanistic interpretability behind it: how these models represent spatial relations, and where in the network that computation happens.
An agent that explores mobile applications on its own and generates test cases, without hand-written scripts or predefined paths.
Curating a dataset of Bangla LLM-generated text and building detection methods for it — an under-resourced setting for existing detectors.
Detecting hallucinated and unsupported content in language-model outputs — developed for CUET Datathon 2026.
Robust detection of AI-generated images and AI-generated video, built at BRAC AI Build Fest. The provenance-first pipeline in the Verilens extension builds on this work.
// projects
Manifest-V3 browser extension that flags AI-generated media and misinformation on X, Instagram, and Facebook. Checks C2PA content provenance first, then falls back to deepfake models and automated fact-checking.
An A2A-protocol evaluator for GAIA benchmark tasks, built on Google ADK. Deterministic scoring, optional LLM-based judging, and multi-agent orchestration for reliable agent evaluation.
A standalone LLM agent for the Build-What-I-Mean benchmark. Two-stage pipeline with speaker-aware pragmatic inference, communicating over the A2A protocol.
Champion project at CUET API Avengers 2025. An event-driven donation platform: FastAPI microservices, Kafka workflows, gRPC service communication, an observability stack, CI/CD, and Kubernetes deployment.
A microservices freelancing platform: API gateway, JWT auth, escrow payments, workspace collaboration, and vector-search matching between clients and freelancers.
A full compiler pipeline written from scratch: symbol table, lexical analysis with Flex, syntax and semantic analysis with ANTLR4, and 8086 assembly code generation with optimisation.
An autonomous maze-solving robot on ATmega32: left-hand-rule search, triple ultrasonic sensors, MPU6050 gyroscope feedback, and PID-assisted movement for reliable navigation on constrained hardware.
// milestones
Global bronze medal on Kaggle.
Top-4% finish in a deep-learning competition.
First place with Team FAT32, for the CareForAll microservices platform.
Hardware and IoT build under a fixed time limit.
My first inter-university competition, and the start of the competitive track.
// contact
The fastest way to reach me is email. I'm open to research collaboration, and to internship or new-grad roles in ML and infrastructure.