Prototypical Mask R-CNN for Vehicle Exterior Damage Detection. Hybrid instance segmentation combining Deep Metric Learning, Squared Euclidean Prototypical head & ResNet-50-FPN backbone on CarDD benchmark.
Hi, I'm Zahidul Hoque 👋
Aspiring Ph.D. Researcher in AI, Computer Vision & Autonomous Systems based in Ilford, London, UK. Recently completed M.Sc. in Data Science and Analytics with Advanced Research at the University of Hertfordshire, following a B.Sc. in Computer Science from Beijing Institute of Technology (BIT).
🔬 Research Domains & Interests
🎓 Key Academic Highlights
- Master's Final Research Project (7COM1039): Formulated a hybrid instance segmentation architecture integrating Mask R-CNN with Deep Metric Learning (replacing standard parametric linear classification with a Squared Euclidean Prototypical Predictor head). Evaluated on the CarDD benchmark, achieving mAP 0.319 (IoU 0.50:0.95), AP@50 0.630, and AR 0.436 across 810 test images.
- Mathematical Foundation: Enforced strict $L_2$ normalisation on 1024-D RoI embeddings and class prototypes, proving empirical pairwise distance $1.91 \le d^2 \le 2.09$ ($\theta \approx 90^\circ$), confirming convergence to an orthogonal latent feature space.
- Undergraduate Foundation: 4-year Computer Science degree at Beijing Institute of Technology (BIT) covering Data Structures & Algorithms, Multivariable Calculus, Linear Algebra, Probability & Statistics, and Operating Systems.
Two-model object detection and segmentation pipeline integrating Detection Transformers (DETR) and Few-Shot Object Detection (FSOD) pipelines for localized exterior vehicle anomalies.
Clean personal portfolio website developed with modern responsive web standards, semantic HTML structure, and Vanilla CSS styles.
Prototypical Mask R-CNN for Vehicle Exterior Damage Detection
University of Hertfordshire • Supervisor: Dr. Joseph Reddington • Academic Module: 7COM1039 (Advanced Research Project)
Core Architectural & Scientific Contributions
1. Novel Hybrid Deep Learning Architecture
Engineered and implemented an end-to-end instance segmentation network integrating Mask R-CNN with Deep Metric Learning, replacing the standard parametric linear classification layer with a Squared Euclidean Prototypical Predictor head.
2. Tackling Extreme Class Imbalance
Formulated a geometric centroid-learning framework to resolve the severe long-tail distribution and gradient dominance inherent in vehicle exterior inspection datasets (CarDD benchmark), eliminating minority class collapse (e.g. flat tires, shattered glass).
3. Hyperspherical Latent Geometry
Enforced strict $L_2$ normalisation on both 1024-D RoI feature embeddings and learnable class prototypes, projecting representations onto a unit hypersphere ($\|p\| = 1$). Mathematically derived that pairwise squared Euclidean distance:
Proved mathematically and empirically that the model converged to an orthogonal feature space that effectively eliminates inter-class visual confusion.
4. Optimization & Mixed-Precision Protocol
Leveraged a ResNet-50-FPN backbone pre-trained on MS-COCO, configured with temperature-scaled softmax
($\alpha = 20$) for rapid monotonic loss convergence, SGD with momentum, and learning rate
step-decay. Utilised Automatic Mixed Precision (torch.cuda.amp) to maximize memory
efficiency on an NVIDIA RTX 3060 (12GB).
5. Data Augmentation & Spatial Fidelity
Engineered synchronized spatial transformations (torchvision.transforms) dynamically
mirroring binary segmentation masks and bounding box coordinates during horizontal flips to maintain
pixel-level spatial fidelity under varying sensor orientations.
6. Future Vision-Language-Action (VLA) Roadmap
Proposed extending metric damage embeddings into the latent space of Multimodal Large Language Models (MLLMs) and VLA frameworks for automated, explainable vehicle safety diagnostics in autonomous mobility ecosystems.
📊 Empirical Evaluation on CarDD Benchmark
| Metric | Value | IoU Threshold / Protocol | Significance / Note |
|---|---|---|---|
| Mean Average Precision (mAP) | 0.319 | IoU 0.50 : 0.95 (COCO standard) | Strong instance segmentation across highly irregular damage shapes |
| AP@50 | 0.630 | IoU 0.50 | Accurate localization and region proposals |
| Average Recall (AR) | 0.436 | 100 detections max | Robust retrieval of rare & long-tail minority defects |
| Evaluation Dataset | 810 test images | CarDD Benchmark | Extreme real-world exterior inspection conditions |
| Prototypes Geometry | $1.91 \le d^2 \le 2.09$ | Unit Hypersphere ($\|p\| = 1$) | Orthogonal class separation ($\theta \approx 90^\circ$) |
🚗 Interactive CarDD Defect Simulator & Inspection Telemetry
Click the defect taxonomy chips below to simulate the Prototypical Mask R-CNN inference pipeline across the 6 CarDD damage categories:
Master of Science (M.Sc.) in Data Science and Analytics with Advanced Research
M.Sc. Thesis Project: Prototypical Mask R-CNN for Vehicle Exterior Damage
Detection.
Academic Supervisor: Dr. Joseph Reddington.
Focus Areas: Advanced Computer Vision, Deep Learning Architectures, Deep Metric
Learning, High-Dimensional Feature Space Geometry, Statistical Machine Learning, Neural Networks.
Bachelor of Science (B.Sc.) in Computer Science
Rigorous 4-year curriculum covering rigorous computational and mathematical theory.
Core Curriculum: Data Structures & Algorithms, Linear Algebra, Multivariable
Calculus, Probability & Mathematical Statistics, Operating Systems, Computer Architecture, Software
Engineering.
Project Associate (Computer Vision & Data Annotation)
- Executed high-precision data annotation, bounding box generation, and segmentation mask labeling for client machine learning and computer vision pipelines.
- Conducted systematic data cleaning, ground-truth quality auditing, and artifact filtering to eliminate label noise, enhancing downstream model training convergence.
Practical model deployment, monitoring, artifact versioning, and pipeline orchestration.
Advanced object-oriented programming, data structures, and algorithmic optimization.
Skills for Employment Investment Program (SEIP), Ministry of Finance, Bangladesh.
🛰️ LiDAR, Camera Perception & Sensor Fusion
Experience with multi-scale feature hierarchies (FPN), RoIAlign, and metric latent spaces directly translates to roadside and vehicle-mounted 3D LiDAR/Camera object detection pipelines (e.g. DINOSTAR, LiGuard) developed at UrbanITY Lab.
🌐 Digital Twins & Sim2Real Gap (UrbanTwin)
Strong mathematical foundation in geometric regularisation, domain adaptation, and feature normalization provides an ideal basis for tackling domain shifts between synthetic digital-twin LiDAR datasets (such as LUMPI, V2X-Real-IC, TUMTraf-I) and real-world sensor streams.
🛡️ Traffic Safety & VRU Protection
Passionate about applying real-time deep learning to vulnerable road user (VRU) detection, smart work zone monitoring, and incident mitigation for connected and automated transportation systems.