About
I'm a co-founder and Applied Scientist at NQB AI, where I build AI systems that solve real-world technical and operational problems.
My work spans two complementary areas. Physical AI connects machine learning with robots, cameras, sensors and other physical systems. Organizational AI applies AI, software and optimization to improve how organizations operate.
My background is rooted in computer vision and robotics. I hold a PhD in computer science from Université Laval, with research spanning object detection, 6D pose estimation, sim-to-real learning and 3D perception. Before NQB AI, I helped develop an industrial robotic vision product at Robotiq and, while working at Aerex Avionics, built AI and software systems for DRDC Valcartier, combining AI, software, sensors and hardware modules in research prototypes and proof-of-concept systems.
Across research and industry, the common thread in my work is turning AI models into systems that perform under real-world constraints.
Experience
I lead and deliver applied AI projects spanning computer vision, robotics, generative AI, optimization and custom software. Depending on the project I work as the primary technical contributor, lead a small technical team, or supervise developers and researchers.
NQB AIBuilt AI and software systems for DRDC Valcartier while employed at Aerex Avionics, combining computer vision, software development, sensor integration and experimental hardware. Work covered object detection, tracking and classification, with visible, infrared and hyperspectral sensing systems. Models ran on servers, PCs and embedded devices in research prototypes. Co-author of a NATO paper.
Thesis "Deep learning for object detection in robotic grasping contexts", supervised by Philippe Giguère (Université Laval) and co-supervised by Abdeslam Boularias (Rutgers University).
Read the thesisCo-developed an industrial pick-and-place vision product on Universal Robots arms, from camera firmware to object detection. The product was inspired by my master's research and I am a co-inventor of its patent (CA2977077C, granted in 2019). Employed from 2015 to 2018, then a further year of collaboration as a PhD intern funded by a Mitacs grant.
RobotiqResearch collaboration in Abdeslam Boularias's lab, on learning object localization and 6D pose estimation from simulation and weakly labeled real images. Published at ICRA 2019.
Rutgers Robot Learning LabRobotic grasping from demonstration: object recognition and matching from 3D point clouds, then control of a Kinova Jaco arm to reproduce the grasp. End-of-master's internship at Robotiq, where the project inspired one of the company's products. Pierre Marchand Excellence Award (2014).
Selected Research & Publications
- Manuscript under review
N. Lapointe, K. Laurent, K. Voyer, J.-P. Mercier, M. R. Karimi Dastjerdi, A. Grenier, C. Croeser, I. Van Der Wat
- FUSION · International Conference on Information Fusion
B. Debaque, H. Perreault, J.-P. Mercier, M.-A. Drouin et al.
Fusing thermal and visible images is a recurring challenge in computer vision, especially when the images of the two modalities are not well registered. This registration problem is traditionally solved by matching descriptors and depends on the richness and discriminating power of the representation. Ensuring that detected features are dense and uniformly distributed is not necessarily guaranteed. More recently, machine learning methods addressed the issue of visible to visible matching, but few address the multi-modality setting. In this paper, we propose to address the special case of thermal-visible image registration with small baseline parallax correction. Our deep homography model is evaluated on an open thermal and visible dataset with two training settings, unsupervised and supervised. Results demonstrate the feasibility of the approach, and performances comparison to state-of-the-art models is evaluated.
- Sensors · MDPI
A. Coluccia et al. (incl. J.-P. Mercier)
Adopting effective techniques to automatically detect and identify small drones is a very compelling need for a number of different stakeholders in both the public and private sectors. This work presents three different original approaches that competed in a grand challenge on the “Drone vs. Bird” detection problem. The goal is to detect one or more drones appearing at some time point in video sequences where birds and other distractor objects may be also present, together with motion in background or foreground. Algorithms should raise an alarm and provide a position estimate only when a drone is present, while not issuing alarms on birds, nor being confused by the rest of the scene. In particular, three original approaches based on different deep learning strategies are proposed and compared on a real-world dataset provided by a consortium of universities and research centers, under the 2020 edition of the Drone vs. Bird Detection Challenge. Results show that there is a range in difficulty among different test sequences, depending on the size and the shape visibility of the drone in the sequence, while sequences recorded by a moving camera and very distant drones are the most challenging ones. The performance comparison reveals that the different approaches perform somewhat complementary, in terms of correct detection rate, false alarm rate, and average precision.
- WACV · Winter Conference on Applications of Computer Vision
J.-P. Mercier, M. Garon, P. Giguère, J.-F. Lalonde
Much of the focus in the object detection literature has been on the problem of identifying the bounding box of a particular class of object in an image. Yet, in contexts such as robotics and augmented reality, it is often necessary to find a specific object instance, a unique toy or a custom industrial part for example, rather than a generic object class. Here, applications can require a rapid shift from one object instance to another, thus requiring fast turnaround which affords little-to-no training time. What is more, gathering a dataset and training a model for every new object instance to be detected can be an expensive and time-consuming process. In this context, we propose a generic 2D object instance detection approach that uses example viewpoints of the target object at test time to retrieve its 2D location in RGB images, without requiring any additional training (i.e. fine-tuning) step. To this end, we present an end-to-end architecture that extracts global and local information of the object from its viewpoints. The global information is used to tune early filters in the backbone while local viewpoints are correlated with the input image. Our method offers an improvement of almost 30 mAP over the previous template matching methods on the challenging Occluded Linemod dataset (overall mAP of 50.7). Our experiments also show that our single generic model (not trained on any of the test objects) yields detection results that are on par with approaches that are trained specifically on the target objects.
- NATO · STO-MP-MSG-SET-183, paper 19
G. Gagné, M. Breton, Y. de Villers, J.-P. Mercier, V. Paquin, J. Giesbrecht
- PhD thesis · Université Laval
J.-P. Mercier
In the last decade, deep convolutional neural networks became a standard for computer vision applications. As opposed to classical methods which are based on rules and hand-designed features, neural networks are optimized and learned directly from a set of labeled training data specific for a given task. In practice, both obtaining sufficient labeled training data and interpreting network outputs can be problematic. Additionally, a neural network has to be retrained for new tasks or new sets of objects. Overall, while they perform really well, deployment of deep neural network approaches can be challenging. In this thesis, we propose strategies aiming at solving or getting around these limitations for object detection. First, we propose a cascade approach in which a neural network is used as a prefilter to a template matching approach, allowing an increased performance while keeping the interpretability of the matching method. Secondly, we propose another cascade approach in which a weakly-supervised network generates object-specific heatmaps that can be used to infer their position in an image. This approach simplifies the training process and decreases the number of required training images to get state-of-the-art performances. Finally, we propose a neural network architecture and a training procedure allowing detection of objects that were not seen during training, thus removing the need to retrain networks for new objects.
- ICRA · International Conference on Robotics and Automation
J.-P. Mercier, C. Mitash, P. Giguère, A. Boularias
This work proposes a process for efficiently training a point-wise object detector that enables localizing objects and computing their 6D poses in cluttered and occluded scenes. Accurate pose estimation is typically a requirement for robust robotic grasping and manipulation of objects placed in cluttered, tight environments, such as a shelf with multiple objects. To minimize the human labor required for annotation, the proposed object detector is first trained in simulation by using automatically annotated synthetic images. We then show that the performance of the detector can be substantially improved by using a small set of weakly annotated real images, where a human provides only a list of objects present in each image without indicating the location of the objects. To close the gap between real and synthetic images, we adopt a domain adaptation approach through adversarial training. The detector resulting from this training process can be used to localize objects by using its per-object activation maps. In this work, we use the activation maps to guide the search of 6D poses of objects. Our proposed approach is evaluated on several publicly available datasets for pose estimation. We also evaluated our model on classification and localization in unsupervised and semi-supervised settings. The results clearly indicate that this approach could provide an efficient way toward fully automating the training process of computer vision models used in robotics.
- Canadian patent granted · CA2977077C
V. Paquin, M.-A. Lacasse, Y. Drolet-Mihelic, J.-P. Mercier
A robotic arm mounted camera system allows an end-user to begin using the camera for object recognition without involving a robotics specialist. Automated object model calibration is performed under conditions of variable robotic arm pose dependent feature recognition of an object. The user can then teach the system to perform tasks on the object using the calibrated model. The camera's body can have parallel top and bottom sides and adapted to be fastened to a robotic arm end and to an end effector with its image sensor and optics extending sideways in the body, and it can include an illumination source for lighting a field of view.
- WACV · Winter Conference on Applications of Computer Vision
J.-P. Mercier, L. Trottier, P. Giguère, B. Chaib-draa
Pick-and-place is an important task in robotic manipulation. In industry, template-matching approaches are often used to provide the level of precision required to locate an object to be picked. However, if a robotic workstation is to handle numerous objects, brute-force template-matching becomes expensive, and is subject to notoriously hard-to-tune thresholds. In this paper, we explore the use of Deep Learning methods to speed up traditional methods such as template matching. In particular, we employed a Single Shot Detection (SSD) and a Residual Network (ResNet) for object detection and classification. Classification scores allows the re-ranking of objects so that template matching is performed in order of likelihood. Tests on a dataset containing 10 industrial objects demonstrated the validity of our approach, by getting an average ranking of 1.37 for the object of interest. Moreover, we tested our approach on the standard Pose dataset which contains 15 objects and got an average ranking of 1.99. Because SSD and ResNet operates essentially in constant time in a Graphics Processor Unit, our approach is able to reach near-constant time execution. We also compared the F1 scores of LINE-2D, a state-of-the-art template matching method, using different strategies (including our own) and the results show that our method is competitive to a brute-force template matching approach. Coupled with near-constant time execution, it therefore opens up the possibility for performing template matching for databases containing hundreds of objects.
- ICRA · International Conference on Robotics and Automation
F.-M. De Rainville, J.-P. Mercier, C. Gagné, P. Giguère, D. Laurendeau
This paper proposes a complete system for robotic sensor placement in initially unknown arbitrary three-dimensional environments. The system uses a novel approach for computing the quality of acquisition of a mobile sensor group in such environments. The quality of acquisition is based on a geometric model of a camera which allows accurate sensor models and simple occlusion computation. The proposed system combines this new metric with a global derivative-free optimization algorithm to find simultaneously the number of sensors and their configuration to sense accordingly the environment. The presented framework compares favourably with current techniques working in two-dimensional environments. Furthermore, simulation and experimental results demonstrate the ability of the system to cope with full three-dimensional environments, a domain still unexplored by previous methods.
Projects
Robotics & Physical AI
Co-developed at Robotiq a vision product for industrial pick-and-place on Universal Robots: camera firmware, calibration, object detection and robot integration. Patented technology (CA2977077C) deployed on factory floors.
View the patentAssistive robotics system for wheelchair users: autonomous grasp replication from a single demonstration using Kinova Jaco, Kinect, ROS, OpenCV and PCL. MSc work, Pierre Marchand Excellence Award (2014), later adapted for industrial use at Robotiq.
Read the thesisCalibration of the transform between a Kinova Jaco arm and a Kinect camera using AR tags, with a documented ROS/RViz workflow (jpmerc/perception3d).
Code on GitHub
Computer Vision & Inspection
Computer vision system for automated sewer defect detection using robotic platforms with fisheye cameras.
Machine learning system developed at NQB AI to assess the degradation severity of mine shaft structural components from inspection data. Paper submitted in 2026.
Creating unique visual signatures for pharmaceutical product labels to authenticate originals and detect counterfeits on cylindrical packaging.
Detection of counterfeits by analyzing halftone printing patterns using engineered features and CNNs.
Computer vision to analyze and predict the evolution of use-wear traces on prehistoric stone tools. Collaboration with Université Laval archaeology laboratories.
Detection & Sensing
Deep learning detectors on visible and infrared imagery with image tiling for small object detection, running on a multi-sensor platform.
Adapted a real-time detection pipeline to process multispectral and hyperspectral imagery across visible and infrared bands.
Multi-camera system with deep learning at the edge and automated report generation, integrated with existing supervision tools.
Consulting work at Aerex Avionics for DRDC Valcartier: object detection, tracking and classification in visible, infrared and hyperspectral imagery, running on servers, PCs and embedded devices in research prototypes. Co-author of a NATO paper.
Qt configuration interface for a Leddar range finder (oversampling, accumulation, LED intensity, thresholds) and integration of the sensor into a ROS stack (jpmerc/leddarUI, jpmerc/leddartech).
Code on GitHubSystem to test a time-of-flight range sensor in deep water (Caribbean Sea). Raspberry Pi with a custom UI for real-time display.
Machine Learning & Optimization
Applied operations research at NQB AI: vehicle routing and schedule construction under real-world business constraints.
Calibration and optimization of hyperspectral plastic scintillation detectors used in radiotherapy dosimetry.
Convolutional LSTMs to predict next frames in fluid dynamics simulations. Automated pipeline to generate synthetic datasets from CFD simulations for training predictive models.
Deep learning and differentiable 3D rendering to predict analysis outcomes across viewpoints and parameters. Compared neural and deterministic interpolation approaches.
ML pipeline to predict material penetration outcomes from experimental data, with analysis of model generalization limits.
Simulator, clustering and classification to detect group movement patterns from sensor tracking data, with synthetic dataset generation for activity recognition.
Software & AI Systems
Automated extraction of wiring information from client PDF schematics and interaction with internal systems to streamline the quoting process.
Intelligent voice response system handling real-time phone calls with specialized AI agents (receptionist, product advisor, support). Built on Asterisk PBX with intelligent call queuing, human handoff, AI voicemail and automated email summaries.
SaaS platform for AI-driven creation and editing of images and 3D models, integrating generative models into production workflows.
Intelligent chatbots using retrieval augmented generation and function calling for context-aware answers drawn from internal knowledge bases.
Evaluated SaaS solutions and built a proof of concept using LLMs (Claude, GPT, Gemini) to automatically resize Adobe InDesign advertisements to multiple formats by transforming exported HTML5. Implemented a generation, capture and critique feedback loop.
Automated email response system for a rental company, parsing reservation requests and generating context-aware replies.
Automated generation of structured product catalogs from noisy, unstructured database entries using NLP techniques.
Semi-automated annotation tool with object detection and tracking for label propagation. Cross-sensor affine transformations enable multi-modality label transfer.
Control software for electro-optic system simulation and hardware-in-the-loop testing.
Interactive visualization tool for electrical wiring schematics, enabling operators to navigate wire by wire with precise path highlighting on PDF diagrams.
